Stirling-PDF
Leading#1 in Open sourcehigh confidence
The de-facto open-source PDF toolkit: 50+ local operations — merge, split, OCR, convert, sign, redact, compress — fully self-hosted. The recognized OSS PDF leader (~88k stars).Our read
Why it ranks #1
The recognized all-in-one open-source PDF leader at about 87.9k GitHub stars, with active daily development and unusually broad self-hosting community adoption.Here is the catch
self-hosters own upgrades, storage and security
large feature surface can make the UI feel busy
resource-heavy OCR and conversion jobs require capable hardware
Does this well
keeps sensitive documents on infrastructure you control
replaces dozens of single-purpose PDF web tools
large active community and frequent releases
Pricing
Checked by hand on 2026-07-23. Prices in this category change often — if this looks wrong, it probably is.
Key features
Merge, split and reorganize PDFsOCR, conversion and compressionSigning, redaction and form toolsDocker-based self-hosting
Sources we read
Quick facts
More in this area
The rest of the Open source column.- 2MarkItDown (Microsoft)Microsoft's document-to-Markdown converter (PDF, Office, images, audio) built to feed clean text to LLMs; the highest-starred tool in this space (~169k stars).
- 3MinerU (OpenDataLab)High-accuracy PDF-to-Markdown/JSON extraction; the standout for complex academic papers and CJK layouts (~76k stars).
- 4Docling (IBM / DS4SD)IBM's document parser (PDF, DOCX, PPTX) with strong layout and table understanding; a default in the LlamaIndex/LangChain ecosystem (~64k stars).
- 5marker (Datalab)Fast, high-quality PDF/EPUB/DOCX to Markdown conversion with tables, math, and images; a general-purpose favorite (~38k stars).
- 6pdfplumberPrecise text, table, and layout extraction from born-digital PDFs; the go-to Python library for table scraping (~11k stars).
- 7PyMuPDFFast low-level PDF/XPS parsing, rendering, and editing (MuPDF bindings); the performance workhorse under many pipelines (~10k stars).