Pdf Inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
About
Fast Rust library for PDF classification and text extraction. By default it detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. Native Rust and CLI consumers can opt into selective OCR. Includes bindings for Python, Node.js, and browser WebAssembly.
Open Source Health
- Stars
- 19,072
- Forks
- 1,274
- License
- MIT
- Last commit
- 29 days ago
Resources & Links
Related Categories
Vendor
Firecrawl
Web data API for AI
Quick Links
Open Source
More by Firecrawl
Related Products
Anydoc
Convert documents (doc, docx, odt, rtf, epub, pdf, presentations, spreadsheets, csv) to GitHub-Flavored Markdown
More from this vendor
Xournalpp
Xournal++ is a handwriting notetaking software with PDF annotation support. Written in C++ with GTK3, supporting Linux (e.g. Ubuntu, Debian, Arch, SUSE), macOS and Windows 10. Supports pen input from devices such as Wacom Tablets.
Top category match
Ng2 PDF Viewer
Angular 5+ component for rendering PDF
Top category match
Go Wkhtmltopdf
Golang commandline wrapper for wkhtmltopdf
Top category match
jsPDF
Client-side JavaScript PDF generation for everyone.
Top category match