Tika
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
About
Apache Tika(TM) detects and extracts metadata and text from over a thousand file types. As of 4.0.0 it emits Markdown by default — output shaped for LLM and RAG pipelines — parses in crash-isolated forked processes, and adds vision-language-model parsers (Claude, Gemini, OpenAI) for documents OCR can't read.
Open Source Health
- Stars
- 4,055
- Forks
- 969
- License
- Apache-2.0
- Last commit
- 28 days ago
Resources & Links
Related Categories
Vendor
Apache
Publisher of Apache Jena, Fuseki and Apache Kafka
Quick Links
Open Source
More by Apache
Related Products
Xournalpp
Xournal++ is a handwriting notetaking software with PDF annotation support. Written in C++ with GTK3, supporting Linux (e.g. Ubuntu, Debian, Arch, SUSE), macOS and Windows 10. Supports pen input from devices such as Wacom Tablets.
Top category match
Ng2 PDF Viewer
Angular 5+ component for rendering PDF
Top category match
Go Wkhtmltopdf
Golang commandline wrapper for wkhtmltopdf
Top category match
jsPDF
Client-side JavaScript PDF generation for everyone.
Top category match
PDF.js
PDF Reader in JavaScript
Top category match
