Tika
Document CollaborationOffice Suites

Tika

The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).

Open Source

About

Apache Tika(TM) detects and extracts metadata and text from over a thousand file types. As of 4.0.0 it emits Markdown by default — output shaped for LLM and RAG pipelines — parses in crash-isolated forked processes, and adds vision-language-model parsers (Claude, Gemini, OpenAI) for documents OCR can't read.

Open Source Health

Not enough history
Stars
4,055
Forks
969
License
Apache-2.0
Last commit
28 days ago
Java

Related Categories