LLM Aided OCR

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

LLM OrchestrationOpen Source AI & Machine LearningOpen Source

About

The LLM-Aided OCR Project is an advanced system designed to significantly enhance the quality of Optical Character Recognition (OCR) output. By leveraging cutting-edge natural language processing techniques and large language models (LLMs), this project transforms raw OCR text into highly accurate, well-formatted, and readable documents.

Open Source Health

Not enough history
Stars
2,996
Forks
214
License
Not stated
Last commit
2 months ago
Python

Related Categories

Using LLM Aided OCR?

Track its cost next to the rest of your stack and get a reminder before it renews.

Add to my stack

Vendor

Jeff Emanuel

Building in NY

View vendor profile

More by Jeff Emanuel