doc2text
Detect text blocks and OCR poorly scanned PDFs in bulk. Python module available via pip.
About
Detect text blocks and OCR poorly scanned PDFs in bulk. Python module available via pip.
Open Source Health
- Stars
- 1,278
- Forks
- 101
- License
- MIT
- Last commit
- 6 years ago
Related Categories
Vendor
Dr. Joe Sutherland
Emory Professor and Director of AI Center. , grad. , alum. Former Amazon and Cisco
Quick Links
Open Source
Related Products
OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Top category match
ZenlessZoneZero-Auto
绝区零 | ZenlessZoneZero | 零号空洞 | 自动战斗 | 自动化 | 图片分类 | OCR识别
Top category match
PaddleOCR2Pytorch
PaddleOCR inference in PyTorch. Converted from [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
Top category match
MinerU
A practical document parsing tool for converting PDF, images, DOCX, PPTX, and XLSX into Markdown and JSON
Top category match
Deepseek OCR App
A quick vibe coded app for deepseek OCR
Top category match