A OCR labeling tool - made for docTR
-
Updated
Oct 4, 2026 - Python
A OCR labeling tool - made for docTR
Cookiecutter template for Python 3 packages
Modular OCR pipeline for live-camera documents/screens. Auto-aligns perspective, blocks obstructions, and extracts text using interchangeable deep-learning models.
A local-first engine for turning scanned textbooks into structured RAG databases. Features layout-aware OCR, deterministic math cleanup, and pedagogical role-indexing using LanceDB and CLIP vision.
Production-grade OCR evaluation and optimization platform for benchmarking PaddleOCR, DocTR, and Tesseract with preprocessing optimization, statistical comparison, experiment tracking, and interactive HTML reports.
To associate your repository with the doctr topic, visit your repo's landing page and select "manage topics."