CodeBrowser / github / ocrmypdf/OCRmyPDF
ocrmypdf/OCRmyPDF
# OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
$ git clone https://github.com/ocrmypdf/OCRmyPDF.git
stars
34,240
forks
2,364
language
Python
license
Mozilla Public License 2.0
What is ocrmypdf/OCRmyPDF?
OCRmyPDF is a Python tool that adds searchable text layers to scanned PDF files using optical character recognition (OCR). It processes image-based PDFs through Tesseract to extract text, making documents full-text searchable while preserving the original layout and image quality. Developers use it to automate document digitization workflows, build search-enabled document systems, or integrate OCR capabilities into larger applications.
Topics
#image-processing #ocr #pdf #python #tesseract
Activity
100 open issues · last updated Jul 21, 2026