PDFWorkshop

Split-screen OCR workbench for scanned PDFs. By esaruoho — see Contributions to PDFWorkshop.

Combines three OCR engines — Tesseract, Gemini, and GLM-OCR (MLX auto-bootstrap on Apple Silicon, Ollama elsewhere) — in a split-screen workbench where the scan and the recognised text sit side by side. Includes a headless batch OCR worker fed by a Syncthing queue, and exports searchable PDFs.

Built for the archival problem at the heart of MERLib: thousands of image-only scanned documents that need accurate, verifiable OCR before they can be read, searched, or cross-referenced.

Contribution surface (pulled 2026-06-01)

Related

In this section