Upload a document to begin
Upload target document(s)
Multiple files supported · PDF · PNG · JPG · WEBPUpload a document to begin
Upload target document(s)
Multiple files supported · PDF · PNG · JPG · WEBPEverything you need to know about DocumentChimp, OCR, templates, and privacy.
DocumentChimp is designed from the ground up to keep your documents and data private.
Your PDFs, images, and extracted text never leave your device. There is no backend server that receives or stores your documents.
No analytics, no fingerprinting, no third-party cookies. We don't know who you are and we don't want to.
Documents and templates exist only in your browser's memory. Refreshing or closing the tab clears everything immediately.
PDF rendering (PDF.js) and OCR (Tesseract.js) both run in your browser via WebAssembly. No computation happens on a remote machine.
Templates are only shared when you explicitly download them as JSON files. No data is ever sent anywhere automatically.
DocumentChimp is built on well-known open-source libraries. You can audit everything with your browser's developer tools.
Plain-language terms for using DocumentChimp Workbench. By using this site you agree to the terms below.
DocumentChimp is provided free of charge, as-is, and without warranty of any kind. You may use it for personal or commercial purposes.
You are responsible for the documents you process and for ensuring you have the right to use them. Do not use DocumentChimp to process content you do not have permission to handle.
You agree not to use DocumentChimp for unlawful purposes, to infringe intellectual property rights, or to process material that is illegal in your jurisdiction.
DocumentChimp is provided without any guarantee of accuracy, availability, or fitness for a particular purpose. We are not liable for any loss arising from its use, including OCR errors or missed extractions.
Features may be added, changed, or removed at any time. These terms may be updated; continued use of the site after changes means you accept the revised terms.
Questions about these terms? Email documentchimp@gmail.com.
Practical articles about OCR template extraction, batch processing, and getting the best results from DocumentChimp.
A step-by-step walkthrough for drawing regions on an invoice, naming them clearly, and saving the layout so it can be reused on dozens of similar documents. Covers the difference between the Create Template and Extract tabs, why region names matter, and how to export a template for backup.
How the template identifier works, why naming conventions in your filenames matter, and how to configure a batch of mixed invoices, receipts, and purchase orders so each document automatically finds the right template without manual selection.
Best practices for large runs: grouping files by template, using the queue in stages, monitoring progress with the run counter, and downloading results as JSON for downstream processing. Includes tips for keeping memory usage reasonable during multi-page PDF runs.
Why scan quality dominates OCR results, how padding around a region affects recognition, when to use higher-resolution inputs, and what to do when a region returns empty. Practical advice drawn from real invoice and receipt extraction tests.
A field-by-field reference for the exported template JSON: name, identifier, pageCount, documentSize, and the region array with its coordinates. Explains how imported templates are renamed, why templates store coordinates rather than values, and how to hand-edit a template file safely.
How regions are tagged with the page they belong to, how to draw on different pages of the same PDF, and what happens during extraction when a template spans several pages. Includes guidance on structuring templates for long documents such as contracts or statements.
An explanation of what it means for OCR and PDF rendering to run entirely in your browser, why that keeps confidential paperwork off any server, and how to verify the behavior yourself using the browser's developer tools and network panel.
Short, practical habits: naming regions consistently, drawing slightly larger regions than the text, using the zoom control before drawing, previewing a template against a sample file, keeping a backup of exported templates, and other quick wins gathered from daily use.