Take it step by step
- Start with an upright, sharply focused scan.
- Choose the available language pack and explicitly download it if needed.
- Run OCR and inspect several pages, including small text and tables.
- Download and search the exported PDF for known phrases.
A situation to try.
Scan a synthetic page containing an invoice number and two similar characters, such as O and 0. Check recognition against the visible image.
Use the tool’s synthetic example first where available. The example here explains a decision; it is not a reported benchmark result.
Before you call it done.
- Verify names, dates and amounts.
- Check rotated pages and reading order.
- Confirm text is searchable after download.
Know the limits.
Recognized text is not an authoritative transcription. Poor scans and handwriting can produce incorrect results.
This tool’s current boundary: English printed-text OCR. Review recognized names and numbers. Searchable font embedding is limited to Latin text.
Evidence and scope
Technical source guidance and task-specific output checks. These do not establish that every file, browser or device passes.
Use the checks above on your own exported result. A source explains the format or mechanism; it is not proof that this export preserved every feature.
Technical reference: Tesseract.js: OCR engine ↗
Choose the right tool for this step.
Document scanner
Straighten, crop and clean up photographed pages.
On your deviceWhy use it: Prepare a sharper upright page first.
Before you start: Poor input limits recognition.
Searchable OCR
Recognize scan text and add a searchable text layer.
On your deviceWhy use it: Add recognized text to a scan.
Before you start: Verify amounts, dates and names against the image.
Preparation tools process locally. Official service links take you to the authority’s website. Inspect any exported copy before sharing it.