OCR / image to text

Processed in this browser

File limit: 32 MB

This tool never uploads your input.

This step may download a library in this browser the first time you use it.

English-first Tesseract (eng). Do not use this for ID cards or CJK documents. The model downloads on first use in this tab.

Drop a file

Drop a file here or choose one.

Sample workflow: drop a photo of printed text, then Process. The OCR model downloads on first use.

Run English-first OCR on a screenshot or scan in this tab. The Tesseract model downloads once. This is not a passport reader, not a 99% accuracy claim, and not the right tool when the PDF already has a text layer.

You need the words off a screenshot or English scan and you do not want that image on an OCR SaaS. That is this page: Tesseract in the tab, English (eng) first.

If the PDF already lets you highlight text, skip OCR and use extract PDF text. If you only needed a picture of the page, use PDF to JPG.

What OCR can and cannot read

Works often: printed English, clear screenshots, high-contrast scans.

Fails often: handwriting, low-res phone photos of glossy paper, stamps, dense tables, and languages we did not load.

This is not Google Cloud Vision and not a structured ID or invoice parser. We will not lead with passport MRZ or “证件 OCR.”

How to use it

  1. Drop a still (screenshot, JPEG, PNG). Typical ceiling about 32 MB.
  2. Wait for the English model on first use, then recognize in this tab.
  3. Copy the text. Clean spacing with whitespace cleaner if needed.

Honest limits

  • Accuracy is “good enough to start typing,” not a court transcript.
  • No searchable-PDF export.
  • HEIC may need HEIC to JPG first.

Local is why you would use a browser OCR — not a claim we are the only private tool on the internet.

FAQ

Which languages work?

The engine is English-first (eng). Do not bring a 身份证 or dense CJK scan expecting a Chinese model. Handwriting and ornate seals fail often.

Why is the PDF text empty if I screenshot it?

If you can already select text in the PDF, extract PDF text instead. OCR is for pixels — photos of pages and screenshots.

Does the image get uploaded?

Recognition runs in the tab after the language data is fetched. Your photo is not the thing we want on a server; the model download is a one-time script/WASM fetch, not your document archive.

Can you make a searchable PDF?

No. You get text you can copy. We do not write a hidden text layer back into the PDF.

Will tables stay aligned?

Rarely. Tesseract-class OCR is line-oriented. Rebuild tables by hand.

Is this HIPAA-certified?

No. Local processing is a privacy habit, not a compliance program.

Related tools