Skip to content
anyutil.io

Image to text (OCR)

Pull the text out of a photo or a scan.

Drop a photo, scan or PDF with text

or click to choose (.jpg, .jpeg, .png, .webp, .bmp, .pdf). Up to 10 MB per file.

The right language is what separates “naive” from “naïve”.

Line breaks inside a paragraph come from the page width, not the author.

Recognition runs in your browser, the document is never uploaded. The first run downloads about 10 MB of engine and language data, after that it starts instantly.

More about Image to text (OCR)

How to pull text out of an image or a scan

Upload a photo, a screenshot or a scanned PDF, choose the language and start recognition. The text appears in a box, ready to copy or download as TXT. You can also fix it up right there.

Recognition runs on Tesseract compiled to WebAssembly, that is, right in your browser. A contract or a receipt never leaves your computer, which is no small detail for documents with personal data in them.

  • Handles English and Czech, accents included.
  • For a scanned PDF it renders the pages and recognises them one by one, up to twenty.
  • The first run downloads about 10 MB of engine and language data, after that there is no waiting.

How to get a better result

Recognition lives and dies by the source image. Shoot the page square on from above, not at an angle, and watch the focus. A shadow across half the page, or a handheld shot in dim light, turns the result into an unreadable mess.

Frequently asked questions

Does it read handwriting?

No. Tesseract is built for printed text. Handwritten notes come out right only by accident.

Why does the text have typos?

When confidence is low, a sharper scan, more contrast or the right language choice all help. Text with accents recognised using the English model comes out stripped of them.

Are the layout and tables preserved?

No, the output is plain text. Columns and tables collapse into lines, so they usually need rebuilding by hand.

Does it work on an ordinary PDF with text?

It does, but there is no point. If the PDF already contains text (you can search it), use the PDF to Word tool, the result will be more accurate.

Related tools