Image to text (OCR)
Pull the text out of a photo or a scan.
Drop a photo, scan or PDF with text
or click to choose (.jpg, .jpeg, .png, .webp, .bmp, .pdf). Up to 10 MB per file.
The right language is what separates “naive” from “naïve”.
Line breaks inside a paragraph come from the page width, not the author.
Recognition runs in your browser, the document is never uploaded. The first run downloads about 10 MB of engine and language data, after that it starts instantly.
- Runs right in your browser
- Free, no sign-up
More about Image to text (OCR)
How to pull text out of an image or a scan
Upload a photo, a screenshot or a scanned PDF, choose the language and start recognition. The text appears in a box, ready to copy or download as TXT. You can also fix it up right there.
Recognition runs on Tesseract compiled to WebAssembly, that is, right in your browser. A contract or a receipt never leaves your computer, which is no small detail for documents with personal data in them.
- Handles English and Czech, accents included.
- For a scanned PDF it renders the pages and recognises them one by one, up to twenty.
- The first run downloads about 10 MB of engine and language data, after that there is no waiting.
How to get a better result
Recognition lives and dies by the source image. Shoot the page square on from above, not at an angle, and watch the focus. A shadow across half the page, or a handheld shot in dim light, turns the result into an unreadable mess.
Frequently asked questions
Does it read handwriting?
No. Tesseract is built for printed text. Handwritten notes come out right only by accident.
Why does the text have typos?
When confidence is low, a sharper scan, more contrast or the right language choice all help. Text with accents recognised using the English model comes out stripped of them.
Are the layout and tables preserved?
No, the output is plain text. Columns and tables collapse into lines, so they usually need rebuilding by hand.
Does it work on an ordinary PDF with text?
It does, but there is no point. If the PDF already contains text (you can search it), use the PDF to Word tool, the result will be more accurate.