Extract text from an image (OCR)

Recognize the text in a photo, a screenshot or a scanned PDF and copy it or download it as .txt. Works in English and Italian. Everything happens in your browser.

🔎 Tap here or drop an image or a PDF

JPG, PNG, WEBP or scanned PDF

⬇️ Download .txt

🔒 Fully private: the image and the recognition happen on your device, nothing is uploaded.

What OCR is

OCR (optical character recognition) turns text «drawn» inside an image into real, editable text you can copy. It's perfect for scanned documents, photos of pages, screenshots, signs or labels.

Tips for a good result

Use sharp, well-lit images with straight text and good contrast between the writing and the background. Pick the right language for the best accuracy. The first run downloads the recognition model (only once), then it stays in memory. If the photo of the document is skewed, it is worth first going through Straighten a document.

What makes recognition fail

Four things, in order of severity. Resolution: below a certain size letters do not have enough pixels to be told apart, and the rule of thumb is that text should be at least twenty pixels tall. Skew: a few degrees off straight lowers accuracy a great deal, because the software looks for horizontal lines.

Then contrast, that is how far the writing stands out from the background: a faded photocopy or a photo with the shadow of a hand across it gives dreadful results. And finally the typeface: decorative fonts, italics and handwriting are another problem entirely, one traditional tools do not even attempt.

How to photograph a document so that it works

In diffuse light, never with flash and never with a lamp to one side: shadows and reflections are the number one cause of bad recognition. The phone should be held parallel to the page, not tilted, because perspective squeezes the far lines and makes them illegible.

If the sheet is on a table, it helps to lay it on a surface of a different colour and fill the frame with the document alone: what surrounds it is useless and confusing. And if the photo has already been taken crooked, before recognising the text it should be straightened: it is the single step that changes the most results.

What to check in the text that comes out

The misreadings are always the same, and are found quickly if you know where to look: 0 and O, 1 and l and I, 5 and S, 8 and B, rn read as m. These are errors a distracted reread does not catch, because the text stays plausible.

Where they really matter is in numbers: a bank account, a tax code, an amount, a date. There a wrong character is invisible and does damage, so every numeric string should be compared by hand against the image. Running text, by contrast, can be corrected by reading.

The limits, stated plainly

The handwriting is not recognised by classic tools, nor by this one: it needs a model trained for the purpose, and even those err far more than on printed text. Tables lose their structure: cells become lines of text in a row, and rebuilding the columns is a separate job.

And the layout does not survive: newspaper columns, footnotes and boxes come out mixed in the order the software meets them. What you get is the content, not the document.

Nearby tools

If the photo is crooked, you first need Straighten a photographed document. For a PDF that already has text inside there is Extract text from a PDF, and to search across many PDFs Search a word across many PDFs. On the resulting text work Word counter and Fix broken characters.