Why Your OCR Output Is Garbage (and How to Fix It)
Resolution, contrast, and page segmentation decide whether OCR gives you clean text or noise. What to change, in the order that actually matters.
The first time I fed a screenshot into Image to Text (OCR) I got back something like Pre:eren<es. My conclusion was that browser OCR is a toy. That was wrong. I retook the same screenshot at 200% browser zoom and the same engine read every word.
Nothing about the software had changed. The letters just had three times as many pixels.
Pixels per letter is most of the game
Tesseract wants a capital letter roughly 30 pixels tall. A 300 DPI scan of 10 pt type gives it about 42, which is why flatbed scans come out at 98–99% of characters right. A screenshot of a web page at 100% zoom on a 1080p monitor gives it about 11. At that size the crossbar of an e and the gap in a c are the same one or two pixels, and the model is guessing.
So before you touch any other setting:
- Take the screenshot at 200% browser zoom, or on a high-density display.
- Photograph a page square-on and fill the frame with it. Don’t shoot the whole book.
- If the image is all you have, leave the upscaling option on. It helps, because resampling restores some of the stroke shape — but it can’t invent detail that was never captured.
The setting almost nobody changes
Page segmentation mode is the second lever, and it matters more than people expect. The default runs full layout analysis: find the columns, find the blocks, read them in order. On a book page that’s exactly right. On a screenshot with six words scattered across a dialog box, Tesseract hunts for columns that don’t exist, decides most of the image is a picture, and sometimes hands back an empty string.
Switch to sparse text for anything screen-shaped. No layout analysis, just find words wherever they are. For a single line — a serial number, a licence plate — pick the single-line mode and skip the guessing altogether.
Contrast, and the dark-mode trap
Tesseract binarises internally and it does that well, so cranking contrast is rarely the fix people hope for. Two exceptions are worth knowing.
Light text on a dark background reads badly. Inverting it costs nothing and usually solves the whole problem in one click. And a photo of a faded thermal receipt is a genuine contrast problem: the ink and the paper are two shades of the same grey, and stretching that range is the only thing that separates them.
JPEG is the other quiet troublemaker. Re-saving a screenshot as JPEG scatters ringing artefacts around every glyph edge, and OCR reads them as punctuation. PNG for anything with text in it.
The line breaks aren’t a bug
OCR reports one line of text per line on the page. It has no way to know which of those breaks the author put there and which ones were just the right margin. That’s why raw output looks shredded.
Turning on “join wrapped lines” merges a line into the next unless it ended on a full stop, a colon, or a closing quote, or the next line opens a bullet. Paragraphs flow again, lists keep their shape. Dehyphenation handles the rest: a word split as recog- / nition goes back together.
When to stop trying
Cursive handwriting won’t work. The standard language files have no handwriting model, and no setting changes that. Tables come out visually aligned but not as real cells. And if your PDF already has a text layer — anything out of Word, a browser, or invoicing software — skip OCR entirely and copy the text, which is exact rather than 98% right.
For everything else, the order is: more pixels, then the right segmentation mode, then invert if needed. Run your worst file through the OCR tool and check the confidence number it prints — under 75, go back and fix the image rather than the text. A quick pass through the spell checker catches most of what’s left.