Content

Turning an image into text, and where the machine gets it wrong

Lifting text out of a picture runs inside the browser. The worry is not the moment it fails outright, but the moment it reads almost correctly and the page still looks clean.

You photograph a page of a book, a notice on a board or a page of a contract, and then you need the words out of it to paste somewhere else. Retyping costs fifteen minutes a page, and one wrong digit in an account number costs a good deal more than that.

Lifting text out of a picture has a name in the trade, OCR, and it has run inside a browser for some years now. What is worth saying is not that it works, but that it fails in a way that slips past people.

How a machine reads

The machine does not read in the sense a person does. It breaks the page into blocks, each block into lines, each line into clusters of marks, then compares every cluster against shapes it has learned and picks whichever letter fits best.

A three-stage diagram from left to right, a page split into two blocks, one block split into lines, and one line split into eight small squares each holding an abstract mark, the last outlined in coral
Page into blocks, block into lines, line into clusters, and only at the last cluster does the guessing start

Because every step is a guess, the result always carries a number saying how sure it was. The tool here takes the mean confidence per word, and on a clean scan that number usually sits in the eighties or above.

A hurried photograph of a slightly curved page drops it to around fifty. At that level the text that comes back is a mix of right and wrong, and the wrong part does not announce itself.

The dangerous errors are the ones that look fine

If the machine returns a line of nonsense you know at once it has failed and you retype it. The trouble is the places where it read almost correctly: a zero becoming a capital O, a one becoming a lowercase l, a decimal point becoming a comma inside a figure.

Pairs of easily confused characters side by side, zero and capital O, one and lowercase l, full stop and comma, each pair joined by a coral dot, with a confidence bar underneath
Three pairs that swap places for each other, and a confidence bar whose coral stretch is the part you reread yourself

Errors like that leave a page that reads cleanly, so nobody goes back over it. That is why the tool says so when confidence falls below its threshold rather than quietly handing over something that looks acceptable.

The safe rule is short. The more the text matters, the more it needs a human pass, and the things that always need checking are account numbers, identity numbers, dates, proper nouns and any figure with a currency attached.

Why Vietnamese is harder than English

English has twenty-six letters, while in Vietnamese every vowel can carry a tone mark and a diacritic on top of that, so the machine has far more distinct shapes to tell apart.

The hardest pair is hoi against nga. Both marks are small, both sit in the same place above the letter, and they differ only in the curve of a stroke, so a slightly soft photograph leaves the machine guessing on probability rather than seeing the shape.

Which is why the language has to be set before reading. Run a Vietnamese page through the English pack and every diacritic disappears, leaving you a page stripped of tone marks rather than a page with occasional errors.

Handwriting, essentially not

This recogniser is built for printed type, meaning shapes that repeat identically from page to page. Handwriting differs from person to person, and one person does not write the same word twice the same way.

What comes back from a handwritten page is usually broken enough that correcting it takes longer than typing it. For a notebook or a filled-in form, retyping is the fast route.

Quality belongs to the shooting, not the reading

Nearly every poor result traces back to the picture that went in, so twenty extra seconds while shooting saves several minutes of correction later.

  • Shoot square on, lens parallel to the paper, never at an angle.
  • Flatten the page; the curve near a book's spine is where letters distort most.
  • Give it enough light, and keep the shadow of your hand or phone off the page.
  • Frame the sheet closely rather than including half the desk.
  • If the shot came out on a slant, flatten it into a scan first and feed that in.

That flattening step is worth more than people expect. It removes the trapezoid and evens the light so the paper reads white, and both of those lift the confidence figure noticeably.

Two panels compared, on the left in coral a tilted phone over a skewed page with a shadow across it, on the right in mint a phone held parallel above a flat evenly lit page
Shot at an angle the page skews into a trapezoid and a shadow lands on the text; held parallel, every line stays straight

Where the text goes next

The words land in a box you can edit on the spot, which is deliberate, because the quickest place to fix a wrong letter is where you are already looking rather than after Word is open.

From there are three ways out. Copy it straight into whatever needed it, download a .txt when plain words are all you want, or download a .docx if it has to be sent on or formatted further.

One confusion is common here. If what you actually need is the whole picture placed inside a document rather than the text pulled out of it, that is a different job, and the piece on putting images into Word covers that branch.

Does the picture go to a server

Most text-reading sites on the web take your picture and send it to their servers to process. For a restaurant menu that is fine; for a contract, a payslip or an identity document it deserves a moment's thought.

The tool here runs the recogniser inside the browser, with both the core and the language pack served from this site's own origin. The picture goes nowhere, and no third party learns what language of document you happen to be reading either.

The cost is that the first run is slow. The core weighs about 3.7 MB and each language pack adds roughly 1.4 to 2.8 MB, and they are fetched only when you first press read in a session, so expect a few extra seconds on that first page.

Text in a video is a different job

People often look for a way to pull text out of a video and assume it is the same tool, when in fact two quite separate technologies are involved.

If the words you want are spoken aloud, that is speech recognition rather than character recognition, and an image reader does nothing for you. If the words are on screen instead, a slide or a table of figures, then capturing that frame and feeding the picture in works exactly as it does for paper.

When not to use it

There are cases where machine-read text is the wrong choice from the start, and knowing them in advance saves the attempt.

  • Papers being submitted to an authority, which want a certified copy rather than the words.
  • Anything handwritten.
  • Multi-column tables, since the text comes back line by line and the column structure usually collapses.
  • A screenshot of a PDF that already has a text layer, where selecting and copying inside the PDF does the job.

One page, thirty seconds

Shoot square on, give it light, set the language to Vietnamese, press read and check the numbers. That is enough to move a page of text out of a photograph and into wherever it was needed without retyping a word.

And if the confidence figure comes back low, do not try to repair that text. Taking the photograph again properly is usually much faster.

Lift text from an image

Read next

Read next

Browse frames, read the guide, or open the booth to pick a frame and shoot.

All articles