You photograph a page of a book, a notice on a board or a page of a contract, and then you need the words out of it to paste somewhere else. Retyping costs fifteen minutes a page, and one wrong digit in an account number costs a good deal more than that.
Lifting text out of a picture has a name in the trade, OCR, and it has run inside a browser for some years now. What is worth saying is not that it works, but that it fails in a way that slips past people.
How a machine reads
The machine does not read in the sense a person does. It breaks the page into blocks, each block into lines, each line into clusters of marks, then compares every cluster against shapes it has learned and picks whichever letter fits best.
Because every step is a guess, the result always carries a number saying how sure it was. The tool here takes the mean confidence per word, and on a clean scan that number usually sits in the eighties or above.
A hurried photograph of a slightly curved page drops it to around fifty. At that level the text that comes back is a mix of right and wrong, and the wrong part does not announce itself.
The dangerous errors are the ones that look fine
If the machine returns a line of nonsense you know at once it has failed and you retype it. The trouble is the places where it read almost correctly: a zero becoming a capital O, a one becoming a lowercase l, a decimal point becoming a comma inside a figure.
Errors like that leave a page that reads cleanly, so nobody goes back over it. That is why the tool says so when confidence falls below its threshold rather than quietly handing over something that looks acceptable.
The safe rule is short. The more the text matters, the more it needs a human pass, and the things that always need checking are account numbers, identity numbers, dates, proper nouns and any figure with a currency attached.
Why Vietnamese is harder than English
English has twenty-six letters, while in Vietnamese every vowel can carry a tone mark and a diacritic on top of that, so the machine has far more distinct shapes to tell apart.
The hardest pair is hoi against nga. Both marks are small, both sit in the same place above the letter, and they differ only in the curve of a stroke, so a slightly soft photograph leaves the machine guessing on probability rather than seeing the shape.
Which is why the language has to be set before reading. Run a Vietnamese page through the English pack and every diacritic disappears, leaving you a page stripped of tone marks rather than a page with occasional errors.
Handwriting, essentially not
This recogniser is built for printed type, meaning shapes that repeat identically from page to page. Handwriting differs from person to person, and one person does not write the same word twice the same way.
What comes back from a handwritten page is usually broken enough that correcting it takes longer than typing it. For a notebook or a filled-in form, retyping is the fast route.
Quality belongs to the shooting, not the reading
Nearly every poor result traces back to the picture that went in, so twenty extra seconds while shooting saves several minutes of correction later.
- Shoot square on, lens parallel to the paper, never at an angle.
- Flatten the page; the curve near a book's spine is where letters distort most.
- Give it enough light, and keep the shadow of your hand or phone off the page.
- Frame the sheet closely rather than including half the desk.
- If the shot came out on a slant, flatten it into a scan first and feed that in.
That flattening step is worth more than people expect. It removes the trapezoid and evens the light so the paper reads white, and both of those lift the confidence figure noticeably.
Where the text goes next
The words land in a box you can edit on the spot, which is deliberate, because the quickest place to fix a wrong letter is where you are already looking rather than after Word is open.
From there are three ways out. Copy it straight into whatever needed it, download a .txt when plain words are all you want, or download a .docx if it has to be sent on or formatted further.
One confusion is common here. If what you actually need is the whole picture placed inside a document rather than the text pulled out of it, that is a different job, and the piece on putting images into Word covers that branch.
Does the picture go to a server
Most text-reading sites on the web take your picture and send it to their servers to process. For a restaurant menu that is fine; for a contract, a payslip or an identity document it deserves a moment's thought.
The tool here runs the recogniser inside the browser, with both the core and the language pack served from this site's own origin. The picture goes nowhere, and no third party learns what language of document you happen to be reading either.
The cost is that the first run is slow. The core weighs about 3.7 MB and each language pack adds roughly 1.4 to 2.8 MB, and they are fetched only when you first press read in a session, so expect a few extra seconds on that first page.
Text in a video is a different job
People often look for a way to pull text out of a video and assume it is the same tool, when in fact two quite separate technologies are involved.
If the words you want are spoken aloud, that is speech recognition rather than character recognition, and an image reader does nothing for you. If the words are on screen instead, a slide or a table of figures, then capturing that frame and feeding the picture in works exactly as it does for paper.
When not to use it
There are cases where machine-read text is the wrong choice from the start, and knowing them in advance saves the attempt.
- Papers being submitted to an authority, which want a certified copy rather than the words.
- Anything handwritten.
- Multi-column tables, since the text comes back line by line and the column structure usually collapses.
- A screenshot of a PDF that already has a text layer, where selecting and copying inside the PDF does the job.
One page, thirty seconds
Shoot square on, give it light, set the language to Vietnamese, press read and check the numbers. That is enough to move a page of text out of a photograph and into wherever it was needed without retyping a word.
And if the confidence figure comes back low, do not try to repair that text. Taking the photograph again properly is usually much faster.
