Make a scan searchable
Read the words on a scanned page and put them in the file, without changing how it looks.
Drop files here
Or paste them, or pick them from this device. Nothing is uploaded: the work happens in this tab.
Files in this tab
Nothing is loaded yet.
Options
Advanced
How finely the page is drawn before it is read. 300 suits most scans.
Which pages to read: '1-3, 7, 9-'.
Read the page
Add a file to get started.
What this does
A scanned page is a photograph of words: you cannot search it, select it or have it read aloud. This finds the words in the picture and writes them into the file underneath the picture of them, drawn in a mode that paints nothing — so the page gains its words without gaining a mark, and looks exactly as it did. Before reading, a crooked page can be straightened and a faint one sharpened, neither of which changes the page you get back.
What it will not do
Every tool here says what it cannot do, in its own words, before you rely on it.
- Pagecraft performs no recognition of its own: the reading is done by an engine the surrounding product supplies. This site supplies Tesseract, offered as a download you choose and keep, and it can read English, German, French, and Spanish. How well it reads them is its answer rather than ours — clean printed text comes out well, a photograph of a crumpled receipt does not — and a language this site has no model for is refused by name rather than read badly.
- The page is not changed: the words go underneath the picture of them. A bad scan stays a bad scan, and the recognition is only as good as it.
- Confidence says how sure the engine is, which is not the same as whether it is right.
- A page carrying so much as a stamp counts as having text and is left alone. Read it anyway and it will carry both its own words and the recognised ones.
- On a crooked page the words are written along the crooked lines, so that selecting one lands on it. Every word is still found by a search, but a reader listing the text of such a page reads it in the order the words sit at rather than the order they were written in.
Questions people ask
Can this site read my scan right now?
Yes, once the engine and the language you want are on this device. Both are offered above the controls with their sizes before anything is fetched, and both stay afterwards, so the second scan needs no network at all. It reads English, German, French, and Spanish; for anything else, the operation runs from the command line and from the SDK with an engine you choose.
Will it change how my scan looks?
No. The words go underneath the picture in invisible text. Straightening and sharpening are done to a copy used for reading, not to the page: a crooked page comes back exactly as crooked, and exactly as searchable.
Is my file uploaded anywhere?
No. The work happens in this browser tab, on your own machine. There is no server in this product to send a file to, which is why the promise is checkable rather than something you have to take on trust: open your browser's network panel and watch it do the work without making a request.
Is there a size limit, or a limit on how many files?
None that we impose. The limit is your own machine's memory, because that is where the work happens — a few hundred megabytes is comfortable on a laptop, and less on a phone. There is no daily task counter, no queue and no account.
What does it cost, and what is the catch?
Nothing, and there is no paid tier that unlocks features: the whole product is the free one. That is affordable because running it costs us nothing — your computer does the work — so there is no per-file cost to recover, no advertising and no data to sell.