PDF to text
Take the words out of a PDF, in the order they are read.
Drop files here
Or paste them, or pick them from this device. Nothing is uploaded: the work happens in this tab.
Files in this tab
Nothing is loaded yet.
Options
Which pages to take the text from.
The words, page by page.
Advanced
Extract the text
Add a file to get started.
What this does
This takes the words out of a PDF in the order a person reads them, which is not the order the file draws them in. The page is cut along the gaps that have no text in them, biggest gap first, so a heading spanning two columns comes before both of them and a two-column page is not read straight across. You can have it as plain text, or as JSON that says where on the page each run of text was.
What it will not do
Every tool here says what it cannot do, in its own words, before you rely on it.
- A scanned page holds a picture of words, not words. Run OCR over it first, and this will read the layer OCR adds.
- Reading order is worked out from where the text sits on the page. It is right for columns, headings and tables, and it can be wrong on a page laid out like a poster.
- Columns are recognised from three lines or more of them. Two short columns side by side are read across, the way a table is.
- Text drawn sideways — a rotated watermark, a turned table header — is kept whole and read after the upright text, rather than being folded into a line it crosses.
Questions people ask
Nothing came out. What went wrong?
The page is almost certainly a scan — a picture of words rather than words. Run it through Make a scan searchable first, and this will read the text layer that adds.
Does it keep columns and tables in the right order?
Columns, yes: they are found from the geometry and read one at a time. Tables are read row by row. A page laid out like a poster, with text at angles or in scattered blocks, has no reading order to find, and it will show.
Is my file uploaded anywhere?
No. The work happens in this browser tab, on your own machine. There is no server in this product to send a file to, which is why the promise is checkable rather than something you have to take on trust: open your browser's network panel and watch it do the work without making a request.
Is there a size limit, or a limit on how many files?
None that we impose. The limit is your own machine's memory, because that is where the work happens — a few hundred megabytes is comfortable on a laptop, and less on a phone. There is no daily task counter, no queue and no account.
What does it cost, and what is the catch?
Nothing, and there is no paid tier that unlocks features: the whole product is the free one. That is affordable because running it costs us nothing — your computer does the work — so there is no per-file cost to recover, no advertising and no data to sell.