Find and remove personal data in a PDF
Find the personal data in a document and take it out of the file for good.
What this does
A list of names is not what anybody came for. What a disclosure needs is a file with the names gone from it — out of the content stream, out of the pictures under the boxes, out of the structure tree and out of the document details. So what is found here is handed straight to Redact, which is the one thing in this product that performs a removal and then goes looking for the words again in the file it wrote, failing rather than handing you a document that looks redacted and is not.
What it will not do
Every tool here says what it cannot do, in its own words, before you rely on it.
- Removal is done by Redact, which takes the text out of the file and then looks for it again in what it wrote. It is real removal, and it is checked.
- A finding whose text does not appear in the document as written is reported and not removed. A model that paraphrased a name would otherwise cause a redaction of nothing at all, reported as a success.
- What is found is what a model found. It is a first pass and not a compliance sign-off: read the list before you accept it, and read the redacted file afterwards.
- A scanned page holds a picture of words. Nothing here can read it, and nothing here can redact it. Run OCR over it first.
- Redaction rewrites the whole document: its metadata is stripped, a protected file comes back in the clear, and what it no longer refers to is dropped.
Questions people ask
Is the text really gone, or just covered up?
Gone. The removal is done by the Redact tool, which deletes the text-drawing instructions and then searches the file it wrote for the same words. If it finds one, it fails and writes nothing rather than hand you a file that looks safe.
Can I stop it removing something?
Yes. Anything you list under “never remove” is left alone, and you see the full list of what was found before the removal runs.
What exactly is sent to the provider?
The text on the screen under “what will be sent”, and nothing else of your document. It is not a summary of the request — it is the request. Nothing is sent until you press send, and the page cannot send it before showing it to you, because the code refuses.
What does this cost?
Pagecraft charges nothing and never will. Your provider charges you for what you send it, at their rates, on your own account — and you are shown the token count and an estimate in dollars before you send anything. A model on your own machine costs nothing at all.
Do I have to pay for an API key?
No. Install Ollama or LM Studio, run a model on your own computer, and point Pagecraft at it: no key, no bill, and your document never leaves the device. That path is first-class here rather than a degraded fallback.