How to make a scanned PDF searchable
· 6 minute read
A scanned PDF is a photograph of text. How OCR reads it, what makes it accurate, and how to turn a scan into a PDF you can search and copy from.
You scan a contract, open the file, press Ctrl+F and search for a name. Nothing is found, though you can see the name on the page. The reason is that the file contains a picture of the page and no text at all. Searching, copying and screen readers all need real text, so a scanned document is closed to every one of them until something reads the letters off the picture.
What OCR does
Optical character recognition, usually shortened to OCR, looks at the picture, finds the lines of text, recognises each character and writes the result out as text. A good tool then places that text invisibly under the picture. The page looks exactly the same as before, but now there is a layer of real words behind it, so a search finds them and a selection can copy them.
What makes recognition accurate
Almost everything depends on the scan. The software is reading a photograph, and a clear photograph reads well.
- Resolution. 300 dots per inch is the standard for text. Lower than about 150 and small print starts to fail.
- Straightness. A page that is skewed or photographed at an angle reads worse. Lay it flat and shoot it from above.
- Contrast. Dark text on a plain light background is ideal. Shadows, creases, coloured paper and show-through from the other side all cause errors.
- Language. Recognition uses a model of the language. A French document read as English comes out as nonsense, so pick the language the document is written in.
- Type. Printed text reads well. Handwriting generally does not.
The steps
- If you have photographs rather than a PDF, put them into one with JPG to PDF.
- Open the OCR tool, choose the file and choose the language.
- Run it. A long document takes a while, because every page is read in turn.
- Save the result and test it by searching for a word you can see on the page.
- If the file is very large, run it through Compress PDF afterwards.
Check the result before you rely on it
OCR makes mistakes, and you will not see them, because the text is invisible. A number read as a different number, or an "l" read as a "1", looks right on the page and is wrong in the search layer. For anything where accuracy matters, such as amounts, dates and names, check the figures against the picture, not against the text layer.
OCR and privacy
Many OCR services work by uploading your pages to a server. Scanned documents are exactly the kind that carry sensitive material: IDs, invoices, medical letters. The recognition can instead run on your own computer, which is slower on a modest machine but keeps the pages where they are. If you plan to remove names from a scan, make it searchable first and then use Redact, which needs real text to find the words.
Tools mentioned
- OCR PDF — Make a scan searchable by reading the words off it.
- Compress PDF — A smaller file that still reads like the original.
- JPG to PDF — Photos and scans gathered into one document, each page its image.
- Redact — Take words out of the file, not just paint over them.