How to Make a PDF Searchable, Even a Scan
To make a PDF searchable, run OCR on it so the words in the scanned images become real text. PDFDEN's free OCR PDF tool reads each page in your browser and adds an invisible text layer on top. The PDF looks exactly the same, but you can now search it and copy from it.
OCR PDF: Make a scanned PDF searchable. Free, and your file never leaves your device.
Open OCR PDFSteps
Open OCR PDF and drop in the scanned PDF.
Choose the document language: English, Spanish, French, German, Italian, Portuguese or Dutch.
Start the OCR. It takes a few seconds per page, so long documents need a while.
Download the searchable PDF. You can also save the recognized text as a .txt file.
Why a scanned PDF is not searchable
A scanner saves each page as a photo. You can read the words, but to the computer it is just pixels, so Ctrl+F finds nothing and you cannot select a line.
OCR (optical character recognition) reads the shapes of the letters and turns them into text. PDFDEN places that text as an invisible layer exactly over the original page. You still see the scan, while search, copy and screen readers use the hidden text.
A quick test: try to select a word. If you can't, the PDF needs OCR. If you can, it already has a text layer and running OCR again adds nothing.
Phone scanning apps and office scanners sometimes run OCR for you. Check before you process a file twice.
Get better OCR results
OCR is only as good as the scan. A few things help:
- Scan at 300 DPI where you can. Low-resolution scans produce more mistakes.
- Straighten pages first. Rotate PDF fixes sideways or upside-down pages, which OCR reads poorly.
- Pick the right language, so accented letters are recognized.
Handwriting and decorative fonts are hard for any OCR engine. Expect clean printed text to come out well and handwritten notes to come out patchy.
OCR PDF: Make a scanned PDF searchable. Free, and your file never leaves your device.
Open OCR PDFWhat to do with the searchable text
Once the text layer is there, the rest of the toolkit can use it. PDF to Text exports it to a .txt file. Compare PDF can diff two scanned versions of a contract. Redact PDF can search for a name and box every match.
The searchable PDF is about the same size as the scan, since the text layer is small. If it is too big to send, run it through Compress PDF.
Make a PDF searchable without Acrobat
Built-in options are limited. Recent versions of macOS can recognize text in images through Live Text, but that helps you copy text on screen and does not save a text layer into the PDF.
Desktop OCR is usually part of paid PDF software. The free PDF OCR tool runs the Tesseract engine in your browser, so the scan is never uploaded, which matters for the IDs, statements and records people tend to scan.
Troubleshooting OCR results
Search misses a word you can see. OCR may have misread a letter, turning "invoice" into "lnvoice". Search for part of the word instead, or rescan the page at a higher resolution.
Accented letters come out wrong. The wrong language was selected. Run OCR again with the right one from the list.
The text comes out in odd order. Multi-column layouts and tables can confuse the reading order. The search still works, but copied text may need tidying.
It is taking a long time. OCR runs page by page on your own computer. Keep the tab open and let the OCR tool finish. To test settings, extract a few pages with Extract PDF Pages and run those first.