Make a scanned PDF searchable (OCR)
Add a hidden text layer to scanned pages, so you can search, select and copy their text. The pages look exactly the same.
Processed on our server, then deleted
Drop a scanned PDF here
Or drag it anywhere on this page. Up to 10 MB.
Sent to our server for this one job, then deleted. Never stored or logged.
One PDF up to 10 MB, up to 10 pages per job, English text. Processed on our server, then deleted. 10 free server jobs a day per person, shared by all server tools. Bigger jobs: the PDF OCR API (paid per call).
How to make a PDF searchable
- Choose a scanned PDF up to 10 MB, or drag it onto the page.
- Check the pages to read: all of them for short files, or up to 10 at a time.
- Press Make searchable, then download the PDF. You can also copy or save the recognised text.
What OCR does to a scanned PDF
A scanner or phone app saves each page as a picture. You can read it, but your computer can't: search finds nothing and you can't copy a sentence. Optical character recognition (OCR) looks at the picture, finds the letters and words, and writes them into the PDF as invisible text placed exactly over the words you see.
Our server renders each chosen page at 200 to 300 dpi, reads it with the open-source Tesseract engine, and lays the text over the original page without changing it. The recognised text is also shown here, so you can copy it straight away.
Tips
- Scan at 300 dpi in black and white or grey for the best accuracy.
- Pages upside down or sideways? Turn them first with Rotate PDF (in your browser).
- Got a photo or screenshot instead of a PDF? Use Image to text.
Questions
Is my PDF uploaded?
Yes. Unlike most tools on this site, this one runs on our server: your PDF is sent over HTTPS, read by the OCR engine, and the searchable copy is sent straight back to your browser. The file exists on the server only while the job runs, in memory and in a private temporary folder that is deleted the moment the job ends. We don't store the PDF, the result or the recognised text, and we never log file contents.
What does OCR change in my PDF?
Nothing you can see. Each page is read by the open-source Tesseract OCR engine, and the words it finds are laid over the page as invisible text in the right places. The original page stays underneath, untouched, so the PDF looks the same but becomes searchable, selectable and copyable.
Why only 10 pages at a time?
Reading a page takes a few seconds of server time, and each job has to finish within about a minute. For a longer document, run the tool again on the downloaded file with the next pages, for example 11-20. Pages already done are kept as they are.
Which languages work?
English, and other languages written in the same Latin letters without accents, with lower accuracy. Accented letters and other scripts aren't recognised reliably yet.
My pages already have text. What happens?
Pages that already contain text are skipped, so a PDF made by a word processor stays as it is. Tick Read pages that already have text to OCR them anyway, for example when a scan has a bad text layer.
How accurate is it?
Clean, straight scans at normal resolution come out very well. Photos taken at an angle, faint or blurry print, handwriting and complex tables give more mistakes. Check names and numbers before relying on them.
Is there a limit?
Files up to 10 MB, up to 10 pages per job, and 10 free server jobs a day per person, shared by all of our server tools. Only jobs that succeed count, and the count resets at midnight UTC. Bigger jobs: the PDF OCR API, paid per call.
What is the human check?
Before a job runs, Cloudflare Turnstile checks that a person, not a bot, is using the tool. It is usually invisible; sometimes it asks you to tick a box.