Converting
How to Convert a Scanned Document to a Searchable PDF
Convert a scanned document to a searchable PDF: from paper, from scanner JPGs or from an image-only PDF, with the right settings and OCR at every step.
Hamza M.
PDF Workflow Editor
A scan is a photograph of a page. It looks like a document and behaves like a picture: you cannot search it, you cannot copy a line out of it, and a screen reader finds nothing on it. To convert a scanned document to PDF properly, you need two things: the pages in one PDF file, and OCR, which adds the text layer that closes that gap.
How do I convert a scanned document to a PDF?
It depends on what you have. Paper: scan it with your phone in a browser scanner and save a searchable PDF in one go. Scanner JPGs: import them into the scanner or convert them to PDF. An image-only PDF: run it through an OCR tool to add a text layer. All three run free in the browser.
| You have | Do this | Result |
|---|---|---|
| Paper pages | Scan PDF with your phone camera, Recognize text (OCR) ticked when saving | One searchable PDF |
| JPG or PNG files from a scanner or phone | Import them into Scan PDF, or use JPG to PDF and then OCR | One searchable PDF |
| TIFF files from an office scanner | TIFF to PDF, then OCR | One searchable PDF |
| A PDF where the text cannot be selected | OCR PDF | Same pages, now searchable |
From paper: scan straight to a searchable PDF
If the document is still on paper, skip the separate steps. The Scan PDF tool on your phone does capture, straightening, cleanup and OCR in one pass:
- 1Open Scan PDF in your phone’s browser and tap Open camera.
- 2Hold the phone over each page; it detects the edges and captures automatically. Turn to the next page and it captures again.
- 3On the review screen, fix any page with Crop or Rotate, and choose a filter: Auto color for most documents, B&W for plain printed text.
- 4Tap Save PDF, choose a page size (A4 or Letter for documents that will be printed), leave Recognize text (OCR) ticked and pick the document language.
- 5Tap Save as PDF and download it. Text recognition runs on the phone; the language model downloads once, and nothing else is fetched or uploaded.
From scanner images: the full workflow
If a flatbed or office scanner has already given you image files, work through these steps in order:
- 1Check the scan settings. 300 DPI grayscale is right for ordinary printed text.
- 2Convert the images to PDF if your scanner produced JPGs or TIFFs rather than a PDF. Importing JPGs into Scan PDF also straightens skewed pages and evens out the lighting on the way.
- 3Merge the pages into one document if they came out as separate files.
- 4Remove blank pages and rotate anything sideways.
- 5Run OCR to add a text layer.
- 6Compress last, once everything else is settled, and gently.
Scanner settings that matter
| Setting | Recommended | Why |
|---|---|---|
| Resolution | 300 DPI | Below 200 the OCR accuracy falls away; above 400 adds size, not accuracy |
| Colour mode | Grayscale | Much smaller than colour, with no loss on black-and-white originals |
| Format | PDF or JPEG | TIFF is enormous and rarely needed |
| Straightening | On | Skewed text is the single biggest cause of OCR errors |
What OCR actually does
OCR reads the picture, recognises the shapes as characters, and writes an invisible text layer positioned over the visible image. The page still looks like the scan, because the original picture is untouched, but a search now finds words, and selecting text copies real characters.
It is not perfect. Clean printed text at 300 DPI is recognised very well; a faint fax, a handwritten note or a dense table with faint rules will produce errors. Choose the right language before running it, because the engine uses it to decide which characters to expect.
Why the order matters
OCR reads the image as it is when you run it. Compress first and OCR reads a degraded picture, which produces more mistakes than necessary. Do every lossless step first, run OCR, then compress. A light compression keeps the text layer, but an aggressive one can turn pages back into plain images; if the converted file must be small, pick a smaller file size when you create it rather than squeezing it hard afterwards, and repeat the search test after compressing.
Privacy for sensitive scans
Scans are often the most sensitive files people have: tax papers, medical letters, ID. Every step above, including the OCR, runs in your browser on your device, so the pages are never uploaded to a server. If you need to send the result, the scanner can add a password when saving, or you can protect the finished PDF separately.
Do it now
OCR PDF
Extract text from scanned documents. It runs in this browser tab — your file is not uploaded anywhere.
Open OCR PDFFrequently asked questions
How do I convert a scanned document to a searchable PDF?+
Run OCR on it. If you are starting from paper, scan it with a phone scanner that has OCR built into the save step; if you already have a scanned PDF, use an OCR PDF tool to add the text layer.
What DPI should I scan at?+
300 DPI for text. Higher adds file size without helping recognition; lower starts to cost you accuracy.
Can OCR read handwriting?+
Not reliably. It is built for printed characters. Neat block capitals sometimes work; ordinary cursive does not.
Does OCR change how my scan looks?+
No. The text layer is invisible and sits behind the image. The page looks exactly the same.
Should I OCR before or after compressing?+
Before. OCR on a compressed image makes more mistakes. Compress gently afterwards and search the file again to make sure the text layer survived.