Slate PDF

Converting

How to Convert a Scanned Document to a Searchable PDF

Convert a scanned document to a searchable PDF: from paper, from scanner JPGs or from an image-only PDF, with the right settings and OCR at every step.

HM

Hamza M.

PDF Workflow Editor

Updated October 4, 20265 min readReviewed October 4, 2026

A scan is a photograph of a page. It looks like a document and behaves like a picture: you cannot search it, you cannot copy a line out of it, and a screen reader finds nothing on it. To convert a scanned document to PDF properly, you need two things: the pages in one PDF file, and OCR, which adds the text layer that closes that gap.

How do I convert a scanned document to a PDF?

It depends on what you have. Paper: scan it with your phone in a browser scanner and save a searchable PDF in one go. Scanner JPGs: import them into the scanner or convert them to PDF. An image-only PDF: run it through an OCR tool to add a text layer. All three run free in the browser.

You haveDo thisResult
Paper pagesScan PDF with your phone camera, Recognize text (OCR) ticked when savingOne searchable PDF
JPG or PNG files from a scanner or phoneImport them into Scan PDF, or use JPG to PDF and then OCROne searchable PDF
TIFF files from an office scannerTIFF to PDF, then OCROne searchable PDF
A PDF where the text cannot be selectedOCR PDFSame pages, now searchable

From paper: scan straight to a searchable PDF

If the document is still on paper, skip the separate steps. The Scan PDF tool on your phone does capture, straightening, cleanup and OCR in one pass:

  1. 1Open Scan PDF in your phone’s browser and tap Open camera.
  2. 2Hold the phone over each page; it detects the edges and captures automatically. Turn to the next page and it captures again.
  3. 3On the review screen, fix any page with Crop or Rotate, and choose a filter: Auto color for most documents, B&W for plain printed text.
  4. 4Tap Save PDF, choose a page size (A4 or Letter for documents that will be printed), leave Recognize text (OCR) ticked and pick the document language.
  5. 5Tap Save as PDF and download it. Text recognition runs on the phone; the language model downloads once, and nothing else is fetched or uploaded.

From scanner images: the full workflow

If a flatbed or office scanner has already given you image files, work through these steps in order:

  1. 1Check the scan settings. 300 DPI grayscale is right for ordinary printed text.
  2. 2Convert the images to PDF if your scanner produced JPGs or TIFFs rather than a PDF. Importing JPGs into Scan PDF also straightens skewed pages and evens out the lighting on the way.
  3. 3Merge the pages into one document if they came out as separate files.
  4. 4Remove blank pages and rotate anything sideways.
  5. 5Run OCR to add a text layer.
  6. 6Compress last, once everything else is settled, and gently.

Scanner settings that matter

SettingRecommendedWhy
Resolution300 DPIBelow 200 the OCR accuracy falls away; above 400 adds size, not accuracy
Colour modeGrayscaleMuch smaller than colour, with no loss on black-and-white originals
FormatPDF or JPEGTIFF is enormous and rarely needed
StraighteningOnSkewed text is the single biggest cause of OCR errors

What OCR actually does

OCR reads the picture, recognises the shapes as characters, and writes an invisible text layer positioned over the visible image. The page still looks like the scan, because the original picture is untouched, but a search now finds words, and selecting text copies real characters.

It is not perfect. Clean printed text at 300 DPI is recognised very well; a faint fax, a handwritten note or a dense table with faint rules will produce errors. Choose the right language before running it, because the engine uses it to decide which characters to expect.

Why the order matters

OCR reads the image as it is when you run it. Compress first and OCR reads a degraded picture, which produces more mistakes than necessary. Do every lossless step first, run OCR, then compress. A light compression keeps the text layer, but an aggressive one can turn pages back into plain images; if the converted file must be small, pick a smaller file size when you create it rather than squeezing it hard afterwards, and repeat the search test after compressing.

Privacy for sensitive scans

Scans are often the most sensitive files people have: tax papers, medical letters, ID. Every step above, including the OCR, runs in your browser on your device, so the pages are never uploaded to a server. If you need to send the result, the scanner can add a password when saving, or you can protect the finished PDF separately.

Do it now

OCR PDF

Extract text from scanned documents. It runs in this browser tab — your file is not uploaded anywhere.

Open OCR PDF

Frequently asked questions

How do I convert a scanned document to a searchable PDF?+

Run OCR on it. If you are starting from paper, scan it with a phone scanner that has OCR built into the save step; if you already have a scanned PDF, use an OCR PDF tool to add the text layer.

What DPI should I scan at?+

300 DPI for text. Higher adds file size without helping recognition; lower starts to cost you accuracy.

Can OCR read handwriting?+

Not reliably. It is built for printed characters. Neat block capitals sometimes work; ordinary cursive does not.

Does OCR change how my scan looks?+

No. The text layer is invisible and sits behind the image. The page looks exactly the same.

Should I OCR before or after compressing?+

Before. OCR on a compressed image makes more mistakes. Compress gently afterwards and search the file again to make sure the text layer survived.

Keep reading

HM

Hamza M. · PDF Workflow Editor

Hamza has spent a decade working with print and digital documents — from prepress production to everyday PDF cleanup — and now writes practical, tested guides for Slate PDF.

Read more about how this site is run