Slate PDF

Converting

How to Copy Text From a Scanned PDF

You cannot select text in a scan because there is none. Here is how OCR turns the picture into characters you can copy, search and edit.

June 19, 20262 min read

You open a PDF, try to select a paragraph, and your cursor draws a rectangle instead of highlighting words. Nothing is broken — the file simply has no text in it. What you are looking at is a photograph of a page.

Make the text selectable

  1. 1Open the OCR tool and add your scanned PDF.
  2. 2Run it. Every page image is read and an invisible text layer is written over it.
  3. 3Download the result and try selecting a paragraph — it should now highlight.
  4. 4Search for a word you know is in the document to confirm the layer is there.

What OCR does to the file

Nothing you can see. The scanned image stays exactly as it was, and the recognised characters are placed invisibly on top, each positioned over the word it came from. That is why selecting text in an OCRed scan highlights in the right place even though the words you see are still part of the picture.

What hurts accuracy

ProblemEffect
Low resolution (under 200 DPI)Characters blur together; accuracy drops sharply
Skewed pagesLines are hard to segment — the biggest single factor
Faint or faxed originalsThin strokes disappear entirely
HandwritingGenerally not recognised at all
Unusual or decorative fontsHigher error rates than plain text faces
Dense tables with faint rulesColumn boundaries confuse the layout analysis

If the result is poor, the fix is almost always at the scanner rather than in the software: rescan at 300 DPI, straight, in greyscale.

The quick alternative for one paragraph

If you only need a few lines and OCR feels like overkill, most phones will now read text out of a photograph directly — point the camera at the screen, select the text, copy. It is crude, and for a couple of sentences it is faster than anything else.

After OCR

Once there is a text layer, the whole toolkit opens up: the document is searchable, convertible to Word or Markdown, and readable by screen readers. It is worth running on anything going into long-term storage, because a scan you cannot search is a scan you will never find again.

Do it now

OCR PDF

Extract text from scanned documents. It runs in this browser tab — your file is not uploaded anywhere.

Open OCR PDF

Frequently asked questions

Why can I not select text in my PDF?+

Because it is a scan — an image of a page rather than text. Run OCR to add a text layer.

Does OCR change how the page looks?+

No. The image is untouched and the recognised text is invisible.

How accurate is OCR?+

On clean printed text at 300 DPI, very. On faint, skewed or handwritten material, unreliable enough that you must check.

Can I OCR a photo I took of a page?+

Yes. Convert it to PDF first, and use your phone's scanner mode when capturing so the page is cropped and squared.

Keep reading