Skip to content
Pditor

How to extract text from a scanned PDF with OCR

A scan is just a picture of words. OCR reads the picture and gives you the words back.

Updated June 7, 2026

Do it now with OCR — free, no signup. Open →

Scan vs text: the quick test

If you can’t select a single word in a PDF, it’s a scan — a photograph of a page, with no text underneath. That’s why search finds nothing and copy-paste comes up empty. OCR (optical character recognition) looks at the image, recognizes the shapes as letters, and hands you back real, selectable text.

Get the best accuracy

OCR quality depends mostly on the input. A few things make a big difference:

  • Resolution — 300 DPI scans read far better than a phone photo.
  • Straightness — de-skew crooked pages before running OCR.
  • Language — telling the recognizer the right language lets it use the correct alphabet and dictionary.

Pditor runs a fast local recognizer first and, on pages where it isn’t confident, falls back to an AI vision model to recover faint print or handwriting that the first pass missed.

Steps

  1. Confirm the PDF is a scan (nothing selects).
  2. Open the OCR tool and add the file.
  3. Pick the document language.
  4. Run OCR — difficult pages use the AI fallback automatically.
  5. Proofread the result, paying attention to names and numbers.

What to do with the text

Once you have the recognized text you can search it, copy it into another document, or feed it to other tools. If you need a fully editable document rather than just the text, follow up with PDF → Word to rebuild the layout as paragraphs and headings you can edit directly.

Frequently asked questions

How do I make a scanned PDF searchable?
Run it through OCR. The recognizer reads the image of each page and produces the underlying text, which you can then search and copy. In Pditor, pick the document's language first for the best accuracy.
How accurate is OCR?
On a clean, straight, printed scan, modern OCR is highly accurate. Accuracy drops with low resolution, skew, faint print, or handwriting. Pditor falls back to an AI vision model on low-confidence pages to recover difficult text.
Can OCR read handwriting?
Printed text is most reliable. Neat handwriting can be read via the AI vision fallback, but messy or cursive writing is hard for any tool — always proofread the result.
Why can't I select text in my PDF?
Because the page is an image, not text — typical of anything scanned or photographed. There's nothing to select until OCR reconstructs the text layer.
→ Open OCR