Guide

How to OCR a scanned PDF

Updated September 2026 · Make a picture-of-a-page searchable and selectable. OCR is not the same job as converting to Word, and it is not a reason to upload a private file.

If you cannot highlight a word, you are looking at pixels, not text. Optical character recognition (OCR) reads those pixels and stores a text layer (or a new file) so Find, copy, and screen readers have something to work with. Accuracy follows the scan: contrast, language, skew, and whether the page is upright. Rotate sideways pages before you OCR — see How to rotate or reorder PDF pages.

OCR vs PDF to Word

Opening a scan in Word can run OCR as a side effect and dump the result into .docx. That is useful for editing. It is the wrong output if a court, a school, or a vendor asked for a searchable PDF. Run Recognize Text / OCR / Make Searchable in a PDF app and export PDF when that is the deliverable.

Mac Preview can rotate and reorder. It does not turn a scan into a hidden text layer. Plan on Word, a PDF app your workplace already licenses, or a local OCR tool — not on Preview’s thumbnail view.

Keep the original scan. OCR is a new file (or a new layer you cannot perfectly undo). Names, amounts, and legal terms still need a human pass. When a 5 becomes an S, you fix it from the untouched image, not from yesterday’s “searchable” export.

Desktop first for secrets

IDs, payroll, health records, tax packets, unpublished research, and unsigned contracts should not go to a random OCR website. Use software that runs on a machine you control.

  1. Copy the scan. Confirm pages are upright and in order.
  2. In your PDF app, look for Recognize Text, OCR, Enhance Scans, or Make Searchable. In Word: File → Open the PDF, then save as .docx only if you wanted Word; otherwise use the PDF app’s searchable-PDF export.
  3. Set the document language. Mixed English/another script usually needs the right language pack or a second pass on those pages.
  4. Prefer “searchable image” / keep original appearance over “text only,” unless you truly want a reflowed document.
  5. Save a new name such as minutes-searchable.pdf. Do not overwrite the scan.

If the app offers downsample-during-OCR, leave it off until you have a good text layer. Shrinking the image in the same pass hides whether a bad word came from OCR or from a crushed photo.

Browser OCR, with friction

In-browser OCR is convenient on a borrowed PC. Treat the convenience as the product, not “unlimited free recognition.” Typical friction — none of it is a bug you can search away:

Use those tools only for files you would paste into a public chat. After download, confirm you can select a word you know is on page 1, then search a word from the last page. If either fails, the job did not finish — do not assume the pretty preview meant a text layer was written.

Scan quality that OCR can use

A clean 6-page scan beats OCR on a 30-page packet of blanks, duplicates, and thumb-over-the-lens frames. Delete junk pages first if the leftover file is what you will keep.

Search, size, and other pitfalls

Stop when Find works on a few known phrases and a couple of numbers (invoice totals, dates, case IDs). Perfect OCR on a faded photocopy is not a realistic finish line. Keep the original scan next to the searchable export.