Guide
How to OCR a scanned PDF
If you cannot highlight a word, you are looking at pixels, not text. Optical character recognition (OCR) reads those pixels and stores a text layer (or a new file) so Find, copy, and screen readers have something to work with. Accuracy follows the scan: contrast, language, skew, and whether the page is upright. Rotate sideways pages before you OCR — see How to rotate or reorder PDF pages.
OCR vs PDF to Word
- OCR a PDF when you want to keep a PDF: search inside the scan, copy a sentence, or meet a “searchable PDF” portal rule. The pages should still look like the original.
- PDF to Word when you need to rewrite paragraphs, restyle a letter, or reuse a table in an editable document. That path guesses Word structure. It is a different checklist: PDF to Word options.
Opening a scan in Word can run OCR as a side effect and dump the result into .docx. That is useful for editing. It is the wrong output if a court, a school, or a vendor asked for a searchable PDF. Run Recognize Text / OCR / Make Searchable in a PDF app and export PDF when that is the deliverable.
Mac Preview can rotate and reorder. It does not turn a scan into a hidden text layer. Plan on Word, a PDF app your workplace already licenses, or a local OCR tool — not on Preview’s thumbnail view.
Keep the original scan. OCR is a new file (or a new layer you cannot perfectly undo). Names, amounts, and legal terms still need a human pass. When a 5 becomes an S, you fix it from the untouched image, not from yesterday’s “searchable” export.
Desktop first for secrets
IDs, payroll, health records, tax packets, unpublished research, and unsigned contracts should not go to a random OCR website. Use software that runs on a machine you control.
- Copy the scan. Confirm pages are upright and in order.
- In your PDF app, look for Recognize Text, OCR, Enhance Scans, or Make Searchable. In Word: File → Open the PDF, then save as
.docxonly if you wanted Word; otherwise use the PDF app’s searchable-PDF export. - Set the document language. Mixed English/another script usually needs the right language pack or a second pass on those pages.
- Prefer “searchable image” / keep original appearance over “text only,” unless you truly want a reflowed document.
- Save a new name such as
minutes-searchable.pdf. Do not overwrite the scan.
If the app offers downsample-during-OCR, leave it off until you have a good text layer. Shrinking the image in the same pass hides whether a bad word came from OCR or from a crushed photo.
Browser OCR, with friction
In-browser OCR is convenient on a borrowed PC. Treat the convenience as the product, not “unlimited free recognition.” Typical friction — none of it is a bug you can search away:
- Account wall or email capture after the first file.
- Page cap, file-size cap, or a queue that throttles large scans.
- Watermark, “preview only,” or a download that is a different page count than you uploaded.
- Language locked to one default, so accented names and non-Latin text come out as noise.
Use those tools only for files you would paste into a public chat. After download, confirm you can select a word you know is on page 1, then search a word from the last page. If either fails, the job did not finish — do not assume the pretty preview meant a text layer was written.
Scan quality that OCR can use
- Aim near 300 dpi for typed text. Far below that, stems of letters collapse. Far above (phone “max quality” into a 40 MB page) slows OCR without a matching accuracy jump — crop and rescan rather than feeding a giant, blurry photo.
- Straighten before OCR. A few degrees of skew is normal; a diamond-shaped page is not.
- High contrast: black type on white. Dark mode screenshots, gray-on-gray photocopies, and photos taken under a shadow fail in predictable ways (thin letters vanish; stamps become blobs).
- One column is easier than a magazine spread. Multi-column pages often concatenate lines across columns in Find.
- Handwriting, stamps, and signatures are drawings. OCR may ignore them or invent words. Do not treat a recognized signature line as legally meaningful.
A clean 6-page scan beats OCR on a 30-page packet of blanks, duplicates, and thumb-over-the-lens frames. Delete junk pages first if the leftover file is what you will keep.
Search, size, and other pitfalls
- Find still misses a word you can see. Wrong language, decorative font, or the engine wrote garbage aligned to the image. Check that page at 200% zoom against the original. Fix critical strings by hand in a PDF editor, or re-OCR that page only.
- The file got larger. Common. You still have the image, plus a text layer (and sometimes a second raster). If a portal has a size cap, OCR first so search works, then a mild compress — from the searchable file, after you verified Find. Crushing the scan before OCR is how you lock in muddy type.
- You can select text but copy-paste is out of order. Column or table layout. Fine for search; poor for reuse. Convert that page with the Word checklist if you need a table, not another OCR pass.
- Hidden text does not match what you see. Some “enhance scan” modes replace the image with a cleaned rendering. Read the page, do not only test Find.
- Encrypted source. Unlock a working copy with a password you already know, OCR that, then protect the output if you still need a lock. See password-protect a PDF you control.
Stop when Find works on a few known phrases and a couple of numbers (invoice totals, dates, case IDs). Perfect OCR on a faded photocopy is not a realistic finish line. Keep the original scan next to the searchable export.