Why your scanned PDF is not searchable, and how to fix it
You open a PDF, press Ctrl+F, type a word you can plainly see on the page — and nothing is found. Or you try to copy a paragraph and get nothing, or a converter says the file has no text. The document is not broken. It is a scan, and a scan is a photograph of paper. Understanding the difference explains a lot of confusing behaviour, and points to the fix.
Two kinds of PDF that look the same
A PDF made on a computer — exported from Word, printed to PDF from a browser, generated by a bank — stores its text as characters: this letter, in this font, at this position. Your PDF reader can search those characters, select them and copy them, and a screen reader can read them aloud.
A PDF made by a scanner or a phone camera stores each page as an image. You see words because your eyes read the picture; the file itself contains only pixels. To the computer there is no "INVOICE" on the page, only a pattern of dark and light dots that happens to look like one.
How to tell which one you have
Try to highlight a sentence with your mouse, or press Ctrl+F and search for a word you can see. If the text highlights and the search finds it, the PDF has real text. If the selection draws a box over the whole page, or nothing highlights at all, the page is an image.
Some documents are mixed: a report typed on a computer with a signed page scanned in at the end. Check the pages that matter to you, not only the first one.
What OCR does
OCR — optical character recognition — is software that looks at the picture of a page and works out which letters are there. A good OCR tool then adds those letters to the PDF as an invisible text layer, lined up exactly over the words in the image. The page looks the same as before, but now search, selection and copying work, because there is real text for them to find.
That text layer is also what lets other tools work with a scan: converting it to Word or Excel, summarising it, or finding every place a name appears so it can be redacted.
How accurate it is, and what to check
On a clean, straight scan of printed text, modern OCR gets nearly every word right. Accuracy falls with blur, shadows, crooked pages, low resolution, small print and decorative fonts, and handwriting is generally not read reliably at all. Choosing the right language matters too: a Filipino word read with an English-only model often comes out as a near miss.
Because the text layer is invisible, its mistakes are invisible. After OCR, search the document for the names, dates and amounts that matter and make sure they are found. A zero read as the letter O is the kind of error that hides until someone searches for an account number and it is not there.