How to Extract Text From a Scanned PDF
A scanned PDF is a stack of page images, so there is no text to select or copy. To get the words out you need OCR. The free PDF to Text tool does both jobs in your browser: it reads the text layer of digital pages directly and runs OCR only on pages that are pictures.
Why can't I copy text from a PDF?
Usually because the page is a scan or a photo saved as PDF. The file shows letters, but underneath there are only pixels. Some PDFs also have copying disabled by their author. If you can't highlight individual words, the page needs OCR before the text can be copied.
How do you tell if a PDF is scanned?
Open it and try to select a sentence. If single words highlight, the page has a text layer. If the whole page selects like one picture, or nothing selects at all, it is a scan. Many PDFs mix both: typed pages with a few scanned attachments at the end.
How do you OCR a PDF for free?
- Open the toolGo to PDF to Text.
- Add the PDFChoose a file of up to 50 MB.
- Pick the OCR languageSelect the language of the document so scanned pages are read correctly.
- ExtractText-layer pages come out instantly; scanned pages are recognized one by one.
- Copy or downloadCopy the result or save it as a text file.
How do you clean up the extracted text?
Scans often break sentences at the end of every printed line and split words with hyphens. If you work with the text in code, a few lines of JavaScript join them back into paragraphs. Paste the OCR result into the text variable and run it in your browser console or Node.js.
const text = `paste the OCR result here`;
const cleaned = text
.replace(/-\n(?=\p{Ll})/gu, '') // re-join hyphen-split words
.replace(/(?<!\n)\n(?!\n)/g, ' ') // single line breaks become spaces
.replace(/[ \t]{2,}/g, ' '); // collapse repeated spaces
console.log(cleaned);Is it safe to use with private documents?
Yes. The PDF is processed by your own browser and never uploaded to a server, which makes it a good fit for contracts, invoices and medical letters. Better scans give better text, so see what really affects OCR accuracy before you rescan anything.
What if I only have photos of the pages?
Then skip the PDF step and use the image tools directly: copy text from an image for one page, or Multi Image OCR for up to 20 pages at once.
Extract text from your PDF
Text layer first, OCR for scans. All in your browser.