OCR PDF: Scanned PDF to Searchable Text
Recognise the text in scanned PDFs and photos. Get a searchable PDF that looks exactly like the original, plus plain text and a Word file. Recognition runs on your device, so documents are never uploaded.
Results appear here. Files are processed on your device and never uploaded.
How it works
- Add a scanDrop a scanned PDF, or a photo of a page taken with your phone. Choose the language the document is written in.
- RecognizeEach page is rendered at high resolution and read by Tesseract, an open-source OCR engine, running inside your browser.
- Get a searchable PDFAn invisible text layer is placed exactly over the words, so the page looks unchanged but you can search, select and copy.
- Export the textDownload plain text or a Word document, or copy the text straight from the box.
What OCR is and when you need it
A scanner or phone camera produces a picture of a page. To a computer that picture is just coloured pixels: you cannot search it, copy a sentence from it or have a screen reader read it aloud. Optical character recognition (OCR) looks at the shapes of the letters and turns them back into real text. A searchable PDF combines both worlds. The original image stays visible, so signatures, stamps and layout are preserved, while an invisible layer of text sits on top of each word.
That makes old contracts, receipts, letters and book chapters useful again. Press Ctrl+F to find a clause, paste a table of figures into a spreadsheet, or feed the text to a translation tool or an AI assistant. Many document management systems and email searches also index the hidden text, so the files become findable later.
Getting accurate results
- Pick the right language. The engine uses a language model to decide between similar letters. English is built in; other languages download once (a few megabytes) and are then cached by your browser.
- Use a sharp, straight scan. 300 DPI is ideal. For phone photos, lay the page flat, fill the frame and avoid shadows and glare.
- Check the confidence score. The tool reports an average confidence. Above 85% usually means very few errors; below 60% suggests a blurry scan or the wrong language.
- Printed text works best. Handwriting, decorative fonts and text over busy backgrounds are recognised poorly by every OCR engine.
Privacy
Scanned documents are often sensitive: IDs, medical letters, bank statements. This tool never sends them anywhere. The OCR engine is downloaded to your browser like any other part of the page, and recognition happens on your own processor. Only the language data for languages other than English is fetched from a public CDN, and that request contains no part of your document.
Frequently asked questions
Will the searchable PDF look different from my scan?
No. The original page images are kept untouched. The recognized text is added as an invisible layer, which is the same technique used by Adobe Acrobat and professional scanners.
How long does OCR take?
Usually one to three seconds per page on a laptop, and a little longer on phones. The first run also loads the engine, which takes a few seconds.
Can it read handwriting?
Only neat, printed-style handwriting, and not reliably. The engine is trained on printed text.
My PDF already has text. Do I need OCR?
If you can already select text in the PDF, it has a text layer and OCR is not needed. Use PDF to Word to extract it instead.