PDF to Text OCR
Extract text from scanned PDFs in your browser with pdf.js rendering and Tesseract.js OCR. Private, free and no uploads.
🔒 Runs entirely in your browser — nothing is uploadedExtract text from scanned PDFs privately
Scanned PDFs often look like normal documents, but the pages are really pictures. That means you cannot select a paragraph, search for a number or paste a section into another app. This PDF to Text OCR tool converts those image-only pages into editable text in your browser. It is useful for old scans, paper forms, receipts, book pages, signed documents and any PDF where copy and search do not work because the text is trapped inside pixels.
The process is fully client-side. pdfjs-dist loads your local PDF file and renders each page to an HTML canvas, much like a browser PDF viewer displays a page on screen. Tesseract.js then reads that canvas and returns the recognized words. The tool repeats the process page by page, adds a clear separator between pages and places the combined text in an editable text area. Your PDF and extracted text stay on your device; there is no upload, account or server queue.
Tips for better PDF OCR
OCR accuracy depends on the quality of the scan. Straight pages with dark text on a light background work best. Blurry pages, heavy shadows, handwriting, decorative fonts and low resolution scans can reduce accuracy. If the source document is available, scan it again at a higher resolution before running OCR. For camera scans, keep the paper flat, crop away the table or background and make sure the page is evenly lit.
Choose the language that matches most of the document. The first time you use a language, the browser may need to download the OCR engine and language model, so the first run can take longer. After that, browser caching usually makes future jobs faster. Always review the output before using it for invoices, legal documents, names, addresses or numbers, because OCR can confuse similar characters such as O and 0 or I and 1.
Copy, edit and archive the result
When recognition finishes, you can edit the result directly in the text area. Page separators make it easier to compare the extracted text with the original PDF and remove headers, page numbers or scanning artifacts. Use the copy button for quick notes, email drafts and searches, or download a plain text file when you want to archive the OCR output next to the source PDF. For very long scanned documents, run the tool while keeping the tab open and avoid switching to power-saving modes that may pause browser work.
How to use
- Add a PDFDrop one scanned PDF onto the upload box, or click to choose it.
- Choose languageSelect the OCR language that best matches the document.
- Run OCREach page renders to a canvas and is recognized locally in sequence.
- Copy or downloadReview the combined text, then copy it or save it as a .txt file.