We use cookies.This website uses essential cookies to operate core features. With your consent, we also use analytics cookies to understand traffic and improve the service. For more details, see our .
Was this tool helpful to use?
Your feedback helps us make it better
Add a searchable, selectable text layer to scanned PDFs. Supports Simplified and Traditional Chinese, English, and Japanese, with automatic deskewing and page rotation.
Choose filesUpload PDF File
PDF (Supported formats: .pdf)
Max 2.5 MB • Max 5 files
Choose the main languages used in the document.
Correct slight tilting in scanned pages.
Detect text orientation and rotate pages correctly.
Run OCR again even when the PDF already has a text layer.
Upload files and configure options, then click start processing
Overview
Understand what the tool solves, how it works, and the boundaries of its data.
PDFs created with scanners or phone cameras often contain only page images, so you cannot select text or search for keywords in a PDF reader. PDF OCR analyzes each page image and adds a hidden text layer over the original content. This makes names, reference numbers, contract terms, and book titles searchable and copyable, while helping screen readers access the text. The output remains a PDF—it is not converted to Word, and the original layout is not intentionally rearranged.
Four language options are available: Simplified Chinese + English, Traditional Chinese + English, English, and Japanese + English. Choosing the option that matches the document helps reduce errors involving similar-looking characters and punctuation. Auto deskew corrects slightly tilted scans, while auto rotate fixes upside-down or sideways pages based on text orientation. “Force OCR” is intended for files with an existing but inaccurate text layer, though it may increase processing time.
Guide
Follow the workflow and verify inputs and outputs with practical examples.
Select the document’s primary language instead of enabling every language. For Simplified Chinese documents, choose “Simplified Chinese + English”; for Traditional Chinese contracts, choose “Traditional Chinese + English”; and for Japanese documents, choose “Japanese + English.” Enable deskew if the scanned page edges are tilted, and keep auto rotate enabled when page orientations vary. Standard image-only PDFs do not require Force OCR. Use it only when text can already be selected but search results are clearly incorrect, or when the text layer is misaligned with the page image.
After processing, open the file in a PDF reader and search for a clearly visible word. Then copy a passage containing numbers, dates, and punctuation to check the result. OCR accuracy can be affected by resolution, compression, stains, handwriting, vertical text, complex tables, and overlapping stamps. Always verify critical information manually.
Use cases
See how the tool fits into real work and everyday tasks.
Adding a text layer to scanned paper contracts lets you search by customer name or contract term. Processed books and handouts allow passages to be copied into notes. Invoices, receipts, and shipping documents can be OCR-processed before manual entry or further data extraction. For Japanese manuals, the Japanese + English option can typically recognize katakana, kanji, and model numbers together. For large archives, test a sample from each batch before processing everything to avoid rework caused by an incorrect language setting.
The purpose of an OCR PDF is to preserve the original pages while making them searchable. It does not create a fully editable document. To modify paragraphs, tables, or image placement, convert the PDF to Word afterward. If you only need plain text, use a PDF-to-text tool. Clear scans are generally more reliable than low-resolution photos, so ensure that small text remains legible and that each page has sufficient contrast.
Q&A
Find concise answers to common questions and confusing cases.
Blur, skew, mixed fonts, and background noise can reduce accuracy. Try using the original scan and enabling deskew. Critical fields should still be checked manually.
The tool is designed to preserve the original page image, and the added text layer is usually invisible. For complex pages, compare the output with the original file.
The current process is designed primarily for printed text, so handwriting accuracy is not guaranteed. Cursive writing, faint ink, and corrections can further reduce accuracy.
OCR must render and recognize each page individually. Files with more pages or higher-resolution scans take longer, and an initial cold start may also increase the wait time.
Pages with existing text are skipped by default. Enable Force OCR only after confirming that the current text layer is inaccurate, as duplicate text layers can interfere with copying.
Notes
Review scope, result limitations, and important precautions before use.
OCR output should not be used as the sole basis for financial, medical, legal, or identity-verification decisions. Compare amounts, decimal points, dates, identification numbers, and negative wording against the original image. A text layer can improve search and copying, but it does not automatically create proper heading structure, reading order, or alternative text. OCR alone is therefore not a complete accessibility remediation.
When processing scans containing personal information, restrict access to download links and output files. Keep the original scans so disputed results can be reviewed later. Password-protected files must be lawfully unlocked before processing. If pages are badly damaged, cropped, or overexposed, the software cannot recover characters that are no longer visible.
Related
Discover related tools, collections, and available API capabilities.