How to OCR a PDF to Word in 3 Steps
- Upload the scanned PDF in the converter above (up to 100 MB).
- Convert. The converter first looks for a text layer. If the PDF has none, it renders the first 10 pages at 300 DPI and reads them with Tesseract OCR.
- Download the DOCX. OCR output is plain paragraphs: the text is editable, but the original layout, tables and images are not rebuilt.
The converter on this page reads English. The PDF to Word converter page (cleverutils.com/pdf-to-docx) has a language menu with 14 OCR languages: English, French, German, Spanish, Italian, Portuguese, Russian, Ukrainian, Polish, Dutch, Turkish, Arabic, Japanese and Korean.
What Is OCR?
Optical Character Recognition (OCR) is a technology that converts images of text into machine-readable, editable text. When you scan a paper document, the scanner creates a photograph of each page. OCR software analyzes that photograph, identifies individual characters, and outputs the corresponding text.
The OCR process typically involves several steps:
- Image preprocessing: Straightening skewed pages, removing noise, adjusting contrast, and binarizing the image (converting to black and white)
- Text detection: Identifying regions of the image that contain text vs. images, borders, or blank space
- Character recognition: Analyzing individual character shapes and matching them against known letter patterns
- Post-processing: Applying dictionary matching and language rules to correct common recognition errors
Scanned vs Native PDFs
Understanding the difference between scanned and native PDFs is crucial for choosing the right conversion approach:
| Feature | Native (Digital) PDF | Scanned PDF |
|---|---|---|
| Created by | Export from Word, browser print, etc. | Scanner, camera, fax machine |
| Content | Structured text data | Images of pages |
| Text selectable? | Yes | No |
| Searchable? | Yes | No (without OCR) |
| OCR needed? | No — text extracted directly | Yes — required for text extraction |
| Conversion accuracy | Exact (text is copied, not recognized) | Depends on scan quality |
Quick test: Open the PDF and try to select text with your mouse. If you can highlight individual words, it is a native PDF. If clicking selects the entire page as a single image, it is a scanned PDF that needs OCR. Native files skip OCR, so read how to convert a native PDF to Word without losing formatting instead.
Factors That Affect OCR Accuracy
OCR accuracy varies dramatically based on input quality. Here are the key factors:
Scan Resolution (DPI)
Resolution is the single most important factor. Higher DPI means more pixel information for the OCR engine to work with:
- 150 DPI: Minimum for OCR. Works for large, clear fonts; small text produces many errors.
- 300 DPI: Recommended standard. Good balance of file size and accuracy on clean text. Our OCR renders pages at 300 DPI.
- 600 DPI: Best for very small text and dense documents. Larger files, slower processing.
Image Quality
Beyond resolution, several image quality factors affect OCR results:
- Contrast: High contrast between text and background produces best results. Faded text on aged paper is harder to recognize.
- Alignment: Straight, properly aligned pages produce better results than skewed or rotated scans. Most OCR engines include deskewing, but starting straight is better.
- Noise: Speckles, smudges, coffee stains, and scanner artifacts reduce accuracy. Clean originals scan better.
- Shadows: Book spines create shadows in the gutter margin. Flatbed scanning or using a document camera reduces this issue.
Font and Text Characteristics
Not all text is created equal for OCR purposes:
- Standard fonts (Times New Roman, Arial, Helvetica) — highest accuracy
- Decorative fonts (script, ornamental) — lower accuracy
- Small text (below 8pt) — needs higher DPI to compensate
- Bold text — generally good; very heavy weights may merge characters
- Colored text on colored backgrounds — reduced contrast lowers accuracy
Improving OCR Results
If your initial OCR results are unsatisfactory, try these preprocessing steps before conversion:
- Rescan at higher DPI: If you have access to the original document, rescan at 300 or 600 DPI.
- Straighten skewed pages: Use your scanner's auto-deskew feature or straighten images before OCR.
- Increase contrast: If the original is faded, adjust the scanner's brightness and contrast settings to darken the text and lighten the background.
- Remove noise: Use despeckle filters to clean up scanner artifacts and paper texture.
- Crop margins: Removing large blank margins, binding holes, and edge artifacts helps the OCR engine focus on the actual content.
Best practice: Scan documents in color at 300+ DPI even if the original is black and white. Color scans preserve more information for the preprocessing stage, even though OCR ultimately works on the binarized image.
Multi-Language OCR
Modern OCR engines support dozens of languages, including those with non-Latin scripts (Chinese, Japanese, Korean, Arabic, Cyrillic, Devanagari). Key considerations for multi-language documents:
- Language selection: Specifying the correct language improves accuracy, because the OCR engine uses language-specific dictionaries and character sets.
- Mixed-language documents: Documents containing multiple languages (common in academic papers) may need multiple OCR passes or a multi-language configuration.
- Right-to-left scripts: Arabic and Hebrew require OCR engines with proper bidirectional text support.
- CJK characters: Chinese, Japanese, and Korean have thousands of characters with subtle differences, requiring specialized recognition models.
Handwriting Recognition Limitations
While OCR technology has advanced significantly, handwriting recognition remains challenging:
- Printed-style handwriting: Neat, separated block letters are partly recognized, with frequent errors.
- Cursive handwriting: Connected letters are extremely difficult for standard OCR engines and usually come out unreadable.
- Individual variation: Unlike machine-printed text, each person's handwriting is unique, making pattern matching unreliable.
- Mixed content: Documents with both printed text and handwritten annotations are best processed in two steps — OCR the printed text, then manually transcribe the handwriting.
Other Free Ways to OCR a PDF
- Google Docs: upload the PDF to Google Drive, right-click it and choose Open with → Google Docs. Drive runs OCR and opens the text, which you can download as .docx. If you only need plain text rather than a DOCX, see how to extract text from a PDF.
- Windows 11: the Snipping Tool’s Text actions copies text from a screenshot of a page, and Microsoft PowerToys includes Text Extractor (Win+Shift+T). For a single page saved as a JPG or PNG, our image to text OCR tool reads the text online.
- Mac: Live Text in Preview lets you select and copy text in scanned pages and images (macOS Monterey and later).
- Tesseract: the same open-source engine our converter uses runs on your computer:
tesseract page.png output -l eng.