Resolution, OCR, and File Formats for Scanning
Scanning at the highest available resolution feels like the safe choice and is usually the wrong one. It produces files several times larger, slows every batch, and past a certain point adds nothing an OCR engine or a human reader can use. For ordinary business documents the answer is almost always 300 dpi, and knowing why lets you defend it.
The short version
- 300 dpi is the working default for text documents and the usual baseline for reliable OCR.
- 200 dpi is acceptable for clean, large-print internal documents where storage matters more.
- 600 dpi and above is for photographs, fine detail, and very small print — not general paperwork.
- Searchable PDF keeps the page image and adds a hidden text layer. It is what most offices should produce.
What resolution actually buys you
Resolution is dots per inch — how finely the scanner samples the page. Doubling it roughly quadruples the data, because the increase applies in both directions.
Colour mode changes file size more than resolution does
| Mode | Use for |
|---|---|
| Black and white (bitonal) | Clean printed text. Dramatically smaller files and often the best OCR input, because the engine works on shapes rather than shades. Poor for anything with photographs, highlighter, or faint pencil. |
| Greyscale | Documents with faded text, stamps, pencil annotations, or photographs where colour is not meaningful. A good compromise and often the safest default for mixed archival material. |
| Colour | When colour carries meaning — signatures in blue ink to prove an original, colour-coded forms, marketing material, photographs. Largest files by a wide margin. |
| Auto colour detection | The scanner picks per page. Excellent for mixed batches, and worth enabling by default on most business intake. |
For a clean printed page, black and white at 300 dpi produces a small file and excellent OCR. Colour at 600 dpi produces a file many times larger with no readability gain and, on some documents, slightly worse OCR because the engine has more noise to interpret. Match the mode to the content, not to caution.
What OCR does, and where it fails
Optical character recognition converts the shapes of characters into machine-readable text. Modern engines are very good on clean printed material and considerably less good on everything else.
| OCR handles well | OCR struggles with |
|---|---|
| Clean printed text at 300 dpi | Handwriting, especially cursive |
| Standard business fonts | Decorative or very small fonts |
| Straight, well-lit pages | Skewed, creased, or shadowed pages |
| Good contrast, black on white | Faded thermal receipts, carbon copies |
| Simple single or multi-column layout | Complex tables and forms |
| One known language | Mixed languages on one page |
Even excellent OCR on clean documents makes occasional errors, and a small per-character error rate becomes a meaningful per-page one. That is fine for making an archive searchable. It is not fine for automatically extracting a figure into an accounting system without a human check. Decide which of those you are doing before you rely on the output — and never discard the page image, because the image is the record and the text layer is only an index into it.
Choosing a file format
| Format | When to use it |
|---|---|
| Searchable PDF | The default for most business scanning. Keeps the page image exactly as scanned and adds an invisible text layer, so the document looks right and is findable. Multi-page, universally readable. |
| Image-only PDF | When OCR is not wanted or would not work — handwritten notes, drawings. Faithful but not searchable. |
| PDF/A | The archival variant. Embeds fonts and forbids features that may not render in future software. Use for anything with a long retention obligation, and expect somewhat larger files. |
| TIFF | Long-established in records management and imaging systems. Excellent for bitonal archival, widely supported by document management software, but not directly searchable. |
| JPEG | Photographs. Avoid for text — the compression blurs character edges and degrades OCR. |
The distinguishing feature of PDF/A is that it is self-contained — nothing it needs to render lives outside the file. For records with a multi-year retention requirement that is worth the extra size, because the failure mode you are guarding against is opening the file in a decade and finding it renders differently.
Sensible defaults
| Document type | Suggested settings |
|---|---|
| General office paperwork | 300 dpi, auto colour detection, searchable PDF |
| Invoices and receipts | 300 dpi greyscale — thermal receipts are often faint, and greyscale preserves what bitonal drops |
| Contracts and signed originals | 300 dpi colour, PDF/A — colour evidences a wet signature |
| Archival records with retention rules | 300 dpi, greyscale or colour, PDF/A |
| Engineering drawings and fine print | 600 dpi, greyscale |
| Photographs | 600 dpi colour, JPEG or TIFF |
| Bulk internal documents, storage constrained | 200–300 dpi bitonal, searchable PDF |
Getting better OCR results
- Scan straight. Skew is the most common cause of poor recognition. Enable auto-deskew and square the stack in the feeder.
- Use 300 dpi. Below that, character edges degrade and accuracy falls off faster than the file size saving justifies.
- Set the right language in the OCR settings. An engine expecting the wrong language guesses badly.
- Clean the glass and feed strip. A speck produces a line down every page, and OCR reads that line as characters — see roller and glass maintenance.
- Prefer bitonal for clean printed text, greyscale for anything faded or annotated.
- Do not over-compress. Aggressive compression softens character edges and quietly costs accuracy.
- Spot-check a sample before committing a large batch. Search the output for a word you know is on the page.
Standards and manufacturer resources
OCR engines and supported formats vary by the software bundled with each scanner. Check the documentation for your specific model and capture application.
Not sure what settings your documents need?
Tell us what you are digitizing and whether it has to be searchable or archival. We will suggest resolution, colour mode, and format — and flag anything where OCR will not give you what you are hoping for.
Still stuck? Talk to someone who works on this hardware daily.
Tell us the model, what you have already tried, and what you are seeing. We will tell you whether it is a setting, a consumable, or a part — and we will say so if you do not need to buy anything.