Explainer

Resolution, OCR, and File Formats for Scanning

Updated 5 min read

Scanning at the highest available resolution feels like the safe choice and is usually the wrong one. It produces files several times larger, slows every batch, and past a certain point adds nothing an OCR engine or a human reader can use. For ordinary business documents the answer is almost always 300 dpi, and knowing why lets you defend it.

The short version

  • 300 dpi is the working default for text documents and the usual baseline for reliable OCR.
  • 200 dpi is acceptable for clean, large-print internal documents where storage matters more.
  • 600 dpi and above is for photographs, fine detail, and very small print — not general paperwork.
  • Searchable PDF keeps the page image and adds a hidden text layer. It is what most offices should produce.

What resolution actually buys you

Resolution is dots per inch — how finely the scanner samples the page. Doubling it roughly quadruples the data, because the increase applies in both directions.

Scan resolution against file size and usefulness 200 dpi suits clean large text at small file size. 300 dpi is the reliable default for text and OCR. 600 dpi quadruples file size relative to 300 and benefits only fine detail and photographs. Resolution, file size, and what it is good for 200 dpi Clean, large text Smallest files OCR less reliable 300 dpi The default Reliable OCR ~2× the data of 200 600 dpi Fine detail, photos Very small print ~4× the data of 300 Above 600 Rarely justified for ordinary paperwork Slows every batch Going higher than the document needs costs storage, transfer time, and scanning throughput — with no gain in readability. Scan at the resolution the smallest text on the page requires, not at the maximum the scanner offers.
Relative sizes are approximate and depend on colour mode and compression. The ranking holds.

Colour mode changes file size more than resolution does

Mode Use for
Black and white (bitonal) Clean printed text. Dramatically smaller files and often the best OCR input, because the engine works on shapes rather than shades. Poor for anything with photographs, highlighter, or faint pencil.
Greyscale Documents with faded text, stamps, pencil annotations, or photographs where colour is not meaningful. A good compromise and often the safest default for mixed archival material.
Colour When colour carries meaning — signatures in blue ink to prove an original, colour-coded forms, marketing material, photographs. Largest files by a wide margin.
Auto colour detection The scanner picks per page. Excellent for mixed batches, and worth enabling by default on most business intake.
Bitonal at 300 dpi beats colour at 600 for plain text

For a clean printed page, black and white at 300 dpi produces a small file and excellent OCR. Colour at 600 dpi produces a file many times larger with no readability gain and, on some documents, slightly worse OCR because the engine has more noise to interpret. Match the mode to the content, not to caution.

What OCR does, and where it fails

Optical character recognition converts the shapes of characters into machine-readable text. Modern engines are very good on clean printed material and considerably less good on everything else.

OCR handles well OCR struggles with
Clean printed text at 300 dpi Handwriting, especially cursive
Standard business fonts Decorative or very small fonts
Straight, well-lit pages Skewed, creased, or shadowed pages
Good contrast, black on white Faded thermal receipts, carbon copies
Simple single or multi-column layout Complex tables and forms
One known language Mixed languages on one page
OCR accuracy is never 100%, and you should plan for that

Even excellent OCR on clean documents makes occasional errors, and a small per-character error rate becomes a meaningful per-page one. That is fine for making an archive searchable. It is not fine for automatically extracting a figure into an accounting system without a human check. Decide which of those you are doing before you rely on the output — and never discard the page image, because the image is the record and the text layer is only an index into it.

Choosing a file format

Format When to use it
Searchable PDF The default for most business scanning. Keeps the page image exactly as scanned and adds an invisible text layer, so the document looks right and is findable. Multi-page, universally readable.
Image-only PDF When OCR is not wanted or would not work — handwritten notes, drawings. Faithful but not searchable.
PDF/A The archival variant. Embeds fonts and forbids features that may not render in future software. Use for anything with a long retention obligation, and expect somewhat larger files.
TIFF Long-established in records management and imaging systems. Excellent for bitonal archival, widely supported by document management software, but not directly searchable.
JPEG Photographs. Avoid for text — the compression blurs character edges and degrades OCR.
Pick PDF/A for anything you must keep

The distinguishing feature of PDF/A is that it is self-contained — nothing it needs to render lives outside the file. For records with a multi-year retention requirement that is worth the extra size, because the failure mode you are guarding against is opening the file in a decade and finding it renders differently.

Sensible defaults

Starting points. Test against your own documents before committing a large batch.
Document type Suggested settings
General office paperwork 300 dpi, auto colour detection, searchable PDF
Invoices and receipts 300 dpi greyscale — thermal receipts are often faint, and greyscale preserves what bitonal drops
Contracts and signed originals 300 dpi colour, PDF/A — colour evidences a wet signature
Archival records with retention rules 300 dpi, greyscale or colour, PDF/A
Engineering drawings and fine print 600 dpi, greyscale
Photographs 600 dpi colour, JPEG or TIFF
Bulk internal documents, storage constrained 200–300 dpi bitonal, searchable PDF

Getting better OCR results

  1. Scan straight. Skew is the most common cause of poor recognition. Enable auto-deskew and square the stack in the feeder.
  2. Use 300 dpi. Below that, character edges degrade and accuracy falls off faster than the file size saving justifies.
  3. Set the right language in the OCR settings. An engine expecting the wrong language guesses badly.
  4. Clean the glass and feed strip. A speck produces a line down every page, and OCR reads that line as characters — see roller and glass maintenance.
  5. Prefer bitonal for clean printed text, greyscale for anything faded or annotated.
  6. Do not over-compress. Aggressive compression softens character edges and quietly costs accuracy.
  7. Spot-check a sample before committing a large batch. Search the output for a word you know is on the page.

Standards and manufacturer resources

OCR engines and supported formats vary by the software bundled with each scanner. Check the documentation for your specific model and capture application.

Not sure what settings your documents need?

Tell us what you are digitizing and whether it has to be searchable or archival. We will suggest resolution, colour mode, and format — and flag anything where OCR will not give you what you are hoping for.

Still stuck? Talk to someone who works on this hardware daily.

Tell us the model, what you have already tried, and what you are seeing. We will tell you whether it is a setting, a consumable, or a part — and we will say so if you do not need to buy anything.