Skip to content
FileBelow

When a scanned document is too large to upload

A scan is not a photograph, and the advice that works for photographs will make your text unreadable. The rules are close to the opposite.

Last reviewed

Why scans are so big

Scanners default to high resolutions — 300 or 600 DPI — because they are built for archiving and reproduction. A single A4 page at 600 DPI is around 35 million pixels, which is three times a typical phone photograph, for a picture of some text.

Worse, scanners often default to lossless formats, or to PDFs containing lossless images. Paper is not flat and evenly lit, so a scan is full of faint speckle and paper texture, and lossless compression stores every bit of it faithfully. A ten-megabyte scan of a one-page letter is entirely normal.

The rule is the opposite of the one for photos

Compressing a photograph to a tight limit is best done by giving up resolution and keeping the encoder quality high. Nobody misses pixels in a picture of a landscape, and heavy quality reduction produces visible blocking.

For a scan that logic inverts, because resolution is legibility. Text is made of thin strokes, and once a stroke is thinner than a pixel it stops existing. Reduce a scan far enough and the letters do not soften, they disappear — and a document nobody can read has failed at its only job, however small the file is.

So: keep the pixels, spend the quality. Text tolerates a surprising amount of JPEG compression before it becomes hard to read, because the strokes are still in the right places even when the edges get noisy. It tolerates almost no downscaling at all.

What to do, in order

1. Crop the page

This is the biggest single win and almost nobody does it. A photographed or scanned document usually includes a margin of scanner lid, desk, or the edge of the next page. Every one of those pixels costs bytes and carries no information.

Cropping tight to the page routinely removes a third of the file before anything is compressed at all. The cropper takes about ten seconds.

2. Reduce to what you actually need

600 DPI is an archival setting. For a document somebody will read on screen or print once, 200 to 300 DPI is ample — that is roughly 1700 to 2500 pixels across an A4 page. Going below about 1200 pixels across is where ordinary body text starts to become genuinely hard.

If your scanner is set to 600, rescanning at 300 is better than shrinking afterwards. If rescanning is not an option, resize to a target width rather than letting a compressor decide.

3. Convert to greyscale if it is black text on white

Colour information costs a substantial share of the file and a typed document has none worth keeping. Most scanning software has a greyscale or black-and-white mode, and using it at scan time typically halves the size — sometimes far more for pure black and white.

4. Then compress to the limit

Now set the number the form gave you in the compressor. FileBelow will search for the best-looking file under it, and because you have already removed the waste, it has far less to give up.

Set a minimum width under Advanced options if legibility is critical. That tells the search it may not scale below your floor and must find the size saving in encoder quality instead — which, for text, is exactly the right trade.

Photographs of documents

If you photographed the page with a phone rather than scanning it, everything above applies and two more things matter more than any setting.

Light. Even, indirect light with no shadow across the page. A shadow forces the compressor to spend bytes describing a gradient that carries no information, and it makes the text harder to read at any file size.

Angle. Straight on, from directly above. A page photographed at an angle is a trapezoid with the far edge smaller and blurrier, and no compression setting fixes perspective.

Where the pages have their own instructions

If the scans are photographs, JPG to PDF covers assembling them; if they are screenshots or exports, PNG to PDF covers the same job with a warning about text sharpness that matters more for that kind of file.

If it needs to be a PDF

Many forms want one file rather than several images. Images to PDF assembles them in your chosen order, in your browser, with a quality setting that applies to every page.

One genuine warning: a photographed or scanned page is a picture of text. It cannot be searched, selected or read aloud, and it takes far more space than the words would. If the document exists as a file anywhere — a word processor document, an email, a web page — exporting that directly to PDF produces something smaller, sharper and more useful in every respect. Scanning is the answer when the original is genuinely on paper, and only then.

Need to actually do this?

Set a minimum width under Advanced options so the search spends quality rather than pixels. Free, no sign-up, and the image is processed on your own device.

Compress a scan without losing the text

Or browse every explanation on this site, and the requirements published by named destinations.