September 29, 2026 · 7 min read
Shrink a Scanned Document Without Making It Unreadable
Government and bank portals often cap uploads at a few megabytes. How to get a scanned contract or ID under the limit while keeping the text legible.
A four-page scanned contract comes out at 18 MB. The portal accepts 3 MB. You compress it hard, it goes through, and three weeks later you get a letter saying the document was illegible and you need to submit it again. That round trip is the real cost, so it's worth understanding what compression actually does to a scan before you pick a setting.
Why scans are so much bigger than documents
A PDF exported from Word is mostly vector text — instructions for drawing letters, which cost almost nothing. A scan is a photograph of a page. Every letter is pixels, and at 600 DPI in colour that is an enormous amount of data for what is, visually, black marks on white paper.
- Scanning at 600 DPI when 200–300 is the legible standard for text.
- Colour scanning a black-and-white page, which triples the data for nothing.
- Phone scanner apps saving at full camera resolution.
- One PDF per page, then merged, each carrying its own overhead.
Fix it at the scanner before you compress anything
Rescanning takes two minutes and beats any compression, because you avoid throwing away detail twice. If you still have access to the paper, do this first.
| Document | Scan at | Mode |
|---|---|---|
| Plain text contract, form, letter | 200–300 DPI | Greyscale |
| ID card, passport page, driving licence | 300 DPI | Colour — it is usually required |
| Anything with a stamp, seal or signature in colour | 300 DPI | Colour |
| Page with photographs or diagrams | 300 DPI | Colour |
| Handwritten notes, faint carbon copies | 300–400 DPI | Greyscale, raise contrast |
- Phone scanning works well if you use a scanner app rather than the camera: Adobe Scan, Microsoft Lens and the built-in scanner in Apple Notes and Google Drive all flatten perspective and clean the background.
- Shoot in even daylight, no flash, page flat, whole page in frame with a small margin.
- Do not crop into the edges of an ID document. Many authorities reject a scan with a corner or the border missing.
The trade-off nobody explains
PDF compression comes in two fundamentally different flavours, and knowing which one you're using is the whole game.
- Lossless re-saving cleans up how the file is structured and re-packs it. Nothing visible changes and any real text stays selectable. The saving is modest, sometimes only a few percent — but it costs you nothing.
- Rasterising or re-encoding renders each page to an image at a chosen resolution. This is where the big reductions come from, and it is also where text gets soft and searchable text stops being searchable, because there is no longer any text in the file — only a picture of it.
In Skrubly's compressor that is the difference between the levels, and we'd rather say it plainly than let you find out after submitting: Light re-saves the PDF losslessly and keeps the text layer intact. Balanced and Strong rasterise each page — around 120 DPI and 96 DPI respectively — which shrinks a bulky scan dramatically but produces an image-only PDF. For a scan that was already an image, that's often a fair trade. For a PDF with real text in it, it is a downgrade you can't undo.
Try Light first, then step up only if you're still over the limit.
Open the File CompressorA working order of attack
- Rescan at 300 DPI greyscale if you can. Very often you're already under the limit.
- Try lossless first. Free and harmless.
- Still over? Compress in one moderate step and then open the result and read the smallest text on the page — a footnote, a reference number, a date of birth.
- Only if that passes, go harder. Check again after each step, not at the end.
- If nothing works, split the submission: portals that cap at 3 MB per file usually allow several files.
What "unreadable" looks like in practice
Judge the result on the details that matter to whoever reads it, not on the page as a whole at 25 percent zoom. Zoom to 100 percent and check: small print and footnotes, digits in reference numbers and dates, the difference between 8 and 3, accented characters, signatures, and the machine-readable lines on an ID document. Those lines are the first thing to smear, and they're often the first thing the other side checks.
If the text has to stay searchable
Some portals ask for a searchable PDF, or you may simply want to find a clause later. That means OCR: a pass that recognises the letters in the scan and stores them as a text layer behind the image. Most scanner apps do it automatically, Adobe Acrobat does it, and free options exist on the desktop. Run OCR before heavy compression — recognition on a mushy image is much worse — and if you rasterise afterwards, expect to lose the text layer and have to do it again.
| Tool | Good for | Note |
|---|---|---|
| Skrubly compressor | Quick size reduction in the browser | Nothing leaves your device |
| Adobe Acrobat | OCR and fine control | Paid for most features |
| Ghostscript | Batch work on a desktop | Command line, very effective |
| Scanner app OCR | Getting a searchable scan in one step | Quality depends on the original |
Privacy, because this is exactly the sensitive category
Contracts, bank statements and ID documents are the files you should be most careful about handing to a random website. A browser-based tool that processes locally never sends the document anywhere; a server-based one does, and you are trusting its retention policy. Whichever you use, check the file's own properties before submitting: PDFs carry an author name, a title and the software that produced them, and scans of ID documents sometimes carry the scanner model or the phone's details. Clearing those fields takes seconds.
Clear author, title and software fields from a PDF before you submit it.
Document metadata guide