Skip to content

Compress a scanned PDF

Scanned PDFs — especially from copy-shop scanners or phone apps — are notorious for size bloat. A 10-page scan can easily exceed 50 MB because each page is a full-resolution image.

Tool

⚡ Open the tool

Free · No account · Files deleted in 1 hour

Why this works

Our scan-optimised compression detects image-only pages, applies smart downsampling and re-encodes with JBIG2 / JPEG2000 where supported. Typical reductions are 70–95% with text still clearly readable.

Scanned PDFs are size hogs because of how they\'re built. A born-digital PDF stores text as text — the actual character codes, taking roughly 1–2 KB per page. A scanned PDF stores each page as a high-resolution photograph of paper, taking 3–6 MB per page in colour at typical scanner settings. A 20-page scanned contract that exists as text would weigh 50 KB; as a scan it weighs 80–120 MB. Same content, 2,000× the file size.

Default scanner settings are the root cause. Office multifunction scanners typically default to 300 DPI in colour, which is appropriate for photographs but enormously overkill for documents that contain only text. Phone scanner apps (Adobe Scan, CamScanner, Microsoft Lens) default to similar settings or even higher to maximise OCR accuracy. Copy-shop self-service scanners are often set to maximum quality by default because operators can\'t predict whether customers want archival photos or text documents.

What the scan-optimised compression actually does, in order: detect image-only pages (no embedded text layer) versus mixed pages (some text, some image) and apply different treatment to each; downsample images from typical scanner output (300 DPI) to 150 DPI — imperceptible on screens, modest impact when printed; convert colour scans to grayscale when colour adds no information (most office documents); apply JBIG2 compression for monochrome pages — a specialised compression algorithm designed for scanned text that achieves 50–80% better ratios than generic JPEG; apply JPEG2000 compression for greyscale and colour pages where supported; strip duplicate page resources (e.g. recurring scanner-bed shadow patterns across pages).

A practical sequence for getting scanned PDFs to manageable sizes. First, run OCR if you haven\'t already — OCR adds an invisible text layer but doesn\'t bloat the file much (typically <100 KB total per page), and once you have OCR\'d text, downstream compression can be more aggressive on the image layer without losing searchability. Then run scan-optimised compression with the Recommended setting — typical reduction is 60–80% with all text readable. If you need to hit a specific size cap (1 MB for visa portals, 10 MB for email), escalate to Aggressive or also convert to grayscale.

Where compression hits limits. Photographs of architectural drawings, blueprints, or detailed diagrams can\'t compress as aggressively because every pixel encodes information — expect 30–50% reduction rather than 80%. Documents with red-ink notations, highlighter marks, or coloured stamps lose information when converted to grayscale; keep colour mode for these. Hand-annotated documents with complex sketches may compress poorly if the annotations are essential to the document\'s meaning.

For very long scanned documents (50+ pages), consider splitting first. A 100-page scan compressed and sent as one file is often unwieldy regardless of the byte count; splitting into logical sections (chapters, sections, monthly statements) makes the recipient\'s life easier and lets you re-send only the affected section if a correction is needed later.

How it works

  1. 1
    Open the compress tool
    Launch the tool with the "Scanned document" preset auto-selected. The preset enables JBIG2 and JPEG2000 codecs and tunes downsampling for scanned content.
  2. 2
    Upload the scan
    Drop the heavy scanned PDF in. We handle 25 MB on the free tier, 1 GB on Pro. Even multi-hundred-megabyte raw scans are accepted.
  3. 3
    Choose colour or grayscale
    Grayscale shrinks file size by roughly 65% on typical scans — use it whenever colour adds no information. Keep Colour mode if signatures, red-ink notes, or coloured stamps are essential.
  4. 4
    Pick a compression level
    Recommended is right for almost everything. Aggressive pushes further for portal-cap compliance. Lossless preserves print quality but yields smaller reductions.
  5. 5
    Run OCR after compression if needed
    For searchable output, run OCR on the compressed result. The text layer adds little size while making the document Cmd-F findable and screen-reader accessible.
Who this is for

Real-world uses

Legal teams

Compress evidence bundles and discovery production for e-filing systems with size caps (often 10 MB per filing).

Healthcare admins

Shrink scanned medical records to fit through HIPAA-compliant transfer systems with bandwidth limits.

Archivists

Convert bulky historical document scans into manageable archive-friendly PDFs without sacrificing legibility.

Bookkeepers

Receipt batches from a phone-scanner app routinely arrive at 100+ MB — compress before filing into the accounting system.

Lawyers and paralegals

Contract scans for client review need to fit into email or document-management upload caps.

Real-estate agents

Disclosure-packet scans need to attach to MLS listings and offer documents without overwhelming recipients' inboxes.

FAQ

Common questions

Will small handwriting still be legible?

Yes at Recommended settings. Default downsampling preserves text legibility down to 8-point handwriting in most cases. Aggressive compression pushes lower and may soften the smallest annotations — verify before sharing.

Should I compress before or after OCR?

OCR first, then compress. Compressing first can introduce artefacts that lower OCR accuracy. After OCR, the text layer is in the file as data and survives compression even at Aggressive settings.

What is JBIG2 and why does it matter?

JBIG2 is a compression standard specifically designed for scanned text and line art. It achieves 50–80% better compression ratios than generic JPEG on monochrome scanned pages. Most modern PDF readers (Acrobat, Preview, Chrome, Firefox) support it natively.

Will my colour scan look different in grayscale?

Yes — colours map to brightness equivalents (bright yellow becomes light gray, deep red becomes dark gray). Text remains crisp. Only use grayscale if colour adds no informational value to the document; keep colour for documents with red-ink corrections, highlights, or coloured stamps.

How do I deal with scanner-bed shadows?

Use the Crop tool first to trim the shadow bands, then compress. Cropped pages compress more efficiently because there's less pixel data to encode.

Does this work on multi-page colour scans from a phone app?

Yes — phone-app scans (Adobe Scan, CamScanner, Microsoft Lens, Apple Notes scan) all produce standard PDFs that work with this tool. Typical 10-page phone scans drop from 50–100 MB to 5–15 MB at Recommended settings.