Pdf91 — Working with PDF files outside paid software | pdf91.com

Shrink scans, keep text

Pdf91 — Working with PDF files outside paid software | pdf91.com

A 20 MB scanned PDF is almost always a stack of full-page images, not text. Re-encode those images to JPEG at quality 75–85 while keeping 300 dpi, and each 2–5 MB page drops to 200–600 KB — sharp enough to print and to OCR.

2008PDF 1.7 published as an open standard, ISO 32000-1:2008.2011PDF/A-2 — including the A-2b archival profile — published as ISO 19005-2:2011.2017PDF 2.0 published as ISO 32000-2:2017, the first full revision of the standard.
One A4 page scanned at 300 dpi is 2480 × 3508 pixels — about 8.7 megapixels — before any compression touches it.

01

Compress a PDF to a target size

The full arithmetic behind file size: pixel counts, codecs and per-page budgets. Read it when a form gives you a hard limit and you need to hit it deliberately.

02

Page sizes and the DPI a print shop wants

A4 and US Letter in points, millimetres, inches and pixels, tied to the 150/300/600 dpi tiers. Read it before sending artwork to print.

03

PDF versions, PDF/A and what archives reject

PDF 1.7 and 2.0, the A-1 and A-2b profiles, and the exact features that fail validation. Read it before submitting to an archive.

Why a 20 MB scan is an image problem, not a text problem

A scanned PDF is a stack of full-page photographs, and those photographs are essentially the whole file.

A full A4 page at 300 dpi is 2480 × 3508 pixels, roughly 8.7 megapixels. Stored with lossless Flate compression — the default in many scanners — one photographic page lands at 2–5 MB, so ten pages make a 20–50 MB file before anything else is added.

Blur after compression is almost never caused by the JPEG quality number. It is caused by throwing away pixels: downsample a page below about 200 dpi and small letterforms lose their edges, and OCR accuracy drops sharply into confusable fragments.

The fix is to change the codec, not the resolution. Re-encoding the same pixels to DCT (JPEG) at quality 75–85 typically takes a 2–5 MB page down to 200–600 KB with no visible change at normal viewing distance.

Six jobs, one toolbox

Compress

Re-encode, don't downsample

Keep 300 dpi and switch page images from Flate to JPEG quality 75–85. A 2–5 MB scanned page drops to 200–600 KB with no visible change.

Split

Divide at 10 MB, not at the quality dial

Mail gateways and upload forms stop around 20–25 MB. Parts of 10 MB or less pass without touching a single pixel.

Merge

Container edits never blur text

Merging, reordering and extracting pages shuffles whole pages without re-encoding images, so quality is untouched — and so is per-page size.

Convert

PDF to image and back

A full A4 page at print resolution is 2480 × 3508 pixels. Convert at that size and the round trip stays sharp; convert smaller and the pixels are gone for good.

Print

150 / 300 / 600: the three numbers

150 dpi proofs on screen, 300 dpi prints photos, 600 dpi carries small line art and halftone text. Full-bleed A4 artwork is 216 mm wide before trimming.

Archive

A-2b, not A-1b

PDF/A-1 forbids attachments along with JavaScript and encryption; PDF/A-2 permits attachments, which is why archives ask for A-2b.

Blurry text after compression is almost always lost pixels, not a low JPEG quality number — keep 300 dpi and turn the quality dial, not the resolution dial.
Re-encoding a scanned page from Flate to JPEG at quality 75–85 takes it from 2–5 MB down to 200–600 KB.

Settings and what they do to one A4 page

What each setting does to a single scanned A4 page
SettingValue to useWhat one page becomesWhen to choose it
Kept resolution300 dpi2480 × 3508 px, about 8.7 megapixelsAnything that will be printed or OCR'd
Screen-proof resolution150 dpiReadable on screen onlyA quick check copy
Line-art resolution600 dpiSmall line art and halftone text stay intactDrawings and fine print
Image codec: JPEG (DCT)Quality 75–85200–600 KB per pageThe email-friendly copy
Image codec: FlateLossless2–5 MB per pageThe master file you keep
Downsampled scanBelow ~200 dpiOCR accuracy drops sharplyAvoid for anything you must read
Split parts10 MB or less eachPasses 20–25 MB gateways and formsWhen the whole file is too big
What each setting does to a single scanned A4 page

What people ask after the first attempt

I compressed my PDF and now the text looks fuzzy. What went wrong?
The tool almost certainly downsampled the pages instead of re-encoding them. Blur comes from lost pixels: keep 300 dpi — never below about 200 dpi — and lower the JPEG quality instead, since quality 75–85 keeps letterforms crisp.
The file is still over the limit after compression. Now what?
Split it. Gateways and forms typically stop at 20–25 MB, and the standard workaround is parts of 10 MB or less. Splitting shuffles whole pages, so nothing gets blurrier.
Will zipping the PDF make it smaller?
Effectively no. The images inside a PDF are already compressed, so a zip wrapper saves almost nothing. Re-encode the images to JPEG quality 75–85 instead — that is where the two-to-five-fold difference lives.
An archive rejected my PDF. What fails validation?
JavaScript, encryption and — under PDF/A-1 — embedded attachments. PDF/A-2 keeps the JavaScript ban but allows attachments, which is why archives specify A-2b; remove the banned features and re-validate.

Sources and standards

The standards documents named above are the sources behind every rule quoted on this page.

Why a 20 MB scan isan image problemThe procedure: sixsteps from 20 MBPicking the qualitysetting withoutWhy text goesblurry — the namedWhen compression isnot enough: splitIf the file isgoing to a printer
A map of this guide's sections