Pdf91 — Working with PDF files outside paid software | pdf91.com
File size is pixel count multiplied by codec choice; control those two numbers and any target becomes arithmetic.
Start with where the bytes live. In a scanned PDF each page is a full-page image: A4 at 300 dpi is 2480 × 3508 pixels, roughly 8.7 megapixels. Stored with lossless Flate compression, one photographic page weighs 2–5 MB, so a ten-page document is 20–50 MB before anything else is added.
The single most effective change is the codec. Re-encoding the same pixels as JPEG (DCT) at quality 75–85 takes that page down to 200–600 KB — a two-to-five-fold reduction with no visible difference at normal viewing distance. The resolution, and with it the sharpness of the text, does not change.
On narrow screens, swipe or scroll the plate sideways.
To hit a target, work out a per-page budget: divide the target size by the page count. Two megabytes over ten pages is 200 KB per page, which points at the quality-75 end of the range; over five pages it is 400 KB per page, comfortably inside quality 85. If the budget falls below 200 KB per page, compression alone will not get you there cleanly.
Three moves waste the effort. Downsampling below about 200 dpi destroys the letterforms you were trying to keep and sends OCR accuracy off a cliff. Re-saving an already-JPEG page stacks artefacts on artefacts. And wrapping the PDF in a zip file changes almost nothing, because the images inside are already compressed.
When the arithmetic says the target is unreachable — a hundred-page scan that must fit in 2 MB — switch strategy from compress to split. Parts of 10 MB or less pass the 20–25 MB ceilings of mail gateways and upload forms, and every page inside them stays at full resolution.
Further reading