PDF size has almost nothing to do with the number of pages. What matters is what kind of content those pages are made of. To show the spread, we built four test documents:
| Document | Pages | Size | Per page |
|---|---|---|---|
| Text report (text, tables, vector charts) | 20 | 96 KB | ~5 KB |
| Scanned document (300 dpi, grayscale) | 6 | 2.9 MB | ~490 KB |
| Slide deck with full-bleed photos | 8 | 4.0 MB | ~500 KB |
| Public 245-page technical report (text + charts) | 245 | 22.6 MB | ~95 KB |
In our test set a scanned page costs about a hundred times more than a typed page. That ratio explains most “why is this PDF 40 MB?” moments.
1. Scans: every page is a photograph
A scanner (or a phone scanning app) doesn’t capture text — it captures a picture of the paper. An A4/Letter page at 300 dpi is roughly 2,500 × 3,300 pixels, about 8 million pixels per page, even if the page holds three lines of text. Paper texture and scanner noise make it worse: noise is detail the image encoder has to keep.
How to tell: try to select a word. If you can’t (and there was no OCR), the page is an image.
What works: lowering resolution. Text stays readable at 120–150 dpi on screen; 300 dpi is only needed for print. This is the case where compression gives the biggest wins — see our measurements.
2. Photos placed at full camera resolution
When a 4000-pixel photo is dropped into a slide that shows it at 13 inches wide, the PDF stores all 4000 pixels. Presentation and word-processing apps often embed the original file untouched. Eight such slides are already several megabytes.
What works: downsampling images to what the page actually needs (about 150 pixels per inch for screen) and re-encoding them as JPEG.
3. Embedded fonts
To look identical everywhere, PDFs embed fonts. Usually only the used characters are embedded (a “subset”), which is small. But some exporters embed full font files — and a full CJK (Chinese/Japanese/Korean) font can be 5–20 MB on its own.
How to tell: in Acrobat or Preview, look at Document Properties → Fonts. Fonts without “(Embedded Subset)” are embedded in full.
What works: re-saving with a tool that subsets fonts (Ghostscript does). Browser-side rasterizing removes fonts entirely — at the cost of selectable text.
4. Leftovers: edit history, thumbnails, metadata
A PDF that was edited and saved many times can keep old versions of objects (“incremental updates”), page thumbnails, form data, and XMP metadata. Each is small, but they add up in documents that passed through several tools.
What works: any full re-write of the file (“Save As”, Ghostscript, or a compressor) drops the dead objects.
5. Vector graphics that are too detailed
Charts and diagrams are usually tiny as vectors. The exception is exports from CAD, GIS or plotting libraries with hundreds of thousands of points — some maps contain millions of path segments, and they are slow to open as well as large.
What works: simplifying the drawing at the source, or rasterizing that one page.
Which compressor for which case
| Your PDF is mostly… | Browser (kompakt web) | Desktop (Ghostscript) |
|---|---|---|
| Scans | Large reduction | Large reduction |
| Photos / slides | Large reduction | Large at low, small at medium |
| Typed text, tables, charts | Usually gets bigger | Small reduction, keeps text |
| Embedded full fonts | Removes them (text becomes image) | Subsets them, keeps text |
Rule of thumb: if you can’t select the text, compress it in the browser. If you can, and you need to keep it selectable, use the desktop app.
Next: what quality you lose at each setting, and how the browser compressor works under the hood.