Most “compress PDF online” sites work the same way: you upload the file, a server runs Ghostscript or a commercial SDK on it, and you download the result. That is a reasonable design — and also a non-starter for a contract, a medical report or a bank statement you would rather not hand to a stranger’s server.
kompakt’s web compressor does the whole job in the browser tab. The file is read with the File API and never leaves the device; after the first visit the page works offline. This article explains how, with the trade-offs stated plainly.
The engine is open source: github.com/xronocode/kompakt-web — MIT, with tests and the benchmark used for the numbers below.
The pipeline
PDF file ──► pdf.js parses and renders page N ──► <canvas> (pixels)
│
canvas.toBlob('image/jpeg', q)
│
new PDF ◄── streaming writer ◄── JPEG bytes
- Parse and render. pdf.js — the PDF engine inside Firefox — opens the file and draws one page at a time onto an HTML canvas at the chosen resolution.
- Encode. The browser’s built-in JPEG encoder turns the canvas into a JPEG at the chosen quality.
- Write. A small hand-written PDF writer appends the JPEG as an image object, adds a page that draws it, and moves on. At the end it writes the cross-reference table and trailer.
Then the canvas is released and the next page starts. Only one page is ever held as pixels: memory use is the input file, one rendered page, and the growing output — not every page at once.
Why a custom writer instead of a PDF library
The first version used pdf-lib to assemble the output. It works, but it builds the whole document as a tree of objects in memory and serializes it at the end — fine for 20 pages, painful for 500 pages of 300 dpi images on a phone.
The output we need is extremely simple: every page is one image. So the writer only has to emit four kinds of objects — catalog, page tree, page, image — and remember the byte offset of each for the cross-reference table. It writes bytes as it goes and hands the browser a list of chunks (new Blob(chunks)), so nothing is ever concatenated into one giant buffer. It is about 70 lines and has no dependencies.
Resolution and quality: what the settings mean
| Setting | Render resolution | JPEG quality | Letter page in pixels |
|---|---|---|---|
| low | 72 dpi | 0.50 | 612 × 792 |
| medium | 120 dpi | 0.65 | 1020 × 1320 |
| high | 200 dpi | 0.75 | 1700 × 2200 |
Three bugs we found by measuring
Writing this article, we ran the compressor against a fixed test set and inspected the output files with pdfinfo, pdfimages and Ghostscript instead of eyeballing them. That turned up three real bugs, all fixed on 10 October 2026:
- The dpi labels were wrong. The render scale was computed as
dpi / 96— the CSS pixel density — but a pdf.js viewport at scale 1 is 72 pixels per inch (one pixel per PDF point). So “150 dpi” really rendered at 112 dpi, “300” at 225 and “72” at 54. The scale is nowdpi / 72, and the presets were re-tuned to 72 / 120 / 200 dpi, which keep the old file sizes at medium while making low actually legible and high actually sharp. - The page size changed. The writer used the image size in pixels as the page size in points, so a Letter page compressed at medium came out as 13.3 × 17.2 inches (and shrank to 6.4 × 8.3 inches at low). On screen nobody notices; in print or when inserting the page into another document, everyone does. Pages now keep their original size and the image is scaled onto them.
- The output was technically malformed. A missing newline produced
endstreamimmediately followed byendobj. Viewers silently repair this, but strict validators, print workflows and document-management systems may reject the file. Every output now passespdfinfoand Ghostscript without errors.
Measured results
Full numbers are in the quality comparison; the short version, for the browser compressor at medium:
| Test file | Before | After (medium) | Time |
|---|---|---|---|
| Scanned document, 6 pages | 2.9 MB | 256 KB (−91%) | ~1.4 s |
| Photo deck, 8 slides | 4.0 MB | 1.4 MB (−66%) | ~2.3 s |
| Long report, 245 pages, mostly text | 22.6 MB | would grow — original kept | ~10 s |
| Text report, 20 pages | 96 KB | would grow — original kept | ~1.2 s |
Chrome on a Mac. Phones are several times slower but give the same sizes, except where the mobile pixel cap kicks in.
The honest limitations
- Text stops being text. Every page becomes a picture. You can’t select, search or copy text in the result, and screen readers can’t read it. For contracts you need to search later, use the desktop app.
- Text-only PDFs get bigger. A typed page is a few kilobytes of instructions (“draw these glyphs here”). The same page as a 150 dpi JPEG is tens of kilobytes. kompakt detects this and tells you to keep the original instead of handing you a worse file.
- JPEG is the only codec. Browsers can encode JPEG, PNG and WebP, and PDF can’t embed WebP. For black-and-white scans, the ideal codecs (JBIG2, CCITT G4) aren’t available in the browser at all — which is one reason the desktop app (Ghostscript) beats the browser on scans.
- Phones have a pixel budget. Mobile Safari limits canvas size and can reload a tab that runs out of memory. On touch devices kompakt caps a page at about 4 megapixels, so “high” on a phone renders at roughly 200 dpi for a Letter page.
- Forms, links and annotations are flattened into the image. Fillable forms stop being fillable.
When to use which
Scans and photo-heavy PDFs you only need to read or send: browser. Documents whose text must stay selectable, or whose size comes from fonts: desktop app.
Related: quality at each setting, with crops · why PDFs get large.