How kompakt compresses PDFs inside your browser

No server, no upload: pdf.js renders each page to a canvas, the canvas becomes a JPEG, and a 70-line writer streams the JPEGs into a new PDF. What that buys you, what it costs, and the bugs we found measuring it.

Mikhail Yevdokimov · 2026-10-10 · 9 min read

Most “compress PDF online” sites work the same way: you upload the file, a server runs Ghostscript or a commercial SDK on it, and you download the result. That is a reasonable design — and also a non-starter for a contract, a medical report or a bank statement you would rather not hand to a stranger’s server.

kompakt’s web compressor does the whole job in the browser tab. The file is read with the File API and never leaves the device; after the first visit the page works offline. This article explains how, with the trade-offs stated plainly.

The engine is open source: github.com/xronocode/kompakt-web — MIT, with tests and the benchmark used for the numbers below.

The pipeline

PDF file ──► pdf.js parses and renders page N ──► <canvas> (pixels)
                                                     │
                                       canvas.toBlob('image/jpeg', q)
                                                     │
                     new PDF ◄── streaming writer ◄── JPEG bytes
  1. Parse and render. pdf.js — the PDF engine inside Firefox — opens the file and draws one page at a time onto an HTML canvas at the chosen resolution.
  2. Encode. The browser’s built-in JPEG encoder turns the canvas into a JPEG at the chosen quality.
  3. Write. A small hand-written PDF writer appends the JPEG as an image object, adds a page that draws it, and moves on. At the end it writes the cross-reference table and trailer.

Then the canvas is released and the next page starts. Only one page is ever held as pixels: memory use is the input file, one rendered page, and the growing output — not every page at once.

Why a custom writer instead of a PDF library

The first version used pdf-lib to assemble the output. It works, but it builds the whole document as a tree of objects in memory and serializes it at the end — fine for 20 pages, painful for 500 pages of 300 dpi images on a phone.

The output we need is extremely simple: every page is one image. So the writer only has to emit four kinds of objects — catalog, page tree, page, image — and remember the byte offset of each for the cross-reference table. It writes bytes as it goes and hands the browser a list of chunks (new Blob(chunks)), so nothing is ever concatenated into one giant buffer. It is about 70 lines and has no dependencies.

Resolution and quality: what the settings mean

SettingRender resolutionJPEG qualityLetter page in pixels
low72 dpi0.50612 × 792
medium120 dpi0.651020 × 1320
high200 dpi0.751700 × 2200

Three bugs we found by measuring

Writing this article, we ran the compressor against a fixed test set and inspected the output files with pdfinfo, pdfimages and Ghostscript instead of eyeballing them. That turned up three real bugs, all fixed on 10 October 2026:

  1. The dpi labels were wrong. The render scale was computed as dpi / 96 — the CSS pixel density — but a pdf.js viewport at scale 1 is 72 pixels per inch (one pixel per PDF point). So “150 dpi” really rendered at 112 dpi, “300” at 225 and “72” at 54. The scale is now dpi / 72, and the presets were re-tuned to 72 / 120 / 200 dpi, which keep the old file sizes at medium while making low actually legible and high actually sharp.
  2. The page size changed. The writer used the image size in pixels as the page size in points, so a Letter page compressed at medium came out as 13.3 × 17.2 inches (and shrank to 6.4 × 8.3 inches at low). On screen nobody notices; in print or when inserting the page into another document, everyone does. Pages now keep their original size and the image is scaled onto them.
  3. The output was technically malformed. A missing newline produced endstream immediately followed by endobj. Viewers silently repair this, but strict validators, print workflows and document-management systems may reject the file. Every output now passes pdfinfo and Ghostscript without errors.

Measured results

Full numbers are in the quality comparison; the short version, for the browser compressor at medium:

Test fileBeforeAfter (medium)Time
Scanned document, 6 pages2.9 MB256 KB (−91%)~1.4 s
Photo deck, 8 slides4.0 MB1.4 MB (−66%)~2.3 s
Long report, 245 pages, mostly text22.6 MBwould grow — original kept~10 s
Text report, 20 pages96 KBwould grow — original kept~1.2 s

Chrome on a Mac. Phones are several times slower but give the same sizes, except where the mobile pixel cap kicks in.

The honest limitations

When to use which

Scans and photo-heavy PDFs you only need to read or send: browser. Documents whose text must stay selectable, or whose size comes from fonts: desktop app.

Related: quality at each setting, with crops · why PDFs get large.