What compression actually does to the text in your PDF — Recto
Recto

Notes

What compression actually does to the text in your PDF

17 August 2026 · 5 min read

Compress a PDF, open it, and it looks fine. Then you try to find a word in it and nothing happens.

Two documents that look identical

Open a PDF and try to select a line of text. If a cursor appears and words highlight, the document contains real text: character codes, fonts, and positions. Search works, copy works, a screen reader can read it, and the file is small because letters are cheap to store.

If instead you get a rectangular selection box over an image, there is no text in the document at all. There is a picture of text. It looks the same on screen and it is a fundamentally different object — unsearchable, uncopyable, invisible to assistive technology, and considerably larger.

Documents arrive in both states. Anything exported from a word processor has real text. A photograph of a page does not, unless something has run OCR over it and written an invisible text layer underneath.

How compression destroys one and not the other

The crude way to shrink a PDF is to render every page to a bitmap and compress the bitmaps. Applied to a photograph, this is exactly right — the page was already an image and you are simply storing it more efficiently.

Applied to a text page it is vandalism. The letters become pixels. The searchable layer, if there was one, is discarded along with everything else. You have made the file smaller by converting a useful document into a picture of a useful document, and there is no way back short of re-running OCR and hoping.

The damage is easy to miss because the visual result is close to the original. You will not notice until months later, when you need to find something.

Looking before touching

The alternative is to inspect each page before deciding what to do with it. Reading a PDF's content stream tells you what is actually being drawn: how many glyphs, how many images and what area they cover, how many painted vector paths.

Those signals separate the cases cleanly. Glyphs present means real text, so leave the page alone. No glyphs, no vector art, and a substantial image means the page is essentially a photograph, so rebuild it freely. The interesting cases are in between, and the right instinct there is caution: when the evidence is unclear, copy the page across untouched and accept the larger file.

Recto works this way, and pages it leaves alone are carried byte-for-byte — not re-encoded at high quality, not re-saved, literally the same bytes moved into the new document. A page that is not modified cannot be damaged.

Checking for yourself

After compressing anything you care about, open the result and try to select a sentence on a page that had text before. Then search for a word you can see. Thirty seconds, and it tells you whether the tool understood what it was working on.

It is worth doing once with any new tool, on a document you can afford to lose. What you learn applies to everything you run through it afterwards.

searchable pdf · pdf text layer · compress pdf

All notes