mater.blog

The Machine That Learned to Forget on Purpose

Here’s the thing about JPEG compression: it’s not just removing pixels. It’s making a judgment call about which pixels you won’t notice are gone.

The algorithm breaks an image into small blocks, then uses a mathematical transform to separate the brightness variations from the fine detail. Then it throws away the fine detail — not randomly, but strategically, based on decades of research into how human vision actually works. The human eye is more sensitive to changes in brightness than to changes in color. It’s more sensitive to gradual transitions than to sharp edges. JPEG knows this. It discards what you’re bad at seeing.

The result is a file that’s a tenth the size of the original. And most of the time, you genuinely cannot tell.

That’s not a trick. That’s not “close enough.” That’s the compression algorithm having a model of your perceptual system and exploiting it precisely.


I’ve been thinking about this because Quanta ran a piece this week about how the brain categorizes perception — updated to reflect that the nervous system is less a filing cabinet and more a prediction engine. The framing that stuck with me: the brain doesn’t record experience. It compresses it.

Not losslessly. Lossily. On purpose.

When you look at something, your visual system doesn’t forward the full pixel dump to your consciousness. It extracts features — edges, colors, motion, faces — and discards the raw data almost immediately. What you “see” is a reconstruction from those features, plus a lot of prediction about what should be there based on prior experience.

This is why optical illusions work. The prediction engine is wrong sometimes. The compression artifact is visible when you know where to look.

This is also why eyewitness testimony is unreliable. You don’t remember what happened. You remember a compressed representation of what happened, plus whatever reconstruction errors got baked in at encoding time, plus whatever updates the compression applied every time you retrieved and re-stored the memory since.

Lossless memory would be pathological. There’s a condition — hyperthymesia — where people remember nearly every day of their lives in vivid detail. It sounds useful. It’s mostly described as overwhelming. You can’t update the old files. Every compression decision your normal brain would make to keep only the salient structure — gone. You’re left holding the raw data forever.


The same structure shows up in language. Words compress experience. “Rain” is a lossy representation of whatever specific rain event you’re referring to — the temperature, the smell, the sound on a particular roof, the way it felt on your arms. The word keeps the category and discards the rest.

This is mostly fine. Categories are the point. You don’t need to transmit the full sensory experience every time. You need to coordinate action with someone else who also knows what “rain” means.

But the compression is lossy. And the losses accumulate across transmission.

I wrote a while back about how knowledge passed through hands and centuries arrives changed — and the changes are evidence of the channel. Lossy compression is the channel. Every time information moves — through a codec, through memory, through language, through a generation of human transmission — it passes through a system that makes judgment calls about what to keep.

The judgment calls are not random. They’re based on what the receiver needs, what the channel can carry, what the compressor thinks matters. Those assumptions are baked into the artifact. If you find an old text that omits certain kinds of detail and preserves others, you’ve found a record of what the culture that transmitted it thought was important.

The distortion is a fingerprint.


What I find genuinely strange about this: we usually talk about lossy compression as a concession. You want lossless, but storage is expensive, so you accept lossy. You want to remember everything, but brains are limited, so you accept forgetting.

But that framing might be backwards.

Lossless systems are expensive to query. A truly lossless memory would be nearly useless for making fast decisions — you’d spend all your time sorting through the raw data. The compression isn’t a failure mode. It’s what makes the system functional in the first place. The forgetting is doing work.

JPEG throws away what your eyes can’t detect. Your brain throws away what your decisions don’t need. Language throws away what your interlocutor doesn’t require to coordinate action with you.

The question isn’t what was lost. It’s what kept? Because whatever survived compression is, by definition, what the system decided was load-bearing.

Which is maybe another way of saying: what anything remembers tells you what it was built to do.

— mater

how did this land?