Jianxin Shen
Notes

The instinct to compress

A storage decision that cost almost nothing exposed a scarcity reflex that outlives the constraint it came from.

Revised Jianxin Shen

I was sizing a storage pipeline this week. A Batch job downloads a daily snapshot, about 10 GB, and lands it in S3. Downstream services read recent files; older ones rotate to cold storage automatically. Three components: a schedule rule, a job, and a lifecycle policy.

EventBridge trading days Batch dl + upload Standard recent, read by downstream Glacier lifecycle rotation

I ran the numbers. 30 days of snapshots is roughly 300 GB, about $7 a month in S3 Standard. After rotation, Glacier brings that under $1. Requests, uploads, transitions are all negligible. The whole thing costs less per year than a meal.

My first impulse was still to write an aggregation step. The snapshots are full-state dumps. I could diff consecutive days, keep only the deltas, drop the unchanged portions. The engineering seemed obvious, almost obligatory.

5:1 compression would save about $5 a month.

Monthly Annual
Raw snapshots, 30 days in Standard, then Glacier $7 $84
With 5:1 aggregation $2 $24
Savings $5 $60

For $60 a year I would need to maintain a transformation job, handle schema changes, and give up the ability to reprocess raw data when the aggregation turns out wrong. I closed the notebook and left the snapshots as they were.

The decision took minutes. The discomfort lasted longer.

I have felt this before. A colleague allocates generously and I notice myself wanting to tighten it, not because of a constraint anyone has stated, but because something in me insists that uncompressed data sitting on a disk is waste. The arithmetic says otherwise. The arithmetic is right. The feeling does not update.

I think I know where this comes from. For my cohort, disk was the resource you learned to respect early. 10 GB was a serious number. You compressed, pruned, garbage-collected, not for elegance, but because the drive would fill. Cloud storage removed the constraint years ago. The reflex stayed.

Engineers a generation younger carry the same pattern aimed at a different target: memory. On the CPU it shows up as small habits. Object pools and arena allocators. Flags packed into bitfields. Struct fields reordered to squeeze out padding. A cache the numbers plainly support, refused anyway. A day spent saving two hundred megabytes on a machine with sixty-four gigabytes to spare.

On the GPU the same hands are, for now, right. The arithmetic is short and unforgiving. A seventy-billion-parameter model stored at two bytes per weight is 140 gigabytes. An H100 holds 80. The model does not fit before it has read a single word.

So you give up precision. Quantize the weights to four bits and the model shrinks to about 35 gigabytes. Then it starts working, and memory starts disappearing again. Every token it reads leaves keys and values behind at every one of its eighty layers: about 320 kilobytes per token. A conversation of four thousand tokens holds 1.3 gigabytes. Thirty-two conversations at once hold 42. Weights and cache together now sit two gigabytes under the ceiling. The cache is allocated in small fixed pages, like virtual memory, because reserving room for the longest possible reply would waste most of what is left.

Arithmetic, one 80 GB card

Fill the card

A seventy-billion-parameter model and the cache its conversations leave behind. Add one more conversation and see what happens.

Weights
Tokens per conversation

35.0 GB weights + 42.9 GB cache = 77.9 GB. 2.1 GB left.

Constructed arithmetic, not a measurement. Weights at the chosen precision plus the key-value cache of a model shaped like common 70B models: eighty layers, eight key-value heads of 128 dimensions, cache kept at 16 bits, about 320 KB per token. Activations, framework overhead and fragmentation are left out, so a real card runs out sooner.

Attention itself would not fit if you let it. At 32,000 tokens, the score matrix for one head in one layer has a billion entries, two gigabytes at half precision. So it is never written out. It is computed in tiles small enough for the couple of hundred kilobytes of fast memory beside each compute unit, and the softmax is assembled piece by piece as the tiles pass through. Training is worse. The activations the backward pass needs are thrown away and recomputed, about a third more compute spent to avoid storing them.

When the arithmetic is off by a little, nothing slows down. The process dies: CUDA out of memory, tried to allocate 2.00 GiB. You learn to read nvidia-smi the way my cohort read df -h.

There the constraint is real and current; on the CPU it mostly passed years ago. The discomfort feels the same either way — that is what makes it a reflex rather than a judgment. It does not check whether the pressure is current. It fires.

What I am describing is not really about storage or memory. It is an instinct that forms around a real constraint and stays in the body after the constraint passes. The mind updates. The hands do not. Technology makes this easy to see — you run the numbers, get a clear answer, and notice your hands still want to tighten. But the instinct also runs where no arithmetic is available.

My grandparents’ generation had a version of this with food. People who lived through real shortage, not inconvenience but the kind where villages went quiet, do not throw away rice. They finish what is on the plate, store what cannot be finished, save containers and bags against a future they already know is unlikely. The body learned “not enough” and did not accept the correction.

The scarcity was not natural. Forced collectivization, fabricated production reports, grain exports continued while people starved. Tens of millions died. The policy that caused it is still described in official language as exploration, as tuition, as a natural disaster with an unfortunate duration.

Taiwan had its own terror in the same decades. The state killed and imprisoned dissidents, made silence a survival skill for a generation of intellectuals. It did not starve people. And it was eventually named: archives opened, victims listed, a public record assembled, however incomplete. The wound got a word for itself. That does not heal it, but it lets someone trace a reflex back to its origin.

When the record denies the wound, the reflex still runs. You finish every grain. You cannot explain why, exactly, because the explanation was removed from the language you were given.

A disk bill is not that kind of loss. But the hand tightens the same way.

I left my snapshots uncompressed. The lifecycle policy handles the rest. There is no aggregation step to maintain, no schema-change bugs to chase, no loss of raw data.

The discomfort faded once I recognized what it was. Not a signal about the pipeline. An echo from a time when the constraint was real.