The Memory Wall is Now a Memory Bill

Arthur Rasmusson Director of AI Infrastructure at Lightbits Labs
Arthur Rasmusson
Director of AI Architecture
October 11, 2026

Context windows grew roughly a hundredfold in three years. High-bandwidth memory (HBM) per GPU did not. That gap is the whole story of inference economics right now, and most of the industry still describes it in the wrong units. It gets filed as a capacity problem, something you solve once by buying a bigger card. It is not. It is a recurring bill, charged per conversation, paid in GPU-seconds.

What is the KV cache? When a model reads your prompt, every token it attends to produces a key and a value vector at every layer. Those vectors are the model’s working memory of the conversation so far: they are what lets token 900,000 attend to token 12. Together they are the KV cache. It is not the model weights, which are fixed and shared across every request; it is per-session state, it grows linearly with context length, and at long context it dwarfs the weights. Throw it away and the only way to get it back is to run the entire prompt through the model again. That re-run is called prefill, and it is the single largest avoidable cost in inference today.

The Ladder Nobody Budgets For

Memory in a serving node is a ladder, and each rung is roughly an order of magnitude larger and an order of magnitude slower than the one above it:

TierTypical capacity per nodeRole today
HBM~100 GBWhere the KV cache lives, or doesn’t
DRAM~1 TBWeights staging, not cache
NVMe~10 TBCheckpoints, datasets
Fabriceffectively unboundedModel distribution

Look at where the KV cache sits. It lives on the top rung, the smallest one, and when it no longer fits there, it is discarded. Not spilled, not tiered. Discarded. The three rungs below it, holding a hundred times the capacity, are not in the conversation at all.

Every other layer of the stack learned this lesson decades ago. Databases page to disk. Filesystems have a page cache. CPUs have L1, L2, L3, and then RAM. Inference serving, uniquely, has exactly one tier and a delete key.

What Rebuilding Actually Costs

This matters because rebuilding is not cheap, and it is not linear the way people assume. Attention is quadratic in sequence length; prefill time grows accordingly. From the benchmark ladder on open.lightbits.ai, the cost to rebuild a session from cold:

ContextTime to rebuild via prefill
100K tokens7.4 seconds
1M tokens2 minutes
10.5M tokens1 hour 42 minutes

That last row is the one that reframes the problem. An hour and forty-two minutes of eight-GPU time, to return a session to exactly the state it was in when the user closed the tab. Nothing was learned. No tokens were generated for the user. The fleet simply paid to get back to where it already was.

At 100K it looks like a latency annoyance. At 10.5M it is not a latency problem at all; it is a capacity decision. You are choosing to spend an hour and forty-two minutes of your most expensive asset on recomputation, or you are choosing not to offer the session.

The tax is recurring, which is the part that hurts

illustration depicting the recurring GPU inefficiency tax solved by KV cache
The loop that makes this a bill rather than a purchase. Nothing in it produces a token the customer pays for.

A capacity problem you pay once. This you pay every time.

Consider what returning traffic actually looks like. A coding agent resumes a session and replays its trajectory. A support conversation picks up after lunch. A document analysis tool opens the same 400-page contract for the fourth reviewer this week. In each case, the model has seen this exact prefix before, often minutes ago, often on the same node, and in each case it computes the whole thing again from scratch, because the only place that state could have lived was HBM, and HBM was needed for someone else.

This is why the framing matters. If you treat it as capacity, you buy more HBM, and you discover that HBM per dollar is improving far more slowly than context lengths are growing. You cannot buy your way up a curve that steep. If you treat it as a recurring bill, you start asking a different question: why am I paying this more than once?

A Storage Problem Wearing a GPU Costume

Here is the thesis, stated plainly: the KV cache is a storage problem that has been misfiled as a compute problem.

It has every property of a storage workload. It is large. It is written once and read many times. It has strong locality: the same prefixes recur constantly. It is perfectly cacheable, because it is deterministic: the same prompt through the same model produces the same cache, bit for bit. And it has a natural hierarchy available to it, three rungs of it, sitting idle.

The reason it has not been treated that way is mechanical rather than philosophical. Moving a KV cache off the GPU and bringing it back fast enough to matter requires the movement to be invisible against attention’s own timeline: the block has to be in flight before the kernel asks for it. That is an RDMA and prefetch problem, and it is tractable. When it is solved, the tier ladder opens up and the cache stops being something you throw away.

What that looks like in practice is the subject of the rest of this series. The short version: a 10.5M-token session that costs an hour and forty-two minutes to rebuild can be restored in seconds, with bit-identical output. GPU-hours that would have gone into recomputation become GPU-hours you can sell.

Why the Gain Grows With the Context

The size of the saving is not a fixed multiple. It tracks sequence length because of the same quadratic term that creates the problem in the first place.

Prefill computes attention between every pair of tokens in the prompt, so its cost carries a term that grows with the square of the context. Restoring a cache does not work that way: a KV cache is linear in tokens, so twice the context is twice the bytes to read back. Recompute is superlinear, restore is linear, and the gap between them widens with every additional token of context.

That is why results improve as sessions get longer, the opposite of the usual pattern where an optimization washes out at scale. Published joint results with FarmGPU and ScaleFlux put time-to-first-token improvements between 100x and 280x across models from 131K to 1M tokens, and beyond 1000x at 10M tokens.

At those lengths, prefill is not one component of a turn; it is very nearly the whole turn, so the improvement a user waits through and the improvement in GPU-seconds consumed are the same order of magnitude. This doesn’t help at the other end of the scale: a short prompt with a long answer is mostly decode; decode is untouched, and there is little for a cache tier to recover. This is long-context infrastructure, and the returns track context length.

Pricing: What You Already Pay for Twice

The useful first move is not architectural; it is arithmetic: find out what rebuilding currently costs you. Take your traffic mix, your typical context lengths, and your reuse share, and price the prefill you are paying for more than once.

The calculator at open.lightbits.ai does exactly this. Give it your fleet and your workload and it will show you the recomputation line as a cost, in dollars and in GPU-hours, alongside what it would cost not to pay it twice.

Methodology note. The calculator is a simulation parameterized by measured benchmarks, not a measurement of your fleet. Treat its output as a hypothesis worth testing, and validate with a proof of value before committing capital.

Or, if you would rather not click through a form, hand the problem to an agent.

Next in this series: Why token economics matter more than GPU count. Why $/GPU/hr is the wrong unit, and what replaces it.

About the writer
Arthur Rasmusson Director of AI Infrastructure at Lightbits Labs
Arthur Rasmusson
Director of AI Architecture