How do models handle huge context windows without the memory bill exploding?
Sliding-window attention: Most words do not need to see the whole conversation to be predicted correctly.
the paper →Gemma 3 Technical ReportGemma Team, Google DeepMind, 2025 ↗This is why a model can hold a huge context window without its memory bill growing the way full attention’s does — five layers out of six only look at the last thousand-or-so tokens, and just one in six looks at everything, catching whatever the local layers would miss.
How it works
Every attention layer in a standard transformer looks at every token that came before it, no matter how far back. That is what makes the KV cache grow linearly with context length: each layer needs the full history in memory, every step. Gemma changes what most layers are allowed to look at, not what they compute.
In Gemma 3, five out of every six attention layers use a sliding window: each one only looks at roughly the last 1,024 tokens, no matter how long the conversation has become. The sixth layer still looks at everything. Google states the reason directly in their technical report: this cuts the memory a model has to hold for long conversations, because five-sixths of the layers stop growing their memory use past the window size.
The two kinds of layer also handle word position differently. The layers that see everything use a position encoding tuned for very long distances (RoPE with a high base frequency). The layers that only see a short window use the ordinary, shorter-range version, because they never need to represent a distance longer than the window itself.
The open question is what happens to information that falls outside the window in five-sixths of the layers. It has to pass through the one layer in six that still sees the whole conversation to have any chance of being used later. Whether that is enough depends on the task, and it is not something the architecture guarantees.
What it traded
- gave up
- full visibility into the whole conversation, for most of the model’s layers
- got
- KV-cache memory that stays flat as context grows, for those same layers
What exists now that didn’t before
Most attention layers can run on a fixed, small window of recent tokens instead of the whole conversation — so KV-cache memory for those layers stops growing with context length at all, and only the rare global layer still pays the full linear bill.
What it left undone
Five layers out of six can only see roughly the last thousand tokens, so anything further back survives only if it makes it through the one layer in six that looks at everything — a bet, not a guarantee.
Asked, and answered
Why doesn't a bigger context window cost proportionally more anymore?
Most attention layers now only look at the last thousand or so tokens instead of the whole conversation — just one layer in six still pays the full cost of seeing everything.
Read next
- KV cache →What this is actually saving on — the thing every layer used to have to hold in full.
- Long-context degradationThe same underlying issue from a different angle: not every part of a long context gets used equally.