Why does a long conversation get slower and cost more than a short one?
the paper →Fast Transformer Decoding: One Write-Head is All You NeedShazeer, 2019 ↗Nine pages, and the one that names the problem: decoding is bound by reloading keys and values, not by arithmetic. Multi-query attention is the fix proposed here.Previously computed state is worth more than recomputing it.
Why does a long conversation with a chatbot feel slower than a short one? It is not thinking harder. Before writing each word, the model reads back a notebook containing everything already said — the whole notebook, every word, every time. That notebook is the KV cache. It is why your conversation has a price, and why a machine with 640 GB of memory serves fifty people rather than five thousand.
On the machineThis lands on HBM, station 3 of 10 on the path the constraint took through the hardware. The second occupant of HBM, and the only one that grows with every conversation. See it on the drawing →The step waits for the longer one. Right now that is memory traffic — and what grows with the conversation is no longer the wait. It is how much of the machine you are holding, which is the row below.
Read as far as you want. Each level assumes the one above it and nothing more.
Every word the model has already read contributes two vectors to attention: a key and a value. They are fixed the moment that word exists, and nothing later changes them. The KV cache is simply those vectors, held in the GPU’s memory so each new word can attend over the past without rebuilding it from scratch. It is memory spent to delete arithmetic, and the bill arrives in bandwidth.