Why did transformers replace RNNs?
Transformer: Parallelism is what turns money into capability.
the paper →Attention Is All You NeedVaswani et al., NeurIPS 2017 ↗The idea that let a model read a whole sentence at once instead of one word at a time. It is why training got fast enough to build ChatGPT and Claude — and every conversation since has come with a growing memory bill nobody has found a way to erase.
The room it came out of
In 2016 Google had just rebuilt Translate on eight stacked layers of recurrent networks. It translated better than anything before it and it was brutally expensive, because a recurrent network reads one word at a time and a warehouse of chips can only wait. Eight people set out to make that cheaper. The paper is benchmarked on English-to-German and English-to-French; there is not a word in it about chatbots.
How it works
Before 2017, a model that read text read it the way you do: one word, then the next, carrying a summary forward. That is a recurrent network, and it has a fatal property for anyone with a warehouse full of chips — step ten cannot begin until step nine has finished. You can own ten thousand processors and the sequence will still be walked single file.
The Transformer's move was to delete that dependency. Instead of passing a summary forward, every word looks at every other word directly, all at once, and the model works out how much attention each one deserves. Nothing waits for anything. A sequence that took a thousand sequential steps to read now takes one very wide step, and a very wide step is precisely what a GPU is for.
That is the whole reason the last decade happened. Not that attention is a cleverer way to read — it is that attention is a *parallel* way to read, and parallelism is the only thing that converts money into capability. Every scaling law, every hundred-million-dollar training run, every argument about compute budgets sits downstream of a decision to make the arithmetic wide instead of deep.
The bill came due at the other end. If every word must see every other word, then the model has to hold something about every word it has already read — and that state grows with the conversation, without limit, forever. In 2017 nobody minded: sequences were a few hundred words and the state was a rounding error. It is not a rounding error now, and almost everything else in this atlas is a response to it.
So the honest summary is not "the Transformer made models better." It is that the Transformer traded a constraint nobody could pay for a constraint everybody could — and then spent ten years discovering how expensive the second one turned out to be.
What it traded
- gave up
- the ability to process a sequence one step at a time, cheaply
- got
- the ability to train on the whole sequence at once
What exists now that didn’t before
Every word now has to see every other word — so remembering a conversation means holding a slice of it in memory for every word already said, and that pile only grows.
Asked, and answered
What did a paper about translating sentences have to do with chatgpt?
Everything — it's the same architecture. The 2017 paper was about making translation faster to train, not chatbots; the trick that did it (read the whole sentence at once instead of one word at a time) turned out to be what made every model since possible.
Read next
- KV cache →The first and largest bill: what the model has to hold about every word it has read.
- Mamba →The counter-argument. Keep a fixed summary after all, and see what it costs you.
- The GPUThe machine the parallel decision was made for, and the reason it was the right decision.