How a chatbot writes one word
Follow one prompt through an open frontier model, drawn in 3D, from the electricity that moves it to the word on your screen. Why the reply comes one word at a time, why memory, not maths, is the expensive part, and what each word costs in electricity and water.
trying to answer →What becomes scarce when intelligence becomes cheap?
pick a question; the machine answers it
What happens between pressing enter and the first word appearing?
This is DeepSeek V3, an open model whose makers published its numbers. You type “The capital of France is”. Follow it from the electricity that moves it to the word on your screen, in seven steps.
Everything here is inference: using a finished model. Training wrote the recipe once, over months. Inference follows it once per word, for every person asking, which is why it now uses most of the world’s AI chips.
Where does the electricity go?
Electricity reaches the chip and flips billions of tiny switches. Every flip ends up as heat.
Each chip draws about 700 W, a small microwave running flat out, all day, and there are thousands of them in a hall. Almost none of that ends up as anything but heat, which is why the building around the chips is mostly plumbing.
≈700 W per GPU
Each H800 draws up to about 700 W. That power drives the transistors that do the arithmetic and move the numbers, and all of it leaves as heat for the cooling loop on sheet 2 to carry away.
How does a sentence become numbers?
Your words are cut into tokens, pieces of words, and each token becomes a list of numbers.
Each token gets an address: a point in a space with 7,168 directions instead of three. Words used in similar ways end up near each other, so “Paris” sits closer to “Rome” than to “carrot”.
7,168 numbers per token
Each token number is swapped for a list of 7,168 numbers, looked up in a learned table.
What is the model, physically?
The model itself is a huge set of numbers learned in training. Only some are used for each token.
The model is a recipe written in numbers: 671 billion of them, about 671 GB at one byte each. It is fixed once training ends. Answering you never changes it; that is what inference means.
671 billion held · 37 billion used
The 671 billion numbers were set by reading 14.8 trillion tokens and learning to predict the next one.
What does the chip actually do with your words?
The chip multiplies your numbers by the model’s numbers and adds them up, layer after layer.
A GPU is a field of calculators, each doing one tiny sum at a time: multiply two numbers, add to a running total. If all 8 billion people on Earth did one sum a second, without sleeping, they would need about 3 days to match one second of this chip. One token takes about 9 sums from every person alive.
why this paper matters Mixture of experts 2021 →The step where each token looks back at the earlier ones to decide what matters: each layer lets the newest token check every earlier one before the experts do their sums.
why this paper matters Transformer 2017 →≈74 billion operations per token
Each new token looks back at every earlier one, through a cache of what was already worked out.
Why is memory the expensive part?
Most of the energy goes into fetching the model’s numbers from memory, not into multiplying them.
The calculators are fast; the bookshelf is slow. For every word, the chip has to read the part of the recipe it needs, 37 GB, from the memory stacked beside it. That read alone takes about 11 milliseconds, however fast the sums are, so the calculators spend most of their time waiting.
why this paper matters KV cache 2019 →≈0.38 J to write · ≈0.08 J to read
The reply comes one token at a time, and each one needs the model’s numbers fetched again.
How does it pick the next word?
The last layer scores every possible next token, and one is picked and turned into text.
129,280 scores
The picked token is turned back into text and sent to you, and the loop runs again.
Why does the reply come one word at a time?
Then all of it runs again for the next token, one at a time, until the reply is done.
Like a writer who can only see what is already on the page: each new word is added to the prompt and the whole journey runs again. A 500-token reply is 500 trips through all 61 layers.
≈20 tokens/s · 500 tokens ≈ 189 J
The picked token is turned back into text and sent to you, and the loop runs again.
that was one question. ask another?