Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Visualization · 26 Sept 2026

How a chatbot writes one word

Follow one prompt through an open frontier model, drawn in 3D, from the electricity that moves it to the word on your screen. Why the reply comes one word at a time, why memory, not maths, is the expensive part, and what each word costs in electricity and water.

trying to answer →What becomes scarce when intelligence becomes cheap?

pick a question; the machine answers it

SHEET 3 · ONE TOKEN, FROM THE WATT TO THE WORDDEEPSEEK V3 · 671B PARAMETERS, 37B USED PER TOKEN · SERVED ON H800 GPUS · READ 1 → 761 LAYERS · 256 EXPERTS, 8 USED≈700 W IN · FROM SHEET 2GPU · H800BILLIONS OF SWITCHESReported: board power of an H800, the same class of part as an H100 SXM.≈700 WRALL OF IT LEAVES AS HEATCACHED PROMPTS · 56%YOUR PROMPT“Why is the sky blue?”TOKENIZERWhyistheskyblue?6 TOKENS · 129,280 POSSIBLEEMBEDDING7,168 NUMBERS PER TOKENPREFILL · WHOLE PROMPT AT ONCEInference: 700 W ÷ (73,700 ÷ 8 tokens/s per GPU), from DeepSeek's reported prefill throughput. GPU power only.≈0.08 J PER PROMPT TOKENIFP8 · 1 BYTE / NUMBERHBM80 GB · 3.35 TB/SFETCHED FOR EVERY TOKENBATCHING · SHARED FETCHESOUTPUT HEAD129,280 SCORES, 1 PICKEDDRAFTS THE NEXT TOKENTHE REPLYThe sky looks bluebecause air scattersshort blue lightmore than red ▍Company claim: DeepSeek's reported average output speed for each user.20–22 TOKENS/SCAFTER LAUNCH · PRICE CUTSDECODE · THEN AGAIN, ONE TOKEN AT A TIMEInference: 700 W ÷ (14,800 ÷ 8 tokens/s per GPU), from DeepSeek's reported decode throughput; about 5× a prompt token. GPU power only.≈0.38 J PER WRITTEN TOKENIElectricity reaches the chip and flips billions of tiny switches. Every flip ends up as heat.1Electricity reaches the chipand flips billions of tinyswitches. Every flip ends upas heat.≈700 W per GPUtoken: A piece of a word. Common words are one token; rare ones are several.2Your words are cut intotokens, pieces of words, andeach token becomes a list ofnumbers.7,168 numbers per tokentoken: A piece of a word. Common words are one token; rare ones are several. training: Adjusting the model’s numbers, over months, until its guesses come out right. Done once, then the model is frozen.3The model itself is a hugeset of numbers learned intraining. Only some are usedfor each token.671 billion held37 billion usedlayer: One stage the numbers pass through; V3 has 61, one after another.4The chip multiplies yournumbers by the model’snumbers and adds them up,layer after layer.≈74 billion operationsper tokenMost of the energy goes into fetching the model’s numbers from memory, not into multiplying them.5Most of the energy goes intofetching the model’s numbersfrom memory, not intomultiplying them.≈0.38 J to write≈0.08 J to readtoken: A piece of a word. Common words are one token; rare ones are several. layer: One stage the numbers pass through; V3 has 61, one after another.6The last layer scores everypossible next token, and oneis picked and turned intotext.129,280 scorestoken: A piece of a word. Common words are one token; rare ones are several.7Then all of it runs againfor the next token, one at atime, until the reply isdone.≈20 tokens/s500 tokens ≈ 189 JTHE LIFE OF ONE MODELH800 GPU-HOURS, DRAWN TO SCALE · DEEPSEEK V3 AND R1PRETRAININGReported by DeepSeek; see the stage for its source.2,664KRCONTEXT EXTENSIONReported by DeepSeek; see the stage for its source.119KRPOST-TRAINING (V3)Reported by DeepSeek; see the stage for its source.5KRR1: LEARNING TO REASONReported by DeepSeek; see the stage for its source.≈180KRONE DAY OF SERVINGInference: 226.75 servers × 8 GPUs × 24 hours, from DeepSeek’s reported day.≈43.5KI→ SERVING PASSES ALL OF PRETRAINING IN ≈61 DAYS

What happens between pressing enter and the first word appearing?

This is DeepSeek V3, an open model whose makers published its numbers. You type “The capital of France is”. Follow it from the electricity that moves it to the word on your screen, in seven steps.

Everything here is inference: using a finished model. Training wrote the recipe once, over months. Inference follows it once per word, for every person asking, which is why it now uses most of the world’s AI chips.

1 of 7

Where does the electricity go?

Electricity reaches the chip and flips billions of tiny switches. Every flip ends up as heat.

Each chip draws about 700 W, a small microwave running flat out, all day, and there are thousands of them in a hall. Almost none of that ends up as anything but heat, which is why the building around the chips is mostly plumbing.

≈700 W per GPU

Each H800 draws up to about 700 W. That power drives the transistors that do the arithmetic and move the numbers, and all of it leaves as heat for the cooling loop on sheet 2 to carry away.

2 of 7

How does a sentence become numbers?

Your words are cut into tokens, pieces of words, and each token becomes a list of numbers.

Each token gets an address: a point in a space with 7,168 directions instead of three. Words used in similar ways end up near each other, so “Paris” sits closer to “Rome” than to “carrot”.

7,168 numbers per token

Each token number is swapped for a list of 7,168 numbers, looked up in a learned table.

3 of 7

What is the model, physically?

The model itself is a huge set of numbers learned in training. Only some are used for each token.

The model is a recipe written in numbers: 671 billion of them, about 671 GB at one byte each. It is fixed once training ends. Answering you never changes it; that is what inference means.

671 billion held · 37 billion used

The 671 billion numbers were set by reading 14.8 trillion tokens and learning to predict the next one.

4 of 7

What does the chip actually do with your words?

The chip multiplies your numbers by the model’s numbers and adds them up, layer after layer.

A GPU is a field of calculators, each doing one tiny sum at a time: multiply two numbers, add to a running total. If all 8 billion people on Earth did one sum a second, without sleeping, they would need about 3 days to match one second of this chip. One token takes about 9 sums from every person alive.

why this paper matters Mixture of experts 2021 →

The step where each token looks back at the earlier ones to decide what matters: each layer lets the newest token check every earlier one before the experts do their sums.

why this paper matters Transformer 2017 →

≈74 billion operations per token

Each new token looks back at every earlier one, through a cache of what was already worked out.

5 of 7

Why is memory the expensive part?

Most of the energy goes into fetching the model’s numbers from memory, not into multiplying them.

The calculators are fast; the bookshelf is slow. For every word, the chip has to read the part of the recipe it needs, 37 GB, from the memory stacked beside it. That read alone takes about 11 milliseconds, however fast the sums are, so the calculators spend most of their time waiting.

why this paper matters KV cache 2019 →

≈0.38 J to write · ≈0.08 J to read

The reply comes one token at a time, and each one needs the model’s numbers fetched again.

6 of 7

How does it pick the next word?

The last layer scores every possible next token, and one is picked and turned into text.

129,280 scores

The picked token is turned back into text and sent to you, and the loop runs again.

7 of 7

Why does the reply come one word at a time?

Then all of it runs again for the next token, one at a time, until the reply is done.

Like a writer who can only see what is already on the page: each new word is added to the prompt and the whole journey runs again. A 500-token reply is 500 trips through all 61 layers.

≈20 tokens/s · 500 tokens ≈ 189 J

The picked token is turned back into text and sent to you, and the loop runs again.