Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Visualization · 25 Sept 2026

How electricity becomes intelligence

An AI data centre is a mill that turns electricity into words. Follow the power from the grid to the chip and back out as heat, and see why each part of the mill, in turn, became the thing everyone was short of, until the shortage reached the power grid itself.

trying to answer →What becomes scarce when intelligence becomes cheap?Where does value move when intelligence becomes commoditized?

pick a question; the machine answers it

SITE PLAN · ONE DATA HALLpower enters left · heat leaves right · scale approximateReported: LBNL interconnection queue studies show multi-year waits with high withdrawal rates.INTERCONNECTION: MULTI-YEAR QUEUERSUBSTATIONInference: campus scale for a build of this class. This is site load across several halls, not the load of the one hall drawn here — that one needs about 59 MW at the meter. The gap is more halls, not conversion loss.≈100 MW CAMPUS · ONE HALL DRAWNIReported: large power transformer lead times are widely quoted in years, not months.TRANSFORMER LEAD TIME: YEARSREstablished: distribution inside a site runs at medium voltage before final step-down.MEDIUM VOLTAGEESWITCHGEAR · UPSSTANDBY GENERATIONEstablished: standby plant covers outages. It does not raise the capacity the utility has granted.STANDBY, NOT CAPACITYEDATA HALL416 racks at roughly 118 kW. A hall of this footprint drew a few megawatts when it was built for 10 kW racks; the building did not change, the thing inside it did.≈49 MW ITIR1R2R3R4R5The drawing shows one representative bay. Racks of this class are reported near 120 kW; what decides how many you run is the power the utility granted, not the budget.416 RACKS × ≈118 kWISPINECDU ROW — FACILITY WATER TO RACK WATERHOTRETURNHEAT REJECTIONEstablished: energy in equals energy out. Every watt of compute is a watt this plant must reject.HEAT OUT = IT LOADEDETAIL BDETAIL B · ONE RACK, FRONT ELEVATIONthe unit of purchase · each tray carries two of the board on sheet 1MANIFOLDInference: two accelerator boards per tray across nine trays, a common dense arrangement.18 BOARDSI72 accelerators at roughly 1200 W each, plus hosts, switches, fans and pumps, divided by what the delivery path loses.≈118 kWIEstablished: conventional enterprise racks sat in the 5-15 kW band.vs ≈10 kW a decade agoECDUONE TRAY = SHEET 1, TWICEBUS BAR — POWER ARRIVES HERECOOLANTRACK LEAD TIME: QUARTERS. CONNECTION LEAD TIME: YEARS.

Why can’t the richest companies on earth buy enough electricity?

A mill. Water turns the wheel and the wheel grinds flour. Here electricity flips billions of tiny switches, the switches do sums, and the sums pick the next word of a reply. Every watt goes in as power and comes out as heat. On DeepSeek’s published numbers, one kilowatt-hour, a kettle boiling for half an hour, writes roughly 4 million tokens, building and cooling included. This is the story of that mill, part by part.

$3,800

the electricity that $30,000 chip burns in five years. Power is the cheap part.

It is also the part nobody can get. This is how a problem inside a chip the size of a postage stamp walked out of the chip, out of the building and into the power grid, and how every step of that walk made somebody rich.

1 of 7

Why is a graphics chip running AI?

A graphics chip was built to draw video games: millions of pixels, each needing the same small sum at the same moment, so it has thousands of slow calculators instead of one fast one. In 2012 a small team trained an image-recognition network on two gaming cards and beat everything. A neural network is the same kind of work.

On record: AlexNet (Krizhevsky, Sutskever and Hinton, 2012) was trained on two NVIDIA GTX 580 graphics cards.

Chips used to get faster for free: each generation fitted more transistors into the same power. That stopped in the mid-2000s. Since then every jump in speed has been paid for in electricity. Then ChatGPT arrived, in November 2022, and everyone wanted the jump at once.

The chip maker, for as long as everyone’s software is written for its chips.

NVIDIA rises 24% in a day, adding about $184bn, on guidance of $11bn for the quarter.

Next: Why does a very expensive chip spend its time waiting?

2 of 7

Why does a very expensive chip spend its time waiting?

Picture a chef who chops faster than anyone can carry ingredients to the counter. The chip got faster at arithmetic than memory could feed it numbers. Every word an AI writes means reading the whole model out of memory again, to do a little arithmetic with it. So the chip mostly waits.

So memory moved closer: stacked into towers, stood millimetres from the chip, joined by thousands of very short wires instead of a few long ones. More than ten times the bandwidth. It had been on sale since 2015, and almost nobody had wanted it.

The same problem was attacked in the code, by papers that each found a way to make the chip wait less. Why each one mattered →

Three companies in the world make this memory, and at first only one of them could make it well enough.

TrendForce reports NVIDIA's HBM3 came at first from SK Hynix alone.

Next: Why is an AI chip really a sandwich?

3 of 7

Why is an AI chip really a sandwich?

A single chip can only be so big: the machines that print them expose one fixed-size field at a time. So the industry started gluing several chips and their memory towers onto one slab of silicon wiring and calling the sandwich a chip. That gluing step used to be the cheap afterthought at the end of the line.

52–78 weeks

the wait for that gluing step, with lines sold out into 2027

The chipmaker that owns most of it, and behind it an insulating film made by Ajinomoto, a company better known for seasoning.

TSMC's chairman says the shortage is of its CoWoS packaging capacity, not of AI chips, and will last about 18 months.

Next: Why can’t air cool it any more?

4 of 7

Why can’t air cool it any more?

A 300-watt chip with a fan on it was solved for twenty years. Now it is 1,200 watts through the same few square centimetres, and the heat has to climb out through a stack of layers, pulled apart here so you can see them. Each layer resists a little. Four times the heat makes every little resistance four times worse.

Air cannot carry that away, so water came back, as it once did in mainframes: a cold plate on every chip, pipes to every rack. It brings plumbing, leaks, and a building that has to have water in the first place.

The makers of cold plates and cooling units, and whoever owns a building that already has water.

NVIDIA announces GB200 NVL72, a rack that ships liquid-cooled.

Next: Why did the whole cabinet become the computer?

5 of 7

Why did the whole cabinet become the computer?

For forty years you bought computers one server at a time, because one box was big enough to hold a useful machine.

Now seventy-two chips have to act as one, talking all the time, and that only works if they sit centimetres apart. So the thing you buy is a whole cabinet, drawing over 100 kilowatts where one used to draw about 10.

The companies that build these cabinets, but only while what they sell is scarce. Selling more than ever is not the same thing.

Supermicro falls 33% in a day after its auditor resigns.

Next: Why are data centres now built around electricity?

6 of 7

Why are data centres now built around electricity?

At 10 kilowatts a cabinet, the power a building lost converting electricity was a rounding error. At 120, the same few percent is a room of heavy copper and transformers the building was never drawn with.

So a new data centre starts from one number, how much power can I get, and everything else is drawn around it. Two to four years from land to first power, at roughly ten million dollars a megawatt.

Whoever already holds land with power, permits and water. Any one is easy to find; all four at once is rare.

Next: Why can’t you just buy more power?

7 of 7

Why can’t you just buy more power?

When a big data centre needed 30 megawatts, the utility had it spare. Nobody queued.

4–7 years

to connect a big new load to the grid. A new AI model comes out every few months.

Utilities, power plants, and whoever already holds the right to connect.

Constellation rises 22% on a 20-year, 835 MW deal to restart a Three Mile Island reactor for Microsoft.

Next: So where does it go next?

So where does it go next?

Nothing here was ever fixed. It only moved: chip, memory, package, heat, cabinet, building, grid. Each fix made the next thing scarce, and whoever held the scarce thing got paid. The next constraint is already outside the fence: turbines, transmission lines, and a public argument about who the power is for.

DeepSeek: NVIDIA falls 17%, $589bn, the largest one-day loss in US market history. Vistra falls 28%; Constellation and GE Vernova more than 20%.