Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Why this paper matters · 3 of 12

Why do some AI models have experts?

Mixture of experts: What you can afford to hold and what you can afford to run stopped being the same number.

the paper →Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityFedus, Zoph and Shazeer, 2021 ↗

This is why a model that behaves like it has hundreds of billions of parameters can still answer as fast and as cheap as one a tenth the size — most of it sits idle on any given word. Only a handful of experts fire per token, so it runs like a small model and has to be held in memory like a huge one.

The room it came out of

The idea dates to 1991 and sat unused for thirty years, because nobody had a reason to want a model with far more parameters than it runs. Google revived it in 2021 for a purely economic reason: the scaling laws wanted parameters and the budget refused to pay for the arithmetic, and this is the one architecture where those two can be separated.

How it works

A dense model puts every token through every parameter. A mixture-of-experts model does not: each layer holds many parallel sub-networks, a small router picks two or three of them per token, and the rest sit idle. A model with six hundred billion parameters might do the arithmetic of a thirty-billion one.

That sounds like a straightforward win and it is not, because the idle experts are still in memory. The router cannot know in advance which expert the next token will want, so all of them have to be resident and ready. Arithmetic scales with the experts you use; memory scales with the experts you have.

Which is why a mixture-of-experts model has two sizes and press releases quote the flattering one. "Active parameters" is what it costs to run a token. "Total parameters" is what it costs to hold the model at all, and it is the number that decides how many machines you need before you serve anybody. A reader who sees one figure for an MoE model is being shown whichever suits the argument.

There is a third cost that only appears at scale. If the experts are spread across machines — and at frontier size they must be — then every token's chosen experts have to be fetched across the network, every layer, in a pattern nobody can predict. That is the all-to-all traffic on the Networking layer, and it is a large part of why the fabric inside a rack became a product people argue about.

So the honest framing is not that sparsity made models cheaper. It is that sparsity moved the cost off arithmetic and onto memory and interconnect, which is the same shape as everything else in this atlas, and it happened to move it toward the two resources that were already the tightest.

What it traded

gave up
memory — every expert must be resident whether or not it runs
got
arithmetic — only a small fraction of the model touches any given token

What exists now that didn’t before

Every expert has to be resident whether or not it runs, so the saving in arithmetic arrived as a bill in capacity and interconnect.

Asked, and answered

How can a huge model answer as fast as a small one?

Only a handful of its internal 'experts' actually fire for any given word — the model has to be held in memory like a giant one, but it computes like a small one, because most of it sits idle on every single token.

Read next