Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Explainer · 26 Sept 2026

Why this paper matters

Twelve papers that moved where the constraint in AI sits, from the Transformer in 2017 to the models of 2025. Each one opens with the question it answered.

trying to answer →What becomes scarce when intelligence becomes cheap?

A paper matters when it changes what is scarce. Each one below solved a real problem, and each solution moved the cost somewhere else, usually into memory. Read them in order and you watch the same few constraints get attacked again and again.

  1. 1Vaswani et al. · 2017Attention Is All You NeedWhy did transformers replace RNNs?Parallelism is what turns money into capability.3-minute brief · Transformer
  2. 2Shazeer · 2019Fast Transformer Decoding: One Write-Head is All You NeedWhy does a long conversation get slower and cost more than a short one?Previously computed state is worth more than recomputing it.explainer, runs in the page · lands on HBM
  3. 3Fedus, Zoph and Shazeer · 2021Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityWhy do some AI models have experts?What you can afford to hold and what you can afford to run stopped being the same number.3-minute brief · Mixture of experts
  4. 4Dao et al. · 2022FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessWhy is attention slow when the GPU is barely doing any arithmetic?Moving data can cost more than computing with it.explainer, runs in the page · lands on HBM
  5. 5Gu & Dao · 2023Mamba: Linear-Time Sequence Modeling with Selective State SpacesWhy does a model reread the whole conversation instead of just remembering it?A fixed memory can only keep what it chose to keep.explainer, runs in the page · lands on HBM
  6. 6Kwon et al. · 2023Efficient Memory Management for LLM Serving with PagedAttentionWhy does a GPU with free memory still refuse new requests?Reserved memory is spent memory.explainer, runs in the page · lands on HBM
  7. 7DeepSeek-AI · 2024DeepSeek-V3 Technical ReportHow did DeepSeek train a frontier model so cheaply?A constraint you cannot buy your way past becomes a research agenda.3-minute brief · DeepSeek
  8. 8Lieber et al., AI21 Labs · 2024Jamba: A Hybrid Transformer-Mamba Language ModelWhy does Jamba matter?AI21’s open-weight model, and the hybrid mamba.left predicted before it existed: one attention layer for every seven Mamba layers, with mixture-of-experts layered on top, holding 256k of context in a fraction of the memory a pure-attention model would need.a note · Jamba
  9. 9Ma et al., Microsoft Research · 2025BitNet b1.58 2B4T Technical ReportWhy does BitNet matter?Microsoft’s open-weight model, trained from scratch at ternary precision —a note · BitNet
  10. 10DeepSeek-AI · 2025DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningCan a small AI model learn to reason like a huge one?Reasoning ability can be trained against a checkable reward, not copied from a hand-written example of the reasoning itself.3-minute brief · DeepSeek-R1
  11. 11Kimi Team, Moonshot AI · 2025Kimi K2: Open Agentic IntelligenceWhy does Kimi K2 matter?Moonshot AI’s open-weight model —a note · Kimi K2
  12. 12Gemma Team, Google DeepMind · 2025Gemma 3 Technical ReportHow do models handle huge context windows without the memory bill exploding?Most words do not need to see the whole conversation to be predicted correctly.3-minute brief · Sliding-window attention

Every page opens with the question the paper answered, not the paper’s own title. Four are full explainers you can run, three levels deep, each pointing at the part of the machine it binds on in How electricity becomes intelligence. The rest are three-minute briefs, and a few are still only notes, marked as such.