Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Why this paper matters · 7 of 12

How did DeepSeek train a frontier model so cheaply?

DeepSeek: A constraint you cannot buy your way past becomes a research agenda.

the paper →DeepSeek-V3 Technical ReportDeepSeek-AI, 2024 ↗

This is why a lab locked out of the best chips could still ship a model that rivalled OpenAI’s and Google’s, and why it rattled the assumption that more capital always wins — by refusing to treat memory as somebody else’s problem. Its latent attention is the most aggressive attack yet on how much state a single token has to carry.

The room it came out of

A Chinese lab under export controls in 2024, unable to buy its way out of a memory constraint at any price. So it attacked the constraint instead, and found a compression of attention state that had been sitting in plain sight for seven years while thousands of better-funded researchers read the same code. Nobody looks hard at a wall they can afford to pay somebody to move.

How it works

DeepSeek is worth a page here for a reason that has nothing to do with benchmark scores. It is the clearest case in the last few years of a lab that could not obtain compute on the terms its competitors could, and therefore had to treat a constraint as a research problem rather than a purchasing one.

Its most cited contribution is multi-head latent attention. Where grouped-query attention reduces how many key and value heads you store, latent attention changes what you store: the keys and values are compressed into a much smaller shared latent representation and reconstructed on read. It is the same idea as the rest of this thread — spend arithmetic to save state — pushed considerably further than anyone else had been willing to push it.

That move is in this atlas as an intervention, not as an anecdote. Switch it on in the model and state per token falls by about four times against grouped-query attention alone, which at long context is the difference between a machine that serves a hundred conversations and one that serves four hundred.

The strategic read is the interesting part, and it generalises past this one lab. Constraints do not slow research down evenly — they redirect it. A lab with abundant memory optimises what it is already doing; a lab without it goes looking for a different architecture. Several of the most-copied efficiency ideas of the last decade came from whoever had the least of the resource in question, and that is a pattern worth carrying into whatever the next scarce thing turns out to be.

The caution, stated plainly: reported training costs are not audited, the comparison figures circulated in the press mixed several different things together, and this atlas has no way to check any of them. What it can check is the architecture, which is published, and the architecture is genuinely different.

What it traded

gave up
architectural simplicity, and a great deal of engineering effort
got
state per token roughly an order of magnitude below the frontier standard

What exists now that didn’t before

A public demonstration that a hardware constraint can be attacked rather than paid around, which cost the industry its assumption that frontier capability follows capital in a straight line.

What it left undone

Six hundred billion parameters still have to sit somewhere, which moves the bill from bandwidth to capacity and to the fabric between machines.

Asked, and answered

How did deepseek train a frontier model so cheaply?

They found a way to compress how much state each token carries in memory — a technique that had been sitting in public, published research for seven years, just never pushed this hard, because nobody with easier options had needed to.

Read next