Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 21 Jul 2026

What Becomes Scarce After Intelligence?

Nuclear reactors and downloadable models look like opposite strategies. They’re two sides of one wager on a question nobody will say out loud: does intelligence have a ceiling?

First published on Substack, 21 Jul 2026.

Two things happened this year that look like they belong to different industries.

In the first, the most valuable companies on earth became nuclear utilities. Microsoft restarted Three Mile Island. Amazon locked in nuclear power through 2042. All told, hyperscalers committed around 9.8 gigawatts of nuclear to AI in about a year, and sovereign wealth funds poured roughly $120 billion into the buildout. The bet underneath all of it: whoever delivers the most power wins. Build the biggest machine. Out-electrify everyone.

In the second, over two weeks this July, the floor fell out of the model business. Five frontier-adjacent open models shipped in a single window — Moonshot’s Kimi K3 at 2.8 trillion parameters, an American entry under a fully open license, DeepSeek’s V4, GLM-5.2, MiniMax’s M3. Kimi K3 hit number one on a coding leaderboard within hours, past a closed frontier model, at a fifth of the price, with weights you can download and run yourself. The gap between the best open model and the best closed one is now measured in months, not years. The bet underneath this one: intelligence is about to be free, and the value drains out of the model and onto whatever hardware you happen to own.

These look like opposite strategies. One spends hundreds of billions making intelligence bigger. The other gives it away. They are not opposite. They are two sides of a single wager, and almost nobody placing it will name the question they’re betting on.

The question is: does intelligence have a ceiling?

The bet nobody names

Here is what that one question decides.

If the demand for intelligence caps if, for most of what people actually want, some model is eventually “good enough” and getting smarter stops mattering then the efficiency camp wins everything. Good-enough intelligence commoditizes, races toward free, and runs on a box you own. The value stops living in the model and moves to whoever delivers it cheapest per watt. And the half-trillion-dollar nuclear buildout becomes the most expensive stranded asset in history: power plants built for a demand that plateaued.

If demand is uncapped if every gain in capability just unlocks a new appetite, the way cheap steel never sated the hunger for steel but multiplied it then brute force runs for a decade. The frontier keeps pulling away, the last increment of intelligence is always worth paying for, and the downloadable models are a footnote chasing a line that never stops moving. The reactor-builders win, and the efficiency camp spent its genius optimizing a commodity nobody cared to own.

Same coin. Everyone in AI has bet their capital on one side of it. And the tell that this is a real bet, not a rhetorical one, is that the smartest money is loudly and expensively betting both sides at once which is not conviction. It’s a hedge against a question the industry hasn’t admitted it’s asking.

So which way does the coin land? We have exactly one prior. It ran for four billion years.

The one time this experiment was run

Biology hit a hard intelligence ceiling, and we know precisely what happened underneath it.

The brain runs on about 20 watts, a fixed cap, set by what blood can deliver and what a skull can shed. And under that ceiling, evolution did not build a bigger, hotter brain. It couldn’t: brute force was never on the menu, because a brain runs on foraged food with no wall to plug into, and there’s no mutation-sized step from a chemical membrane to a digital logic gate. So it was forced into architecture and every trick it found is one the AI industry is now scrambling to copy. It fires only a few percent of its neurons at once, because a single spike is so costly that firing more would burn the brain’s entire energy budget sparsity as a power bill. It fuses memory and computation in the same place, so it never pays to shuttle data across a gap. It spends most of its energy about twenty-seven to one over computation not on thinking, but on moving information down the wire.

The verdict of the only run we have is unambiguous: under a fixed ceiling, architecture beats brute force. So if intelligence has a ceiling, biology already named the winner. The efficiency camp is betting with four billion years of precedent behind it.

And here’s the part that turns dusty precedent into this week’s headline.

The brain’s tricks are shipping, with model names

Why did five near-frontier models suddenly become things you can run at home? Because they are built on the brain’s first two tricks.

They’re sparse. Kimi K3 carries 2.8 trillion parameters but activates only about 50 billion per token a “mixture of experts,” hundreds of specialists of which a handful speak per word. That is the brain’s sparsity rendered in silicon: don’t light up the whole network, light up the sliver you need. It’s the entire reason a trillion-scale model can run on something that isn’t a data center.

And they win on locality. The reason these models run on a Mac and starve a traditional GPU rig isn’t compute it’s memory. Inference is bound by how fast you can move the active weights to the processor, not by arithmetic; the bottleneck is bytes, not math. The machines that run them well are the ones with memory fused close to the compute unified memory, shortening the wire. Which is the brain’s second trick, exactly: keep memory and computation in the same place.

Neuron, mixture-of-experts-on-a-laptop, in-memory chip the same move at three scales: collapse the distance between where the data sits and where it’s used. The efficiency camp isn’t praying the architecture shows up. It shipped this month, with version numbers.

The tell in the venture money

And the money that has no near-term reason to move is moving the same way. Between early 2024 and 2026, over $5 billion went into non-GPU, brain-shaped computing hardware against maybe $50 million of revenue a 110-to-1 bet on architecture over brute force, years before it can pay. The largest seed in the category’s history, $475 million, went to a founder whose whole pitch is the ceiling itself: we cannot produce enough energy to keep scaling, so the only way forward is to compute more per joule. That’s not a trade chasing this quarter. It’s nine figures wagered on the belief that the ceiling is real.

Where the bull case breaks

Now the honest part because the version of this argument that got reach this year (”intelligence goes free, the model layer dies, the giants are dead men walking”) bets the whole thesis on an assumption it never prices.

Look again at what the open models actually did this month. They took the middle, not the top. Every practitioner guide converges on the same routing rule: send the cheap, high-volume work to open weights, and keep the hardest, highest-stakes, most ambiguous work on the closed frontier. The gap closed to months on ordinary tasks and stayed double digits on the hardest ones. So intelligence didn’t commoditize. The legible part of it did the well-posed, gradeable, common work while the illegible frontier kept its margin.

Which means the coin isn’t heads-or-tails. The ceiling isn’t a single wall; it’s a line, and the real question is where it sits. Below the line, intelligence is already free and running on your own silicon. Above it, the frontier keeps its rents. The entire trillion-dollar fight is over where that line settles and both camps are pretending it’s a binary they’ve already won.

The second ceiling

Energy was never the only wall, either. The brute-force bet needs three inputs compute, energy, and data and the third is draining in step with the first two. The stock of high-quality human text runs out between now and the early 2030s, and most new web text is already machine-generated, so each fresh scrape drinks its own exhaust. Feed a model its own output in a closed loop and it collapses. The energy ceiling is the buyable one enough capital annexes another gigawatt. The data ceiling is stubborner: you cannot mint fresh, uncontaminated human experience with money. Two ceilings, closing at once and biology cleared both with the same move, doing more per unit of the scarce thing, whether the scarce thing was a joule or a token.

What to actually watch

So don’t ask who’s right. Ask which uncertainty you’d rather own.

The efficiency camp is the convex bet: it wins outright if intelligence is capped, and it wins eventually even if it isn’t because runaway demand still meets a ceiling somewhere, and when it does, the efficiency gap is the only prize left to take. The brute-force camp only wins in the uncapped world, and only until it hits the wall biology hit. It also carries its own fat tail: it needs the demand to actually arrive. A reactor switched on in 2032 is a decade-long wager that the load shows up strand that demand, and the power plant is the thing that blows up.

There’s a single number that reports the score in real time: whether the frontier holds its margin or the open middle keeps climbing into it. Watch where the line sits, and which way it’s moving. That’s the whole game and everyone has already bet on it, most of them without ever saying so.

Evolution ran this experiment once and filed the result. Half the industry is betting it holds. Half is betting this time is different. They just never told you that was the bet.