Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 9 Jun 2026

The Rate-Limiting Step

How AI keeps solving the wrong bottleneck

First published on Substack, 9 Jun 2026.

An engineer recently embedded 685 million pieces of text in about thirty-two minutes. Eight rented A100 GPUs, the spot market, total cost around seven dollars. The number that should interest you isn’t the price. It’s what he learned on the way there.

Once you go past a couple of GPUs, he found, the GPUs stop being the problem. A single A100 chews through a batch faster than the surrounding code can tokenize the next one and hand it over. So the expensive thing the scarce thing, the thing the entire industry is mortgaging its future to buy more of sits idle, waiting to be fed. Almost all his engineering went into the feeding. He had purchased compute. His bottleneck was logistics.

This is not a one-off war story. When Snowflake’s engineers profiled embedding workloads, they found the actual GPU inference accounted for roughly a tenth of total compute time. The rest went to the unglamorous business of moving data into position tokenizing, serializing, queuing. They rewrote the plumbing and got many times the throughput on the same silicon. The math was never the constraint.

Bottlenecks hide, and they move

Eliyahu Goldratt built a management philosophy on a single uncomfortable observation. In The Goal, his 1984 business novel, he argued that every system has exactly one binding constraint at any moment, and that nearly everything we do to improve the system is wasted unless it improves that one thing. His sharpest line: an hour saved at a non-bottleneck is not a saving at all. It’s a mirage.

Biology had the idea first. A metabolic pathway runs only as fast as its slowest enzyme the rate-limiting step. Flood the cell with every other input and the pathway will not speed up. Henry Ford hit the same wall: the constraint on his line was never welding speed but the staging of subassemblies arriving at the welder. Speed up the welder and you produce a larger pile of half-finished cars.

The pattern is always the same. We optimize the part of the system we can see the part that is expensive, measured, equipped with a vendor and a quarterly earnings call. The real constraint sits one step upstream, unmeasured, doing its quiet throttling.

In AI, the visible thing is the GPU. So for two years we have optimized the GPU.

Where the constraint actually sits

Watch a model generate a single token. To produce that one token, the hardware must read every weight in the model out of memory. For a seventy-billion-parameter model in half precision, that is roughly 140 gigabytes of data movement per token. On an H100, whose memory bandwidth tops out near 3.35 terabytes per second, that transfer alone costs about forty-two milliseconds. No additional arithmetic capacity lowers that number. The chip is not thinking too slowly. It is reading too slowly.

This is what the FLOPS-counters miss. At the generation step, inference is memory-bandwidth-bound, not compute-bound. The tensor cores sit mostly idle, starved, waiting on memory the silicon-scale version of the engineer’s GPU twiddling its thumbs at seven dollars an hour. And the gap is widening. Over the past decade, compute on AI chips grew something like eightyfold while memory bandwidth grew only seventeenfold. We have been pouring money into the half of the machine that was never the limit.

The most credible confirmation comes from the company least likely to volunteer it. Anthropic’s own engineers, having handed a growing share of their coding to Claude, discovered that as they pushed more code through the organization, human code review became the new bottleneck. They relieved one constraint and it relocated immediately downstream. The firm building the frontier is rediscovering Goldratt inside its own walls.

And once you start looking for bottleneck misidentification, you begin seeing the same pattern at larger and larger scales.

The same error, one altitude up

Now lift the camera from the data center to the State Department.

For several years, American policy toward Chinese AI rested on a single assumption: that the chokepoint was hardware. Deny the most advanced chips and you deny the frontier. A small yard, a high fence and the fence was built around the silicon. Then DeepSeek released a frontier-class model trained around the restrictions, leaning on algorithmic efficiency where it could not lean on raw hardware. The market lost roughly six hundred billion dollars of Nvidia’s value in a session. The fence had been built around the yard where the constraint used to live.

Here the honest version splits in two, and I won’t pretend otherwise. One reading, Anthropic’s, is that the fence works and China stays close only by routing around it smuggling chips, renting offshore compute, and distilling American models. Another, Sangeet Paul Choudary’s, is that China was never running the intelligence race at all, having bet instead that intelligence will become abundant and that the durable advantage lies in the physical systems that turn intelligence into output. Both can be partly true: route around the frontier where you can, and build the layer that wins even if you can’t. What neither reading supports is the comfortable belief that controlling the chips controls the outcome. The constraint already moved.

What software can and cannot compress

So where does the constraint go from here, and can we get ahead of it?

Here is the distinction that matters, and the one I have to be careful with. Software compresses some kinds of coordination almost effortlessly. The friction of moving money, once a constraint, dissolved into an API. The friction of provisioning a server collapsed into a cloud console. And the memory wall I just described the one supposedly limiting inference Google’s researchers recently published a compression technique called TurboQuant that shrinks the relevant memory footprint roughly sixfold with negligible accuracy loss. A constraint that looked physical turned out to be partly an algorithm waiting to be written.

But not all coordination yields like that. The friction of building a battery supply chain, of accumulating the production know-how to make ten million electric drivetrains, of synchronizing thousands of suppliers into a vehicle that rolls off a line that compresses slowly, because the learning is embodied in machines, relationships, and processes that cannot be rewritten in an afternoon. Digital coordination is soft. Physical-deployment coordination is sticky. The mistake is to treat them as one thing.

This is why I will not tell you that physical coordination is the final constraint, the place the migration stops. Electricity once looked foundational, then semiconductors, then software, then distribution, then data, then compute. Every generation believed it had located the permanent bottleneck, and every generation was eventually wrong. The most I will claim is that physical-deployment coordination currently appears to be the layer hardest to compress with either money or cleverness which is exactly why it is worth watching, and exactly why a power positioning there today is making an intelligent bet rather than a sure one.

The discipline

Goldratt’s framework never lets you rest. Relieve the binding constraint and a new one becomes binding; the system does not become unconstrained, only constrained somewhere else. The engineer who fixes the data feed will find the bottleneck has migrated to the network, then to power, then to something he hasn’t thought of yet.

So the durable question is never what is the bottleneck. It’s where has it moved to now and the discipline is to keep asking even when you think you’ve found the answer. The worst thing I could do is end this essay confident that I have finally located the real constraint. That confidence is the precise error I opened with. The fence keeps getting built around the wrong yard, not because the builders are foolish, but because the yard keeps moving and the building takes longer than the moving.

The engineer with the seven-dollar job understood this in miniature. He didn’t ask how to make the GPU faster. He asked what the GPU was waiting for. That inversion from optimizing the visible thing to interrogating the invisible one is the entire discipline. Everything else is just a larger pile of half-finished cars.