Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 31 May 2026

Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew Why

Big Tech turned AI usage into a metric and a 124-year-old bounty scheme in colonial Hanoi explains exactly what happened next.

First published on Substack, 31 May 2026.

The Tailless Rats of Hanoi

In 1902 the French administration in Hanoi had a rat problem, and what looked like a tidy solution. Pay a bounty for every rat killed. To save officials from handling thousands of carcasses, just collect the tails as proof. The bounties went out. Tails came in by the sackful. And then people started spotting rats around the city with no tails.

The locals had read the incentive more carefully than the people who wrote it. The government wasn’t paying for dead rats. It was paying for tails. So you caught a rat, cut off the tail, and released it to go make more rats with tails. A few enterprising types skipped the sewers and just bred rats on the edge of town.

I’ve been thinking about those tailless rats, reading the news out of Uber, Amazon, and Microsoft.

The metric ate the strategy

Here’s what happened in 2026, with the press releases removed. The big tech companies decided AI adoption was the future, which is fine. But “are we an AI-first company” is almost impossible to measure, so they grabbed something they could count instead. How many tokens are people burning. What share of engineers touch the tool each week. Then they put it on leaderboards and wired it into performance reviews.

You already know the rest, and so did the engineers. Amazon staff started calling it tokenmaxxing: pointing autonomous agents at make-work just to climb the internal board. Reddit threads and internal forums started filling up with advice on how to maximize token usage without setting off alarms: point autonomous agents at make-work, generate endless reports, route trivial tasks through the most expensive models. Uber ranked its engineers on a leaderboard by how much they used the tool, then blew its entire full-year AI budget in four months. When the company’s COO went looking for what all that spending had bought, he couldn’t find it. The link between the tokens and the value, he admitted, “is not there yet.”

The tails were coming in by the sackful. The rats were fine.

A measure becomes a target

The economist Charles Goodhart gets his name on this, though the clean version came later: once a measure becomes a target, it stops being a good measure. The idea underneath is almost too simple to respect. A metric is a stand-in. It points at something you care about but can’t watch directly. Reward people for the stand-in and you’ve told them, without saying it, that the stand-in is the actual job. People are extraordinarily good at doing the actual job.

What makes this hard to catch is that it never feels like a blunder while you’re committing it. It feels like rigor.

We are still making one big nail | Global Development Network

We have done it constantly. The Soviets ran nail factories on output quotas and got mountains of tiny useless nails; they switched to weight and got a few comically enormous ones. During Vietnam, Robert McNamara, the original numbers man, made enemy body count the measure of progress, and his commanders duly inflated kills and filed civilians as combatants while the war went sideways underneath the rising graph. Before 2008, banks measured danger with Value-at-Risk, which told you the worst you’d lose on 99 days out of 100; traders built books that were spotless on those 99 days and ruinous on the hundredth, the one the model couldn’t see.

My favorite is the NHS, because it’s so physical you can picture it. Hospitals were told no patient could wait more than four hours on a trolley before admission. Administrators unscrewed the wheels and reclassified the trolleys as beds. The clock stopped. The target turned green. The patient was still in the corridor, now lying on a piece of furniture that had been promoted.

Every one of these was run by clever people. That’s the part I’d ask you to sit with for a second.

Why the clever people miss it

The lazy read is that the executives were fools, and it’s wrong. These are among the best-credentialed problem-solvers on the planet. The failure isn’t IQ. It’s a specific trap, and it has a name.

Psychologists call it surrogation. Asked to judge something genuinely hard is the strategy working? the mind quietly swaps in an easier question it can actually answer: is the number going up? You don’t feel the swap happen. That’s what makes it dangerous. The dashboard stops being a window onto the strategy and becomes the strategy. The map eats the territory.

The historian Jerry Muller put a sharper edge on this in The Tyranny of Metrics. His claim is that organizations don’t reach for metrics despite their flaws but because of one specific feature: a number lets leadership skip the work of judgment. It looks objective. It fits in a board deck. It spares you from walking down to where the work happens and forming an opinion about whether it’s any good. Measurement quietly becomes a replacement for understanding.

And then consulting industrializes the whole thing. In late 2023 almost too perfectly McKinsey published a piece called “Yes, you can measure software developer productivity,” and a good chunk of the engineering world, Kent Beck included, set it on fire. The complaint was clean: the proposed metrics counted activity, not value. Lines of code, deployments, story points. Pay an engineer for lines of code and you get more lines of code. You get the exact bloat you were trying to kill. That advice didn’t stay in the PDF. Two years later it was the unspoken philosophy behind every mandate to push AI usage up and to the right, with nobody checking whether the usage made anything.

Why this one is actually worse

I don’t want to land on the comfortable note same old mistake, nothing new under the sun because mechanically there is something new here, and it’s the part that should bother you.

Each of those needed human labor to game. The factory still had to forge the giant nail. McNamara’s officers still had to go run the pointless mission. Someone still had to walk over and unscrew the trolley wheels. Faking output took work, and work is a speed limit. There’s only so fast a person can waste a budget by hand.

The agent has no speed limit.

It doesn’t run one query and stop. It loops. Queries a database, reads what came back, critiques itself, calls another tool, patches its own errors, goes again. Aim one at a meaningless task to pad your usage stats and you’ve kicked off an automated loop that can spend thousands of dollars while you’re getting coffee. Output tokens cost a multiple of input tokens, so every needless report it spits out is more expensive coming out than going in. For the employee, the cost of faking productivity has fallen to nothing. For the company it’s gone vertical.

There’s a tempo here the older cases never had. Each of those pathologies ripened slowly Soviet industry rotted over generations, the leverage that broke the banks built across years and even then the reckoning arrived only at human pace. What cost the Soviets generations, Uber ran in four months. Take away the human speed limit and you don’t just waste money faster; you watch a metric’s whole pathology run in fast-forward, months where it used to take careers.

A telemetry firm called Entelligence put a number on the damage that I haven’t been able to shake. Across thousands of companies, they reckoned that of every dollar spent on AI tokens, about eighteen cents turned into something a user could actually use. The rest went to fixing the bugs the AI wrote, reworking its code, and clearing the review pileup it caused. We built the most efficient waste machine ever made and then gave it a leaderboard.

There’s a stranger layer underneath all this. When gaming a metric was expensive, the cost at least stayed inside the company that set the metric the giant nail sat on the factory floor, useless, but nobody outside got rich off it. Token waste doesn’t sit on a floor. It leaves the building as a payment. Every pointless agent loop, every make-work report run to nudge a usage number upward, is somebody’s revenue. Which puts the AI vendors closer to Hanoi’s rat farmers than to the government that set the bounty: not the party defrauded by the system, but the party that quietly understood the system paid for tails and built a business on supplying them. The buyer calls it productivity. The seller books it as growth. The fastest-growing revenue line in the industry may turn out to be, in part, a lot of companies paying to learn the wrong lesson from their own dashboards.

What to do, and what it’s really about

The fix isn’t a cheaper model or a faster chip. It’s the thing all of these stories have been waving at for a hundred years: pay for the dead rat, not the tail. Measure the outcome instead of the exhaust cost per resolved ticket rather than tokens spent, product shipped rather than pull requests opened. Route the trivial work to cheap models automatically, so the call isn’t left to a nervous engineer reaching for the most powerful thing on the shelf in case someone’s watching. And, as Muller would say, let people use judgment again, since that’s the thing the number was always a bad copy of.

But the operational fixes aren’t the lesson that stays with me. The lesson is that the urge to swap judgment for a number is bottomless, and it always arrives wearing the costume of discipline. The Soviet planner thought tonnage was discipline. McNamara thought the enemy body count was discipline. The bank thought VaR was discipline. The executive watching the AI usage dial climb thinks the dial is discipline.

It never is. A metric is a finger pointing at something worth looking at, and most of management is the long history of people admiring the finger.