Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 5 Jun 2026

The Speed of Being Wrong

Whether AI levels you up or quietly hollows you out comes down to a single variable almost no one is naming. It is not your skill.

First published on Substack, 5 Jun 2026.

Two findings, published a year apart, that you have almost certainly seen quoted but never in the same room.

The first: when researchers gave GitHub Copilot to nearly five thousand developers across Microsoft, Accenture, and a Fortune 100 firm, junior engineers gained 21 to 40% in output while seniors gained 7 to 16%. AI as the great leveler. This finding gets quoted to argue that AI compresses talent and democratizes skill. (Cui, Demirer, Jaffe et al., Management Science, 2026)

The second: when a different team handed a GPT-4 business mentor to a few hundred entrepreneurs in Kenya, high performers improved 15 to 25% and low performers got worse roughly 8 to 10% worse than if they’d had no AI at all. AI as the great divider. This finding gets quoted to argue the opposite: that AI rewards the already-capable and punishes the rest. (Otis, Clarke, Delecourt, Holtz, Koning, Harvard Business School Working Paper 24-042)

Both are real. Both are carefully run. They point in opposite directions.

So either AI’s effect on skill is random which no one believes or there is a hidden variable doing the work, and the variable is not what everyone is staring at. It is not skill level. Skill is the outcome the two studies disagree about, not the cause of the disagreement.

The cause is in the environments. And once you locate it, it explains not just these two findings but most of what has been confusing about AI and labor for the past three years.


The developers had a compiler. The entrepreneurs did not.

Write bad code and the machine tells you in seconds red text, failed build, no argument. Make a bad strategic decision and the market tells you in months, tangled up with variables you never controlled. One environment corrects error quickly and honestly. The other corrects it slowly, noisily, or not at all.

That gap does most of the explaining. Call it the speed of being wrong.

When the loop is fast and honest, AI is a leveler. A junior who accepts a flawed suggestion gets caught before the mistake compounds, so the floor of competence rises toward the ceiling. The referee on the field makes the weaker player nearly as safe as the stronger one. When the loop is slow or absent, AI is an amplifier. No referee. The strong performer uses the tool to sharpen judgments she can already check against her own experience. The weak performer asks it for things he cannot evaluate, cannot fully execute, and will not discover are wrong until it is far too late. The Kenyan entrepreneurs who did worse weren’t given worse advice. They were given confident advice in an environment that wouldn’t tell them, in time, that they’d misused it.

The variable, stated as a test you can run on any task in about ten seconds: How fast, and how honestly, does this work tell you that you got it wrong?

Fast and honest, AI commoditizes the work and flattens the people doing it. Slow and noisy, AI widens the gap between those who already know and those who only sound like they do.


The reason this is worth more than another “AI helps juniors” headline is that it cuts across the categories we normally use, splitting apart jobs that intuition files together.

Take two people who both “pick financial winners.” A day trader and a venture capitalist. Intuition groups them. Run the test and they fly apart. The day trader finds out he was wrong within minutes, by a P&L that does not care about his thesis. Fast, honest loop. That is why day trading was substantially replaced by algorithms long before ChatGPT arrived, and why a human trader with AI is being pushed toward a machine-defined par where the edge evaporates. The venture capitalist finds out whether a seed check was smart in seven to ten years, by which point the signal is hopelessly confounded with luck, follow-on rounds, and the macro cycle. AI can sharpen a great partner’s sourcing and memo-drafting but cannot rescue a mediocre one and — critically — cannot be checked against outcomes on any timeline that matters. The partner’s taste stays a premium because the world refuses to grade it quickly.

Same surface activity. Opposite loop structure. Opposite AI fate. The test predicts it; “investing” as a category predicts nothing.

It cuts inside a single profession too. A radiologist hunting a nodule works in a relatively fast loop - scan, biopsy, pathology, answer. A psychiatrist titrating treatment over months works in a slow, confounded one. The framework predicts that AI would be strong and leveling in the first and carry real hazards in the second, which is roughly what the deployment evidence suggests — though the picture in both fields is still developing. The point isn’t the specific prognosis for any specialty. The point is that “medicine” as a unit tells you almost nothing, while “how quickly and honestly does this subspecialty grade its errors” tells you quite a lot.


Here is where the standard career advice breaks down.

Fast feedback feels like the safe place to be. Errors get caught, AI obviously helps, the work goes smoothly. Nearly all of the “learn these tools or get left behind” advice points you toward exactly these tasks, because that’s where AI’s value is clearest and most immediate.

That visible value is the kiss of death. A fast, honest feedback loop is also, structurally, the same thing as legibility. And legibility is the precondition for commoditization. To measure a task precisely enough that the loop runs fast is to measure it precisely enough to route to something cheaper — a smaller model, a junior in a lower-cost market, a script. Enterprise AI infrastructure is already built around exactly this triage: semantic routers that send legible, fast-feedback queries to cheap open-weight models while reserving expensive frontier reasoning for the rare genuinely hard problem. The same classification is coming for people.

The durable position is not the task where AI helps you most. It is the task where AI is too dangerous to deploy unsupervised, where the loop is slow, the errors are expensive, and someone has to own an irreversible call and carry the consequence. The places AI feels risky are the places human judgment holds a premium. The places it feels magical are the places being priced toward zero.

Standard advice is pointing people toward the kill zone and calling it an opportunity.


There is another layer, and this is the one I find genuinely unsettling.

AI’s value depends on a fast feedback loop. And there is real evidence that AI’s proliferation slows feedback loops down.

The clearest example comes from medicine. A 2025 study in The Lancet Gastroenterology & Hepatology tracked endoscopists doctors hunting for polyps during colonoscopies before and after they began using an AI detection tool. With the AI active, they caught more. But when those same doctors later worked without it, their unassisted detection rate had fallen from 28% to 22%. This is not the straightforward deskilling story the feedback loop still ran, the scope and biopsy and pathology all happened. But the loop now closed through the machine instead of through the human. The physician’s own internal error-correction stopped getting exercised and atrophied. Whether this dynamic generalizes to other fast-loop professions we don’t yet fully know. But the mechanism is not obviously medical: offload the correction to a tool reliably enough, and the human stops running the circuit that made them competent.

At the organizational level, the pressure runs a different way. Microsoft’s own randomized study of over 6,000 Copilot users across 56 firms found workers completing documents faster and spending less time on email and initiating 11% more documents. Every one of those needs to be read and verified by someone downstream. The concern and it remains a concern rather than a settled finding, because the systemic evidence isn’t yet in is that verification burden rises faster than generation savings, lengthening the effective feedback loop for the organization even as individual output metrics look healthy. If that pattern holds at scale, AI would be quietly converting fast-loop environments into slow-loop ones, shifting the whole map over time.

AI degrades the precondition for its own value. Whether this shows up clearly in the macro data over the next decade is genuinely uncertain. But the mechanism is worth watching precisely because it’s invisible in the short run.

The labor math has a related trap, and it’s the one firms are most likely to spring on themselves.

When the legible, fast-feedback tasks get handled by AI, the rational firm response is to thin the junior tier. The Burning Glass Institute documented a 37% drop in entry-level postings in New York alone between 2022 and 2024. Friebel, Huang, Li, Shukla, and Zhang modeled the same shift this spring: the pyramid becomes a diamond as firms freeze junior hiring, and under a learning shock the diamond may be permanent. On a spreadsheet, this looks like discipline.

What disappears from the spreadsheet is that the junior tier was the apprenticeship. Not because of the output it produced, but because of the loop it ran through the person producing it. The rote discovery, the doomed first drafts, the code that didn’t compile until it did this was the mechanism by which tacit judgment got built, and it only works by doing the thing badly until you stop. Automate it for immediate efficiency and you stop manufacturing the judgment that slow-loop, AI-resistant work requires. The developer communities documenting this as a “tragedy of the commons” have it right: individual productivity metrics rise while the shared stock of deep competence that reviews and maintains the AI output is not being replenished.

Here is the asymmetry that the spreadsheet can’t see. Routing tokens from an expensive model to a cheap one is reversible the frontier capability still exists, you’ve stopped overpaying. When Klarna discovered its AI-first customer-service bet had quietly degraded quality, it could hire humans back. The capability had been dormant, not destroyed. But the apprenticeship layer that produces senior judgment doesn’t sit dormant when you automate it away. The seniors who would have been created simply don’t exist.

You can route tokens back. You cannot route experience back.

A hundred years ago, Frederick Taylor installed the first formal feedback loop on human labor: the stopwatch. Measure the worker, correct the worker, optimize the output. He told the founding story of the whole method a pig-iron handler whose daily output he claimed to have nearly quadrupled through precise measurement. The number became scripture. Subsequent historians Wrege and Perroni (1974) and others found the worker’s real name was Henry Noll and the famous figures were, in the load-bearing details, fabricated.

What matters here is not that Taylor lied, though he did. It is what a feedback loop with a corrupted signal actually produces. It runs. It corrects. It optimizes. And everything it optimizes is pointed at a number that doesn’t reflect reality. A corrupted loop is not just ineffective it is more dangerous than no loop, because it manufactures confident, measured, authoritative wrongness and calls it evidence.

We are now installing AI-driven measurement across nearly every kind of knowledge work: git commits, email response rates, documents completed, issues resolved per hour. Each of these is a feedback loop of a kind. Some are honest. Many are fast but shallow proxies that capture the legible surface of work while the slow, judgment-intensive, hard-to-measure parts go unrecorded and eventually get defunded.

The question worth asking about any feedback loop you’re operating inside, whether you installed it or inherited it, is not whether it is fast. It is whether it is honest. Whether the thing it is correcting you toward is the thing that actually matters, or whether it is, like Taylor’s stopwatch, a clean measurement of the wrong number.

That question does not change when the next model drops. Capability is what changes with model releases. Loop structure changes with how work is organized, which moves more slowly and is mostly under human control. Which is either reassuring or alarming, depending on who in your organization is currently deciding what gets measured.