The Speed of Being Wrong
Whether AI levels you up or quietly hollows you out comes down to a single variable almost no one is naming. It is not your skill.
Two findings, published a year apart, that you have almost certainly seen quoted but never in the same room.
The first: when researchers gave GitHub Copilot to nearly five thousand developers across Microsoft, Accenture, and a Fortune 100 firm, junior engineers gained 21 to 40% in output while seniors gained 7 to 16%. AI as the great leveler. This finding gets quoted to argue that AI compresses talent and democratizes skill. (Cui, Demirer, Jaffe et al., Management Science, 2026)
The second: when a different team handed a GPT-4 business mentor to a few hundred entrepreneurs in Kenya, high performers improved 15 to 25% and low performers got worse roughly 8 to 10% worse than if they’d had no AI at all. AI as the great divider. This finding gets quoted to argue the opposite: that AI rewards the already-capable and punishes the rest. (Otis, Clarke, Delecourt, Holtz, Koning, Harvard Business School Working Paper 24-042)
Both are real. Both are carefully run. They point in opposite directions.
So either AI’s effect on skill is random which no one believes or there is a hidden variable doing the work, and the variable is not what everyone is staring at. It is not skill level. Skill is the outcome the two studies disagree about, not the cause of the disagreement.
The cause is in the environments. And once you locate it, it explains not just these two findings but most of what has been confusing about AI and labor for the past three years.
The developers had a compiler. The entrepreneurs did not.
Write bad code and the machine tells you in seconds red text, failed build, no argument. Make a bad strategic decision and the market tells you in months, tangled up with variables you never controlled. One environment corrects error quickly and honestly. The other corrects it slowly, noisily, or not at all.
That gap does most of the explaining. Call it the speed of being wrong.
When the loop is fast and honest, AI is a leveler. A junior who accepts a flawed suggestion gets caught before the mistake compounds, so the floor of competence rises toward the ceiling. The referee on the field makes the weaker player nearly as safe as the stronger one. When the loop is slow or absent, AI is an amplifier. No referee. The strong performer uses the tool to sharpen judgments she can already check against her own experience. The weak performer asks it for things he cannot evaluate, cannot fully execute, and will not discover are wrong until it is far too late. The Kenyan entrepreneurs who did worse weren’t given worse advice. They were given confident advice in an environment that wouldn’t tell them, in time, that they’d misused it.

The variable, stated as a test you can run on any task in about ten seconds: How fast, and how honestly, does this work tell you that you got it wrong?
Fast and honest, AI commoditizes the work and flattens the people doing it. Slow and noisy, AI widens the gap between those who already know and those who only sound like they do.
The reason this is worth more than another “AI helps juniors” headline is that it cuts across the categories we normally use, splitting apart jobs that intuition files together.
Take two people who both “pick financial winners.” A day trader and a venture capitalist. Intuition groups them. Run the test and they fly apart. The day trader finds out he was wrong within minutes, by a P&L that does not care about his thesis. Fast, honest loop. That is why day trading was substantially replaced by algorithms long before ChatGPT arrived, and why a human trader with AI is being pushed toward a machine-defined par where the edge evaporates. The venture capitalist finds out whether a seed check was smart in seven to ten years, by which point the signal is hopelessly confounded with luck, follow-on rounds, and the macro cycle. AI can sharpen a great partner’s sourcing and memo-drafting but cannot rescue a mediocre one and — critically — cannot be checked against outcomes on any timeline that matters. The partner’s taste stays a premium because the world refuses to grade it quickly.
Same surface activity. Opposite loop structure. Opposite AI fate. The test predicts it; “investing” as a category predicts nothing.
It cuts inside a single profession too. A radiologist hunting a nodule works in a relatively fast loop - scan, biopsy, pathology, answer. A psychiatrist titrating treatment over months works in a slow, confounded one. The framework predicts that AI would be strong and leveling in the first and carry real hazards in the second, which is roughly what the deployment evidence suggests — though the picture in both fields is still developing. The point isn’t the specific prognosis for any specialty. The point is that “medicine” as a unit tells you almost nothing, while “how quickly and honestly does this subspecialty grade its errors” tells you quite a lot.
Here is where the standard career advice breaks down.
Fast feedback feels like the safe place to be. Errors get caught, AI obviously helps, the work goes smoothly. Nearly all of the “learn these tools or get left behind” advice points you toward exactly these tasks, because that’s where AI’s value is clearest and most immediate.
That visible value is the kiss of death. A fast, honest feedback loop is also, structurally, the same thing as legibility. And legibility is the precondition for commoditization. To measure a task precisely enough that the loop runs fast is to measure it precisely enough to route to something cheaper — a smaller model, a junior in a lower-cost market, a script. Enterprise AI infrastructure is already built around exactly this triage: semantic routers that send legible, fast-feedback queries to cheap open-weight models while reserving expensive frontier reasoning for the rare genuinely hard problem. The same classification is coming for people.
The durable position is not the task where AI helps you most. It is the task where AI is too dangerous to deploy unsupervised, where the loop is slow, the errors are expensive, and someone has to own an irreversible call and carry the consequence. The places AI feels risky are the places human judgment holds a premium. The places it feels magical are the places being priced toward zero.
Standard advice is pointing people toward the kill zone and calling it an opportunity.
There is another layer, and this is the one I find genuinely unsettling.
AI’s value depends on a fast feedback loop. And there is real evidence that AI’s proliferation slows feedback loops down.
The clearest example comes from medicine. A 2025 study in The Lancet Gastroenterology & Hepatology tracked endoscopists doctors hunting for polyps during colonoscopies before and after they began using an AI detection tool. With the AI active, they caught more. But when those same doctors later worked without it, their unassisted detection rate had fallen from 28% to 22%. This is not the straightforward deskilling story the feedback loop still ran, the scope and biopsy and pathology all happened. But the loop now closed through the machine instead of through the human. The physician’s own internal error-correction stopped getting exercised and atrophied. Whether this dynamic generalizes to other fast-loop professions we don’t yet fully know. But the mechanism is not obviously medical: offload the correction to a tool reliably enough, and the human stops running the circuit that made them competent.
At the organizational level, the pressure runs a different way. Microsoft’s own randomized study of over 6,000 Copilot users across 56 firms found workers completing documents faster and spending less time on email and initiating 11% more documents. Every one of those needs to be read and verified by someone downstream. The concern and it remains a concern rather than a settled finding, because the systemic evidence isn’t yet in is that verification burden rises faster than generation savings, lengthening the effective feedback loop for the organization even as individual output metrics look healthy. If that pattern holds at scale, AI would be quietly converting fast-loop environments into slow-loop ones, shifting the whole map over time.
AI degrades the precondition for its own value. Whether this shows up clearly in the macro data over the next decade is genuinely uncertain. But the mechanism is worth watching precisely because it’s invisible in the short run.
The labor math has a related trap, and it’s the one firms are most likely to spring on themselves.
When the legible, fast-feedback tasks get handled by AI, the rational firm response is to thin the junior tier. The Burning Glass Institute documented a 37% drop in entry-level postings in New York alone between 2022 and 2024. Friebel, Huang, Li, Shukla, and Zhang modeled the same shift this spring: the pyramid becomes a diamond as firms freeze junior hiring, and under a learning shock the diamond may be permanent. On a spreadsheet, this looks like discipline.
What disappears from the spreadsheet is that the junior tier was the apprenticeship. Not because of the output it produced, but because of the loop it ran through the person producing it. The rote discovery, the doomed first drafts, the code that didn’t compile until it did this was the mechanism by which tacit judgment got built, and it only works by doing the thing badly until you stop. Automate it for immediate efficiency and you stop manufacturing the judgment that slow-loop, AI-resistant work requires. The developer communities documenting this as a “tragedy of the commons” have it right: individual productivity metrics rise while the shared stock of deep competence that reviews and maintains the AI output is not being replenished.
Here is the asymmetry that the spreadsheet can’t see. Routing tokens from an expensive model to a cheap one is reversible the frontier capability still exists, you’ve stopped overpaying. When Klarna discovered its AI-first customer-service bet had quietly degraded quality, it could hire humans back. The capability had been dormant, not destroyed. But the apprenticeship layer that produces senior judgment doesn’t sit dormant when you automate it away. The seniors who would have been created simply don’t exist.
You can route tokens back. You cannot route experience back.
A hundred years ago, Frederick Taylor installed the first formal feedback loop on human labor: the stopwatch. Measure the worker, correct the worker, optimize the output. He told the founding story of the whole method a pig-iron handler whose daily output he claimed to have nearly quadrupled through precise measurement. The number became scripture. Subsequent historians Wrege and Perroni (1974) and others found the worker’s real name was Henry Noll and the famous figures were, in the load-bearing details, fabricated.
What matters here is not that Taylor lied, though he did. It is what a feedback loop with a corrupted signal actually produces. It runs. It corrects. It optimizes. And everything it optimizes is pointed at a number that doesn’t reflect reality. A corrupted loop is not just ineffective it is more dangerous than no loop, because it manufactures confident, measured, authoritative wrongness and calls it evidence.
We are now installing AI-driven measurement across nearly every kind of knowledge work: git commits, email response rates, documents completed, issues resolved per hour. Each of these is a feedback loop of a kind. Some are honest. Many are fast but shallow proxies that capture the legible surface of work while the slow, judgment-intensive, hard-to-measure parts go unrecorded and eventually get defunded.
The question worth asking about any feedback loop you’re operating inside, whether you installed it or inherited it, is not whether it is fast. It is whether it is honest. Whether the thing it is correcting you toward is the thing that actually matters, or whether it is, like Taylor’s stopwatch, a clean measurement of the wrong number.
That question does not change when the next model drops. Capability is what changes with model releases. Loop structure changes with how work is organized, which moves more slowly and is mostly under human control. Which is either reassuring or alarming, depending on who in your organization is currently deciding what gets measured.