AI Has Passed Every Exam. It Has Never Had an Idea.
The machine has passed every exam we can build and discovered nothing.
The machine has passed every exam we can build and discovered nothing. Both are true. What dissolves the contradiction is the one thing no test has ever measured in machines, or in us.
Here are two facts about AI in 2026 that you have almost certainly seen, but never had to hold in the same hand.
The first: the machine passes everything. Humanity’s Last Exam, a wall of graduate-level questions built specifically to be too hard, it clears. Gold at the International Math Olympiad. The bar exam, the medical boards, competition-grade code. We keep building harder tests, grandly named, designed by PhDs to find the ceiling, and the machine keeps stepping over them. On any exam a human can be given, it is now roughly the best test-taker alive. Frontier models gained 30 percentage points in a single year on Humanity’s Last Exam alone.
The second is harder to state cleanly than it was even a year ago. AI systems now genuinely contribute to scientific discoveries. AlphaFold, which won the 2024 Nobel Prize in Chemistry, GNoME, which predicted 2.2 million crystal structures of which outside labs have physically synthesized 736, FunSearch and others have shown that much. But notice the shape of those discoveries. In every case, a human decided the question was worth asking. The machine answered it, often spectacularly. I could not find an example where it independently decided which question mattered. No AI has posed a question nobody thought to ask, found an anomaly nobody was looking at, or made the kind of leap that turns a field inside out. It answers our questions superbly. It has never, on its own, decided which question was worth asking.
Both true. Sit in the discomfort for a second, because the discomfort is the whole essay. The best test-taker in history, and it has never had an idea. If intelligence were one thing, that shouldn’t be possible.
So what mechanism would reliably produce this outcome, over and over, regardless of how good the underlying system gets?
The easy dissolves don’t work
Everyone reaches for a comfortable way out of that contradiction, and both exits are blocked.
The first exit: it’s not really thinking, it’s just autocomplete. That feels good and explains nothing, because “just autocomplete” doesn’t clear Humanity’s Last Exam. Whatever it’s doing, dismissing it doesn’t survive the scoreboard.
The second exit: give it time and scale, discovery is coming. Maybe. But that’s a promise, not an explanation, and it dodges the actual question which is why the gap has this particular shape. Why superhuman on every answer and silent on every question? Scale doesn’t explain a structural asymmetry. It just promises to wash it away later.
Around halfway through writing this essay, Tom Zahavy at Google DeepMind published a position paper called LLMs Can’t Jump. For about ten minutes I thought I’d been beaten to the idea. Then I realized we were asking different questions. His argument runs through Peirce: models handle induction and deduction, but not the abductive leap that invents an explanation the data never contained. He has since clarified that this is a personal position rather than DeepMind’s, and that scaling might prove him wrong.
Zahavy asks whether the machine can make the leap. I found myself asking something slightly stranger: suppose it did. How would our benchmarks know?
So there’s a hidden variable, the way there always is when two careful observations point in opposite directions. And the variable isn’t in the machine. It’s in the tests.
A benchmark is a question someone already chose
Here is the thing that dissolves the paradox, and once you see it you can’t unsee it.
Every benchmark, every exam, every one of those grandly-named tests, has the same structure: someone writes the question, and the machine finds the answer. The question is given. It arrives pre-selected, pre-formatted, flagged as important, with a known answer sitting in a locked drawer so the thing can be graded.
And that means every test we have ever built measures exactly one half of intelligence, the finding of answers and structurally cannot measure the other half: the choosing of the question. A test can’t measure question-choosing, because a test is a chosen question. The container can’t weigh the thing that decides what goes in the container.
That sounds like wordplay until you look at where the great leaps actually came from, and notice that the choosing was always the hard part.
Einstein didn’t answer the question. He found it.
In 1887, two physicists named Michelson and Morley ran an experiment expecting to measure how the Earth’s motion changed the speed of light. They got nothing. Light moved at the same speed no matter which way you chased it. A clean, baffling null result and they published it.
Then it sat there. For eighteen years, in the open, in the literature, available to every physicist alive. The anomaly wasn’t hidden. The data was on the shelf.
In 1905 a patent clerk picked it up and did the thing nobody else had done. He didn’t find a cleverer answer to the question everyone was asking how does the Earth’s motion affect light? He decided that question was wrong. He found the existing picture of absolute time and space intolerable in a way his colleagues, staring at the same data, did not. The leap wasn’t the mathematics of special relativity; the math is undergraduate now. The leap was upstream of the math, the decision that simultaneity itself was the thing to doubt. Everyone had the answer-shaped problem. Einstein found the question-shaped one.
The physicist David Deutsch puts the general version cleanly: the people who changed science didn’t extrapolate from what was known. They stepped outside the space of existing explanations and proposed one that wasn’t logically waiting there to be derived. Darwin, Turing, Einstein none of them was doing better test-taking. They were deciding, against the grain of everyone around them, which question deserved a life.
The test we built for genius, and what we had to hand it
Now the part that should make you put the coffee down.
Researchers at DeepMind have proposed what they call the Einstein test the sharpest benchmark yet for real machine genius. The design: feed an AI everything known before the breakthrough, and see if it can independently produce the leap. Give it the physics of the era and see if relativity falls out.
It’s a beautiful idea. And look at what we had to do to build it.
We curated the dataset around the discovery. We pre-selected the era, the field, the anomaly, and pointed the machine straight at it: here is everything before 1905, now derive relativity. We built a test for the one man whose genius was deciding what to look at and to make it gradeable, we did the looking for it. I don’t think this is a flaw in the benchmark. I think it’s unavoidable. The moment you need a score, someone has to decide what counts as the question.
The Einstein test skips the part that made Einstein. It measures whether the machine can answer the question once a human has already found it. Which is the same thing every other benchmark measures, wearing a lab coat. Even our test for genius is, underneath, an answer key, because an answer key is the only kind of test we know how to write.
The measurement cannot contain the thing it’s trying to measure.
Why every test we build has this hole
There’s a deeper reason our benchmarks all share this blind spot, and it’s the quiet assumption running the entire debate.
We have graded answer-finding for three thousand years. Exams, vivas, boards, olympiads, we are extremely good at scoring whether someone can produce the right answer to a set question, because that’s the part of the mind we long ago learned to make legible. We have never once built a reliable test for curiosity, for the decision that a boring answer is intolerable, that a settled question isn’t settled, that this problem and not that one deserves a decade. We can’t score it because we’ve never figured out how to make it legible, even in each other.
The closest anyone has come is instructive MEDIQ strips patient information out of medical question-answering so a model has to decide at each turn whether it knows enough or needs to ask. Prompting models to ask questions dropped diagnostic accuracy by 11.3 percent. CLAMBER found models fail at clarifying questions because they cannot assess the boundaries of their own knowledge. Curiosity isn’t just unmeasured. It’s penalized.
So our benchmarks inherit our own blindness. They measure the half of the mind we can see and are silent on the half we can’t and then we’re startled when the machine sails past us, without noticing that it’s sailing past us on the only stretch we ever learned to measure. The machine looks like it’s overtaking human intelligence. It’s overtaking the articulable part of human intelligence, which was always the smaller, cheaper part.
What struck me wasn’t the benchmark itself. The same optimization pressure reappears after training. In deployment, reliability is the product, so inference is made as deterministic as possible, even at significant computational cost: Thinking Machines Lab traced residual nondeterminism at temperature zero to GPU batching and fixed it at roughly 60 percent slower inference. In academia unstable creativity benchmark are treated as a problem to be engineered away with tighter rubrics. Different incentives, same direction. That’s probably an essay of its own.
Three independent systems end up selecting for the same thing. Benchmarks determine what gets measured. Training determines what gets optimized. Products determine what survives deployment. None of them explicitly sets out to eliminate question-finding. Together, they may.
Where I’ll stop short
Let me refuse the too-easy version, because it’s the mirror of the hype and I don’t believe it.
I am not telling you a machine can never ask a real question. Two honest problems with that claim. First, the Einstein test is leakier than it looks the anomaly is sitting in the pre-1905 data, so a sufficiently good search might surface relativity without any genuine leap, which means even this test doesn’t cleanly separate asking from answering. Second, machines already produce things nobody designed: novel protein structures, game moves no human would play, research systems that propose their own next experiments. Those are real. But so far they are a straight-A student’s answer to a question the rules already posed extraordinary answer-finding inside a space we defined, not the leap of deciding the space was wrong.
So the honest claim isn’t “the machine can’t.” It’s stranger and more unsettling than that: we have no instrument that would tell us if it could. We have never built a test that measures wanting-to-ask, because a test is a thing you’re handed. So on the single capability that would actually mark the arrival of a new kind of mind, we are flying completely blind and we’ve mistaken our own blindness for the machine’s limit.
The half we never learned to grade
There’s a tool hiding in all of this, and you can run it on your own life before you finish the page.
Ask of any test you’re inside a benchmark, a KPI, an exam, a performance review: did someone hand me this question, or did I decide it was the question? The first is answer-work, valuable, gradeable, and exactly the kind of work now being handed to machines. The second is question-work, and it has always been the scarce thing, in science, in a career, in a life. We just never learned to measure it, so we quietly built a civilization that rewards the half we could score.
We are now doing it again, at the scale of datacenters. We will keep building harder exams, and the machine will keep clearing them, and every time it does, someone will announce that genius has arrived. It hasn’t. It won’t show up on the scoreboard, because the scoreboard was only ever built to grade answers, and the thing we’re waiting for was never an answer.
We built a test for Einstein and handed it Einstein’s question. The genius was never in the derivation. It was in deciding the question was worth a life and that is the one exam we have never known how to write, for a machine or for ourselves.