AI benchmarks and why they mislead
I ran one evaluation twice against one model and changed a single thing: the wording of the instruction wrapped around the questions. Not the questions. The instruction.
The score moved further than the gap I had been using it to settle.
That gap was two points between two models, quoted in a launch post, and I had spent an hour that morning taking it seriously. After the second run I could not tell you which model was better, and I had produced that uncertainty myself in four minutes with a text edit.
This is not the argument that benchmarks are fake. Benchmarks are among the more honest objects in this industry. A benchmark is a fixed list of questions, a fixed list of correct answers, and a script that counts. Everything that goes wrong happens to the number afterwards, on its way to you.
The measurement is narrow and the claim is not
Somebody chooses a task. Somebody assembles a few hundred items. Somebody runs a model over those items under specific conditions and records a percentage.
That percentage is a genuine measurement of one narrow thing, and the paper it came from usually says exactly how narrow, in plain language, near the front.
Then it becomes a bar on a chart. The chart goes in a launch post. The launch post gets compressed on social media into “better at reasoning.”
No one committed fraud in that sequence. A narrow measurement became a general claim by being repeated, and the general claim is what people make purchasing decisions on.
Two points is inside the noise
Start here, because it invalidates most benchmark conversations before any of the deeper problems get a chance to.
The things that move a score, at roughly the magnitude of the differences vendors advertise: how the instruction is phrased, how many worked examples sit in front of the question, the temperature, and whether the grading script accepts an answer wrapped in a polite sentence or demands a bare token. That last one is worth several points on its own in my experience, and it belongs to the evaluation code rather than to the model.
Then there is sampling. Leave any randomness switched on, run the identical model over the identical set twice, and the number moves by itself. Plenty of results get published as a single run, which is a coin toss presented as a measurement.
So when a post says one model beat another by 1.3 points, part of what you are reading is a comparison of two evaluation setups written by different people for different purposes.
My working rule: if I can shift the number more by rewriting the prompt than the two models differ from each other, the two models do not differ.
The answers may already be in the model
This is the mechanism that is hardest to rule out and it quietly weakens everything built on top of a public score.
Benchmark items are published. They live on the open web, in repositories, in papers, on forums where people argue about individual questions and paste the answers underneath. Training corpora are scraped from the open web. So there is a permanent possibility that some fraction of any public test is already inside the model, and that the score partly measures recall of an answer rather than the ability the task was designed to probe.
I am not accusing anyone of cheating and I will not name a model. Contamination is the default outcome of scraping the internet at scale. It is what happens when nobody does anything in particular.
What has changed is the difficulty of ruling it out. When training sets were a few hundred gigabytes, searching them for overlap was a real option. When the set is most of the reachable web plus whatever was licensed, “held out” degrades toward “hoped.”
It also flatters models exactly where you would most want the number to be trustworthy. The hard items are the ones people wrote up and argued about in public, which is precisely how they ended up in a scrape. The boring questions nobody blogged about stay clean.
The lab answer is a private test set nobody outside can see. That helps and it trades one problem for another, because you cannot audit a set you are not allowed to look at, and a private set stops being private once it has been run against enough vendors who keep logs.
My own check is cheap. Take a puzzle a model handles perfectly. Change the names and numbers, keep the structure identical, ask again. Sometimes it sails through, which tells you something real. Sometimes it falls apart in a way that only makes sense if it had met the original wording.
You are shown the charts they won
A launch post contains the benchmarks where the model came first. The ones where it came fourth exist and will not be in the post.
Nobody is lying. I write comparisons and I choose which cases to show. Anyone would.
The consequence is that a vendor chart approximates a best case bound rather than an estimate of typical behaviour. And lifting a number from one company’s post to sit beside a number from another company’s post is worse than useless, because the run conditions differ and neither post tells you how.
What happens to a number once it is worth money
Once a benchmark is public and commercially load bearing, it becomes a target, and the mechanics are dull. You train on data resembling the task. You tune output formatting until the grader stops rejecting correct answers on technicalities. You choose decoding settings suited to that shape of question. You run the evaluation several times and publish the better run.
Each of those is defensible alone. Together they lift the number faster than they lift the ability underneath it, and eventually the benchmark saturates: everyone crowds into the top few points and the spread stops meaning anything.
There is a quieter version. I once read a couple of hundred items from a public test set by hand and found answers marked correct that I would have marked wrong. That is normal for a set built at speed. It means that once a model sits within two points of the ceiling, some of what is being scored is agreement with the mistakes in the answer key.
Most benchmarks have a useful life of two or three years. Several famous ones are past theirs and still being quoted at me.
Tidy questions, ugly inputs
Benchmark items have been edited. Someone wrote them, someone checked them, and someone discarded the ambiguous ones, because ambiguous items make automatic grading impossible.
The input on my desk this morning was a supplier quote sent as a photograph of a screen with two line items obscured by a reflection.
There is a second half people skip. Benchmark questions are phrased by someone who knows the right vocabulary for the task. Your users do not. They ask for the wrong thing using the wrong word and expect the model to work out what they meant, and no public set grades that.
A score describes performance on the clean, well posed version of your problem. The clean version is the one you never needed help with.
What I still use them for
Ruling models out. The floor carries far more information than the ceiling. A model near the bottom of a serious coding evaluation is genuinely not writing your migration script, and knowing that in advance saves an afternoon.
Direction over time. Model against model at one instant is noise. This year against three years ago on the same fixed task is real signal, and close to the only way to discuss the rate of progress without hand waving.
Regression detection, which is the one I actually depend on. Providers update models behind an unchanged name. I had a classifier’s category distribution shift over about ten days with nothing changed on my side and nothing logged, because a model developing a new favourite label raises no error anywhere. A fixed set of cases, rerun on a schedule, caught it. It does not need to be a famous benchmark. It needs to be fixed and yours.
And shortlisting, which I will admit to rather than posture about. Public numbers take twenty candidates to four in ten minutes. The four are candidates, not a ranking.
The replacement is small and it is not a leaderboard
Twenty of your own real inputs, with the correct answers written down before anything runs. That is the whole thing. It takes an afternoon, most of which happens before you open the tool, and the procedure has its own piece on this site, so I am not repeating it here.
The point that belongs in this argument is narrower. No leaderboard can be improved into that. A public benchmark cannot contain your inputs, because your inputs are private, specific and mildly embarrassing, which is the exact property that makes them worth testing on.
Where this runs out
For about a year I went too far the other way, treated every published number as marketing, said so loudly, and stopped reading. That cost me. I missed a real capability shift on a task class I care about for a couple of months, because the first evidence of it arrived as a boring chart and I had decided charts were noise.
I also cannot rule out contamination on my own sets. A good few of my test inputs were pulled off public pages years ago and I have no way to check whether a model has read them. Neither do you.
And I have nothing useful for someone who cannot run their own evaluation at all. If you are choosing a model for a task you do not control and cannot test, public numbers and vendor charts are what you have, and my advice shrinks to reading them as a floor rather than a forecast. What I can offer instead is comparisons run on my own cases with the answers written down first, published with the point where each model broke rather than the point where it won, and those are here.