Why LLM benchmarks mislead you (and what to do)
every time a new language model launches, the vendor posts a table. their model is in bold, the competitors are slightly lower, and the numbers have names like MMLU, HumanEval and GSM8K. if you are picking a model for a real project, it is tempting to read that table like a car review and buy the one with the biggest number.
i run small AI-assisted projects out of Singapore, and i have made that mistake. the model that topped the chart did not always do my actual job best. this article explains what benchmarks are, why the scores are less useful than they look, and what i do instead. you do not need any machine learning background to follow it.
what it is
a benchmark is a fixed set of questions with known answers. you give the questions to a model, count how many it gets right, and report a score. that is the whole idea. it is the same as a school exam, with the same strengths and the same weaknesses.
a few you will see constantly:
- MMLU: multiple-choice questions across 57 subjects, from law to elementary maths. the original paper is Measuring Massive Multitask Language Understanding.
- HumanEval: 164 small Python programming problems, introduced in the paper Evaluating Large Language Models Trained on Code. a solution counts as correct if it passes hidden unit tests.
- GSM8K: grade-school maths word problems.
- Chatbot Arena: a different style. real people see two anonymous answers to the same prompt and vote for the better one. the votes are turned into a ranking, described in Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.
“misleading” does not mean “fake”. the scores are real measurements of something. the problem is the gap between what they measure and what you need.
how it works
to see why the gap exists, follow a score from start to finish.
step one, someone writes the test. the questions are frozen so that every model sees the same thing. that is good for fairness and bad for realism. your customers do not send you tidy multiple-choice questions.
step two, someone runs the model on the test. here a lot of hidden choices matter:
- the prompt wording: the same model can score noticeably differently depending on how the question is phrased or how many worked examples are shown first.
- the sampling settings: temperature and similar settings change how random the answers are. i cover this in temperature, top-p and sampling explained.
- the scoring rule: does the model have to output exactly “B”, or is a sentence containing B fine? small parsing rules shift results.
- the number of attempts: some reports give the model several tries and count a success if any one passes. others count one try only.
two vendors can both say “MMLU” and not be running the same experiment. a number in a launch post is usually the vendor’s own run, set up in the way that flatters their model. that is not necessarily dishonest, it is just human.
step three, the score gets compressed into one figure. a single percentage hides everything about how the model fails. a model that is right 90 percent of the time with obvious errors is very different from one that is right 90 percent of the time with confident, plausible errors. the table shows both as the same number.
step four, the test gets old. this is the part newcomers rarely hear about. test questions are public. models are trained on huge scrapes of the internet. if the questions, or close copies of them, ended up in the training data, the model has effectively seen the exam. researchers call this contamination. it is hard to prove for any one model, and it does not require anyone to cheat on purpose. a public dataset gets copied into forums, blog posts and repositories, and a web scrape picks it all up.
there is also a pressure effect. once a benchmark becomes the number everyone quotes, teams tune toward it. when a measure becomes a target, it stops being a good measure. that idea is old, and it applies here.
finally, many benchmarks saturate. when top models all score in the high 80s or 90s, the remaining differences are small, and some of the remaining “errors” are bad questions or wrong answer keys in the test itself. the chart still shows a ranking, but the ranking is mostly noise at the top.
why it matters
four reasons i care, in terms of actual money and time.
reason 1, you can pay for capability you do not need. a frontier model costs more per token than a small one. if your task is classifying support emails into five categories, a cheaper model may do it equally well. a benchmark gap on graduate-level exam questions says nothing about that. see what is a token and why your bill depends on it for how those costs add up.
reason 2, you can pick the wrong model for the job. coding benchmarks test short, self-contained functions. real coding work means reading a messy repository, following a house style and not breaking other files. a model can look excellent on HumanEval and still be frustrating in your codebase. the reverse also happens.
reason 3, the things that hurt you are not on the chart. latency, how often the output format breaks, how the model behaves when it does not know, how it handles your language or your industry’s jargon. i work with Singapore English, Singlish-flavoured customer messages and some Mandarin, and none of the headline benchmarks tell me much about that.
reason 4, it affects trust. if you quote a benchmark to your boss or client as proof the tool will work, and it then fails on their real data, that costs you credibility. a number that sounds scientific carries weight it has not earned.
common misconceptions
“a higher score means a better model.” only for that test, run that way. a two-point gap on MMLU is well within the range that prompt wording and scoring rules can move. treat small gaps as ties.
“human preference rankings fix everything.” Chatbot Arena style voting is a genuinely useful complement, because real people write the prompts and judge the answers. but the voters are a self-selected crowd, the prompts skew toward what curious internet users ask, and people tend to reward longer, friendlier, better-formatted answers even when they are not more accurate. it measures what a broad crowd likes. it does not measure what your support queue needs.
“if a benchmark is old and public, it is safe.” the opposite is closer to the truth. the longer a test has been public, the more chance it has leaked into training data and the more teams have tuned for it. a fresh private test is worth more than a famous public one.
“the vendor’s numbers are lies.” mostly no. i have no evidence of fabricated scores, and i would not claim it. the issue is selection and setup. vendors pick the benchmarks where they look good and run them under favourable settings. you would do the same on your own launch day. read the table as marketing that happens to contain real measurements.
where to go from here
what i do now is simple and takes an afternoon. i collect 50 to 100 real examples of the task, write down what a good answer looks like, run two or three candidate models on them, and read the outputs myself. i keep the set private so it cannot leak into anyone’s training data. that small private test has told me more than any leaderboard.
if you want to go further, these are the topics i would read next:
- build your own test set: building an eval set from support tickets walks through turning real messages into a repeatable test. if those tickets contain customer details, read up on handling personal data first, and the privacy wire blog is a good place to start.
- decide between changing the prompt and changing the model: when fine-tuning beats prompting covers when extra training is worth the effort.
- understand the settings that move your results: temperature, top-p and sampling explained shows why one model can look good or bad depending on configuration.
- check what you will pay: what is a token and why your bill depends on it helps you cost out the model you end up choosing.
the full list of explainers is on the blog index.
benchmarks are a useful first filter. they can tell you a model is in the right neighbourhood. they cannot tell you it is the right one for you. the only test that does is your own work, run on your own data.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-04.