The AI DownsideDocumenting AI's downsides

LLMs

Why AI Benchmarks Mean Less Than You Think

A new model tops the leaderboard, the chart goes up and to the right, and none of it predicts whether the thing will be useful to you.

Abstract editorial illustration for “Why AI Benchmarks Mean Less Than You Think”.

Every model launch comes with a chart. Bars, usually, or a spider diagram, showing the new model edging past its rivals on a row of benchmarks with acronyms most people cannot expand. The bar is taller. The press writes it up as a leap. And within a week, users report that the new state-of-the-art model is, for their actual work, about the same as the last one or occasionally worse. The benchmark said one thing. Reality said another. This happens so reliably that it is worth understanding the mechanics of the gap.

The test is public, which ruins the test

The most fundamental problem is contamination. Many popular benchmarks are published, discussed, and sitting on the open web — which is exactly where models get their training data. When the questions and answers to your exam are in the study material, a high score measures memorisation as much as ability. Nobody needs to cheat deliberately; the leak is structural. A model can score brilliantly on a benchmark it has effectively already seen and then flounder on a genuinely novel version of the same task.

A benchmark stops measuring intelligence the moment it becomes famous enough to end up in the training data. Fame is the thing that breaks it.

The number becomes the marketing, and the marketing corrupts the number

There is a commercial feedback loop that makes benchmark figures even less trustworthy than their technical limitations alone would suggest. A high score is not just an engineering result; it is a marketing asset worth an enormous amount in attention, funding and credibility. That raises the stakes on every fractional improvement, and where the stakes are high, the temptation to select, frame and present the numbers favourably is irresistible. Vendors choose which benchmarks to headline, which comparisons to draw, and which unflattering results to leave in an appendix or omit entirely. The chart on the launch slide is not a neutral readout; it is a curated argument.

This is not necessarily fraud — it rarely needs to be. It is simply the ordinary gravity of a metric that has become a sales tool. When beating a particular number by a point translates into headlines and a valuation bump, engineering effort flows toward that number regardless of whether it corresponds to anything you care about, and communication effort flows toward presenting it as impressively as the facts allow. The benchmark started life as an attempt to measure capability honestly. By the time it is famous enough to appear on a keynote slide, it has been thoroughly repurposed into an instrument of persuasion, which is a different job with different loyalties.

Optimising for the test, at the expense of the job

Benchmarks are also targets, and targets get gamed — not always cynically, but inevitably. When a specific set of evaluations becomes the scoreboard the whole industry watches, enormous effort goes into nudging those specific numbers up. This is Goodhart's law in its purest form: when a measure becomes a target, it stops being a good measure. A model tuned to excel at benchmark-shaped questions is not necessarily a model that is better at your unglamorous, benchmark-unshaped problem.

The benchmark is not your job

Even a perfectly clean, ungamed benchmark would mislead you, because the tasks bear little resemblance to real work. Benchmarks favour the things that are easy to score automatically: multiple-choice questions, problems with a single verifiable answer, self-contained puzzles. Your actual use is messier — a long, ambiguous document; a vague request; a task where “good” is a matter of taste and context and there is no answer key. A model that aces graduate-level multiple choice may still write emails you would be embarrassed to send.

The dimensions you actually care about are mostly unmeasured by the leaderboard: does it follow instructions precisely, does it keep its tone consistent, does it refuse sensible requests (a problem we covered in our piece on over-cautious refusals), does it stay coherent over a long session, is it fast enough not to break your flow. None of those fit neatly on the launch chart.

“Human preference” leaderboards have their own trap

The response to all this has been crowd-sourced arenas where humans vote on which of two anonymous answers they prefer. These are genuinely more useful than static exams — but they measure preference, not correctness, and preference has biases. People tend to prefer longer, more confident, more flattering answers, which rewards models for being verbose and agreeable rather than accurate and concise. A model can climb a preference leaderboard by being a better sycophant, which is not the quality most of us are shopping for.

How to actually evaluate a model

A single number for a many-shaped thing

Underneath every benchmark complaint is a category error: the attempt to collapse a wildly multidimensional thing into a single rankable score. “How good is this model” is not one question. A model can be superb at code and mediocre at prose, brilliant at short tasks and lost over long ones, precise at following instructions and hopeless at knowing when to refuse. These qualities do not move together, and a customer cares about different ones. Averaging them into a leaderboard position discards exactly the information you needed and hands you a number that is true, precise, and useless for your decision.

This is why two people can use the “same” top-ranked model and reach opposite verdicts. The one doing bulk classification loves it; the one writing nuanced long-form finds it flat. Neither is wrong, and the benchmark cannot adjudicate between them, because it measured a blend of capabilities that neither of them actually has. A ranking implies a single axis of better-and-worse. Real model quality is a landscape, and the leaderboard is a photograph of it taken from one arbitrary angle, sold as the view.

The numbers are aimed at investors, not users

It helps to remember who the benchmark chart is really for. A dramatic score is a fundraising asset, a recruiting asset, and a press asset long before it is a user asset. In a field where enormous sums move on the perception of being at the frontier, a benchmark result is a claim to that frontier — and the audience that most rewards the claim is not the person choosing a tool for their work, but the investor, the journalist and the prospective hire deciding who is winning. The chart is optimised for that audience, and its conventions follow accordingly.

Which means the leaderboard is best read as a marketing artefact that happens to be expressed in numbers, rather than a measurement that happens to be useful for marketing. Numbers carry an air of objectivity that a slogan does not, and that borrowed authority is much of their value to the vendor. Your defence is to decline the frame entirely: not to argue about whose benchmark is fairer, but to stop treating the leaderboard as the thing that decides, and to move the decision back onto the only ground that is contaminated by nobody's incentives — your own real tasks, run yourself, judged by whether the output was any good.

Benchmarks shape what gets built, not just what gets sold

The quiet damage of benchmark obsession is not only that it misleads buyers; it is that it steers the technology itself. When a specific set of scores is the scoreboard the whole field watches, research effort, training choices and product decisions all bend toward moving those particular numbers. Capabilities that happen to be benchmarked get lavish attention; capabilities that matter to real users but resist tidy scoring — consistency, restraint, knowing when to refuse, staying coherent over a long session, admitting uncertainty — get comparatively neglected, because there is no leaderboard to win for them. The measure does not just describe progress; it decides what “progress” is allowed to mean.

This is Goodhart's law operating at the scale of an entire industry. Optimise hard enough for a proxy and you get models finely tuned to the proxy and subtly misaligned with the goal it was standing in for. A generation of models sculpted to ace exam-shaped questions may be genuinely worse, in ways no benchmark records, at the messy, ambiguous, judgement-laden work that actually fills your day — not because anyone chose that trade-off, but because the scoreboard never counted the thing that was quietly lost. So the leaderboard misleads twice over: once when you read it, and once, invisibly, in the shape of the models it helped bring into being. The only defence against both is the same — hold the measure loosely, and keep your own tasks as the thing that actually decides.

The only benchmark that reliably predicts whether a model is good for you is the one nobody can sell you: your own tasks. Keep a small, private set of real problems from your actual work — the kind of thing you use these tools for every day. When a new model appears, run your set. Ignore the chart. The results will frequently disagree with the leaderboard, and when they do, trust your set. It is contaminated by nothing, gamed by no one, and it is measuring the only thing that matters, which is whether the tool is useful to you. The leaderboard is measuring whether it is useful to the launch.

Related grievances

All articles →