The Benchmark That Matters Is Your Own Work
Public model benchmarks measure someone else's problem. A weekend spent building a private test bench turned up three things no leaderboard would have shown.
Choosing a model has become a spec-sheet exercise. Somebody posts a chart, a leaderboard shifts, a vendor claims a new frontier, and teams pick the model at the top of the list. Then they ship it into their own product and discover the ranking had almost nothing to do with the work they actually do.
That gap is worth closing, and closing it is cheaper than most teams assume. A single web page, a few hours, and an API key are enough to run the same prompt and the same context against three models at once and watch what comes back — with tokens, latency, and cost per model in plain sight.
The tool that came out of it is live at jcornelius.com/claude-bench. It is free, it runs entirely in your browser, and it uses your own API key. But the tool is not the interesting part. What building it exposed is.
The models do not take the same request
The first assumption to die was the simplest one: that “run the same prompt against all three” is a thing you can do.
It is not. Claude Haiku 4.5, Sonnet 4.6, and Sonnet 5 accept meaningfully different request shapes. Haiku takes a fixed thinking budget. The Sonnets take adaptive thinking. Effort levels exist on the Sonnets and not on Haiku, and the two Sonnets support different sets of them. Temperature was removed on the newest model. Send one identical request to all three and two of them return an error before a single token comes back.
This matters well beyond a test bench. Any code that treats a model family as interchangeable parts is carrying an assumption that quietly expired. Model choice is not a config value you swap. It is a capability boundary, and boundaries need handling.
Price does not follow the version number
The second surprise was the pricing table. Sonnet 4.6 costs $3 per million input tokens and $15 per million output. Sonnet 5 — newer, more capable — costs $2 and $10. The upgrade is a third cheaper.
Anyone budgeting on the intuition that newer means pricier would have built the wrong forecast. Anyone who standardized on an older model last quarter to save money is now paying a premium for less. Neither team would notice, because neither has a reason to look. The assumption feels too obvious to check.
The general form of this: costs that move in the direction you expect get audited. Costs that move backward do not.
The real difference was not quality
Here is the finding that changed how the work gets done.
The models were asked to write a restaurant special — a short piece of marketing copy, the kind of task that runs a hundred times a month. All three produced good copy. The output was close enough that picking a winner on quality alone would have been a coin flip.
What separated them was discipline. One model returned the copy and stopped. Another wrapped the deliverable in a warning about missing information, then closed with a paragraph explaining its own creative choices, under a heading, below a horizontal rule. Useful commentary in a conversation. Useless when the output is meant to be pasted into a menu.
The instinct was to strip that commentary out in the interface. That instinct was wrong twice over. A bench that edits what the models return is not measuring anything, and hiding the difference would have hidden the actual result — verbosity discipline was the sharpest distinction between these three models on this task.
The fix belonged upstream, in a system prompt the bench did not yet have. One instruction — return the artifact, nothing else, no preamble, no sign-off, no commentary on your own choices — and the gap mostly closed.
Which means the honest conclusion of the first real test was not “use this model.” It was “the prompt was underspecified, and the model comparison was measuring that.”
What this is really about
Most teams evaluating AI right now are reading someone else’s benchmark and treating it as an answer to their own question. It is not. A leaderboard measures performance on a curated task set that resembles nobody’s actual product. Your evaluation set is your own work — your prompts, your context, your output format, your tolerance for a model that talks too much.
The good news is that building the thing that answers your question takes an afternoon, not a quarter. And the moment it exists, the conversation changes shape. It stops being an argument about which model is best and becomes a specific, cheap, repeatable question: on this task, with this context, what does each option cost, how long does it take, and does the output arrive in a form anyone can use?
That question has an answer. The leaderboard one never did.
Two habits are worth carrying out of this beyond model selection. Check the assumptions that feel too obvious to check, because those are the ones nobody audits. And never build the version of the dashboard that flatters the result — a measurement that cleans up its own inputs is not a measurement.
The bench is here. Bring your own key, paste in a task you actually run, and see what your own work says.