july 28, 2026 · 2 min read · llm, evaluation, benchmarks
Your eval dashboard is measuring nothing
Public benchmarks are saturated. The three measurements teams shipping agents actually rely on, and what each one costs you.
SWE-bench Verified is saturated. Claude Opus 5 hits 96.0 percent, GPT-5.6 Sol 96.2, Fable 5 95. GPQA Diamond is nearly there: GPT-5.6 Sol at 94.6 and Gemini 3.1 Pro at 94.3, against a human PhD baseline of 65. OpenAI itself recommended SWE-bench Pro in February 2026 as the harder replacement. In July they audited it, found roughly 30 percent of the 731 public tasks broken, and retracted their own recommendation. If your eval dashboard is any of these, you are measuring nothing.
Nobody notices because the numbers keep looking good. Public benchmarks leak into training data, every new frontier model tops them by a few points at release, and the dashboard stays green while users complain. The gap is that public benchmarks measure the model, not your task.
The teams shipping agents in production in 2026 rely on three different measurements.
Regression gating on a golden set
The Braintrust, LangChain and Confident AI pattern. Collect a few hundred to a few thousand labelled examples from real traffic, calibrate a judge model against human labels, run the whole set on every release, and block the deploy on regression. Braintrust ships this as a GitHub Action that fails the PR when scores drop below a threshold. Confident AI's DeepEval treats LLM outputs like unit tests you assert on. The signal is real; the calibration decays whenever the base model changes.
Production trace review
Langfuse and Arize. Instrument every model call in production, sample failures and low-confidence outputs, send them to a human review queue, and promote reviewed cases back into the golden set. Langfuse (acquired by ClickHouse in January 2026, with the MIT core kept) is the open-source lane. Arize AX shipped continuous production trace review as a first-class feature this year. The eval set grows with the product instead of being frozen at launch.
Task-specific rubric evals
RAGAS for retrieval; DeepEval and Promptfoo for hallucinations and adversarial red-teaming. Faithfulness, context precision, answer relevance, refusal rate: reference-free scores that catch shortcomings a generic LLM judge misses. Promptfoo alone ships 500-plus built-in attack vectors on the security side.
Choosing
There is no best one. You are choosing between how much you trust your own labels, how far into production you can afford to catch failures, and whether the metric survives a model version bump.
Public benchmarks still have one honest use in 2026: a coarse floor. Below a certain GPQA score a model isn't worth evaluating further. Above it, you learn nothing new about how it will behave on your task.
Originally posted on LinkedIn.