Skip to content
Sadiq Khan

july 28, 2026 · 2 min read · llm, evaluation, benchmarks

Your eval dashboard is measuring nothing

Public benchmarks are saturated. The three measurements teams shipping agents actually rely on, and what each one costs you.

SWE-bench Verified is saturated. Claude Opus 5 hits 96.0 percent, GPT-5.6 Sol 96.2, Fable 5 95. GPQA Diamond is nearly there: GPT-5.6 Sol at 94.6 and Gemini 3.1 Pro at 94.3, against a human PhD baseline of 65. OpenAI itself recommended SWE-bench Pro in February 2026 as the harder replacement. In July they audited it, found roughly 30 percent of the 731 public tasks broken, and retracted their own recommendation. If your eval dashboard is any of these, you are measuring nothing.

Nobody notices because the numbers keep looking good. Public benchmarks leak into training data, every new frontier model tops them by a few points at release, and the dashboard stays green while users complain. The gap is that public benchmarks measure the model, not your task.

The teams shipping agents in production in 2026 rely on three different measurements.

Regression gating on a golden set

The Braintrust, LangChain and Confident AI pattern. Collect a few hundred to a few thousand labelled examples from real traffic, calibrate a judge model against human labels, run the whole set on every release, and block the deploy on regression. Braintrust ships this as a GitHub Action that fails the PR when scores drop below a threshold. Confident AI's DeepEval treats LLM outputs like unit tests you assert on. The signal is real; the calibration decays whenever the base model changes.

Production trace review

Langfuse and Arize. Instrument every model call in production, sample failures and low-confidence outputs, send them to a human review queue, and promote reviewed cases back into the golden set. Langfuse (acquired by ClickHouse in January 2026, with the MIT core kept) is the open-source lane. Arize AX shipped continuous production trace review as a first-class feature this year. The eval set grows with the product instead of being frozen at launch.

Task-specific rubric evals

RAGAS for retrieval; DeepEval and Promptfoo for hallucinations and adversarial red-teaming. Faithfulness, context precision, answer relevance, refusal rate: reference-free scores that catch shortcomings a generic LLM judge misses. Promptfoo alone ships 500-plus built-in attack vectors on the security side.

Choosing

There is no best one. You are choosing between how much you trust your own labels, how far into production you can afford to catch failures, and whether the metric survives a model version bump.

Public benchmarks still have one honest use in 2026: a coarse floor. Below a certain GPQA score a model isn't worth evaluating further. Above it, you learn nothing new about how it will behave on your task.


Originally posted on LinkedIn.

← All writing

Contact

sadiqkhan795@gmail.com

Say hello. I read everything.