may 12, 2026 · 1 min read · agents, harness, research
You might be optimising the wrong layer
A paper reached the top of SWE-bench with 12% fewer tokens by evolving the harness, not the model or the prompt.
Perhaps you are optimising the wrong thing.
Better model. Bigger context window. Smarter prompt. And yet the agent still fails at the exact moment it matters.
A recent paper explains why. The researchers took the same model to the top of SWE-bench using 12% fewer tokens: no fine-tuning, no new architecture, no prompt rewriting. They evolved the harness.
The harness is the invisible layer between your AI and the world. Not the model, not the prompt, but the scaffolding: how tools are defined, how memory is structured, how information flows to and from the model. Most people never touch it. They treat it as plumbing. The paper treated it as the product.
What they found was striking. Performance gains came almost entirely from tools, middleware and memory structures, not from prose instructions or system prompts. Prompts tell the model what to do. Harnesses determine what the model can do. There's a reason the same prompt produces wildly different results across setups.
The other finding hit harder. Harness improvements transferred across three different model families with zero retraining: +5 to +10 percentage points, just by importing a better-structured harness. Structural knowledge generalises. Text doesn't.
Which means the real leverage isn't in your prompt library. It's in the architecture of the layer your agents actually operate inside. Most teams are building on a foundation they've never questioned.
Paper: Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses.
Originally posted on LinkedIn.