july 20, 2026 · 2 min read · llm, prompt-caching, multi-agent, cost
Parallel agents don't share a prompt cache
Fan out five agents at t=0 and you pay for five cache writes. How Anthropic, OpenAI and Gemini cache differently, and the one move that works for all three.
Fire five Claude agents in parallel with the same system prompt and none of them share the cache with each other. Anthropic's docs say it directly: cache entries become available only after the first response begins, and are not available for concurrent parallel requests. Fan out at t=0 and you write five cache entries in parallel while paying full input tokens on every branch.
Nobody notices in development because the hits show up on retries within five minutes. In production, under real concurrency, cache hits collapse toward zero on the first request of each branch, and the input token bill scales linearly with your fan-out width. Every branch pays the cache-write premium (1.25x standard input for the 5-minute tier, 2x for the 1-hour tier), so five agents means five separate cache writes at t=0.
The three frontier providers each cache differently.
Anthropic: explicit cache_control blocks
Default 5-minute TTL, with a 1-hour tier if you set ttl: "1h". Cache writes cost 1.25x standard input for the short TTL and 2x for the long one; reads cost 0.10x, a 90 percent discount. Since February 2026 the cache is scoped per workspace, not per organisation. Parallel writes still race until the first response starts, so the standard trick is to warm the shared prefix with one sequential request before dispatching the batch.
OpenAI: implicit prefix caching
Automatic, no flag. As of May 29, 2026 the default retention is up to 24 hours (it used to be minutes). Cache reads on the GPT-5.x family are billed at 10 percent of standard input. The cache is keyed on the exact token prefix, so any drift in the invariant part of the prompt (a rotating date string, a shuffled few-shot order) silently kills the hit rate.
Google Gemini: paid context caching
You create a cached content object with a TTL you choose (default 1 hour, no documented maximum). Storage is billed per token per hour, roughly $1 to $4.50 per million per hour depending on the model. Reads on Gemini 2.5 and newer are 90 percent off. The cache is durable across requests and users inside a project, so a fan-out actually amortises the write cost.
The move that works everywhere
There is no best one. You are choosing between whether the cache is per request or per content, how long it lives, and how much of your prompt layout you are willing to freeze.
The engineering move is the same in all three: hoist the invariant prefix (role, tools, constraints, few-shot examples) to the top of the prompt, put the per-agent variance in the tail, and warm the shared prefix with one sequential request before dispatching the parallel batch. Skipping the warm-up call is the mistake I see most often.
Originally posted on LinkedIn.