What is prompt caching?
Cache reads can cut input cost by up to 90%. About 4 minutes.
Published 2026-09-19 · Updated 2026-09-19
The first time you send a long prompt, the model reads every token at full price. The second time, it shouldn't have to. Prompt caching lets providers remember chunks of your prompt and bill them at a steep discount on the next request.
What gets cached
Prefixes, not random spots. A cache stores the early part of your request — the big system prompt, pasted documents, tool definitions — anything that stays identical between calls. Change a character too early and the cached path breaks: everything after that point is re-read at full price.
How the pricing works
Cached reads are cheap; the first request pays a write premium:
- Cache read: up to 90% cheaper than base input on Claude models. OpenAI discounts cached input by 50–90% depending on the model.
- Cache write: costs a little more than plain input (1.25× on Claude) — the provider stores the tokens for you.
- TTL: caches are short-lived, minutes not days. An active agent session keeps hitting them; a request hours later starts cold.
When it pays off
Caching rewards repetition: chat with a long fixed system prompt, agents calling the same tools all day, RAG pipelines reusing the same document library. One-off one-line prompts get nothing — there is nothing to reuse.
Stowly's calculator prices cache tiers in — open it, pick a model, and use the cache controls to see the gap between a cold and a warm prompt. Every rate is cited on Sources, and the exact math is in Methodology.