Cut cost and latency by making your context packet's stable content — system rules, schemas, few-shot examples — a byte-identical prefix on every call, with volatile retrieval appended after it. When the prefix matches, the model skips re-processing it and you pay roughly a tenth of the price, with a shorter time-to-first-token.

Quick Answer: Mark the unchanging front of your prompt as a cache breakpoint. A cache hit costs about 10% of the base input price and cuts prefill latency — but only if the stable content comes first, in byte-identical order, and nothing volatile (a timestamp, a random ID, a reordered JSON key) sneaks in ahead of it.

What Prompt Caching Actually Is, and Why It's a PM Decision

Prompt caching lets a model provider skip re-processing the part of your prompt it has already seen, provided the bytes match exactly. It is not a cost switch an infra team flips after launch — it is a direct consequence of how you assemble the context packet, which makes it a design decision, not an implementation detail.

Most teams treat caching as something an engineer configures after the fact: add a cache_control marker, ship it, move on. That undersells it. Whether a packet can cache well is determined earlier, by whether you separated its stable and volatile parts before anyone wrote a line of API-calling code.

This piece is about prompt caching context design — the upstream choices that determine whether a cache breakpoint ever has a chance to hit. Our complete guide to context engineering covers that separation in full; this is the one lever it unlocks at the API layer.

It's also worth distinguishing this from the broader context vs. prompt engineering question. Caching itself is an infrastructure mechanic — a provider-side optimization triggered by matching bytes. Deciding what's stable enough to freeze, and in what order it renders, is a context-engineering decision. Done well, it's context cost optimization in its most literal form: the same information, restructured so a repeat call doesn't pay full price for it twice.

Why this matters at volume

A single call where caching saves a few cents isn't interesting. A support-triage agent, a coding assistant, or an eval harness that sends a structurally identical system prompt thousands of times a day is different — the stable prefix is the same packet, paid for from scratch, every time, unless someone deliberately restructured it to be reusable.

Order the Packet: Stable Prefix First, Volatile Retrieval Last

Caching is a prefix match, not a fuzzy-similarity match: the provider compares bytes from the start of the prompt and stops at the first difference. Put what never changes — system instructions, output schemas, few-shot examples, tool definitions — first, and push everything that varies per call to the end, after the breakpoint.

This ordering constraint is the whole game. If one volatile token (today's date, a session ID, an unsorted tool list) sits ahead of your stable content, the cache misses on every request, even when 95% of the packet is identical. The anatomy of a context packet walks through these layers — instructions, schemas, examples, retrieved knowledge, working memory — in more depth; caching is what happens when you render those layers in the right sequence and mark where the stable ones end.

A layout that caches well

A packet built for caching typically renders in this order:

  1. Tool definitions — static, versioned, sorted deterministically
  2. System instructions — the role, constraints, and output format, as frozen text
  3. Few-shot examples — fixed demonstrations that don't vary by user
  4. (cache breakpoint goes here)
  5. Retrieved context — documents or search results pulled per query
  6. The user's message — the one thing guaranteed to differ every time

Everything above the breakpoint should be identical, byte for byte, across requests. Everything below it is expected to change and isn't worth trying to cache.

What counts as "stable" — borrow from JTBD

Deciding what belongs above the line is easier with a borrowed frame. Jobs-to-be-Done thinking asks what job a product is hired to do, independent of who's using it or when. Ask the same question of your packet: what instructions define the job the model is hired to do on every single call, regardless of which user, which query, or which retrieved document shows up?

That job-defining scaffolding — the role, the schema, the constraints, the worked examples — is what belongs in the stable prefix. The specific situation the model is reasoning about this time belongs after it.

The Cost and Latency Math: Cache Hit vs. Cache Miss

A cache write costs more than an uncached call; a cache hit costs much less. Anthropic's published pricing puts a 5-minute cache write at 1.25x the base input price and a 1-hour write at 2x, while a cache hit runs about 0.1x — so a 5-minute cache breaks even at two calls to the same prefix, and a 1-hour cache needs three.

Call typePrice multiplier*What's happening
Uncached1x base input priceFull prefill, every token processed from scratch
Cache write (5-minute TTL)1.25x base input pricePays a one-time premium to populate the cache
Cache write (1-hour TTL)2x base input priceHigher premium, survives longer gaps between calls
Cache hit~0.1x base input priceModel skips re-processing the matched prefix

*Anthropic's published multipliers for Claude models. Other providers price the same mechanic differently — see the comparison below.

Why the break-even math works out this way

Two calls on a 5-minute-TTL cache cost 1.25 + 0.1 = 1.35x combined, versus 2x for two uncached calls — cheaper after just the second call. Three calls on a 1-hour-TTL cache cost 2 + 0.2 = 2.2x, versus 3x uncached — the larger write premium needs one more repeat to pay off. Below those call counts, caching is a net cost, not a saving.

A worked example

Say a coding agent's stable prefix — tool schemas, system rules, few-shot examples — runs 6,000 tokens, and gets called 50 times a day in bursts against a 5-minute TTL cache. At a representative $3-per-million-token input rate, an uncached run costs $0.018 per call, or $0.90 across 50 calls.

With caching, the first call in each burst writes the cache at $0.0225 (1.25x); the remaining 49 read it at roughly $0.0018 each (0.1x). Total cost drops to about $0.11 for the same 50 calls — an ~88% reduction on that portion of the request. This is illustrative math, not a benchmark from a live deployment, but it shows why the multiplier table above compounds quickly once reuse clears the break-even point.

Latency, not just cost

The other half of the payoff is time-to-first-token (TTFT). Skipping re-processing of a matched prefix means the model doesn't run full prefill on tokens it has already seen — only on the new, volatile suffix. Because the system prompt is typically the single largest stable block in the packet, it's also where the latency win is most noticeable: cache the system prompt and the model skips prefill on everything above the breakpoint, not just a handful of tokens.

Academic work on this exact mechanic backs up the direction, if not the exact numbers. A Yale University systems paper on modular attention reuse (Prompt Cache, Gim et al.) reported time-to-first-token reductions of several-fold on GPU inference, and considerably more on CPU-only setups, when a large shared prefix could be reused instead of reprocessed. Production numbers vary by prefix size and provider, but a cache hit is consistently faster, not just cheaper.

How other providers price the same mechanic

The prefix-match idea is now common across major providers, but the pricing shape differs enough to matter when estimating savings:

ProviderMechanicCost shape
Anthropic (Claude)Explicit cache_control breakpoints you place in the prompt~1.25x-2x write premium once; ~0.1x on reads
OpenAI (GPT models)Automatic — no breakpoints, kicks in above roughly 1,024 tokensRoughly a 50% discount on the cached portion, no separate write charge
Google (Gemini)Explicit context caching; cached content stored as a resourcePer-hour storage fee for the cache, plus a reduced per-token rate on hits

The practical difference: Anthropic and Gemini require you to decide where the stable prefix ends; OpenAI's automatic caching removes that decision but also removes your control over it. Either way, the underlying design work — separating stable from volatile — is identical.

The Rule of Thumb: When Does Reuse Rate Justify Restructuring?

Restructure a packet for caching once its stable prefix will be reused at least three to five times within the cache's TTL window. Below that, the write premium can erase the savings; high-volume, high-repetition surfaces — a support bot's system prompt, a coding agent's tool schemas — clear that bar almost immediately, while one-off analyses rarely do.

If the stable prefix won't be reused at least three times before its cache expires, restructuring the packet for caching is engineering effort spent on a saving that doesn't materialize.

Use call frequency, not raw traffic, as the input to this decision:

  • 1-2 calls per prefix, ever — skip caching. The write premium won't be recovered.
  • 3-10 calls within 5 minutes — a 5-minute TTL cache pays for itself.
  • 10+ calls, with gaps over 5 minutes — switch to a 1-hour TTL cache.
  • Continuous, high-frequency calls — any TTL works; consider pre-warming the cache with a zero-output request at startup so the first real call is already a hit.

Reuse rate isn't uniform across a product

The same context packet gets called differently depending on where it sits in the product. An agent invoked at every step of a mapped customer journey — onboarding, activation, renewal — accumulates repeat calls at wildly different rates at each stage. Map where in the journey the packet actually fires before deciding whether restructuring it for caching is worth the engineering time; a packet reused constantly at one touchpoint might be a one-off at another.

Why Caches Silently Break (and How to Audit It)

A cache that never hits fails silently — there's no error, just a cache_read_input_tokens value of zero on every request. The most common cause is a byte in the prefix that shouldn't be there: a timestamp, a UUID, or non-deterministic serialization sitting ahead of the breakpoint.

Common silent invalidators to grep for in your prompt-assembly code:

  • datetime.now() or similar timestamp calls inside the system prompt
  • Randomly generated request or session IDs placed early in the packet
  • JSON.stringify() / json.dumps() without sorted keys, so field order shifts between calls
  • A tool list rebuilt per user or per request instead of a fixed, versioned list
  • Conditional system-prompt sections (if flag: prompt += "...") that create a different prefix per flag combination

Verify, don't assume

Every response carries a usage object with cache_creation_input_tokens, cache_read_input_tokens, and input_tokens. If cache_read_input_tokens stays at zero across repeated, structurally identical requests, something upstream is rebuilding the prefix differently each time — diff the actual rendered bytes between two calls rather than guessing.

Caching is a system, not a switch

Moving a cache breakpoint doesn't just change one call's cost — it changes what's easy to patch later (a stable prefix is harder to update quickly), how stale a cached instruction can get before anyone notices, and how much retrieval you can afford to push into the volatile suffix. Systems thinking is a useful frame here: treat the packet as a system with feedback loops, not a static string optimized once and forgotten.

Where Prodinja Fits: Caching as a Context-Design Decision

Prodinja's Context Engineering tool, in the Studio, has you separate a context packet into fixed and dynamic parts and set a token budget for each layer before you ever wire up an API call. That's not a caching feature in itself — Prodinja is a prototype product, not a model-serving runtime — but it's exactly the decomposition a cache breakpoint requires, made explicit at design time instead of discovered later through a zeroed-out cache_read_input_tokens field in production.

Treating the fixed/dynamic split as a first-class design step, rather than something bolted onto the API call afterward, is what turns caching from an infra footnote into a cost lever a PM can actually reason about — how much of the packet is fixed, how often it repeats, and whether that reuse rate clears the break-even bar covered above.

Key Takeaways

  • Caching is a prefix match, not a similarity match — a single volatile byte ahead of the breakpoint invalidates everything after it.
  • Order the packet deliberately: stable content (system rules, schemas, few-shot examples, tool definitions) first, volatile retrieval and the user's message last.
  • The cost math has a break-even point: roughly two repeat calls for a 5-minute cache, three for a 1-hour cache, using Anthropic's published multipliers of ~1.25x/2x on writes and ~0.1x on reads.
  • Caching also cuts latency, not just cost — a hit skips prefill on the matched prefix, shortening time-to-first-token.
  • Reuse rate determines whether restructuring is worth it — three to five calls within the TTL window is the rough threshold; below that, don't bother.
  • Silent cache misses are common — audit for timestamps, unsorted JSON, and non-deterministic tool lists in the stable prefix, and verify hits via the usage object rather than assuming.
  • Separating fixed from dynamic content is a context-engineering decision, not an infra afterthought — the same split tools like Prodinja's Context Engineering studio ask you to make explicitly.

Frequently Asked Questions

What is prompt caching and how is it different from prompt engineering?

Prompt caching is an infrastructure mechanic: a model provider skips re-processing a prompt prefix it has already seen, if the bytes match exactly. Prompt engineering is about wording a single instruction well; caching only works if the structure of the packet — what's stable versus volatile — was designed with reuse in mind, which is a context-engineering concern, not a wording one.

How much does prompt caching actually save on cost?

On Anthropic's published pricing, a cache hit costs roughly 10% of the base input price, against a one-time write premium of 1.25x (5-minute TTL) or 2x (1-hour TTL). Net savings depend on reuse: two calls on a 5-minute cache already come out cheaper than paying full price twice, and savings compound with every additional hit inside the TTL window.

Does prompt caching reduce latency, or only cost?

Both. A cache hit skips the prefill computation on the matched prefix, which shortens time-to-first-token — how long a user waits before the first token streams back. The cost and latency benefits come from the same mechanic, so a packet worth restructuring for cost savings is usually also worth it for responsiveness.

How long does a cached prompt last before it expires?

It depends on the TTL chosen at the breakpoint — commonly 5 minutes or 1 hour on Anthropic's API, with the longer window costing a higher write premium. If real traffic has gaps longer than 5 minutes but shorter than an hour, the 1-hour TTL usually wins even though each write costs more, because it avoids repeatedly re-paying the write premium.

Is prompt caching worth setting up for a low-traffic feature?

Usually not, unless the stable prefix will be called at least three to five times within its TTL window. Below that reuse rate, the one-time write premium can exceed what an uncached call would have cost, so a feature that fires once per session with long gaps between calls is a poor caching candidate — restructuring it adds engineering effort without a cost or latency payoff.