Blog
→
AI Spend Management

Prompt caching explained: lower AI cost and latency

Written by:
Prompt caching explained: lower AI cost and latency
Share on XShare on LinkedInShare on FacebookShare by email
In this article
Ready to get started?
Get full visibility into your spend in minutes.
Get started for free

Prompt caching

Prompt caching lowers AI cost and latency by reusing the expensive part of an LLM request: the unchanged prompt prefix. If your product keeps sending the same system instructions, tool definitions, schemas, long documents, or conversation history, you are paying the model to process the same material again and again. Prompt caching turns that repeated work into something reusable.

That sounds technical, but the impact is commercial. Faster cache reads can improve time to first token. Fewer repeated input charges can improve feature margin, which is central to AI unit economics for SaaS. Multi-turn apps become more predictable instead of quietly growing more expensive with every user turn. For founders and finance teams, that matters a lot more than a clever infra diagram.

The catch is simple. Prompt caching only works when the reusable prefix stays stable. Change the wrong tool, insert a timestamp too early, rewrite old messages, or place the cache boundary after dynamic content, and your hit rate drops. This guide explains what prompt caching does, how it works, when it actually saves money, what breaks it, and how to measure the result in a way the business can use.

Why teams care about it

LLM applications tend to get more expensive as they get more useful. The first version might send a short system prompt and one user message. The next version adds tool definitions, output schemas, policy text, retrieved context, prior conversation, and a layer of developer instructions. Then the product works better, but every request gets heavier. In chat and agent workflows, each turn can force the model to reprocess a large amount of old context before it even starts on the new question.

Prompt caching attacks exactly that problem. Instead of recomputing the same prefix every time, the provider reuses previously processed state for the unchanged portion and only spends full compute on the new suffix. That usually means three practical wins: lower repeated input cost, lower latency, and cleaner scaling behavior in long-running sessions.

It also solves a business problem that often arrives later than the engineering problem. AI spend is rarely painful on day one. It becomes painful when the product succeeds, usage grows, and the shared prompt layers become large enough to drag margins down. Prompt caching is one of the few optimizations that can help users feel a faster product while helping finance see a tighter cost curve.

What prompt caching actually means

Prompt caching is the reuse of previously processed prompt context, usually the stable prefix at the start of a request. Think of it as caching the setup, not the whole exchange. If the same instructions, tools, schemas, documents, and earlier messages appear again in the same order, the model provider can reuse prior work and avoid processing that part from scratch.

That is different from response caching, where you return the same final answer for the same request. Response caching works when both the question and desired answer should be identical. Prompt caching is more flexible. It lets you reuse the expensive shared context while still generating a fresh answer for the new user input at the end.

It is also related to KV caching, but not identical from a user perspective. KV caching usually describes the lower-level reuse of internal model state for earlier tokens. Prompt caching is the API or platform feature you work with: matching rules, breakpoints, token accounting, TTL, privacy behavior, and pricing. In practice, teams often use the term prompt caching to describe the whole feature, even when the underlying mechanism is a form of KV-state reuse.

One more distinction matters. Prompt caching is usually about exact or near-exact prefix reuse, not semantic similarity. If you want to reuse answers for similar questions, that is closer to semantic or response-level caching. If you want the model to stop reprocessing the same prompt scaffolding, prompt caching is the right tool.

What usually belongs in the cached prefix

The best candidates are the parts of a request that stay stable across many calls. That often includes system or developer instructions, tool definitions, function schemas, policy text, few-shot examples, long reference documents, large repo summaries, and earlier conversation turns that should remain unchanged. The current user question, current timestamp, request-specific metadata, or one-off retrieved snippets usually belong after the cache boundary.

The important detail is that providers typically match the full rendered prefix, not just the visible user text. If your request includes hidden instructions, tool metadata, structured-output settings, or multimodal content, those may also be part of the cacheable context. That is why cache hits can fail even when the last user message looks almost identical.

How it works inside an LLM request

The first request writes the cache

On the first eligible request, the provider processes the prompt normally and creates a cache entry for a chosen prefix. Some platforms do this automatically. Others let you mark an explicit cache breakpoint. Either way, the first write is the expensive part because the model still has to ingest the prefix in full before anything can be reused later.

Many providers also enforce a minimum cacheable length. If the stable prefix is too short, the platform may skip caching entirely. That detail matters because a short repeated prompt may never produce meaningful savings, even if the feature exists. It also means you should not assume that any repeated prefix is worth optimizing. Size and reuse frequency both matter.

Later requests read from it

When a later request arrives with the same eligible prefix, the provider tries to find a matching cached entry. If it finds one, the platform reuses the cached prefix state and processes only the new suffix. That is the core economic model of prompt caching. You pay more for the initial write, then less for repeated reads, while also reducing latency because the model can skip part of the work.

In multi-turn conversations, this can be powerful. A long setup with tools, safety instructions, and prior history may be reused across many turns while only the newest message changes. In coding assistants, that might mean the model does not keep reprocessing the same repo guidance and tool catalog. In document workflows, it might mean the model does not keep rereading the same contract or policy pack for every follow-up question.

Matching happens on the rendered context

This is the part many teams miss. Cache matching usually depends on the full rendered request, not the simplified version you see in your application code. If the order of tools changes, if a schema changes shape, if a developer message gets edited, or if a generation setting that affects the rendered context changes, the reusable prefix may no longer match. The user can ask the same question twice and still miss the cache because the hidden setup changed.

That is why stable prompt assembly matters. Providers often treat the request as a structured sequence of components such as tools, system or developer instructions, messages, documents, and related settings. If something earlier in that sequence changes, the cacheable prefix can break. It is less about your intent and more about byte-level or structure-level consistency.

Some platforms also limit how they search for cache matches. They may inspect specific eligible boundaries rather than every possible token position, and very long conversations can make automatic reuse less reliable unless you place deliberate checkpoints. In other words, prompt caching is not magic. It is a matching system with rules, limits, and tradeoffs.

Automatic and explicit caching

Automatic mode keeps setup simple

Automatic caching is the easiest place to start. The provider decides where the cacheable prefix ends and tries to reuse the longest stable portion of the request. If your prompt structure is already clean, automatic mode can deliver real savings with very little extra engineering work. That is why it is often the right default for early experiments or straightforward chat applications.

The downside is control. Automatic mode may follow the shape of the request as it evolves, and in long or messy conversations it can drift toward content that is not actually the best shared prefix. That can reduce hit rates when several requests share a stable setup but differ in the dynamic material that comes right after it.

Explicit breakpoints give you control

Explicit caching lets you define where the reusable prefix should stop. This is useful when you know exactly which part of the prompt stays stable and which part changes often. A classic example is a grading workflow with a fixed rubric, a fixed output schema, and a different student answer on every request. The correct breakpoint is after the rubric and schema, not after the volatile answer.

Some providers also let you place more than one breakpoint. That helps when different prompt segments change at different speeds. You may have a mostly fixed global instruction set, a tool bundle that changes occasionally, and a conversation branch that changes frequently. Separate checkpoints can preserve partial reuse instead of forcing an all-or-nothing result.

Where to place the breakpoint

The rule is simple: put the breakpoint after the stable material and before the volatile material. Shared instructions should come first. Stable tools and schemas should come next. Long reference context that many requests reuse should come before the cache boundary. User-specific details, current turn content, request IDs, and timestamps should come after it.

A lot of cache failures come from violating that sequence. Teams build a prompt in the order that feels natural to the application, then wonder why the hit rate is poor. A clever prompt that rewrites its opening on every turn may feel dynamic, but it is terrible for caching. Stable prefix first. Dynamic suffix last. That one discipline fixes a surprising number of expensive mistakes.

Why cache hits fail

Changes that break the prefix

Prompt caching is usually strict. Small changes near the top of the request can invalidate reuse for everything that follows. If you add one tool, change the order of two tools, rename a function argument, or adjust a JSON schema, the provider may treat the prefix as different. The same goes for hidden instruction blocks, system text, few-shot examples, or multimodal attachments included before the breakpoint.

Tool and schema changes

Tools are a frequent source of misses because they sit early in the rendered context and tend to change quietly. A new tool description, different tool order, modified schema, or changed tool-choice setting can all change the prompt prefix enough to bust the cache. This matters in agents and coding assistants because those systems often evolve quickly. If your tool layer is unstable, the cache will be unstable too.

Settings and hidden instructions

Cache misses can also come from settings that are easy to overlook. Model choice, reasoning configuration, structured-output settings, verbosity controls, developer messages, and other provider-specific options may affect the rendered context. Even when users see the same text in the UI, the request reaching the model may no longer match the previous cached version.

Conversation rewrites and compaction

Conversation structure matters just as much. Appending new turns usually preserves the old prefix better than rewriting existing messages. If your application edits earlier messages, injects fresh timestamps into the opening block, or compacts history by rewriting prior context, you may destroy reuse. Summarization and compaction can still be worthwhile, but they change the economics. Once the prefix changes, the cache has to be rebuilt.

Common placement mistakes

The most common mistake is putting the cache boundary after content that changes every request. That guarantees cache writes on volatile material and reduces the chance that later requests can reuse anything. Another frequent problem is trying to cache content that is below the provider's minimum length. In that case, the feature may silently do nothing.

Mode changes can also matter. If you begin with automatic caching and later switch to explicit-only behavior, earlier checkpoints may not be reused the way you expect. Long conversations create their own edge cases too. Some systems only search a limited set of eligible boundaries or recent checkpoints, so a stable shared prefix from much earlier in the session may stop getting reused unless you placed a deliberate breakpoint there.

There is also a timing issue. The first request usually has to finish writing the cache before later requests can read it. If several identical first-time requests arrive at once, they can all miss and all pay the full price. That is one reason prewarming or controlled rollout matters in high-traffic workflows.

When prompt caching saves money

The write-versus-read tradeoff

Prompt caching is not free money. The first write often carries a premium because the provider has to process the prefix and create the reusable entry. The savings appear when later requests read that prefix at a lower effective cost. The more often the same long prefix is reused inside the cache lifetime, the better the economics get.

This is why prompt caching shines in repeated workflows with a large shared setup. A twelve-thousand-token support policy pack reused across hundreds of conversations is a strong candidate. A short one-off prompt used twice is not. The same logic applies to latency. If the prefix is large, skipping its reprocessing can noticeably reduce time to first token. If the prefix is tiny, users may not feel much difference.

A simple break-even way to think about it

Ignore the details of any LLM pricing comparison for a moment and think in three variables: shared prefix size, reuse count, and hit rate. Large prefix plus frequent reuse plus high hit rate is where prompt caching wins. Small prefix plus rare reuse plus unstable request assembly is where it disappoints. If you have to pad or expand the prompt just to cross a minimum cacheable threshold, run the math carefully. More tokens only help if the later reads save enough to justify the extra write cost.

A practical example makes this clearer. Imagine an evaluation system with a large fixed rubric, a fixed schema, and a new answer to grade each time. The rubric and schema are perfect cache candidates because they stay stable while the submission changes. The first request pays the write cost. The next hundred requests likely benefit. Now compare that with a real-time assistant where each request pulls fresh data, changes tools, and rewrites earlier history. Same feature, very different economics.

When it is not worth it

The key is to treat caching as a measured optimization within a FinOps for AI framework, not a default badge. If it reduces cost per successful request, improves latency on a user-visible path, and holds steady as the product scales, keep it. If it adds complexity without meaningful savings, move on.

Lifetime, routing, and prewarming

TTL and retention

Prompt caches do not live forever. Providers usually apply a time-to-live and may also expire entries after inactivity. Some support short-lived caches designed for active sessions. Others offer longer retention options at different pricing or policy terms. The right TTL depends on how often the shared prefix is reused. A short TTL is enough for dense conversation bursts. A longer TTL helps when usage is spread across a wider window.

Retention settings also matter for privacy and governance. Enterprise customers should verify how prompt caches behave under their provider's data processing terms, zero-retention options, and regional controls. If your application handles sensitive data, treat cache lifetime as part of your production design, not an afterthought.

Cache location and routing keys

Some providers store cache state on specific machines, shards, or regions. That means cache reuse is not only about matching the prefix, but also about landing on infrastructure that can access the right cached state. This is why certain platforms expose routing or prompt cache keys. A stable key can improve the odds of reuse, reduce noisy cross-user probing, and make spend attribution cleaner.

Not every provider exposes this control, and the details vary. Still, the principle is useful: stable routing improves stable economics. If your cache works beautifully in a local test and poorly in production, routing behavior may be part of the reason.

Prewarming shared context

Prewarming means creating the cache before real user traffic needs it. If your app knows that a large prompt prefix will be reused soon, you can often send a no-output or minimal-generation request first, let the provider write the cache, and then serve later requests faster and cheaper. This is especially useful for scheduled jobs, launches, daily agent sessions, and evaluation runs with predictable shared context.

Prewarming does not fix a bad prompt structure. If the supposedly shared prefix changes between the warm-up call and the live request, you still miss. It also does not always help under heavy concurrency if the first live requests arrive before the warm-up has completed. But when the context is truly stable, prewarming can smooth both latency and spend from the very first meaningful request.

Best use cases for prompt caching

Coding assistants and agents

Coding copilots, debugging agents, and workflow agents are strong prompt caching candidates because they usually carry a lot of setup. There may be persistent instructions, tool catalogs, policy constraints, repo summaries, environment context, and prior turns that remain relevant for many follow-up actions. If that shared layer stays stable, caching can reduce the repeated input burden dramatically. If the agent keeps changing tools and rewriting context, the benefit falls fast.

Document-heavy workflows

Prompt caching also works well in document analysis, contract review, internal policy Q&A, and long-context retrieval workflows. The common pattern is simple: a large body of reference material stays the same while users ask different questions about it. Instead of paying the full price to reprocess the same long document for every follow-up, the provider can reuse the stable prefix and only handle the new question.

Support, evaluation, and structured generation

Support assistants, grading systems, classifiers, and structured generation pipelines often reuse the same instruction block, few-shot examples, and output schema across many requests. That is almost ideal prompt caching territory. The setup is heavy, the suffix is short, and the workload repeats at scale. In these cases, caching is not just a technical optimization. It can materially change the cost per ticket, per evaluation, or per generated result.

Where it works poorly is just as important. Real-time data lookups with highly personalized context, one-off creative prompts, or workflows that constantly change the opening prompt are harder to optimize. The reuse has to be real, not theoretical.

How to measure prompt caching properly

The metrics that matter

Do not judge prompt caching by intuition alone. Measure cache write tokens, cache read tokens, uncached input tokens, hit rate, time to first token, and realized cost per request; converting tokens to dollars can help you quantify the impact. Total token counts by themselves can be misleading because a successful cache read may lower the amount of newly processed input while still relying on a large underlying context. Teams often track these with advanced AI usage dashboards.

You should also compare before and after on the same workflow, not across unrelated traffic. The real question is not whether caching exists. The question is whether it lowered the unit cost of a specific path that matters to your product.

Tie savings to business outcomes

This is where many teams stop too early. Engineering sees a better hit rate and calls it a win. Finance still sees a large provider invoice and asks what changed, especially when enforcing AI budget limits for product teams. To make prompt caching useful beyond infrastructure, you need attribution. Which feature got cheaper? Which workflow became faster? Which customer segment generates the most cacheable traffic? Which team is shipping prompts that hurt reuse?

That is the difference between observability and control. If you can connect cached-token savings to requests, features, customers, and teams, the optimization becomes commercially useful. It stops being API trivia and becomes margin data. For fast-moving companies building on AI, that is the level that matters in AI FinOps. Spend should be visible in real time, not reconstructed from a monthly puzzle.

A practical rollout approach

Start with one high-volume workflow that already has a large shared prefix. Keep the opening prompt stable. Decide whether automatic mode is enough or whether you need explicit breakpoints. Measure write tokens, read tokens, hit rate, latency, and cost before and after. Then test failure cases on purpose by changing tools, schemas, or early instructions so you understand what breaks reuse in your environment.

Once the path works, roll the pattern into other AI features. Keep prompts modular. Avoid unnecessary edits to the prefix. Document which settings are allowed to change without harming the cache. Most importantly, report the savings in business terms. If caching improved the economics of onboarding, support, or your copilot feature, show that clearly. That is how a good engineering choice becomes an operating advantage.

FAQs about prompt caching

What does prompt caching do?

Prompt caching reuses previously processed prompt context so the model does not have to fully reprocess the same prefix on every request. In practice, that usually means lower repeated input cost and lower latency for workflows that reuse large instructions, tool definitions, long documents, or conversation history.

Is prompt caching faster?

Usually, yes. When the provider can reuse a large cached prefix, the model has less new input to process before it starts generating. That often improves time to first token and sometimes total response time. The effect is strongest when the shared prefix is large and the cache hit rate is high.

Is prompt caching the same as KV caching?

No, but they are closely related. KV caching typically describes the lower-level reuse of internal model state for earlier tokens. Prompt caching is the feature you interact with at the API or platform level, including matching rules, cache lifetime, billing behavior, and explicit breakpoints when supported. Related mechanism, different level of abstraction.

Does prompt caching change the model's answer?

Not by itself. Prompt caching changes how the model processes repeated context, not the intent of the prompt. If the effective input is the same, caching should not alter the meaning of the request. That said, outputs can still vary because language models are not always perfectly deterministic, and any hidden change in the rendered prompt can still lead to a different result.

Can prompt caching help in multi-turn chats?

Yes, and that is one of its strongest use cases. Multi-turn chats often resend system instructions, tools, and previous turns on every request. If that earlier context remains stable, caching can reuse it across turns and reduce the repeated input burden. The benefit drops when the app rewrites earlier history or frequently changes the tool layer.

What usually causes a cache miss?

The usual causes are changes early in the prompt. That includes tool definition changes, schema edits, modified system or developer instructions, different settings that affect the rendered request, timestamps placed before the cache boundary, rewritten conversation history, or a shared prefix that is too short to qualify. Small changes can have large effects.

Does Claude Code use prompt caching?

Anthropic has supported prompt caching for supported Claude APIs and models, but whether Claude Code or any Claude-based tool uses it in your workflow depends on the client, model, and prompt assembly logic. The safe answer is to verify it in the product or API documentation and then confirm it with actual cache-read metrics rather than assumption.

Do cached tokens still count toward limits or cost?

Often, yes, but differently. Providers may bill cache writes, cache reads, and uncached input at different rates, and some still include cached context in rate-limit or context-window logic. That is why you should read the platform's token accounting fields carefully. A cache hit is cheaper than a full reprocess, but it is not always free or invisible.

Prompt caching is one of the cleanest ways to strengthen AI spend management without cutting product quality. Keep the shared prefix stable, place boundaries deliberately, monitor read-versus-write economics, and connect the result to features, customers, and teams. Do that well, and prompt caching stops being a neat model trick and becomes a real operating lever.

You deserve
financial clarity.
Get full visibility into your company’s finances in minutes.
Get started for free