// specific build playbooksby updated 6 min read

Anthropic Prompt Caching: How the Pricing Works and What It Saves

Prompt caching bills repeated input at a fraction of the normal price, and what it saves on a real bill depends on how much of each prompt repeats and how often. How the pricing works, what breaks the cache, and how to estimate savings from your own traffic.

Prompt caching lets the Claude API reuse the start of a prompt it has already processed, and it bills the reused tokens at a fraction of the normal input price. The discount applies only to the cached part of the input, and only on requests that find the cache warm. What it saves on a real bill depends on three things: how much of each prompt repeats, how often requests arrive before the cache expires, and how much of the bill is output.

Prices, multipliers and limits are according to Anthropic's prompt caching documentation as of October 2026. They vary by model and they change, so check the current page before you budget.

How it works

You mark a point in the prompt with cache_control. The API caches everything up to and including that block, in a fixed order: tool definitions first, then the system prompt, then the messages. On the next request, if everything up to that point is byte-for-byte identical, the API reads it from the cache instead of processing it again.

Because the match is exact, order matters. Stable content goes first and anything that changes per request goes last. A timestamp near the top of a system prompt means nothing after it ever matches.

You can set up to four breakpoints in one request, so sections that change at different rates can cache separately: tools that never change, reference documents that change daily, a conversation that grows every turn. The breakpoints themselves cost nothing.

There's also an automatic mode, a single cache_control field at the top level of the request, which places the breakpoint on the last block it can cache. It suits multi-turn conversations. It works against you when every request ends with something unique, such as a retrieved record or a one-off question, because each request then writes a cache entry that nothing reads back. In that case put an explicit breakpoint at the end of the shared part.

What it costs

According to Anthropic's documentation, cache pricing is set as multiples of each model's base input price:

  • A cache write with the default five-minute lifetime costs 1.25 times the base input price.
  • A cache write with the one-hour lifetime costs twice the base input price.
  • A cache read costs a tenth of the base input price on most models. A few of the newest models price reads lower still.
  • Tokens after the last breakpoint cost the normal input price.
  • Output tokens are billed as usual. Caching only changes the price of input.

Each read also resets the cache's lifetime at no extra charge, so a prefix used at least every five minutes stays cached for as long as the traffic keeps coming. The lifetime counts from the start of a request, and generation time counts against it. If a response takes four minutes to stream, the next request has about a minute to start. On the Claude API, cache reads also don't count against input-token rate limits on most models, so caching can raise throughput as well as lower cost.

The break-even comes quickly. With the five-minute cache, one write and one read cost 1.35 times the base price, against 2 times for sending the same prefix twice uncached, so caching is ahead from the second request. With the one-hour cache, the higher write price needs a third request to come out ahead: 2.2 times against 3.

Estimate the savings on your own traffic

You need four numbers, from your logs or a sample of requests:

  1. The share of each prompt that's a stable prefix: tools, system prompt, fixed documents and examples.
  2. The hit rate, meaning the share of requests that find the prefix already cached.
  3. Average input and output tokens per request.
  4. The base input and output prices for your model.

Then work out the input cost of a request three ways: uncached, on a hit, and on a miss. Here's the arithmetic with made-up round numbers. A request has 10,000 input tokens, 9,000 of them a stable prefix. Nine requests in ten find the cache warm, and reads cost a tenth of the input price. Counting one uncached input token as one unit:

  • Uncached: 10,000 units.
  • A hit: 9,000 × 0.1 + 1,000 = 1,900 units.
  • A miss, which writes the cache: 9,000 × 1.25 + 1,000 = 12,250 units.
  • Blended at nine hits in ten: 0.9 × 1,900 + 0.1 × 12,250 = 2,935 units.

In that example, input cost falls to under a third of the uncached figure. Output stays the same, so if output is a large share of your bill, the total falls by much less than the input line does. Run the same arithmetic with your own four numbers before you promise anyone a saving.

Five minutes or one hour

Choose by the gap between requests that share a prefix, measured start to start.

  • Under five minutes: use the default. Every read refreshes it for free, and the one-hour write price buys nothing.
  • Between five minutes and an hour: this is where the one-hour cache is worth its higher write price. Typical cases are a user who replies after twenty minutes, or a side task that runs longer than five minutes between reads.
  • Over an hour: neither one helps directly. Accept the cold start, or warm the cache on a schedule.

To warm the cache, Anthropic documents sending the request with max_tokens set to 0. The API writes the cache and returns without generating anything, so you pay for the cache write and no output tokens. Each warm-up is still a write, so warming many prefixes on speculation can cost more than it saves.

What breaks it

A broken cache fails quietly. Requests keep succeeding and the bill is just higher. Common causes:

  • A changing value near the top. A timestamp, request ID or user name in the system prompt changes the prefix on every request.
  • Tool definitions that change. Tools come first, so editing, adding or reordering them invalidates the whole cache.
  • Unstable serialization. If your code builds JSON with keys in a different order from one request to the next, the bytes differ even when the content is the same.
  • Settings that are part of the prompt. Turning web search or citations on or off changes the system prompt. Adding or removing images, or changing tool_choice or the thinking settings, invalidates the cached messages.
  • A prefix under the minimum. Each model has a minimum cacheable length, from 512 tokens on the newest models up to 4,096 on some others. Below it, nothing is cached and no error comes back.
  • Too many prompt variants. If every customer or every repository gets its own system prompt, each one is a separate cache entry, and a variant used less often than the cache lifetime misses nearly every time. Put the shared instructions first, cache them, and put the per-customer part after the breakpoint.
  • Separate workspaces. On the Claude API, caches are kept per workspace, so the same prompt sent from two workspaces is cached twice.

How to measure it

Every response reports three numbers in its usage block:

  • cache_creation_input_tokens: tokens written to the cache on this request.
  • cache_read_input_tokens: tokens read from the cache.
  • input_tokens: tokens after the last breakpoint, billed at the normal price.

Total input is the sum of the three. When caching works, input_tokens on its own looks small, so don't read it as the size of the prompt. If both cache fields are 0, nothing was cached, often because the prefix is under the minimum. If writes stay high on every request while reads stay low, something in the prefix is changing between requests. Compare the raw bytes of two consecutive requests to find it.

Log the three numbers per request and chart the hit rate by day. A drop usually follows a change to the code that builds prompts. Add a test that sends the same request twice and checks that the second one reads from the cache, so a regression gets caught before it reaches the bill.

What to do first

One, find the stable part of your prompts: tool definitions, the system prompt, reference documents and examples. Move anything that changes per request to the end.

Two, put a breakpoint at the end of the stable part, or turn on automatic caching for multi-turn conversations.

Three, check cache_read_input_tokens on a second identical request. If it's 0, fix that before anything else.

Four, log the usage fields and work out your real hit rate. Then decide whether the one-hour cache is worth it from the gaps in your own traffic.

anthropicprompt cachingcostclaudelong-tail
// go deeper

Worked steps, FAQs and resources on this topic.

read the guide
// the_list

Get the next one.

Field notes from real deployments, written the way this one was. No schedule promised and nothing sold to you. Unsubscribe whenever.

// keep reading

Related posts

// ready to ship?

Let's build yours.

Reading is the easy part. We do the work. Tell us what's broken and we'll tell you straight up whether we can help.