What Prompt Caching Means When Your AI Coding Bill Arrives

What Prompt Caching Actually Does

Cache Reads Cost One Tenth of Fresh Input

The Rate That Matters
  • ● Cache hit billed at 0.1x base input
  • ● Opus 5 input $5 falls to $0.50 per MTok
  • ● A write costs 1.25x before it saves

Turn caching on if your assistant resends the same repository context every turn, and leave it off for one-shot questions. A cache hit on the Claude API is billed at 0.1x the base input rate, so Claude Opus 5 input drops from $5 to $0.50 per million tokens as of Aug 2026. OpenAI publishes the same shape, with cached input on gpt-5.5 at $0.50 against $5.00 standard.

The catch is that storing the cache is not free. A write costs more than ordinary input, so caching only pays once the same content is read back.

That single sentence explains most surprising bills. Developers see the 90 percent discount advertised and assume it applies to the whole request, then wonder why the invoice barely moved.

What Actually Sits in the Cache

Caching works on a prefix, not on ideas. The provider hashes the front of your prompt and reuses the processed form when the next request starts with exactly the same bytes.

In a coding session that prefix is usually three things stacked in order. The system instructions, the tool definitions your client sends, and whatever project files the assistant loaded at the start.

Everything after the last cached breakpoint is fresh input at the ordinary price. Your new question, the diff you just pasted, and the tool results from this turn all land outside the cache.

Order matters more than size. One changed character near the top invalidates every cached token below it, which is why a rotating timestamp in a system prompt can quietly destroy the whole saving.

The Write Premium and When It Pays Back

Anthropic publishes three prices per model rather than one. Base input, a cache write, and a cache read, and the write is the one people forget.

A 5-minute cache write is billed at 1.25x the base input price, a 1-hour write at 2x, and a read at 0.1x. The documentation states the break-even plainly, since caching pays off after one cache read on the 5-minute duration and after two reads on the 1-hour duration.

So a single question against a large file is worse with caching than without it. You pay 1.25x to store context you never read again.

An interactive session is the opposite case. Ten turns against the same loaded files pay one write and nine reads, which is where the headline discount finally shows up on the invoice.

Published Cache Rates Side by Side

Three Prices, Not One
  • ● 5 minute write 1.25x base input
  • ● 1 hour write 2x base input
  • ● Cache hit 0.1x base input

These are list prices from the official pricing pages, in US dollars per million tokens, as of Aug 2026. Confirm current pricing on the official site before you budget against them, because model rates move.

Model Base input 5m cache write 1h cache write Cache hit
Claude Opus 5 $5.00 $6.25 $10.00 $0.50
Claude Sonnet 5 $2.00 $2.50 $4.00 $0.20
Claude Haiku 4.5 $1.00 $1.25 $2.00 $0.10
Claude Fable 5 $10.00 $12.50 $20.00 $1.00
OpenAI gpt-5.5 $5.00 not published separately not published separately $0.50
OpenAI gpt-4o $2.50 not published separately not published separately $1.25

Two things stand out in that table. The Claude cache read is a flat tenth of base input across every current model, and the OpenAI discount is not flat at all.

The gpt-4o row is the warning. Its cached input is half the standard price rather than a tenth, so a caching strategy tuned on a newer model does not transfer backwards to an older one.

The Five Minute Window Is Shorter Than a Coffee Break

The default Claude cache entry lives 5 minutes, and every hit refreshes it for another 5 minutes. That refresh is the useful part, because an active session keeps its own cache alive at no extra cost.

Stop typing and the window closes. A build that runs 8 minutes, a meeting, or a long read of the assistant output all end the entry, and the next request pays for the write again.

The 1-hour option exists for exactly this pattern, at 2x base input for the write. It suits a day of work against one large codebase where gaps between prompts are measured in tens of minutes.

Neither duration is a subscription. You are billed per write, so an hour cache written five times in a morning costs five writes rather than one.

What the Discount Looks Like in a Real Session

Anthropic publishes a worked example that is easier to reason about than the multipliers. A one-hour session on Opus 5 consuming 50,000 input tokens and 15,000 output tokens costs $0.25 for input and $0.375 for output.

Run the same session with 40,000 of those input tokens served from cache, and the input side splits in two. The 10,000 fresh tokens cost $0.05, the 40,000 cache reads cost $0.02, and the input line falls from $0.25 to $0.07.

That is a 72 percent cut on input and roughly a quarter off the whole session once output is counted. Output tokens never cache, which sets a hard floor on how much caching can save you.

The floor is the number worth internalising. If your work is generation heavy rather than context heavy, caching moves the smaller half of the bill, a point our breakdown of per-seat against usage-based pricing applies to the plan you sit on.

Where the Cache Quietly Breaks

Four patterns account for most vanished discounts, and none of them raise an error. Your requests keep working and simply cost more.

Editing near the top of the context is the first. Reordering loaded files, or letting the client inject the current time into the system prompt, changes the prefix and invalidates everything after it.

Changing tools is the second. Tool definitions sit in the prefix on most clients, so adding one server mid-session rewrites the cache for every later request.

Switching models is the third, since a cache belongs to one model. The fourth is simple idleness, which the 5-minute window punishes harder than anyone expects.

How the Two Providers Frame It Differently

Anthropic makes caching explicit. You add a cache control field, either once at the top level for automatic breakpoint management or on individual blocks for fine control, and the usage response itemises write and read tokens.

OpenAI applies caching automatically on supported models with no write premium published. The trade is that the discount varies by model, at 90 percent on gpt-5.5 and 50 percent on gpt-4o.

Neither approach is better in the abstract. Explicit control rewards a client that knows which part of the prompt is stable, and automatic caching rewards everyone else.

What both share is the prefix rule. Whatever the billing label, the saving only arrives when the front of the request is byte-identical to last time.

Which Caching Setup Fits Your Workflow

Match Cache to Session
  • ● Short edits - 5 minute cache
  • ● All-day refactor - 1 hour cache
  • ● Batch jobs - stack both discounts

Quick one-off question against a single file: Skip caching. You would pay a 1.25x write for context that is never read back, which makes the request more expensive rather than less.

Interactive session on one codebase for an hour: The 5-minute cache, left to refresh itself. Continuous prompting keeps the entry alive and every turn after the first is billed at 0.1x on the shared prefix.

All-day refactor with long gaps between prompts: The 1-hour cache at 2x write. It needs two reads to break even, and a full day of returning to the same context clears that easily.

Automated review or migration jobs: Stack the discounts. The Batch API takes 50 percent off input and output, and Anthropic states that batch and caching multipliers combine.

Team running a shared assistant on a subscription: Measure before optimising. Plans such as Claude Pro at $20 per month meter usage rather than tokens. Caching then buys more work inside the cap rather than a smaller bill, a distinction our notes on Claude Code pricing set out in detail.

Three Habits That Waste It

Reloading the project on every prompt is the most common. If your client rebuilds context from scratch each turn, it writes a new cache each turn and never reads one, so you pay 1.25x for the privilege.

Sprinkling volatile text through the prefix is the second. Timestamps, request identifiers, and a git hash injected into the system prompt each guarantee a miss.

Model hopping is the third. Bouncing between a cheap model for edits and an expensive one for reasoning is often correct on quality grounds, and it does mean each model keeps its own cache and its own write cost.

Fixing all three is usually a client configuration change rather than a code change. Look at what your assistant sends before the first user message, since that block is the one worth stabilising.

What to Check Before You Blame the Model

Start with the usage numbers rather than the invoice total. Both APIs report cache creation and cache read tokens per request, and the ratio between them tells you immediately whether the cache is working.

A healthy interactive session shows one large write followed by many reads. A broken one shows a write on every request and reads near zero.

If reads are missing, compare two consecutive raw requests byte for byte from the top. The first difference is your invalidation point, and it is almost always something the client inserted rather than something you typed.

Then decide whether the fix is worth it. Caching is only a lever on repeated context. If your sessions are short, varied, and generation heavy, a cheaper model or a smaller loaded context will save more. Our comparison of local against cloud assistants works through that same trade from the other direction.

Loading less context is the other half of that trade, and it only works if you know what the window is actually holding. What context window size really means for coding work walks through where the number changes a task and where it only changes the invoice.

A cache miss shows up as latency before it ever shows up on an invoice. Long sessions are why an AI coding assistant slows down covers the mid-afternoon lag from the other side, where the fix is a fresh session rather than a cheaper prefix.

FAQ

Does prompt caching make every request cheaper?

Only the repeated part. Caching stores a prefix of the prompt, so the system instructions, tool definitions, and files you keep resending get the discount. Anything new in that turn is billed at the ordinary input rate.

How much cheaper is a cache hit?

A cache hit is billed at 0.1x the base input rate on the Claude API, so Opus 5 input falls from $5 to $0.50 per million tokens as of Aug 2026. OpenAI lists cached input on gpt-5.5 at $0.50 against $5.00 standard, the same 90 percent cut. Confirm current pricing on the official site.

Why did my cost jump after a short break?

Almost certainly the cache expired. The default Claude cache entry lives 5 minutes and each hit refreshes it, so a lunch break or a long build ends the window and the next request pays for a fresh write.

Is the one-hour cache worth the higher write price?

A one-hour cache write is billed at 2x the base input rate rather than 1.25x, which means it needs at least two later reads before it saves money. It suits a long session against one large codebase, not a quick edit.

Do subscription plans benefit from caching?

Not directly. Subscription plans such as Claude Pro at $20 per month bill by usage allowance rather than per token, so caching shows up as more work inside the same limit rather than as a smaller invoice.

Sources


Some links may be affiliate links. We may earn a commission at no extra cost to you.

This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.

Comments