Skip to content

Blog · September 24, 2026

Updated · By , AI-assisted.

Prompt caching, explained

Prompt caching lets an AI provider reuse its work on the start of a prompt it has just read, and charge a fraction of the input price for it. A chat resends the whole conversation with every message, so caching keeps long chats cheaper to serve. In llmwise it never changes what a message costs you.

What prompt caching is

A model doesn't remember your chat between messages. Each time you send one, the whole conversation goes back to it: the instructions, the tools it may use, every earlier message, and your new one. The start of that prompt is the same as last time, with a little more added at the end.

Prompt caching lets the provider keep its work on that repeated start for a short while. When the next request begins with exactly the same text, it reuses the work instead of reading it again. Those tokens then cost a fraction of the normal input price.

Why it matters in a chat

In a chat, input adds up fast. A typical message in our pricing is 4,000 tokens in and 700 out, because the model rereads the conversation each time. The longer the chat, the more of each message is repeated text. Caching is what keeps a long conversation from costing a full re-read on every message.

What does a 10-message chat send? A worked example

Take a chat of 10 questions of 300 tokens, each answered in 700. Every message sends the whole chat so far, so the prompt grows by 1,000 tokens a turn.

Tokens sent to the model, message by message
MessageTokens sentEarlier messages among them
13000
21,3001,000
54,3004,000
109,3009,000

Across the chat that's 48,000 input tokens, and 45,000 of them (94%) are earlier messages read again. On Claude Sonnet 5, at $2.00 per million input tokens, the input costs $0.0960 without a cache.

With Anthropic's five-minute cache, writing to the cache costs 25% more than plain input and reading from it 10% of the price, per its Anthropic's prompt caching docs (read October 1, 2026). Prompts under 1,024 tokens aren't cached, so the first message is read in full. The same chat's input then costs about $0.0315: 33% of the uncached price. The replies cost the same either way.

What breaks a cache

  • Changing the beginning

    Editing an early message, or changing the instructions at the top, changes the start of the prompt. Everything after the change is new to the provider: OpenAI's prompt caching guide says reuse needs the whole prefix to match.

  • Switching models

    Each model keeps its own cache. The first message after a switch is read in full; the ones after it can be cached again.

  • Waiting

    Providers keep a cache for a short time, often minutes: five by default in Anthropic's prompt caching docs, renewed each time it's used. Come back to a chat the next day and the first message is read in full.

  • Short prompts

    Providers only cache a prompt past a minimum length, so a short chat is simply read in full each time. It's cheap anyway.

How llmwise uses it

llmwise keeps the start of each chat's prompt the same from one message to the next, so the model's provider can cache it. The instructions hold only the chat's own settings, and the tools go in a fixed order. What changes per message goes at the end.

You don't have to do anything, and it never changes your price. Each message has a fixed price shown before you send it; caching lowers what serving your chat costs us, not what you're charged.

What you can do: keep chats lean

  • Start a new chat for a new topic. Once a chat sent to the model passes 64k tokens, each reply counts as 2 (as 4 past 128k), and the message box says so: “Long chat — each reply counts as 2. Start a new chat to save messages.”

  • Attach the pages you need rather than the whole archive: attached files are part of the conversation the model rereads, and a lean chat stays under that limit longer.

  • Start on an everyday model, and move up when an answer isn't good enough.

More in LLM cost optimization.

Bar chart: Tokens sent with each message of one chat. Message 1: 300 tokens; Message 2: 1,300 tokens; Message 3: 2,300 tokens; Message 4: 3,300 tokens; Message 5: 4,300 tokens; Message 6: 5,300 tokens; Message 7: 6,300 tokens; Message 8: 7,300 tokens; Message 9: 8,300 tokens; Message 10: 9,300 tokens.
A chat of 300-token questions and 700-token answers: each message sends the whole chat so far.

Prompt caching: questions

Does prompt caching change what I pay in llmwise?

No. A message's price is fixed before you send it. Caching makes a chat cheaper for us to serve; it never changes what you're charged or what you see.

Does a cached prompt change the answer?

No. It only saves the provider from re-reading text it has just read: the model gets the same prompt either way.

A price per message that caching can't move

In llmwise a message counts the same whether or not the provider cached your prompt, so there's nothing to tune.