Prompt caching: reusing the start of a prompt
Prompt caching is a feature where a model server saves the computed state of a prompt's opening section so later requests that start with the same text can skip that work.
What prompt caching is
Before a model writes its first token, it has to process every token of the prompt. For a long roleplay, that is most of the work. But consecutive requests share almost all of their prompt: the same system prompt, the same card, the same history, plus one new message. Prompt caching keeps the processed state of that shared opening part, so the next request only processes what is new. Providers that offer it often bill cached input tokens at a lower rate and return the first token sooner.
How it works
While processing a prompt, a transformer computes key and value vectors for every token at every layer, the KV cache. The cache for a token depends on all tokens before it, so it can only be reused for an identical prefix. If the first 50,000 tokens of a new request are byte-for-byte the same as a recent one, the server can load their KV cache and start from token 50,001. Change one token near the start, and everything after it must be recomputed.
Some APIs cache automatically for long enough prompts. Anthropic-style APIs use explicit cache_control markers on content blocks. Caches expire after a short idle period, typically minutes.
What breaks it in roleplay
- Lorebook entries near the top that change from turn to turn.
- Macros that change every request, such as the current time or random picks, in the system prompt.
- Context trimming. When old messages drop off the front of the history, the prefix after the fixed blocks changes every turn.
- Author's note and summary insertion at a depth inside the history shifts every message after it.
- Vector memory inserting different retrieved text each turn.
Put stable content first and dynamic content as late as possible. Swipes and Continue are ideal for caching because the prompt is identical.
Common mistakes
Assuming caching applies on every model and provider; assuming the cache lasts indefinitely; and measuring savings on short prompts, where many providers do not cache at all because there is a minimum length.
Prompt caching on Wild West API
Whether a request benefits from caching, and how cached tokens are billed, depends on the model and its upstream. Check the docs and the usage object in responses, which shows cached token counts when they are reported. Structuring prompts with stable content first costs nothing and helps wherever caching is available.
FAQ
Why is my prompt not being cached?
Something near the start of the prompt changes each request, such as a lorebook entry, a timestamp macro, or trimmed history, or the prompt is below the provider's minimum length.
Does prompt caching change the model's output?
No. It reuses identical computation, so the output distribution is the same as without caching.