Wild West API

Summarization: keeping long roleplays on track

Summarization is a memory technique where older parts of a chat are condensed into a short recap that is inserted into the prompt so the model keeps track of the story.

What summarization is

Every chat eventually outgrows what you want to send per request, either because it passes the context window or because sending all of it costs too much. Messages that fall out of the prompt are gone as far as the model is concerned. A summary keeps the gist of them: who did what, what changed, what is unresolved, in a few hundred tokens.

How it works in SillyTavern

The Summarize extension periodically asks a model to write or update a summary of the chat so far. Settings include:

  • Source: the main API you are chatting with, or a separate local model.
  • Update frequency: every N messages or every N words.
  • Summary prompt: the instruction used to write it. You can rewrite it to focus on plot events, relationships, or open threads.
  • Injection position and depth: where the summary goes in the prompt, for example before the main prompt or at a depth in the chat.
  • Manual editing: you can correct or rewrite the summary at any time.

Each update is an extra API request that sends chat history, so automatic summaries add cost.

What it does well and badly

Summaries are good at the shape of a story: major events, current goals, relationship status. They are bad at specifics: the exact wording of a promise, the color of a cloak, a running joke. They can also introduce errors, and once an error is in the summary, the model treats it as fact. A model writing summaries of its own story may also smooth over or omit things it handles poorly.

Good practice

  • Read the summary now and then and fix mistakes by hand.
  • Use a structured summary prompt: characters and status, location, key events, open threads.
  • Keep it short. A 2,000-token summary partly defeats the point.
  • Pair it with a lorebook for fixed facts and vector memory for recalling specific old passages.
  • Write your own summary at the start of a new chat when continuing a long story.

Summarization with Wild West API

With 512K to 1M token context windows, many chats never need summaries for space. They can still pay off on cost: a summary plus the last 30K tokens of chat is far cheaper per turn than resending 400K tokens. If you use the main API as the summary source, each summary run is a normal billed request.

FAQ

Does summarization cost extra?

If the summary is generated with your main API, yes. Each update is a separate request that sends part of the chat history.

Why does the summary get facts wrong?

Summaries are written by a model and can drop or invent details. Edit the summary directly when you spot an error, since the model will trust it.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.