Wild West API

Context window: how much an LLM can see at once

The context window is the maximum number of tokens a model can take into account in one request, covering both the input and the reply.

What the context window is

A model has no memory between requests. Everything it knows about your story at a given moment is whatever text is in the current request: the system prompt, the character card, any lorebook entries, the chat history, and its own reply as it writes it. The context window is the size limit on all of that combined, measured in tokens.

How it works

Transformers process all tokens in the window together through attention, which is how a later token can refer back to an earlier one. Positional encoding schemes set how far the model was trained to look, and methods such as RoPE scaling extend that range. Running past the window is not possible: the API rejects the request or the frontend drops older text first.

Frontends like SillyTavern build each prompt from scratch. They add the fixed parts, reserve space for max tokens, then fill the remaining budget with as much recent chat history as fits. Older messages simply fall out. That is why a long roleplay "forgets" early events, and why tools like summarization and vector memory exist.

Why bigger is not free

  • Cost. The whole context is resent and billed as input every turn. A 200K token chat costs 200K input tokens per reply and per swipe.
  • Speed. Long prompts take longer to process before the first token appears. Prompt caching can help when the start of the prompt stays the same.
  • Attention quality. A model that accepts 1M tokens does not use every token equally well. Details buried in the middle of a long context are recalled less reliably than recent text or the very start.

Sensible settings

Set your frontend's context size to no more than the model's limit. For most roleplay, 16K to 64K tokens of history is plenty, and it keeps cost and latency low. Go larger when the story truly needs it, for instance a long campaign where you want the model to see earlier chapters verbatim. Keep important facts in places that are always sent, such as the card or a constant lorebook entry, rather than trusting that they are still in history.

Context windows on Wild West API

outlaw-1, glm-5.3-outlaw, glm-5.3-flash-outlaw and mimo-v2.6-flash-outlaw accept up to 1M tokens, and qwen3.8-27b-outlaw accepts 512K. See the model list. In SillyTavern, raise the Context Size slider (and unlock it if needed) to use more than the default, but remember every token sent is billed.

FAQ

Does a bigger context window mean the model remembers everything?

It can see more text, but recall of details deep in a long context is weaker than recall of recent text. Keep key facts in the card or lorebook.

Does the context window include the reply?

Yes. Prompt tokens plus generated tokens must fit inside the window.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.