Max tokens: the cap on response length
Max tokens is the upper limit on how many tokens the model is allowed to generate in a single response.
What max tokens is
max_tokens is a ceiling on output length. The model stops when it decides it is done or when it reaches this number, whichever comes first. It is a limit, not a target: setting 1000 does not make the model write 1000 tokens. To get longer or shorter replies you have to ask for that in the prompt or show it in the first message and examples.
How it relates to the context window
The context window has to hold the prompt and the reply together. If a model has a 128K window and your prompt uses 126K, it only has room for about 2K tokens of output, whatever max tokens says. Frontends like SillyTavern reserve the max tokens amount when they decide how much chat history to include, so raising "Max Response Length" also shrinks the history sent each turn.
On models with 1M token windows this squeeze is rarely the issue, but cost and speed still scale with what you send.
Reasoning models and cut-off replies
On reasoning models, thinking tokens are generated output too, and they usually count toward the cap. A low max tokens can be used up entirely by thinking, leaving an empty or truncated answer. If you use a model that thinks, give it a much higher limit than you would for a direct reply.
When a response hits the limit, the API reports a finish reason of length (OpenAI style) or max_tokens (Anthropic style). That is how you tell a cut-off reply from a finished one. SillyTavern has a "Continue" button to extend a cut-off message.
Sensible values for roleplay
- 200 to 400: short, chatty turns.
- 500 to 1000: typical narrative roleplay.
- 2000 or more: long-form scenes, or any reasoning model.
A common mistake is setting max tokens very low to force short replies. The model does not know about the cap, so it writes a long reply and gets cut mid-sentence. SillyTavern's "trim incomplete sentences" option hides this, but asking for shorter replies in the prompt is the real fix.
Max tokens on Wild West API
max_tokens is a standard OpenAI parameter. Newer OpenAI docs also use max_completion_tokens, but most OpenAI-compatible APIs accept max_tokens. On the Anthropic-compatible /v1/messages endpoint, max_tokens is a required field. Output tokens are billed, so the cap is also a cost ceiling per request.
FAQ
Does max tokens make the model write longer replies?
No. It only sets a maximum. Ask for the length you want in the prompt, and set max tokens comfortably above it.
Why is my reply cut off mid-sentence?
The reply reached max tokens. Raise the limit, use Continue in SillyTavern, or ask for shorter replies in the prompt.