Wild West API

Fixing "maximum context length" (context_length_exceeded)

This 400 means the prompt plus the reply you reserved add up to more tokens than the model can hold. It is about the request, not an outage, so retrying unchanged will fail again.

What the error says

HTTP 400
{
  "error": {
    "message": "This model's maximum context length is 524288 tokens. However, you requested 530112 tokens. Please reduce the length of the messages.",
    "type": "invalid_request_error",
    "code": "context_length_exceeded",
    "param": "messages"
  }
}

The numbers come from the model's own refusal. On /v1/messages the same case reads prompt is too long: N tokens > M maximum, the wording Claude Code looks for to compact a conversation.

Limits per model

qwen3.8-27b-outlaw has a 512K context window. The other models (outlaw-1, glm-5.3-outlaw, glm-5.3-flash-outlaw, mimo-v2.6-flash-outlaw) have about 1M. The window has to hold the prompt and the reply together.

How the API sizes the reply

If you send no max_tokens, the API picks one that fits: the window minus the estimated prompt, minus a small safety margin (2 percent of the window, at least 64 tokens), capped at a default ceiling. So requests without max_tokens rarely hit this error. If you do send max_tokens, it is used as given (only lowered when your balance cannot afford it), so a large reply reservation on top of a large prompt can overrun the window.

Step by step fix in SillyTavern

  1. Open the AI Response Configuration panel.
  2. Set Context Size to no more than the model's window. If you raised it past the default, the Unlocked context size option allows it, so check it carefully.
  3. Keep Max Response Length reasonable. Context Size should be at least prompt plus response, so lower one if the other is large.
  4. Trim big lorebook or World Info budgets and long example dialogue if the chat is huge.
  5. Retry.

In code, either drop max_tokens and let the API size it, or compute max_tokens from the window minus your prompt. See max tokens.

Telling it apart from similar errors

Other upstream failures are reported as a 503 Model overloaded, which is worth retrying; this one is a 400 and is not. A reply that stops with finish_reason: "length" did not hit the window, it hit max_tokens; see replies cut short.

FAQ

Am I charged when this happens?

No output is produced. The model refuses before generating, so the request settles without a charge for output.

Why does it happen at a smaller size than the slider shows?

Token counts differ between tokenizers, and the chat template adds tokens around each message. Leave some slack below the advertised window.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.