Wild West API

Fixing replies cut short with finish_reason "length"

A reply that ends with finish_reason "length" stopped because it reached its token limit. The question is where that limit came from.

Where the limit comes from

The API sends the model a max_tokens on every request, chosen like this:

  1. If you set max_tokens, that is the starting point.
  2. If you did not, the API uses the room left in the model's context window after your prompt, minus a small margin, held under a default ceiling of 32,768 tokens.
  3. Then it lowers that number to what your balance and key caps can pay for. This is the part people miss: with a low balance, replies get shorter with no error, and they end on finish_reason: "length".

Because the limit is set before generation, a reply never runs past what you can afford and never gets cut off mid-stream by billing. It is just shorter.

How it shows up

In SillyTavern, the reply stops mid-thought, and Continue picks it up. In code, choices[0].finish_reason is "length". On the reasoning models (outlaw-1, glm-5.3-outlaw, glm-5.3-flash-outlaw) thinking counts toward the same limit, so the visible part can be much shorter than the number suggests.

Step by step fix

  1. Check your balance and any key cap in the dashboard. If it is low, top up; full-length replies come back immediately.
  2. Raise Max Response Length in SillyTavern, or max_tokens in code. See max tokens.
  3. For reasoning models, allow extra room for thinking.
  4. Use Continue in SillyTavern for very long scenes rather than one huge reply.

Telling it apart from similar errors

A stop with no finish_reason, or a dropped connection, is a stream cut off, not a length cap. When the balance cannot pay for even a minimal reply, the request is refused with a 402 instead; see spend limit reached. An empty reply with length on a reasoning model is covered in empty response.

FAQ

Will I get a warning when my balance starts shortening replies?

No error is raised. The reply is simply shorter and ends on finish_reason length. Keep an eye on the balance in the dashboard.

If I set no max_tokens, how long can a reply be?

Up to the room left in the context window after the prompt, capped at the default ceiling, and then by what the balance can pay for.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.