Fixing replies cut short with finish_reason "length"
A reply that ends with finish_reason "length" stopped because it reached its token limit. The question is where that limit came from.
Where the limit comes from
The API sends the model a max_tokens on every request, chosen like this:
- If you set
max_tokens, that is the starting point. - If you did not, the API uses the room left in the model's context window after your prompt, minus a small margin, held under a default ceiling of 32,768 tokens.
- Then it lowers that number to what your balance and key caps can pay for. This is the part people miss: with a low balance, replies get shorter with no error, and they end on
finish_reason: "length".
Because the limit is set before generation, a reply never runs past what you can afford and never gets cut off mid-stream by billing. It is just shorter.
How it shows up
In SillyTavern, the reply stops mid-thought, and Continue picks it up. In code, choices[0].finish_reason is "length". On the reasoning models (outlaw-1, glm-5.3-outlaw, glm-5.3-flash-outlaw) thinking counts toward the same limit, so the visible part can be much shorter than the number suggests.
Step by step fix
- Check your balance and any key cap in the dashboard. If it is low, top up; full-length replies come back immediately.
- Raise Max Response Length in SillyTavern, or
max_tokensin code. See max tokens. - For reasoning models, allow extra room for thinking.
- Use Continue in SillyTavern for very long scenes rather than one huge reply.
Telling it apart from similar errors
A stop with no finish_reason, or a dropped connection, is a stream cut off, not a length cap. When the balance cannot pay for even a minimal reply, the request is refused with a 402 instead; see spend limit reached. An empty reply with length on a reasoning model is covered in empty response.
FAQ
Will I get a warning when my balance starts shortening replies?
No error is raised. The reply is simply shorter and ends on finish_reason length. Keep an eye on the balance in the dashboard.
If I set no max_tokens, how long can a reply be?
Up to the room left in the context window after the prompt, capped at the default ceiling, and then by what the balance can pay for.