Fixing a streamed reply that stops mid-sentence
A stream that stops mid-sentence either hit its token limit, lost its connection, or stalled. Each one leaves different clues.
First check: did it end on length?
Most cut-off replies are not errors. If the final chunk carries finish_reason: "length", the reply reached max_tokens. That limit is either your Max Response Length, or a lower number the API chose because your balance or key cap could only afford that much. See replies cut short.
How the API handles a live stream
- Keepalives. While the model is quiet (for example while it reasons), the API sends an SSE comment line every 15 seconds. OpenAI-compatible parsers ignore it, and it keeps proxies and clients from closing an idle connection.
- Stall cutoff. If the model sends nothing for two minutes, the API closes the stream. It is measured from the last chunk, so a long healthy reply is never cut by this.
- You hang up. Pressing Stop, closing the tab or a client timeout cancels the request, and generation upstream stops as well.
- The upstream stream breaks. A broken connection from the model mid-reply ends your stream too.
What gets billed
You pay for what was actually generated and streamed. On a Stop or disconnect, the request is billed from the output streamed so far. If the model stalls and sends nothing at all, it is not billed. The worst-case hold placed before the call is released when the request settles.
Step by step fix
- Check
finish_reason. In SillyTavern, a reply that ends cleanly at the limit is a length cap; raise Max Response Length or use Continue. - Check your balance in the dashboard. A low balance shortens replies silently.
- If it stops at random points with no finish reason, look at what sits between you and the API: VPNs, corporate proxies or mobile connections that drop long requests.
- Make sure any proxy in front of SillyTavern allows long-lived streaming responses and does not buffer them.
- Retry with a short message to see whether it is size-related.
Background on streaming.
Telling it apart from similar errors
A request that fails before any text appears is a status error such as 503 or 502, not a cut-off. A non-streaming request that never returns is covered in request timeout.
FAQ
Am I charged for a reply I stopped?
Yes, for the part that streamed before you stopped. Generation upstream is cancelled when you disconnect, so it does not keep running.
Why do I see lines starting with a colon in the raw stream?
Those are SSE comment keepalives sent every 15 seconds during silence. Standard parsers ignore them.