Wild West API

Fixing a streamed reply that stops mid-sentence

A stream that stops mid-sentence either hit its token limit, lost its connection, or stalled. Each one leaves different clues.

First check: did it end on length?

Most cut-off replies are not errors. If the final chunk carries finish_reason: "length", the reply reached max_tokens. That limit is either your Max Response Length, or a lower number the API chose because your balance or key cap could only afford that much. See replies cut short.

How the API handles a live stream

  • Keepalives. While the model is quiet (for example while it reasons), the API sends an SSE comment line every 15 seconds. OpenAI-compatible parsers ignore it, and it keeps proxies and clients from closing an idle connection.
  • Stall cutoff. If the model sends nothing for two minutes, the API closes the stream. It is measured from the last chunk, so a long healthy reply is never cut by this.
  • You hang up. Pressing Stop, closing the tab or a client timeout cancels the request, and generation upstream stops as well.
  • The upstream stream breaks. A broken connection from the model mid-reply ends your stream too.

What gets billed

You pay for what was actually generated and streamed. On a Stop or disconnect, the request is billed from the output streamed so far. If the model stalls and sends nothing at all, it is not billed. The worst-case hold placed before the call is released when the request settles.

Step by step fix

  1. Check finish_reason. In SillyTavern, a reply that ends cleanly at the limit is a length cap; raise Max Response Length or use Continue.
  2. Check your balance in the dashboard. A low balance shortens replies silently.
  3. If it stops at random points with no finish reason, look at what sits between you and the API: VPNs, corporate proxies or mobile connections that drop long requests.
  4. Make sure any proxy in front of SillyTavern allows long-lived streaming responses and does not buffer them.
  5. Retry with a short message to see whether it is size-related.

Background on streaming.

Telling it apart from similar errors

A request that fails before any text appears is a status error such as 503 or 502, not a cut-off. A non-streaming request that never returns is covered in request timeout.

FAQ

Am I charged for a reply I stopped?

Yes, for the part that streamed before you stopped. Generation upstream is cancelled when you disconnect, so it does not keep running.

Why do I see lines starting with a colon in the raw stream?

Those are SSE comment keepalives sent every 15 seconds during silence. Standard parsers ignore them.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.