Streaming responses from an LLM API
Streaming is an API mode where the model's reply is sent piece by piece while it is being generated, instead of all at once at the end.
What streaming is
Without streaming, you send a request and wait until the full reply is finished before seeing anything. With streaming, text appears within a second or two and keeps arriving as the model writes. Total generation time is the same; the waiting just feels much shorter. For roleplay this also lets you stop a reply that is going the wrong way and swipe, without paying for the rest of it.
How it works
Streaming uses server-sent events (SSE) over a normal HTTP response that stays open. On an OpenAI-style endpoint you set "stream": true, and the server sends lines like:
data: {"choices":[{"delta":{"content":"The door"}}]}
data: {"choices":[{"delta":{"content":" creaks"}}]}
data: [DONE]Each chunk holds a small delta to append. The last chunk carries a finish_reason, and [DONE] ends the stream. Anthropic-style /v1/messages streams use named events instead (message_start, content_block_delta, message_stop and so on), so clients need the parser that matches the endpoint they call. Tool calls also stream, with the function arguments arriving as JSON fragments you must join before parsing.
What breaks streaming
- Buffering proxies. A reverse proxy or corporate network that buffers responses makes the whole reply arrive at the end.
- Timeouts. Long generations, especially from reasoning models that think before writing, can sit silent for a while. Some clients give up too early.
- Partial JSON. Code that tries to parse each network chunk as a full event breaks when an event is split across reads. Buffer by line.
- Usage numbers. Token counts may only arrive in a final chunk, or only if requested, so a client that stops reading at the first
finish_reasoncan miss them.
Streaming in SillyTavern and on Wild West API
In SillyTavern, streaming is a checkbox on the sampler panel. Turning it off is a useful debugging step: a non-streamed request returns one complete error message, which is easier to read than a stream that dies halfway.
On Wild West API, stream is a standard parameter on /v1/chat/completions, and the Anthropic-compatible /v1/messages endpoint streams in Anthropic's event format. Note that the API does not answer CORS today, so a web page calling it directly from the browser will fail. SillyTavern works because its own server makes the request.
FAQ
Does streaming cost more?
No. You are billed for the same input and output tokens either way. Stopping a stream early can save output tokens that would otherwise have been generated.
Why does streaming show nothing and then the whole reply at once?
Something between you and the API is buffering the response, or the client has streaming turned off.