Fixing request timeouts on long generations
A timeout usually means your client gave up waiting, not that the API failed. It is most common with streaming turned off and long replies.
Why it happens
With streaming off, nothing is sent back until the whole reply is finished. Reasoning models can think for a long time before writing, and a long reply on a big prompt can take minutes. Any client, proxy or SDK with a shorter timeout gives up first. In SillyTavern this can look like a generic network error; in the OpenAI SDKs it raises APITimeoutError, which they retry by default, sending the whole request again.
With streaming on, the API sends a keepalive comment every 15 seconds while the model is quiet, so idle-connection timeouts in proxies and clients do not fire. That is the main reason to stream long replies.
What the API does when you give up
When the client disconnects, the request is cancelled and the cancel is passed on to the model, so generation stops. On a stream, you are billed for what was streamed before the disconnect. A non-streaming request that is cancelled before the answer arrives returns nothing to you.
On the server side, a stream is closed if the model sends nothing at all for two minutes. That limit is measured from the last chunk, so a reply that keeps producing tokens is not cut by it.
Step by step fix
- Turn streaming on. In SillyTavern, enable Streaming in the AI Response Configuration panel. In code, set
stream: true. - If you need non-streaming, raise the client timeout. For the Python SDK, pass
timeout=600to the client; the Node SDK takes atimeoutoption in milliseconds. - Be careful with automatic retries on timeout: each retry is a new request and is billed for what it generates.
- If a proxy sits in front of your app, make sure its read timeout is longer than your longest expected reply, and that it does not buffer streaming responses.
- Lower Max Response Length if you do not need very long replies.
See streaming and the docs for examples.
Telling it apart from similar errors
A 502 or 503 is a status the API sent back; see 502 and 503. A stream that starts and then stops is covered in stream cut off.
FAQ
Is a timed-out request charged?
Generation stops when you disconnect. A stream is billed for what it sent before the disconnect.
Why does streaming help if the reply takes the same time?
Bytes keep arriving, including keepalives every 15 seconds, so no idle timeout fires, and you see the text as it is written.