Fixing an empty response in SillyTavern
An empty reply with no error means the request succeeded but no visible text came back. On this API there are a few specific reasons for that.
Reasoning used up the reply
outlaw-1, glm-5.3-outlaw and glm-5.3-flash-outlaw always reason before they answer. The thinking is streamed as reasoning, separate from the reply text, and it counts against max_tokens. With a small Max Response Length, the model can spend the whole budget thinking and stop before it writes any reply. The response ends with finish_reason: "length" and empty content.
Turning reasoning off does not help on these three models: the API removes reasoning_effort: "none" and similar switches for them, because the models reject the request otherwise.
The reasoning is there but hidden
If SillyTavern is set not to show or request reasoning, a reply where the model only thought can look blank. Turning on reasoning display in the AI Response Formatting settings shows whether the model was thinking.
Other causes
- The model closed the stream without output. Rare. A stream that ends with no usage and no text served is not billed.
- Low balance. The API lowers
max_tokensto what your balance can afford, so a nearly empty account can leave too little room after reasoning. See replies cut short. - A prompt that ends on an assistant turn or a prefill the model treats as complete.
- Stop strings that match at the very start of the reply, so SillyTavern trims everything.
Step by step fix
- Raise Max Response Length. For the reasoning models, 2000 tokens or more leaves room to think and answer.
- Turn on reasoning display so you can see what the model did.
- Check custom stop strings in the formatting settings.
- Check your balance in the dashboard.
- Try
mimo-v2.6-flash-outlaworqwen3.8-27b-outlaw, which do not always reason, to confirm the cause. - Regenerate.
Telling it apart from similar errors
An empty reply is a 200. If there is an error toast, find that error instead, such as 503 or context length exceeded. A reply that starts and stops mid-sentence is covered in stream cut off.
FAQ
Am I charged for an empty reply?
You pay for tokens the model generated, including reasoning. A stream that produced no output and reported no usage is not billed.
Why does the same chat work on one model and come back empty on another?
The always-reasoning models spend part of max_tokens thinking. A budget that suits a direct model may be too small for them.