Wild West API

Fixing an empty response in SillyTavern

An empty reply with no error means the request succeeded but no visible text came back. On this API there are a few specific reasons for that.

Reasoning used up the reply

outlaw-1, glm-5.3-outlaw and glm-5.3-flash-outlaw always reason before they answer. The thinking is streamed as reasoning, separate from the reply text, and it counts against max_tokens. With a small Max Response Length, the model can spend the whole budget thinking and stop before it writes any reply. The response ends with finish_reason: "length" and empty content.

Turning reasoning off does not help on these three models: the API removes reasoning_effort: "none" and similar switches for them, because the models reject the request otherwise.

The reasoning is there but hidden

If SillyTavern is set not to show or request reasoning, a reply where the model only thought can look blank. Turning on reasoning display in the AI Response Formatting settings shows whether the model was thinking.

Other causes

  • The model closed the stream without output. Rare. A stream that ends with no usage and no text served is not billed.
  • Low balance. The API lowers max_tokens to what your balance can afford, so a nearly empty account can leave too little room after reasoning. See replies cut short.
  • A prompt that ends on an assistant turn or a prefill the model treats as complete.
  • Stop strings that match at the very start of the reply, so SillyTavern trims everything.

Step by step fix

  1. Raise Max Response Length. For the reasoning models, 2000 tokens or more leaves room to think and answer.
  2. Turn on reasoning display so you can see what the model did.
  3. Check custom stop strings in the formatting settings.
  4. Check your balance in the dashboard.
  5. Try mimo-v2.6-flash-outlaw or qwen3.8-27b-outlaw, which do not always reason, to confirm the cause.
  6. Regenerate.

Telling it apart from similar errors

An empty reply is a 200. If there is an error toast, find that error instead, such as 503 or context length exceeded. A reply that starts and stops mid-sentence is covered in stream cut off.

FAQ

Am I charged for an empty reply?

You pay for tokens the model generated, including reasoning. A stream that produced no output and reported no usage is not billed.

Why does the same chat work on one model and come back empty on another?

The always-reasoning models spend part of max_tokens thinking. A budget that suits a direct model may be too small for them.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.