Wild West API

Reasoning models: LLMs that think before answering

A reasoning model is a language model trained to write out a step-by-step thinking process before giving its final answer, which tends to improve accuracy on hard problems.

What a reasoning model is

Ordinary chat models start writing the answer immediately. Reasoning models are trained, usually with reinforcement learning, to first produce a block of thinking: working through the problem, checking options, correcting themselves. Only then do they write the answer. Examples include the DeepSeek R1 line, Qwen3 models in thinking mode, and many recent GLM and frontier models. Some models can switch thinking on or off per request.

How the thinking is returned

Providers handle it in different ways:

  • A separate field, such as reasoning_content or reasoning on the message or stream delta.
  • Inline tags in the text, such as <think>...</think>, which the client has to strip.
  • Anthropic-style thinking content blocks on /v1/messages.
  • Hidden entirely, with only a token count reported.

SillyTavern can display reasoning in a collapsible block and has settings to auto-parse <think> style tags and to request reasoning from supported sources.

Cost and length

Thinking tokens are output tokens, and they are billed as output whether or not you see them. They also count toward max tokens on most APIs: if the limit is 300 and the model thinks for 400, you get no answer. Give reasoning models a much higher limit, often several thousand tokens. Expect a longer pause before the visible reply starts; with streaming, the thinking may stream first.

What it means for roleplay

  • Better: tracking complex plots, keeping rules and stats consistent, puzzles, mysteries, and following detailed instructions.
  • Worse or mixed: slower turns, higher cost, and sometimes stiffer prose, since the model plans rather than flows.
  • Do not send old thinking back. Most providers recommend leaving previous turns' reasoning out of the history. Keeping it wastes context and can make the model copy its old thoughts. SillyTavern does not include parsed reasoning in the prompt by default.
  • Samplers. Some reasoning models are tuned for specific temperature ranges and can loop at very low temperature; check the model's recommended settings.

Reasoning on Wild West API

Whether a given model thinks, and how the thinking is returned, depends on the model; check the models page and the docs. If a model returns reasoning, budget max tokens for it and make sure your client does not paste old reasoning into the next request.

FAQ

Do I pay for thinking tokens?

Yes. Reasoning tokens are generated output and are billed as output, even when the thinking is hidden.

Why did a reasoning model return an empty answer?

It probably used all of max tokens on thinking. Raise max tokens well above the length of answer you expect.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.