Wild West API

Temperature: how it changes what the model writes

Temperature is a sampling setting that makes a language model's word choices more predictable when lowered and more varied when raised.

What temperature is

Every time a model writes a token, it first produces a score (a logit) for every token in its vocabulary. Those scores are turned into probabilities with a softmax, and one token is drawn from that distribution. Temperature is a number applied in the middle of that step. Lower values make the likely tokens even more likely. Higher values flatten the distribution so less likely tokens get picked more often.

It does not make the model smarter, more creative in any deep sense, or more willing to write certain content. It only changes how the model picks between options it already considered plausible.

How it works under the hood

The logits are divided by the temperature before the softmax. At 1.0 nothing changes. At 0.5 the gaps between scores double, so the top token dominates. At 1.5 the gaps shrink and the tail of unlikely tokens gains probability mass. At or near 0, most backends switch to greedy decoding and always take the single most likely token.

Because temperature reshapes the whole distribution, it interacts with truncation samplers such as top P and min P. Whether temperature runs before or after truncation changes the result. Local backends like llama.cpp and KoboldCpp let you set the order; SillyTavern exposes a "temperature last" option for text completion backends. With a hosted chat completions API, the order is decided by the server.

What it does to roleplay

  • Too low (under about 0.5): replies get safe and samey. The character reuses the same phrases, swipes look nearly identical, and the story stalls.
  • Middle (about 0.7 to 1.0): the usual range for fiction. Prose varies without losing the thread.
  • High (above about 1.2): more surprising word choices, but also wrong names, broken grammar and drifting plot, unless a truncation sampler like min P trims the junk tokens first.

Sensible values and common mistakes

Start at 0.8 to 1.0 for roleplay and adjust in steps of 0.1. For rules-heavy games, stat tracking or tool calling, go lower (0.3 to 0.6) so the model follows formats. The OpenAI spec accepts 0 to 2.

Common mistakes: raising temperature to fix repetition (a repetition penalty or DRY usually works better), changing temperature and top P at the same time and not knowing which helped, and copying a preset tuned for one model onto another. Models differ in how peaked their distributions are, so the same number behaves differently across them.

Temperature on Wild West API

temperature is a standard OpenAI parameter, so you send it in the request body of /v1/chat/completions and it is passed along with the request. In SillyTavern it is the Temperature slider on the sampler panel when you connect through Chat Completion with a custom OpenAI-compatible source.

{
  "model": "glm-5.3-outlaw",
  "temperature": 0.9,
  "messages": [{"role": "user", "content": "Continue the scene."}]
}

FAQ

What temperature is best for roleplay?

Most people land between 0.8 and 1.0. Go lower if the model loses track of names and facts, higher if replies feel repetitive and you also use a truncation sampler like min P or top P.

Does temperature 0 make output fully deterministic?

Usually close, but not guaranteed. Hosted APIs can still vary slightly because of batching and floating point differences on the server.

Is higher temperature more uncensored?

No. Temperature only changes how tokens are picked. Whether a model refuses depends on its training, not on this setting.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.