Wild West API

Top K sampling explained

Top K is a sampler that only lets the model choose its next token from the K most probable candidates.

What top K is

Top K is the simplest truncation sampler. Sort the vocabulary by probability, keep the top K tokens, discard everything else, renormalize, and sample. With top_k = 40 the model can only ever choose among its 40 best guesses for each position.

How it behaves

The problem is that K is fixed while the model's confidence is not. When the next token is obvious (the second half of a name, a closing bracket), 40 candidates is far too many and can let nonsense through if temperature is high. When the next token is wide open (the first word of a description), 40 might be too few and cut off good options. Top P and min P both adapt to the shape of the distribution, which is why most current presets lean on them and treat top K as a safety net or turn it off.

Top K is cheap to compute, and in local inference it is often applied first so the later samplers work on a short list instead of the full vocabulary of 100,000 or more tokens.

Effect on roleplay

A low K (under about 20) makes characters sound repetitive and predictable. A high K, or K turned off, gives more variety but depends on other samplers to remove junk. In practice, top K matters most when you run high temperature without any other truncation.

Sensible values

  • 0 (or -1 in some backends): disabled. A good choice when you use min P.
  • 40: the long-standing llama.cpp default, a harmless safety cap.
  • 100 or more: barely does anything on most text.
  • 1: greedy decoding, always the top token.

A common mistake is setting top K to 1 for "consistency" in a roleplay and then getting loops, since greedy decoding is the most repetition-prone setting there is.

Top K on Wild West API

top_k is not a standard OpenAI chat completions parameter. It is frontend and backend specific: you can include it in the request body, but the API may ignore it unless the upstream model server supports it. For portable settings use temperature and top_p. If you connect through SillyTavern's custom OpenAI-compatible source, extra sampler fields can be added, but do not assume they change anything until you have tested the output.

FAQ

What does top K 0 mean?

In most backends 0 disables top K, so all tokens remain candidates and other samplers do the filtering. Some backends use -1 for the same thing.

Do I need top K if I use min P?

Usually not. Min P already removes the low probability tail, so most people set top K to off.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.