Wild West API

Min P sampling explained

Min P is a truncation sampler that removes every token whose probability is less than a set fraction of the most likely token's probability.

What min P is

Min P is a way of throwing away unlikely tokens before one is picked. You give it a number such as 0.05. Any token whose probability is below 5 percent of the top token's probability is removed, and the model samples from what is left.

It became popular in local model communities (llama.cpp, KoboldCpp, text-generation-webui) because it adapts to how confident the model is, which fixed tokens counts and fixed probability mass do not.

How the cutoff scales

The threshold is min_p × p(top token). That multiplication is the whole idea:

  • If the model is confident and the top token has probability 0.90, a min P of 0.05 sets the cutoff at 0.045. Only a handful of strong alternatives survive.
  • If the model is unsure and the top token has only 0.10, the cutoff falls to 0.005, so many reasonable options stay in play.

Compare top K, which keeps the same number of tokens regardless of confidence, and top P, which can keep a long tail of junk when the distribution is flat. Min P cuts hard when the answer is obvious (a name, a closing quote) and loosely when many words would work (a description).

What it does to roleplay output

Min P is the reason people can run temperature above 1 without the text falling apart. The high temperature spreads probability around for more varied prose, and min P removes the garbage tokens that would otherwise appear: misspelled names, random symbols, a sudden switch of language. Where samplers run in order, putting temperature after min P (SillyTavern's "temperature last") keeps the cutoff based on the model's original confidence.

Sensible values

  • 0.02 to 0.05: light filtering, good default for creative writing.
  • 0.05 to 0.1: tighter, useful with high temperature or a small or heavily quantized model.
  • Above 0.2: output gets close to greedy and repetitive.
  • 0 disables it.

When you use min P, you can usually set top P to 1.0 and top K to 0 (off) so they do not stack on top of it in ways that are hard to reason about.

Min P on Wild West API

min_p is not part of the OpenAI chat completions spec. It is a frontend and backend specific parameter. You can add it to the request body, and SillyTavern may send it for some connection types, but the API may ignore it unless the upstream model server supports it. If you need guaranteed behavior, rely on the standard parameters (temperature, top_p, the penalties) and treat min P as a bonus when it is honored. See the docs for request details.

FAQ

What is a good min P value?

0.05 is a common starting point for roleplay. Lower it toward 0.02 for more variety, raise it toward 0.1 if you see broken words or names.

Should I use min P and top P together?

You can, but it is easier to tune one. Most people who use min P set top P to 1.0 and top K to off.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.