Wild West API

Top P sampling, also called nucleus sampling

Top P, or nucleus sampling, limits each token choice to the smallest group of candidates whose combined probability reaches a set threshold.

What top P is

Top P is a truncation sampler. The model sorts candidate tokens from most to least likely, adds up their probabilities in order, and stops once the total reaches P. Only the tokens inside that group (the nucleus) can be picked. A top P of 0.9 means the next token is drawn from whichever tokens make up the first 90 percent of probability mass.

It was introduced in a 2019 paper on neural text degeneration and is one of the few samplers in the OpenAI API spec, so almost every hosted endpoint accepts it.

How it works

Say the top candidates have probabilities 0.50, 0.25, 0.10, 0.05, and the rest share 0.10. With top P 0.9 the sampler keeps the first four tokens (0.50 + 0.25 + 0.10 + 0.05 = 0.90), renormalizes them to sum to 1, and samples. With top P 0.75 only the first two survive.

The weakness is flat distributions. When the model is unsure, reaching 90 percent of the mass can require hundreds of tokens, including bad ones. That is the gap min P was designed to close. On peaked distributions top P behaves well.

Effect on roleplay

Lowering top P makes the prose more conservative: fewer odd words, but also more of the model's favorite phrases. Raising it toward 1.0 lets more variety through. Because temperature changes the shape of the distribution, the two settings interact strongly. A higher temperature flattens the curve, which means top P has to include more tokens to reach the same total.

Sensible values and mistakes

  • 1.0: off. Use this if you control randomness with temperature or min P instead.
  • 0.9 to 0.95: a light trim of the tail, a common default.
  • 0.7 to 0.85: noticeably safer, can feel flat in long stories.

OpenAI's own guidance is to adjust temperature or top P, not both at once, because the effects overlap. A frequent mistake is leaving top P at a low value from an old preset and then wondering why raising temperature does very little.

Top P on Wild West API

top_p is a standard OpenAI parameter, so it goes in the body of a /v1/chat/completions request next to temperature. In SillyTavern it is the Top P slider on the Chat Completion sampler panel.

{"model": "outlaw-1", "temperature": 0.9, "top_p": 0.95, "messages": [...]}

FAQ

What does top P 1.0 mean?

Top P 1.0 keeps every token, so the sampler is effectively off and only temperature and any other samplers shape the output.

Is top P better than min P?

Min P handles flat distributions better because its cutoff scales with the top token's probability. Top P is more widely supported because it is part of the OpenAI spec.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.