Wild West API

Presence penalty in the OpenAI API

Presence penalty is an OpenAI API setting that lowers the score of any token that has already appeared, by the same amount no matter how often.

What presence penalty is

Presence penalty is one of the two standard repetition controls in the OpenAI API, alongside frequency penalty. It answers a yes or no question for each token: has this token appeared yet? If so, its score goes down by a fixed amount. The intended effect is to push the model toward words and topics it has not used.

How it works

OpenAI documents the adjustment as a subtraction from the logit: the presence penalty is subtracted once for any token whose count is above zero, and the frequency penalty is subtracted once per occurrence. Values range from -2.0 to 2.0. Positive values discourage repeats; negative values encourage them.

Because the penalty is subtracted from a logit rather than dividing it, its strength is on a different scale from llama.cpp style repetition penalty. A presence penalty of 0.5 is moderate, while a repetition penalty of 0.5 would mean something else entirely. Implementations also differ on whether tokens in the prompt count or only tokens in the reply being generated, so the same value can feel different across providers.

Effect on roleplay

A small positive value helps a scene move on instead of circling the same detail, such as a character who keeps returning to the same object, smell or gesture in every paragraph. Too much and the model starts avoiding necessary words: character names, pronouns, recurring locations. Since the penalty is flat, it does not care whether a word appeared once or twenty times, which makes it a gentle tool rather than a loop breaker.

Sensible values and mistakes

  • 0: off, the default.
  • 0.1 to 0.5: mild push toward novelty, fine for long chats.
  • 1.0 and up: often degrades fluency.

Mistakes include stacking a high presence penalty, a high frequency penalty and a repetition penalty at once, and using negative values by accident when moving a slider. Change one penalty at a time and keep the others at zero. Judge the result over several swipes, not one, since a single reply can look better or worse by chance.

Presence penalty on Wild West API

presence_penalty is a standard OpenAI parameter, so you send it in the /v1/chat/completions body and SillyTavern exposes it on the Chat Completion sampler panel. How strongly a given model responds still depends on the upstream implementation.

{"model": "glm-5.3-outlaw", "presence_penalty": 0.3, "messages": [...]}

FAQ

What is the difference between presence and frequency penalty?

Presence penalty applies the same reduction once a token has appeared at all. Frequency penalty grows with each additional occurrence.

Can presence penalty be negative?

Yes, the OpenAI range is -2.0 to 2.0. Negative values make the model more likely to reuse tokens, which is rarely what you want in fiction.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.