Wild West API

Repetition penalty: what it does and how to set it

Repetition penalty is a sampler that makes tokens already present in the recent context less likely to be chosen again.

What it is

Language models like to repeat themselves, especially in long chats. The same sentence opener, the same gesture, the same descriptive phrase show up turn after turn. Repetition penalty pushes back by lowering the score of any token that has already appeared within a recent window of text.

How it works under the hood

The common form comes from the CTRL paper and is what llama.cpp, KoboldCpp, text-generation-webui and Hugging Face transformers implement. For each token that has appeared in the window, its logit is divided by the penalty if positive and multiplied by it if negative. Either way the token becomes less likely. A penalty of 1.0 does nothing; 1.1 is a moderate push.

Two details matter:

  • It is binary per token. A token that appeared once gets the same penalty as one that appeared fifty times. Counting is what frequency penalty adds.
  • It works on tokens, not ideas. It penalizes "the", commas, quotation marks and your character's name just as much as an overused phrase.

Most backends add a range setting (how many recent tokens to scan) and sometimes a slope that weakens the penalty for older tokens.

What it does to roleplay

Mild settings break loops and reduce stock phrases. Strong settings cause the classic symptoms of overcorrection: the model avoids your character's name and uses awkward synonyms, punctuation gets strange, it switches to rarer words, and dialogue formatting breaks because quote marks and asterisks are penalized. If you see a character suddenly called "the woman" or "the brunette" everywhere, the penalty is probably too high.

Sensible values

  • 1.0: off.
  • 1.03 to 1.1: typical for roleplay with a range of 1024 to 2048 tokens.
  • Above 1.15: likely to damage names and formatting.

Modern alternatives target the actual problem better. The DRY sampler penalizes repeated sequences rather than single tokens, so it leaves names and punctuation alone. Many current presets use DRY with repetition penalty at 1.0.

Repetition penalty on Wild West API

repetition_penalty is not a standard OpenAI parameter. It is frontend and backend specific, so it may be ignored by the API unless the upstream model server supports it. The standard alternatives are frequency_penalty and presence_penalty, which use a different (subtractive) formula and a different scale, so do not copy a 1.1 repetition penalty into those fields.

FAQ

What is a good repetition penalty for roleplay?

Around 1.05 to 1.1 over the last 1024 to 2048 tokens. If names or punctuation start breaking, lower it or switch to DRY.

Is repetition penalty the same as frequency penalty?

No. Repetition penalty divides the logit once for any token present, while frequency penalty subtracts an amount that grows with how many times the token appeared, and they use different scales.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.