Wild West API

Frequency penalty in the OpenAI API

Frequency penalty is an OpenAI API setting that lowers a token's score in proportion to how many times it has already appeared.

What frequency penalty is

Frequency penalty is the counting partner of presence penalty. Instead of a single flat reduction, it subtracts an amount for every time a token has already been used. A word used once is nudged down a little; a word used ten times is pushed down a lot. It targets the verbal tics that build up over a long reply.

How it works

In OpenAI's formulation, the frequency penalty multiplied by the token's count so far is subtracted from that token's logit. Values run from -2.0 to 2.0. Since the penalty keeps growing with each repeat, even a modest value can become strong on common words in a long output. Function words like "the", "and" and "her" are the most repeated tokens in any text, so they absorb most of the penalty.

Implementations vary on what text is counted. Some count only the reply being generated, others include the prompt. That is one reason a value that works on one provider feels too strong or too weak on another.

Effect on roleplay

Low values reduce repeated adjectives and stock phrases within a single long reply. High values make prose strange: the model starts dodging ordinary words, uses odd synonyms, and drops pronouns. Frequency penalty is also blind to multi-token phrases, so it is weak against the problem most roleplayers actually have, the same sentence pattern reappearing across many turns. The DRY sampler handles that better where it is available.

Sensible values

  • 0: off.
  • 0.1 to 0.4: light touch for long replies.
  • Above 0.7: expect fluency problems in long outputs.

Keep frequency penalty and presence penalty modest and change one at a time, testing on a long reply where the effect has room to build up. If you are also sending a backend specific repetition penalty, leave it at 1.0 while you tune these, or you will not know which one is causing an effect.

Frequency penalty on Wild West API

frequency_penalty is a standard OpenAI parameter. Put it in the body of a /v1/chat/completions request, or set the Frequency Penalty slider in SillyTavern's Chat Completion panel. Each upstream model may respond to it with different strength, so test on the model you actually use.

FAQ

What frequency penalty should I use for roleplay?

Start at 0 and go up to around 0.2 or 0.3 only if a single long reply keeps reusing the same words.

Why does a high frequency penalty make text weird?

Common words like 'the' and pronouns repeat constantly, so a high count-based penalty pushes the model away from ordinary grammar.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.