Wild West API

Stop sequences: telling the model where to end

A stop sequence is a string that, when the model produces it, immediately ends the response so the text stops at that point.

What stop sequences are

A model keeps generating until it emits its end-of-turn token, hits max tokens, or produces one of your stop sequences. A stop sequence is plain text you supply. The moment the output contains it, generation halts, and the stop text itself is normally not included in what you get back.

How they work

The server checks the decoded text as it is produced. Because tokens and strings do not line up neatly, servers compare against the text, not the token list, and may hold back a few characters while streaming until they know whether a partial match completes. In the OpenAI chat completions spec the field is stop, which takes a string or a short array of strings (up to 4 in OpenAI's own API; other providers set their own limits).

{"stop": ["\nUser:", "\n{{user}}:"]}

Why roleplay frontends use them

The classic failure in roleplay is the model writing the user's lines. It finishes the character's reply, then starts a new line with the user's name and keeps going. SillyTavern adds stop strings such as a newline followed by the user's name and a colon so generation ends right there. It also allows custom stopping strings, which are useful for:

  • Stopping before a model starts a new scene header or out-of-character note.
  • Ending at a formatting marker like ### or a closing tag.
  • Cutting off a model that continues into a fake next turn in a group chat.

Common mistakes

  • Stop strings that are too common. A single newline or a period as a stop will cut almost every reply short.
  • Exact matching. User: does not match user: or User :.
  • Escaping. In JSON a newline is \n. In a UI field, check whether the frontend expects a literal newline or the escape.
  • Relying on stops instead of instructions. A stop sequence trims the symptom; it does not stop the model from wasting tokens on the user's turn before the match, or from describing the user's actions in the narration.

Stop sequences on Wild West API

stop is a standard OpenAI parameter and goes in the /v1/chat/completions body. On the Anthropic-compatible /v1/messages endpoint the equivalent field is stop_sequences. SillyTavern fills these for you from its stopping string settings.

FAQ

Why does my reply cut off early?

One of your stop strings probably appears naturally in the text, for example a short word or a single newline. Make stop strings more specific.

Is the stop sequence included in the output?

Normally no. The response ends just before the stop text.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.