Prefill: starting the model's reply for it
Prefill is the technique of supplying the opening text of the model's reply so that the model continues from those words instead of starting fresh.
What prefill is
Normally the model writes its whole reply from the first token. With prefill, the request ends with the start of the assistant's turn already filled in, and the model picks up mid-reply. If the prefill is *Mara, the reply will begin with Mara doing something, in asterisk action format.
How it works
On Anthropic's Messages API, prefill is an official feature: if the last message in the messages array has the assistant role, the model continues that text. On OpenAI's chat completions API there is no standard prefill; a trailing assistant message is treated differently by different servers. Some open model servers can continue the final assistant message, others start a new turn after it. In text completion setups, prefill is trivial, because you control the raw prompt and simply end it partway into the assistant's turn.
{"messages": [
{"role": "user", "content": "Describe the tavern."},
{"role": "assistant", "content": "The tavern smells of"}
]}The response contains only the continuation, so frontends usually stitch the prefill back on for display.
Uses in roleplay
- Format. Forcing the reply to start in character, in a given format, or with a status block.
- Point of view. Starting with the character's name keeps a reply from opening with the user's actions.
- Structured output. Starting with
{to get JSON. - Getting past refusals on filtered models by making the reply begin as if it had already agreed. That use is fragile and unnecessary on uncensored models.
SillyTavern has "Start Reply With" (for all APIs) and an "Assistant Prefill" field for Claude-style connections.
Common mistakes
Prefilling with trailing whitespace, which some models handle badly; very long prefills that box the model in; and assuming prefill works the same on every endpoint. Reasoning models are a special case: prefilling the answer can conflict with their thinking step, and some APIs reject prefill when extended thinking is on.
Prefill on Wild West API
Wild West API offers both /v1/chat/completions and the Anthropic-compatible /v1/messages. Whether a trailing assistant message is continued depends on the upstream model server, so test with a short prefill and check that the reply continues your text rather than starting a new message. Because the models are uncensored, you will mostly want prefill for format and voice, not to dodge refusals.
FAQ
Does OpenAI's chat completions API support prefill?
Not as a documented feature. Anthropic's Messages API does. Open model servers vary in how they treat a final assistant message.
Is prefill the same as a jailbreak?
No. Prefill is a general technique for steering the start of a reply. It has been used to get past refusals, but its main uses are format and voice.