Instruct templates and chat formats
An instruct template is the pattern of special tokens and markers that wraps each message so the model can tell system, user and assistant turns apart.
What an instruct template is
Under the hood, a model only ever sees one long sequence of tokens. Chat models were trained with that sequence formatted in a particular way: special tokens mark where a turn begins, whose turn it is, and where it ends. The instruct template, also called a chat template or prompt format, is that pattern. Each model family has its own.
Common formats
ChatML, used by many models including Qwen:
<|im_start|>system You are Mara.<|im_end|> <|im_start|>user Hello.<|im_end|> <|im_start|>assistant
Llama 3 uses <|start_header_id|>user<|end_header_id|> headers and <|eot_id|> to end turns. Mistral models use variants of [INST] ... [/INST]. Others, such as GLM models, have their own role tokens. The model's tokenizer configuration usually includes the official template as a Jinja script.
Why the right one matters
With the wrong template, the model is reading a format it never saw in training. Typical symptoms: replies that never end, the model writing the user's next turn, stray tags in the output, worse instruction following, or a sudden shift in tone. The end-of-turn token is especially important; if the template does not include the right one as a stop, generation runs on.
When you need to set it
Only in text completion. If you send raw prompts to a local backend, SillyTavern's Advanced Formatting panel builds the prompt with the instruct template you select, and you must match it to the model. SillyTavern ships presets for common formats and can try to auto-select one from the model name.
In chat completion, the server applies the model's built-in template to your messages array. Your frontend's instruct template is not used, and you should never type role tokens into messages yourself.
Common mistakes
- Editing the instruct template while connected via Chat Completion and expecting it to change anything.
- Using a generic Alpaca template (
### Instruction:) for a model trained on ChatML. It may half work, which makes the problem harder to spot. - Missing the template's stop tokens in stop sequences, so the model continues into an imaginary next turn.
- Putting
<|im_start|>style tokens in a card or system prompt when using a chat API.
On Wild West API
Wild West API uses chat-style endpoints, so the server formats each model's turns with its own template. In SillyTavern, connect through Chat Completion and you can ignore instruct template settings entirely.
FAQ
Do I need an instruct template for chat completion?
No. The server applies the model's template. Instruct templates only matter when you send raw text completion prompts.
What is ChatML?
A widely used chat format that wraps each turn in <|im_start|>role and <|im_end|> tokens, used by Qwen and many other models.