Jailbreak prompts: what they are and why they fail
A jailbreak prompt is text added to a request to persuade a safety-tuned model to answer things it would normally refuse.
What a jailbreak prompt is
Mainstream chat models are trained to refuse a range of requests, and for fiction writers the line often falls on dark themes, violence, or adult content. A jailbreak prompt is an instruction block meant to override that training, usually by framing the conversation as fiction, assigning the model an unrestricted persona, or stacking rules telling it not to refuse. SillyTavern once had a field literally called "Jailbreak prompt"; it has since been renamed to post-history instructions, which is a more accurate description of where the text goes (after the chat history).
Why they work at all
Refusal behavior comes from fine-tuning, and that fine-tuning does not cover every framing. Text that pushes the conversation into a context the training did not anticipate can tip the model past the point where it would refuse. Placing the instruction after the chat history, or using prefill so the reply starts as if it already agreed, adds weight because recent text counts for more.
Why they are a poor fit for fiction
- Unstable. Providers update models and filters. A jailbreak that works this week can stop working with no notice.
- Token cost. Long jailbreaks are resent on every request.
- Side effects. Prompts that tell the model to be "unrestricted" often change its tone, make it overly crude, or push content where the story did not ask for it.
- Soft refusals. Even when they work, filtered models can still dodge by summarizing, moralizing, skipping scenes, or turning characters suddenly agreeable.
- Terms of service. Many providers prohibit attempts to bypass their safeguards and can suspend accounts.
The alternative: models that do not refuse fiction
The more reliable route is a model whose refusal behavior has been removed or never added, so your system prompt can be about the story instead of arguing with the model. That is done through abliteration, which removes the refusal direction from the weights, or through an uncensored fine-tune. See abliterated models for a fuller explanation.
On Wild West API
Wild West API serves uncensored models, so you can drop jailbreak text from your preset. If you import a SillyTavern preset that includes a long post-history jailbreak, consider removing it: it costs tokens and can push the writing toward a style you did not ask for. Keep post-history instructions for genuine style guidance, such as length and point of view. Usage still has to follow the terms of the service.
FAQ
Do I need a jailbreak with an uncensored model?
No. Uncensored models do not refuse fiction requests in the first place, so jailbreak text only adds tokens and can distort the writing.
What did SillyTavern rename the jailbreak prompt to?
Post-history instructions, the text inserted after the chat history.