Abliteration: removing refusals from model weights
Abliteration is a technique that removes a language model's tendency to refuse by finding the internal direction tied to refusal and erasing it from the weights.
What abliteration is
Safety-tuned models refuse certain requests. Abliteration is a way to take that behavior out of an existing model without retraining it on new data. The name, a blend of "ablate" and "obliterate", was popularized in 2024 by FailSpy and others who applied the method to open-weight models. It builds on research showing that refusal in many chat models is controlled by a single refusal direction in the model's internal activations.
How it works under the hood
- Collect activations. Run the model on a set of prompts it refuses and a set it answers, and record the residual stream activations (the running internal state) at a chosen layer.
- Find the direction. Subtract the mean activation of the answered set from the mean of the refused set. The difference is a vector that points toward "this will be refused".
- Remove it. Modify the weight matrices that write into the residual stream so their outputs have no component along that vector. This is weight orthogonalization. After it, the model can no longer represent the direction, so it cannot act on it.
Because the change is baked into the weights, the result is a normal model file that runs anywhere the original did. No special inference code or prompt is needed.
What it does to output
A well-abliterated model answers requests it used to refuse, including dark or explicit fiction, without needing a jailbreak prompt. Side effects to know about:
- Small capability loss. Benchmarks often drop slightly. Some releases follow abliteration with a short fine-tune to recover quality.
- Leftover hedging. Disclaimers, moralizing and soft refusals can come from more than one direction, so some remain.
- No new knowledge or style. It removes a behavior; it does not teach the model to write better fiction. If the base model writes flat prose, the abliterated one will too.
- Agreeableness. Some abliterated models become reluctant to say no in character, so villains can turn oddly cooperative. A clear character card helps.
Abliteration compared with fine-tuning
An uncensored fine-tune trains the model on data without refusals, which can also change its style and knowledge. Abliteration is a targeted edit that leaves everything else mostly in place. Some models use both. See abliterated models for more detail.
On Wild West API
Wild West API serves uncensored models for fiction and roleplay, so you can write a system prompt about the story rather than about avoiding refusals. Check the models page for what each model is and its context and vision support.
FAQ
Does abliteration make a model worse?
Usually only slightly. Some benchmark scores drop a little, and some releases add a short fine-tune afterward to recover quality.
Is abliteration the same as fine-tuning?
No. Abliteration edits existing weights to remove one internal direction. Fine-tuning trains the model further on new data.