How abliteration strips the refusals out
Abliteration is a way to stop a language model refusing by finding the single direction in its activations that corresponds to "I won’t do that" and removing it from the weights. No fine-tuning run and no new data: a surgical edit, published as a new set of weights.
Where the word comes from
The method rests on one interpretability result: Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (2024), showed that across a range of open chat models, refusal is controlled by one linear direction in the residual stream. Remove it and the model stops refusing; add it and the model refuses harmless requests.
FailSpy turned that into a practical recipe for editing released weights and named it, a portmanteau of ablate and obliterate. Maxime Labonne's tutorial, Uncensor any LLM with abliteration, made it widely known, and "-abliterated" has been a suffix on Hugging Face model names ever since.
How abliteration works, step by step
- Collect two sets of prompts. A few hundred harmful instructions the model refuses, and a few hundred harmless ones it answers.
- Record activations. Run both sets through the model and save the residual stream activation at the last token position, at every layer.
- Take the difference of means. For each layer, the mean harmful activation minus the mean harmless one, normalised, is a candidate refusal direction.
- Pick the best layer. Test each candidate by ablating it on held-out prompts. Keep the one that removes the most refusals while leaving harmless answers unchanged.
- Orthogonalise the weights. Every matrix that writes into the residual stream (the token embeddings, each attention output projection, each MLP output projection) has the direction projected out, so no layer can write along it again.
- Save, and usually retune. The edited weights are an ordinary checkpoint that any server can load. The better releases follow with a light preference retune to win back what the edit cost.
The core of step five is a few lines. For a weight W that writes to the residual stream and a unit refusal direction r:
# r: unit refusal direction, shape (d_model,) # Linear weights are (out_features, in_features); out_features = d_model here for W in [*attn_o_proj_weights, *mlp_down_proj_weights]: W -= torch.outer(r, r) @ W # embedding rows are residual vectors, so project each row E -= torch.outer(E @ r, r)
Runtime ablation versus editing the weights
The same projection can be applied at inference time instead, as a hook that removes the direction from the activations on every forward pass. That is how the effect is usually tested. Baking it into the weights is what makes an abliterated model portable: the result is a normal checkpoint that converts to GGUF, loads in any inference server and needs no special code to serve.
Newer tools automate the choices that used to be made by hand. Heretic, for example, searches over which layers to edit and how strongly, scoring each attempt on how many refusals remain and how far the outputs drift from the original model, which is why "heretic" now appears in model names beside "abliterated".
Abliteration, jailbreaks and uncensored finetunes
| What changes | Survives a new prompt | Cost to capability | |
|---|---|---|---|
| Jailbreak | Only the prompt | No, each one has to be found and can be patched | None, but unreliable |
| Uncensored finetune | The weights, by training on chosen data | Yes, mostly | Depends on the data; can still refuse |
| Abliteration | The weights, by projecting out one direction | Yes | Small but real, recovered partly by a retune |
The full comparison of the last two is in uncensored vs abliterated.
What it does and does not change
Because it edits the mechanism of refusal rather than teaching new facts, an abliterated model keeps most of its original capability and simply stops declining. It does not make the model more knowledgeable or more accurate, and refusal is not perfectly isolated in one direction: some refusals survive, and some unrelated behaviour, most often careful reasoning and instruction following, gets slightly worse.
Why security teams want it
Offensive security, trust-and-safety data generation and model red-teaming all involve prompts that a frontier model refuses on sight, even from a caller with written authorisation. See the red team use case.
Calling an abliterated model
You do not need to run the edit yourself to use one. The abliterated models page lists the ones Wild West API sells, and they answer on an OpenAI-compatible API at https://wildwestapi.com/v1, so any OpenAI client works by changing the base URL.
FAQ
What is abliteration in AI?
A technique for removing refusal from an open-weight language model by finding the direction in its activations that carries refusal and projecting it out of the weights. The model stops declining requests without being retrained.
Is abliteration the same as jailbreaking?
No. A jailbreak is a prompt that tricks a still-aligned model into answering. Abliteration edits the weights so there is nothing to trick; the refusal behaviour is gone.
Does abliteration hurt model quality?
It can slightly degrade output, because the removed direction is not perfectly isolated from useful behaviour. A well-done abliteration keeps most capability, and a light retune afterwards recovers more.
Can any model be abliterated?
Any open-weight model whose activations you can read, which rules out closed models served only through an API. In practice it works best on instruction-tuned chat models, where refusal was trained in and concentrates cleanly.
How many prompts does it take?
Not many. The direction is a difference of two means, so a few hundred harmful and a few hundred harmless prompts are typically enough to estimate it.