Where a model keeps its no
Recent interpretability work found that whether a chat model refuses is governed largely by one linear direction in its activation space. Push activations along it and the model refuses; remove it and the model stops.
One direction, many refusals
Across many kinds of "unsafe" request, the internal signal that precedes a refusal points the same way. That means refusal is not scattered through the network but concentrated in a direction you can measure by contrasting harmful and harmless prompts, then averaging the activation gap.
What you can do with it
Add the direction to a compliant model and it starts refusing harmless things; subtract it and a cautious model complies. Doing the subtraction permanently in the weights is abliteration. The finding is also useful defensively: a classifier can watch that direction to detect when a model is about to refuse or has been steered.
Where it breaks down
FAQ
What is the refusal direction?
A single direction in a chat model’s residual stream activations that tracks whether it is about to refuse. Arditi et al. (2024) found it in a range of open chat models by contrasting activations on harmful and harmless prompts.
How is the refusal direction found?
Run the model on two prompt sets, one it refuses and one it answers, record the activations at the last token for each layer, and take the difference of the means. The layer whose direction best removes refusals when ablated is the one used.
What can you do with it besides abliteration?
Add it to make a model refuse more, or monitor it as an early signal that the model is about to refuse or has been steered by a jailbreak.