Wild West API

Where a model keeps its no

Recent interpretability work found that whether a chat model refuses is governed largely by one linear direction in its activation space. Push activations along it and the model refuses; remove it and the model stops.

One direction, many refusals

Across many kinds of "unsafe" request, the internal signal that precedes a refusal points the same way. That means refusal is not scattered through the network but concentrated in a direction you can measure by contrasting harmful and harmless prompts, then averaging the activation gap.

What you can do with it

Add the direction to a compliant model and it starts refusing harmless things; subtract it and a cautious model complies. Doing the subtraction permanently in the weights is abliteration. The finding is also useful defensively: a classifier can watch that direction to detect when a model is about to refuse or has been steered.

Where it breaks down

Linear is an approximation The refusal direction explains a large share of refusal behaviour, not all of it. Some refusals survive removal, and heavy edits can bleed into unrelated behaviour. It is a strong effect, not a clean switch.

FAQ

What is the refusal direction?

A single direction in a chat model’s residual stream activations that tracks whether it is about to refuse. Arditi et al. (2024) found it in a range of open chat models by contrasting activations on harmful and harmless prompts.

How is the refusal direction found?

Run the model on two prompt sets, one it refuses and one it answers, record the activations at the last token for each layer, and take the difference of the means. The layer whose direction best removes refusals when ablated is the one used.

What can you do with it besides abliteration?

Add it to make a model refuse more, or monitor it as an early signal that the model is about to refuse or has been steered by a jailbreak.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.