Wild West API

The refusal direction in language models

The refusal direction is a single direction in a language model's internal activation space that, when present, makes the model refuse a request.

What the refusal direction is

Inside a transformer, each token is represented at each layer by a long vector of numbers called the residual stream. Researchers studying chat models found that whether a model refuses is largely controlled by one direction in that space. The 2024 paper "Refusal in Language Models Is Mediated by a Single Direction" by Arditi and colleagues showed this across many open chat models of different sizes.

How it was found

The method is a difference of means. Feed the model many prompts it refuses and many it answers. At a given layer and token position, average the activations for each group, then subtract one average from the other. The resulting vector is the candidate refusal direction. Researchers test candidates from different layers and pick the one with the strongest effect.

Two experiments confirm it matters:

  • Removing it (projecting it out of the activations) makes the model answer prompts it previously refused.
  • Adding it to activations on harmless prompts makes the model refuse things it would normally answer, such as a recipe request.

That two-way result is what makes it a cause of refusal rather than a side effect.

Why it matters

Refusal training shapes behavior, but it ends up stored in a compact way. That finding is the basis of abliteration: modify the weights so they cannot write anything along the refusal direction, and refusals mostly disappear while other abilities stay largely intact. It also explains why jailbreak prompts sometimes work: certain framings keep the activation along this direction below the point where the model refuses.

Limits and caveats

  • "Single direction" is an approximation. Removing it removes most hard refusals, but warnings, moralizing and topic-steering can involve other features.
  • The direction is found from a particular set of prompts. Behavior on categories that were not represented can differ.
  • Removing a direction can slightly degrade unrelated tasks, because the space is shared.
  • It explains refusal behavior; it says nothing about what the model knows or how well it writes.

Why roleplay users care

For fiction writers, refusals and soft refusals are the main friction with mainstream models: scenes skipped, villains softened, stories redirected. Models with the refusal direction removed, or trained without refusals in the first place, avoid that friction. Read more on abliterated models, and browse the models Wild West API serves.

FAQ

Who discovered the refusal direction?

It was described in the 2024 paper 'Refusal in Language Models Is Mediated by a Single Direction' by Arditi and coauthors, building on earlier work on linear features in model activations.

Can the refusal direction be added back?

Yes. Adding the direction to activations makes a model refuse even harmless requests, which is one of the experiments that showed it causes refusals.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.