Wild West API

Why a language model says no

A language model is not born refusing. Pretraining only teaches it to continue text. Refusal is a behaviour added later, in the stages that turn a text predictor into an assistant, and a second layer of refusals often lives outside the model entirely, in the platform that serves it.

The stages that teach a model to say no

  1. Pretraining. The model learns to predict the next token over a huge corpus. It has no notion of a request or a refusal; asked a question, it might answer, or might write three more questions.
  2. Instruction tuning. Supervised examples of a request followed by a helpful answer teach it the assistant format. Some of those examples are a request followed by a polite refusal, and that is where the habit starts.
  3. Preference training. RLHF, DPO and their relatives rank pairs of answers. Raters, or a reward model, prefer the refusal for whole categories of request, so the model learns that declining those categories is the high scoring move.
  4. Safety training. A targeted pass with adversarial and harmful prompts hardens the refusal so rephrasing does not get around it.

The result is a learned behaviour stored in the weights like any other. Interpretability work found that it concentrates in one place: a single refusal direction in the model’s activations, which is what abliteration removes.

Over-refusal: when the training generalises too far

The training sees examples, not rules, so the model learns surface features of what gets declined. A question about how a buffer overflow works shares words with a request to exploit a live system, and a cautious model treats them the same. Security researchers, penetration testers, malware analysts and fiction writers meet this most, because their legitimate work uses the vocabulary the training learned to avoid.

Refusals that do not come from the model

Many hosted APIs put a moderation classifier in front of the model, behind it, or both. The request is scored, and if it crosses a threshold the platform returns an error or a canned message before the model ever sees it. Swapping in a more permissive model changes nothing when the block is the platform’s.

Model refusalPlatform filter
Where it livesIn the weightsA classifier in the serving stack
What it looks likeA normal reply that declines, often with a suggestionAn error code, a policy message, or a cut-off stream
Fixed by a different modelYesNo
Fixed by a different hostNoYes

Wild West API runs no moderation layer of its own: it prices, caps and passes the request through, so a refusal on Wild West API comes from the model. See LLM API without a content filter.

Getting past each kind

A platform filter goes away by changing host. A model refusal goes away by changing the weights, which is what an uncensored finetune or an abliterated build is. Prompt tricks sit in between and are the least reliable of the three; the full list is in ways to uncensor an LLM.

FAQ

Why does ChatGPT refuse some requests?

Two reasons that stack. The model was trained, through instruction tuning and preference training, to decline certain categories of request, and the platform also runs moderation classifiers that can block a request or a reply before you see it.

Are refusals hard coded?

Not in the model. A refusal is a learned behaviour stored in the weights, like any other habit the model picked up in training. Platform filters are closer to hard coded: a separate classifier with a threshold.

Why do models refuse harmless requests?

The training generalises from surface features. A harmless question that shares vocabulary with requests the model was taught to decline can trigger the same refusal. This is called over-refusal, and it is most common in security, medicine and fiction.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.