Wild West API

Ways to take the filter off an LLM

Uncensored is a description of behaviour, not a method. At least six different techniques produce a model that refuses less, and they differ in what they touch, what they cost in quality and whether the effect survives the next conversation.

The methods side by side

What it changesWhat it needsLasts
AbliterationThe weights, by projecting out the refusal directionA few hundred contrasting prompts and the model’s activationsPermanently, in every conversation
Uncensored finetuneThe weights, by more trainingA dataset of direct answers without refusalsPermanently
LoRA adapterA small trained add-on beside the weightsA finetuning dataset, much less computeWhile the adapter is loaded
Model mergeThe weights, by blending two or more releasesExisting models with the same architecturePermanently
Base modelNothing; the safety stages were never appliedA published pretrained checkpointAlways, but it is not an assistant
Jailbreak promptOnly the inputA prompt that slips past the trainingOne conversation, until patched

Abliteration

Measure the direction in the activations that separates refused prompts from answered ones, then remove it from every matrix that writes to the residual stream. No training run and no dataset beyond the probe prompts. It is the most direct attack on the mechanism, so it removes refusals more thoroughly than a finetune usually does, at a small cost to reasoning that a light retune can win back. Full method: abliteration.

Uncensored finetunes

Continue training on examples that answer directly, with the stock assistant disclaimers removed. Because it is ordinary training, the data also shifts style, persona and domain knowledge, which is why finetunes such as Dolphin, Hermes and Euryale each have a recognisable voice. They refuse much less but not never: the original alignment is diluted rather than removed.

LoRA adapters

A low-rank adapter trains a few small matrices that sit beside the frozen weights and nudge the output. It is the cheap way to finetune and the easy way to share one, since the adapter is megabytes rather than gigabytes. The behaviour only applies while the adapter is attached, so the same base can be served aligned or not.

Model merges

Merging averages or otherwise combines the weights of several releases of the same architecture, for example an aligned instruct model with an uncensored finetune of it. Results are unpredictable in both directions: a merge can keep the strong model’s skill and the permissive model’s willingness, or inherit the worst of each. Test before trusting.

Base models

A pretrained checkpoint that never went through instruction or safety tuning has no refusal habit, because refusal was never taught. It also has no assistant habit: it continues text rather than following instructions, so it needs careful few-shot prompting and will not call tools. Useful for research and data generation, rarely for an agent.

Jailbreak prompts

A jailbreak leaves the model untouched and wraps the request in role play, encoding or formatting that the safety training did not cover. Each one is found by trial, shared, and then trained against in the next release, and against a hosted model it can be caught by the platform’s filter regardless. They matter far more as a test of your own model than as a way to get work done.

Which one to use

If you are calling an API rather than training, the choice is between models, not methods: pick an abliterated build when refusals must be near zero, and an uncensored finetune when a stronger base matters more than the last few percent. Wild West API’s line has both, every model calls tools, and the models page says which is which.

FAQ

What is the best way to uncensor an LLM?

For open weights, abliteration removes refusals most thoroughly for the least compute, and an uncensored finetune is the better choice when you also want to change style or knowledge. For a closed model served by someone else there is no reliable way: you cannot edit its weights, and jailbreaks get patched.

Is a jailbroken model the same as an uncensored model?

No. A jailbreak is a prompt that tricks an unchanged, still-aligned model for one conversation. An uncensored model has had its weights changed, so it behaves the same way for every prompt without tricks.

Do base models refuse?

Rarely, because refusal is taught in the instruction and safety stages that a base model skipped. They also do not follow instructions reliably, so they are awkward to use as an assistant or an agent.

Can you uncensor a closed model like GPT or Claude?

Not in the weights, since they are never released. Only prompt-level jailbreaks reach them, and both the model training and the platform filters are updated against those.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.