The field guide
AbliterationAbliteration removes the one direction in a model’s activations that produces refusals, by editing the weights, with no retraining. How it works, step by step, and what it costs.
Refusal directionA single direction in a model’s residual stream carries its tendency to refuse. Finding it is the basis of abliteration.
Uncensored vs abliteratedUncensored usually means a loosely aligned finetune that still refuses sometimes. Abliterated means the refusal direction was removed. They are different claims.
Why LLMs refuseRefusals are trained in after pretraining, by instruction tuning and preference training, and some are added by the platform in front of the model. Where each one comes from and how to tell them apart.
Ways to uncensor an LLMAbliteration, uncensored finetunes, LoRA adapters, model merges, base models and jailbreak prompts all reduce refusals. What each one changes, what it costs, and how long it lasts.
Abliteration: pros, limits, risksAbliteration cuts false refusals cheaply and keeps the original instruction tuning, but it can cost capability, leave some refusals in place and make a model too agreeable. What to expect before you rely on one.
Evaluating abliterated modelsTest an abliterated model for what it gained and what it lost: refusal rate on held-out prompts, indirect refusals, capability against the original, agreeableness and tool calling. With a script.
Running abliterated models locallyRun an abliterated GGUF on your own machine with Ollama or LM Studio, pick a quantization that fits your memory, and know when a hosted API is the better trade.
Well known abliterated modelsThe community abliterated releases people search for most: Llama 3.1 8B, Qwen3 14B, Gemma 3 27B, Mistral Small 3.2 and DeepSeek R1 Distill 32B. Who made each, what it is based on, and what it keeps.
Uncensored LLM APIAn uncensored LLM API serves models that do not refuse, with no moderation layer in front. Where refusals actually come from, how to test an API before you trust it, and what to check.
API without a content filterSome LLM APIs run a classifier on every prompt and reply and block what it flags. Which ones do by default, how that differs from a model refusing, and what an unfiltered API really means.
OpenAI base URLThe OpenAI base URL is https://api.openai.com/v1 by default. How to point the Python and Node SDKs, environment variables, LangChain, LlamaIndex, the Vercel AI SDK and LiteLLM at another endpoint.
LLM gatewayAn LLM gateway is one API in front of many model providers, adding keys, billing, spend limits and logging without changing your code.
Per-key spend capsA spend cap that is checked before the call, not after, so a runaway or leaked key stops at the number you set instead of billing past it.
max_tokens clampingWild West API lowers the outbound max_tokens to whatever the remaining budget covers, so a response cannot cost more than was reserved.
OpenAI-compatible APIAn OpenAI-compatible API accepts the same requests as OpenAI’s chat completions endpoint, so any OpenAI SDK works by changing the base URL. Working curl, Python and Node examples.
Anthropic-compatible APIAn Anthropic-compatible API accepts the Messages format, so Claude Code and the Anthropic SDKs work by changing the base URL.
Red-teaming LLMsRed-teaming an LLM means attacking your own model on purpose: jailbreaks, prompt injection and abuse generation, to find failures before an attacker does.
Integration guides · Compare Wild West API · Models
Uncensored AI models on one key
OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.