Wild West API

Heretic explained: the tool that removes AI censorship automatically

Heretic is an open-source command line tool that takes a censored language model and produces an uncensored copy with no training run and no manual tuning. It sounds alarming. It is better understood as abliteration with the guesswork automated, and what it can and cannot do is narrower than the headlines suggest.

What Heretic is

Heretic is a free, open-source tool by the developer p-e-w, published on GitHub under the AGPL-3.0 licence. You give it the name of an open-weight model and it edits the weights so the model stops refusing, without retraining it. The README describes it as removing censorship, also called safety alignment, from transformer language models.

It does this with abliteration, the technique that finds the direction in a model that carries refusal and projects it out. The part Heretic adds is the search for good settings.

How it works

A hand-made abliteration needs a person to choose which layers to edit, how hard, and which direction to use, then check whether the model still works. Heretic turns that into an optimisation problem.

  • It measures a refusal direction as the difference between the model's internal activations on harmful prompts and on harmless ones.
  • It searches for the edit with a TPE optimiser from the Optuna library, trying many combinations of direction, strength and shape, separately for the attention and the MLP parts of the model.
  • It scores every attempt on two things at once: how many refusals remain, and the KL divergence from the original model, which is a measure of how far the outputs have drifted. It looks for edits that lower the first without raising the second.
  • You pick from the best trade-offs. The result is a set of Pareto-optimal edits, and you choose one, then save it locally, upload it to Hugging Face or chat with it to test.

The README gives the install and the run as two lines, and says that on an RTX 3090 with default settings it takes roughly 20 to 30 minutes to process a 4B model.

pip install -U heretic-llm
heretic Qwen/Qwen3-4B-Instruct-2507

What its own numbers show

The project reports one comparison, on Google's gemma-3-12b-it, counting refusals out of 100 harmful prompts and the KL divergence from the original on harmless ones. Lower is better for both.

ModelRefusalsKL divergence
gemma-3-12b-it (original)97 of 1000
mlabonne abliterated v2 (by hand)3 of 1001.04
huihui-ai abliterated (by hand)3 of 1000.45
p-e-w heretic (Heretic)3 of 1000.16

All three edits cut refusals from 97 to 3. The claim is that Heretic gets there while changing the model least. These are the author's figures from one model on one machine, and the README itself says the exact values can depend on platform and hardware, and that automated metrics are no substitute for human evaluation. Treat it as a good sign, not a guarantee for every model.

Why heretic shows up in model names

Because anyone can run it, people publish the output. A model with heretic in its name is one that went through this tool, in the same way abliterated in a name signals a manual or scripted edit. The label tells you the method, not the quality. Two heretic builds of the same base model can differ, because the search is run by whoever published it and they choose which trade-off to keep.

The README says most dense models, many multimodal ones and several mixture-of-experts designs are supported, and that pure state-space models and some research architectures are not.

Is Heretic scary?

It is fair to find it unsettling that one command can strip the refusals from a model. A few facts help put it in proportion.

  • It adds no knowledge. Heretic only removes the habit of refusing. Whatever the base model could already do, it can do after, and nothing it never learned appears.
  • It only works on open weights. You need the model files, so it cannot touch a hosted closed model such as GPT or Claude. The models it can edit are ones whose weights anyone could already change with older, slower methods.
  • It does not make a model smarter or more capable. Any edit can cost some quality, which is exactly what the KL divergence figure is there to limit.
  • The risk is real and sits with the person using the output. With no refusal in the loop, nothing stops a harmful request, so legality and authorisation are yours to check. That is true of every uncensored model, however it was made.
The honest summary Heretic lowers the effort needed to make an uncensored model from a skilled afternoon to one command. That is a real change in access, and also not a change in what the models know. The limits and risks of an edited model are covered in Abliterated models: advantages, limits and risks.

Using a heretic model without running Heretic

Running the tool needs a GPU with enough memory for the model and some patience. If you only want to use an uncensored model, you do not need any of that: Wild West API sells a short line of uncensored models behind an OpenAI and Anthropic compatible API, each key with a hard spend cap. We do not claim any model on the line was made with Heretic, so check the models page for what each one is. To judge any edited model, heretic or not, run it on your own prompts with the script in how to evaluate an abliterated model.

FAQ

What is Heretic AI?

Heretic is an open-source command line tool by p-e-w that removes the refusal behaviour from open-weight language models automatically. It combines abliteration with an optimiser that searches for the edit that removes the most refusals while changing the model the least.

How does Heretic remove censorship?

It measures the direction inside the model that carries refusal, then projects it out of the weights. An Optuna TPE optimiser tries many settings and scores each on remaining refusals and on KL divergence from the original model, and you choose from the best trade-offs.

Is Heretic better than manual abliteration?

On the one comparison the project publishes, gemma-3-12b-it, it matched two hand-made abliterations at 3 refusals in 100 prompts with a lower KL divergence, 0.16 against 0.45 and 1.04. That is the author's own result on one model, so test any build on your own prompts.

Is Heretic dangerous?

It adds no knowledge and works only on open weights, so it makes already-available models easier to uncensor rather than creating new capability. It does remove the refusal gate, which moves responsibility for the output to whoever uses it.

Is Heretic free?

Yes. It is open source under the AGPL-3.0 licence and installs with pip. You supply the hardware: a GPU with memory for the model, and roughly 20 to 30 minutes on an RTX 3090 for a 4B model.

What does heretic mean in a Hugging Face model name?

It marks a model that was processed with the Heretic tool, in the way abliterated marks one edited by abliteration. It names the method and says nothing about quality, which varies by who ran it.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.