Run abliterated weights on your own machine
Abliterated models run in the same local tools as any open-weight model. The work is in picking the right file: a GGUF quantization small enough for your memory and large enough to keep the quality the abliteration did not already cost.
Pick the GGUF, not the full weights
Publishers release full weights as Safetensors, usually BF16, at about two bytes per parameter: a 24B model is close to 48GB. Desktop runtimes want a GGUF conversion instead, published either by the same author or by a quantizer such as bartowski, and named with the original model plus -GGUF.
Choosing a quantization
| Quant | Roughly, per billion parameters | Use it when |
|---|---|---|
| Q8_0 | about 1.1GB | You have memory to spare and want output closest to the original |
| Q6_K | about 0.8GB | Near-lossless at a useful saving |
| Q4_K_M | about 0.6GB | The usual starting point: big saving, small quality loss |
| Q3_K_M / IQ3 | about 0.5GB | The model only just does not fit at Q4 |
| Q2_K / IQ2 | about 0.35GB | Last resort; quality falls off sharply |
Add room for the context: the KV cache grows with every token of context you allow, so start at 8K and raise it only when you need to. An abliterated model has already given up a little quality, which is a reason to stay at Q4_K_M or above rather than squeeze lower.
Ollama
Ollama can pull a GGUF directly from Hugging Face by repository and quant tag, or wrap a file you downloaded with a two-line Modelfile.
# Pull a GGUF straight from Hugging Face and chat with it ollama run hf.co/mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated-GGUF:Q4_K_M # Or wrap a .gguf you downloaded yourself # Modelfile: # FROM ./model.Q4_K_M.gguf # PARAMETER num_ctx 8192 ollama create my-abliterated -f ./Modelfile ollama run my-abliterated
LM Studio
In the app, search Discover for the GGUF repository, check the publisher, download the quant you want and pick it from the model menu in Chat. The bundled lms command does the same from a terminal.
# LM Studio's CLI (open the app once first) lms get mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated-GGUF lms ls --llm lms load <model_key> --context-length 8192 lms chat
Local or hosted
| Local | Hosted (Wild West API) | |
|---|---|---|
| Hardware | Your GPU or unified memory sets the ceiling | None |
| Model size | Whatever fits, usually 8B to 32B | Bases larger than most machines can load |
| Context | Limited by memory, often 8K to 32K in practice | Up to a million tokens on the current line |
| Tool calling for agents | Depends on the model and the runtime | Every model on the line, tested |
| Offline | Yes | No |
| Cost | The hardware and the power | Per token, under a hard cap per key |
Local wins when the data must never leave the machine or there is no network. Hosted wins when you need a bigger model, a long context or an agent that calls tools reliably. The Ollama alternative page covers the switch, and any OpenAI client points at https://wildwestapi.com/v1 by changing one URL.
FAQ
Can I run an abliterated model on my computer?
Yes. Download a GGUF quantization of it and load it in Ollama, LM Studio or llama.cpp. An 8B model at Q4_K_M needs around 5GB of memory plus context, so most recent laptops can run one.
What quantization should I use?
Start at Q4_K_M. Go up to Q6_K or Q8_0 if you have the memory, and only drop to Q3 or Q2 if nothing else fits, since quality falls off quickly below Q4.
How much VRAM does an abliterated model need?
The same as the original model, since abliteration does not change the size. At Q4_K_M, allow roughly 0.6GB per billion parameters plus a few GB for context.
Do abliterated models work offline?
Run locally, yes. Once the GGUF is downloaded nothing needs the network.