Wild West API

Quantization: smaller models, lower precision

Quantization is the process of storing a model's weights with fewer bits per number so the model uses less memory and runs faster, at some cost in accuracy.

What quantization is

A model's weights are billions of numbers. Trained models are usually stored at 16 bits per weight (FP16 or BF16), so a 27 billion parameter model needs about 54 GB just for weights. Quantization rounds those numbers to fewer bits: 8 bits halves the size, 4 bits quarters it. That is what makes it possible to run large models on consumer GPUs, and it also makes serving cheaper and faster for providers.

How it works

Weights are grouped into small blocks, often 32 to 256 values. Each block stores a scale (and sometimes an offset) at higher precision, and each weight is stored as a small integer that is multiplied by the scale when used. Smarter schemes spend more bits on the weights that matter most, using importance estimates from sample data (an "imatrix" in llama.cpp) or methods such as GPTQ and AWQ that minimize the error each layer introduces.

The KV cache can be quantized separately, which saves memory on long contexts at a small quality cost.

GGUF and the quant names

GGUF is the single-file model format used by llama.cpp and tools built on it, such as KoboldCpp, Ollama and LM Studio. It bundles weights, tokenizer and metadata. Its quant names describe bit width and method:

  • Q8_0: 8-bit, nearly indistinguishable from full precision.
  • Q6_K, Q5_K_M: very close to full quality.
  • Q4_K_M: the common sweet spot for local use.
  • IQ3, Q3_K, IQ2: smaller, with noticeable degradation, more so on smaller models.

Other ecosystems use formats like EXL2, AWQ, GPTQ and FP8.

What it does to roleplay

Down to about 5 or 4 bits, most people cannot tell the difference in casual chat. Below that, problems appear: forgotten details, broken formatting, more repetition, worse instruction following, garbled names. Larger models tolerate heavy quantization better than small ones, so a 70B model at 3 bits often beats a 13B at 8 bits. Low quants also make the probability distribution noisier, which is one reason samplers like min P help on local setups.

Quantization and a hosted API

When you use an API you do not pick the quantization; the provider runs the model in whatever format it serves. That is one reason the same named model can behave a little differently between providers. On Wild West API you choose a model by ID and the serving format is handled upstream. If you want to compare against a local GGUF, test with identical prompts and settings, since samplers and templates also differ between setups.

FAQ

Which quantization is best?

For local use, Q4_K_M or Q5_K_M is the usual balance of size and quality. Use Q6_K or Q8_0 if you have the memory; go below 4 bits only when you must.

Does quantization make a model dumber?

A little, and more at lower bit widths. Down to about 4 to 5 bits the loss is small for most uses.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.