Quantisation without the hype: what Q4 really costs you
How to read names like Q4_K_M and Q8_0, work out whether a model fits before you download it, and measure the quality loss on your own task instead of trusting a leaderboard.
You open a model page and find a dozen files of the same model: Q2_K, Q4_K_M, Q8_0, fp16. Someone online says 4-bit is basically lossless. A leaderboard shows the quantised version barely behind the original. You have 16 GB of memory and a task that matters. Which file do you download?
The honest answer is "it depends on your task", which sounds like a dodge. It isn't, once you know what quantisation changes and how to check it yourself in an afternoon.
What quantisation actually does
A model is billions of numbers, its weights, usually released at 16 bits each. Quantising stores each weight with fewer bits. Take a small block of weights, find the largest, and pick a scale so that it maps to the largest integer you can store: 7 for 4 bits. Divide every weight by the scale, round it, and store the small integer. When the model runs, it multiplies back.
Here is one weight from a block whose largest value is 0.97. At 4 bits the scale is 0.97 ÷ 7, about 0.139. The weight 0.44 becomes 3, and comes back as 0.416. The error can never exceed half a step, about 0.069 here. On that same block the worst error is 0.004 at 8 bits and 0.440 at 2 bits, where most weights collapse to −1, 0 or 1 times the scale.
Three things change:
- Memory. Fewer bits per weight, so a smaller file and less RAM.
- Speed. Decoding reads every weight for every token, so fewer bytes means faster tokens. Llama 2 7B on an Apple M1 decoded 7.92 tokens a second at 8 bits and 14.19 at 4 bits, in llama.cpp's benchmarks. Reading the prompt barely changed.
- Quality. The rounding loses precision for good. How much that matters depends on the bits, the model and the task.
Reading the names
Four families turn up on model pages:
| Format | What it is | Where you meet it |
|---|---|---|
| GGUF k-quants: Q8_0, Q6_K, Q5_K_M, Q4_K_M | llama.cpp's block formats, each with its own scale per small block | Ollama, llama.cpp and LM Studio on laptops |
| GPTQ, AWQ | 4-bit weights with a scale for each group of 128, tuned on sample data | GPU servers such as vLLM |
| INT8 | 8-bit weights, sometimes with 8-bit maths too | GPU servers |
| FP8 | 8-bit floating point, run natively by recent NVIDIA GPUs | Data-centre serving, not a CPU laptop |
Two details save confusion. First, the number after the Q is roughly the bits, and the letters after it name a mix. A "4-bit" Q4_K_M file is really about 4.9 bits per weight, because it keeps some sensitive layers at 6 bits and every block stores its own scale. That is why an 8B model at Q4_K_M is 4.9 GB, not 4.0.
Second, on Ollama the default tag is usually Q4_K_M, and tags such as -q8_0 or -fp16 give you bigger, more precise files. Read the tag, not just the family name: qwen3:4b now points to a thinking version that reasons at length before answering, while qwen3:4b-instruct answers directly.
Will it fit? Do the sum first
The weights need parameters × bits per weight ÷ 8 bytes. With parameters in billions, the answer is in gigabytes. Then add the overhead: the KV cache, which grows with your context; about 0.5 GB for the runtime; and 3 to 4 GB for your system and apps.
Here is Llama 3.1 8B, which has 8.03 billion parameters, at an 8,192-token context. Its KV cache at 16 bits is 128 KiB per token, so 1.07 GB at that length.
| Format | Bits per weight | Weights | + KV cache + runtime | Total | 16 GB laptop? |
|---|---|---|---|---|---|
| 16-bit | 16 | 16.06 GB | 1.07 + 0.5 GB | 17.63 GB | No |
| Q8_0 | 8.5 | 8.53 GB | 1.07 + 0.5 GB | 10.11 GB | Yes |
| Q6_K | 6.6 | 6.62 GB | 1.07 + 0.5 GB | 8.20 GB | Yes |
| Q5_K_M | 5.7 | 5.72 GB | 1.07 + 0.5 GB | 7.30 GB | Yes |
| Q4_K_M | 4.9 | 4.92 GB | 1.07 + 0.5 GB | 6.49 GB | Yes |
| Q2_K | 3.2 | 3.21 GB | 1.07 + 0.5 GB | 4.79 GB | Yes |
"Fits" here means the total stays under about 12.5 GB, leaving 3.5 GB for the system, and totals are worked out before rounding. The bits per weight are llama.cpp's measured averages, and the results match the real downloads: Ollama's Q4_K_M file for this model is 4.9 GB, its Q8_0 file 8.5 GB and its 16-bit file 16.1 GB.
On an 8 GB laptop with about 4.5 GB free, none of these fit. A 4B model does: qwen3:4b-instruct needs 2.5 GB of weights, about 0.6 GB of cache at 4,096 tokens and 0.5 GB of runtime, roughly 3.6 GB.
One trap: quantising the weights does not shrink the KV cache. Gemma 3's technical report shows the same 4.7 GB cache at 32K tokens with 16-bit or 4-bit weights. In Ollama, OLLAMA_KV_CACHE_TYPE=q8_0 halves the cache, usually with no noticeable loss.
What the quality numbers say
llama.cpp measured each format on Llama 3 8B against the 16-bit original:
| Format | Perplexity change | Same top token |
|---|---|---|
| Q8_0 | +0.04% | 97.7% |
| Q6_K | +0.35% | 96.0% |
| Q4_K_M | +2.8% | 91.9% |
| Q3_K_M | +10.5% | not given |
| Q2_K | +56% | 71.1% |
Perplexity is how surprised the model is by real text, so lower is better. "Same top token" is how often its first choice matches the original's. The pattern: 8 bits is nearly free, 4 bits costs a little, and below 4 bits quality falls off a cliff. Q2_K fits on a 16 GB laptop, but it is a much worse model.
Two larger studies add detail. Red Hat's engineers ran over half a million evaluations on quantised Llama 3.1 models: FP8 was effectively lossless, and 4-bit weights kept about 96% of the 8B model's score on harder tasks, but a small 1.5B reasoning model kept only 93.5%. Cohere's researchers found that automatic scores understate the damage. Languages in non-Latin scripts suffered most, and maths degraded fastest. For Japanese, a 1.7% drop on automatic tests was a 16% drop in human ratings.
Why a leaderboard can't choose for you
Perplexity is useful for comparing formats, but it is blind to many of the errors your users would notice. Benchmark scores are averages over someone else's tasks, mostly in English. The research says the damage is uneven: worst for small models, reasoning, maths and languages other than English. If your users write in Yoruba, Hausa or Amharic, a leaderboard measured in English tells you very little about your case.
So the test that counts is yours. Here is the procedure I'd follow:
- Write 20 real questions per language, each with the answer you would accept.
- Fix the rule in advance. For example: keep the smaller file only if it passes in every language.
- Run both files at temperature 0 with the script below.
- Grade by hand: right, or wrong in a way that matters.
- Keep the smallest file that passes. If memory is tight, also try a smaller model at 8 bits against a larger one at 4.
import csv, json, urllib.request
BASE = "http://127.0.0.1:11434/v1/chat/completions"
TAGS = ["qwen3:4b-instruct", "qwen3:4b-instruct-2507-q8_0"] # Q4_K_M and Q8_0
QUESTIONS = "questions.txt" # one real question per line, in every language your users use
def ask(tag, question):
body = {"model": tag, "temperature": 0, "max_tokens": 400,
"messages": [{"role": "user", "content": question}]}
req = urllib.request.Request(BASE, json.dumps(body).encode(),
{"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=600) as resp:
return json.load(resp)["choices"][0]["message"]["content"]
with open(QUESTIONS, encoding="utf-8") as f:
questions = [line.strip() for line in f if line.strip()]
with open("to_grade.csv", "w", newline="", encoding="utf-8") as out:
writer = csv.writer(out)
writer.writerow(["question", "model", "answer", "correct (y/n)"])
for q in questions:
for tag in TAGS:
writer.writerow([q, tag, ask(tag, q), ""])
print(f"{tag:30} {q[:50]}")
print("Done. Open to_grade.csv and mark every answer right or wrong.")
It uses only the standard library and talks to Ollama's OpenAI-compatible endpoint, so you can swap in any two tags you have pulled, or two different models.
questions.txt: 20 questions your users actually ask, in the languages they ask them in, with the answer you would accept noted beside each. That file will tell you more about which quantisation to ship than any leaderboard.Learn it properly
This post is a slice of Track 3 of the AI Study Group, Inference engineering basics. It's free and laptop-sized, and it covers memory budgets, the KV cache and quantisation with every number checked against its source. The track draws on Rohit Ghumare's open-source course AI Engineering from Scratch (MIT).
Sources
- llama.cpp quantize README: file sizes and bits per weight for each format
- llama.cpp perplexity README: quality tables for each GGUF format on Llama 3 8B
- Performance of llama.cpp on Apple Silicon: decode speed at Q4_0 and Q8_0
- Llama 3.1 tags on Ollama: real download sizes by format
- Give me BF16 or give me death? Accuracy-performance trade-offs in LLM quantization (Red Hat, 2024)
- How does quantization affect multilingual LLMs? (Cohere, 2024)
- AWQ: activation-aware weight quantization (2023): how 4-bit GPU formats protect the weights that matter
- Gemma 3 technical report (2025): KV cache memory with 16-bit and 4-bit weights
- Ollama FAQ: context length and KV cache quantisation