edidiong umana · writing
home
Engineering notes5 min read

Quantisation without the hype: what Q4 really costs you

How to read names like Q4_K_M and Q8_0, work out whether a model fits before you download it, and measure the quality loss on your own task instead of trusting a leaderboard.

You open a model page and find a dozen files of the same model: Q2_K, Q4_K_M, Q8_0, fp16. Someone online says 4-bit is basically lossless. A leaderboard shows the quantised version barely behind the original. You have 16 GB of memory and a task that matters. Which file do you download?

The honest answer is "it depends on your task", which sounds like a dodge. It isn't, once you know what quantisation changes and how to check it yourself in an afternoon.

What quantisation actually does

A model is billions of numbers, its weights, usually released at 16 bits each. Quantising stores each weight with fewer bits. Take a small block of weights, find the largest, and pick a scale so that it maps to the largest integer you can store: 7 for 4 bits. Divide every weight by the scale, round it, and store the small integer. When the model runs, it multiplies back.

Here is one weight from a block whose largest value is 0.97. At 4 bits the scale is 0.97 ÷ 7, about 0.139. The weight 0.44 becomes 3, and comes back as 0.416. The error can never exceed half a step, about 0.069 here. On that same block the worst error is 0.004 at 8 bits and 0.440 at 2 bits, where most weights collapse to −1, 0 or 1 times the scale.

Three things change:

  • Memory. Fewer bits per weight, so a smaller file and less RAM.
  • Speed. Decoding reads every weight for every token, so fewer bytes means faster tokens. Llama 2 7B on an Apple M1 decoded 7.92 tokens a second at 8 bits and 14.19 at 4 bits, in llama.cpp's benchmarks. Reading the prompt barely changed.
  • Quality. The rounding loses precision for good. How much that matters depends on the bits, the model and the task.

Reading the names

Four families turn up on model pages:

FormatWhat it isWhere you meet it
GGUF k-quants: Q8_0, Q6_K, Q5_K_M, Q4_K_Mllama.cpp's block formats, each with its own scale per small blockOllama, llama.cpp and LM Studio on laptops
GPTQ, AWQ4-bit weights with a scale for each group of 128, tuned on sample dataGPU servers such as vLLM
INT88-bit weights, sometimes with 8-bit maths tooGPU servers
FP88-bit floating point, run natively by recent NVIDIA GPUsData-centre serving, not a CPU laptop

Two details save confusion. First, the number after the Q is roughly the bits, and the letters after it name a mix. A "4-bit" Q4_K_M file is really about 4.9 bits per weight, because it keeps some sensitive layers at 6 bits and every block stores its own scale. That is why an 8B model at Q4_K_M is 4.9 GB, not 4.0.

Second, on Ollama the default tag is usually Q4_K_M, and tags such as -q8_0 or -fp16 give you bigger, more precise files. Read the tag, not just the family name: qwen3:4b now points to a thinking version that reasons at length before answering, while qwen3:4b-instruct answers directly.

Will it fit? Do the sum first

The weights need parameters × bits per weight ÷ 8 bytes. With parameters in billions, the answer is in gigabytes. Then add the overhead: the KV cache, which grows with your context; about 0.5 GB for the runtime; and 3 to 4 GB for your system and apps.

Here is Llama 3.1 8B, which has 8.03 billion parameters, at an 8,192-token context. Its KV cache at 16 bits is 128 KiB per token, so 1.07 GB at that length.

FormatBits per weightWeights+ KV cache + runtimeTotal16 GB laptop?
16-bit1616.06 GB1.07 + 0.5 GB17.63 GBNo
Q8_08.58.53 GB1.07 + 0.5 GB10.11 GBYes
Q6_K6.66.62 GB1.07 + 0.5 GB8.20 GBYes
Q5_K_M5.75.72 GB1.07 + 0.5 GB7.30 GBYes
Q4_K_M4.94.92 GB1.07 + 0.5 GB6.49 GBYes
Q2_K3.23.21 GB1.07 + 0.5 GB4.79 GBYes

"Fits" here means the total stays under about 12.5 GB, leaving 3.5 GB for the system, and totals are worked out before rounding. The bits per weight are llama.cpp's measured averages, and the results match the real downloads: Ollama's Q4_K_M file for this model is 4.9 GB, its Q8_0 file 8.5 GB and its 16-bit file 16.1 GB.

On an 8 GB laptop with about 4.5 GB free, none of these fit. A 4B model does: qwen3:4b-instruct needs 2.5 GB of weights, about 0.6 GB of cache at 4,096 tokens and 0.5 GB of runtime, roughly 3.6 GB.

One trap: quantising the weights does not shrink the KV cache. Gemma 3's technical report shows the same 4.7 GB cache at 32K tokens with 16-bit or 4-bit weights. In Ollama, OLLAMA_KV_CACHE_TYPE=q8_0 halves the cache, usually with no noticeable loss.

What the quality numbers say

llama.cpp measured each format on Llama 3 8B against the 16-bit original:

FormatPerplexity changeSame top token
Q8_0+0.04%97.7%
Q6_K+0.35%96.0%
Q4_K_M+2.8%91.9%
Q3_K_M+10.5%not given
Q2_K+56%71.1%

Perplexity is how surprised the model is by real text, so lower is better. "Same top token" is how often its first choice matches the original's. The pattern: 8 bits is nearly free, 4 bits costs a little, and below 4 bits quality falls off a cliff. Q2_K fits on a 16 GB laptop, but it is a much worse model.

Two larger studies add detail. Red Hat's engineers ran over half a million evaluations on quantised Llama 3.1 models: FP8 was effectively lossless, and 4-bit weights kept about 96% of the 8B model's score on harder tasks, but a small 1.5B reasoning model kept only 93.5%. Cohere's researchers found that automatic scores understate the damage. Languages in non-Latin scripts suffered most, and maths degraded fastest. For Japanese, a 1.7% drop on automatic tests was a 16% drop in human ratings.

Why a leaderboard can't choose for you

Perplexity is useful for comparing formats, but it is blind to many of the errors your users would notice. Benchmark scores are averages over someone else's tasks, mostly in English. The research says the damage is uneven: worst for small models, reasoning, maths and languages other than English. If your users write in Yoruba, Hausa or Amharic, a leaderboard measured in English tells you very little about your case.

So the test that counts is yours. Here is the procedure I'd follow:

  1. Write 20 real questions per language, each with the answer you would accept.
  2. Fix the rule in advance. For example: keep the smaller file only if it passes in every language.
  3. Run both files at temperature 0 with the script below.
  4. Grade by hand: right, or wrong in a way that matters.
  5. Keep the smallest file that passes. If memory is tight, also try a smaller model at 8 bits against a larger one at 4.
import csv, json, urllib.request

BASE = "http://127.0.0.1:11434/v1/chat/completions"
TAGS = ["qwen3:4b-instruct", "qwen3:4b-instruct-2507-q8_0"]   # Q4_K_M and Q8_0
QUESTIONS = "questions.txt"    # one real question per line, in every language your users use

def ask(tag, question):
    body = {"model": tag, "temperature": 0, "max_tokens": 400,
            "messages": [{"role": "user", "content": question}]}
    req = urllib.request.Request(BASE, json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=600) as resp:
        return json.load(resp)["choices"][0]["message"]["content"]

with open(QUESTIONS, encoding="utf-8") as f:
    questions = [line.strip() for line in f if line.strip()]
with open("to_grade.csv", "w", newline="", encoding="utf-8") as out:
    writer = csv.writer(out)
    writer.writerow(["question", "model", "answer", "correct (y/n)"])
    for q in questions:
        for tag in TAGS:
            writer.writerow([q, tag, ask(tag, q), ""])
            print(f"{tag:30} {q[:50]}")
print("Done. Open to_grade.csv and mark every answer right or wrong.")

It uses only the standard library and talks to Ollama's OpenAI-compatible endpoint, so you can swap in any two tags you have pulled, or two different models.

Do this todayWrite questions.txt: 20 questions your users actually ask, in the languages they ask them in, with the answer you would accept noted beside each. That file will tell you more about which quantisation to ship than any leaderboard.

Learn it properly

This post is a slice of Track 3 of the AI Study Group, Inference engineering basics. It's free and laptop-sized, and it covers memory budgets, the KV cache and quantisation with every number checked against its source. The track draws on Rohit Ghumare's open-source course AI Engineering from Scratch (MIT).

Sources