edidiong umana · writing
home
Engineering notes6 min read

Why your local model is slow: prefill, decode and the memory wall

A long pause, then a slow trickle: two different problems with two different fixes. How to measure each one, and how to work out your laptop's speed limit with a pencil before you buy anything.

You pull a model with Ollama, ask it a question, and wait. Nothing happens for a few seconds. Then the words arrive, slower than you can read them. It is tempting to blame "the model" or "my laptop" and move on.

But the pause and the trickle are two different jobs inside the model. They have different bottlenecks, and they need different fixes. Here is how I break it down, with numbers you can check on your own machine.

Two halves of every answer

Every answer runs in two phases.

  • Prefill reads your whole prompt in one pass, all the tokens together. It keeps what it worked out about each token in the KV cache, and scores the first token of the answer.
  • Decode writes the answer one token per pass. Each pass adds the newest token, reuses the KV cache, and runs the model again, until it reaches a stop token or your limit.

Prefill has many tokens to process at once, so it keeps the chip's maths units busy. It is compute-bound, and it grows with the length of your prompt. Decode does very little maths per pass, but it still has to read every weight in the model from memory. It is memory-bound, and it grows with the length of the answer. The difference is big enough that research systems such as DistServe and Splitwise run the two phases on separate GPUs.

Ollama shows you both. Run ollama run qwen3:4b-instruct --verbose and read the block under the answer. The prompt eval lines are prefill. The eval lines are decode, and the eval rate is the pace of the stream. load duration is neither: it is the model being loaded into memory, which takes seconds the first time and is near zero while the model stays loaded.

Measure what users feel

Three numbers describe what a person using your app experiences:

  • Time to first token (TTFT): from sending the request to the first token arriving. It includes any queue, the network and prefill.
  • Time per output token (TPOT): the average gap between tokens after the first. Tokens per second is the same number upside down: 20 ms per token is 50 tokens a second.
  • End-to-end latency: the whole answer, which is TTFT + TPOT × (tokens − 1).

Benchmark tools don't all define these the same way, so check before you compare two numbers. And report the median and p95 rather than the average. Nine answers in 1 second and one in 9 seconds average 1.8 seconds, a wait nobody actually had.

This script times any OpenAI-compatible server, using only the standard library. It defaults to Ollama, which serves that API under /v1:

import json, sys, time, urllib.request, uuid

MODEL = sys.argv[1] if len(sys.argv) > 1 else "qwen3:4b-instruct"
BASE = sys.argv[2] if len(sys.argv) > 2 else "http://127.0.0.1:11434/v1"  # 127.0.0.1, not localhost
PROMPT = "In about 120 words, explain why the sky is blue."

def one_run():
    body = {"model": MODEL, "stream": True, "max_tokens": 150, "temperature": 0,
            "stream_options": {"include_usage": True},
            # a fresh start on every prompt, so a cached prefix can't flatter the TTFT
            "messages": [{"role": "user", "content": f"[{uuid.uuid4().hex[:8]}] {PROMPT}"}]}
    req = urllib.request.Request(BASE + "/chat/completions", json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    t0, stamps, usage = time.perf_counter(), [], None
    with urllib.request.urlopen(req, timeout=600) as resp:
        for raw in resp:                       # server-sent events, one "data:" line each
            line = raw.decode("utf-8").strip()
            if not line.startswith("data:") or line[5:].strip() == "[DONE]":
                continue
            chunk = json.loads(line[5:])
            usage = chunk.get("usage") or usage
            if any((c.get("delta") or {}).get("content") for c in chunk.get("choices", [])):
                stamps.append(time.perf_counter())
    if len(stamps) < 2:
        sys.exit("Too few tokens streamed. Is the server running and the model pulled?")
    n = (usage or {}).get("completion_tokens") or len(stamps)
    return stamps[0] - t0, (n - 1) / (stamps[-1] - stamps[0]), n

one_run()                                      # warm-up: may include loading the model
for i in range(5):
    ttft, rate, n = one_run()
    print(f"run {i + 1}: TTFT {ttft * 1000:6.0f} ms   decode {rate:5.1f} tok/s   {n} tokens")

It avoids three traps that fake the numbers. It throws away a warm-up run, which can include loading the model. It puts a random tag at the start of each prompt, so the engine can't reuse a cached prompt and hand you a flattering first token. And it uses 127.0.0.1 and time.perf_counter(): on Windows, "localhost" tries IPv6 first, which added about 2 seconds to every request on one test machine, and before Python 3.13 time.time() only ticks every 15.6 ms.

Why decode is slow: the memory wall

To write one token, the model passes it through every layer and uses every weight once. So every weight has to travel from memory to the processor, for every single token. The maths is tiny by comparison, about two operations per weight, and the processor finishes long before the next weights arrive. It spends most of its time waiting.

That gives a rule you can work out with a pencil:

fastest decode (tokens per second) ≈ memory bandwidth ÷ bytes read per token

The bytes read per token are the model file plus the KV cache so far. Inference guides from Baseten, Databricks and NVIDIA all build on this rule.

The KV cache deserves its own sentence. Attention looks back at every earlier token, so in every layer each token gets a key and a value, and the cache keeps them instead of recomputing the whole conversation for each new token. Its size per token is 2 × layers × KV heads × head size × bytes per number. For Llama 3.1 8B at 16 bits that is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, or 128 KiB: 1.07 GB at 8K tokens and 17.2 GB at 128K, more than three times the 4-bit weights. A long context costs you twice, in memory and in speed, because the cache is read for every token too.

A worked example

Take qwen3:4b-instruct, a 2.5 GB download, on a laptop with DDR4-3200 memory.

  1. Find the bandwidth. 3,200 MT/s × 8 bytes = 25.6 GB/s for one memory channel. One stick is one channel. Two matching sticks usually run as two channels: 51.2 GB/s.
  2. Divide by the model. 25.6 ÷ 2.5 = 10.2 tokens a second on one stick. 51.2 ÷ 2.5 = 20.5 on two.
  3. Add the KV cache. Qwen3-4B needs 144 KiB of cache per token. At 4,096 tokens of context that is 144 × 1,024 × 4,096 bytes, about 0.60 GB, read for every token as well. On two sticks: 51.2 ÷ (2.5 + 0.6) = 16.5 tokens a second.
  4. Expect less than the ceiling. Real speeds land at roughly 70 to 85% of it. Peak bandwidth is a best case, and attention, unpacking the 4-bit numbers and sampling take time too.

The measurements agree. llama.cpp's community benchmark ran Llama 2 7B at 4 bits, a 3.8 GB file, on many machines. An Apple M1 at 68 GB/s decoded 14.19 tokens a second, 80% of its ceiling of about 17.8. The same model at 8 bits, 7.2 GB, decoded only 7.92 tokens a second on the same chip. Half the bytes nearly doubled decode. Prefill barely moved, 108.21 against 107.81 tokens a second, because it was never waiting on memory.

What actually speeds it up

  • Fewer bytes. A 4-bit file instead of an 8-bit one roughly doubles decode. So does a smaller model.
  • More bandwidth. A second matching memory stick gives you dual channel. A graphics card is better still, if the whole model fits on it. Run ollama ps while it answers: PROCESSOR should read 100% GPU or 100% CPU, because a split runs part of every token at the slower speed.
  • Not more threads. Johannes Gäßler, a llama.cpp developer, found that about five threads saturate dual-channel memory, and that more threads than cores can slow things down.
  • For the pause, a shorter prompt. In llama.cpp's benchmarks the M1 prefills only about 7.6 times faster than it decodes, so a 1,000-token prompt waits about 9 seconds for its first token. An RTX 3060 gets there in about half a second. On a laptop without a strong GPU, long chat histories and pasted documents are what make the pause long.

If you are choosing a laptop for local models, buy bandwidth, not gigahertz. A used laptop with an RTX 3060 Laptop GPU (336 GB/s) can decode models that fit in its 6 GB several times faster than a new thin laptop without a graphics card (51 to 90 GB/s).

Do this todayWork out your ceiling: your bandwidth in GB/s divided by your model's file size in GB. Then run ollama run qwen3:4b-instruct --verbose, ask for a 200-word answer, and compare the eval rate with it. Well under 70% means something is wrong: check ollama ps for a CPU and GPU split, and Task Manager for a single memory stick.

Learn it properly

This post is a slice of Track 3 of the AI Study Group, Inference engineering basics. It's free, runs on an 8 GB laptop, and goes on to memory budgets, quantisation, batching and caching, with every number checked against its source. The track draws on Rohit Ghumare's open-source course AI Engineering from Scratch (MIT).

Sources