Arithmetic intensity and the roofline
Why decode is memory-bound and prefill compute-bound.
Concept
A GPU step can be slow for two different reasons: it has too much arithmetic to do, or too many bytes to move from memory. The roofline model says a step takes whichever is longer:
- arithmetic time = FLOPs ÷ the chip's FLOP rate;
- memory time = bytes moved ÷ the chip's memory bandwidth.
Their ratio for a given step, FLOPs per byte, is its arithmetic intensity. The chip has a balance point too, its ridge point: peak FLOP rate ÷ peak bandwidth. Below the ridge a step is memory-bound: the arithmetic units wait for data. Above it the step is compute-bound.
With the H100 figures in the simulator this site borrows (989 TFLOP/s dense BF16 and 3.35 TB/s, derated to 55% and 80% of peak), the ridge sits at 203 FLOP per byte. For the A100 it is about 105.
Now look at what a decode step does. To produce one token per sequence it must read every weight of the model once (the whole layer stack and the output head) plus each sequence's KV cache. For Llama-3-8B at batch 1 with 2,048 tokens of context that is 15.28 GB read, for about 16 billion FLOPs: an intensity of 1.1 FLOP/B, two hundred times below the ridge. The step takes 6.201 ms, almost all of it waiting on memory: about 161 tokens per second, with the arithmetic units nearly idle.
Prefill is the opposite. A 2,048-token prompt multiplies every weight by 2,048 token vectors, so each byte read feeds about two thousand FLOPs (2,081.7 FLOP/B). It is compute-bound: 59.033 ms on one H100.
Move the sliders below and watch the decode point climb as the batch grows.
Interactive
Prefill and decode on the roofline
Roofs are the simulator's derated peaks (55% of peak FLOP/s, 80% of peak bandwidth; hardware.py). Llama-3-8B on 1× H100-SXM. Points are the cost model's steps; their height excludes the fixed 0.5 ms step overhead.
Concept
That last observation is the economic engine of LLM serving. Batch 64 sequences together and one decode step still reads the weights once, but now produces 64 tokens. The step takes 12.514 ms instead of 6.201, so throughput rises from about 161 to about 5,114 tokens per second while each user's per-token latency only doubles. Intensity rises to 32 FLOP/B: still memory-bound, because the KV cache grows with the batch and has to be read too.
So decode wants large batches (chapter 4), which need memory for their caches (chapter 5) and fewer bytes per step (chapter 8). Prefill needs none of that; it already saturates the chip. That difference is also why some systems run the two phases on separate machines, which the later chapters on disaggregation cover.
Maths
For a step with FLOPs and bytes on a device with FLOP rate and bandwidth (both derated):
memory-bound when . The simulator adds a fixed ms per step (scheduling and kernel launches) and caps power at the board's TDP.
For a transformer with matmul parameters, layers and width , a decode step over batch with total context costs
with bytes per weight and KV bytes per token: the weights once, one embedding row per sequence, and the cache. Prefill over tokens is , .
Code
// src/lib/inference/costModel.ts (excerpt), ported from hardware.py
decode(ctx: number, batch: number): StepCost {
const flops = 2 * matmulParams(m) * batch + 4 * m.n_layers * m.d_model * (ctx + batch);
const nbytes = weightBytesRead(m, batch) + (ctx + batch) * kvBytesPerToken(m);
return time(flops, nbytes); // max(compute, memory) + overhead, TDP-capped
}
tests/unit/inference/costModel.test.ts checks this port against 1,284
steps written by the Python simulator (exactly, except the documented
cube-root tolerance in the power-capped branch) and against every row of
its results.md §1.