LLM Inference Explained
How a model actually generates text, one token at a time.
A forward pass makes one prediction. Serving a model means running that pass thousands of times a second for many users at once. These chapters cover what real inference systems do about it: the KV cache, the roofline, batching, paged memory, FlashAttention, speculative decoding, quantisation, parallelism, the metrics that judge them, and disaggregated prefill and decode, with a live simulator. Every interactive runs tested code in your browser.
Chapters
Fourteen chapters, from the generation loop to disaggregated serving. Toggle the Concept, Maths and Code layers to choose your depth.
See the chapters →
The KV cache, live
A real tiny decoder generating in your browser, its cache filling column by column, checked against recomputation at every step.
Watch it fill →
The roofline
Why decode waits on memory and prefill on arithmetic, with H100 and A100 figures from a tested simulator.
Open the roofline →
The companion sites: the Transformer Decoder Explainer walks through the forward pass itself, LLM Architectures Explained shows how real models’ designs differ, GPU Kernels Explained shows how a GPU executes the maths, Numerics Explained shows the number formats and quantisation behind it, and Systolic Arrays Explained shows the matrix hardware of TPUs. The source of this one is on GitHub.