llm-inference-explained

LLM Inference Explained

How a model actually generates text, one token at a time.

A forward pass makes one prediction. Serving a model means running that pass thousands of times a second for many users at once. These chapters cover what real inference systems do about it: the KV cache, the roofline, batching, paged memory, FlashAttention, speculative decoding, quantisation, parallelism, the metrics that judge them, and disaggregated prefill and decode, with a live simulator. Every interactive runs tested code in your browser.

recompute every stepwith the KV cachetokens generated →
Generating 24 tokens with a tiny decoder, computed on the server for this page. With the KV cache it does 7% of the arithmetic of recomputing everything each step, and its logits differ from the recomputed ones by exactly 0. See why

The companion sites: the Transformer Decoder Explainer walks through the forward pass itself, LLM Architectures Explained shows how real models’ designs differ, GPU Kernels Explained shows how a GPU executes the maths, Numerics Explained shows the number formats and quantisation behind it, and Systolic Arrays Explained shows the matrix hardware of TPUs. The source of this one is on GitHub.