/learn
Inference, chapter by chapter
Read in order: each chapter builds on the one before it. Toggle layers (Concept / Maths / Code) inside any chapter to choose how deep to go. New to transformers? Start with the Transformer Decoder Explainer, which walks through one forward pass.
Prefill, decode, sampling and stopping: why generation is sequential.
What is cached, why it is exact, and how big it gets.
Why decode is memory-bound and prefill compute-bound.
- 04Batching
Static, dynamic, continuous and chunked prefill.
PagedAttention, block tables, prefix sharing and preemption.
FlashAttention's IO-aware tiling; split-K for long contexts.
Draft, verify, and the expected tokens per step.
Fewer bytes per weight and per cached value, faster decode.
Tensor, pipeline and expert parallelism and what they send.
TTFT, TPOT, ITL, goodput, tail latency and Little's law.
Prefill stalls decode on a shared GPU; separate pools stop it.
Hand-off size, link bandwidth, streaming, compression, mixed pools.
Pools, devices, links, load and SLOs: run the simulator yourself.
Small models, low load and slow links, and where to go next.