/about
About this site
LLM Inference Explained is the companion to the Transformer Decoder Explainer. That site shows what happens inside one forward pass of a decoder. This one shows what it takes to serve one: generating token by token, and doing it fast for many users at once.
Where the numbers come from
- The live tiny model in chapters 1 and 2 is the Transformer Decoder Explainer’s pure-TypeScript transformer (verified against PyTorch to 1e-5), vendored here and extended with a KV cache. Its tests prove cached generation gives bit-identical logits to full recomputation. Its weights are random, so its text is gibberish; the arithmetic is real.
- Hardware and step times (H100, A100, Llama-3-8B and 70B, links) come from the cost model of Disaggregated_Inference_Sim, ported to TypeScript and checked in CI against fixtures that the Python package writes, and against its recorded results.
- The live simulator in chapters 11–14 is Disaggregated_Inference_Sim’s own JavaScript engine, copied byte for byte from a pinned commit by a script that records the commit and the file’s SHA-256. CI runs it against fixtures the Python package writes at the same commit (every request’s timestamps must match) and against the rows of its
results.mdthat the chapters quote. Its workloads are the ones Python recorded. - Everything else (batching, paging, kernels, speculative decoding, parallelism, metrics) is small, tested code in
src/lib/inference/. Where a model is simplified or a coefficient is illustrative, the chapter says so. - Papers are cited by arXiv ID and were checked against arXiv. Claims about serving engines (vLLM, SGLang, TensorRT-LLM and others) describe published designs and change often; follow the links for the current state.
How it is built
Next.js 14 (App Router) with strict TypeScript, Tailwind, MDX chapters rendered on the server with KaTeX, and D3 for chart scales. The design system is copied from the Transformer Decoder Explainer so the two sites look like one. There is no database and no sign-in: every page is static, and the interactives compute in the browser. Vitest covers the maths; Playwright checks every chapter at desktop and phone widths in light and dark mode.
Source, tests and the README: github.com/BrendanJamesLynskey/llm-inference-explained.
Going further
The LLM Inference Simulators slide series covers how to model all of this in a simulator, with a glossary every chapter here links into. Start at the chapters.