llm-inference-explained

/about

About this site

LLM Inference Explained is the companion to the Transformer Decoder Explainer. That site shows what happens inside one forward pass of a decoder. This one shows what it takes to serve one: generating token by token, and doing it fast for many users at once.

Where the numbers come from

How it is built

Next.js 14 (App Router) with strict TypeScript, Tailwind, MDX chapters rendered on the server with KaTeX, and D3 for chart scales. The design system is copied from the Transformer Decoder Explainer so the two sites look like one. There is no database and no sign-in: every page is static, and the interactives compute in the browser. Vitest covers the maths; Playwright checks every chapter at desktop and phone widths in light and dark mode.

Source, tests and the README: github.com/BrendanJamesLynskey/llm-inference-explained.

Going further

The LLM Inference Simulators slide series covers how to model all of this in a simulator, with a glossary every chapter here links into. Start at the chapters.