Concept
A serving system is judged by what users feel and what operators pay for. The standard vocabulary:
- TTFT, time to first token: from the request's arrival to its first output token. Queueing plus prefill. It is what makes a chat interface feel responsive.
- TPOT, time per output token: for one request, the average gap between its tokens after the first, (last token time − first token time) ÷ (output tokens − 1). It sets reading speed.
- ITL, inter-token latency: every individual gap. TPOT averages a request's gaps; ITL keeps them, so a single 200 ms stall, invisible in the TPOT, shows up in the ITL tail.
- End-to-end latency: arrival to last token.
- Throughput: output tokens (or requests) per second, over all users.
- Goodput: requests per second that met both SLOs, a TTFT target and a TPOT target. Throughput counts everything served; goodput counts only what was served well enough. DistServe frames serving capacity this way, as the highest request rate at which a target fraction of requests meet their SLOs.
Why percentiles, and why p99. Averages hide the tail, and the tail is what users remember. Means also move when a few requests are very slow. Report p50 for the typical experience and p90/p99 for the tail. Tails compound: if a page needs 100 independent requests and each is slow 1% of the time, 63% of page loads include at least one slow one.
Little's law ties the averages together: the average number of requests in the system, , equals the arrival rate times the average time each spends there, . It needs no assumptions about arrival patterns or service times, only a stable system. Use it as a sanity check on any simulator or load test: the widget measures directly, by sweeping arrivals and departures, and compares it with .
Push the arrival rate up and watch TTFT's tail break its SLO first: as the server saturates, requests queue for prefill. Tighten the TPOT SLO and watch goodput fall below throughput.
Interactive
Percentiles, SLOs, goodput and Little's law
300 requests on one H100 running Llama-3-8B (the batching chapter's simplified scheduler, max batch 32). Percentiles interpolate linearly, as Disaggregated_Inference_Sim's do.
Maths
For request with arrival , token times :
Percentiles interpolate linearly between order statistics: for sorted and , (numpy's default, and the simulator's).
Goodput .
Little's law: over a window that starts and ends empty, with and the mean time in system. The tail: .
Code
// src/lib/inference/metrics.ts (excerpt)
export function percentileSorted(s: ArrayLike<number>, p: number): number {
if (s.length === 0) return NaN;
const k = ((s.length - 1) * p) / 100;
const lo = Math.floor(k);
const hi = Math.ceil(k);
return s[lo]! + (s[hi]! - s[lo]!) * (k - lo);
}
The same function as _percentile_sorted in Disaggregated_Inference_Sim's
metrics.py, so numbers from this site and from the simulator are
computed the same way.