llm-inference-explained
← /learn · 10

Serving metrics

TTFT, TPOT, ITL, goodput, tail latency and Little's law.

Concept

A serving system is judged by what users feel and what operators pay for. The standard vocabulary:

  • TTFT, time to first token: from the request's arrival to its first output token. Queueing plus prefill. It is what makes a chat interface feel responsive.
  • TPOT, time per output token: for one request, the average gap between its tokens after the first, (last token time − first token time) ÷ (output tokens − 1). It sets reading speed.
  • ITL, inter-token latency: every individual gap. TPOT averages a request's gaps; ITL keeps them, so a single 200 ms stall, invisible in the TPOT, shows up in the ITL tail.
  • End-to-end latency: arrival to last token.
  • Throughput: output tokens (or requests) per second, over all users.
  • Goodput: requests per second that met both SLOs, a TTFT target and a TPOT target. Throughput counts everything served; goodput counts only what was served well enough. DistServe frames serving capacity this way, as the highest request rate at which a target fraction of requests meet their SLOs.

Why percentiles, and why p99. Averages hide the tail, and the tail is what users remember. Means also move when a few requests are very slow. Report p50 for the typical experience and p90/p99 for the tail. Tails compound: if a page needs 100 independent requests and each is slow 1% of the time, 63% of page loads include at least one slow one.

Little's law ties the averages together: the average number of requests in the system, LL, equals the arrival rate λ\lambda times the average time each spends there, WW. It needs no assumptions about arrival patterns or service times, only a stable system. Use it as a sanity check on any simulator or load test: the widget measures LL directly, by sweeping arrivals and departures, and compares it with λW\lambda W.

Push the arrival rate up and watch TTFT's tail break its SLO first: as the server saturates, requests queue for prefill. Tighten the TPOT SLO and watch goodput fall below throughput.

Interactive

Percentiles, SLOs, goodput and Little's law

300 requests on one H100 running Llama-3-8B (the batching chapter's simplified scheduler, max batch 32). Percentiles interpolate linearly, as Disaggregated_Inference_Sim's do.

Scheduler
00.50.90.991SLOTTFT (ms), 0 – 75000.50.90.991SLOTPOT (ms), 0 – 38
TTFT p50 / p99
31 / 137 ms
TPOT p50 / p99
7.6 / 11.1 ms
ITL p99 / max
48.7 / 233 ms
Throughput
688 tok/s
SLO met
100.0%
both TTFT and TPOT
Goodput
5.90 req/s
offered 6 req/s
L (measured)
5.664
time-average in system
λ·W
5.664
arrival rate × mean latency

Maths

For request rr with arrival ara_r, token times τr,1<⋯<τr,nr\tau_{r,1} < \dots < \tau_{r,n_r}:

TTFTr=τr,1−ar,TPOTr=τr,nr−τr,1nr−1,ITLr,i=τr,i−τr,i−1.\text{TTFT}_r = \tau_{r,1} - a_r,\qquad \text{TPOT}_r = \frac{\tau_{r,n_r} - \tau_{r,1}}{n_r - 1},\qquad \text{ITL}_{r,i} = \tau_{r,i} - \tau_{r,i-1}.

Percentiles interpolate linearly between order statistics: for sorted x(0)≤⋯≤x(n−1)x_{(0)} \le \dots \le x_{(n-1)} and k=(n−1) p/100k = (n-1)\,p/100, qp=x(⌊k⌋)+(x(⌈k⌉)−x(⌊k⌋))(k−⌊k⌋)q_p = x_{(\lfloor k\rfloor)} + (x_{(\lceil k\rceil)} - x_{(\lfloor k\rfloor)})(k - \lfloor k\rfloor) (numpy's default, and the simulator's).

Goodput =∣{r:TTFTr≤T1, TPOTr≤T2}∣ / (alast−afirst)= |\{r : \text{TTFT}_r \le T_1,\ \text{TPOT}_r \le T_2\}| \,/\, (a_{\text{last}} - a_{\text{first}}).

Little's law: over a window [0,H][0, H] that starts and ends empty, L=1H∫0HN(t) dt=1H∑rWr=λWL = \frac{1}{H}\int_0^H N(t)\,dt = \frac{1}{H}\sum_r W_r = \lambda W with λ=n/H\lambda = n/H and WW the mean time in system. The tail: P(any of N slow)=1−0.99NP(\text{any of } N \text{ slow}) = 1 - 0.99^{N}.

Code

// src/lib/inference/metrics.ts (excerpt)
export function percentileSorted(s: ArrayLike<number>, p: number): number {
  if (s.length === 0) return NaN;
  const k = ((s.length - 1) * p) / 100;
  const lo = Math.floor(k);
  const hi = Math.ceil(k);
  return s[lo]! + (s[hi]! - s[lo]!) * (k - lo);
}

The same function as _percentile_sorted in Disaggregated_Inference_Sim's metrics.py, so numbers from this site and from the simulator are computed the same way.