Trade-offs: when not to disaggregate
Small models, low load and slow links, and where to go next.
Concept
Disaggregation removes one problem, prefill stalling decode, and adds three: a KV hand-off on every request, a pool ratio that has to match the workload, and more machinery (two schedulers, a router, a transfer engine). Whether that trade pays depends on the model, the load and the link. The interactive sweeps the offered load from 2 to 14 req/s and runs both clusters at every point.
Small models. Llama-3-8B fits on one H100, and its prefills are short (59.0 ms for 2,048 tokens), so a colocated decode step is only ever delayed by a short prefill. On two colocated H100s it meets both SLOs for 100.0% of requests at every load from 2 to 14 req/s; so does 1P1D. What disaggregation changes is the balance between the two latencies: at 8 req/s TPOT p99 improves from 12.2 ms to 8.8 ms, while TTFT p99 worsens from 287.6 ms to 353.6 ms, because one prefill instance now serves every prompt (results.md §11). If the colocated cluster already meets its SLOs, disaggregation buys margin on one SLO by spending it on the other.
High load and the wrong ratio. One Llama-3-70B prefill instance can sustain about 7.49 req/s of these prompts. Below that, 1P1D wins clearly; above it, prompts queue without bound: at 10 req/s 1P1D meets 0.0% of SLOs, worse than colocated (3.8%), whose two instances both prefill. The fix is the ratio, not the idea: 2P1D meets 100.0% at every load from 2 to 10 req/s (results.md §4). A fixed ratio is only right for one workload mix; production systems resize pools or route adaptively.
Slow links. Over 25 GbE, Llama-3-70B's 671.1 MB hand-offs fill the link at about 4.7 req/s. Disaggregation then meets fewer SLOs than colocated serving at every load from 2 to 12 req/s (at 14 both fail): at 5 req/s it meets 49.4% of SLOs against colocated's 89.7%, with the link 95.1% busy. Compression (chapter 12) or a faster link is a prerequisite, not an optimisation.
Low load. With few requests in flight there are few prefills to interfere with, and two half-idle pools spend most of their energy on static power whichever way they are arranged: at 0.5 req/s static power is 54% of the disaggregated cluster's energy (results.md §6). In this model the energy comparison at low load still favours disaggregation (Llama-3-70B at 2 req/s: 4.190 J per output token disaggregated, 5.100 colocated), so the argument against it there is cost and complexity, not joules.
Alternatives. Chunked prefill (chapter 4) bounds decode stalls on colocated GPUs without moving any KV cache, and is the usual first step. Disaggregation is now offered by several serving engines; their designs and maturity differ and change quickly, so read their documentation rather than a summary: vLLM's disaggregated prefill, SGLang's PD disaggregation, TensorRT-LLM's disaggregated serving and NVIDIA Dynamo.
Concept
A checklist. Disaggregate when:
- decode-tail SLOs (TPOT, ITL) are the binding constraint, and chunked prefill has not fixed them;
- prompts are long relative to outputs, so prefills are long stalls;
- the KV link has headroom: hand-off bytes × request rate well below its bandwidth (chapter 12's ), or compression gets it there;
- the load is high and steady enough to keep both pools busy, and you can size them (or resize them) to the workload;
- the hardware split pays: a cheaper or more power-efficient decode part, or more prefill FLOPs where TTFT matters.
Simulate before you build: the right ratio, link and batch limits depend on the prompt and output length distributions and on the SLOs, which is the job DistServe gives its simulator, and the job of the one in this series.
Where to go next.
- Inference Trade-offs Explained widens this chapter's one trade-off to every serving lever: batching, paged KV, prefix caching, disaggregation, parallelism, quantisation, speculative decoding and hardware, each measured against goodput, latency, cost and energy on five workloads, with a Pareto explorer and a chapter per lever (for this one, disaggregation and the pool split).
- The LLM Inference Simulators series (eleven decks) builds the simulator used here, from the discrete-event kernel to power, hot-spots, validation and acceleration.
- The Fourier Optics for Inference series asks whether optical hardware could serve as a prefill pool or carry the hand-off, and mostly finds that it cannot.
- The Simulation Engineering Toolkit covers the engineering around simulators, including PPA trade-offs (area, cost and performance per dollar of pools).
- Interview_Simulation has questions and worked answers on these chapters' material.
Maths
For a disaggregated cluster with prefill and decode instances, the sustainable request rate is bounded, roughly, by each stage in turn:
prefill capacity, decode capacity (mean batch , output tokens per request, mean step ) and link capacity. For the 70B defaults the simulator's analytic bounds are prefill 7.49 req/s, decode 28.8 and KV link 74.5: prefill is the bottleneck, so a second prefill instance, not a second decode instance, is what raises the knee.
Code
// src/components/interactive/TradeoffWidget.tsx (excerpt)
const { colocated, disagg } = configs({ ...DEFAULT_CONTROLS, model, link });
for (const r of WORKLOAD_RATES) {
out[`c${r}`] = colocated; // two colocated instances
out[`d${r}`] = disagg; // one prefill + one decode instance
}
Each of the eighteen points is a full 800-request run of the engine on the recorded workload for that load.