01 The Big Picture
Two tokens cost the same on your invoice but exercise the hardware in opposite ways. Prefill is a fat parallel matmul burst; decode is a thin sequential crawl that streams every weight from HBM once per token. Every serving strategy — chunked prefill, disaggregation, speculative decoding, continuous batching — is a scheduling answer to that mismatch.
Doc 06 established the anatomy: prefill processes your prompt tokens in one parallel pass (compute-bound), then decode emits one token at a time, each pass reading the entire model + KV cache (memory-bandwidth-bound, doc 08). Doc 14 covered which token to pick (temperature, top-p). If that doc is the probability layer, this doc is the dispatch layer: the geometry of how passes are grouped, split, interleaved, and overlapped on silicon. Sampling distributions are out of scope here — the loop that produces them is not.
02 What — Dispatch Is a Scheduling Problem
The two phases have opposite personalities, and batching amplifies the opposition:
| Property | Prefill | Decode |
|---|---|---|
| Work shape | Many tokens, one pass — massive parallelism | One token per pass per sequence — serial |
| Bottleneck | FLOPs (compute-bound) | HBM bandwidth (memory-bound) |
| Prefers | Parallelism across chunks / sequences | Large batch (amortize weight reads) |
| User-facing SLO | TTFT (time to first token) | ITL (inter-token latency) / TPOT |
Because the personalities conflict, what you batch together matters as much as how much. A decode token and a prefill chunk have incompatible resource footprints; naive co-batching makes both worse. The strategies below are all ways of arranging passes so each phase runs in its favored regime.
03 Why — TTFT vs ITL, Competing Budgets
Serving SLOs pull in opposite directions:
04 How — The Speculative Draft/Verify Loop
Decode memory-boundness has a loophole: one target pass over L tokens costs the same bandwidth as a pass over 1 token — the weights dominate. So guess several tokens at once (draft), verify them all in a single parallel target pass, and keep the verified prefix. Watch one round:
The loop then repeats from the extended sequence. The guarantee, proven in §07: the tokens that come out are distributed exactly as if the target model had generated them alone — regardless of how bad the draft is. Speedup, not correctness, is what α buys.
05 Prefill Strategies
5.1 Chunked prefill (Sarathi-style)
Instead of one giant prefill burst, split the prompt into Δ-token chunks and interleave chunks with the ongoing decode batches of other requests. Chunking time-slices between the compute-hungry newcomer and the bandwidth-cadence of existing streams: TTFT for the new request rises slightly (chunks serialize), but ITL variance for everyone else collapses from "spiky cliff every prefill" to "small hill every step." The Δ tradeoff is exactly the algebra in §03. Ragged/varlen kernels (doc 22 systems note; FlashAttention variable-length) make it real: within one kernel launch, each row attends only to its own valid tokens, so a batch can hold [chunk A, decode row B, decode row C] with no padding waste.
5.2 Prefix / prompt caching
The cheapest prefill is the one you skip: reuse KV for the longest stable prefix. That was the entire subject of doc 07 — hash-keyed blocks, prefix-first prompt anatomy, the ~10× cheaper rate. In strategy terms: it converts TTFT from compute to memory restore.
5.3 Prompt-adaptive batching & multi-LoRA
Schedulers group requests by prompt shape (long-prefill streams together, decode streams together) so batch composition matches the hardware regime. With adapters, multiple LoRA users sharing the same base prefix share one prefill pass and one KV prefix — batching by shared prefix turns N prefills into 1 + N tiny decodes.
5.4 Disaggregated prefill/decode (Splitwise · DistServe · Mooncake)
Stop co-batching entirely: dedicate a prefill pool (GPU-starved for FLOPs, cheap on memory) and a decode pool (memory-rich, KV-heavy), shuffling KV between them over high-bandwidth interconnect (Mooncake's KV-transfer layer). Why it works: prefill wants parallel FLOPs throughput, decode wants maximal HBM per sequence — different hardware gratings. Why co-location fails: batch composition theory — every co-batched prefill chunk raises decode ITL (§03) while riding "spare" FLOPs decode doesn't even use. Disaggregation buys each pool its natural regime at the cost of KV-transfer plumbing and loss of fine-grained load mixing.
06 Decode Strategies
6.1 Baseline autoregressive (greedy / sampled)
One forward pass, one token, update KV, repeat. For which token (temperature, top-p, penalties) see doc 14 — this doc only cares that each pass is memory-bound and mostly idle compute.
6.2 Beam search
Keep b candidate hypotheses, expand each step, prune to top-b. Cost multiplies the decode batch by b: O(b·s) passes for a b-token answer. For open-ended chat this is dead — likelihood-optimal text is bland, degenerate, and repetitive, and nobody wants b slightly-different essays. It survives where the output space is structured and scoring is honest: machine translation, constrained parsing, program synthesis with test-time reordering. Treat it as a search-layer tool, not a decoding default.
6.3 Speculative decoding — the headline
The draft/verify loop from §04. The key insight: a parallel target pass over γ tokens costs the same bandwidth as a pass over one, so used-for-verification FLOPs are nearly free. Variants of "where the draft comes from":
Self-speculative variants: use the model itself at lower cost — early-exit/layer-skip drafts (run only the first k layers, cheap "shadow" logits), or skip modules (e.g. MoE experts) during drafting. The model's own shallowness is the draft model.
Beyond serial speculation: lookahead/parallel decoding expands several hypotheses per step off reuse patterns in the prompt; Jacobi/fixed-point decoding reformulates generation as solving x = f(x) iteratively, "peeling" tokens from wrong guesses; diffusion-based decoding drafts a block in parallel and denoises it — full landscape in doc 25.
6.4 Continuous batching — the serving layer beneath
Not a per-user strategy but the substrate: Orca-style iteration-level scheduling lets sequences join/leave the batch at every forward step, so a finished sequence frees its slot mid-step instead of blocking the batch until its slowest member finishes. Connection: decode prefers large batches because one weight-read serves all rows (bandwidth amortization — AI rises toward the compute roof); prefill prefers parallelism across chunks precisely because each chunk is already roofline-bound. Chunked prefill (5.1) is what lets both share a scheduler without ITL carnage.
07 The Speculative Math
| α (draft agreement) | E[accepted], γ=4 | η = E/γ | S with c=0.1 |
|---|---|---|---|
| 0.3 | 1.43 | 0.36 | 1.02× |
| 0.5 | 1.94 | 0.48 | 1.38× |
| 0.7 | 2.77 | 0.69 | 1.98× |
| 0.9 | 4.10 | 1.02 | 2.93× |
Worked example: α = 0.7, γ = 4 → E = (1 − 0.7⁵)/0.3 ≈ 2.77 accepted tokens per round; only ~28% of drafts reach position 5, so η ≈ 0.69 — with a draft at 10% of target cost (c = 0.1) that's S ≈ 2.77/1.4 ≈ ~2× real speedup. Note the curve shape: E grows toward a ceiling, so η = E/γ falls as γ grows for weak drafts — the sweet spot pairs γ with α (α=0.9 tolerates γ=6–8; α=0.3 wants γ=2–3). Below α ≈ 0.3 the verify pass pays for nothing — speculation becomes overhead. This is why aligned/distilled drafters matter more than bigger ones.
08 Mental Models
The draft model is the CPU's branch predictor: guess a whole pipeline of branches (γ tokens), execute the verification "speculatively" in one pass, roll back on mispredict (rejection) — the architecture guarantees you never execute the wrong path to completion. Prediction accuracy α is branch-predictor hit rate; γ is pipeline depth. Lets you reason about: why wrong drafts cost only time, not correctness.
Disaggregation = Splitwise's insight: a prep station (prefill — chopping, FLOPs) and a line (decode — plating, cadence) share a kitchen badly when mixed; separate them and pass plates (KV) across. Chunked prefill is slicing the prep work so the line never idles. Lets you reason about why dedicated pools beat one general pool.
09 Common Misconceptions
"Speculative decoding changes the output distribution." No — it is mathematically exact (rejection sampling in §07). Bad drafts cost speed, never fidelity. It is not "mostly-right tokens slip through."
"A bigger draft model is always better." The objective is E/(γ·c+1): a huge draft raises accuracy slightly but raises c a lot. A draft at 30% of target cost can easily be slower than autoregressive. Alignment (token-level agreement) beats size.
"Beam search gives better answers than sampling." For open-ended generation, top-likelihood text is measurably degenerate — beams find the boring optimum. It's a structured-task tool (translation, constrained decoding), not a quality dial (doc 14).
"Chunked prefill slows everything down." It trades a small TTFT increase for the elimination of ITL cliffs for every other stream — for streaming workloads it's strictly better per-agent.
"Continuous batching is a decoding strategy I can pick." It's the scheduler floor beneath all strategies — modern engines have it by default. Your strategy choices are chunking, disaggregation, and speculation on top.