01 The Big Picture
Doc 10 sketched the optimization stack and name-dropped GQA and MLA in one paragraph each. Here is the full design space, with the arithmetic that separates them.
Everything in inference loops back to one tensor: the KV cache. It is the only model state that grows with your conversation, the only state that must be read every decode step (doc 08's wall), and the reason a 32K-context conversation can cost more memory than the 70-billion weights underneath it. The weights are fixed; the cache is a design variable — and an entire subfield of architecture research is the question: what is the smallest tensor we can store per token such that attention can still reconstruct the keys and values it needs?
The answers form a spectrum: store everything (MHA), share copies (MQA, GQA), compress and regenerate (MLA), share across layers (YOCO), store less history (sliding windows), waste less of what you store (PagedAttention), or store fewer bits per element (KV quantization). Each point on the spectrum trades a different currency: memory, bandwidth, quality, or serving complexity.
02 What the Cache Is — a Tensor, Not a Dict
The KV cache is not a key-value store in the programmer's sense. It is a dense 4-D tensor with shape [L, n_kv, s, d_head] per K and per V — a fixed-size slab of VRAM addressed by (layer, head, position, dimension). "Key" and "value" mean attention keys and values, not hash-map keys and values.
Three consequences follow from it being a tensor:
n_kv, the head dimension d_head, and the layer count L are baked into the weights. You cannot convert a trained MHA model to GQA at serve time — the query heads have learned to expect their private keys. (MQA "uptraining" exists but is a retraining, not a switch.)[d_head, s] instead of [s, d_head] — so the score computation S = Q·Kᵀ is a clean GEMM over contiguous memory, and V stays [s, d_head] for the value-weighted sum. Contiguous runs along the s axis are what let FlashAttention-style kernels read whole histories with coalesced, high-bandwidth transactions. Every trick in this doc — paging, quantization, eviction — must preserve or explicitly pay for that contiguity.2 · n_h · d_head elements. Every design below is an answer to "can we store fewer than that?" — and at what cost to the attention output.03 Why — The Wall, Times Three
Doc 08 established that decode is memory-bandwidth-bound: the GPU reads weights and the KV cache from HBM for every single token. Cache design matters because three multipliers compound:
Context length × batch size × cache size ⇒ throughput. Shrink the per-token cache and every factor improves at once: more sequences fit in VRAM (bigger effective batch), each decode step reads fewer bytes (faster tokens), and the freed capacity can hold other users' caches — which is precisely what makes prompt-caching economics possible (doc 07). A provider whose cache is 8× smaller can keep 8× more prefixes resident, raising cache-hit rates and deepening the cached-token discount you see on the price card.
That is the "why": the KV cache is the scarce resource of serving, and the designs in this doc are its compression algorithms.
04 How — One Request, Three Designs
Watch the same 8-token request write its cache under MHA, GQA, and MLA, one decoder layer at a time.
The three rows are the spine of this doc; everything that follows either interpolates between them (MQA, GQA's knob), pushes further (cross-layer sharing), or changes how the bytes are stored rather than how many (paging, quantization, eviction).
05 The Taxonomy
MHA — the baseline
Multi-Head Attention. Every query head owns a private K and V. Full representational capacity — each head learns its own "address book" — at the highest memory cost: bytes/token = 2 · L · n_h · d_head · s · b.
✓ Pros
Maximum quality headroom; simplest kernels; no restore compute; every attention implementation supports it.
✗ Cons
Biggest cache — the 1.0× every other design is measured against. At long context, memory and decode bandwidth dominate everything.
Where used: original Transformer, GPT-2, BERT-era encoders, Llama-1 — and modern small models where the cache is small enough not to matter.
MQA — one shared address book
Multi-Query Attention (Shazeer, 2019): all query heads read from a single shared K/V head. The cache shrinks exactly ÷n_h — 64× for a 64-head model — because n_kv = 1.
✓ Pros
Maximum savings; also shrinks the K/V weight matrices and makes decode GEMMs fatter (better GPU utilization).
✗ Cons
Measurable quality drop — all queries share one view of history, costing diversity exactly where attention needs it.
Where used: PaLM, Falcon, StarCoder.
GQA — the interpolation knob
Grouped-Query Attention: partition the n_h query heads into groups; each group shares one K/V head. n_kv is a continuous dial between MQA (n_kv = 1) and MHA (n_kv = n_h), with the memory ratio simply n_kv / n_h. Llama-2/3 70B: 8 KV heads serving 64 query heads → 4 KB/token/layer instead of 32 KB, at near-MHA quality. This is the industry's default compromise.
MLA — compress, then regenerate
Multi-head Latent Attention (DeepSeek-V2/V3): instead of sharing whole heads, compress all K and V jointly into one low-rank latent vector per token. See the derivation in §06. ~28× smaller than MHA at comparable width, at near-lossless quality — the price is per-position restore compute and custom kernels.
Cross-layer / YOCO-style shared KV
"You Only Cache Once." Stack a few shared-KV attention layers whose cache is reused by many downstream layers: the layer-count factor in the sizing formula partially collapses. The extreme version caches once for the whole model. Trade: layers no longer build progressively richer keys, which constrains architecture design — but the savings multiply with everything else.
Sliding-window + global hybrid layers
Mistral, Gemma: most layers attend only to the last w tokens (4K–8K), a few layers attend globally. Local layers cap the cache at w per layer and make the attention matrix O(s·w) instead of O(s²); the global layers — far fewer — preserve long-range recall. Per-layer storage stops growing with s for the majority of the model.
PagedAttention — waste nothing (vLLM)
Pre-allocation forces reserving s_max per sequence; real sequences rarely fill it, and the fragmentation + reservation waste historically discarded 60–80% of KV memory. PagedAttention stores the cache in fixed-size blocks (classic block size: 16 tokens) with a page table — OS virtual memory for tensors:
It stores the same bytes — it's the design that makes the other designs' savings actually bankable, and its copy-on-write blocks are what make prefix sharing (doc 07) and parallel sampling cheap.
KV quantization, eviction, offload
FP8 / INT4 KV: halve or quarter b. The subtlety: V tolerates coarse quantization, but K does not — softmax divides by the temperature and is acutely sensitive to outliers in specific K channels, so K needs per-channel scales (one scale per head-dimension) while V gets away with per-token scales. Eviction (H₂O, SNAP): drop tokens with low accumulated attention mass, keeping "heavy hitters" + recency — the cache stops being a lossless record. Prefix caching + NVMe offload: move cold prefixes down GPU HBM → CPU DRAM → NVMe, keyed by token-prefix hash — doc 07's mechanism, enabled by small caches.
06 The Math
One formula to size them all
Worked table — one 70B-class sequence, s = 32K, fp16
Model shape: L = 80, d_head = 128, n_h = 64 (so d_model = 8192). Plug in:
| Design | n_kv / latent | Elems / token / layer | Bytes / token / layer | GB @ 32K ctx | vs MHA |
|---|---|---|---|---|---|
| MHA | n_kv = 64 | 2·64·128 = 16,384 | 32 KB | 2·80·32768·128·64·2 ≈ 85.9 GB | 1× |
| GQA (8 groups) | n_kv = 8 | 2·8·128 = 2,048 | 4 KB | ≈ 10.7 GB | ⅛× |
| MQA | n_kv = 1 | 2·1·128 = 256 | 0.5 KB | ≈ 1.34 GB | 1/64× |
| MLA | d_c + d_rope = 576 | 576 | ≈1.1 KB | 80·32768·576·2 ≈ 3.0 GB | ~1/28× |
Read the last two rows together: MLA's cache is bigger than MQA's per token, yet delivers near-MHA quality — it pays 576 elements to keep a compressed copy of all heads' information instead of 256 elements that literally share one head. That is the recurring theme: the knob isn't "how few bytes," it's "how much information survives the compression."
MLA derivation sketch
MHA keeps, per token t, n_h separate k-vectors and v-vectors — 2·n_h·d_h numbers, most of them redundant across heads. MLA's bet: the joint K/V content of a token lives in a much lower-dimensional subspace. So project it down once, at write time:
The elegance is in the absorption trick: attention scores need qᵀ·k_t = qᵀ·W_UK·c_t = (W_UKᵀ·q)ᵀ·c_t. Fold W_UK into the query projection at kernel time and you never materialize k_t at all — you score directly against the 512-d latent. What remains is one small matmul per position per layer on the restore path: compute and kernel complexity bought with a ~28× memory discount. The capacity check is one line:
07 Engineering Takeaways — Which to Pick When
| Situation | Pick | Why |
|---|---|---|
| Short context, quality-critical, small model | MHA | Cache is small enough that capacity buys real quality |
| General-purpose long-context serving | GQA, n_kv ≈ n_h/8 | The industry equilibrium: ~⅛ memory, ~no quality loss |
| Extreme batch / code models, quality slack exists | MQA | 1/n_kv = 1 savings; Falcon/StarCoder accepted the trade |
| Very long context at fleet scale | MLA | ~28× + quality — if you can pay the kernel-complexity tax |
| Cache doesn't fit VRAM at all | Paging + quantization + offload | Orthogonal to the above — multiplies with every row |
~0.1× cached-input price line is this entire doc, expressed as money.08 Mental Models
The KV cache is attention's main memory: every query must fetch from it before it can think. KV-cache design is then register-file compaction — the same game RISC architects play of keeping the hottest state in the fewest bytes, because every access is on the critical path. Lets you reason about: why cache bytes convert 1:1 into decode latency, and why compression here beats almost any other optimization.
MHA stores the full address book (all heads' K,V). MLA stores a compressed archive (the latent) plus instructions to regenerate each page on demand — paying a small unzip cost (W_UK·c) each time a page is read. Lets you reason about: the memory-vs-compute trade, and why restore cost shows up in kernel engineering, not in model quality.
64 query heads commuting to the same history. MHA gives each its own car (private K/V); MQA puts everyone in one bus (one shared head, crowded and slow-witted); GQA runs 8 carpools — nearly as cheap as the bus, nearly as comfortable as private cars. Lets you reason about: n_kv as a continuous dial, not a binary choice.
09 Common Misconceptions
"GQA hurts quality as much as MQA." The quality cliff is near MQA (n_kv = 1); at n_kv ≈ 8 the loss is small enough that it became the default for every major open model. The dial's ends are very different from its middle.
"MLA is just fancy GQA." No. GQA shares whole heads but keeps every dimension of the heads it keeps. MLA compresses across the dimension too, jointly for K and V, into one latent — a fundamentally different (low-rank, restorable) representation, with the restore matmul and decoupled-RoPE machinery GQA doesn't have.
"KV quantization is like weight quantization." Weights are static and forgiving. K vectors feed a softmax that is sharply sensitive to outliers in specific channels — hence per-channel scales for K, per-token for V, and quality degradation that appears only at long range, exactly where you wanted the savings.
"PagedAttention makes the cache smaller." It stores the identical bytes; it stops wasting 60–80% of the allocation on fragmentation and over-reservation. The savings are real but come from accounting, not compression — it composes with, rather than replaces, GQA/MLA/quantization.
"Sliding window means the model can't remember anything long-range." In hybrid designs (Mistral, Gemma), only most layers are local; dedicated global layers still attend across the whole context. You lose long-range per-local-layer, not long-range capability overall — though retrieval-heavy tasks do measurably degrade versus full attention.