01 The Big Picture
Docs 01–02 covered pre-training: trillions of tokens, one objective, one model. But a raw base model only completes text — ask it a question and it may answer with three more questions. The gap between "text completer" and "assistant you can ship" is closed entirely in post-training.
Post-training is not one technique. It is a pipeline of three stages, each shrinking the space of behaviors: supervised fine-tuning (SFT) teaches form from demonstrations; preference tuning (RLHF/RLAIF/DPO) teaches judgment from comparisons; parameter-efficient adaptation (LoRA/QLoRA) makes the whole thing cheap enough to own. For an engineer with a systems background, this doc is the mapping from those ML terms onto things you already know: gradient descent on a loss, a constrained-optimization problem with a leash term, and a low-rank factorization that trades a d² matrix for two d·r rectangles.
02 The Three Post-Training Stages
03 Why Post-Training Exists
Why can't pre-training learn all of this? Two reasons, one economic and one structural:
Not every behavior fits in weights cheaply
Pre-training sees each behavior pattern in maybe 0.001% of its tokens. To make "always cite uncertainty" dominate, you'd need to upweight it by orders of magnitude across the whole corpus — retraining everything. Post-training concentrates gradient signal: 100K tokens of only the behavior you want moves the policy a hundred times more per FLOP.
Distribution collapse & safety
A pure MLE policy is a mimic: it reproduces the training distribution, including its toxic tail, and it drifts off-distribution the moment you deploy it. Preference tuning doesn't add new knowledge — it sculpts the output distribution, pushing probability mass from plausible-but-bad completions toward plausible-but-good ones. Safety behavior is suppression, and suppression needs a ranking signal, not more text.
The systems framing: pre-training buys you capability (predicting text well requires world knowledge). Post-training buys you control (choosing which of the capabilities you already paid for actually surface). Control is cheap — which is exactly why LoRA, in section 05, can deliver it with 0.4% of the parameters.
04 How It Works — The RLHF Pipeline
Step through the full pipeline, from base model to merged adapter.
Stage 2 in equations
Preference is learned, not given. A reward model rm(x,y) — usually the SFT model with its LM head swapped for a scalar head — is trained on human pairwise choices with a Bradley–Terry logistic loss:
Then the policy is optimized against it — with a leash. Pure reward maximization degenerates: the policy finds adversarial text the reward model scores absurdly high (reward hacking) and drifts far from fluent language. The standard fix (PPO/InstructGPT) is to make the objective the reward penalized by KL divergence from the reference policy πref — the frozen SFT snapshot:
In practice the reward-per-token is clipped for stability — the PPO-style surrogate:
DPO: the shortcut that deletes the reward model. The KL-constrained optimum has a closed form: π* ∝ πref·exp(r/β), which rearranges to r = β·log(π/πref) + β·log Z. Substitute that into the Bradley–Terry loss and the intractable partition function cancels in the pairwise difference — leaving a classification loss directly on the policy:
That is the whole trick: the reward model was only ever a proxy for preferences, and the pairs themselves contain enough signal to rank the policy's own implicit reward. No RL loop, no reward model to hack, one stable supervised loss. This is why most open-weight post-training recipes (Zephyr, Llama-family community tunes) use DPO or its variants (IPO, KTO, ORPO) rather than full PPO.
05 LoRA Math — Low-Rank Adaptation
Every stage above updates weights. LoRA changes what gets updated: not W, but a rank-r approximation of ΔW.
The observation: fine-tuning updates have low intrinsic rank. The ΔW you'd learn adapting a 4096-dim model to medical QA lives (approximately) in a low-dimensional subspace — task adaptation is a low-dimensional signal riding on high-dimensional pre-trained features. So parameterize it directly:
Parameter count
full: d²
LoRA: 2·d·r
d=4096, r=16: 131K vs 16.7M per matrix — 0.4% trainable. Optimizer state (Adam keeps 2 moments per param) shrinks by the same ratio; that, plus gradients, is where the GPU memory actually goes.
Why init B = 0
Initialize A with Gaussian noise, B with zeros. Then ΔW = B·A = 0 at step 0, so W′ = W exactly — training starts from the pristine pre-trained model with no random perturbation of the function. (Init A=0 too, and neither side would ever learn.)
Scaling law note
ΔW is scaled by α/r at runtime, so you can sweep r without retuning α. Common practice: r = 8–64, applied to attention q,v projections (sometimes all linear layers). Rank saturates fast — beyond ~64, extra rank mostly buys noise fit.
Where adapters merge — and why there's zero latency. During training, forward passes compute h = Wx + (BA)x, keeping B and A separate (small, trainable, cheap optimizer state). But after training, BA is just a d×d matrix. Compute it once and fold it in:
No extra FLOPs at inference, no added latency, no architectural difference. Deploy-time choice: either merge (single set of weights, zero overhead) or keep adapters separate and hot-swap them per request — many adapters sharing one frozen base is exactly how multi-tenant serving works (batch different customers' LoRAs in one pass).
QLoRA — one consumer GPU, a 65B model
QLoRA adds memory engineering on top: quantize the frozen base to 4-bit NF4 (NormalFloat — a quantization alphabet matched to the normal distribution of weights, plus double-quantization of the quantization constants), keep the LoRA adapters in bfloat16, and page optimizer state between CPU and GPU on memory spikes. Result: fine-tuning a 65B model in ~48 GB instead of ~780 GB. The math is unchanged — the gradients still flow through the dequantized weights into B and A; only the base's storage format changed.
06 The Engineering Decision Table
Three levers change model behavior. They differ in what they can express and what they cost — and they stack:
| Prompt engineering | RAG | Fine-tuning (LoRA) | |
|---|---|---|---|
| What changes | Instructions in context | Retrieved facts in context | Weights (adapters) |
| Buys you | Behavior control, per request | Fresh, private knowledge | Format, tone, domain skill |
| Doesn't buy | New skill or knowledge | New behavior patterns | Up-to-date facts (weights lag) |
| Cost | Engineering time only | Infra + retrieval quality | 1 GPU, hours–days, then near-zero |
| Update cadence | Instant | Index refresh | Retrain (minutes for LoRA) |
| Latency tax | More input tokens (doc 07) | More input tokens | Zero after merging |
Output style/structure is the product (strict schemas, house tone, tool-call discipline); a base or small model must adopt a niche skill; you're distilling a big model into a cheap one for a narrow task — then ship it to the edge (doc 20).
The problem is missing or stale facts (that's RAG — fine-tuning memorizes a snapshot, then confidently hallucinates on updates); a better system prompt would do (always try that first — doc 03); you lack a few hundred clean examples.
07 Mental Models
Pre-training is 20 years of reading everything: raw capability, no professionalism. SFT is onboarding — copying how the seniors write emails and run meetings. RLHF is the performance review: not "here's how it's done" but "this draft was better than that one." Lets you reason about: why 50K review comments (RLHF) can change behavior more than 50B more reading tokens (pre-training) — feedback compresses to judgment.
The base model is a shared runtime library, frozen and loaded once per GPU. Each LoRA adapter is a small plugin compiled against it — trained independently, hot-swappable, mergeable into one binary when stable. Lets you reason about: why you can serve dozens of customized models per GPU, and why "fine-tuning" no longer implies owning the weights.
08 Common Misconceptions
"Fine-tuning teaches the model new facts." Weakly, and unreliably. SFT is still next-token prediction — it teaches mappings and formats. Facts memorized via fine-tuning are brittle, order-sensitive, and confidently wrong after their expiry date. Fresh knowledge belongs in context (RAG), not weights.
"RLHF makes the model smarter." It makes the model better-aligned: same capabilities, different distribution over outputs. A base model often scores higher on raw benchmarks; the tuned model is the one that follows instructions instead of completing them.
"DPO replaces RLHF, so RL is obsolete." DPO removes the reward model for pairwise-preference objectives. Frontier labs still use RL variants (with verifiable rewards — unit tests, math checkers) where the signal is richer than "A better than B." DPO is a shortcut for one specific feedback type, not a general replacement.
"LoRA's small parameter count means small learning." The update ΔW = BA lives in the full d×d space — its effective rank is r, but its reach is the whole layer. For task adaptation, r = 16 routinely matches full fine-tuning; the bottleneck is data quality, not rank.
"QLoRA's 4-bit base degrades the fine-tune." The frozen base's forward pass runs through dequantized NF4 weights; gradients to the bf16 adapters are unaffected in kind. NF4 was designed so quantization error is information-preserving for normally-distributed weights — measured quality loss is typically within noise for adapter training.