AI Learning Series · Part 17

Post-Training: Fine-Tuning, RLHF & LoRA

Pre-training builds a next-token predictor; it does not build an assistant. Everything that happens after — the alignment, the personality, the safety — is a second training campaign run on the same weights.

Transformers
→
Foundations
→
Post-Training
→
Test-Time Compute

01 The Big Picture

Docs 01–02 covered pre-training: trillions of tokens, one objective, one model. But a raw base model only completes text — ask it a question and it may answer with three more questions. The gap between "text completer" and "assistant you can ship" is closed entirely in post-training.

Post-training is not one technique. It is a pipeline of three stages, each shrinking the space of behaviors: supervised fine-tuning (SFT) teaches form from demonstrations; preference tuning (RLHF/RLAIF/DPO) teaches judgment from comparisons; parameter-efficient adaptation (LoRA/QLoRA) makes the whole thing cheap enough to own. For an engineer with a systems background, this doc is the mapping from those ML terms onto things you already know: gradient descent on a loss, a constrained-optimization problem with a leash term, and a low-rank factorization that trades a d² matrix for two d·r rectangles.

02 The Three Post-Training Stages

SFT — Supervised Fine-Tuning. Take ~10–100K high-quality prompt→response pairs (written or curated by humans), and run ordinary maximum likelihood on the responses: maxθ 𝔼(x,y)∼D[log πθ(y|x)]. Identical machinery to pre-training — cross-entropy on next tokens — just on a tiny, curated dataset. This clones the *format*: answer in prose, use markdown, refuse politely, call tools with valid JSON.
Preference tuning — RLHF / RLAIF / DPO. SFT clones what demonstrations look like, but can't rank two responses against each other. Collect (prompt, better response, worse response) triples and optimize the policy against them — with a learned reward model (RLHF/RLAIF) or directly from the preference pairs (DPO). This is where helpfulness-vs-harmfulness tradeoffs and tone live.
PEFT — Parameter-Efficient Fine-Tuning (LoRA / QLoRA). Freeze the base weights entirely and train a low-rank update matrix instead. This turns fine-tuning from a multi-GPU cluster job into something one engineer runs on one GPU over a weekend — and it's what makes your own domain adaptation economically viable.
🧭
Naming, decoded: RLHF = reinforcement learning from human feedback (humans do the ranking). RLAIF = the same, but a strong model ("AI feedback") does the ranking — 100× cheaper labels, at the cost of the judge's own biases. DPO = direct preference optimization, a closed-form shortcut that removes the RL loop entirely (see 04).

03 Why Post-Training Exists

Why can't pre-training learn all of this? Two reasons, one economic and one structural:

Not every behavior fits in weights cheaply

Pre-training sees each behavior pattern in maybe 0.001% of its tokens. To make "always cite uncertainty" dominate, you'd need to upweight it by orders of magnitude across the whole corpus — retraining everything. Post-training concentrates gradient signal: 100K tokens of only the behavior you want moves the policy a hundred times more per FLOP.

Distribution collapse & safety

A pure MLE policy is a mimic: it reproduces the training distribution, including its toxic tail, and it drifts off-distribution the moment you deploy it. Preference tuning doesn't add new knowledge — it sculpts the output distribution, pushing probability mass from plausible-but-bad completions toward plausible-but-good ones. Safety behavior is suppression, and suppression needs a ranking signal, not more text.

The systems framing: pre-training buys you capability (predicting text well requires world knowledge). Post-training buys you control (choosing which of the capabilities you already paid for actually surface). Control is cheap — which is exactly why LoRA, in section 05, can deliver it with 0.4% of the parameters.

04 How It Works — The RLHF Pipeline

Step through the full pipeline, from base model to merged adapter.

BASE MODEL next-token predictor · pre-trained "Q: capital of France?" → "A: what a fine question." (completes, doesn't answer) SFT max 𝔼[log π(y|x)] · demos learns FORMAT from ~50K demonstrations REWARD MODEL −log σ(r(y_w) − r(y_l)) learns JUDGMENT from human pairwise rankings RL AGAINST THE REWARD ∇J = 𝔼[∇log π(a|y) · A] (PPO clip + KL leash) r = r_m(x,y) − β·KL[π‖π_ref] π_ref = frozen SFT snapshot — the leash anchor DPO SHORTCUT −log σ(β·Δ log[π/π_ref]) on pairs — no RM, no RL loop the reward is defined implicitly by the preference data LoRA ADAPTER: W′ = W + BA all of the above trains only B and A (r ≪ d) — frozen base weights, merged at deploy time

Stage 2 in equations

Preference is learned, not given. A reward model rm(x,y) — usually the SFT model with its LM head swapped for a scalar head — is trained on human pairwise choices with a Bradley–Terry logistic loss:

ℒ_RM = −log σ( r(x, y_w) − r(x, y_l) )

Then the policy is optimized against it — with a leash. Pure reward maximization degenerates: the policy finds adversarial text the reward model scores absurdly high (reward hacking) and drifts far from fluent language. The standard fix (PPO/InstructGPT) is to make the objective the reward penalized by KL divergence from the reference policy πref — the frozen SFT snapshot:

r(x,y) = r_m(x,y) − β · KL[ π_θ(·|x) ‖ π_ref(·|x) ]
∇J = 𝔼[ ∇log π_θ(a|y) · A ] where A = advantage (reward − baseline)

In practice the reward-per-token is clipped for stability — the PPO-style surrogate:

ℒ = 𝔼[ min( ρ·A, clip(ρ, 1−ε, 1+ε)·A ) ] with ρ = π_θ(a|y) / π_old(a|y)

DPO: the shortcut that deletes the reward model. The KL-constrained optimum has a closed form: π* ∝ πref·exp(r/β), which rearranges to r = β·log(π/πref) + β·log Z. Substitute that into the Bradley–Terry loss and the intractable partition function cancels in the pairwise difference — leaving a classification loss directly on the policy:

ℒ_DPO = −log σ( β·( log[π_θ(y_w|x)/π_ref(y_w|x)] − log[π_θ(y_l|x)/π_ref(y_l|x)] ) )

That is the whole trick: the reward model was only ever a proxy for preferences, and the pairs themselves contain enough signal to rank the policy's own implicit reward. No RL loop, no reward model to hack, one stable supervised loss. This is why most open-weight post-training recipes (Zephyr, Llama-family community tunes) use DPO or its variants (IPO, KTO, ORPO) rather than full PPO.

05 LoRA Math — Low-Rank Adaptation

Every stage above updates weights. LoRA changes what gets updated: not W, but a rank-r approximation of ΔW.

The observation: fine-tuning updates have low intrinsic rank. The ΔW you'd learn adapting a 4096-dim model to medical QA lives (approximately) in a low-dimensional subspace — task adaptation is a low-dimensional signal riding on high-dimensional pre-trained features. So parameterize it directly:

W′ = W + ΔW, ΔW = B·A with A ∈ ℝ^(r×d), B ∈ ℝ^(d×r), r ≪ d

Parameter count

full: d²
LoRA: 2·d·r

d=4096, r=16: 131K vs 16.7M per matrix — 0.4% trainable. Optimizer state (Adam keeps 2 moments per param) shrinks by the same ratio; that, plus gradients, is where the GPU memory actually goes.

Why init B = 0

Initialize A with Gaussian noise, B with zeros. Then ΔW = B·A = 0 at step 0, so W′ = W exactly — training starts from the pristine pre-trained model with no random perturbation of the function. (Init A=0 too, and neither side would ever learn.)

Scaling law note

ΔW is scaled by α/r at runtime, so you can sweep r without retuning α. Common practice: r = 8–64, applied to attention q,v projections (sometimes all linear layers). Rank saturates fast — beyond ~64, extra rank mostly buys noise fit.

Where adapters merge — and why there's zero latency. During training, forward passes compute h = Wx + (BA)x, keeping B and A separate (small, trainable, cheap optimizer state). But after training, BA is just a d×d matrix. Compute it once and fold it in:

W′ = W + B·A → h = W′x — one matrix multiply, identical to the original layer

No extra FLOPs at inference, no added latency, no architectural difference. Deploy-time choice: either merge (single set of weights, zero overhead) or keep adapters separate and hot-swap them per request — many adapters sharing one frozen base is exactly how multi-tenant serving works (batch different customers' LoRAs in one pass).

QLoRA — one consumer GPU, a 65B model

QLoRA adds memory engineering on top: quantize the frozen base to 4-bit NF4 (NormalFloat — a quantization alphabet matched to the normal distribution of weights, plus double-quantization of the quantization constants), keep the LoRA adapters in bfloat16, and page optimizer state between CPU and GPU on memory spikes. Result: fine-tuning a 65B model in ~48 GB instead of ~780 GB. The math is unchanged — the gradients still flow through the dequantized weights into B and A; only the base's storage format changed.

# the whole "training loop" mental model for a LoRA run for batch in sft_data: h = quantize_4bit(W_frozen(x) + B(A(x))) # base NF4, adapters bf16 loss = cross_entropy(h, batch.labels) loss.backward() # grads flow to B, A only optimizer.step() # 2·d·r params, not d² # deploy: W' = W + B@A → one plain Linear, zero added latency

06 The Engineering Decision Table

Three levers change model behavior. They differ in what they can express and what they cost — and they stack:

Prompt engineeringRAGFine-tuning (LoRA)
What changesInstructions in contextRetrieved facts in contextWeights (adapters)
Buys youBehavior control, per requestFresh, private knowledgeFormat, tone, domain skill
Doesn't buyNew skill or knowledgeNew behavior patternsUp-to-date facts (weights lag)
CostEngineering time onlyInfra + retrieval quality1 GPU, hours–days, then near-zero
Update cadenceInstantIndex refreshRetrain (minutes for LoRA)
Latency taxMore input tokens (doc 07)More input tokensZero after merging
✓ Fine-tune when

Output style/structure is the product (strict schemas, house tone, tool-call discipline); a base or small model must adopt a niche skill; you're distilling a big model into a cheap one for a narrow task — then ship it to the edge (doc 20).

✗ Don't fine-tune when

The problem is missing or stale facts (that's RAG — fine-tuning memorizes a snapshot, then confidently hallucinates on updates); a better system prompt would do (always try that first — doc 03); you lack a few hundred clean examples.

07 Mental Models

Grad school vs. onboarding

Pre-training is 20 years of reading everything: raw capability, no professionalism. SFT is onboarding — copying how the seniors write emails and run meetings. RLHF is the performance review: not "here's how it's done" but "this draft was better than that one." Lets you reason about: why 50K review comments (RLHF) can change behavior more than 50B more reading tokens (pre-training) — feedback compresses to judgment.

A review teaches taste, not knowledge — preference tuning can't add facts your model never read.
Low-rank = plugin DLL

The base model is a shared runtime library, frozen and loaded once per GPU. Each LoRA adapter is a small plugin compiled against it — trained independently, hot-swappable, mergeable into one binary when stable. Lets you reason about: why you can serve dozens of customized models per GPU, and why "fine-tuning" no longer implies owning the weights.

A plugin can't change the runtime: rank r ≪ d bounds how far behavior can move; deep capability shifts still need full fine-tunes.

08 Common Misconceptions

"Fine-tuning teaches the model new facts." Weakly, and unreliably. SFT is still next-token prediction — it teaches mappings and formats. Facts memorized via fine-tuning are brittle, order-sensitive, and confidently wrong after their expiry date. Fresh knowledge belongs in context (RAG), not weights.

"RLHF makes the model smarter." It makes the model better-aligned: same capabilities, different distribution over outputs. A base model often scores higher on raw benchmarks; the tuned model is the one that follows instructions instead of completing them.

"DPO replaces RLHF, so RL is obsolete." DPO removes the reward model for pairwise-preference objectives. Frontier labs still use RL variants (with verifiable rewards — unit tests, math checkers) where the signal is richer than "A better than B." DPO is a shortcut for one specific feedback type, not a general replacement.

"LoRA's small parameter count means small learning." The update ΔW = BA lives in the full d×d space — its effective rank is r, but its reach is the whole layer. For task adaptation, r = 16 routinely matches full fine-tuning; the bottleneck is data quality, not rank.

"QLoRA's 4-bit base degrades the fine-tune." The frozen base's forward pass runs through dequantized NF4 weights; gradients to the bf16 adapters are unaffected in kind. NF4 was designed so quantization error is information-preserving for normally-distributed weights — measured quality loss is typically within noise for adapter training.

🗺️
The dot connects like this: pre-training (docs 01–02) bought capability; post-training converts it into a steerable product; adapters (LoRA) make steering a per-team artifact instead of a per-lab project. The next lever is on the inference side: instead of moving more probability mass into good answers at training time, spend compute at query time — doc 18, test-time compute. And when you distill a tuned model down to run on a phone, that's doc 20, small models on device. The prompt side of this same control problem — behavior without touching weights at all — is doc 03, prompt & context engineering.