AI Learning · End to End

The Shifting Paradigm of Code

From static code to static + dynamic. From CPU cycles to tokens as currency. What, how, and why — connecting software approaches down to hardware bottlenecks.

Knowledge graph — the whole series at a glance
Foundations 1–5 Systems 6–10 Practice 11–16 Perception Frontier 17–25 · News Security & Standards 26–27

Drag to rotate · scroll to zoom · hover a node for its title · click to open the doc

Theory track — concepts in order
00 · Prehistory

Pre-ML Era: Classical NLP

Rule-based systems and statistical NLP — what came before learning, and why it hit a wall.

01 · Foundations

Supervised Learning → Neural Nets

ML as function search, loss and gradient descent, neural network anatomy — with an animated forward pass.

02 · The Engine

Transformers → Inference

Attention, training, MoE, FlashAttention — plus an animated prefill/decode loop showing why the KV cache exists.

03 · The Math

Mathematics of AI

The linear algebra, calculus, and probability that make the previous two docs precise.

04 · The Craft

Prompt & Context Engineering

Tokens as currency: harnessing AI, saving tokens, fighting hallucination.

Practice track — building with AI
11 · Retrieval

Embeddings & RAG

Meaning as geometry: vector search, HNSW, the animated RAG pipeline, and where retrieval fails.

12 · Action

Agents End to End

The tool-use loop animated, MCP, and the design patterns that survive production.

13 · Measurement

Evaluation

Why benchmarks mislead, the animated eval-building loop, and the LLM-judge problem.

14 · The Dice

Sampling & Decoding

Temperature, top-p, grammar-constrained JSON, speculative decoding — the animated sampling funnel.

15 · Vision

Multimodal & Diffusion

How models see (VLMs) vs how they paint (diffusion) — animated denoising, and why image models can't spell.

16 · Trust

Safety & Alignment

The alignment stack, jailbreaks vs prompt injection (animated attack walkthrough), and defense in depth.

Frontier track — 2026 & beyond
17 · Adaptation

Post-Training: RLHF, DPO & LoRA

How weights get their behavior — SFT, preference optimization with the KL leash, DPO's shortcut, and the LoRA math.

18 · Thinking

Test-Time Compute & Reasoning

Thinking tokens billed as output, RLVR, best-of-n math, and when extra thinking isn't worth it.

19 · The Window

Long Context & Memory

1M-token windows: positional math, ring attention, RAG vs agentic memory, and what each approach costs.

20 · The Edge

Small Models & On-Device AI

The bandwidth wall on a phone: 4-bit math, NPUs, tokens-per-second from LPDDR, hybrid edge→cloud routing.

21 · The Frontier

SOTA LLMs & MoE Architectures

DeepSeek-V3, Mixtral, Qwen3, GLM — the gate math, expert capacity, and why total vs active params is the new spec.

22 · The Cache

KV Cache Types

MHA → MQA → GQA → MLA — the design space of attention memory, with the unified sizing formula.

23 · The Harness

Agent Harness Architectures

Orchestrator-workers, memory banks, skills-as-progressive-disclosure — each mapped to the bottleneck it relieves.

24 · The Dispatch

Prefill & Decoding Strategies

Chunked prefill, disaggregated serving, speculative decoding with the exact acceptance-rate speedup math.

25 · Challengers

New Architectures Landscape

Mamba & SSM hybrids, RWKV, BitNet b1.58, byte-level & diffusion LMs, JEPA world models, JEV decision models.

Live · News

Latest AI Enhancements

The moving target, dated: decision models, MoE everywhere, KV innovation wave, agent standards — linked into deep dives.

Security & Standards track — trust and handshakes
26 · Security

AI Security Engineering

Direct vs indirect prompt injection, threat models, the guardrail stack as AND-gates with correlation, agent hardening, red teaming as a regression loop.

27 · Standards

Standardization: MCP → Governance

MCP & A2A protocols, interface convergence, NIST AI RMF & EU AI Act — and why standards are context engineering at scale (N×M → N+M).

Perception track — the engineer's lens
Perspective

The Developer Perspective

How the paradigm of coding is shifting for working engineers.

Harness

The Developer Harness

Skills, rules, workflows, memory banks, subagents — software answers to the context problem.

Hardware

CPU vs GPU

Serial genius vs parallel army; memory bandwidth as the real bottleneck — with live data-traffic animation.

Observability

AI as an Observability Stack

Agents and models as runtime dependencies: telemetry ↔ quality ↔ semantics ↔ governance planes, SLI/SLO math, evals as production monitors.

Big picture

AI & Human Evolution

Knowledge transfer, genes to GPUs — the long arc.

Systems track — software meets hardware
06 · The Pipeline

Inference Anatomy

Prefill vs decode, KV cache math, batching — animated request lifecycle from Send to streamed tokens.

07 · The Economics

Context Caching & Cost

Tokens as currency: animated cache hit vs miss, cache-friendly prompt anatomy, the three token price classes.

08 · The Silicon

GPU Memory Hierarchy

The bandwidth wall, arithmetic intensity, and an animated naive-vs-FlashAttention walkthrough.

09 · Capstone

Software Context Solutions

Skills, rules, workflows, memory banks, subagents — each mapped to the hardware bottleneck it relieves.

10 · The Map

The Inference Optimization Stack

Model → Memory → Runtime → Cluster: MoE, SSMs, GQA, quantization, PagedAttention, kernel fusion, tensor parallelism, Splitwise — one animated mental model.