AI Learning Series · Latest News

Latest AI Enhancements

The moving frontier, as of this writing. This is the series' living front page: what changed recently, why it matters to engineers, and which deep-dive doc explains the math behind it.

Foundations
→
Frontier 21–25
→
Latest News
📅 Everything here is framed "as of this writing" — no dates baked in. Numbers are hedged as publicly reported unless reproducible. Treat each item as a signal, not a verdict.

01 The Big Picture — Five Axes to Watch

Model news is noisy. Axis-level news is durable. Almost every credible announcement of the recent period lands on one of five axes — and each axis maps to a hardware or software constraint this series has already derived.

Sparsity. Frontier recipes went dense → Mixture-of-Experts as the default: huge total parameter counts, small active counts per token. The FLOPs bill stopped being proportional to the parameter count.
New memory formats. GQA, multi-head latents, paged KV, FP8/INT4 KV — the KV cache became an engineered data structure, not a byproduct. Memory engineering is the frontier (doc 08's bandwidth math, weaponized).
Test-time compute. Reasoning models spend "thinking tokens" as billable output; quality became a dial you pay per token, not a fixed property of the checkpoint.
Decision readouts. A new class of model reportedly emits calibrated probabilities directly instead of decoding text — skipping autoregressive generation entirely. The newest and least-verified axis.
Harness standards. MCP and harness patterns standardized the software around models. The "static code → static + dynamic" shift of this series, arriving as an interoperability layer.
🧭
How to read this page: the scoreboard (02) gives one-line math per trend, the frontier map (03) is one animated picture, the deep-dive list (04) points into the series docs where each trend's math is derived. When a headline confuses you, find its axis first.

02 The Frontier Scoreboard

Each trend × its mathematical essence × where the series derives it.

TrendOne-line math essenceSeries doc
Test-time compute (o1/R1 style)Quality ∝ thinking tokens; tokens billed as output (3–5× input price)Doc 18 · Test-Time Compute
MoE everywhere671B total / 37B active ≈ 5.5% of params per tokenDoc 21 · SOTA LLMs & MoE
KV-cache innovation waveKV bytes = 2 · layers · kv-heads · head-dim · seq · bytes; shrink any factorDoc 08 · GPU Memory · Doc 22 · KV Cache Types
Alternative architecturesSSM state O(1) vs attention O(n²); 1.58-bit ternary weightsDoc 02 · Transformers · Doc 25 · New Architectures
Decision models (JEV)Probabilities read from hidden states; state reuse Q·S → S tokensItem 01 below
Agent standards & harnessingTool schemas are cached prefix tokens; harness = context managerDoc 12 · Agents · 09 · Context Solutions · Doc 27 · Standardization
Speculative & prefill schedulingE[accepted draft tokens] × cheap verify vs expensive decodeDoc 14 · Decoding · 10 · Inference Stack
On-device & small models4-bit weights: bytes/param ÷ 4; NPU SRAM beats HBM latencyDoc 10 · Inference Stack · Doc 20 · On-Device AI
Long-context battlesAttention O(n²) compute, O(n) KV memory — and fidelity decays mid-contextDoc 06 · Inference · 11 · RAG
Post-training everywhereRLHF/DPO commoditized; RLVR rewards = checkable answersDoc 17 · Post-Training

03 One Anim — The 2026 Frontier Map

Step through the axes in the order they arrived. Each stop is one bet about where the marginal dollar of compute buys the most capability.

Dense baseline: every token pays for every parameter Dense era FLOPs ∝ params Sparsity 671B total → 37B active MoE becomes default Memory formats GQA · MLA · paged · FP8 KV KV bytes engineered down Test-time compute thinking tokens = output $ quality becomes a dial Decision readouts p·C_miss > (1−p)·C_esc probabilities, not prose Harness standards MCP · skills · memory banks the software frontier Left wins = cheaper compute per unit of capability. Right wins = cheaper coordination per unit of capability.

04 The Deep-Dive List

One mini-card per trend: what it is, why engineers care, the math where it's natural, and where the series derives it properly.

01Decision Models — JEV (TypeSafe)

What it is (as of this writing): JEV is reportedly a decision model — a model that outputs typed decisions with calibrated probabilities directly, e.g. Yes: 80% / No: 20%, instead of generating text and hoping the prose parses. The only detailed public source is a Medium piece by Bijit Ghosh, "Inside JEV: Architecture of a Decision Model" — the architecture is not officially disclosed, so treat everything below as a report about an emerging class, not a verified spec.

The reported mechanism: the architecture was reconstructed black-box from ~10k API calls. A causal transformer encodes the shared context once and reuses its KV/state across Q question branches — for context length S, the per-question state work reportedly drops from ~Q·S to S tokens. Each branch ("route this ticket" / "estimate urgency" / "flag for review") attends to the shared context plus only its own options; attention masks isolate sibling questions so they can't leak into each other. Probabilities are read directly from hidden states — no autoregressive decode loop at all. Training reportedly targets calibration (RLCD — Reinforcement Learning for Calibrated Decisions, plausibly via a log-loss or Brier-style objective), and the reported MMLU calibration error is ≈ 0.031, concentrated in high-confidence predictions.

// Calibration: Expected Calibration Error, binned by confidence ECE = Σₘ (|Bₘ| / n) · | acc(Bₘ) − conf(Bₘ) | // Bₘ = predictions in confidence bin m; 0 = perfectly honest // Agentic routing: escalate only when expected cost of silence wins escalate ⟺ p · C_miss > (1 − p) · C_escape // escalation cost 1, missed-urgent cost 9: // 9p > 1 − p ⟺ p > 0.1 → escalate above 10% "urgent" probability

Why engineers care: this is expected-cost routing as a first-class primitive. An agent harness can branch on a calibrated probability instead of regex-parsing prose — and the routing inequality turns a model output directly into an escalation budget. It slots into the agent-loop economics of doc 12 and the context-reuse strategy of doc 09: one encoded context, many decision branches, is exactly the cache-friendly shape doc 07 preaches.

⚠️
Frankly flagged: new product, unpublished architecture, single public source. The durable takeaway is the class — "decision readout instead of decode" — not this specific vendor. Watch for independent calibration replications before betting routing logic on it.

02Test-Time Compute & Reasoning Models

What: o1/R1-style models emit thousands of hidden "thinking" tokens before the visible answer. Why it matters: thinking tokens are billed as output — the 3–5× price class from doc 07 — so accuracy is now a per-token purchase. RLVR (reinforcement learning with verifiable rewards) trains the thinking; budget forcing caps it. Math hook: E[answer quality] rises roughly with thinking-token budget, with diminishing returns — your job is finding the knee of that curve per task class. Deep dive: doc 18.

03MoE Everywhere

What: DeepSeek-V3 (publicly reported 671B total / 37B active parameters), Llama-4, Qwen3-MoE, GLM-4.5 — sparsity is now the default frontier recipe. Why it matters: capability scales with total parameters while per-token cost tracks active ones (≈5.5% here) — but only if routing keeps experts balanced, hence aux-loss-free balancing schemes, and only if memory fits, hence MLA. Math hook: cost per token ∝ active params, memory ∝ total params. Deep dive: dense-vs-sparse grounding in doc 02; the MoE doc is doc 21 · SOTA LLMs & MoE.

04The KV-Cache Innovation Wave

What: GQA became standard, MLA compresses K/V into low-rank latents, paged KV eliminates fragmentation, FP8/INT4 KV halves-or-quarters the bytes. Why it matters: the KV cache is the memory wall of doc 08 — KV bytes = 2 · layers · kv-heads · head-dim · seq · bytes-per-elem — and every term of that product is now an engineering target. Math hook: shrink any factor, shrink batchable concurrency proportionally. Deep dive: doc 08; the dedicated KV doc is doc 22 · KV Cache Types.

05Alternative Architectures

What: a wave of non-vanilla-attention designs, as of this writing mostly at the hybrid or niche stage:

ArchitectureKey mathWhat it trades
Mamba/SSM hybrids (Jamba, Zamba)Recurrent state, O(n) in sequence lengthGives up exact recall at distance; gains linear cost
RWKVAttention folded into RNN-style state updatesConstant-memory inference for weaker parallel scoring
Gated DeltaNetDelta-rule state updates with gatingCompression loss vs full attention fidelity
BitNet b1.58Ternary weights {−1, 0, 1}: log₂3 ≈ 1.58 bitsExtreme efficiency for lower per-weight precision
Byte-latent transformersPatches bytes adaptively, no tokenizerPatching complexity for tokenizer robustness
Diffusion LMs (LLaDA, Mercury)Parallel denoising, not left-to-right decodeSpeed for changed sampling/edition semantics
JEPA world modelsPredict in representation space, not token spaceAbstraction for direct pixel/token prediction
Titan / test-time memoryLearned memory updated at inference timeExtra state machinery for long-horizon recall

Why it matters: each row re-attacks the O(n²) attention or 16-bit weight assumptions derived in doc 02. None has displaced the transformer yet; all are worth tracking. Deep dive: doc 25 · New Architectures Landscape.

06Agent Standards & Harnessing

What: MCP (Model Context Protocol) as a tool-interoperability standard, plus harness patterns — orchestrator-workers, skills, memory banks. Why it matters: this is the "software" frontier of the series: the static + dynamic shift, where tool schemas become cached prefix tokens (doc 07) and the harness becomes a context manager deciding what earns a slot in the window. Math hook: every standardized tool schema is amortized cache input, not fresh input. Deep dives: doc 12, doc 09; the standards doc is doc 27 · Standardization.

07Speculative & Prefill Scheduling

What: draft-verify decoding (EAGLE, Medusa) uses a cheap drafter whose guesses a strong model verifies in parallel; chunked prefill interleaves prompt reading with decoding; disaggregated serving (Splitwise, DistServe, Mooncake) splits prefill and decode onto different machines. Why it matters: this is the serving-layer SLO handshake — trading verification compute for bandwidth-bound decode speed. Math hook: speedup ≈ E[tokens accepted per draft] × drafter-cheapness − verify overhead; chunked prefill trades TTFT against per-token latency (ITL). Deep dives: doc 14, doc 10; the serving doc is doc 24 · Prefill & Decoding Strategies.

08On-Device & Small Models

What: NPUs in consumer silicon, 4-bit quantization as a shipping default, WebGPU bringing inference to the browser. Why it matters: edge became a first-class substrate — private, offline, zero marginal API cost. Math hook: 4-bit weights cut the bytes-per-param term of the memory equation by 4×; NPU SRAM wins on latency where HBM bandwidth wins on throughput (the CPU-vs-GPU asymmetry of Perception · CPU vs GPU). Deep dive: doc 10 covers quantization; the edge doc is doc 20 · Small Models & On-Device AI.

09Long-Context Battles

What: a public arms race toward 1M–2M-token windows — versus the quieter math of attention fidelity. Why it matters: attention compute is O(n²) and KV memory is O(n) per token forever, and empirically retrieval fidelity sags mid-context ("lost in the middle"). Ring attention shards the n² across devices; RAG keeps n small and retrieves. Math hook: a 2M-token window doesn't repeal doc 06's per-decode-step KV read — long context permanently taxes output speed. Deep dives: doc 06, doc 11 · RAG; the long-context doc is doc 19 · Long Context & Memory.

10Post-Training Everywhere

What: RLHF/DPO/LoRA fine-tuning became commodity skills, while the differentiator moved to RLVR for reasoning and "verifier-as-data" — the checker defines the curriculum. Why it matters: when anyone can LoRA a checkpoint, the moat is the reward signal you can compute cheaply and honestly. Math hook: RLVR rewards are checkable predicates (unit tests, exact match), so advantage estimates carry less label noise than preference pairs. Deep dive: doc 17.

05 How to Track a Moving Target

You cannot subscribe to "the frontier." You can subscribe to signal types — each with a different reliability/cost profile:

📄 arXiv & tech reports

Highest information density, highest reading cost. Filter by what changes an equation you already know — a new KV layout, a new routing loss — not by headline model names. Cross-check claims against the appendix tables, not the abstract.

🔄 Provider changelogs

The ground truth for what you can actually ship: pricing shifts, cached-token discounts, context windows, tool-calling APIs. A pricing change is a hardware disclosure in disguise — read doc 07's economics off every new price list.

📊 Benchmarks

Weakest signal, most viral. A benchmark score is a sample from one eval distribution under unknown serving conditions. Trust trends across independent evals more than any single leaderboard jump — and check calibration (ECE-style), not just accuracy.

✓ Do

File each headline under one of the five axes; ask "which bottleneck did this attack?"; hedge every number with its source; re-check this page's claims against primary sources before building on them.

✗ Don't

Rewrite production routing on a single-source report (see item 01); compare benchmarks across different serving configurations; assume a demo latency is a served latency; date-stamp knowledge you can't refresh.

06 Mental Models

Weather report, not encyclopedia

A news page like this one is a forecast with a freshness half-life. The axes (sparsity, memory, test-time compute, readouts, harnesses) are climate — they change on year scales. Specific models, prices, and benchmark numbers are weather — they change weekly. Invest your reasoning in climate. Lets you reason about: which parts of this doc to memorize (none of the numbers, all of the axes).

Even axes drift: "decision readouts" barely existed as a class before it appeared — keep one eye on axis-zero, "new axis appears."
Scoreboard of bets

Every trend is a bet that one bottleneck (FLOPs, KV bandwidth, decode latency, coordination cost) is now the binding constraint. MoE bets on FLOPs; KV formats bet on bandwidth; test-time compute bets that tokens can buy accuracy; harness standards bet that coordination, not capability, is scarce. Lets you reason about: why two credible announcements can point in opposite directions — they're attacking different walls.

Bottlenecks rotate: when one axis wins, the frontier's constraint moves to the next — today's winning bet re-prices tomorrow's.

07 Common Misconceptions

"News = product demos." A demo is a sample under curated conditions. Benchmarks ≠ serving reality: latency, cost, calibration, and failure modes only show up under your traffic. The gap between leaderboard and production is exactly docs 07 + 08's economics and memory math.

"A bigger context window means the model remembers everything in it." Window size is an advertisement about capacity, not about fidelity. Attention quality degrades mid-context (doc 06), and every decode step still pays the full KV read (doc 08) — retrieval (doc 11) often beats stuffing.

"Decision models will replace LLMs." The reported JEV class replaces decoding for structured classification and routing — one readout instead of hundreds of generated tokens. Generation, synthesis, and code remain autoregressive territory. Complement, not replacement — and an unverified one at that.

"MoE means the model got smaller." Total parameters grew (671B); only the per-token compute shrank (37B active). You still pay the memory bill for all experts — which is precisely why the KV/memory-format axis (item 04) had to follow it.

"As of this writing" hedging is journalism, not engineering. It's the opposite: production systems that hard-code model behaviors break on the next changelog. Hedged claims + five-axis filing is how you keep an architecture alive across model swaps.

🔗
Where to go next: item 02's math is fully derived in doc 18; the memory math behind items 03–04 lives in doc 08; and the harness patterns of item 06 are built end-to-end in doc 12 and doc 09.