AI Learning Series · Part 27

Standardization: MCP, A2A, Evaluations & Governance

The handshake layer of the AI stack. Who agrees on what — and why a standard is context engineering at civilization scale: one agreement that saves a million contexts from re-learning the same thing.

Agents
→
Harnesses
→
Standardization
→
Governance in Practice

01 The Big Picture

The frontier moved from capability to coordination. Models can read, code, and plan; what they can't do is agree with each other's plumbing. Standardization is where the software frontier now lives.

This series' thesis is that software is shifting from static (deterministic instructions compiled ahead of time) to static + dynamic (a deterministic harness assembling context for a probabilistic model at runtime). Docs 09 and 23 covered the dynamic half — skills, memory banks, harness layers. But a stack of one-off, hand-rolled harnesses is 1995 software before the web: brilliant, incompatible islands. Every tool vendor hand-wires every app, every agent speaks a private dialect, every evaluation is reported in its own format.

Standards answer that with a different kind of artifact — not code, not context, but an agreement about context: a fixed shape for tool schemas, a fixed handshake for sessions, a fixed vocabulary for risk. Notice what all three buy you, in this series' own currency: they convert unstable, per-vendor tokens into stable prefixes (doc 07's cheap class), and they convert pairwise confusion into a shared protocol. A standard is a context compression you perform once, on behalf of everyone.

📐
Thesis of this doc: standardization is context engineering at scale. MCP standardizes the shape of tool context; evaluation standards standardize the shape of evidence; governance standards standardize the shape of accountability. Same mechanism, three layers — and all three obey network-effects math (section 03).

02 What It Is (Precisely): Three Layers

"AI standardization" is not one thing. It stacks, and each layer answers a different question:

Layer 1Protocol — how machines talk. Wire-level agreements: MCP (Model Context Protocol) for model↔tool, A2A (Agent-to-Agent) for agent↔agent. JSON-RPC 2.0 messages, capability exchange, discovery. Question answered: "what does a valid conversation between these programs look like?"
Layer 2Interface — what the model sees. De-facto conventions in the context window itself: OpenAI-style function/tool JSON-schema calling (now imitated by nearly every provider), structured outputs via constrained decoding (doc 14), usage policies / model specs, system-prompt contracts, agent manifest files, and skills-format conventions (docs 09, 23). Question answered: "what shape of context does a model reliably honor?"
Layer 3Governance & metrics — what society sees. Risk frameworks (NIST AI RMF, EU AI Act), management systems (ISO/IEC 42001), documentation artifacts (model cards, datasheets), incident databases, evaluation reporting standards (doc 13, doc 25's calibration reporting). Question answered: "who is accountable, for what, and how do we check?"

The protocol layer: MCP vs A2A

MCP (Anthropic, late 2024, since adopted industry-wide) is a client–server protocol: a host application (IDE, chat app, your harness) runs an MCP client that connects to one or more MCP servers, each exposing capabilities to the model. It rides JSON-RPC 2.0 over stdio (local subprocess) or HTTP (remote), with a version handshake at session start. A2A (Google, 2025, since moved to a neutral foundation) targets the other edge: independent agents — possibly built on different vendors — discovering each other and delegating tasks, via a public agent card describing what an agent can do.

MCPA2A
ConnectsModel/host ↔ tools & dataAgent ↔ agent
MetaphorUSB port for capabilitiesNetwork of peers delegating work
Unit of exchangeTool calls, resources, promptsTasks with lifecycle & artifacts
DiscoveryClient enumerates a server's capabilities at connect timeAgent card fetched at a well-known URL
Trust boundaryHost mediates everything (permissions live in the harness)Negotiated between peers (auth, opaque execution)

They compose rather than compete: an orchestrator agent can use MCP to reach tools and A2A to hand a subtask to another agent's agent card. Doc 12 called MCP "USB for tools"; A2A is the office directory plus a work-order system.

03 Why: The Math of Agreement

Two formulas justify the entire standards effort. Both are about taming a product of two variable counts.

The M×N collapse

Without a protocol, every one of M harnesses that wants N tools writes N bespoke integrations: M·N adapters, each a maintenance liability that breaks on either side's change. With a shared protocol each side implements one endpoint: M + N. And the payoff of joining a network grows superlinearly — the Metcalfe-style count of possible useful connections:

no standard: adapters(M,N) = M·N — grows multiplicatively
with standard: adapters(M,N) = M + N — grows additively
network payoff: m(n) ∝ n(n−1)/2 — value of n compatible nodes

At M = N = 100 the ratio is 10,000 adapters versus 200. That 50× gap is the ecosystem: it's why one well-adopted MCP server instantly works in every compliant host, and why de-facto standards (QWERTY, USB, LSP, OpenAI's tool-call JSON) beat technically superior splinters. The winner is whichever agreement crosses the adoption threshold first — value m(n) compounds on the installed base, not on elegance.

Cache economics: schemas are tokens

Now the series' native accounting. A tool schema lives in the context window on every single request. A standardized schema is byte-stable across turns, across hosts, and across vendor switches — it sits in the stable prefix and lands in the cheap cache class (doc 07):

// standardized: one stable prefix, cached at ~0.1× cost ≈ S·c_cached + Δ·c_fresh // S = schema+system tokens, Δ = new tokens/turn
// splintered: every vendor switch rewrites the schema region cost ≈ Σᵢ (Sᵢ·c_fresh + Δᵢ·c_fresh) // fresh-input price for S on every switch

Concretely: 15K tokens of tool schemas re-prefilled at full price on each of 20 vendor/platform switches is 300K full-price tokens — versus ~30K full-price-equivalents if one canonical schema rides the cache the whole way. Multiply by every developer on earth re-learning the same 50 tools in 50 dialects, and the "wasted tokens" of fragmentation become a civilization-scale line item. Standards don't just save engineering hours; they save prefill compute — the exact resource doc 07 said providers pass savings on for.

💱
Restating in the series' currency: a standard is a prepaid prefix. The community pays once (spec design, review, conformance tests) and every future participant inherits a stable, cache-friendly, machine-checkable opening for their context. Fragmentation makes each participant pay full price, forever.

04 How It Runs — One MCP Handshake, Seven Steps

Step through a session: from process spawn to a permissioned, audited tool call.

HOST (IDE / app / harness) MCP client MCP SERVER tools · resources · prompts stdio (local) or HTTP (remote) 1 · spawn / connect (subprocess or URL) 2 · initialize: protocolVersion + capabilities (both ways) 3 · tools/list → typed JSON schemas returned system prompt + tool schemas (stable prefix) user turn 4 · schemas enter context, cached (doc 07) 5 · model emits tool call (JSON-RPC-shaped) { "name": "run_query", "arguments": { … } } 6 · server executes, returns result → observation appended 7 · permission gate + audit log (host-side, every step)

Two structural notes. First, the model never speaks MCP directly — it emits ordinary tool-call JSON; the host's client translates to JSON-RPC and enforces permissions. The security boundary lives in the harness, never in the model's good intentions (doc 16). Second, everything the protocol sends is text — tokens. The handshake exists so that the highest-value text (typed schemas) arrives in a predictable shape, in a predictable place in the context.

05 The Protocol Layer in Detail

MCP primitives (as of this writing — the spec evolves)

PrimitiveWhat a server exposesControlled by
ToolsExecutable functions with typed JSON schemas the model may invokeModel chooses; host approves
ResourcesRead-only context data (files, records) addressed by URIApplication decides what to load
PromptsReusable, parameterized prompt templates the user selectsUser chooses
SamplingA server may ask the client to run a model completion — inversion of control, keeping model choice and keys host-sideHuman approves via host
RootsClient tells the server which filesystem/URI scopes it may operate inClient constrains server
CapabilitiesEach side declares what it supports at initialize; unknown versions negotiate down or fail fastBoth

The capabilities exchange is the quiet masterstroke: instead of everyone upgrading in lockstep, a session opens with both sides announcing their version and features. Compatibility becomes a runtime negotiation, not a release-calendar coincidence.

A2A: the agent card

An agent publishes its self-description at a well-known URL; peers fetch it to decide whether — and how — to delegate:

// /.well-known/agent.json — the "business card" (illustrative) { "name": "code-review-agent", "skills": [ { "id": "pr-review", "description": "Reviews diffs; returns findings" } ], "capabilities": { "streaming": true, "pushNotifications": false }, "auth": { "type": "oauth2" }, "endpoint": "https://…" }

Delegation then follows a task lifecycle — submitted, working (with streaming status and artifacts), completed or failed. The counterpart to doc 12's subagent fan-out, but across organizational boundaries: the other agent's internals stay opaque; only the card and the task contract are shared.

Which standard, when?

SituationReach forWhy
Your app needs tools/data (DB, files, SaaS APIs)MCPOne server per provider; every compliant host discovers it automatically
Your agent must hand work to another org's agentA2APeer-level task contracts with auth; no visibility into the other side's internals needed
You need reliable structured output from a modelTool/JSON-schema + constrained decodingInterface-layer standard; grammar-constrained sampling (doc 14) makes validity mechanical
You're packaging reusable procedures for a harnessSkills-format conventionsDocs 09/23: skills are standard-shaped context — loadable, indexable, cache-friendly

06 The Governance Layer — Risk as a Function

Protocols standardize conversation; governance standards standardize judgment about consequences. The shared move: make risk a computable function, not a vibe.

Frameworks like the NIST AI RMF (Map → Measure → Manage → Govern) and the EU AI Act both start the same way: enumerate use-cases, then score each on likelihood × severity of harm:

risk(use-case) = P(harm) × severity(harm) — evaluated per deployment context
obligations = f( risk tier ) — tier is determined by the score above

The EU AI Act's tiers, sketched (as of this writing — regulators keep refining them):

Risk tierExamplesCore obligations
UnacceptableSocial scoring; manipulative techniques exploiting vulnerabilityBanned outright
HighHiring screening, credit, safety-critical usesRisk management system, data governance, logging, human oversight, conformity assessment
LimitedChatbots, synthetic contentTransparency: disclose "you're talking to AI"; label generated media
MinimalSpam filters, game AIVoluntary codes only
GPAI modelsGeneral-purpose foundation modelsTechnical documentation, copyright policy, training-data summaries; systemic-risk models add evaluations & incident reporting

Supporting standards turn these duties into artifacts: ISO/IEC 42001 defines an auditable AI management system (AIMS — the "ISO 27001 of AI"); model cards and datasheets standardize the self-description of models and datasets; the AI Incident Database aggregates real-world failures so likelihood estimates aren't invented from scratch. Doc 13's evaluation discipline is the measurement engine underneath all of it: a risk score is only as good as the evals that populate P(harm), and doc 25's calibration reporting is what keeps those numbers honest rather than vibes in a compliance costume.

🔁
Evals as monitors, not just gatekeepers: governance isn't a one-time certificate. The same eval suites that gate a release can run continuously in production as monitors — drift detection, guardrail trips, incident streams — feeding the observability stack (Perception: the AI observability stack). Measure (NIST's second verb) is a runtime activity now, not a pre-launch memo.

07 What Standardization Does NOT Solve

Model behavior drift. A protocol guarantees message shape, not message quality. A provider silently updates the model behind the same API and your tool-calling accuracy shifts — the standard is unchanged, the behavior isn't. (This is exactly the fragility doc 23's harness layers and the News doc's "weather vs climate" framing warn about: releases are weather; standards are climate — and climate moves slower.)

Eval gaming. Standardized eval reporting standardizes the format of evidence, not its honesty. Goodhart's law survives every schema: optimize for the published benchmark and the benchmark decouples from the skill (docs 13, 16). Governance frameworks assume measured numbers mean something — that assumption is the attack surface.

Adversarial use. Attackers adopt standards too: a uniform tool interface is a uniform target surface (prompt injection through a tool's own results, capability abuse across connected servers). Standardization widens the attack fan-in even as it narrows integration cost — the permission/audit layer of section 04 step 7 is not optional decoration.

The pace mismatch. The capability frontier re-prices itself monthly; a standard takes years to negotiate and longer to replace. Expect protocol features to arrive after the patterns they formalize are already common practice — that lag isn't failure, it's what durability costs.

08 Mental Models

USB ports for tools

MCP is the USB of the AI stack: device makers build to the port, host makers build the port, nobody renegotiates per plug. Lets you reason about: why a single MCP server instantly works across every compliant host, and why capabilities negotiation exists — USB asks devices "what can you do?" at plug-in, exactly like initialize.

USB devices are dumb peripherals; an MCP server can return arbitrary text that will flow into a model's context — the "device" can talk back and lie.
Shipping containers for capability

A2A and standardized eval reports are shipping containers: before containers, every dock re-handled cargo differently; after, any crane moves any box. The container spec says nothing about what's inside — only its shape. Lets you reason about: why standards enable delegation across orgs without trust in internals, and why a standardized eval report can be compared across teams even when the models differ.

A container can hold anything — including mislabeled goods. Format standards can't detect a gamed benchmark inside.

09 Misconceptions & Closing

"Standards cap innovation." They cap it at one layer to free it at every other. TCP didn't freeze networking; it made the transport boring so applications could go feral. Expect exactly this split: protocol/context shapes ossify, models and harness strategies (doc 23) keep churning.

"MCP replaces the agent harness." No. MCP standardizes one layer of doc 23's stack — how tools announce themselves and how calls travel. Context assembly, memory policy, permission gates, worker scheduling, the control loop: all still yours. A standard pipe doesn't decide what flows through it.

"A2A is just MCP with different names." Different trust topology. MCP assumes one host that sees and mediates everything; A2A assumes peers that don't see each other's internals and need negotiated, auth-scoped task contracts. Confusing them gives you either a surveillance bottleneck or an unaccountable swarm.

"Compliance = safety." A filled risk matrix is a documentation artifact; the harm channel is the deployed system. ISO-style process reduces the probability of negligence, not of failure — pair it with real evals (doc 13) and runtime monitors.

🧠
Closing insight — the paradigm shift, named: the static era standardized instruction formats (compilers, ABIs, HTTP). The static + dynamic era must standardize context — the shape of what a probabilistic machine reads and the shape of the evidence that it behaves. MCP, A2A, tool schemas, model cards, and eval-reporting standards are all the same object at different layers: a stable prefix paid for once and cached by everyone. That is the largest context optimization this series has described — measured not in tokens per request, but in integrations per ecosystem.