mpl vs MCP: An Audit and Contract Layer for Agent Protocols
MCP defines how agents talk. mpl checks what they said against a contract, scores it and records it. Why production agent systems may need both.
The protocol-stack question
MCP (Model Context Protocol) and A2A (Agent-to-Agent) are the two common protocols for agent-to-tool and agent-to-agent communication. They are well designed, widely adopted, and answer the question they set out to answer: how do agents exchange messages?
They do not set out to answer the question production teams end up asking: did that message meet the contract, how good was the exchange, and can you show what happened afterwards? That is the gap mpl is built for.
The short version: MCP is a transport. mpl is a contract, quality and audit layer. They stack. mpl runs as a sidecar between your agent and your MCP server, validating, scoring and recording what passes through.
What MCP is
MCP is a JSON-RPC-based protocol that lets an LLM application call tools exposed by an MCP server. The server advertises its tools, with names, descriptions and JSON Schema for their inputs (and optionally outputs); the client invokes them; results come back:
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "calendar.create",
"arguments": {
"title": "Quarterly Review",
"start": "2026-03-15T14:00:00Z",
"end": "2026-03-15T15:00:00Z"
}
},
"id": "req-1"
}
MCP moves messages and describes tool interfaces. Whether a particular message is semantically right, how good the agent’s output was, and how to prove later what was sent are left to the implementations on either side.
What mpl is
mpl is a protocol layer that sits between your agents and the underlying transport — MCP for client-server, A2A for peer-to-peer, or plain HTTP. At every hop it validates the payload against a contract, scores it, applies policies and writes an audit record:
┌─────────────────────────────┐
│ Your Agent Logic │
├─────────────────────────────┤
│ mpl (this layer) │
│ Contracts · Quality · │
│ Policies · Proofs │
├─────────────────────────────┤
│ MCP (client-server) or │
│ A2A (peer-to-peer) or HTTP │
└─────────────────────────────┘
mpl ships with:
- A contract registry — versioned semantic types (stypes)
such as
org.calendar.Event.v1, each backed by a Draft 2020-12 JSON Schema. Pre-built contracts ship under theorg.*,data.*,eval.*andai.*namespaces; you add your own. - Quality measurement — six Quality-of-Message metrics
(
schema_fidelity,instruction_compliance,groundedness,determinism,ontology_adherence,tool_outcome) composed into profiles such asqom-basic,qom-strict-argcheck, or your own. - A policy engine — rules that enforce organisational constraints, such as requiring a given quality profile for stypes matching a pattern.
- Audit records — every message carries a BLAKE3 hash of its
canonicalised payload (
sem_hash), a provenance block (who emitted it and why), a quality report and a timestamp. mpl produces the records; you choose the append-only store they go to.
The core is written in Rust (the mpl-protocol, mpl-proxy,
mplx and registry crates), with SDKs for Python and
TypeScript. It is MIT-licensed and self-hosted.
The dimensions
| Dimension | MCP | mpl |
|---|---|---|
| Layer | Transport and tool interface (JSON-RPC) | Contract, quality and audit layer |
| Solves | How agents talk to tools | Whether a message met its contract, and evidence that it did |
| Adoption | Broad; the common agent-to-tool standard | Early; Phase 3 (conformance suite, A2A hardening) in progress |
| Schema validation | JSON Schema on tool inputs (and optional outputs) | Versioned stype contracts on every message |
| Quality measurement | Not in scope | Six metrics, profile-driven |
| Policy enforcement | Not in scope beyond authorisation | Policy engine |
| Audit records | Not in scope | BLAKE3 content hash, provenance and quality report per message |
| Regulatory mapping | Not in scope | Primitives that map to common asks; not certifications |
| Mode: observe | n/a | Transparent mode — log, never block |
| Mode: enforce | n/a | Production/strict mode — invalid payloads and policy denials rejected at the proxy |
| Transports | Itself | MCP, A2A, HTTP |
| Language SDKs | Many official and community SDKs | Python, TypeScript (Rust core) |
mpl publishes no per-message latency figure, and we are not going to invent one. It is an extra hop doing schema validation, hashing and scoring; measure the cost on your own traffic in transparent mode before you decide to enforce.
When to use which
Use MCP alone when:
- You are prototyping, and a bad message is cheap.
- The agent’s actions are reversible, such as read-only queries on your own data.
- A single agent calls a couple of tools — mpl would be overhead with no return, and a guardrail library will cover the realistic failure modes more cheaply.
Use mpl on top of MCP when:
- The agent takes actions that are hard to reverse: sending messages, making purchases, changing settings.
- Someone — an auditor, a regulator, an incident review — will ask months later what an agent sent and why, and your logs cannot answer.
- You want malformed requests rejected before they reach your server: every calendar event must have a title, a start and an end.
- You want a quality signal you can trend, such as whether outputs are grounded in the sources they cite.
- Several agents or teams need one contract and one policy language across their traffic.
Where mpl loses
- Contract authoring is real work. Semantic types have to be written by someone who understands the domain, and a wrong contract is worse than none because it is trusted.
- Quality scores are experimental. Reducing an exchange to a number invites the usual pathology of a metric under pressure. We treat QoM as a trend to watch more than a threshold to enforce.
- It is early. MCP has a broad ecosystem; mpl’s conformance suite and A2A hardening are still in progress.
- It is not a guardrail. Guardrails ask “is this safe?”; mpl asks “did this meet the contract, and can you prove it?”. If you need content filtering, run one alongside.
A two-minute setup
mpl runs as a sidecar proxy in front of your existing MCP server. The agent’s code does not change; you point it at the proxy instead of the server. From the mpl quickstart:
# 1. install the CLI
cargo install mplx
# 2. run the proxy in transparent (observe-only) mode
mpl proxy http://your-mcp-server:8080
# dashboard at http://localhost:9080
# 3. learn contracts from live traffic, then approve them
mpl schemas generate
mpl schemas approve --all
# 4. switch to enforcement
mpl proxy http://your-mcp-server:8080 --mode production
# invalid payloads and policy denials are now rejected at the proxy
To call it from application code, install an SDK:
pip install mpl-sdk (Python) or npm install @mpl/sdk
(TypeScript).
Regulatory asks: evidence, not compliance
Teams usually reach for mpl because of an obligation — SOX, GDPR, HIPAA, the EU AI Act — and it is worth being exact about what it does there. mpl produces primitives that map to common asks:
| Regulatory ask | Primitive mpl produces |
|---|---|
| Tamper-evident records (e.g. SOX) | Per-payload BLAKE3 content hash in every audit record |
| Data-handling proof (e.g. GDPR) | Consent references and policy enforcement |
| Accuracy controls (e.g. HIPAA) | Quality thresholds on clinical-domain stypes |
| Transparency (e.g. EU AI Act) | Quality scores and provenance on each message |
These are evidence, not certifications. mpl does not make a system compliant with any of these regimes; your compliance programme decides what the obligations require and whether the evidence satisfies them. What mpl changes is whether the evidence exists at all — MCP, by design, does not produce it.
What to read next
- Formalising Prompts as First-Class Research Objects — the prompt side of the same problem
- mpl site — features, quickstart and FAQ
- mpl repository
- MCP specification