The Missing Authorization Layer in Agent Payments
Every agent payment protocol solved 'who paid' and nobody solved 'who said they could.' Working on dispute evidence across three protocols is how I found the gap.
You've seen agent payment demos. An AI agent browses an API marketplace, hits an HTTP 402, pays with crypto, gets the data, moves on. Clean. Automatic. The demo always works.
Now imagine the agent was yours. You told it to grab some market data. You expected a few cents. It found a premium tier for $5.00 per call and ran 40 requests before you checked your wallet. The payments were authentic. The signatures were valid. The settlement was final. Nobody committed fraud.
So what went wrong? And which part of the payment protocol can you point at to prove it?
None of them. That's the gap.
The assumption nobody questioned
Every agent payment protocol today is built on one assumption: the entity signing the payment is the entity authorizing the payment. x402, the most active protocol in the space (18 chains, SDKs in four languages, Coinbase and Cloudflare on the steering committee), works like this: server returns a 402 with payment terms, the agent constructs a signed payment, a facilitator verifies the signature and settles the transaction. Three parties, clean separation of concerns.
The assumption holds when a human opens their wallet and taps "pay." It stops holding the moment an AI agent acts on their behalf. The agent has the signing key. The agent picks the payment. The facilitator verifies the cryptography. Everyone did their job. But nobody in the flow ever checked whether the human behind the agent said "yes, spend up to this much, with these merchants, before this date."
Authentication without authorization. The payment system knows who paid. It has no idea who said they could.
I found this by looking somewhere else
This wasn't where I started. I started with dispute evidence.
The question was simpler: when an agent payment goes wrong, how does a third party verify the claim mechanically, without calling either side? I proposed an evidence triple across three protocols (ACK, MPP, x402): embed the authorization scope, the payment receipt, and a structured delta describing the mismatch. A resolver extracts the named fields, confirms the values differ. No interpretation, no callbacks.
Four other developers engaged with the x402 thread (#3500) and each one added a failure class I hadn't considered.
The first pointed out that in an agent overspend scenario, the offer and receipt match perfectly. The agent paid exactly what the server asked. The thing it violated, "fetch market data under $0.10," exists in neither artifact. It lives in the authorization scope the user delegated to the agent, which is invisible to every party in the forward flow. Without that scope artifact, the dispute is just one side's word against the other.
The second added a different failure mode: what if the offer itself was wrong? The agent acted on a 402 response where the pay_to address had been swapped. Valid signature, correct settlement, agent within budget, but the counterparty wasn't who the endpoint historically settled to. That's a bad-information problem. And it requires a third party to attest to settlement history, which is a different verification model than comparing two fields.
The third raised the concurrency case: two payments that each pass individually but together exceed the budget. Payment A is $80 against a $150 budget. Payment B is $80 against the same budget. Both are in-scope in isolation. The violation is only visible when you have the full set.
The fourth brought delivery conformance: everything matches (offer, receipt, scope) but the actual content delivered doesn't match what was promised. Three days of forecast data paid for, two delivered.
Each contribution added a layer. Together they shaped a tiered reason code model: mechanical codes that a resolver evaluates by extracting fields and comparing (scope exceeded, budget exceeded), attested codes that require a third party to sign (counterparty mismatch, offer drift), and state-dependent codes that need the full receipt set to evaluate (aggregate budget exceeded).
But every tier had the same prerequisite. Every single failure class needed one artifact that doesn't exist: the authorization scope. The thing the user told the agent it was allowed to do.
Nobody has both halves
I surveyed every implementation I could find. The pattern was the same everywhere.
ACK (Agent Commerce Kit) has the most complete delegation model. A grant is a JWT signed by the owner, carrying the agent's identifier, constraints, audience, and expiry. The relying party verifies it offline, no callbacks needed. But ACK has no budget enforcement at the protocol level. The grant says "up to $100" and nothing prevents the agent from spending it across 50 separate payments that individually look fine.
The Budget Reservation Protocol (proposed in the AP2 ecosystem) has budget enforcement. Four verbs: authorize, commit, refund, query. The authorize call is an atomic check-and-decrement, no separate "check remaining balance" that two agents could race on. But it has no delegation. It knows about "principal" and "agent" as roles but doesn't define how that relationship is established or verified.
x402's SpendControls cap per-payment amounts on the client side. The server and facilitator can't see them. They can't be verified by a third party. They can't constrain recipients, categories, or aggregate spend.
One system has delegation without budget enforcement. Another has budget enforcement without delegation. A third has client-side limits invisible to everyone else. The Budget Reservation Protocol spec puts it plainly: "A delegation without a budget has no spending limit. A budget without a delegation has no proof of authorization."
Both halves exist in isolation. Nobody connected them.
Where delegation fits without changing the protocol
x402 has an extension model designed for exactly this kind of additive capability. Every payment message (PaymentRequired from the server, PaymentPayload from the client, VerifyResponse from the facilitator) carries an extensions field. Nine extensions already ship this way: offer-and-receipt, payment identifiers, auth hints, and others. No core spec changes needed.
A delegation extension works like this: the principal issues a JWS grant to the agent containing the agent's address, constraints, audience (which servers this grant works with), the asset denomination, and an expiry. The server advertises that it accepts or requires delegation in its 402 response. The agent attaches the grant in its payment payload. The facilitator extracts the grant on /verify, checks the principal's signature, confirms the agent's address matches the payer's address (a simple comparison, not DID resolution), verifies the constraints, and returns the result.
Servers that don't understand delegation ignore the extension field. Agents without grants pay directly. Backward compatible at every layer.
The constraint model splits into two categories. Well-known constraints (maximum amount, allowed recipients, expiry, allowed networks) that the facilitator must understand and enforce. And audience-scoped constraints (namespaced, like vendor:acme/category) that the facilitator passes through to the server, because the server, not the facilitator, has the business relationship with the principal. This avoids a deadlock where generic facilitators would need to understand every custom constraint any principal might invent.
The edges that matter
Three design decisions shaped the proposal more than the core mechanism.
Privacy. A grant reveals the principal's identity, the agent relationship, and the spending constraints. For enterprise B2B that's fine. For consumer use it's a problem. The solution uses infrastructure that already exists: the facilitator already sits between the agent and the server. The grant goes to the facilitator. The facilitator verifies it and returns an attestation to the server. The server learns "this payment is authorized" without learning by whom or under what limits. The principal controls the disclosure level per-grant: full visibility, attestation only, or just a boolean. No zero-knowledge proofs, no new cryptography. Just routing.
Multi-chain denomination. x402 runs on 18 chains. A grant that says "max $10" means nothing if it doesn't specify which asset on which chain. USDC on Base has 6 decimals. A hypothetical 18-decimal token on another chain would interpret the same integer as a vastly different amount. The grant carries its own denomination (CAIP-19 identifier), and the facilitator converts at verify time. Without this, the constraint is uninterpretable, and two honest implementations would disagree on whether a payment exceeds the limit.
Who holds the evidence in a dispute. In a scope violation, the agent is the party that exceeded its authority. The agent has no incentive to produce the grant that proves it was wrong. But the principal issued the grant. The principal retains it. The facilitator retains a hash of the grant (not the contents) to confirm it was the grant used in that specific transaction. At dispute time, the principal presents their copy and the facilitator's hash links it to the payment record. The evidence is in the hands of the right party.
What this completes
The delegation grant is the authorization-scope artifact that the dispute taxonomy needs. With it, the evidence triple is complete: grant (what the principal allowed), offer (what the server promised), receipt (what was settled). A dispute resolver verifies the grant against the receipt. The hash reference connects them without coupling the dispute extension to the delegation extension's internals.
Without delegation, dispute evidence can prove "the receipt doesn't match the offer." That's a payment error. With delegation, it can prove "the agent exceeded its authority." That's the problem agent payments actually have.
The delegation proposal is live on x402. The dispute evidence thread that started this is still active. The next research threads are session binding (how multiple payments relate to one agreement) and the refund pipeline (how dispute evidence connects to remediation when the dominant payment scheme has no refund path).
Agent payment protocols solved authentication years ago. Authorization is the next layer. And until it exists at the protocol level, every agent with a signing key is spending on trust.
Sources
- x402 Delegation Extension Proposal (Issue #3693) - the delegation grant proposal for x402, introducing principal-issued JWS grants with constraint enforcement at the facilitator layer
- Dispute Evidence for Agent Commerce (Issue #3500) - the original thread on structured dispute evidence across x402, ACK, and MPP that surfaced the delegation gap
- Dispute Evidence for Agent Commerce - the evidence triple model (authorization scope, payment receipt, structured delta) for machine-verifiable dispute resolution
Frequently Asked Questions
Built by Trio, a fintech-native engineering partner helping teams build the next generation of financial technology and infrastructure.
Subscribe to Ledger Drift for high-signal insights into how modern fintech is built, from systems to code to teams.