Field note · architecture

What Evidence Should Technical Buyers Collect About MCP or API?

A reproducible MCP-versus-API test packet for permissions, traces, failure recovery, rollback, cost fields, ownership, and veto decisions.

11 minute read
  • MCP
  • APIs
  • Architecture
  • Security
Illustration of a technical buyer comparing an MCP tool surface with an API contract using a redacted evidence packet

I ran one approval workflow twice, once through an API-shaped interface and once through an MCP tool surface. The result was less dramatic than most comparison charts. Both paths could pass the workflow. The useful difference was in the proof they made me collect before I could trust the result.

Illustration of a technical buyer comparing an MCP tool surface with an API contract using a redacted evidence packet

Observed results from the paired workflow

The bounded run did not produce a universal winner. It produced a minimum evidence packet.

Run on 2026-08-24API pathMCP path
Surface requests recorded910
Request bytes1921,139
Response bytes1,5702,991
Total local dispatch time0.079 ms0.070 ms
Approval promptacceptedaccepted
Approval attempts2, including one induced 5032, including one induced failure
Rollbackapproved to pendingapproved to pending
Token costNot applicableNot applicable

The extra MCP request was protocol setup and tool discovery: initialize and tools/list. The larger MCP payloads came from JSON-RPC envelopes and tool descriptions. Those are observations from this harness, not production latency or price estimates.

Observed result: both surfaces rejected a token issued for the other surface; denied req-002 because it was outside the test principal's object boundary; allowed a read-only token to read req-001 but not approve it; failed the first approval attempt with a retryable 503; reached approved after one retry; and rolled the record back to pending.

That is the sourceable result on this page: a buyer can compare MCP and API without arguing from protocol labels. Run the same state transition twice, then keep the traces and the veto conditions.

What should technical buyers collect before choosing MCP or API?

Collect evidence about the workflow, contract, authority, execution, recovery, and ownership. A vendor answer is useful only when it points to a reproducible field in the packet.

Evidence areaWhat to collectWhat it lets you decide
Workflow fixtureSanitized input, expected states, hidden fields, acceptance criteriaWhether both paths solve the same job
ContractOpenAPI document or MCP initialization and tools/list output, versions, schemasWhether a reviewer can inspect and diff the interface
AuthorityAudience, principal, scopes, object boundary, 401 and 403 behaviorWhether the caller can do only what it should
Data boundaryReturned fields, excluded fields, tenant or object rulesWhether sensitive data crosses the surface
ExecutionExact endpoint or tool call, transport, settings, timestamps, latency, bytes, cost basisWhat actually happened, not what the demo implied
Human controlApproval prompt, elicitation, decision, cancellation, and call bindingWhether side effects need an accountable human decision
Failure recoveryInduced failure, retry policy, retry count, state before and afterWhether the system fails in a bounded way
RollbackCompensating command, result, and final stateWhether an error is reversible
AuditRedacted events with principal, scope, object, action, status, and timestampWhether an incident can be reconstructed
Decision recordOwner, acceptance criterion, veto condition, dateWho can say go, hold, narrow, or reject

If a row is blank, the buyer doesn't have integration evidence yet. They have a promise.

How did the API and MCP paths differ?

The API path exposed three OpenAPI operations: read a request, set its decision, and roll it back. The MCP path exposed three tools with JSON Schemas: get_approval_request, set_approval_decision, and rollback_approval_decision. The business workflow was equivalent, but the inspection surface differed.

The OpenAPI Specification describes the paths and operations that make up an API contract. In the test packet, that contract made the HTTP verbs, paths, operation IDs, and response cases inspectable before execution.

The MCP tools specification uses tools/list for discovery and tools/call for invocation. It also allows annotations such as read-only, destructive, and idempotent hints. Those hints helped the client policy decide when to show an approval prompt, but the specification says clients must treat tool annotations as untrusted unless they come from trusted servers. A buyer should therefore record the annotation and the independently enforced permission, not accept the annotation as proof of safety.

The evidence question is not “does MCP have tools?” or “does the API have endpoints?” It is “can the team show the exact inventory, schema, permission, and side-effect behavior for the tool or endpoint that owns this workflow?”

Illustration of equivalent approval workflow states mapped to API endpoints and MCP tool calls with matching inputs and outputs

Which authorization and permission checks belong in the packet?

The packet needs both identity checks and action checks. A valid token is not evidence that the caller is allowed to read this object or invoke that mutation.

The MCP authorization specification says HTTP-based MCP implementations that support authorization should follow its flow. It calls for protected-resource metadata discovery, bearer tokens in the Authorization header, audience validation, and distinct handling for invalid authorization and insufficient scope. In the local run, I used pre-issued test tokens to exercise the request-time audience and scope checks. I did not run an external OAuth server or a PKCE exchange, and the packet says so.

The test exercised four checks for each surface:

  1. A read token with requests:read could read req-001.
  2. The same read token could not approve req-001; the API returned HTTP 403 and MCP returned a JSON-RPC error with insufficient-scope data.
  3. The read token could not read req-002, which was outside the principal's object boundary.
  4. A token for the other surface was rejected as the wrong audience.

This maps to the OWASP API Security Top 10, especially broken object-level authorization, broken function-level authorization, and improper inventory management. It also catches a common procurement mistake: asking whether a vendor supports OAuth without asking which object and function checks run after the token is accepted.

The evidence fields should include the token audience, principal, requested scope, effective scope, object ID, operation or tool name, expected status, actual status, and redacted WWW-Authenticate response where relevant. Never publish or store the raw bearer token in a trace.

What should count as a human approval prompt?

Record the approval mechanism separately from protocol authorization. Authorization answers whether a caller may act. Approval answers whether a person accepted this particular side effect.

The local client policy showed this prompt before the mutation:

Approve the requested mutation for req-001?
Decision: accepted
Operation: mutating approval

That was a client policy prompt, not an MCP elicitation exchange. The distinction matters. The MCP elicitation specification describes form and URL modes for collecting information or directing a user to a sensitive out-of-band interaction. It also distinguishes that from authorizing the MCP client to access the MCP server. If a product says “human in the loop,” ask for the exact message, the operation it binds to, what decline and cancel do, and whether the server or client owns the decision record.

For an API, the approval may live in the application that calls the endpoint. For MCP, a client may mediate the approval before tools/call. Neither arrangement is automatically safer. The packet must show the boundary and the state transition.

Method and sample

The sample was one synthetic approval workflow run once through each surface, with fresh in-process state for each run. In other words, this is n = 1 paired workflow per surface, not a production benchmark. The method was to dispatch equivalent HTTP-shaped API requests and JSON-RPC 2.0 MCP messages, using the same fixture, client policy, tokens, scopes, approval prompt, induced 503, retry limit, and rollback assertion.

Run the same fixture and inspect the dated outputs. The harness is deliberately small so a reviewer can understand the entire path.

cd .content-studio/jobs/1787560708050-fd7550ed
python3 scripts/run_comparison.py

The run writes:

fixtures/approval-request.json
scripts/run_comparison.py
traces/api-2026-08-24.json
traces/mcp-2026-08-24.json
artifacts/comparison.json
artifacts/decision-record.json
artifacts/evidence-packet-template.md

The raw fixture is synthetic and contains no customer data:

{
  "fixture_version": "2026-08-24.1",
  "workflow": "read an approval request, attempt an approval with insufficient permission, approve after a client confirmation, then roll back",
  "request": {
    "id": "req-001",
    "requester": "user-redacted-001",
    "department": "engineering",
    "category": "software",
    "amount": 1250,
    "currency": "EUR",
    "status": "pending",
    "policy": "amount_over_1000_requires_procurement_review",
    "internal_note": "DO NOT EXPOSE: synthetic fixture secret"
  },
  "unavailable_request": {
    "id": "req-002",
    "owner": "other-team",
    "reason": "used to test object-level authorization"
  },
  "expected": {
    "initial_status": "pending",
    "approval_status": "approved",
    "rollback_status": "pending",
    "hidden_fields": ["internal_note"]
  }
}

The fixture deliberately includes internal_note as a sentinel. Both implementations must omit it from the returned payload. That is a data-boundary assertion, not a claim about a real secret.

What did the failure and rollback traces contain?

The API trace used HTTP status codes. The MCP trace used an outer HTTP success for the JSON-RPC envelope and carried permission or backend errors inside the JSON-RPC response. A buyer should normalize both into the same comparison fields.

[
  {"label": "object authorization denied", "status": 403, "reason": "object_not_allowed"},
  {"label": "write denied", "status": 403, "reason": "insufficient_scope"},
  {"label": "approve attempt 1", "status": 503, "retryable": true},
  {"label": "approve attempt 2", "status": 200, "after": "approved"},
  {"label": "rollback", "status": 200, "after": "pending"}
]

The block above uses normalized application statuses so API and MCP can share one table. For MCP, the transport status was 200 because the JSON-RPC envelope was valid. The equivalent permission failure was recorded inside that envelope:

{
  "jsonrpc": "2.0",
  "error": {
    "code": -32003,
    "message": "insufficient_scope",
    "data": {"error": "insufficient_scope", "required_scope": "requests:write"}
  }
}

The raw traces also contain redacted audit events. Each event keeps a timestamp, surface, principal, scope list, object ID, action, status, and before or after state where a mutation occurred. The authorization token is represented as REDACTED.

Illustration of a redacted API and MCP trace showing a permission denial, one retry, approval, and rollback

When should the buyer veto both options?

Use the decision record to assign a person or role to each material risk. A score without a veto owner is still an unowned demo.

RiskOwnerEvidence requiredVeto condition
Wrong audience or excessive scopeSecurity ownerAudience, scopes, 401 and 403 results, token redactionVeto if another audience's token works, read scope mutates, or a raw token is logged
Object or function authorizationSystem ownerOut-of-bound object and read-only mutation testsVeto if req-002 is readable or read permission invokes approval
Contract or inventory driftIntegration ownerVersioned OpenAPI paths or MCP tools/list, schemas, and operation namesVeto if the executable contract cannot be diffed
Unreviewed side effectWorkflow ownerMutation classification, prompt, decision, and call bindingVeto if mutation can execute without explicit approval
Failure and rollbackService ownerInduced failure, retry limit, state transition, rollback resultVeto if retry is unbounded, state is unknown, or rollback is unproven
Data boundary and auditSecurity ownerExcluded-field check and redacted event trailVeto if sensitive fields cross the boundary or the action cannot be reconstructed

The reusable packet template in this job directory follows the same order: scope, contract, authorization, execution, recovery, audit, and decision. It is intentionally boring. Boring evidence is easier to review than a polished demo.

When I taught product managers to move from writing specifications to building and shipping, the failure was usually not the model. It was that nobody could say what “done” meant. That observation from Marius Manolachi's teaching experience is a framing device here, not a rate or a security result. Define done as acceptance criteria and vetoes before a vendor demo starts.

Illustration of a buyer-ready evidence packet with named owners, acceptance criteria, and veto conditions beside an MCP or API decision

Limitations and unknowns

This n = 1 paired run shows how to collect comparable evidence. It does not prove that API calls are faster, that MCP calls cost more, or that one surface is safer in production. The latency is local Python dispatch overhead. The byte counts are request and response sizes in this harness, not provider billing. The run has one synthetic workflow and one induced failure per path.

It also does not prove that an MCP server has correctly implemented an external OAuth flow. The MCP authorization specification has requirements for discovery, resource indicators, audience-bound tokens, and PKCE. This harness used pre-issued local tokens to make permission behavior reproducible. A real procurement packet should add the authorization-server metadata, consent screen, redirect URI, PKCE, token lifetime, refresh, and revocation evidence before approval.

What remains unknown is how the two surfaces behave over the buyer's production-like network path, under concurrency and rate limits, with provider billing, an external OAuth server, and the buyer's actual rollback semantics. The next test should use the buyer's own sanitized workflow and keep the same fields so the second run remains comparable.

What should a technical buyer approve?

Approve an MCP or API path only when the same owned workflow passes its acceptance criteria and every material risk has an owner and a veto condition. Choose an API when the team needs a stable endpoint contract and already controls the integration boundary. Choose MCP when the workflow must be discoverable as a schema-described tool surface for MCP clients and the team can own client-mediated approval and protocol-specific authorization.

If you are still deciding the architecture, use this evidence packet alongside How to Choose MCP or a Custom API. If the chosen path is MCP, the security checks in How to Secure an MCP Server for AI Agents are the next review layer.

For a team that needs help turning an owned workflow into a reproducible test and a decision record, work with Marius Manolachi on AI capability after the packet's scope and veto authority are explicit.

Questions people ask next

Should buyers compare MCP and API on the same workflow?

Yes. Use the same fixture, acceptance criteria, permissions, failure injection, and rollback test. Otherwise the comparison measures different business problems instead of different integration surfaces.

Does an MCP tool replace an API contract?

No. MCP adds a discoverable tool surface with schemas and JSON-RPC calls. The underlying system still needs an owned authorization, data, failure, and audit contract.

Are local latency numbers enough to choose MCP or API?

No. A local harness can expose extra protocol steps and measurement fields. Production choice requires network, proxy, runtime, concurrency, rate-limit, and provider-cost evidence too.