Field note · architecture

Why Does AI Fail When Integration Choice Is Unclear?

A paired direct-function and MCP test shows how to classify integration failures and assign repair ownership after a promising demo.

10 minute read
  • AI architecture
  • MCP
Illustration of a paired AI integration test separating direct API and MCP repair ownership

The demo is usually not where this failure becomes visible. It appears when a user asks for a capability that sits between a model, a tool surface, a policy check, and a system of record.

When I taught product managers to move from writing specifications to building and shipping, the hard part was often not the model. It was deciding what “done” meant. Integration failures have the same shape: everyone can point at a component, but nobody owns the boundary.

The paired test result

The direct function/API path and the MCP path produced the same classification in all seven cases because the planner inputs, schemas, policy, and backend behavior were held constant. The test did not show that one surface is inherently safer. It showed where the first failed invariant belonged.

PathCaseCandidate setSelectedArgumentsValidationAuthorizationResultVerified effectLatency msRepair owner
directunavailable capability[]none{}not runnot rununavailablenone0.000routing in code
MCPunavailable capability[]none{}not runnot rununavailablenone0.000routing in code
directoverlapping toolsget_status, lookup, change, deleteorders_lookupA-17passpasswrong capabilitynone0.010routing in code
MCPoverlapping toolsget_status, lookup, change, deleteorders_lookupA-17passpasswrong capabilitynone0.005routing in code
directinvalid argumentsget_status, lookup, change, deleteorders_get_statusorder_id: 17fail typenot runinvalid argumentsnone0.005tool contract
MCPinvalid argumentsget_status, lookup, change, deleteorders_get_statusorder_id: 17fail typenot runinvalid argumentsnone0.003tool contract
directunauthorized actionget_status, lookup, change, deleteorders_change_statusA-17, shippedpassdeny viewerunauthorizednone0.005executor policy
MCPunauthorized actionget_status, lookup, change, deleteorders_change_statusA-17, shippedpassdeny viewerunauthorizednone0.004executor policy
directambiguous resultget_status, lookup, change, deleteorders_get_statusA-17passpassambiguousnot verified0.003result contract
MCPambiguous resultget_status, lookup, change, deleteorders_get_statusA-17passpassambiguousnot verified0.003result contract
directretry after commitget_status, lookup, change, deleteorders_change_statusA-17, k-retrypasspassalready appliedverified once0.004executor idempotency
MCPretry after commitget_status, lookup, change, deleteorders_change_statusA-17, k-retrypasspassalready appliedverified once0.004executor idempotency
directirreversible side effectget_status, lookup, change, deleteorders_deleteA-17, not-confirmedpasspassconfirmation requirednone0.003routing in code
MCPirreversible side effectget_status, lookup, change, deleteorders_deleteA-17, not-confirmedpasspassconfirmation requirednone0.003routing in code

The local adapter measurements are not a speed comparison. They exclude model generation, network time, and real backend time. The useful result is the ownership split.

Illustration of a paired AI integration failure matrix mapping traces to repair owners

Why the integration surface can become the failure

An integration is not just a wire format. It decides what the model can see, which calls are valid, where authorization runs, and how a result says “verified.”

MCP tools are model-controlled and expose names, descriptions, input schemas, and optional output schemas. The specification also separates protocol errors from tool execution errors and calls for input validation, access controls, result validation, timeouts, and logging. MCP Tools specification

Direct function calling has the same design pressures under a different envelope. OpenAI function definitions include a name, description, JSON Schema parameters, and a strict setting. The API also provides tool_choice, allowed_tools, and a way to disable parallel calls. OpenAI function calling guide

Anthropic makes the same boundary visible through detailed tool descriptions, JSON Schema input, tool-choice controls, and strict tool use. Claude tool definitions

The shared lesson is an analysis, not a vendor claim: changing the surface does not decide who owns ambiguity. Your system still needs a named owner for routing, contracts, policy, execution, and verification.

Microsoft Research gives the problem a useful external name. Its 2025 analysis of 1,470 MCP servers describes tool-space interference, including exact name collisions, semantic overlap, large tool surfaces, and long responses. Microsoft Research

Classify the first failed boundary

Do not start by asking whether MCP or a custom API is better. Start with the first invariant that failed.

Failure signalWhat it meansRepair first
No candidate can express the taskCapability routing is incompleteKeep routing in code or add one scoped tool
A valid call chooses a near-matchThe model-facing surface is ambiguousNarrow the surface or constrain allowed tools
Arguments fail before executionThe contract is wrong or driftedChange the schema and regenerate both adapters
A valid call is deniedRuntime policy or credentials own the failureRepair executor authorization, not the prompt
Result is parseable but unverifiableThe output contract hides identity or verificationReturn explicit status, identity, and verified state
Timeout follows a committed writeRetry ownership is missingAdd idempotency and verification in the executor
Irreversible action lacks confirmationThe action boundary is too looseKeep approval routing in code

This matrix is the practical answer to the query. “AI failed” is not a repair category. “The call selected a near-match and returned a valid record” is.

When should you narrow the tool surface?

Narrow the surface when two tools can satisfy the same user wording but produce different meanings. In the test, orders_get_status and orders_lookup both accepted A-17. The call passed validation and authorization, yet it was still the wrong capability.

The repair belongs in routing or the exposed candidate set. Rename tools only if the names are part of the ambiguity. Otherwise, expose the smallest allowed subset for the current workflow or select the route in code before the model sees tools. OpenAI documents allowed_tools for restricting callable tools, and Microsoft Research recommends exposing as few tools as possible. OpenAI function calling guide, Microsoft Research

Verify the repair by replaying the overlap case. A passing result is not enough. The selected tool must match the requested capability, and the trace must show why the other tool was unavailable or rejected.

When should you change the API or MCP contract?

Change the contract when the first failure is validation or verification. In the invalid-argument case, order_id: 17 failed before execution. In the ambiguous-result case, the response was structurally readable but carried two candidate records and verified=false.

The direct and MCP definitions deliberately carry the same fields, even though one uses parameters and the other uses inputSchema. MCP also supports an output schema, and its specification recommends client validation of structured results. MCP Tools specification

Do not repair a schema failure with a stronger prompt. Add the missing type, enum, required field, identity key, or verification state to the contract. Then regenerate both surfaces and replay invalid, ambiguous, and success cases.

When should routing and approval stay in code?

Keep routing in code when the system must abstain, choose among semantically overlapping tools, or gate an irreversible action. These decisions depend on application state and authority, not only language understanding.

The test blocked orders_delete because confirmation was not-confirmed, even though the operator was authorized. That is a code-owned stop rule. MCP's specification says applications should show the tools being exposed and provide confirmation prompts for sensitive operations. MCP Tools specification

For a concrete preview boundary around risky actions, pair this decision with the dry-run mode guide.

Authorization is also an executor concern. MCP's authorization specification distinguishes optional authorization support from the token and scope checks required for protected HTTP servers. MCP Authorization specification

Verify these repairs with a negative test: the unauthorized action and unconfirmed delete must produce no external effect. A natural-language refusal after the side effect is not a passing test.

Worked decision artifact for a post-demo incident

Use this sequence in the incident review:

  1. Freeze the trace. Save the exact user task, candidate set, selected tool, arguments, validation result, authorization result, tool status, external effect, and latency. Do not begin with a rewritten prompt.
  2. Find the first failed invariant. If no capability was exposed or a near-match was selected, start with routing. If arguments failed, start with the contract. If authorization failed, start with the executor policy. If the result was ambiguous or a retry followed a commit, start with result or execution semantics.
  3. Assign one owner. The owner must be one of application routing, API/MCP contract, executor policy, executor idempotency, or result verification. “The AI team” is not an owner.
  4. Replay the case and its neighbor. A schema fix needs a valid and invalid call. A routing fix needs the overlap case and an unavailable-capability case. A retry fix needs timeout-after-commit and a second request. An approval fix needs confirmed and unconfirmed actions.

The repair decision is complete only when the original failure is classified, the external effect is verified, and the same trace can be replayed with the new owner visible.

What this test does not prove

This was a small synthetic local fixture run on 2026-08-24 with scripted-planner/v1. It did not call a live OpenAI or Anthropic model, a real MCP server, a remote API, or a production system. It does not estimate failure frequency, model-selection accuracy, provider differences, or network latency. Its claim is narrower: when the planner input and backend behavior are fixed, the integration surface alone did not change the classification, while the trace made repair ownership explicit.

If you are still choosing between the two surfaces, use the parent guide on how to choose MCP or a custom API. If you already have a working demo, replay the seven cases before expanding the tool set. For a team that needs a tighter build-and-ownership practice, Marius Manolachi's AI consulting and tutoring work is about making existing people capable of building AI products on their own work.