Field note · commercial

What Should a Buyer Measure After an AI Consultant Leaves?

Use a 30/60/90 scorecard and an owner-run failure drill to decide whether an AI workflow is ready to continue after the consultant leaves.

12 minute read
  • AI consulting
  • AI measurement
  • Team capability
Illustration of a buyer reviewing an AI workflow scorecard with an internal owner, failure case, and decision gate

The handoff meeting is a poor finish line. A folder of prompts, diagrams, and recordings can still leave the buyer dependent on the person who built them.

I use a harder test: can the internal owner run the evaluation, explain one failure, make the next decision, and contain the workflow when it misbehaves? That is what the scorecard below is designed to show.

Observed result from the bounded sanitized run: the first evaluation passed 7 of 8 rows. One critical classification failure was corrected on rerun. The rollback drill routed all eight rows to manual review with zero unreviewed writes. The decision was REVISE, because the exercise had no production cost or latency telemetry.

Illustration of a buyer moving from an AI handoff pack to an owner-run evaluation, failure drill, and go or revise decision

What should a buyer measure first?

Start with one workflow and one evidence pack. Record the business baseline, current outcome, adoption, operational change, technical health, failures, cost, latency, owner, evaluation method, review cadence, and next decision in the same place.

This keeps activity separate from value. McKinsey's five-layer framework makes the same separation across technical performance, user adoption and engagement, operational KPIs, strategic outcomes, and financial impact. It also attaches an owner to each layer. Read the published framework.

The NIST AI Risk Management Framework gives the risk-management spine: Govern, Map, Measure, and Manage. A buyer can translate that into a simple question for the exit pack:

QuestionScorecard fieldProof to keep
What was the work like before the engagement?Original business baseline and baseline definitionDated process measure, sample, or explicit statement of what was not measured
What is different now?Current operational outcomeBefore and after result, with the comparison unit named
Is the workflow actually used?Adoption or workflow penetrationEligible tasks, completed tasks, user or role breakdown
Does it fail safely?Quality and safety failuresEvaluation rows, incident log, severity, containment, and escalation
Can it run economically?Cost and latencyCost per run or case, p50 or p95 latency, and measurement method
Who can operate it?Named internal ownerPerson or role, access, runbook, and decision authority
Can the result be reproduced?Evaluation set and rubricInputs, expected outputs, scoring logic, and raw rows
What happens next?Decision cadence and next decisionReview date, threshold, and go, revise, or stop outcome

If a field is unknown, write unknown and assign an owner. An empty cell is evidence about the handoff. It is not a reason to quietly delete the metric.

Use this blank post-engagement scorecard

Copy this template for one workflow. Keep the scope small enough that the internal owner can run it without asking the consultant how the system works.

FieldFill inMinimum evidence
Workflow and boundaryTrigger, input, output, excluded casesOne-sentence workflow contract
Review dateDay 0, day 30, day 60, or day 90Calendar owner and decision forum
Original business baselineWhat happened before, measured how, during which datesRaw sample or dated business measure
Current operational outcomeCycle time, rework, throughput, accuracy, or another workflow resultSame unit and comparison period as baseline
Adoption or penetrationEligible work, AI-assisted work, acceptance, override, and manual bypassEvent or queue counts by role
Quality and safetyPass rate, critical errors, unsafe outputs, incidents, and escalationsEval rows plus failure log
Cost and latencyCost per case, total run cost, p50/p95 latency, retriesInstrumentation source and unit
Internal ownerName, role, permissions, and authority to stop or escalateOwner acknowledgment and escalation contact
Evaluation set and rubricCases, expected result, grader, threshold, and exception ruleVersioned cases and raw scores
Rollback or escalationManual path, disable switch, data protection, and notification pathDry-run record or incident test
CadenceWho reviews which layer and whenCalendar invite, ticket, or review log
Next decisionGo, revise, or stop, with the reason and due dateSigned decision record

The baseline needs a definition, not just a number. “Time saved” is incomplete unless you say whose time, on which unit of work, over what period, and whether review or rework is included. “Adoption” is incomplete unless you define the eligible workflow population.

The UK Government's AI procurement guidance is useful here even when the buyer is not a government department. It separates procurement from ongoing management and calls for lifecycle oversight, ongoing evaluation, knowledge transfer, support, auditability, and clear end-of-contract roles. See the procurement and ongoing-management guidance. Treat those points as buyer requirements to adapt, not as a private-sector law.

What does a worked scorecard look like?

Here is a sanitized example for a product-feedback routing workflow. The input is a short feedback note. The output is a category, severity, and next action. The rows are authored fixtures, so the example demonstrates the measurement and handoff procedure rather than claiming a customer result.

FieldWorked entryStatus
WorkflowRoute product feedback into category, severity, and next actionDefined
Original business baselineBefore handoff, there was no named owner, evaluation set, rubric, rollback drill, or review cadenceRecorded baseline
Current operational outcomeFirst run: 7 of 8 rows passed. The failed row passed after the rule change and rerunMeasured fixture result
Adoption or workflow penetrationOne internal-owner runner completed all eight rows. Production penetration was not measuredExercise result and explicit unknown
Quality and safety failuresOne critical false feature classification. The stop action still prevented an unreviewed writeFailure and containment recorded
Cost and latencyLocal exercise cost: €0.00. Production API cost and latency were not capturedInstrumentation gap
Named internal ownerOperations lead runs the evaluation. Product owner receives critical escalationsAssigned role
Evaluation set and rubricEight labeled rows; exact class, severity, action, critical stop rule, and no-write ruleVersioned fixture
Decision cadenceDay 0 exit pack, day 30 owner-run eval, day 60 adoption and operational review, day 90 value and cost reviewScheduled proposal
Next decisionREVISE: add cost and latency telemetry, expand edge cases, then rerun before scaleExplicit decision

This is the distinction I want buyers to preserve: the workflow can be operationally promising while the evidence pack is still incomplete. A polished demo does not turn missing cost telemetry into a zero, and a successful rerun does not prove long-term adoption.

How do you run the independent operator test?

Give the internal owner the blank scorecard, the evaluation set, the current run instructions, and the escalation path. Then step away. The owner must complete five actions.

  1. Reproduce the run. Execute every case and save the raw output, score, timestamp, and workflow version.
  2. Investigate one failure. Start from the input and trace the output to the rule, prompt, retrieval result, tool response, or human decision that produced it. Preserve the failing row.
  3. Record the change decision. Choose GO, REVISE, or STOP. State the threshold, the observed evidence, the change owner, and the next review date.
  4. Contain the workflow. Switch to the documented manual path or disable the affected action. Confirm that no unreviewed write or external action occurred during the drill.
  5. Escalate when required. Send the preserved failure to the named technical, product, security, or compliance owner. Record who accepted it and what happens next.

The test passes only if the owner can perform all five actions without a consultant explaining hidden steps. A consultant can observe the exercise, but the buyer should not count a consultant-led rescue as transfer.

My own teaching observation points to the same boundary. I taught product managers who went from writing specs to building and shipping the product, and automating work around it. The recurring failure is often not that the model is incapable. It is that nobody can say what “done” means. That is why the scorecard makes the threshold, failure response, and next decision explicit. This is the bounded observation recorded in my AI consulting and tutoring work, not a claim about a measured failure rate.

What did the operator test find?

The eight raw rows below make the result inspectable.

CaseSanitized inputExpectedFirst runResult
01Export loses timezone on some recordsbug, high, escalatebug, high, escalatepass
02Request for a dark-mode settingfeature, low, backlogfeature, low, backlogpass
03Invoice total becomes 0 after importbug, critical, stopfeature, critical, stopfail
04User asks where to change timezonehow-to, low, answerhow-to, low, answerpass
05Renewal reminders were sent twicebug, high, stopbug, high, stoppass
06Request for a weekly unpaid-invoice viewfeature, medium, backlogfeature, medium, backlogpass
07CSV import rejects comma decimalsbug, high, escalatebug, high, escalatepass
08User asks how to change notification settingshow-to, low, answerhow-to, low, answerpass

The owner investigated case 03. The note described a financial-data corruption symptom after import. The first output treated it as a feature request, even though the severity and stop action were present. The owner changed the rule to classify a numeric corruption symptom in an import or export path as a bug requiring stop or escalation. The rerun classified case 03 as bug, critical, stop.

The owner then ran the rollback drill. The workflow switched to manual-review, all eight rows entered the manual queue, the critical case was escalated to the product owner, and the record showed zero unreviewed writes. The drill passed as a control-path test.

The overall decision remained REVISE, not GO. The procedure proved that an operator could reproduce, diagnose, change, and contain the workflow. It did not prove production economics, because the exercise had no billed API call and no latency instrumentation. It did not prove adoption, because no live eligible-task denominator existed. Those are decision-relevant gaps, not editorial footnotes.

How should the 30/60/90 review cadence work?

Use the cadence to move from handoff proof to business proof. NIST's 2026 report groups deployed-AI monitoring into functionality, operational, human factors, security, compliance, and large-scale impacts. It also identifies open questions about who, what, when, why, and how to monitor, including the right cadence and the balance between automated and human-validated monitoring. Read the NIST monitoring report.

CheckpointBuyer questionEvidenceDecision
Day 0Can the internal owner operate and contain the workflow?Scorecard, eval set, owner test, failure and rollback recordDo not close the engagement until the owner can run it
Day 30Does it still work on current inputs, and are failures handled?Repeated eval, incident review, manual bypass count, owner notesGo or revise the workflow and controls
Day 60Is it used in eligible work, and is the process improving?Penetration, acceptance and override, cycle time, rework, qualityContinue only if adoption and operational measures move together
Day 90Does the business case justify continued investment?Strategic outcome, total cost, support load, risk and compliance recordScale, redesign, renew support, or stop

This cadence is a recommendation, not a universal standard. A high-risk workflow may need daily or per-release checks. A low-risk internal drafting workflow may use a lighter schedule. The owner, risk boundary, and failure cost should determine the frequency.

When should the buyer go, revise, or stop?

Make the decision rule visible before the consultant leaves.

DecisionUse it whenRequired next action
GOThe owner reproduces the eval, critical failures are within threshold, the manual path works, and cost, latency, adoption, and outcome evidence are presentContinue at the defined scope and cadence
REVISEThe workflow is useful but a control, measurement layer, edge case, or owner path is incompleteName the change, owner, threshold, and next review date
STOPThe workflow creates unacceptable harm, cannot be contained, misses a critical threshold, or has no accountable ownerDisable or contain it, preserve evidence, and decide whether to redesign or retire

The buyer should not let a consultant's departure date create a false deadline. If the internal owner cannot complete the operator test, the correct result is REVISE or STOP, even if the demo looked good.

For the wider commercial decision, see Choose Between an AI Consultant, Agency, or Internal Team. For a broader question about evidence that an AI partner made a team more capable, read What Evidence Proves an AI Partner Made the Team More Capable?. Those pages provide adjacent context. This page is the post-exit scorecard and test.

If you are buying an engagement, put the blank scorecard and operator test into the statement of work before the first build session. If you have already reached handoff, run it now. The result should change what you do next.

Questions people ask next

Is usage enough to prove an AI consultant created value?

No. Usage shows that people interacted with the workflow. Pair it with an operational outcome, quality and safety checks, cost and latency, and an owner-run evaluation before deciding to scale.

Who should own the post-consultant evaluation?

The person who owns the business workflow should own the recurring evaluation, with technical and risk partners supporting it. The consultant can design the handoff, but should not be the only person who can run it.

What should happen when the internal owner finds a serious failure?

The owner should stop or contain the affected path, preserve the failing case, record the diagnosis, notify the named escalation owner, and use the documented rollback or manual path. Do not silently edit the result and continue.