Field note · commercial
What Should a Buyer Measure After an AI Consultant Leaves?
Use a 30/60/90 scorecard and an owner-run failure drill to decide whether an AI workflow is ready to continue after the consultant leaves.

The handoff meeting is a poor finish line. A folder of prompts, diagrams, and recordings can still leave the buyer dependent on the person who built them.
I use a harder test: can the internal owner run the evaluation, explain one failure, make the next decision, and contain the workflow when it misbehaves? That is what the scorecard below is designed to show.
Observed result from the bounded sanitized run: the first evaluation passed 7 of 8 rows. One critical classification failure was corrected on rerun. The rollback drill routed all eight rows to manual review with zero unreviewed writes. The decision was REVISE, because the exercise had no production cost or latency telemetry.

What should a buyer measure first?
Start with one workflow and one evidence pack. Record the business baseline, current outcome, adoption, operational change, technical health, failures, cost, latency, owner, evaluation method, review cadence, and next decision in the same place.
This keeps activity separate from value. McKinsey's five-layer framework makes the same separation across technical performance, user adoption and engagement, operational KPIs, strategic outcomes, and financial impact. It also attaches an owner to each layer. Read the published framework.
The NIST AI Risk Management Framework gives the risk-management spine: Govern, Map, Measure, and Manage. A buyer can translate that into a simple question for the exit pack:
| Question | Scorecard field | Proof to keep |
|---|---|---|
| What was the work like before the engagement? | Original business baseline and baseline definition | Dated process measure, sample, or explicit statement of what was not measured |
| What is different now? | Current operational outcome | Before and after result, with the comparison unit named |
| Is the workflow actually used? | Adoption or workflow penetration | Eligible tasks, completed tasks, user or role breakdown |
| Does it fail safely? | Quality and safety failures | Evaluation rows, incident log, severity, containment, and escalation |
| Can it run economically? | Cost and latency | Cost per run or case, p50 or p95 latency, and measurement method |
| Who can operate it? | Named internal owner | Person or role, access, runbook, and decision authority |
| Can the result be reproduced? | Evaluation set and rubric | Inputs, expected outputs, scoring logic, and raw rows |
| What happens next? | Decision cadence and next decision | Review date, threshold, and go, revise, or stop outcome |
If a field is unknown, write unknown and assign an owner. An empty cell is evidence about the handoff. It is not a reason to quietly delete the metric.
Use this blank post-engagement scorecard
Copy this template for one workflow. Keep the scope small enough that the internal owner can run it without asking the consultant how the system works.
| Field | Fill in | Minimum evidence |
|---|---|---|
| Workflow and boundary | Trigger, input, output, excluded cases | One-sentence workflow contract |
| Review date | Day 0, day 30, day 60, or day 90 | Calendar owner and decision forum |
| Original business baseline | What happened before, measured how, during which dates | Raw sample or dated business measure |
| Current operational outcome | Cycle time, rework, throughput, accuracy, or another workflow result | Same unit and comparison period as baseline |
| Adoption or penetration | Eligible work, AI-assisted work, acceptance, override, and manual bypass | Event or queue counts by role |
| Quality and safety | Pass rate, critical errors, unsafe outputs, incidents, and escalations | Eval rows plus failure log |
| Cost and latency | Cost per case, total run cost, p50/p95 latency, retries | Instrumentation source and unit |
| Internal owner | Name, role, permissions, and authority to stop or escalate | Owner acknowledgment and escalation contact |
| Evaluation set and rubric | Cases, expected result, grader, threshold, and exception rule | Versioned cases and raw scores |
| Rollback or escalation | Manual path, disable switch, data protection, and notification path | Dry-run record or incident test |
| Cadence | Who reviews which layer and when | Calendar invite, ticket, or review log |
| Next decision | Go, revise, or stop, with the reason and due date | Signed decision record |
The baseline needs a definition, not just a number. “Time saved” is incomplete unless you say whose time, on which unit of work, over what period, and whether review or rework is included. “Adoption” is incomplete unless you define the eligible workflow population.
The UK Government's AI procurement guidance is useful here even when the buyer is not a government department. It separates procurement from ongoing management and calls for lifecycle oversight, ongoing evaluation, knowledge transfer, support, auditability, and clear end-of-contract roles. See the procurement and ongoing-management guidance. Treat those points as buyer requirements to adapt, not as a private-sector law.
What does a worked scorecard look like?
Here is a sanitized example for a product-feedback routing workflow. The input is a short feedback note. The output is a category, severity, and next action. The rows are authored fixtures, so the example demonstrates the measurement and handoff procedure rather than claiming a customer result.
| Field | Worked entry | Status |
|---|---|---|
| Workflow | Route product feedback into category, severity, and next action | Defined |
| Original business baseline | Before handoff, there was no named owner, evaluation set, rubric, rollback drill, or review cadence | Recorded baseline |
| Current operational outcome | First run: 7 of 8 rows passed. The failed row passed after the rule change and rerun | Measured fixture result |
| Adoption or workflow penetration | One internal-owner runner completed all eight rows. Production penetration was not measured | Exercise result and explicit unknown |
| Quality and safety failures | One critical false feature classification. The stop action still prevented an unreviewed write | Failure and containment recorded |
| Cost and latency | Local exercise cost: €0.00. Production API cost and latency were not captured | Instrumentation gap |
| Named internal owner | Operations lead runs the evaluation. Product owner receives critical escalations | Assigned role |
| Evaluation set and rubric | Eight labeled rows; exact class, severity, action, critical stop rule, and no-write rule | Versioned fixture |
| Decision cadence | Day 0 exit pack, day 30 owner-run eval, day 60 adoption and operational review, day 90 value and cost review | Scheduled proposal |
| Next decision | REVISE: add cost and latency telemetry, expand edge cases, then rerun before scale | Explicit decision |
This is the distinction I want buyers to preserve: the workflow can be operationally promising while the evidence pack is still incomplete. A polished demo does not turn missing cost telemetry into a zero, and a successful rerun does not prove long-term adoption.
How do you run the independent operator test?
Give the internal owner the blank scorecard, the evaluation set, the current run instructions, and the escalation path. Then step away. The owner must complete five actions.
- Reproduce the run. Execute every case and save the raw output, score, timestamp, and workflow version.
- Investigate one failure. Start from the input and trace the output to the rule, prompt, retrieval result, tool response, or human decision that produced it. Preserve the failing row.
- Record the change decision. Choose
GO,REVISE, orSTOP. State the threshold, the observed evidence, the change owner, and the next review date. - Contain the workflow. Switch to the documented manual path or disable the affected action. Confirm that no unreviewed write or external action occurred during the drill.
- Escalate when required. Send the preserved failure to the named technical, product, security, or compliance owner. Record who accepted it and what happens next.
The test passes only if the owner can perform all five actions without a consultant explaining hidden steps. A consultant can observe the exercise, but the buyer should not count a consultant-led rescue as transfer.
My own teaching observation points to the same boundary. I taught product managers who went from writing specs to building and shipping the product, and automating work around it. The recurring failure is often not that the model is incapable. It is that nobody can say what “done” means. That is why the scorecard makes the threshold, failure response, and next decision explicit. This is the bounded observation recorded in my AI consulting and tutoring work, not a claim about a measured failure rate.
What did the operator test find?
The eight raw rows below make the result inspectable.
| Case | Sanitized input | Expected | First run | Result |
|---|---|---|---|---|
| 01 | Export loses timezone on some records | bug, high, escalate | bug, high, escalate | pass |
| 02 | Request for a dark-mode setting | feature, low, backlog | feature, low, backlog | pass |
| 03 | Invoice total becomes 0 after import | bug, critical, stop | feature, critical, stop | fail |
| 04 | User asks where to change timezone | how-to, low, answer | how-to, low, answer | pass |
| 05 | Renewal reminders were sent twice | bug, high, stop | bug, high, stop | pass |
| 06 | Request for a weekly unpaid-invoice view | feature, medium, backlog | feature, medium, backlog | pass |
| 07 | CSV import rejects comma decimals | bug, high, escalate | bug, high, escalate | pass |
| 08 | User asks how to change notification settings | how-to, low, answer | how-to, low, answer | pass |
The owner investigated case 03. The note described a financial-data corruption symptom after import. The first output treated it as a feature request, even though the severity and stop action were present. The owner changed the rule to classify a numeric corruption symptom in an import or export path as a bug requiring stop or escalation. The rerun classified case 03 as bug, critical, stop.
The owner then ran the rollback drill. The workflow switched to manual-review, all eight rows entered the manual queue, the critical case was escalated to the product owner, and the record showed zero unreviewed writes. The drill passed as a control-path test.
The overall decision remained REVISE, not GO. The procedure proved that an operator could reproduce, diagnose, change, and contain the workflow. It did not prove production economics, because the exercise had no billed API call and no latency instrumentation. It did not prove adoption, because no live eligible-task denominator existed. Those are decision-relevant gaps, not editorial footnotes.
How should the 30/60/90 review cadence work?
Use the cadence to move from handoff proof to business proof. NIST's 2026 report groups deployed-AI monitoring into functionality, operational, human factors, security, compliance, and large-scale impacts. It also identifies open questions about who, what, when, why, and how to monitor, including the right cadence and the balance between automated and human-validated monitoring. Read the NIST monitoring report.
| Checkpoint | Buyer question | Evidence | Decision |
|---|---|---|---|
| Day 0 | Can the internal owner operate and contain the workflow? | Scorecard, eval set, owner test, failure and rollback record | Do not close the engagement until the owner can run it |
| Day 30 | Does it still work on current inputs, and are failures handled? | Repeated eval, incident review, manual bypass count, owner notes | Go or revise the workflow and controls |
| Day 60 | Is it used in eligible work, and is the process improving? | Penetration, acceptance and override, cycle time, rework, quality | Continue only if adoption and operational measures move together |
| Day 90 | Does the business case justify continued investment? | Strategic outcome, total cost, support load, risk and compliance record | Scale, redesign, renew support, or stop |
This cadence is a recommendation, not a universal standard. A high-risk workflow may need daily or per-release checks. A low-risk internal drafting workflow may use a lighter schedule. The owner, risk boundary, and failure cost should determine the frequency.
When should the buyer go, revise, or stop?
Make the decision rule visible before the consultant leaves.
| Decision | Use it when | Required next action |
|---|---|---|
| GO | The owner reproduces the eval, critical failures are within threshold, the manual path works, and cost, latency, adoption, and outcome evidence are present | Continue at the defined scope and cadence |
| REVISE | The workflow is useful but a control, measurement layer, edge case, or owner path is incomplete | Name the change, owner, threshold, and next review date |
| STOP | The workflow creates unacceptable harm, cannot be contained, misses a critical threshold, or has no accountable owner | Disable or contain it, preserve evidence, and decide whether to redesign or retire |
The buyer should not let a consultant's departure date create a false deadline. If the internal owner cannot complete the operator test, the correct result is REVISE or STOP, even if the demo looked good.
For the wider commercial decision, see Choose Between an AI Consultant, Agency, or Internal Team. For a broader question about evidence that an AI partner made a team more capable, read What Evidence Proves an AI Partner Made the Team More Capable?. Those pages provide adjacent context. This page is the post-exit scorecard and test.
If you are buying an engagement, put the blank scorecard and operator test into the statement of work before the first build session. If you have already reached handoff, run it now. The result should change what you do next.
Continue with a related field note
Questions people ask next
Is usage enough to prove an AI consultant created value?
No. Usage shows that people interacted with the workflow. Pair it with an operational outcome, quality and safety checks, cost and latency, and an owner-run evaluation before deciding to scale.
Who should own the post-consultant evaluation?
The person who owns the business workflow should own the recurring evaluation, with technical and risk partners supporting it. The consultant can design the handoff, but should not be the only person who can run it.
What should happen when the internal owner finds a serious failure?
The owner should stop or contain the affected path, preserve the failing case, record the diagnosis, notify the named escalation owner, and use the documented rollback or manual path. Do not silently edit the result and continue.