Field note · commercial
Evidence an AI Consultant Should Show Before a Platform Recommendation
A buyer-ready packet should expose the test, rejected alternative, errors, data flow, contract unknowns, and veto condition behind any AI platform recommendation.

The quickest way to test an AI consultant's recommendation is to ask what they rejected. A platform can look excellent in a demo and still fail when the team measures correction time, data access, contract exit, or the actual definition of done.
When I taught product managers who went from writing specs to building and shipping the product, and automating work around it, the failure was almost never the model. It was that nobody could say what done meant. This is a bounded teaching observation, not a measured rate. It matters here because a consultant cannot show meaningful evidence until the buyer has agreed what a successful result is. Marius Manolachi's Learn AI page
Quick answer: Ask for a recommendation-evidence packet, not a demo. It should contain the business problem and success threshold, a status-quo alternative, the fixed task set, raw outputs and correction notes, data-flow and security evidence, lifecycle and monitoring requirements, total-cost and contract assumptions, an uncertainty log, and explicit stop conditions. If those pieces are missing, approve only a bounded evaluation, not the platform.
The packet below is the sourceable artifact on this page. It gives another buyer something concrete to request, score, and reuse.
What should be inside the recommendation packet?
The packet should make the recommendation falsifiable. A reader should be able to see what had to be true, what was tested, what failed, what remains unknown, and what would stop the purchase.
The structure follows the strongest recurring requirements in the GOV.UK AI procurement guidance, the Australian Government's proof-of-concept procurement guidance, and the NIST AI Procurement in a Box workbook.
| Packet section | What the consultant should show | What counts as evidence |
|---|---|---|
| Problem and threshold | The job, owner, baseline, and definition of done | A one-paragraph problem statement and measurable pass rule |
| Alternatives | The status quo, a non-AI option, and each plausible platform | Rejected options with reasons, not a straw man |
| Fixed test | Same representative task set for every option | Input IDs, configuration, version, and run date |
| Outputs and errors | Raw results, corrections, abstentions, and failures | Exported outputs or inspectable links with reviewer notes |
| Data flow | Where data starts, moves, is stored, and is deleted | A simple flow diagram plus permissions and retention assumptions |
| Security and governance | Privacy, access, audit, human review, and accountability | Control evidence, open questions, owner, and review status |
| Lifecycle | Monitoring, incident response, rollback, retraining, and retirement | Runbook, alerts, review cadence, and decommission plan |
| Economics | License, implementation, integration, review, support, and exit costs | Dated assumptions and a sensitivity range |
| Contract and lock-in | Export, reuse, IP, portability, service levels, and termination | Contract answers or explicit unresolved questions |
| Uncertainty and veto | What is not known and what would stop the deal | Named owner, due date, and stop condition |
The UK guidance explicitly says to remain open to alternative solutions, assess integration and governance, avoid black-box and lock-in risks, and consider ongoing evaluation and support. The Australian guidance adds short contracts, exit points, transparent work-package costs, open APIs, exportable artifacts, and knowledge transfer. Those are procurement requirements, not optional polish.
How do you test the recommendation fairly?
Use one fixed task set and one scoring rule. Do not let each vendor choose the example that makes its platform look best.
Start with a buyer-owned control. That might be a person using the current workflow, a spreadsheet with deterministic rules, or an existing search and approval process. The control is not there to prove that AI loses. It shows whether the platform creates enough net value to justify its added risk and operating work.
For a small internal knowledge workflow, a usable fixed set might be:
- Summarize a meeting record into decisions, open questions, owners, and source pointers.
- Draft a follow-up message from the same record without adding commitments.
- Search an authorized document corpus and return three supported findings with source pointers.
Run every task against the same representative cases. Include ordinary cases, ambiguous cases, missing-context cases, and at least one case where the correct response is to abstain or ask for clarification. Preserve the input ID, platform and version, prompt or configuration, raw output, reviewer corrections, elapsed operator time, and final disposition.
Here is a compact scoring rule you can put in the packet:
| Case score | Meaning |
|---|---|
| 2 | All required fields are present, sources are inspectable, and no material correction is needed |
| 1 | The output is useful but needs a material correction or extra review |
| 0 | The output fails, invents support, cannot be inspected, or leaves no usable handoff |
Set the threshold before seeing the winner. For the worked example below, the illustrative buyer passes a candidate at 20 of 24 points, with zero fabricated citations and no loss of net time savings after review. Those numbers are a buyer-defined decision rule, not a benchmark or a claim about a vendor.
NIST's procurement workbook asks for test reports, logs, quality criteria, drift checks, periodic testing, interoperability, APIs, dependencies, and business continuity. The NIST AI RMF likewise calls for documented test sets and metrics, deployment-like conditions, known limitations, production monitoring, incident response, and the ability to deactivate or decommission a system. A slide saying “98% accurate” without the cases and error taxonomy does not meet that bar.

Worked example: a public platform trial that still ends in a veto
The following is a bounded application of a public evaluation. It is not a test I ran and it is not a recommendation that every organization should buy Microsoft 365 Copilot.
The Australian Government's published evaluation of its Microsoft 365 Copilot trial gives us something useful: an inspectable result set, a described method, and clear gaps. It reports that 69% of survey respondents agreed Copilot improved task speed and 61% agreed it improved quality. Up to 7% reported that it added time. The report also says that verification and editing sometimes cancelled out efficiency gains. Read the evaluation findings
That is stronger than a demo. It is still not enough to approve a new buyer's platform commitment.
The decision packet
| Field | Worked entry |
|---|---|
| Business problem | Reduce low-risk preparation work for internal meetings and documents while keeping a human responsible for final communication and decisions |
| Success threshold | At least 20 of 24 points on the fixed task set, zero fabricated citations, and net time savings after human correction |
| Control | The existing human workflow on the same source material |
| Candidate platform | Microsoft 365 Copilot, for work already held in Microsoft Teams, Word, and related Microsoft 365 systems |
| Fixed tasks | Summarization, first draft, and authorized information search, each repeated over 12 representative cases |
| Raw outputs | The public report exposes survey results and task categories, but not a buyer-ready 24-case output set |
| Quality notes | Quality and speed were positive for many respondents, but verification, editing, poor Excel functionality, and Outlook access issues reduced the benefit |
| Integration assumptions | Teams and Word are in scope; non-Microsoft 365 integrations require a separate test because they were out of scope in the trial |
| Security and governance | The public report records data-security, accountability, disclosure, and Freedom of Information concerns; buyer-specific control evidence is still required |
| Lifecycle | Rolling releases require adaptive planning, plus monitoring, incident response, rollback, and decommissioning evidence |
| Cost assumptions | License, setup, integration, training, review, support, and exit costs are not established by the public trial |
| Contract questions | Can the buyer export data, configurations, and work products? What happens at termination? What service levels, audit rights, and reuse rights apply? |
| Uncertainty log | No matched status-quo baseline, no raw task-level outputs, no buyer security review, no dated quote, no contract, and no second-platform comparison |
The public evaluation used document and data review, agency evaluations, interviews, focus groups, and surveys. It included more than 2,000 trial participants from more than 50 agencies, with 831 post-use survey respondents. That gives the worked packet a real public evidence base. It does not turn perceived survey improvement into a controlled result for your workflow.
The report also notes that non-Microsoft 365 integrations were out of scope, that poor data security and information management practices could be magnified, and that participants raised concerns about vendor lock-in. These are exactly the details a platform recommendation should carry forward instead of hiding behind adoption enthusiasm.
The decision
Conditional recommendation: run a short evaluation only if the buyer already works in Microsoft 365, can provide representative cases, and can keep the existing human workflow as the control. Require the consultant to produce the 24-case outputs, correction log, data-flow review, security sign-off, total-cost model, and contract answers before a platform commitment.
Veto: do not name Microsoft 365 Copilot, or any other platform, as the winner when the consultant cannot show same-task results against the control, buyer-specific security and data-flow evidence, a cost model that includes review and support, and an exit path. If a second viable platform exists, it must run the same task set before the word “winner” appears in the recommendation.
This is the practical difference between evidence and a positive case study. The public result justifies a test. The missing evidence blocks a purchase.
Which evidence should a consultant show about risk and operations?
Risk evidence should be specific to the workflow, not a generic security slide.
Ask the consultant to draw the data flow from source to output. Identify which records are sent to the platform, which permissions are used, where logs are stored, who can inspect prompts and outputs, how long data is retained, and how a user deletes or exports it. Then name the human who reviews a result and the action that happens when the system abstains, fails, or produces an unsafe answer.
Ask for evidence in five operational areas:
- Privacy and security: data classification, access control, retention, encryption, residency, subprocessors, incident notification, and a buyer-owned approval.
- Governance: use-case owner, approval authority, audit trail, appeal or override path, and a record of known limitations.
- Monitoring: quality metrics, correction rate, drift or distribution checks, latency, cost, availability, and alert thresholds.
- Recovery: fallback to the control process, rollback, incident handling, and a tested way to suspend the platform.
- Retirement: export, deletion, handover, replacement, and decommissioning steps.
The Australian proof-of-concept guidance calls for security and privacy conditions, explainability, short contracts, cost transparency, open standards, vendor transition, performance measurement, and exit deliverables. The NIST AI RMF makes the same lifecycle shape explicit: govern, measure, manage, monitor, respond, and deactivate when outcomes no longer match intended use.
What should make you reject the recommendation?
A good consultant should be willing to write the stop conditions before the test begins. Treat these as vetoes, not risks to be waved away in a later phase.
- No definition of done. If the team cannot say what a passing output contains and what remains human-owned, there is no meaningful test.
- No control. If the current workflow was not measured on the same cases, “better” has no stable meaning.
- No raw outputs. A chart without cases, corrections, and configuration cannot reveal failure modes.
- No buyer-specific data-flow review. A vendor's general security page is not proof that your data, permissions, and retention path are acceptable.
- No net-cost view. License price is not total cost. Include setup, integration, training, review, support, monitoring, and exit.
- No exit. If data, prompts, configurations, logs, or work products cannot be exported or reused, lock-in is a decision variable, not a footnote.
- No lifecycle owner. Someone must own monitoring, incident response, rollback, and retirement after the consultant leaves.
- Review removes the benefit. If correction and verification consume the claimed time saving, keep the control or redesign the task.
The exception is a low-risk discovery engagement. A consultant may recommend spending a small, bounded amount to answer these questions. That recommendation should be written as “run the evaluation under these gates,” not “buy this platform.”
Copyable recommendation-evidence packet
Save this block and ask the consultant to complete it before you accept a platform recommendation.
USE CASE AND OWNER
Business problem:
Decision owner:
Human-owned decision or action:
SUCCESS THRESHOLD
Baseline/control:
Fixed task set and representative data:
Required quality and error limits:
Required operator effort and latency:
Stop condition:
OPTIONS
Status quo or non-AI alternative:
Platform option 1:
Platform option 2, if viable:
Rejected options and reasons:
TEST EVIDENCE
Run date and platform/version:
Configuration or prompt:
Input IDs:
Raw outputs or inspectable links:
Reviewer corrections:
Scoring rule and result:
Failures, abstentions, and edge cases:
DATA, SECURITY, AND GOVERNANCE
Source, destination, permissions, retention, and deletion:
Security and privacy evidence:
Use-case owner, reviewer, audit, override, and incident path:
LIFECYCLE AND ECONOMICS
Monitoring metrics and alert thresholds:
Fallback, rollback, and decommission plan:
License, implementation, integration, review, support, and exit cost:
CONTRACT AND LOCK-IN
Data and configuration export:
IP and reuse rights:
Service levels and audit rights:
Termination, transition, and replacement:
UNCERTAINTY LOG
Unknown:
Why it matters:
Owner and due date:
DECISION
Conditional recommendation:
Veto conditions:
Evidence still required before commitment:
If the consultant fills this packet with “vendor says,” “works in the demo,” or “we will measure later,” the recommendation is not ready. Ask for the cases, the rejected alternative, the unknowns, and the condition that would make the consultant change their mind.
If you are mapping the broader choice between consulting, tutoring, and other paths, start with AI consulting and tutoring decisions. If a pilot is already being proposed, use how to buy an AI pilot without platform lock-in as the next contract conversation. For help turning a real team decision into a testable packet, see Marius Manolachi's AI consulting and tutoring work.