Field note · commercial
How Should a Buyer Release Payment for an AI Experiment?
A buyer-side worksheet and worked payment table for tying AI experiment payments to accepted evidence, capped discovery, and a clear next decision.

I’ve taught product managers to ship instead of only writing specs. The failure was almost never the model. It was that nobody could say what done meant. An AI experiment proposal can hide the same problem behind a kickoff date and a fixed price.
Quick answer: Release payment for an AI experiment only when the supplier turns the work into a bounded, phase-gated learning decision before it starts. Require a baseline, representative cases, evaluation method, acceptance threshold, safety and data gates, spend cap, named reviewer, payment trigger, and next decision. If those are unresolved but worth learning, fund capped discovery. Reject ownerless or unsafe work.
The buyer-side result
The following is a synthetic proposal, created for this article. It is not a client result or a measured pilot.
| Proposal as received | Buyer-side finding |
|---|---|
| Four-week AI support-escalation pilot | The use case may be worth investigating. |
| £12,000 fixed price, 50% at kickoff and 50% at final demo | The payment schedule is clear, but payment is not tied to evidence. |
| Claimed benefit: faster, more consistent escalation packets | Value is stated, but no baseline or acceptance threshold is stated. |
| 100 historical tickets plus 20 shadow-live cases | The sample is named, but its representative case mix is not. |
| “Review results with the team” as the exit language | No observable stopping rule, named reviewer, safety gate, or next decision. |
| Decision | Revise to paid discovery. Do not sign the full pilot. |
That is the sourceable artifact on this page: payment release depends on an accepted evidence packet, not a kickoff or a polished demo. A valuable question can still earn a small, capped discovery phase, but the full experiment payment stays locked until the buyer can make the next decision.
What the payment worksheet must make observable
Copy this table into the proposal review. A field is not complete because someone mentioned it on a call. It is complete when the buyer can point to the evidence, the owner, the cap, and the next payment decision.
| Field | What the proposal must state |
|---|---|
| Learning question | One question the phase is paid to answer, not a promise to “explore AI.” |
| Baseline | The current process, metric, or decision record that the comparison uses. |
| Representative cases | The case types and counts, including rare, incomplete, or high-risk cases. |
| Evaluation method | How outputs will be compared, who will review them, and what evidence is retained. |
| Acceptance threshold | The result that counts as passing, expressed before the phase starts. |
| Mandatory safety or data gates | Data allowed, data prohibited, permissions, human review, and side effects that remain disabled. |
| Maximum spend and exposure | Cash cap plus a time, data, operational, or reputational exposure cap. |
| Named reviewer | The person who can accept, reject, or request a defined revision. |
| Payment trigger | The evidence packet or deliverable that releases the next payment. |
| Sign-revise-stop decision | What happens if the threshold passes, partly passes, or fails. |
This structure follows a useful pattern in the UK government AI procurement guidance: describe the problem and required performance in output-based terms, use proof-of-concepts when they help test the wider requirement, and carry important governance requirements into the contract. The guidance also points buyers toward data assessment and go/no-go points before releasing the next commitment.
The GSA Buy AI guidance makes the same buying moment concrete: use a testbed, sandbox, or pilot before a large purchase, start with a small user group, measure results, understand data flow and storage, and set usage limits. Those are not decorations around a pilot. They define its boundaries.
Worked example: two paid phases with evidence-based payment
The example below keeps the assumptions visible. The names, prices, thresholds, and arithmetic are synthetic.
| Phase | Learning question and baseline | Cases and evaluation | Acceptance, gates, cap, reviewer, payment | Sign, revise, or stop |
|---|---|---|---|---|
| Paid discovery, £2,500 | Can the team define the data boundary, baseline, and evaluation set for escalation drafting? Baseline must be captured on 30 redacted historical cases, including current drafting time and correction time. | 30 cases: 10 routine, 10 incomplete, 10 policy or escalation cases. Two buyer reviewers label expected disposition, evidence source, risk class, and owner. | Accept when all 30 cases have those fields, the baseline is recorded, and the data and safety boundary is approved. Use redacted data only, a read-only sandbox, and no external writes. Synthetic reviewer: Priya Shah, Head of Support. Pay on an accepted evidence pack. Cash cap £2,500. | Sign Phase 2 only if the pack is accepted. Revise the discovery scope if one field is missing but the reviewer can specify the correction. Stop if the supplier cannot produce a usable case set or the data gate fails. |
| Bounded pilot, £5,500 | Does the read-only assistant reduce total operator time without hiding high-risk cases? Synthetic baseline assumption: 18 minutes to draft plus 7 minutes to correct, or 25 minutes per case. | Run the same 30-case set in a sandbox. Compare the manual route and assistant route. Record draft time, correction time, evidence links, escalation decisions, and reviewer overrides. NIST ARIA describes a layered approach using model testing, red teaming, field testing, and measurement trees, which is a useful pattern for separating these checks. | Pass at a 15% reduction in median total operator time, which means 25 × 0.85 = 21.25 minutes or less per case, with zero unauthorized writes, zero unreviewed high-risk dispositions, and evidence links on every accepted output. Keep real customer data out until the data review passes. Synthetic reviewer: Priya Shah. Pay on the signed comparison pack, not the demo. Cash cap £5,500. | Sign the optional build only if every threshold and gate passes. Revise to paid discovery if the result misses the threshold for a defined, recoverable reason. Stop if a safety or data gate fails, the baseline cannot be reproduced, or the reviewer cannot accept the evidence. |
The external phase cap is £2,500 + £5,500 = £8,000. Add a synthetic internal review cap of 24 hours × £75/hour = £1,800. Maximum pre-build exposure is therefore £9,800. The optional build remains a separate decision with its own worksheet.
The phrase “synthetic baseline assumption” matters. The 25-minute baseline is not a result. It only makes the arithmetic reproducible so a real buyer can replace it with a measured baseline before releasing Phase 2 payment.
Why payment needs a stop condition
Inference from the procurement and evaluation evidence: if the supplier cannot make the stop condition observable before work starts, the buyer is funding activity rather than a reversible learning decision.
That inference follows from three separate requirements. Output-based procurement needs a defined problem and required performance. A useful pilot needs a small group and measured results. A serious evaluation needs more than a polished model output, so it separates model behavior, adversarial testing, field use, and measurement. Without a threshold and a reviewer, none of those can release or stop payment.
The NIST ARIA pilot evaluation report describes model testing, red teaming, and field testing, then discusses dialogue annotation, tester questionnaires, and measurement trees. You do not need to copy NIST's exact study design for a small company. You do need to decide which layer answers your question and what evidence would make you stop.
How to use eROI without pretending the probability is known
The paper AI Strategy: How to Choose What AI Product to Implement separates Value if Successful, Likelihood of Success, and Investment Required. That is a better starting point than arguing about a single ROI percentage while the pilot definition is still moving.
For the synthetic proposal:
- Value if Successful: £60,000 of annual capacity, an assumption.
- Likelihood of Success: unknown because no baseline, threshold, or case mix exists yet.
- Investment Required: £8,000 of capped external discovery and pilot phases, before internal review time.
Do not turn “unknown” into a convenient 70%. Ask whether discovery can make the likelihood more knowable at a bounded cost. If a later evidence packet justified a 0.5 scenario, the illustrative expected surplus would be £60,000 × 0.5 - £8,000 = £22,000 before internal review time. That arithmetic is a what-if branch, not a forecast or a reason to sign the original proposal.
If the supplier cannot describe what would change the likelihood estimate, the eROI column is not helping you decide. It is giving uncertainty a neat number.
When paying for discovery is the right exception
Pay for discovery when the unknown itself is the valuable question and the discovery phase has a deliverable, a cap, and an exit. Do not pay for a full pilot merely because the supplier needs time to discover what the pilot is.
A paid discovery phase can be reasonable when:
- The value if successful is large enough to justify learning.
- The buyer can provide a safe, bounded data set or sandbox.
- A named reviewer can accept the discovery artifact.
- The phase will produce a baseline, representative case set, evaluation method, and next decision.
- The buyer can stop without authorizing production integration or a larger build.
The principal exception is a pilot where the learning question is genuinely exploratory and cannot be reduced to one numeric threshold at the start. Even there, the buyer can set a deliverable threshold: a defined case taxonomy, a documented uncertainty range, a safety review, and a recommendation to sign, revise, or stop. Uncertainty can remain. Unobservable work cannot.
What the contract must say before money moves
Put the worksheet into the statement of work and payment schedule. The EU model contractual AI clauses are explicit that they are not a complete contract and do not supply terms for acceptance, payment, delivery times, applicable law, or liability. The clauses must be customized to the context, so do not treat an AI schedule as a substitute for those commercial terms. See the EU model contractual AI clauses for the warning.
At minimum, attach:
- the accepted case list and baseline version;
- the evaluation method and reviewer names;
- the pass threshold and treatment of borderline cases;
- the data, security, permission, and human-review gates;
- the maximum cash and exposure cap;
- the exact evidence packet that triggers payment;
- the sign, revise, or stop authority at the end of each phase;
- the rule that production access or an optional build requires a new approval.
This also keeps the buyer aligned with the parent guide on AI services buying decisions. If the proposal itself is hard to compare before this step, use the AI consulting proposal comparison guide. If the question is whether the commercial case is strong enough at all, see what commercial evidence supports an AI pilot investment.
The decision to use on the next proposal
Ask the supplier to complete the worksheet before you approve the pilot. Then apply this table:
| Proposal state | Decision |
|---|---|
| Every field is observable, the safety and data gates pass, and the payment trigger is evidence-based | Sign the bounded phase. |
| The value is credible, but the baseline, case mix, threshold, or reviewer still needs defined work | Revise to paid discovery. Cap it and specify its deliverable. |
| The supplier refuses a threshold, hides the data boundary, or leaves acceptance ownerless | Reject. The buyer would be paying for motion without a reversible decision. |
If you are a founder or operator reviewing a live proposal, this worksheet is the useful next step. If you want a buyer-side conversation about turning an AI proposal into a capability your team can own, Marius Manolachi's AI consulting and AI tutoring work is designed around making existing people capable of building AI products on their own work.