Field note · commercial
What Should Procurement Ask When AI Tutoring Is Hard to Quantify?
A buyer-ready scorecard for turning a qualitative AI tutoring goal into proxies, pilot evidence, privacy checks, and a stop or renew rule.

An AI tutoring demo can show a fluent explanation. It cannot, by itself, show that a learner will solve a new problem later.
That is the procurement problem. The buyer wants a qualitative change such as independence, confidence, or better reasoning, while the supplier naturally presents usage, satisfaction, or a test result from another setting.
For the wider buying context, see the AI commercial decisions cluster. This page gives you the narrower artifact to attach to an RFP or pilot review.
Quick answer: Procurement should ask the vendor to translate the qualitative goal into an observable learner behavior, at least three proxies, a baseline or safe comparison, and a reversible pilot with privacy and equity checks. Renew only if the delayed independent task improves against the agreed comparator. If no proxy or safe comparison exists, redesign the pilot or do not buy.
Start with the learner change, not the tutor feature
Write the requirement as a change in learner behavior, then ask the supplier how the product could help produce and evidence that change. Do not start with “adaptive hints,” “personalisation,” or another feature label.
The UK government’s AI procurement guidance recommends a clear problem statement instead of detailed solution specifications, with output-based requirements grounded in user needs and required performance. SREB similarly asks education buyers whether a product supports instructional goals and student learning outcomes. (UK AI procurement guidance, SREB AI procurement questions)
Here is the translation I would put in the requirement:
Qualitative goal: Learners become more independent at introductory statistics.
Procureable learner change: Given an unfamiliar statistics problem, a learner chooses a defensible first step and explains the reasoning without a tutor prompt.
When I taught product managers to ship instead of only writing specifications, the failure was almost never the model. It was that nobody could say what “done” meant. The same failure appears here when “independence” is left as a feeling instead of an observable behavior. This is a teaching judgment from Marius Manolachi, not a measured tutoring result.
Worked scorecard for an AI tutoring pilot
The following is the sourceable artifact on this page. It is analysis for a hypothetical first-year statistics use case, not a report of a completed institutional pilot.
| Scorecard field | Buyer specification | Evidence, owner, and decision check |
|---|---|---|
| Intended learner change | On an unfamiliar problem, the learner selects a sound first step and explains why without tutor prompts. | Analysis: Use this as the output-based requirement. The course or assessment lead owns the rubric. This follows the UK problem-statement guidance and SREB’s instructional-goal question. |
| Observable proxies | 1. Blind rubric score for the explanation, scored 0 to 4. 2. Correct first-step selection without a hint. 3. Delayed transfer-task completion two weeks after the last supported session. 4. Tutor prompts per correct solution. | Analysis: Record all four. Do not use session enjoyment or number of turns as the primary outcome. The delayed task is the principal renewal proxy because it tests independence after support is removed. |
| Baseline and comparison | Give every participant a pre-pilot task with the same rubric. Compare AI-tutor participants with business-as-usual support or a delayed-access group that still receives the institution’s required instruction. | Analysis: Freeze the comparison before launch. If the institution cannot define a safe comparator, return pilot redesign or no-buy. A simple pre/post increase without a comparison cannot separate tutoring from practice, maturation, or novelty. SREB asks whether buyers have conducted pilots; UK guidance calls for iterative evaluation and decision points. |
| Evidence the supplier must provide | Study protocol and report; learner level; intervention duration and dose; task type; comparator; outcome definition; subgroup results; attrition; limitations; raw or aggregate pilot export; data dictionary; model-change record; accessibility and safeguarding evidence. | Vendor claim: Any statement that the tutor improves independence remains a vendor claim until the supplier shows conditions that match this use case. SREB asks vendors to disclose capabilities, limitations, data handling, bias mitigation, ownership, retention, and deletion. The tutoring review shows why duration, control, learner level, and outcome must be compared. |
| Privacy and equity checks | Minimize learner data. Record retention, deletion, access, and secondary-use terms. Test accessibility modes. Review participation and outcomes across relevant learner groups where lawful and necessary. Keep a human appeal or escalation route. | Analysis: Name the institutional data protection officer and student support or equity lead before launch. UNESCO frames education AI as a benefit-risk, inclusion, and equity question. The U.S. Department of Education’s 2025 guidance also emphasizes affected stakeholders and privacy in responsible school AI adoption. |
| Pilot instrumentation and owners | Capture a pseudonymous learner key, condition, task version, tutor exposure, hints and prompts, completion, rubric score, time, accessibility mode, and incident or appeal events. | The supplier analytics lead owns the event export and data dictionary. The course lead owns outcome scoring. The data protection officer owns data minimization and retention approval. Procurement owns acceptance. Analysis: Collect only fields needed for the agreed decision. |
| Review cadence | Baseline sign-off before launch; weekly safety and data-quality review; midpoint outcome review; end-pilot decision; quarterly review if renewed; annual compliance and effectiveness review. | SREB asks about feedback systems, annual reviews, and adjustment. UK guidance treats evaluation and support as lifecycle work rather than a one-time award. |
| Commercial rule | No-go or pilot redesign: no observable proxy, no safe comparator, material privacy or equity failure, or missing supplier evidence/export. Renew: safety gates pass, the delayed transfer task improves by at least 10 percentage points versus comparison or the pre-agreed equivalent, and at least two of the other three proxies improve without material subgroup harm. Expand: the renew rule holds for two review cycles, at least 80% of the intended learners complete the agreed minimum dose, and implementation burden remains within the approved budget and owner capacity. | Analysis: These are buyer-side thresholds for a reversible decision, not universal tutoring standards or observed effects. Tie them to the contract and record who can approve an exception. UK guidance explicitly supports impact review and go/no-go points; SREB asks buyers to include exit strategies for underperformance, security breaches, or non-compliance. |
This matrix passes the minimum evidence test: it has one realistic qualitative goal, four measurable proxies, a baseline and comparison plan, named data owners, a supplier evidence request, privacy and equity checks, pilot instrumentation, review cadence, and explicit no-go, renew, and expand thresholds.
Ask for evidence that matches the condition
Ask “under what conditions did this work?” before asking “how large was the improvement?” A study result is useful only when you can compare its learners, dose, task, comparator, and outcome with the proposed pilot.
The 2025 systematic review of AI-driven intelligent tutoring systems in K-12 education covered 26 publications and reports that study designs vary enough that effect sizes are difficult to compare cleanly. It also reports that half of the interventions lasted less than a week, with some lasting only one class period. The review cautions that long-term performance cannot be inferred confidently from such brief interventions. (npj Science of Learning review)
So the RFP should require a vendor evidence table with these columns:
| Ask the supplier to disclose | Why procurement needs it |
|---|---|
| Learner age, level, subject, and prior attainment | A result for one learner group is not automatically a result for yours. |
| Intervention duration, frequency, and minimum dose | A one-class demonstration cannot prove a delayed capability change. |
| Comparison condition | “Improved” is incomplete without knowing what the comparison group did. |
| Task and outcome definition | In-session correctness, confidence, explanation quality, and delayed transfer are different outcomes. |
| Attrition and subgroup results | A headline average can hide who did not participate or benefit. |
| Product version and material changes | A study of one tutor configuration does not automatically describe the version you are buying. |
| Data export and audit evidence | The buyer must be able to reproduce the pilot decision and investigate harm. |
Label the response in your evaluation record. A supplier’s case study is a vendor claim. Your local scorecard result is buyer evidence. Your threshold is buyer analysis. Keeping those categories separate prevents a polished demo from becoming an undocumented acceptance criterion.
Design the pilot before signing the contract
The pilot should be small enough to reverse and specific enough to answer the scorecard. Do not buy a broad platform first and hope the outcome becomes measurable later.
- Freeze the goal and rubric. Publish the learner change, four proxies, minimum dose, comparator, owners, and thresholds in the RFP appendix or pilot order form.
- Collect the baseline. Use a comparable unfamiliar task before tutor access. Store the scoring rubric and task version with the result.
- Choose a safe comparison. Use business-as-usual or delayed access when random assignment is acceptable. If not, use a pre-agreed matched or stepped design and state what it cannot prove. If no safe design is feasible, return pilot redesign or no-buy.
- Instrument support and transfer separately. Log hints, prompts, and supported completion, but score the delayed task without tutor assistance. Otherwise the pilot may measure how well the tutor performs inside the session.
- Review harm and data weekly. The data protection officer and student support or equity lead should be able to pause the pilot for a privacy, accessibility, bias, or safeguarding issue. Procurement should not wait for the end-of-pilot meeting to discover an unsafe condition.
- Make the decision at the agreed checkpoints. Record the evidence, the threshold comparison, unresolved limitations, and the named decision owner. Do not quietly convert a failed pilot into a longer paid trial.
SREB asks buyers to conduct pilots, gather user input, monitor effectiveness and usability, and include exit strategies. UK guidance calls for iterative development, impact assessment, and lifecycle management. Those sources support the procedure. The exact proxies and thresholds above are my analysis for making those principles usable in a tutoring purchase. (SREB AI procurement questions, UK AI procurement guidance)
Pre-agree the commercial decision
The contract should contain three outcomes, not a vague promise to “review the pilot.”
- No-go or pilot redesign: The team cannot observe the intended change, cannot run a safe comparison, cannot obtain the evidence needed to audit the decision, or finds a material privacy, accessibility, equity, or safeguarding failure. Payment should not silently roll into a wider deployment.
- Renew: The safety gates pass, the delayed transfer task meets the agreed threshold, and the supporting proxies show a consistent direction without material subgroup harm. Renew for the bounded use case and keep the quarterly review.
- Expand: The renew condition holds in two review cycles, exposure reaches the pre-agreed minimum dose for at least 80% of intended learners, and the cost and owner workload stay inside the approved envelope. Expansion is a new decision, not an automatic consequence of renewal.
If a buyer cannot fill the proxy or comparison fields, the answer is not a more enthusiastic vendor demo. It is pilot redesign or no-buy. That is the principal exception to the scorecard: some qualitative goals are not ready to purchase because nobody can observe the change safely.
Why a tutoring study does not transfer automatically
Tutoring research can inform the design of a pilot, but it cannot substitute for one. The review evidence includes short interventions, multi-week interventions, different school levels, different controls, and different outcomes. It reports examples where learner level and control condition changed the result, including studies using pre-tests, post-tests, and follow-up tests under different durations. (npj Science of Learning review)
That variation is not a reason to ignore evidence. It is a reason to request the conditions behind the evidence. The buyer should be able to answer five questions before accepting a vendor result:
- Were the learners comparable?
- Was the tutoring dose comparable?
- Was the comparison condition comparable?
- Was the task testing supported performance or independent transfer?
- Was the outcome measured at the same time horizon?
If the answer to several questions is no, record the study as background evidence, not as a forecast.
For a broader pre-procurement rubric, use How to Score an AI Tutor Before Procurement. For help adapting the scorecard to a real learning workflow, Marius Manolachi’s AI consulting and tutoring work is the relevant next step. The decision remains yours: buy only when the goal, observation, comparison, and exit rule are all explicit.