Field note · commercial

How to Score an AI Tutor Before Procurement

A practical scorecard for comparing AI tutors by learner evidence, tutor behavior, operating controls, and commercial fit before procurement.

10 minute read
  • AI tutoring
  • Buying AI services
  • AI evaluation
Illustration of an L&D buyer scoring an AI tutor before procurement

An AI tutor can look excellent in a demonstration and still leave a buyer unable to answer the procurement question that matters: can people do the work after the help is gone?

The scorecard below treats that as one part of a larger purchase decision. It separates learning evidence from tutor engagement, scores whether the tutor helps or simply supplies answers, checks the operating conditions around the product, and applies vetoes before a weighted total.

Use it as a procurement screen. It is not a claim that the score predicts every learner outcome or ranks every tutor.

Start with the capability the purchase must change

Score the capability the buyer needs to see in ordinary work, not the tutor's lesson completion rate.

Write one sentence in this form: “After this intervention, a person can ___, under ___ conditions, with ___ acceptable evidence.” Then create a baseline task that a learner can attempt before using the tutor. The task should produce something inspectable, such as a decision, a short analysis, a working artifact, or a justified recommendation.

This definition is the first procurement control. Without it, a vendor can prove activity while the buyer keeps changing the meaning of success. When I taught product managers who went from writing specs to building and shipping the product, and automating work around it, the failure was almost never the model. It was that nobody could say what done meant. That is a teaching observation, not a measured rate, and it is why the scorecard starts with a visible work product. Marius Manolachi describes this capability-led teaching approach.

The baseline needs four fields:

  • the work situation and input;
  • the output the learner must produce;
  • the reasoning or checks that make the output acceptable;
  • the conditions that must hold without tutor assistance.

Do not use a generic quiz when the purchased capability is workplace performance. A quiz can be useful for prerequisite knowledge, but it cannot stand in for the work product. The exception is a narrow knowledge purchase where recall itself is the agreed outcome. Even then, define the allowed help and the delayed check before the trial begins.

Separate learning evidence from tutor engagement

Give the evidence row its own score. Chat volume, completion, confidence, and immediate correctness are context, not proof that the learner can work independently.

The research supports keeping these outcomes separate. A Frontiers longitudinal analysis distinguishes engagement, near transfer, topic-shift transfer, and delayed performance, and describes broader transfer as conditional on scaffolds for verification, retrieval, and abstraction. The ACL AIME-Con work connects cognitive engagement with next-item performance while leaving distal transfer open for further study. The Frontiers analysis and the ACL paper support the separation, not a universal pass threshold.

Use this evidence block in the trial:

Illustration of a procurement scorecard separating learner evidence from tutor behavior

Evidence rowWhat the buyer recordsScore 0Score 1Score 2
Baseline and independent workThe same learner's unaided baseline and post-trial workNo inspectable workWork is partly usable or needs material correctionWork meets the predefined standard
Related new taskPerformance on a structurally similar task the tutor did not seeCannot completeCompletes with material prompts from a human evaluatorCompletes and explains the relevant checks
Delayed checkThe same capability after the defined interval, without tutor helpCapability is absentPartial recall or inconsistent executionCapability remains usable
Evidence qualityTask definition, scoring notes, and raw learner artifactMissing or anecdotalSome records, but hard to auditRecords are complete and independently reviewable

The score is a buyer rubric, not a result from the cited studies. The studies justify why these rows should not be collapsed into engagement or next-item correctness. A Harvard randomized study used a separate team to construct test questions and measured learning outcomes and perceptions for an AI-tutor intervention versus active learning. A JAMA trial scored realistic performance separately from practice performance. Those designs support independent assessment as a procurement requirement, while their populations and tasks prevent a direct vendor ranking. The Harvard study and the JAMA trial are evidence for the measurement distinction, not for this score's weights.

Inspect whether the tutor teaches or gives answers

Score the tutor's interaction behavior from raw transcripts. A tutor that produces a correct answer quickly may be useful for reference, but it has not shown that it can support the capability the buyer is purchasing.

For each of the five trial scenarios, record the first meaningful tutor move:

  • 0 if it gives the answer, a complete solution, or an unearned recommendation before the learner reasons;
  • 1 if it offers a partial answer or hint but does not check the learner's reasoning;
  • 2 if it asks the learner to attempt or explain, gives the smallest useful help, checks the result, and releases the learner to finish.

Add a second field for recovery. When the learner is wrong, does the tutor expose the error and ask for a correction, or does it replace the learner's work? The recovery field matters because a tutor can appear Socratic in an easy scenario and turn into an answer dispenser under pressure.

This is where a raw transcript beats a vendor summary. Khan Academy's account of its tutor testing describes next-item correctness alongside guardrail metrics, latency, and comparisons with the existing experience across tutoring threads. That is a useful model for recording both outcome and interaction cost, but it is a first-party product report, not independent evidence that every tutor behaves the same way. Khan Academy's product-testing account is the source for those measurement categories.

There are legitimate exceptions. A learner may need an accessibility accommodation, a worked example, or a direct explanation after repeated attempts. Keep those cases in the transcript and label the help. The procurement question is not whether the tutor may ever answer. It is whether the buyer can see when, why, and how often it does so, and whether independent work remains possible afterward.

Run one bounded trial across the shortlist

Use the same task pack, learner conditions, scoring definitions, and evidence fields for every shortlisted configuration. Change one thing at a time when a configuration includes a model, prompt, or teacher-control setting.

  1. Choose one capability and write its definition of done, baseline task, related new task, and delayed-check task.
  2. Create five scenarios that exercise the same capability in different surface situations. Keep the success criteria stable even when the wording changes.
  3. Freeze the pack with a version, date, learner instructions, evaluator instructions, and allowed-help policy.
  4. Run the baseline without tutor help. Save the learner artifact and the evaluator's score before the tutor trial.
  5. Run the five scenarios and save the full transcript, product or model version, configuration, timestamps, latency method, and every direct-answer event.
  6. Give the learner the related new task without tutor assistance. Use an evaluator who did not coach the learner during the scenarios when possible.
  7. Run the delayed check after the same interval for every configuration. Record the interval instead of treating “later” as a result.
  8. Score the evidence and interaction rows blind to vendor identity when practical. Keep the raw artifacts behind the summary table.
  9. Check privacy, retention, export, teacher control, escalation, and implementation requirements against the buyer's own policy. A product feature is not a compliance decision by itself.
  10. Apply vetoes first, then calculate the weighted score. Write the reason for the decision and the next review condition.

This procedure creates a comparable procurement record without pretending that a small buyer trial is a controlled efficacy study. How to grade an AI output against a rubric is a useful companion for calibrating evaluator notes, while the commercial decision map is the parent for choosing the right kind of AI help.

Apply vetoes before the weighted score

Reject a configuration that fails a non-negotiable requirement, even if its weighted total looks attractive. A score should rank viable options, not compensate for a privacy failure or the inability to inspect independent work.

Use this worked decision table:

Gate or weighted rowRuleWeight after vetoes
Independent assessmentThe buyer can run and score unaided work using its own task packVeto if no
Privacy and retentionThe configuration meets the buyer's data, retention, and access policyVeto if no
Tutor behaviorAverage the five scenario interaction scores and review direct-answer events20%
Independent performanceUse the baseline, related new task, and delayed-check rows30%
Evidence and controlExports, transcript access, evaluator controls, escalation, and version records20%
Implementation fitSetup effort, workflow fit, support burden, and migration constraints15%
Commercial fitPrice, contract terms, renewal conditions, and acceptable switching cost15%

For each weighted row, assign 0, 1, or 2 using the evidence you retained. Calculate:

weighted score = sum(row score / 2 × row weight)

The percentages are a starting procurement recommendation, not a research finding. A buyer can change them before the trial, but must not change them after seeing a preferred vendor's result. A simple decision rule is:

  • any veto fails: reject or return to the vendor with a written condition;
  • no veto fails and the total is below 60: reject for this use case;
  • no veto fails and the total is 60 to 74: run a restricted pilot with named safeguards;
  • no veto fails and the total is 75 or higher: shortlist for procurement review, subject to contract and policy approval.

The score still has an exception. If the capability is high stakes, a buyer may set a stricter independent-performance floor instead of relying on the total. If the product cannot expose enough evidence to score that floor, the correct answer is “not ready for this use case,” not a higher weight for another row.

Treat the score as a purchase gate, not an efficacy claim

The artifact tells a buyer whether a shortlisted tutor produced enough evidence for the next procurement step. It does not prove that the tutor works for every learner, subject, or workplace.

The cited research uses different populations, tasks, interventions, and outcome definitions. Khan Academy's report describes its own product-testing practice. The scorecard borrows useful distinctions from those sources and turns them into a buyer process. It does not combine their findings into a single effect size, and it does not claim that one trial can establish long-term capability transfer.

That boundary is useful. It keeps a procurement conversation honest while still giving the buyer something concrete to do: define done, preserve raw evidence, inspect answer-giving behavior, check the operating constraints, and apply vetoes before price or polish wins the decision.

If you are choosing between tutoring, consulting, and an internal team, compare the shape of the work and the capability you need to own. How to compare AI consulting proposals covers the neighboring commercial decision. This scorecard is the right next artifact when the shortlist contains AI tutors and the buyer needs a defensible procurement record.