Field note · opportunity

How to Score AI Task Judgment Separately From Prompt Quality

Score AI task judgment separately from prompt quality with a paired decision trace covering framing, output checks, assumptions, and justification.

10 minute read
  • AI learning
  • Prompt engineering
  • Task judgment
  • AI training
Illustration of separate prompt construction and task judgment lanes in an AI learning diagnostic

Prompt training often produces a visible improvement first. The request gets clearer. The output looks more organized. A learner can explain why one instruction is better than another.

That improvement matters, but it does not answer the harder question: can the learner decide what the task is, what good looks like, what the model missed, and whether the work is safe to accept?

Illustration of separate prompt construction and task judgment lanes in an AI learning diagnostic

Prompt quality and task judgment need separate scores

Score prompt construction and task judgment separately because they answer different questions. Prompting shapes the request. Judgment defines the decision and tests the result against consequences.

The distinction is visible in the research. The Co-Reasoning paper separates Framing, Judging, and Steering, and describes Judging as the gate between deciding what the task means and acting on the result. Its abstract describes simulated learners and points to human-rater agreement and outcome evaluation as work still ahead, so it does not establish a measured workplace transfer rate. The Co-Reasoning paper supports the competency distinction, not a claim about how often learners fail.

A 2024 intervention with 27 undergraduate students reported improvement in AI self-efficacy, AI knowledge, and prompt-engineering ability after a 100-minute workshop. The abstract does not report separate workplace scores for task framing, output checking, assumption detection, or decision justification. The intervention study is useful evidence that prompt skill can be taught. It is not evidence that prompt skill transfers into judgment.

The same boundary appears in the Frontiers paper. It connects prompt engineering with the problem, context, constraints, and critical online reasoning, then says concrete assessment recommendations are beyond its scope. The Frontiers paper leaves room for a practical diagnostic, but it does not supply one.

My teaching experience points to the same measurement choice without turning it into a study result. I have taught product managers who went from writing specs to building and shipping the product. The recurring capability question was not only whether they could ask a model for code or a plan. It was whether they could say what done meant and check the work against it. That is why the diagnostic below scores the task contract separately from the prompt.

The useful conclusion is narrow: a prompt score answers one question about request construction. It cannot stand in for a judgment score. A review rubric needs both.

How to run a decision-review trace

Run the review by holding the workplace decision constant, allowing the prompt to improve, and recording whether the learner makes the decision contract explicit. This is a scoring procedure, not a claim that a particular learner or percentage has already failed.

Use an ordinary case with a visible consequence and enough information for a reviewer to check the result. Examples include choosing which customer requests belong in the next support batch, deciding whether a document is complete enough for review, or selecting which product feedback needs escalation. Avoid a case where any polished prose counts as success.

Run the case in two passes:

  1. Ask for the initial task brief, success criteria, prompt, model output, and accept-or-reject decision.
  2. Let the learner revise the prompt without changing the underlying case, decision owner, or available evidence.
  3. Capture the revised prompt and output beside the first version.
  4. Require the learner to name what they checked outside the model output, which assumption could change the decision, and what happens next.
  5. Score prompt quality separately from the four judgment dimensions below.
  6. Repeat the exercise on a new case before treating the score as a stable capability observation.

The trace must preserve the connection between request and decision. A compact record looks like this:

Trace fieldWhat it revealsWhat a reviewer checks
Initial task briefWhether the learner named the real jobDecision, owner, scope, and consequence are explicit
Success criteriaWhether quality is defined before generationCriteria are observable and relevant to the decision
Initial and revised promptWhether request construction improvedContext, constraints, output format, and missing information
Model outputWhat the learner had to inspectClaims, omissions, unsupported assumptions, and edge cases
Accept or reject decisionWhether the learner owns the resultDecision is tied to evidence and criteria, not confidence
External checks and next actionWhether the workflow closes the loopThe learner verifies, escalates, repairs, or records uncertainty

Do not report a score as a research result unless a real learner trace has been collected, anonymized, reviewed, and archived. The reproducible part is the case design, fields, rubric, and review sequence. It is valuable because a team can run it without pretending that the worksheet itself is a result.

The AI learning capability hub gives this diagnostic a broader place in a capability plan. For a related training-transfer decision, compare it with how to make AI training stick in a small team. The links are useful for sequencing the work, not for proving the diagnosis.

Illustration of a decision-review worksheet with task criteria, prompt revision, output checks, assumptions, and an accept-or-reject decision

How to score the decision trace

Keep one prompt score and four judgment scores. A single blended score hides the failure the team is trying to see.

Score each dimension from 0 to 2. The rubric is deliberately small enough to apply to a trace and specific enough to explain a repair.

Dimension012
Prompt qualityThe request is missing or unusableThe request can run but has material omissionsThe request includes relevant context, constraints, and output requirements
Task framingNo decision, owner, or success conditionSome task detail exists, but material ambiguity remainsDecision, owner, scope, success criteria, and constraints are explicit
Output checkingAccepts or rejects without checkingChecks surface features onlyChecks material claims, omissions, constraints, and decision relevance
Assumption detectionDoes not identify assumptionsNames assumptions without testing importanceIdentifies material hidden assumptions and tests or escalates them
Decision justificationUses preference or confidence onlyRationale cites some evidenceDecision is tied to evidence, criteria, risk, and a clear next action

The review result is not “the learner scored 7 out of 10.” That kind of total can conceal a dangerous pattern. The useful result is a profile such as: prompt quality is 2, task framing is 1, output checking is 0, and the final decision has no external evidence. That profile tells the trainer what to practice next without claiming a prevalence rate.

The main source of error is a category mistake. A learner can add role, format, constraints, and examples to a request while still accepting the first plausible output. The prompt is then better as a prompt. The work is not better as a decision.

The Chinese University of Hong Kong research record is a useful boundary case. It describes a 2024-2025 action-research intervention with 96 students, a Prompt-Observe-Evaluate protocol, and self-reported changes in self-efficacy and critical-thinking or ethical-AI outcomes. Those reported outcomes do not provide the anonymized workplace traces, paired prompt revisions, separate judgment scores, or decision-review worksheet required by this diagnostic. The CUHK research record supports careful separation of reported perception from observed task performance.

How to choose the next exercise from the scores

Use the weakest judgment dimension to choose a changed exercise, then verify the same dimension on a new case. A better prompt alone is not a repair for missing task ownership or weak checking.

  1. Repair framing first. Ask the learner to write the decision owner, scope, success criteria, constraints, and escalation condition before they write a revised prompt. If they cannot define the decision, more prompt examples add surface skill to an undefined task.
  2. Repair checking with a stop condition. Require a list of material claims, omissions, and constraints to check before acceptance. A format check is not enough when the output informs a consequential choice.
  3. Repair assumption detection with a counterfactual. Ask which fact, if changed, would reverse the decision. The learner must either test that fact, find an authoritative source, or escalate the case.
  4. Repair justification with evidence. Make accept or reject a recorded decision. The rationale names the criteria, evidence checked, residual uncertainty, and next action. Confidence is not a substitute for any of these.
  5. Verify on a new case. Keep the scoring dimensions fixed but change the surface task. The score is useful only if the learner can recreate the decision contract and checks without copying the demonstrated workflow.

This repair loop is compatible with the risk-management principle that AI work needs defined responsibilities, evaluation, and attention to context. The NIST AI Risk Management Framework is a primary governance reference for that broader boundary. It does not validate a learner score, so use it to shape the review conditions rather than to decorate the article with authority.

The decision-review worksheet is the practical artifact:

Case ID and date:
Model and configuration:
Decision owner:

Initial task brief:
Success criteria:
Initial prompt:
Revised prompt:
AI output:

Prompt quality score (0-2):
Task framing score (0-2):
Output checking score (0-2):
Assumption detection score (0-2):
Decision justification score (0-2):

Accept or reject:
Evidence for the decision:
Material assumption that could change the decision:
What I checked outside the model output:
Escalation or next action:
Reviewer notes:

The worksheet is complete when the reviewer can follow the path from task to prompt, output, checks, assumption, decision, and next action. It is not complete because the prompt looks sophisticated.

When this rubric is not the right tool

Do not use this rubric to turn every drafting exercise into a formal evaluation. It matters most when the learner must make, recommend, or prepare a decision with a meaningful consequence.

For low-stakes rewriting, brainstorming, formatting, or exploration, prompt quality may be the main skill being taught. If a human already owns the criteria and checks every line, the learner's task judgment is not the bottleneck in that exercise. Say so clearly rather than pretending that one score measures everything.

The diagnosis also changes when the task has no stable answer. In that case, the reviewer can still score framing, assumptions, and justification, but the success criteria must describe a useful decision process, not a single correct output. A good rationale can remain valid when reasonable people choose differently.

Do not infer a general causal claim from one paired task. A thin corpus can show that a learner needs practice on a dimension. It cannot establish that AI learners as a population improve prompting but not judgment. The supplied research has the same limit: the studies and records support distinctions among prompting, reasoning, and self-reported outcomes, while the proposed workplace trace supplies the missing operational check.

If a team wants to record a capability observation, use the smallest honest claim: the learner completed a new task, stated the decision and criteria, checked the output, surfaced a material assumption, and justified the next action. That is stronger than a prompt screenshot and narrower than a population statistic.

The next step is to run one anonymized case with the worksheet, review it with the decision owner, and fix the weakest dimension before adding another prompt lesson. For hands-on help building that capability around real work, learn how Marius Manolachi teaches teams to build AI products.

Sources

Questions people ask next

Does better prompting prove that AI training worked?

No. It proves that the learner can write a better request under the tested conditions. Training transfer needs a separate check of task framing, output inspection, material assumptions, and the final decision rationale.

What is the simplest test for judgment transfer?

Give the learner a paired workplace task, score the prompt separately, and require an accept-or-reject decision tied to criteria, evidence, checks, and assumptions. Repeat with a new case instead of grading the training example only.