Field note · opportunity
How to Score AI Task Judgment Separately From Prompt Quality
Score AI task judgment separately from prompt quality with a paired decision trace covering framing, output checks, assumptions, and justification.

Prompt training often produces a visible improvement first. The request gets clearer. The output looks more organized. A learner can explain why one instruction is better than another.
That improvement matters, but it does not answer the harder question: can the learner decide what the task is, what good looks like, what the model missed, and whether the work is safe to accept?

Prompt quality and task judgment need separate scores
Score prompt construction and task judgment separately because they answer different questions. Prompting shapes the request. Judgment defines the decision and tests the result against consequences.
The distinction is visible in the research. The Co-Reasoning paper separates Framing, Judging, and Steering, and describes Judging as the gate between deciding what the task means and acting on the result. Its abstract describes simulated learners and points to human-rater agreement and outcome evaluation as work still ahead, so it does not establish a measured workplace transfer rate. The Co-Reasoning paper supports the competency distinction, not a claim about how often learners fail.
A 2024 intervention with 27 undergraduate students reported improvement in AI self-efficacy, AI knowledge, and prompt-engineering ability after a 100-minute workshop. The abstract does not report separate workplace scores for task framing, output checking, assumption detection, or decision justification. The intervention study is useful evidence that prompt skill can be taught. It is not evidence that prompt skill transfers into judgment.
The same boundary appears in the Frontiers paper. It connects prompt engineering with the problem, context, constraints, and critical online reasoning, then says concrete assessment recommendations are beyond its scope. The Frontiers paper leaves room for a practical diagnostic, but it does not supply one.
My teaching experience points to the same measurement choice without turning it into a study result. I have taught product managers who went from writing specs to building and shipping the product. The recurring capability question was not only whether they could ask a model for code or a plan. It was whether they could say what done meant and check the work against it. That is why the diagnostic below scores the task contract separately from the prompt.
The useful conclusion is narrow: a prompt score answers one question about request construction. It cannot stand in for a judgment score. A review rubric needs both.
How to run a decision-review trace
Run the review by holding the workplace decision constant, allowing the prompt to improve, and recording whether the learner makes the decision contract explicit. This is a scoring procedure, not a claim that a particular learner or percentage has already failed.
Use an ordinary case with a visible consequence and enough information for a reviewer to check the result. Examples include choosing which customer requests belong in the next support batch, deciding whether a document is complete enough for review, or selecting which product feedback needs escalation. Avoid a case where any polished prose counts as success.
Run the case in two passes:
- Ask for the initial task brief, success criteria, prompt, model output, and accept-or-reject decision.
- Let the learner revise the prompt without changing the underlying case, decision owner, or available evidence.
- Capture the revised prompt and output beside the first version.
- Require the learner to name what they checked outside the model output, which assumption could change the decision, and what happens next.
- Score prompt quality separately from the four judgment dimensions below.
- Repeat the exercise on a new case before treating the score as a stable capability observation.
The trace must preserve the connection between request and decision. A compact record looks like this:
| Trace field | What it reveals | What a reviewer checks |
|---|---|---|
| Initial task brief | Whether the learner named the real job | Decision, owner, scope, and consequence are explicit |
| Success criteria | Whether quality is defined before generation | Criteria are observable and relevant to the decision |
| Initial and revised prompt | Whether request construction improved | Context, constraints, output format, and missing information |
| Model output | What the learner had to inspect | Claims, omissions, unsupported assumptions, and edge cases |
| Accept or reject decision | Whether the learner owns the result | Decision is tied to evidence and criteria, not confidence |
| External checks and next action | Whether the workflow closes the loop | The learner verifies, escalates, repairs, or records uncertainty |
Do not report a score as a research result unless a real learner trace has been collected, anonymized, reviewed, and archived. The reproducible part is the case design, fields, rubric, and review sequence. It is valuable because a team can run it without pretending that the worksheet itself is a result.
The AI learning capability hub gives this diagnostic a broader place in a capability plan. For a related training-transfer decision, compare it with how to make AI training stick in a small team. The links are useful for sequencing the work, not for proving the diagnosis.

How to score the decision trace
Keep one prompt score and four judgment scores. A single blended score hides the failure the team is trying to see.
Score each dimension from 0 to 2. The rubric is deliberately small enough to apply to a trace and specific enough to explain a repair.
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| Prompt quality | The request is missing or unusable | The request can run but has material omissions | The request includes relevant context, constraints, and output requirements |
| Task framing | No decision, owner, or success condition | Some task detail exists, but material ambiguity remains | Decision, owner, scope, success criteria, and constraints are explicit |
| Output checking | Accepts or rejects without checking | Checks surface features only | Checks material claims, omissions, constraints, and decision relevance |
| Assumption detection | Does not identify assumptions | Names assumptions without testing importance | Identifies material hidden assumptions and tests or escalates them |
| Decision justification | Uses preference or confidence only | Rationale cites some evidence | Decision is tied to evidence, criteria, risk, and a clear next action |
The review result is not “the learner scored 7 out of 10.” That kind of total can conceal a dangerous pattern. The useful result is a profile such as: prompt quality is 2, task framing is 1, output checking is 0, and the final decision has no external evidence. That profile tells the trainer what to practice next without claiming a prevalence rate.
The main source of error is a category mistake. A learner can add role, format, constraints, and examples to a request while still accepting the first plausible output. The prompt is then better as a prompt. The work is not better as a decision.
The Chinese University of Hong Kong research record is a useful boundary case. It describes a 2024-2025 action-research intervention with 96 students, a Prompt-Observe-Evaluate protocol, and self-reported changes in self-efficacy and critical-thinking or ethical-AI outcomes. Those reported outcomes do not provide the anonymized workplace traces, paired prompt revisions, separate judgment scores, or decision-review worksheet required by this diagnostic. The CUHK research record supports careful separation of reported perception from observed task performance.
How to choose the next exercise from the scores
Use the weakest judgment dimension to choose a changed exercise, then verify the same dimension on a new case. A better prompt alone is not a repair for missing task ownership or weak checking.
- Repair framing first. Ask the learner to write the decision owner, scope, success criteria, constraints, and escalation condition before they write a revised prompt. If they cannot define the decision, more prompt examples add surface skill to an undefined task.
- Repair checking with a stop condition. Require a list of material claims, omissions, and constraints to check before acceptance. A format check is not enough when the output informs a consequential choice.
- Repair assumption detection with a counterfactual. Ask which fact, if changed, would reverse the decision. The learner must either test that fact, find an authoritative source, or escalate the case.
- Repair justification with evidence. Make accept or reject a recorded decision. The rationale names the criteria, evidence checked, residual uncertainty, and next action. Confidence is not a substitute for any of these.
- Verify on a new case. Keep the scoring dimensions fixed but change the surface task. The score is useful only if the learner can recreate the decision contract and checks without copying the demonstrated workflow.
This repair loop is compatible with the risk-management principle that AI work needs defined responsibilities, evaluation, and attention to context. The NIST AI Risk Management Framework is a primary governance reference for that broader boundary. It does not validate a learner score, so use it to shape the review conditions rather than to decorate the article with authority.
The decision-review worksheet is the practical artifact:
Case ID and date:
Model and configuration:
Decision owner:
Initial task brief:
Success criteria:
Initial prompt:
Revised prompt:
AI output:
Prompt quality score (0-2):
Task framing score (0-2):
Output checking score (0-2):
Assumption detection score (0-2):
Decision justification score (0-2):
Accept or reject:
Evidence for the decision:
Material assumption that could change the decision:
What I checked outside the model output:
Escalation or next action:
Reviewer notes:
The worksheet is complete when the reviewer can follow the path from task to prompt, output, checks, assumption, decision, and next action. It is not complete because the prompt looks sophisticated.
When this rubric is not the right tool
Do not use this rubric to turn every drafting exercise into a formal evaluation. It matters most when the learner must make, recommend, or prepare a decision with a meaningful consequence.
For low-stakes rewriting, brainstorming, formatting, or exploration, prompt quality may be the main skill being taught. If a human already owns the criteria and checks every line, the learner's task judgment is not the bottleneck in that exercise. Say so clearly rather than pretending that one score measures everything.
The diagnosis also changes when the task has no stable answer. In that case, the reviewer can still score framing, assumptions, and justification, but the success criteria must describe a useful decision process, not a single correct output. A good rationale can remain valid when reasonable people choose differently.
Do not infer a general causal claim from one paired task. A thin corpus can show that a learner needs practice on a dimension. It cannot establish that AI learners as a population improve prompting but not judgment. The supplied research has the same limit: the studies and records support distinctions among prompting, reasoning, and self-reported outcomes, while the proposed workplace trace supplies the missing operational check.
If a team wants to record a capability observation, use the smallest honest claim: the learner completed a new task, stated the decision and criteria, checked the output, surfaced a material assumption, and justified the next action. That is stronger than a prompt screenshot and narrower than a population statistic.
The next step is to run one anonymized case with the worksheet, review it with the decision owner, and fix the weakest dimension before adding another prompt lesson. For hands-on help building that capability around real work, learn how Marius Manolachi teaches teams to build AI products.
Sources
- Co-Reasoning: Framing, Judging, and Steering
- Prompt engineering intervention study
- Frontiers prompt-engineering framework
- Chinese University of Hong Kong research record
- NIST AI Risk Management Framework
Continue with a related field note
Questions people ask next
Does better prompting prove that AI training worked?
No. It proves that the learner can write a better request under the tested conditions. Training transfer needs a separate check of task framing, output inspection, material assumptions, and the final decision rationale.
What is the simplest test for judgment transfer?
Give the learner a paired workplace task, score the prompt separately, and require an accept-or-reject decision tied to criteria, evidence, checks, and assumptions. Repeat with a new case instead of grading the training example only.