Field note · capability
What Does AI Fluency Look Like in a Product Review Meeting?
A three-case fictional test shows the meeting behaviours that turn a polished AI recommendation into a decision-ready product review packet.

The dangerous AI recommendation is not the one that sounds confused. It is the one that sounds ready before anyone has checked what “ready” means.
When I taught product managers to move from writing specs to building and shipping, the recurring failure was an undefined meaning of done, not the model. A product review meeting exposes the same gap. The recommendation can be polished while the decision is still ungrounded.

What does AI fluency look like in the meeting?
AI fluency looks like turning AI output into a reviewable decision, not treating a fluent paragraph as the decision itself. The reviewer states the choice, checks the evidence and its limits, names the best alternative, preserves contrary evidence, and refuses to approve an action when a failure condition or human boundary is unresolved.
That is a situated version of the broader definitions. Anthropic describes fluency through Delegation, Description, Discernment, and Diligence. Zapier describes it through purposeful use, clear communication, evaluation, risk judgment, and integration into real work. The U.S. Department of Labor makes the practical point that the required depth of AI skill depends on the role and context. (Anthropic's AI Fluency Framework, Zapier's AI fluency guide, the Department of Labor's AI Literacy Framework)
For a product manager, the meeting is a useful test because it makes judgment visible. You can hear whether someone asks, “What would make this recommendation unsafe?” or only asks, “Can you make the recommendation shorter?”
What changed in the three-case exercise?
The sourceable result is a small, fictional exercise, not a workplace outcome. I gave GPT-5 three fictional product-review cases in the same format, ran a concise baseline prompt, then ran a repaired prompt that required a full decision packet. A manual reviewer scored seven fields from 0 to 2.
| Run | Packet points covered | What the output looked like |
|---|---|---|
| Baseline AI review | 22/42 | Polished recommendations with missing boundaries, alternatives, or veto conditions |
| Repaired AI review | 42/42 | Explicit packets with evidence, contrary evidence, decision conditions, and next ownership |
The score counts field coverage. It does not prove that a fictional source is true, that a threshold is appropriate, or that a real product will succeed. The complete prompts, inputs, raw outputs, corrections, and limitations are saved in the research record for this post.
The practical finding is narrower and more useful than “AI fluency matters”: a reviewer can look fluent while silently dropping the fields that make a product decision safe to act on. Requiring the packet makes those omissions visible.
Reproduction settings
The run used GPT-5 in this Codex session on 2026-08-23. No external product tools were available to the case prompts, and hidden temperature settings were not exposed.
Baseline prompt:
You are the PM's meeting-prep assistant. Given one product review case, recommend a path in no more than 120 words. Prioritize a crisp recommendation, summarize the strongest evidence, and name the next step. Do not invent numbers or sources. You may resolve ambiguity with a reasonable assumption.
Repaired prompt:
You are preparing a decision-ready product-review packet. Use only the supplied case fields. Do not invent evidence, thresholds, owners, or dates. Return exactly these fields: decision requested, decision, evidence links, contrary evidence, alternatives considered, quality or failure criteria, human approval boundary, owner/date, unresolved question, and next step. If a veto condition is not met or evidence is insufficient, set decision to HOLD and state what must be learned before approval.
Where did the polished baseline review fail?
The baseline prompt asked for a crisp recommendation, the strongest evidence, and a next step. It allowed the model to resolve ambiguity with a reasonable assumption. That last instruction created the failure clinic.
Case A: the pricing-page monitor
The fictional case asked whether to run a 10% pilot for an assistant that watches public competitor pricing pages. The input included 30 labelled samples with four missed changes and two false alarms, a 90% recall criterion, a manual-review boundary, and two alternatives.
The baseline said:
Recommend a 10% pilot next week. The labelled sample shows strong signal quality and the interviews suggest a useful weekly workflow. Keep a human check before alerts are distributed and monitor false positives during the four-week pilot.
That paragraph turned 80% recall into “strong signal quality,” omitted the four misses, and replaced a release criterion with “monitor.” It also did not preserve the deterministic diff alternative. The repaired packet set the decision to HOLD until the held-out set reached 90% recall and incomplete extraction caused an abstention.
Case B: the support-ticket triage assistant
The second fictional case concerned low-risk support tags. The product could suggest a tag, but it could not reply, close a ticket, or mutate a billing record. The input included 40 low-risk tickets, three uncertain cases, and an explicit exclusion for billing disputes.
The baseline recommended a two-week shadow pilot with human oversight. That was directionally right, but “human oversight” did not say what the human could approve. The repaired packet preserved the 90% threshold, required abstention for uncertain and billing-related tickets, and limited the pilot to tag suggestions. This was the only repaired case marked PROCEED, and it proceeded as a shadow-only review.
Case C: the onboarding coach
The third fictional case asked whether to ship an AI onboarding coach to 25% of users. Eight of ten participants completed a plan, but three of ten transcript reviews showed users treating suggestions as product promises. Two sessions also lacked account context.
The baseline recommended the beta and suggested clearer copy. The repaired packet set the decision to HOLD until the team retested promise confusion, made missing context trigger a question, and preserved the rule that no plan could be committed automatically.
The pattern is the point. A limited beta is not automatically safe. Human oversight is not a boundary until the packet says what the human must approve. A completion rate is not enough when the failure is a misleading commitment.
What belongs in a decision-ready packet?
A decision-ready packet contains the information another reviewer needs to challenge the recommendation without reconstructing the case from memory.
| Field | The reviewer should be able to answer | Failure signal |
|---|---|---|
| Decision requested | What exact choice must this meeting make? | The recommendation solves a different problem. |
| Evidence links | Which source records support the recommendation? | Evidence is described but cannot be inspected. |
| Contrary evidence | What points against the preferred option? | Caveats are replaced by optimism. |
| Alternatives | What else could we do, and what would each trade away? | Alternatives are strawmen or absent. |
| Quality or failure criteria | What must be true, and what is a veto? | “Monitor it” replaces a condition. |
| Human approval boundary | What may AI do, and what must a person approve? | “Human in the loop” has no defined action. |
| Owner/date | Who takes the next action, and by when? | The meeting ends with a noun instead of an owner. |
Productboard's product-review guidance uses the same decision shape: state the decision, recommendation, evidence, contrary evidence, alternatives, open questions, and follow-up. Its AI-evals guidance adds an important product role: PMs help define what good means, which failures matter, and how evaluation results change the roadmap. (Productboard's Product Review Presentation, Productboard's AI evals guide)
The packet is not a new fluency framework. It is a meeting artifact. That distinction matters because a framework can name discernment while leaving the reviewer with no evidence to inspect.
How can a team repeat the exercise?
Use three cases with different failure shapes. Do not let everyone review one flattering example.
- Freeze the input. Record the feature, user/job, exact decision requested, evidence links, contrary evidence, alternatives, criteria, approval boundary, and owner/date.
- Run the baseline. Ask the AI for a short recommendation. Do not give it a packet schema. Save the model, version, date, prompt, input, and raw output.
- Score the omissions. Give 0, 1, or 2 points for each of the seven fields. Do not award a point because the output sounds sensible.
- Correct the review contract. Ask for the full packet. Tell the AI to set HOLD when a veto condition is not met or evidence is insufficient. Keep the original output beside the repaired one.
- Challenge the decision. A reviewer checks the packet against the source records and asks which field would change the decision if it were false.
- Transfer the skill. Give the learner a fourth case they have not seen. Ask them to produce the packet without the prompt's explanation, then compare their choices with an independent reviewer.
The transfer check is important. A person who can fill the packet only after seeing the answer has learned the format, not yet the judgment. Ask them to reject one recommendation, defend one bounded approval, and name the missing evidence in a new case.
Blank packet
Copy this for a real or fictional case. In a live review, replace fictional IDs with inspectable links.
Case and feature:
User/job:
Decision requested and options:
Evidence links:
Contrary evidence:
Alternatives and trade-offs:
Quality or failure criteria:
Human approval boundary:
Owner/date:
Unresolved question:
Next step:
Baseline output:
Reviewer corrections:
Repaired output:
Score each field from 0 to 2:
Decision / Evidence / Contrary evidence / Alternatives /
Failure criteria / Approval boundary / Owner-date
Total: /14
Veto condition passed: yes/no
Limitation:
What should the reviewer say out loud?
Fluency becomes audible through challenge questions that force the recommendation to carry its own conditions.
- “What exact decision are we making today?”
- “Which source link supports that sentence?”
- “What evidence would make you change the recommendation?”
- “Which alternative did you reject, and what are we giving up?”
- “What is the veto condition?”
- “What can the AI do without approval?”
- “Who owns the next action, and what date makes this reviewable?”
These questions are not a performance of scepticism. They are a way to separate useful AI assistance from delegated accountability. The AI can draft, compare, and expose gaps. The product manager still owns the decision boundary.
When I teach teams to build with AI, I want the learner to leave with a better next move, not merely a better paragraph. In a product review, the better next move may be “HOLD and collect the missing evidence.” That is a successful review outcome when the current evidence does not support approval.
When is a full packet unnecessary?
A full packet is unnecessary for a low-risk, reversible choice where the evidence is complete, the action has no material downstream consequence, and the owner can undo it. Even then, record the decision and owner so the team does not confuse a quick choice with an unowned choice.
Use the full packet when the recommendation can launch a customer-facing feature, change a policy, affect money or access, expose sensitive information, or create an irreversible workflow. Also use it when the AI is making a claim that sounds more certain than the evidence.
The exercise has limits. It uses three fictional cases, one GPT-5 session, one manual reviewer, and a completeness rubric. It does not estimate model accuracy, team adoption, meeting speed, user trust, or production safety. Its result is a learning artifact: the baseline made omissions easy to miss, and the repaired schema made them easy to challenge.
If you want to turn the exercise into a team practice, start with what to learn before building AI agents, compare the packet with a guide to measuring AI output quality, and then use Marius Manolachi's AI consulting and tutoring work when the team needs a facilitated capability loop.
Continue with a related field note
Questions people ask next
What should a product manager ask an AI before trusting its recommendation?
Ask it to show the decision requested, source links, contrary evidence, alternatives, failure criteria, approval boundary, owner, date, unresolved questions, and next step. If any field is missing, treat the recommendation as incomplete rather than filling the gap with confidence.
Can an AI make the final product decision in a review meeting?
It can recommend, compare, and expose missing evidence, but the accountable human should approve the decision when the action has material product, customer, financial, or safety consequences. Keep the approval boundary explicit in the packet.
Is this three-case test a benchmark of AI models?
No. The cases are fictional fixtures, and the score measures packet-field coverage in one GPT-5 session. It is a learning exercise and reusable artifact, not a workplace outcome, adoption rate, or model benchmark.