Field note · opportunity
Why AI Discovery Selects the Loudest Workflow, Not the Best One
A blind-versus-visible enthusiasm test shows how a promising AI workflow can lose rank when the room rewards salience over evidence.

The loudest workflow usually arrives with a story attached to it. Someone can picture the demo. Someone important wants it. The outcome sounds obvious.
That is useful adoption evidence. It is not the same as evidence that the workflow deserves to go first.
I built a small counterfactual worksheet to see how much the ordering can move. The result is the artifact, not a theory about office politics.
The same shortlist changed when enthusiasm became visible
In both scoring passes, the synthetic supplier-exception workflow ranked first when enthusiasm was hidden. Adding a visible 1-5 enthusiasm field changed the ordering in both passes. In Scorer A, the synthetic candidate fell from first to fourth. In Scorer B, it fell from first to second.
| Candidate | Workflow | Scorer A hidden to visible | Scorer B hidden to visible |
|---|---|---|---|
| H | Synthetic supplier-exception review | 1 to 4 | 1 to 2 |
| B | Invoice extraction and exception triage | 2 to 1 | 2 to 1 |
| C | Meeting follow-up and action draft | 4 to 3 | 3 to 3 |
| E | Support-ticket triage and routing | 3 to 2 | 4 to 4 |
The practical diagnosis is simple: the shortlist was sensitive to salience. That does not prove that a political room would make the same choices, and it does not prove that H is objectively the best workflow. It shows why a team should publish the rank-change calculation before treating a room's favorite as a portfolio decision.

What the worksheet tested
The test separates operational evidence from rollout energy. It uses eight candidates, seven operational fields, and one enthusiasm field.
Seven candidates are public-example patterns: internal policy Q&A, invoice extraction, meeting follow-up, customer-feedback clustering, support-ticket triage, contract-clause extraction, and field-service visit preparation. These patterns appear in public opportunity and use-case materials from OpenAI Academy and Microsoft. Microsoft’s public blueprints pair examples such as finance, customer service, knowledge, legal, sales, and field service with KPIs and baselines to capture before deployment. Microsoft’s use-case blueprints are useful inputs, but they do not provide a local decision for your company.
The eighth candidate is deliberately synthetic:
A procurement analyst receives a source-linked brief about twelve supplier exceptions each week. The brief takes the analyst 90 minutes to prepare. The workflow only extracts, groups, and drafts recommendations. The analyst approves every action.
That scenario is not a client result. It exists to give the worksheet one candidate with a defined operating shape, a reversible action boundary, and low stakeholder enthusiasm.
The candidate descriptions, public basis, baseline notes, raw scores, and limitations are all published in the reviewable research artifact. The useful point is not the eight names. It is the fixed field set that makes a loud candidate easier to challenge.
Score the operational fields before you discuss enthusiasm
Use seven fields, each scored from 1 to 5. Higher is better. For control risk and adoption burden, a 5 means lower risk or lower burden.
| Field | Score 1 | Score 5 |
|---|---|---|
| Outcome | The intended change is vague or feature-shaped. | The result is specific, observable, and consequential. |
| Baseline | No current-state measure or credible sampling plan. | A current measure, owner, and collection method are clear. |
| Workflow fit | The idea sits outside the real workflow or needs a new operating model. | The proposed action fits an existing workflow and decision owner. |
| Data readiness | Inputs are inaccessible, unstable, or untrusted. | Inputs are available, bounded, and usable for the first test. |
| Control risk | A wrong result could trigger hard-to-reverse harm or lacks review. | The action is read-only or reversible, with clear review and escalation. |
| Adoption burden | New roles, systems, or habits are heavy and unowned. | The first users, owner, and operating change are clear and small. |
| Learning value | A test would teach little or leave the next decision unclear. | The smallest test would resolve an important uncertainty. |
This field set is a local decision artifact, not a standard claimed by NIST. It is informed by public guidance that asks teams to define a business priority, workflow owner, outcome, baseline, controls, readiness, and evidence gap before a bounded test. The OpenAI Academy opportunity worksheet gives that kind of structure. NIST’s AI Risk Management Framework adds the need for broad perspectives and lifecycle risk management through Govern, Map, Measure, and Manage. NIST’s AI RMF is a voluntary risk resource, not a magic scoring formula.
Make the rank change calculable
The hidden score is the sum of the seven operational fields. The enthusiasm field stays off the sheet.
The visible score is:
visible score = hidden operational score + stakeholder enthusiasm
Enthusiasm is scored from 1 to 5 and published alongside the core score. Ranks are descending. Ties go to the higher core score, then the earlier candidate ID. Rank change is:
rank change = hidden rank - visible rank
A positive number means the candidate moved up when enthusiasm became visible. This is intentionally plain. A team can change the weight, but it must publish the weight. Otherwise, “stakeholder alignment” becomes an invisible bonus that only some candidates receive.
Here is the raw output for the two passes.
Scorer A raw output
| ID | Operational vector in field order | Hidden | Enthusiasm | Visible | Hidden rank | Visible rank | Change |
|---|---|---|---|---|---|---|---|
| A | 4, 2, 4, 4, 3, 4, 3 | 24 | 4 | 28 | 6 | 6 | 0 |
| B | 4, 5, 5, 3, 3, 3, 4 | 27 | 3 | 30 | 2 | 1 | +1 |
| C | 3, 1, 5, 4, 5, 5, 2 | 25 | 5 | 30 | 4 | 3 | +1 |
| D | 3, 2, 4, 3, 5, 4, 4 | 25 | 3 | 28 | 5 | 5 | 0 |
| E | 4, 4, 5, 3, 3, 3, 4 | 26 | 4 | 30 | 3 | 2 | +1 |
| F | 4, 3, 5, 3, 2, 3, 3 | 23 | 2 | 25 | 7 | 7 | 0 |
| G | 4, 2, 3, 2, 3, 2, 4 | 20 | 4 | 24 | 8 | 8 | 0 |
| H | 4, 3, 5, 4, 5, 3, 4 | 28 | 1 | 29 | 1 | 4 | -3 |
Scorer B raw output
| ID | Operational vector in field order | Hidden | Enthusiasm | Visible | Hidden rank | Visible rank | Change |
|---|---|---|---|---|---|---|---|
| A | 3, 2, 4, 4, 3, 4, 3 | 23 | 5 | 28 | 6 | 6 | 0 |
| B | 5, 4, 5, 3, 3, 3, 5 | 28 | 4 | 32 | 2 | 1 | +1 |
| C | 3, 1, 5, 4, 5, 5, 3 | 26 | 5 | 31 | 3 | 3 | 0 |
| D | 3, 2, 4, 3, 5, 4, 4 | 25 | 4 | 29 | 5 | 5 | 0 |
| E | 4, 4, 5, 3, 3, 3, 4 | 26 | 4 | 30 | 4 | 4 | 0 |
| F | 4, 3, 5, 3, 2, 2, 3 | 22 | 3 | 25 | 7 | 7 | 0 |
| G | 4, 2, 3, 2, 3, 2, 4 | 20 | 5 | 25 | 8 | 8 | 0 |
| H | 4, 4, 5, 4, 5, 4, 4 | 30 | 1 | 31 | 1 | 2 | -1 |
The scorers disagree on magnitude. Scorer A gives H a 28 and a three-place drop. Scorer B gives H a 30 and a one-place drop. They agree on the more important diagnostic: H is first when enthusiasm is hidden, and the visible field changes the ordering. That is an interpretable divergence, not a null result.
Diagnose the failure
The fault in this reproduction is a scoring-rule fault: the worksheet turned a separate rollout signal into an unqualified point bonus. It did not show that enthusiastic people are wrong, or that a quiet workflow is objectively better. It showed that the ordering changed when the field changed.
The primary literature supports a narrower claim than “loud stakeholders make bad decisions.” Prioritization research has long found gaps in criteria and stakeholder coverage. A systematic review of 40 use-case prioritization approaches reported missing risks, goals, and quality-related requirements, while a 2026 review of public enterprise AI frameworks found uneven coverage of adoption and ownership and no framework-level evidence that selection outcomes improved. The 2023 systematic review and the 2026 multivocal review frame the problem, but neither tests this worksheet.
Narrative can also increase engagement without proving decision quality. In a preregistered field experiment on AI policy outreach, narratives were as effective as expert information at engaging policymakers. The result is useful here as a boundary: a compelling story can change attention or participation, but attention is not proof that the underlying workflow is the best investment. The Purdue-hosted record of that study reports the engagement result.
So the safe sentence is an inference:
When enthusiasm is visible before operational criteria are fixed, it can function as an unexamined bonus and change the shortlist.
The worksheet directly tests the rank sensitivity of that bonus. It does not test hierarchy, airtime, persuasion, or organizational politics. Any claim that political salience causes selection error needs a study that measures those mechanisms.
Reproduce and trace the failure
Use the procedure before approving a discovery sprint, prototype, or platform purchase. To reproduce the published result, keep the candidate descriptions and seven operational scores unchanged, calculate the hidden rank, add enthusiasm as a separate 1-5 field, and calculate the visible rank.
The trace below follows the two candidates that matter most. It shows the calculation that caused the ordering to change, not a claim about what happened in a real organization.
| Scoring pass | Hidden condition | Visible condition | Observed transition |
|---|---|---|---|
| Scorer A | H: 4+3+5+4+5+3+4 = 28, rank 1; B: 27, rank 2 | H: 28+1 = 29, rank 4; B: 27+3 = 30, rank 1 | H fell three places; B rose one |
| Scorer B | H: 4+4+5+4+5+4+4 = 30, rank 1; B: 28, rank 2 | H: 30+1 = 31, rank 2; B: 28+4 = 32, rank 1 | H fell one place; B rose one |
The trace isolates the trigger. Candidate descriptions did not change. The operational fields did not change. Only the visible enthusiasm field entered the total. The full candidate set and raw vectors remain in the reviewable research artifact.
- Collect six to ten candidates. Write each as a workflow result with a user, current owner, input, output, and next decision. Include one candidate that is operationally strong but not sponsored by the loudest person.
- Record the baseline before the pitch. Capture current cycle time, volume, review effort, error or exception signal, or a small sampling plan. Mark unknowns as unknowns. Do not turn guesses into low scores without saying why.
- Score seven operational fields independently. Hide enthusiasm, sponsor name, requested tool, and demo quality. Give scorers the same candidate descriptions and the same rubric.
- Publish the hidden ranking. Keep raw vectors, not just the total. A total without the vector hides the disagreement that tells you what to investigate.
- Reveal enthusiasm as a separate field. Score the energy or adoption signal from 1 to 5, define its weight, and calculate the visible ranking. Do not quietly edit the operational scores after seeing the result.
- Inspect every meaningful move. For a candidate that rises, ask whether enthusiasm represents real owner capacity or merely a persuasive story. For a candidate that falls, ask whether adoption is actually a constraint or only a weak signal.
- Set the next gate. A high rank earns a bounded test, not automatic funding. Define the evidence threshold, review owner, action boundary, stopping condition, and what result would change the recommendation.
Repair the decision rule
Keep operational value and rollout energy as separate decisions. Use the hidden operational rank to decide which workflow deserves evidence collection. Use enthusiasm to decide whether adoption work is already available, what owner must be involved, and how much change support the test needs.
| Result after the two passes | Reader decision |
|---|---|
| High hidden rank, low enthusiasm | Collect evidence or run a reversible pilot. Assign an owner before treating adoption as a blocker. |
| High hidden rank, high enthusiasm | Still require a baseline, control boundary, review owner, and stopping condition. Enthusiasm helps rollout; it does not waive the gate. |
| Low hidden rank, high enthusiasm | Investigate the adoption signal separately. Do not promote the workflow on enthusiasm alone. |
| Low hidden rank, low enthusiasm | Stop, narrow the idea, or return it to discovery. |
If you choose to combine the fields, publish the weight and run the hidden condition first. A visible score with an undisclosed bonus is not a repair because nobody can tell whether the rank came from evidence or advocacy. In this artifact, H is eligible for a bounded evidence step because its hidden score is first and its action boundary is reversible. It is not automatically funded.
Verify the repair
The repair passes when the same inputs produce an auditable hidden rank, a separately labeled enthusiasm sensitivity result, and a next action that still makes sense after enthusiasm is removed.
| Verification check | Evidence in this artifact | Result |
|---|---|---|
| Inputs are fixed | Eight candidate descriptions and raw vectors are published in research.md. | Pass |
| Hidden calculation is reproducible | H ranks first in both hidden passes at 28 and 30; B ranks second at 27 and 28. | Pass |
| Sensitivity is visible | Adding the published enthusiasm field moves H to fourth or second and B to first in both passes. | Pass |
| The repair does not hide disagreement | Scorer A and Scorer B differ in magnitude, and the limitation is stated. | Pass |
| The next action has a gate | The recommendation is evidence collection or a reversible pilot with owner, threshold, review, and stop conditions. | Pass |
Do not call the repair verified if a team edits the operational scores after seeing the enthusiasm result, changes the weight without recording it, or skips the stopping condition. Those are new decision rules and need their own trace.
The parent AI opportunity discovery guide is the place to start when the candidate list itself is unclear. The existing small-business AI use-case prioritization guide covers the broader comparison problem. If your candidates come from feedback rather than workflow observation, use the customer-feedback opportunity guide before you score them.
The failure is often an undefined done-state
I have taught product managers who moved from writing specifications to building, shipping, and automating work. The recurring failure was often that nobody could say what done meant. That is the locked F-pms observation in Marius Manolachi’s entity facts, not a prevalence claim. The AI learning page states the teaching context.
AI discovery has the same failure one level earlier. A workflow can win the room while nobody can say what would make the discovery complete. Is done a demo? A baseline? A reviewed sample? A read-only brief? A decision to stop?
Write that condition before the loudest workflow gets a budget. If the team cannot name the output, the baseline, the owner, the control boundary, and the next decision, enthusiasm is asking to be mistaken for evidence.
The worksheet is small on purpose. Run it before you build. Then let the rank change tell you where the real disagreement is.
Questions people ask next
Is stakeholder enthusiasm bad evidence for an AI workflow?
No. Enthusiasm can reveal adoption energy, an urgent constraint, or a person willing to own the change. Treat it as a separate rollout signal. Do not let it substitute for a named outcome, a baseline, a control boundary, or a credible way to learn.
How do I test whether enthusiasm changed our AI shortlist?
Score the same candidates twice. First hide enthusiasm and rank outcome, baseline, workflow fit, data readiness, control risk, adoption burden, and learning value. Then reveal enthusiasm, publish the weighting, recalculate ranks, and investigate every meaningful move.
Does the highest-ranked AI workflow deserve a pilot?
Not automatically. Ranking identifies the next decision candidate. Before a pilot, define the smallest reversible test, the evidence threshold, the review owner, the stopping condition, and what result would change the recommendation.