Field note · opportunity
Why AI Opportunity Scores Change After Observation
AI opportunity scores change when observation replaces a process story with the operator's real steps, exceptions, data access, and review burden.

An opportunity score often changes the moment someone watches the work instead of hearing a description of it. The workflow owner remembers the intended path. The operator reveals the path that survives missing data, handoffs, approvals, and exceptions.
That difference matters because the score is supposed to guide a decision, not reward the most optimistic description. The practical response is to keep the scoring rubric fixed, replace assumptions with observed evidence, and record why each dimension moved.
The first score describes a story, not the work
The first score is useful for choosing what to inspect, but it is not evidence that the workflow is ready for AI. It usually combines a workflow owner's memory, a proposed use case, and assumptions about data, ownership, and review.
OpenAI Academy's workflow evaluator starts with the work problem and asks teams to surface manual workarounds, unclear ownership, duplicated steps, approval bottlenecks, and process gaps. The National AI Centre's opportunity guide likewise separates business impact from AI fit. Those are prompts for investigation, not proof that the proposed workflow has a clean automation boundary.
Use the first score as a hypothesis. Write one evidence note and one assumption note for every dimension. If a score has no evidence note, it is a question for observation.
| Before observation | What it means | What observation must test |
|---|---|---|
| “The task is repetitive” | The owner expects recurring work with a stable shape | Whether cases follow the same path and where they branch |
| “The data is available” | The owner knows where some inputs live | Whether the operator can access the right fields at the moment of decision |
| “Review will be light” | The proposed system appears low-risk | Which cases require judgment, approval, or a second system check |
| “The process is standard” | The documented path looks complete | Which workarounds and handoffs are absent from the documentation |
The exception is a workflow that cannot be observed with permission, or where the observed cases are not representative. In that situation, keep the score provisional. Do not turn missing access into a high score.
Observation changes the dimensions that depend on hidden work
Observation changes a score when it reveals a fact that the initial description could not establish. The most important facts are not the visible clicks. They are the conditions that determine whether an AI system can act safely and whether a person can review its result.
Capture five kinds of evidence in the same order every time:
- Recurring demand: how often the work arrives and whether cases have a repeatable shape.
- Exception load: where cases leave the normal path, including workarounds and rework.
- Data access: which systems, fields, attachments, and permissions the operator needs.
- Process stability: whether the steps, owner, policy, and handoffs stay consistent.
- Review burden: where a person must interpret context, approve an action, or catch an error.
This is narrower than “watch the process.” It tells the observer what to write down. A process-mining overview makes the same distinction from another angle: a model built from real activity can differ from a hand-built model of the ideal process. NIST's AI RMF Core also calls for context, assumptions, limitations, business value, and human-oversight requirements before a go or no-go decision.
The exception is a purely informational task with no external action and no sensitive input. Observation still helps, but review burden and permission checks may be lighter. Do not use that exception to generalize a low-risk score to a workflow that changes records, sends commitments, or makes consequential decisions.
Use one fixed scorecard before and after observation
Use the same five dimensions and the same 1-to-5 meanings before and after the observation. A moving rubric can make any opportunity appear to improve.
For this worksheet, 1 means a weak candidate and 5 means a strong candidate. Exception load and review burden are scored in the favorable direction: a 5 means few exceptions or little judgment-heavy review.
| Dimension | 1 | 3 | 5 |
|---|---|---|---|
| Recurring demand | One-off or unpredictable | Recurring but variable | Frequent with a repeatable shape |
| Exception load | Many unique exceptions | Known branches need handling | Few exceptions with clear boundaries |
| Data access | Key inputs are missing or inaccessible | Inputs are split across sources | Required inputs are available and permissioned |
| Process stability | Steps, owner, or policy change often | Stable core with ownership gaps | Stable steps, owner, and handoffs |
| Review burden | Every case needs substantial judgment | Edge cases need review | Most cases have a clear, low-risk review path |
Add the five scores for a total from 5 to 25. The thresholds below are part of the artifact, not a claim about an industry-wide benchmark:
| Total and veto check | Decision | Smallest next action |
|---|---|---|
| 20-25, with no dimension below 3 | Test now | Run a bounded, reversible pilot on representative cases |
| 14-19, with no dimension below 2 | Validate | Shadow the workflow or run a read-only test before automating an action |
| 9-13, or any dimension at 1 | Sequence | Fix the process, ownership, or data access before adding AI |
| 5-8, or review cannot be bounded | Avoid for now | Keep the decision human-owned and record what would need to change |
Treat a score as a decision aid, not a permission slip. NIST's Measure playbook recommends documenting why metrics were selected, their acceptable limits, and how pre- and post-deployment performance will be compared. That is why the worksheet preserves both the number and the reason.
A worked pre/post table makes the score change explainable
The following is an illustrative example constructed to show how the artifact works. It is not a report of an observed company, a measured sample, or a statistic. Replace it with permissioned workflow notes before using the result in a real investment decision.
The stated workflow is “use AI to classify invoice exceptions and route them for resolution.” The initial description makes it look like a good pilot. The observation record exposes the work that the description left out.
| Dimension | Before score | Observed evidence | After score | Why the recommendation changes |
|---|---|---|---|---|
| Recurring demand | 4 | Cases recur, but the reason for an exception varies by supplier | 3 | A classifier can help, but the categories need a bounded fixture set |
| Exception load | 4 | The operator leaves the normal path for missing purchase orders and policy questions | 2 | The first pilot needs explicit abstention and escalation paths |
| Data access | 4 | The operator checks an enterprise system, an attachment, and an email thread | 2 | A single prompt does not have the decision context |
| Process stability | 4 | Ownership changes when an exception crosses finance and procurement | 2 | Routing logic depends on an ownership rule that is not written down |
| Review burden | 3 | The operator verifies supplier identity and policy before routing | 2 | Human review remains part of every consequential case |
| Total | 19 | Observation replaces assumptions with workflow evidence | 11 | Sequence process and data work, then validate a read-only classifier |
The important result is not that the total fell by eight points. Those numbers are illustrative. The useful result is the chain of reasons: split context lowered data access, hidden branches lowered exception fit, and ownership changes lowered process stability. The decision changes from “test an automated router” to “document ownership, assemble representative cases, and validate classification without external action.”
That distinction prevents a common mistake: treating a lower score as a rejection of AI rather than a sequencing decision. The National AI Centre guide lists process changes, templates, rules-based automation, training, and integration as possible responses when business impact is high but AI fit is low.

Run the paired decision procedure before committing budget
The smallest reliable process is a paired assessment with a fixed rubric and a visible decision rule. Complete it before a business case treats the opportunity score as settled.
- Bound the workflow. Name the trigger, input, output, owner, systems, and external effect. “Improve finance” is too broad. “Classify invoice exceptions before a finance reviewer routes them” is bounded.
- Score the stated workflow. Use the five dimensions above. For each score, record what is known and what is assumed. Do not fill an evidence column with a stakeholder's confidence.
- Observe representative cases. Watch the operator complete normal and non-normal cases. Record handoffs, workarounds, rework, missing data, approvals, and human review. Remove identifying details from the published artifact and retain permission evidence separately.
- Re-score without changing the rubric. Put the observed fact next to every changed number. If a score did not change, keep the evidence that confirmed the original assumption.
- Apply the threshold and veto. Choose test, validate, sequence, or avoid. Then write one smallest next action, one owner, and one condition that would change the decision.
The procedure fails if the observer only watches the happy path, changes the scale after seeing inconvenient evidence, or averages away a veto. A low total can still hide a dangerous single dimension. Keep the veto visible.
If the workflow is sensitive, consequential, or not permissioned for observation, stop at a documented assumption set and choose validation or sequence. The score is not ready for a build decision. For the broader prioritization context, start with how to prioritize AI use cases in a small business, then compare the result with how to tell if a business process is ready for AI automation before deciding whether an agent is appropriate.
The article is complete without a sales call. If the team wants help turning its own observation notes into a bounded build decision, Marius Manolachi's AI consulting and tutoring work follows the same capability goal: leave the people doing the work able to explain and improve the system themselves.