Field note · opportunity
Which AI Workflow Should We Improve First When Every Team Has a Backlog?
Choose the first AI workflow to improve across competing team backlogs with a scorecard for queue pressure, handoffs, evidence, and risk.

When every team brings a queue, the first mistake is to compare the size of the queues. A large backlog may be badly defined, hard to review, or impossible to change safely.
I taught product managers who moved from writing specifications to building, shipping, and automating work. The recurring lesson was simple: before choosing what to build, someone has to say what “done” means. That is the same missing decision behind many cross-team AI backlogs.
This page is the narrow decision inside the broader AI opportunity discovery guide. If you still have a list of slogans rather than workflows, start with How to Prioritize AI Use Cases in a Small Business first.

The first workflow should relieve a shared bottleneck
Choose the eligible workflow that is delaying work now, sits on a handoff between teams, and can produce reviewable evidence through a reversible first change.
That choice is narrower than “the most valuable AI opportunity.” You are choosing the first place to learn and improve. OpenAI Academy's workflow matrix points to frequency, repeatability, and reach as useful value signals, while process complexity, input and output readiness, governance, and system dependencies affect effort. Its matrix is a good starting point, but a cross-team backlog needs one extra question: how much waiting does this workflow create for other people?
The rule I use is:
Start with the workflow that has the highest current queue pressure and shared handoff reach, provided one person owns the result, the result can be checked, the first action is safe, and the owner can review it.
That is a decision rule, not a claim that shared work is always the best AI use case. A high-consequence workflow with no safe first move waits, even when its backlog is painful.
Run four gates before scoring
Do not give points to a workflow that cannot yet be operated. Check these conditions first.
| Gate | Pass condition | What to do if it fails |
|---|---|---|
| Workflow boundary | You can name the trigger, output, and current owner. | Rewrite the idea around one workflow or observe it first. |
| Outcome check | The owner can define what gets passed, edited, or rejected. | Collect examples, define a baseline, or write a review rubric. |
| Safe first move | The first version can stay read-only, draft-only, classification-only, recommendation-only, or approval-gated. | Narrow the action or move it to specialist review. |
| Review capacity | A named owner has protected time to review the pilot. | Wait for capacity or reduce the pilot slice. |
These are gates, not low scores. A missing owner cannot be compensated for by a large theoretical benefit.
Microsoft's business-envisioning guidance asks teams to define the problem, business objective, measurement of success, and accountability before comparing business, experience, and technology viability. Microsoft's guidance is written for ISVs, but the questions fit a team backlog: what problem is changing, who cares, how will you know, and who carries the result?
NIST's AI RMF makes the same sequence more explicit for risk. Its Map function establishes context and informs an initial go or no-go decision. Its guidance also says roles and responsibilities should be clear. The NIST AI RMF FAQ states that the framework is voluntary, not a mandatory scorecard. Use the NIST AI RMF Core to expose missing context before a point total hides it.
Use the backlog-first scorecard
Score only workflows that pass all four gates. The complete worksheet is below.
| Dimension | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Queue pressure Q | None or unknown | Occasional delay | Recurring queue | Visible queue blocks downstream work now |
| Handoff reach H | Isolated output | One handoff | Two roles depend on it | The output feeds three or more roles or teams |
| Evidence density E | No cases or check | Anecdotes only | Representative cases and a plausible check | Cases, baseline, expected result, and reviewer exist |
| Reversibility R | Consequence cannot be contained | Approval is possible but weak | Read-only, draft, or approval-gated path is credible | Failure is detectable, containable, reversible, and escalatable |
| Owner capacity O | No review owner | Owner named, no protected time | Owner has bounded review time | Owner and review time are protected for the pilot |
| Dependency cost D | Known and low | One manageable dependency | Several dependencies or access questions | Major integration, policy, data, or procurement unknowns |
Calculate:
Backlog-first score = (2 × Q) + H + E + R + O - D
The double weight on queue pressure is deliberate. This question is about choosing where to improve first when work is already waiting. It is not a universal measure of AI value, and it is not a forecast of return.
Use these reading bands:
- 13-18: first candidate, provided no gate is invalidated.
- 9-12: prepare it or compare it with a stronger candidate.
- 0-8: defer, observe, redesign, or reject.
If no candidate reaches 13, do not lower the threshold just to start something. Improve the best candidate's evidence, boundary, or dependency plan first.
A worked comparison across four team backlogs
The numbers below are fictional inputs. They show how the artifact works, not what usually happens in a company.
| Candidate workflow | Q | H | E | R | O | D | Score | First reading |
|---|---|---|---|---|---|---|---|---|
| Support escalation intake and draft routing | 3 | 3 | 2 | 2 | 3 | 1 | 15 | First candidate |
| Finance invoice exception review | 2 | 1 | 3 | 3 | 2 | 2 | 11 | Prepare or second |
| Product release-note assembly | 1 | 2 | 2 | 3 | 2 | 1 | 10 | Prepare or second |
| Hiring shortlist ranking | 2 | 1 | 1 | 0 | 2 | 1 | 7 | Redesign before scoring |
Support wins because its queue is active and its output crosses team boundaries. Finance has better evidence and safer reversibility, so it may become the next candidate once its system dependency is clearer. Release-note assembly is safe but does not relieve as much waiting. Hiring shortlist ranking does not get rescued by a busy backlog because its proposed action is not a safe first move.
This is the important distinction: a score orders eligible work. It does not turn an unsafe workflow into an eligible one.
Break ties with evidence, then dependency cost
When two workflows have the same score, choose the one that can close its evidence gap faster. If both have similar evidence, choose the one with fewer unresolved dependencies.
Do not break a tie with the louder sponsor or the larger promised market. Microsoft's guidance treats success measurement and accountability as inputs to prioritization, while NIST's Measure function calls for documented testing before deployment and during operation. NIST's Measure guidance supports asking which candidate can produce credible evidence under conditions close to use.
Use this tie-break table:
| Question | Winner if yes |
|---|---|
| Can the owner produce representative cases this week? | That workflow gets the tie. |
| Can a reviewer compare the output with a source or rubric? | That workflow gets the tie. |
| Can the first version stay read-only, draft-only, or approval-gated? | That workflow gets the tie. |
| Does it avoid an integration, policy, or procurement dependency? | That workflow gets the tie. |
| Does it create an artifact another workflow can reuse? | Use it as a sequencing advantage, not as permission to ignore risk. |
If the tie still holds, run a short observation period and record the result instead of arguing from preference.
High backlog pressure is not permission to automate
Wait when the workflow affects employment, payments, access, safety, health, legal position, or another high-impact outcome and the first action cannot be contained. Also wait when the data boundary is unknown, the result cannot be checked before a consequence occurs, or the owner cannot review the cases.
NIST's guidance says risk management should reflect context, risk tolerance, impact, likelihood, and available resources. Its Generative AI Profile also makes clear that implementation depends on the use case and the organization's risk tolerance and resources. The NIST Generative AI Profile is a useful reminder that a generic “best first workflow” answer would be dishonest.
NIST also keeps viable non-AI alternatives in the management conversation. If a deterministic rule, queue redesign, clearer source of truth, or ordinary automation fixes the bottleneck more safely, improve that instead. The question is which workflow to improve first, not which workflow must receive a language model.
Make the first improvement smaller than the backlog
Once one workflow wins, do not automate the whole queue. Choose the smallest slice that can show whether the bottleneck is real and the output is useful.
For the support example, that might mean classifying one request category and drafting a suggested route for a support owner to approve. It does not mean sending replies, issuing refunds, changing account records, or promising delivery dates.
Write five lines before building:
- Current trigger and output.
- One owner and one reviewer.
- Allowed data and forbidden data.
- First action the system may take.
- Stop condition, such as unsupported claims, missed escalations, or review taking longer than the current workflow.
NIST's Measure function calls for documented test methods and results, and its Manage function includes response, recovery, monitoring, and responsibility for disengaging a system that behaves inconsistently with intended use. That makes the stop condition part of the first improvement, not paperwork after launch.
The existing guide on whether a business process is ready for AI automation can help with the next readiness check. If the workflow passes, keep the first version visible and approval-gated until the owner has evidence to widen its action boundary.
The decision to take
Pick the highest-scoring workflow that passes all four gates, with queue pressure weighted twice because the organization is already carrying waiting work. If two candidates tie, choose the one with the smaller evidence gap and lower dependency cost. If none passes, improve the workflow definition or evidence before choosing a tool.
That gives every team a fair hearing without pretending every backlog deserves equal priority. It also leaves a record of why the first workflow won and what would make you change course.
If you want guided help turning this worksheet into a real decision, work with Marius Manolachi through AI consulting or tutoring.