Field note · opportunity

Which AI Workflow Should We Improve First When Every Team Has a Backlog?

Choose the first AI workflow to improve across competing team backlogs with a scorecard for queue pressure, handoffs, evidence, and risk.

9 minute read
  • AI opportunities
  • Workflow prioritization
  • Small teams
Illustration of a team choosing one AI workflow from competing backlogs

When every team brings a queue, the first mistake is to compare the size of the queues. A large backlog may be badly defined, hard to review, or impossible to change safely.

I taught product managers who moved from writing specifications to building, shipping, and automating work. The recurring lesson was simple: before choosing what to build, someone has to say what “done” means. That is the same missing decision behind many cross-team AI backlogs.

This page is the narrow decision inside the broader AI opportunity discovery guide. If you still have a list of slogans rather than workflows, start with How to Prioritize AI Use Cases in a Small Business first.

Illustration of the backlog-first AI workflow scorecard with four gates and six dimensions

The first workflow should relieve a shared bottleneck

Choose the eligible workflow that is delaying work now, sits on a handoff between teams, and can produce reviewable evidence through a reversible first change.

That choice is narrower than “the most valuable AI opportunity.” You are choosing the first place to learn and improve. OpenAI Academy's workflow matrix points to frequency, repeatability, and reach as useful value signals, while process complexity, input and output readiness, governance, and system dependencies affect effort. Its matrix is a good starting point, but a cross-team backlog needs one extra question: how much waiting does this workflow create for other people?

The rule I use is:

Start with the workflow that has the highest current queue pressure and shared handoff reach, provided one person owns the result, the result can be checked, the first action is safe, and the owner can review it.

That is a decision rule, not a claim that shared work is always the best AI use case. A high-consequence workflow with no safe first move waits, even when its backlog is painful.

Run four gates before scoring

Do not give points to a workflow that cannot yet be operated. Check these conditions first.

GatePass conditionWhat to do if it fails
Workflow boundaryYou can name the trigger, output, and current owner.Rewrite the idea around one workflow or observe it first.
Outcome checkThe owner can define what gets passed, edited, or rejected.Collect examples, define a baseline, or write a review rubric.
Safe first moveThe first version can stay read-only, draft-only, classification-only, recommendation-only, or approval-gated.Narrow the action or move it to specialist review.
Review capacityA named owner has protected time to review the pilot.Wait for capacity or reduce the pilot slice.

These are gates, not low scores. A missing owner cannot be compensated for by a large theoretical benefit.

Microsoft's business-envisioning guidance asks teams to define the problem, business objective, measurement of success, and accountability before comparing business, experience, and technology viability. Microsoft's guidance is written for ISVs, but the questions fit a team backlog: what problem is changing, who cares, how will you know, and who carries the result?

NIST's AI RMF makes the same sequence more explicit for risk. Its Map function establishes context and informs an initial go or no-go decision. Its guidance also says roles and responsibilities should be clear. The NIST AI RMF FAQ states that the framework is voluntary, not a mandatory scorecard. Use the NIST AI RMF Core to expose missing context before a point total hides it.

Use the backlog-first scorecard

Score only workflows that pass all four gates. The complete worksheet is below.

Dimension0123
Queue pressure QNone or unknownOccasional delayRecurring queueVisible queue blocks downstream work now
Handoff reach HIsolated outputOne handoffTwo roles depend on itThe output feeds three or more roles or teams
Evidence density ENo cases or checkAnecdotes onlyRepresentative cases and a plausible checkCases, baseline, expected result, and reviewer exist
Reversibility RConsequence cannot be containedApproval is possible but weakRead-only, draft, or approval-gated path is credibleFailure is detectable, containable, reversible, and escalatable
Owner capacity ONo review ownerOwner named, no protected timeOwner has bounded review timeOwner and review time are protected for the pilot
Dependency cost DKnown and lowOne manageable dependencySeveral dependencies or access questionsMajor integration, policy, data, or procurement unknowns

Calculate:

Backlog-first score = (2 × Q) + H + E + R + O - D

The double weight on queue pressure is deliberate. This question is about choosing where to improve first when work is already waiting. It is not a universal measure of AI value, and it is not a forecast of return.

Use these reading bands:

  • 13-18: first candidate, provided no gate is invalidated.
  • 9-12: prepare it or compare it with a stronger candidate.
  • 0-8: defer, observe, redesign, or reject.

If no candidate reaches 13, do not lower the threshold just to start something. Improve the best candidate's evidence, boundary, or dependency plan first.

A worked comparison across four team backlogs

The numbers below are fictional inputs. They show how the artifact works, not what usually happens in a company.

Candidate workflowQHERODScoreFirst reading
Support escalation intake and draft routing33223115First candidate
Finance invoice exception review21332211Prepare or second
Product release-note assembly12232110Prepare or second
Hiring shortlist ranking2110217Redesign before scoring

Support wins because its queue is active and its output crosses team boundaries. Finance has better evidence and safer reversibility, so it may become the next candidate once its system dependency is clearer. Release-note assembly is safe but does not relieve as much waiting. Hiring shortlist ranking does not get rescued by a busy backlog because its proposed action is not a safe first move.

This is the important distinction: a score orders eligible work. It does not turn an unsafe workflow into an eligible one.

Break ties with evidence, then dependency cost

When two workflows have the same score, choose the one that can close its evidence gap faster. If both have similar evidence, choose the one with fewer unresolved dependencies.

Do not break a tie with the louder sponsor or the larger promised market. Microsoft's guidance treats success measurement and accountability as inputs to prioritization, while NIST's Measure function calls for documented testing before deployment and during operation. NIST's Measure guidance supports asking which candidate can produce credible evidence under conditions close to use.

Use this tie-break table:

QuestionWinner if yes
Can the owner produce representative cases this week?That workflow gets the tie.
Can a reviewer compare the output with a source or rubric?That workflow gets the tie.
Can the first version stay read-only, draft-only, or approval-gated?That workflow gets the tie.
Does it avoid an integration, policy, or procurement dependency?That workflow gets the tie.
Does it create an artifact another workflow can reuse?Use it as a sequencing advantage, not as permission to ignore risk.

If the tie still holds, run a short observation period and record the result instead of arguing from preference.

High backlog pressure is not permission to automate

Wait when the workflow affects employment, payments, access, safety, health, legal position, or another high-impact outcome and the first action cannot be contained. Also wait when the data boundary is unknown, the result cannot be checked before a consequence occurs, or the owner cannot review the cases.

NIST's guidance says risk management should reflect context, risk tolerance, impact, likelihood, and available resources. Its Generative AI Profile also makes clear that implementation depends on the use case and the organization's risk tolerance and resources. The NIST Generative AI Profile is a useful reminder that a generic “best first workflow” answer would be dishonest.

NIST also keeps viable non-AI alternatives in the management conversation. If a deterministic rule, queue redesign, clearer source of truth, or ordinary automation fixes the bottleneck more safely, improve that instead. The question is which workflow to improve first, not which workflow must receive a language model.

Make the first improvement smaller than the backlog

Once one workflow wins, do not automate the whole queue. Choose the smallest slice that can show whether the bottleneck is real and the output is useful.

For the support example, that might mean classifying one request category and drafting a suggested route for a support owner to approve. It does not mean sending replies, issuing refunds, changing account records, or promising delivery dates.

Write five lines before building:

  1. Current trigger and output.
  2. One owner and one reviewer.
  3. Allowed data and forbidden data.
  4. First action the system may take.
  5. Stop condition, such as unsupported claims, missed escalations, or review taking longer than the current workflow.

NIST's Measure function calls for documented test methods and results, and its Manage function includes response, recovery, monitoring, and responsibility for disengaging a system that behaves inconsistently with intended use. That makes the stop condition part of the first improvement, not paperwork after launch.

The existing guide on whether a business process is ready for AI automation can help with the next readiness check. If the workflow passes, keep the first version visible and approval-gated until the owner has evidence to widen its action boundary.

The decision to take

Pick the highest-scoring workflow that passes all four gates, with queue pressure weighted twice because the organization is already carrying waiting work. If two candidates tie, choose the one with the smaller evidence gap and lower dependency cost. If none passes, improve the workflow definition or evidence before choosing a tool.

That gives every team a fair hearing without pretending every backlog deserves equal priority. It also leaves a record of why the first workflow won and what would make you change course.

If you want guided help turning this worksheet into a real decision, work with Marius Manolachi through AI consulting or tutoring.