Field note · opportunity

How to Separate AI Opportunity Size From Automation Fantasy

A worksheet for separating AI exposure from usable capacity by checking outcomes, task boundaries, review, exceptions, and cash-realizable value.

12 minute read
  • AI strategy
  • AI opportunities
Illustration of a worksheet separating theoretical AI exposure from usable workflow value

Most AI opportunity claims begin with a percentage. A team hears that a role is highly exposed, or that a tool could remove half of a task, and jumps to a business case. The missing step is the workflow.

The useful question is not “How much of this job can AI touch?” It is “What observable result can improve, under whose review, with which exceptions, and what value remains after the controls?” That distinction is the basis of the AI opportunity discovery parent guide. It also connects this article to the more operational question of how to prioritize AI use cases in a small business.

The result is a decision change, not a bigger savings number

The result of the observation is simple: an AI opportunity becomes credible only after the claim survives outcome definition, workflow boundaries, review, exceptions, and value realization. The process often changes the category of the idea before it changes a savings estimate.

The table below is the sourceable atom of this article. It records three bounded contexts from Marius Manolachi's teaching and product work. The observations are qualitative. They are not a representative sample and do not support a percentage.

Observed contextLocked observationWhat the observation changesDecision after the observation
Product managers learning to build and shipThe people Marius taught moved from writing specifications to building, shipping, and automating work around products. The recurring failure was an undefined “done,” not a missing model.A claimed automation gain is not yet an opportunity if nobody can name the external result that proves it worked.Define the outcome and its acceptance check before comparing tools.
ChatGPT workshop at OrangeThe workshop started from the work attendees already did, not from an agent label.Opportunity discovery belongs at the task boundary. A broad “use AI at work” claim is too large to evaluate.Map one repeated task, its inputs, outputs, review, and consequence of error.
TryUncle product workTryUncle watches the screen and annotates it live. That makes latency and human approval product constraints, not afterthoughts.Apparent autonomy is not the same as usable capacity. Waiting, checking, and approval remain part of the product.Start with a supervised pilot unless the workflow can prove a safe outcome without approval.

The table does not prove that one type of workflow is always better. It gives a practical reclassification rule: if the claim disappears when you name the outcome, task boundary, review, or control point, you have an investigation to do, not an automation case to fund.

Anonymized workflow observation ledger with outcome, review, and exception fields

Method and sample: three bounded contexts, no invented timing

This is a bounded qualitative observation, not a benchmark. I selected three documented contexts that the entity facts authorize: teaching product managers to build and ship, leading a ChatGPT workshop at Orange, and building TryUncle. I kept each observation at the strength of the source. I did not infer a client, task duration, savings rate, population size, or success rate.

For each context, I asked five questions:

  1. What work or product behavior is actually described?
  2. What result would make the work useful to someone outside the AI system?
  3. Which human review or approval remains necessary?
  4. Which exception would change the product boundary?
  5. Does the evidence justify investigating, redesigning, piloting under supervision, or stopping?

The sample is deliberately small because the permitted firsthand record is small. Its value is not statistical generalization. Its value is a check against a common reasoning error: treating a broad capability or an attractive demo as a measured business result.

The method also separates firsthand observation from sourced context. The OECD distinguishes exposure, complementarity, and automation potential rather than treating them as one measure. Eloundou and colleagues estimate potential task exposure, but their paper does not predict adoption timing or a particular company's realized productivity. Those sources help define the question. They are not substituted for the observation table.

Exposure is a screening signal, not usable capacity

Exposure tells you where AI capabilities may touch work. It does not tell you how much capacity a team can use after review, exceptions, coordination, and error handling. Use exposure to choose what to inspect, never to close the business case.

Eloundou and colleagues estimate that around 80% of the US workforce could have at least 10% of tasks affected by language-model capabilities and about 19% could have at least 50% affected. The paper is about potential task exposure, not a promise that those tasks will be automated or that the affected time becomes available. Read the paper before turning the estimate into a local forecast.

The Philadelphia Fed makes a similar boundary explicit: its measure is an intermediate exposure measure that assumes access to an LLM without fully implemented complementary technologies. It also says exposure does not necessarily measure susceptibility to job loss. Its report is useful precisely because it narrows what the number can mean.

The exception is a very low-risk task where a person already reviews every output and the only intended benefit is faster drafting. Even there, record the review step. The capacity is not “the model's output time”; it is the time from input to an acceptable, usable result.

Observe the workflow before choosing automation

Start with one repeated task and its external result. Do not start with an agent, a model, or a role-level exposure percentage.

The first observation pass should capture the task boundary, input, output, frequency, cycle time, review, exception, handoff, and consequence of error. The point is not to create a perfect time-and-motion study. The point is to stop a headline claim from hiding the work that makes the result safe and useful.

Use this ledger:

FieldWhat to recordWhy it changes the opportunity decision
User and jobWho acts, and what decision or deliverable they are trying to completeA role is too broad to automate as one unit
Trigger and inputWhat starts the task and which source material is availableMissing or unstable inputs can dominate the design
Output and external outcomeWhat the system produces and what a person checks in the worldA plausible response is not proof of a useful result
Frequency and cycle timeHow often the task occurs and the full path from input to accepted outputRare work may not repay integration cost
ReviewWho checks, what they inspect, and what approval meansReview time belongs in the opportunity calculation
ExceptionsMissing data, ambiguity, escalation, refusal, or a risky caseExceptions can move a task from automation to assistance
HandoffsPeople, systems, or queues that receive the resultCoordination can erase local time savings
Consequence of errorWhat happens if the result is wrong, late, or incompleteRisk sets the control boundary and pilot size

This is where the Orange observation matters. Starting from existing work gives the team a real input and output to inspect. The exception is a genuinely new product job with no current workflow. In that case, write a proposed outcome and test it with users before estimating automation value.

Gross time, usable capacity, net time, and cash value are different claims

Keep four calculations separate. Combining them into one “AI savings” number makes it impossible to see which assumption failed.

gross theoretical time
  = eligible task volume × current task time

usable capacity
  = gross theoretical time × work that the system can assist

net time after controls
  = usable capacity - review time - exception time - handoff time - rework time

cash-realizable value
  = net time that the team can actually redeploy or avoid paying for
    × an agreed value rate

The formula is a worksheet structure, not a claim that every term should become a precise number on day one. A team can mark a field unknown and run the next observation. It must not silently turn unknown into zero.

The QJE field study by Brynjolfsson, Li, and Raymond reports a 15% average productivity increase for 5,172 customer-support agents using a conversational assistant, with heterogeneous effects. That is a useful example of a measured result in a particular setting. It is not a shortcut to the value of another workflow. The study shows why context and outcome definition matter.

Worksheet separating gross theoretical time, usable capacity, net time, and cash-realizable value

If a workflow is low-risk and fully reviewable, you may stop at usable capacity while deciding whether to pilot it. If the business case depends on headcount reduction, revenue, or avoided cost, continue to cash-realizable value. Time that cannot be redeployed, sold, or used to avoid a real cost is not cash value.

Review and exceptions can change the product category

Review is not a nuisance deduction. It can be the reason the system should be an assistant, a fixed workflow, or a supervised pilot rather than an autonomous agent.

The TryUncle observation makes this concrete. A system that watches a screen and annotates it live must meet a timing expectation, and a person may still need to approve what happens next. Those constraints belong in the product promise. If the promise says “the agent completes the task,” but the observed workflow requires a person to inspect every step, the honest promise is closer to “the agent prepares a next action for approval.”

Classify each exception by what it does to the boundary:

Exception patternReclassificationSafer next step
The output is useful but needs a quick human checkAssistanceDraft or recommend, with explicit approval
The input is often incomplete or ambiguousTriageAsk a clarifying question or route to a person
The cost of a wrong result is highControl-heavy workflowRead-only or supervised pilot with a veto
The result changes external stateAction systemAdd authorization, idempotency, audit, and outcome verification
The task has no stable observable outcomeUnready opportunityDefine the outcome or stop estimating savings

An exception does not automatically kill the idea. It tells you what the product is. The principal exception is a low-consequence task where review is cheap and visible. In that case, a supervised workflow may create real value even when full automation would be fantasy.

Use a decision artifact to choose investigate, redesign, or stop

End the first pass with a decision, not a score. The artifact should tell the team what evidence to collect next and what would block expansion.

Copy this compact version into a working document:

OPPORTUNITY REALITY CHECK

Workflow and user:
External outcome that proves usefulness:
Current input and output:
Current cycle, review, exception, and handoff:

AI-assisted boundary:
Human approval required:
Unacceptable failure:
Outcome check after the model response:

Gross theoretical time: known / unknown
Usable capacity: known / unknown
Net time after controls: known / unknown
Cash-realizable value: known / unknown

Decision:
[ ] investigate the baseline
[ ] redesign as assistance or a fixed workflow
[ ] run a supervised pilot
[ ] stop the opportunity for now

Expansion veto:
Evidence that would change the decision:
Owner of the next observation:

The artifact is useful because it makes a missing field visible. It is not a named framework claiming universal authority. It is a decision record built from the observed contexts in this article and from the distinctions in the primary sources.

Decision artifact with investigate, redesign, supervised pilot, and stop paths

Use “investigate” when the outcome or baseline is unclear. Use “redesign” when the task is valuable but review, ambiguity, or handoffs rule out the original automation claim. Use “supervised pilot” when the outcome is observable and the control path is safe enough to test. Use “stop” when the outcome cannot be checked or the consequence of error exceeds the available controls.

What field studies can and cannot prove

Field studies can show realized outcomes in a defined setting. They cannot make every workflow look like that setting.

The AEA field experiment across 66 firms and 7,137 knowledge workers reports that, among treated workers who used the tool, email time fell by two hours per week in the second half of a six-month experiment. The study is valuable evidence about a particular intervention, population, and measurement design. It is not evidence that a founder's workflow has the same baseline, adoption, review burden, or value rate.

The OECD likewise separates exposure to generative AI from potential complementarity and automation. Its analysis is a reason to ask better questions, not a local opportunity calculation.

If your team has a comparable field study, use its result with the same discipline: name the population, intervention, outcome, time window, and boundary. If it does not, use it as context and run the workflow ledger.

Limitations and what I still do not know

This article has a narrow evidence boundary. The three observations are locked firsthand contexts, not a random sample of companies or workflows. They support a decision procedure and a taxonomy of reclassification. They do not establish average savings, an adoption rate, a payback period, or a causal effect.

I also do not know, from the permitted record, how long any one task took before and after assistance, how often a particular exception occurred, or whether a team converted recovered time into cash value. Those are not missing details to fill with a plausible estimate. They are the next measurements.

The primary-source studies have their own boundaries. Their populations, tools, tasks, and outcomes differ. Model behavior, product interfaces, and policies can change, so any current tool recommendation needs a fresh review. A future version of this article should add a dated workflow ledger only when the observation and raw measurements are actually preserved.

The practical conclusion is therefore modest: use exposure to choose what to inspect, use the ledger to define the opportunity, use review and exceptions to set the product boundary, and use cash-realizable value only when the organization can explain what will change.

How to apply the worksheet this week

Choose one repeated task, write the external outcome in one sentence, and observe the complete path from input to accepted result. Do not start by choosing an agent. Start by recording what a person does, what the person checks, and what happens when the input is incomplete.

Then run one safe assisted version with no external side effect. Record the output, the review, the exception, and the work that remains. If the result is promising, calculate the four value layers separately. If the result is unclear, keep the decision at “investigate.” If the controls are doing most of the work, redesign the promise around assistance.

For the next decision in the sequence, compare this worksheet with how to tell if a business process is ready for AI automation. If the problem is demand rather than workflow value, use how to validate demand for an AI feature before building it.

Marius Manolachi helps existing people build AI products on their own work as an AI consultant and AI tutor. The useful starting point is one real workflow, one observable outcome, and one honest control boundary. The article is complete without a larger automation promise.