Field note · opportunity

Which Operations Task Should a Founder Baseline Before an AI Workshop?

Baseline the recurring operations task that has a clear owner, observable done condition, measurable friction, safe review, and a reversible first test.

10 minute read
  • AI opportunity research
  • AI workshops
Illustration of a founder comparing recurring operations tasks before an AI workshop

The expensive mistake is not choosing the wrong model. It is paying for an AI workshop with three vague tasks and no way to tell which one deserves a first test.

Start with a real recurring operation. Record what triggers it, who owns it, how often it happens, what completion means, and where the work breaks. Then compare that task with at least two alternatives before anyone recommends an agent, workflow, or tool.

Illustration of a founder comparing recurring operations tasks before an AI workshop

The task to baseline first

Baseline the task that repeats often enough to create observable friction, has one accountable owner, ends in a checkable output, and can be tested with a human review gate. If a task fails any of those conditions, it is not the first baseline. It is a discovery problem.

That answer is narrower than “automate the most expensive process.” Cost matters, but a high-cost task with changing inputs, several decision-makers, or an unclear outcome may be a poor first test. A smaller task with stable inputs and a visible output can produce better evidence about whether a workshop changes daily work.

The rule came from a five-source audit described below. The sources agree on starting with real work, recording frequency and friction, naming owners, preserving a baseline, and controlling risk. They do not provide a founder dataset that identifies one universal task category. The defensible conclusion is therefore a selection rule, not a claim that every founder should automate email, support, reporting, or any other named function.

What the five-source audit actually found

The audit sampled the five primary sources already attached to this article. I extracted the fields and controls each source explicitly supplied, then recorded the decision consequence for a founder preparing for a workshop.

SourceObserved contributionWhat it does not establishDecision consequence
OpenAI Academy workshop playbookStart from recurring workflows, real examples, frequency, owners, inputs, outputs, and a small follow-up testIt does not rank a founder's actual tasksBring real task records, not broad AI ideas
OpenAI Academy workflow readiness resourceCompare frequency, reach, friction, rework, stability, dependencies, ownership, and readinessIt does not provide a founder sample or measured savingsAdd a veto for unclear ownership, unstable process, or unresolved dependencies
OpenAI Academy evidence guidancePreserve baseline, unit, period, method, evidence label, and limitation; separate measured, observed, reported, estimated, planned, and unknownIt does not show that an AI intervention improved a taskLabel every worksheet field by evidence status
Microsoft use-case blueprintsPair KPIs with pre-go-live baselines such as volume, median or P90 handling time, cycle time, rework, and hoursIts numerical examples are illustrative, not a founder's observed resultRecord distributions and units before workshop claims about value
NIST AI RMFOrganize risk work through Govern, Map, Measure, and ManageIt does not select a business processKeep human authority, risk, measurement, and escalation in the first-test design

The result is a gap with a practical implication. The sources give enough evidence to design a founder baseline, but none gives permission to fill in missing founder observations. A strategic label is not a measurement. A vendor example is not your counterfactual. A workshop conversation is not evidence of AI effectiveness.

Illustration of an anonymized operations-task comparison matrix

The five gates that decide whether a task is eligible

Apply the gates in order. A veto at an earlier gate prevents a high score at a later one from rescuing the task.

  1. Owner gate. One person can explain the current workflow, supply examples, review the test, and accept or reject the output.
  2. Done gate. The final output has a checkable definition of done. “Helpful,” “strategic,” and “better” are not enough.
  3. Evidence gate. You can record frequency, units, handling time, rework, delay, quality, or another signal with a stated period and method.
  4. Safety gate. A human can review the output before a consequential action, and the first slice has a clear stop condition.
  5. Reversibility gate. The test can run in shadow mode or produce a draft without changing the source system or committing an irreversible decision.

Only tasks that pass all five gates enter the comparison score. This prevents the common failure where a visible executive problem wins because it sounds important, even though nobody can say what the AI output should be or who is allowed to approve it.

The gates are a decision artifact, not a claim that the five-source audit measured task performance. They tell you when a task is ready to generate that evidence.

The baseline fields that matter

Record the smallest set of fields that lets another person reconstruct the work. The field is useful only if its value has a source note or an explicit unknown label.

FieldRecordExample evidence label
TriggerWhat starts one unit of workMeasured from timestamps, or observed from a documented workflow
OwnerPerson accountable for the current outputReported and confirmed by the owner
Frequency and unitsCount per day, week, or monthMeasured over a stated period
Handling timeMedian time and an exception or high-end timeMeasured from a time log, not a memory estimate
HandoffsPeople, queues, systems, and waiting pointsObserved in a workflow trace
Rework or error signalRevisions, returns, escalations, or missed fieldsMeasured or reported, with the definition stated
Definition of doneConditions that make the output acceptableConfirmed by the person who accepts it
DependenciesData, access, policy, approval, or system requirementsObserved or unknown
Risk and reversibilityWhat can go wrong and whether the action can be undoneMapped before any live test
CounterfactualWhat happens if the team does nothing or uses the current processRecorded as the comparison condition

Do not convert unknowns into zeros. If no one has measured rework, write “unknown” and make measuring it the next step. The OpenAI evidence guidance explicitly distinguishes measured, observed, reported, estimated, planned, and unknown evidence. That distinction is what keeps a baseline from becoming a polished guess.

How to compare three candidate tasks

Choose three real recurring tasks, not three possible AI products. A task boundary should fit in one sentence and name the trigger, output, and owner.

For each task, calculate recurring effort for the same period:

units per period × median handling minutes ÷ 60 = median labor hours per period

Keep exception handling separate. If ten weekly requests usually take 12 minutes but one takes 90 minutes, do not hide the 90-minute case in an average. Record the distribution or at least the median and high-end case. The high-end case may reveal a handoff or approval problem that a model cannot solve.

Then compare the tasks using this order:

  1. Remove any task that fails an eligibility gate.
  2. Rank the survivors by recurring effort and visible friction.
  3. Prefer the task with stable inputs and a checkable output when effort is similar.
  4. Prefer a task whose first slice can be reviewed and reversed.
  5. Keep one vetoed candidate and its reason in the record.

This procedure deliberately avoids a universal numeric score. The source audit found criteria that recur across the sources, but it did not find a validated weighting that could justify pretending a 4.2 is meaningfully better than a 3.8. Use a transparent ordering and preserve the raw observations instead.

Worked decision artifact for the workshop

The following is a blank artifact to complete with real records. It is not a fabricated founder result. The decision is publishable only after the cells contain source notes and the owner checks the task boundary.

CandidateOwner and done conditionFrequency and effortRework or riskReversible first sliceDecision
Task A: __________________________________________________Test now / validate
Task B: __________________________________________________Test now / validate
Task C: __________________________________________________Test now / validate

Add a source note to each row: record ID or sanitized file, observation date, period covered, collection method, evidence label, and reviewer initials or role. If a row has no owner or done condition, mark it vetoed rather than scoring it.

Illustration of a human-reviewed decision rule for selecting an AI workshop task

The winner is not “the task with the biggest theoretical savings.” It is the eligible task for which a bounded test can answer a useful question. The vetoed task is not a failure. It is evidence that the workflow needs clarification, measurement, ownership, or risk treatment before an AI intervention.

What the first test should look like

Run the first slice in shadow mode when the task affects customers, money, access, policy, or another consequential outcome. The AI can draft, classify, summarize, or recommend. A named human reviews every output before the existing process changes.

Write four things before the workshop ends:

  • Question: what one uncertainty will the test answer?
  • Success condition: which baseline signal should improve without lowering the acceptance standard?
  • Stop condition: what error, missing evidence, unsafe output, or review burden ends the test?
  • Review date: when will the owner inspect the records and decide whether to continue, change the workflow, or stop?

For example, a team might test whether a recurring draft can reach the current definition of done with fewer revisions while a human retains final approval. That is a testable question. “See whether AI makes reporting better” is not.

The test should preserve the counterfactual. Keep a sample of work under the current process, or compare against a clearly defined prior period when that is the only feasible option. Do not claim causality from a shorter week, a new operator, or a changed workload without recording the alternative explanation.

Illustration of a human-reviewed shadow test for a recurring operations task

What this evidence does not tell us

The observed result in this article is a source audit, not a founder field study. Its sample is five primary documents selected because they were already attached to the assigned research package. It does not estimate which task category is most valuable across founders, measure workshop outcomes, or prove that any AI system saves time.

The audit also has a coverage limit. Microsoft provides concrete baseline examples and illustrative calculations, while the OpenAI resources emphasize workflow discovery and evidence labeling. NIST supplies a risk-management structure. Those are different jobs. Their agreement supports the fields and gates in the worksheet, but agreement among guidance documents is not a measured business result.

What we still do not know is which gate fails most often for founders, how much observation is enough before a workshop, and whether the task that wins the baseline also produces the best learning transfer after the workshop. Those questions require permissioned, anonymized founder records and a bounded follow-up test. They cannot be answered by filling this page with plausible numbers.

How to use the worksheet before booking the workshop

Spend one observation cycle on three recurring tasks. Use existing records where possible: timestamps, drafts, ticket histories, checklists, approval messages, or recurring reports. Sanitize customer and company details. Ask the task owner to confirm the boundary and definition of done.

Bring the completed matrix to the workshop and ask the facilitator to challenge the vetoes. The goal is not to force one task into an AI solution. The goal is to leave with one eligible first test, one named owner, a human approval route, a counterfactual, and a date for reviewing evidence.

If you want help turning the worksheet into a safe first experiment, Marius Manolachi's AI consulting and tutoring work is designed to make existing people capable of building AI products on their own work. The AI opportunity research parent provides the broader opportunity context, and the small-business AI use-case prioritization guide covers the adjacent portfolio decision.

The short answer remains simple: baseline the task you can observe, own, review, and reverse. If you cannot do those things yet, the right workshop outcome is better evidence, not automation.

Questions people ask next

Why baseline a task before attending the workshop?

A workshop can narrow ideas, but it cannot recover a missing owner, unclear outcome, or absent counterfactual after the fact. A short baseline gives the group real work to compare and keeps the first test small.

What if none of the tasks passes the veto rule?

Do not force a winner. Use the workshop to repair the weakest missing condition, such as defining done, finding the owner, or recording the workflow for another cycle.