Field note · opportunity

Which Customer Workflow Should a Founder Observe Before Choosing AI?

A founder-sized observation packet ranks customer workflows by recurrence, consequence, handoffs, reviewability, and reversible action before AI enters the plan.

14 minute read
  • AI opportunities
  • Customer workflows
  • Original research
A founder comparing observed customer workflows before choosing an AI pilot

Founders often start with the most visible complaint, the most repeated task, or the AI demo that looks easiest to sell. None of those tells you which workflow to study first.

The useful starting point is smaller. Watch one customer-facing job move through the business. Look for the handoffs, waiting, exceptions, and review points that a person can actually describe. Then compare that workflow with two alternatives before choosing a model, an agent, or an automation.

A founder mapping three customer-facing workflows before choosing an AI pilot

Which customer-facing workflow should you observe first?

Observe the workflow where a recurring customer job crosses a visible handoff, produces an output someone can review, and ends with a reversible next action. That combination gives you enough repetition to learn, enough structure to compare, and enough control to avoid turning an early experiment into an uncontrolled customer action.

This is a decision rule for where to look first, not a claim that every business should automate the same job. A founder should prefer a workflow whose owner can show a real example, name the current workaround, and say what would count as a useful result. If the owner cannot do that, the workflow is still a discovery problem.

NIST's AI Risk Management Framework puts context mapping before an initial decision about whether an AI solution is appropriate. It asks teams to understand purpose, users, setting, impacts, constraints, and human oversight before they measure or manage a system. GOV.UK's discovery guidance makes the same practical move from another direction: research the service in the context where people use it, rather than designing from assumptions. (NIST AI RMF Core, GOV.UK user research in discovery)

The exception is a workflow with a consequential side effect. If an error can change money, access, eligibility, a legal position, or a customer's safety, observe it, but keep the first action in a human-controlled path. You are learning the boundary, not granting the system authority.

What should you record while watching the workflow?

Record the customer's job first, then record how the business currently completes it. A useful observation card contains the fields below. Unknown is a valid value. A blank that says “not observed” is safer than a confident guess.

FieldWhat to captureWhy it changes the decision
Customer jobWhat the customer is trying to achieve, in their wordsSeparates a real outcome from an internal task label
Entry triggerThe event, request, message, or delay that starts the workShows whether the workflow can be sampled repeatedly
Visible stepsWhat the operator actually reads, decides, changes, and sendsPrevents a solution from hiding the real process
HandoffsPeople, queues, systems, or approvals between stepsExposes lost context and ownership gaps
WaitingWhere the customer or operator pauses, and whyDistinguishes model latency from an organisational delay
ExceptionsInputs or cases that leave the normal pathShows whether a narrow pilot can exist
Current workaroundThe tool, spreadsheet, message, or memory used todayGives you a baseline and a non-AI alternative
Data boundaryWhat can be seen, copied, retained, or sent elsewhereSets the first safe experiment boundary
Decision ownerThe person who can accept, reject, or change the resultPrevents “the team” from becoming an unaccountable owner
Error consequenceWhat happens if the output is wrong or lateTells you how much review and reversibility you need
Review pointWhere a person checks the work before the customer is affectedMakes human oversight part of the workflow, not a promise
Smallest reversible actionThe least consequential next step that would test valueKeeps learning separate from irreversible rollout

Microsoft's intake guidance treats an AI idea as something to structure and prioritize, not as a self-justifying request for an agent. Microsoft's planning guidance also asks teams to connect adoption to business outcomes, roles, governance, and readiness. The OpenAI Academy workflow matrix adds a practical comparison frame for weighing opportunities before implementation. These sources anchor the card. The card's value is the discipline of attaching each field to one observed workflow and one next decision. (Microsoft intake and prioritization, Microsoft plan for AI adoption, OpenAI Academy workflow discovery and prioritization matrix)

Do not turn the card into an interview script that asks, “Where could AI help?” Ask the person to replay the last real case. Let the work reveal the trigger, the handoff, the correction, and the point where the customer waits.

How can you compare three workflows without inventing a statistic?

Score the cases with an ordinal rubric, then show the row-level reasoning. The score is a comparison aid, not a probability, forecast, or return-on-investment estimate.

Give each dimension 0 to 3 points:

Dimension0123
RecurrenceNo repeatable case foundOccasional caseRepeated caseRepeated and easy to sample
Customer consequenceConsequence unclearLow visible effectNoticeable delay or reworkClear outcome that the owner can inspect
VariationNo case boundaryMostly different casesSome stable patternStable job with meaningful exceptions
HandoffsNo handoff observedOne hidden handoffSeveral visible handoffsHandoffs have clear states or owners
ReviewabilityNo reviewer or standardInformal reviewNamed reviewerReviewer and acceptance condition are explicit
Safe action boundaryNo reversible stepReversible step is vagueDraft or recommendation is possibleSmall action is reversible and human-controlled

The total is out of 18. Do not use the total to override a veto. Reject or hold a candidate when the decision owner is missing, the data boundary is unknown, the output would act without a review point, or the error cannot be recovered at reasonable cost. A high score cannot make an unsafe action safe.

This is the page's reusable artifact: an observation card plus a six-dimension comparison matrix and a veto. Another founder can copy the tables, observe three workflows, score the same fields, and see where the ranking came from. The artifact does not pretend that an ordinal score is a measured business result.

A six-dimension workflow comparison matrix with a separate veto row for unsafe actions

What does a bounded observation corpus show?

The following three cards are deliberately bounded. They come from Marius Manolachi's locked teaching and shipping facts, not from a claim about the distribution of founder-led companies. Fields that the source does not establish remain unknown.

Card A: moving from product specs to shipping

The observed job is the move from writing a specification to building and shipping a product. The useful signal is not that a model could write more text. It is that the owner must define what “done” means before a result can be reviewed.

When I taught product managers who went from writing specs to building and shipping the product, and automating work around it, the recurring failure was an undefined “done,” not the model. That observation makes this a strong first workflow to inspect when a founder is choosing where to learn: the acceptance boundary is visible, the owner can review the result, and a draft artifact can be tested before any customer-facing action.

Known fields: customer job, product-work trigger, owner, and review decision. Unknown fields: customer volume, exception rate, data sensitivity, and exact waiting time. The card is useful because it exposes the decision boundary. It is not a claim about those unknowns.

Card B: starting from work people already do

The Orange workshop supplies a different observation. The session started from the attendees' existing work rather than from an agent as the starting point. That is a useful guard against solution enthusiasm: choose a real workflow, then decide whether AI belongs in it.

Known fields: work context and a human learning setting. Unknown fields: the workflow's recurrence, customer consequence, exception rate, and action boundary. This card therefore scores well as an observation principle but cannot, by itself, justify an automation pilot.

The decision is simple: use the real work as the unit of discovery. Do not treat a workshop idea, an attractive prompt, or a vendor demonstration as evidence that a customer-facing workflow is ready.

Card C: live screen annotation with human approval

TryUncle is an agent that watches the screen and annotates it live. That makes latency and human approval product constraints, not afterthoughts. It is a valuable reminder that a workflow can look technically interesting while still needing a narrow action boundary.

Known fields: live visual context, timing constraint, and the need to consider human approval. Unknown fields: customer job distribution, exception rate, data retention, and the exact consequence of a wrong annotation. The card should therefore be observed before it is expanded, and its first test should preserve review rather than remove it.

These cards are qualitative observations, not a sample of three customer businesses. Their purpose is to make the matrix concrete and to show how a locked observation can change the next decision. The matrix below is a worked inference from the available evidence, not a market ranking.

Which candidate wins the worked comparison?

The product-spec-to-shipping workflow ranks first for a first observation or pilot because its review boundary is clearest and its smallest reversible action is easy to name: write the acceptance condition, create a draft, and have the owner inspect it. TryUncle's live annotation workflow is rejected as the first candidate because its timing and human-approval boundary need more observation before action can safely move closer to a customer.

CandidateRecurrenceConsequenceVariationHandoffsReviewabilitySafe boundaryTotalDecision
Product specs to shipping33233317/18Rank first
Start from existing work22223314/18Use as discovery rule
Live screen annotation23322113/18Reject as first action

The numbers are not measurements of frequency or value. They are the worked application of the rubric to the three bounded cards. If you run the packet on real customer workflows, your rows may change. That is the point of showing the method.

The winner still does not mean “build an agent.” It means “observe this workflow in more detail.” The next artifact should contain three to five real cases, a baseline for completion and correction, a named reviewer, the approved data boundary, and a draft-only test. If those fields cannot be completed, the correct decision is HOLD.

A ranked workflow decision artifact showing a winner, a rejected candidate, and a human review boundary

When should a founder choose no AI?

Choose no AI for the first intervention when a clearer source of truth, a rule, a form, a queue, or a process change can solve the observed problem with less uncertainty. AI is not the default winner because the workflow contains language or because a model can produce a plausible draft.

Use four no-AI checks:

  1. Can a deterministic rule route the case without hiding an important exception?
  2. Can a better form collect the missing information at the trigger?
  3. Can one source of truth remove the handoff instead of summarizing it?
  4. Can a queue, owner, or review schedule remove the waiting time?

If the answer to one of these is yes, test that change first or include it as the baseline. NIST explicitly includes viable non-AI alternatives in risk management. A founder who records the baseline makes the AI decision more honest because the comparison is against the work as it could be improved, not only the work as it is today.

The exception is a workflow where the non-AI change fixes the process but not the interpretation problem. If people still need to classify messy language, compare context, or draft a reviewable response after the process is clear, a bounded AI assist may be worth testing. Keep the output as a draft until the owner can show reliable review and recovery.

How should a founder run the first observation sprint?

Run a short, evidence-first sprint with three candidate workflows. The goal is not to finish with an AI architecture. The goal is to leave with one ranked observation target and one explicit rejection.

  1. Choose three real workflows. Pick customer jobs that have occurred recently. Name the customer outcome, not the internal department.
  2. Replay the last case. Watch the operator use the actual message, record, tool, or document. Ask what happened next when the normal path failed.
  3. Fill the observation card. Record the trigger, steps, handoffs, waiting, exceptions, data boundary, owner, consequence, review point, and reversible action. Mark unknowns.
  4. Score before discussing models. Apply the same 0-3 rubric to every row. Write one sentence for each non-zero score.
  5. Apply the veto. Hold any candidate without an owner, data boundary, review path, or recoverable action.
  6. Choose the smallest test. Prefer a draft, recommendation, classification, or retrieval step. Do not let the first experiment send a customer message or change a durable record automatically.
  7. Write the next decision. State what evidence would promote the candidate, keep it on hold, or reject it.

A seven-step observation sprint from customer job replay to a reversible AI decision

The sprint produces a decision packet, not a promise. Keep the raw cards, the scoring sheet, the rejected candidate, the reviewer boundary, and the sample limitations together. A later build should be able to trace its scope back to an observed case.

What should the founder carry into an AI pilot?

Carry one ranked workflow, one rejected alternative, one no-AI baseline, and one reviewable test. The packet is complete when another person can understand why the winner ranked first without hearing the founder's original pitch.

For the ranked workflow, write this next-step contract:

Decision fieldRequired answer
JobThe customer outcome being supported
First AI roleDraft, classify, retrieve, compare, or recommend
Source of truthThe approved records or documents the output may use
ReviewerThe person who accepts, edits, rejects, or escalates it
Success evidenceThe observed result that would justify continuation
Failure evidenceThe case or correction that pauses the test
Action boundaryWhat the system cannot send, change, approve, or delete
Exit decisionContinue, narrow, replace with a non-AI change, or stop

This is where the broader guide to prioritizing AI use cases in a small business becomes useful: it helps place the ranked candidate in the wider opportunity set. Link the canonical AI opportunity page when the reader needs the cluster-level map. The narrower job of this article is the observation packet that comes before the shortlist.

If you want help turning a real workflow into a packet your team can own, Marius Manolachi's AI learning and consulting work follows the same boundary: make existing people capable of building AI products on their own work. The next conversation should start with the workflow cards, not with a preferred model.

Distribution hooks

  • LinkedIn: For founders and product leaders choosing a first AI workflow, show the ranked observation table and the veto that kept a flashy customer-facing automation from becoming the first pilot.
  • Newsletter: For small teams building or buying AI capability, share the anonymized observation card and explain what must be visible before a founder chooses AI over a rule, process change, or better source of truth.
  • X: For AI builders and operators, thread the observed workflow signals: repeatability, variation, handoffs, reviewability, and safe action boundary, with this dataset link as the primary evidence.