Field note · opportunity
Which Customer Workflow Should a Founder Observe Before Choosing AI?
A founder-sized observation packet ranks customer workflows by recurrence, consequence, handoffs, reviewability, and reversible action before AI enters the plan.

Founders often start with the most visible complaint, the most repeated task, or the AI demo that looks easiest to sell. None of those tells you which workflow to study first.
The useful starting point is smaller. Watch one customer-facing job move through the business. Look for the handoffs, waiting, exceptions, and review points that a person can actually describe. Then compare that workflow with two alternatives before choosing a model, an agent, or an automation.

Which customer-facing workflow should you observe first?
Observe the workflow where a recurring customer job crosses a visible handoff, produces an output someone can review, and ends with a reversible next action. That combination gives you enough repetition to learn, enough structure to compare, and enough control to avoid turning an early experiment into an uncontrolled customer action.
This is a decision rule for where to look first, not a claim that every business should automate the same job. A founder should prefer a workflow whose owner can show a real example, name the current workaround, and say what would count as a useful result. If the owner cannot do that, the workflow is still a discovery problem.
NIST's AI Risk Management Framework puts context mapping before an initial decision about whether an AI solution is appropriate. It asks teams to understand purpose, users, setting, impacts, constraints, and human oversight before they measure or manage a system. GOV.UK's discovery guidance makes the same practical move from another direction: research the service in the context where people use it, rather than designing from assumptions. (NIST AI RMF Core, GOV.UK user research in discovery)
The exception is a workflow with a consequential side effect. If an error can change money, access, eligibility, a legal position, or a customer's safety, observe it, but keep the first action in a human-controlled path. You are learning the boundary, not granting the system authority.
What should you record while watching the workflow?
Record the customer's job first, then record how the business currently completes it. A useful observation card contains the fields below. Unknown is a valid value. A blank that says “not observed” is safer than a confident guess.
| Field | What to capture | Why it changes the decision |
|---|---|---|
| Customer job | What the customer is trying to achieve, in their words | Separates a real outcome from an internal task label |
| Entry trigger | The event, request, message, or delay that starts the work | Shows whether the workflow can be sampled repeatedly |
| Visible steps | What the operator actually reads, decides, changes, and sends | Prevents a solution from hiding the real process |
| Handoffs | People, queues, systems, or approvals between steps | Exposes lost context and ownership gaps |
| Waiting | Where the customer or operator pauses, and why | Distinguishes model latency from an organisational delay |
| Exceptions | Inputs or cases that leave the normal path | Shows whether a narrow pilot can exist |
| Current workaround | The tool, spreadsheet, message, or memory used today | Gives you a baseline and a non-AI alternative |
| Data boundary | What can be seen, copied, retained, or sent elsewhere | Sets the first safe experiment boundary |
| Decision owner | The person who can accept, reject, or change the result | Prevents “the team” from becoming an unaccountable owner |
| Error consequence | What happens if the output is wrong or late | Tells you how much review and reversibility you need |
| Review point | Where a person checks the work before the customer is affected | Makes human oversight part of the workflow, not a promise |
| Smallest reversible action | The least consequential next step that would test value | Keeps learning separate from irreversible rollout |
Microsoft's intake guidance treats an AI idea as something to structure and prioritize, not as a self-justifying request for an agent. Microsoft's planning guidance also asks teams to connect adoption to business outcomes, roles, governance, and readiness. The OpenAI Academy workflow matrix adds a practical comparison frame for weighing opportunities before implementation. These sources anchor the card. The card's value is the discipline of attaching each field to one observed workflow and one next decision. (Microsoft intake and prioritization, Microsoft plan for AI adoption, OpenAI Academy workflow discovery and prioritization matrix)
Do not turn the card into an interview script that asks, “Where could AI help?” Ask the person to replay the last real case. Let the work reveal the trigger, the handoff, the correction, and the point where the customer waits.
How can you compare three workflows without inventing a statistic?
Score the cases with an ordinal rubric, then show the row-level reasoning. The score is a comparison aid, not a probability, forecast, or return-on-investment estimate.
Give each dimension 0 to 3 points:
| Dimension | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Recurrence | No repeatable case found | Occasional case | Repeated case | Repeated and easy to sample |
| Customer consequence | Consequence unclear | Low visible effect | Noticeable delay or rework | Clear outcome that the owner can inspect |
| Variation | No case boundary | Mostly different cases | Some stable pattern | Stable job with meaningful exceptions |
| Handoffs | No handoff observed | One hidden handoff | Several visible handoffs | Handoffs have clear states or owners |
| Reviewability | No reviewer or standard | Informal review | Named reviewer | Reviewer and acceptance condition are explicit |
| Safe action boundary | No reversible step | Reversible step is vague | Draft or recommendation is possible | Small action is reversible and human-controlled |
The total is out of 18. Do not use the total to override a veto. Reject or hold a candidate when the decision owner is missing, the data boundary is unknown, the output would act without a review point, or the error cannot be recovered at reasonable cost. A high score cannot make an unsafe action safe.
This is the page's reusable artifact: an observation card plus a six-dimension comparison matrix and a veto. Another founder can copy the tables, observe three workflows, score the same fields, and see where the ranking came from. The artifact does not pretend that an ordinal score is a measured business result.

What does a bounded observation corpus show?
The following three cards are deliberately bounded. They come from Marius Manolachi's locked teaching and shipping facts, not from a claim about the distribution of founder-led companies. Fields that the source does not establish remain unknown.
Card A: moving from product specs to shipping
The observed job is the move from writing a specification to building and shipping a product. The useful signal is not that a model could write more text. It is that the owner must define what “done” means before a result can be reviewed.
When I taught product managers who went from writing specs to building and shipping the product, and automating work around it, the recurring failure was an undefined “done,” not the model. That observation makes this a strong first workflow to inspect when a founder is choosing where to learn: the acceptance boundary is visible, the owner can review the result, and a draft artifact can be tested before any customer-facing action.
Known fields: customer job, product-work trigger, owner, and review decision. Unknown fields: customer volume, exception rate, data sensitivity, and exact waiting time. The card is useful because it exposes the decision boundary. It is not a claim about those unknowns.
Card B: starting from work people already do
The Orange workshop supplies a different observation. The session started from the attendees' existing work rather than from an agent as the starting point. That is a useful guard against solution enthusiasm: choose a real workflow, then decide whether AI belongs in it.
Known fields: work context and a human learning setting. Unknown fields: the workflow's recurrence, customer consequence, exception rate, and action boundary. This card therefore scores well as an observation principle but cannot, by itself, justify an automation pilot.
The decision is simple: use the real work as the unit of discovery. Do not treat a workshop idea, an attractive prompt, or a vendor demonstration as evidence that a customer-facing workflow is ready.
Card C: live screen annotation with human approval
TryUncle is an agent that watches the screen and annotates it live. That makes latency and human approval product constraints, not afterthoughts. It is a valuable reminder that a workflow can look technically interesting while still needing a narrow action boundary.
Known fields: live visual context, timing constraint, and the need to consider human approval. Unknown fields: customer job distribution, exception rate, data retention, and the exact consequence of a wrong annotation. The card should therefore be observed before it is expanded, and its first test should preserve review rather than remove it.
These cards are qualitative observations, not a sample of three customer businesses. Their purpose is to make the matrix concrete and to show how a locked observation can change the next decision. The matrix below is a worked inference from the available evidence, not a market ranking.
Which candidate wins the worked comparison?
The product-spec-to-shipping workflow ranks first for a first observation or pilot because its review boundary is clearest and its smallest reversible action is easy to name: write the acceptance condition, create a draft, and have the owner inspect it. TryUncle's live annotation workflow is rejected as the first candidate because its timing and human-approval boundary need more observation before action can safely move closer to a customer.
| Candidate | Recurrence | Consequence | Variation | Handoffs | Reviewability | Safe boundary | Total | Decision |
|---|---|---|---|---|---|---|---|---|
| Product specs to shipping | 3 | 3 | 2 | 3 | 3 | 3 | 17/18 | Rank first |
| Start from existing work | 2 | 2 | 2 | 2 | 3 | 3 | 14/18 | Use as discovery rule |
| Live screen annotation | 2 | 3 | 3 | 2 | 2 | 1 | 13/18 | Reject as first action |
The numbers are not measurements of frequency or value. They are the worked application of the rubric to the three bounded cards. If you run the packet on real customer workflows, your rows may change. That is the point of showing the method.
The winner still does not mean “build an agent.” It means “observe this workflow in more detail.” The next artifact should contain three to five real cases, a baseline for completion and correction, a named reviewer, the approved data boundary, and a draft-only test. If those fields cannot be completed, the correct decision is HOLD.

When should a founder choose no AI?
Choose no AI for the first intervention when a clearer source of truth, a rule, a form, a queue, or a process change can solve the observed problem with less uncertainty. AI is not the default winner because the workflow contains language or because a model can produce a plausible draft.
Use four no-AI checks:
- Can a deterministic rule route the case without hiding an important exception?
- Can a better form collect the missing information at the trigger?
- Can one source of truth remove the handoff instead of summarizing it?
- Can a queue, owner, or review schedule remove the waiting time?
If the answer to one of these is yes, test that change first or include it as the baseline. NIST explicitly includes viable non-AI alternatives in risk management. A founder who records the baseline makes the AI decision more honest because the comparison is against the work as it could be improved, not only the work as it is today.
The exception is a workflow where the non-AI change fixes the process but not the interpretation problem. If people still need to classify messy language, compare context, or draft a reviewable response after the process is clear, a bounded AI assist may be worth testing. Keep the output as a draft until the owner can show reliable review and recovery.
How should a founder run the first observation sprint?
Run a short, evidence-first sprint with three candidate workflows. The goal is not to finish with an AI architecture. The goal is to leave with one ranked observation target and one explicit rejection.
- Choose three real workflows. Pick customer jobs that have occurred recently. Name the customer outcome, not the internal department.
- Replay the last case. Watch the operator use the actual message, record, tool, or document. Ask what happened next when the normal path failed.
- Fill the observation card. Record the trigger, steps, handoffs, waiting, exceptions, data boundary, owner, consequence, review point, and reversible action. Mark unknowns.
- Score before discussing models. Apply the same 0-3 rubric to every row. Write one sentence for each non-zero score.
- Apply the veto. Hold any candidate without an owner, data boundary, review path, or recoverable action.
- Choose the smallest test. Prefer a draft, recommendation, classification, or retrieval step. Do not let the first experiment send a customer message or change a durable record automatically.
- Write the next decision. State what evidence would promote the candidate, keep it on hold, or reject it.

The sprint produces a decision packet, not a promise. Keep the raw cards, the scoring sheet, the rejected candidate, the reviewer boundary, and the sample limitations together. A later build should be able to trace its scope back to an observed case.
What should the founder carry into an AI pilot?
Carry one ranked workflow, one rejected alternative, one no-AI baseline, and one reviewable test. The packet is complete when another person can understand why the winner ranked first without hearing the founder's original pitch.
For the ranked workflow, write this next-step contract:
| Decision field | Required answer |
|---|---|
| Job | The customer outcome being supported |
| First AI role | Draft, classify, retrieve, compare, or recommend |
| Source of truth | The approved records or documents the output may use |
| Reviewer | The person who accepts, edits, rejects, or escalates it |
| Success evidence | The observed result that would justify continuation |
| Failure evidence | The case or correction that pauses the test |
| Action boundary | What the system cannot send, change, approve, or delete |
| Exit decision | Continue, narrow, replace with a non-AI change, or stop |
This is where the broader guide to prioritizing AI use cases in a small business becomes useful: it helps place the ranked candidate in the wider opportunity set. Link the canonical AI opportunity page when the reader needs the cluster-level map. The narrower job of this article is the observation packet that comes before the shortlist.
If you want help turning a real workflow into a packet your team can own, Marius Manolachi's AI learning and consulting work follows the same boundary: make existing people capable of building AI products on their own work. The next conversation should start with the workflow cards, not with a preferred model.
Distribution hooks
- LinkedIn: For founders and product leaders choosing a first AI workflow, show the ranked observation table and the veto that kept a flashy customer-facing automation from becoming the first pilot.
- Newsletter: For small teams building or buying AI capability, share the anonymized observation card and explain what must be visible before a founder chooses AI over a rule, process change, or better source of truth.
- X: For AI builders and operators, thread the observed workflow signals: repeatability, variation, handoffs, reviewability, and safe action boundary, with this dataset link as the primary evidence.