Field note · opportunity

Which AI Opportunity Should a 20-Person Company Investigate First?

Use a worked six-opportunity scorecard to choose one safe, testable AI investigation for a 20-person company.

11 minute read
  • AI opportunity
  • Small business AI
  • AI strategy
Illustration of a 20-person company choosing one AI opportunity from a weighted decision matrix

Most 20-person companies don’t need six AI pilots. They need one defensible reason to start with one.

Here is my worked answer, followed by the worksheet that produces it.

Illustration of six AI opportunities entering a weighted decision matrix

The worked answer: start with document or inbox triage

For the representative company in this worksheet, investigate document or inbox triage first. It scores 86 out of 100 because it is frequent, mostly reversible, based on authorized information, easy to keep in draft mode, and suitable for a measurable human-reviewed test.

That is a situated choice, not a universal best AI use case. If your company has stable historical demand and a real forecasting owner, the same worksheet makes forecasting the winner. If the first action is irreversible or high consequence, a high score cannot rescue it from a veto.

Microsoft’s business-envisioning guidance starts with the problem, opportunity, business objective, success measure, and accountable owner, then compares business, experience, and technology viability. The Microsoft guidance is written for ISVs, but its discipline fits a small-company shortlist.

When I taught product managers who moved from writing specs to building and shipping, the recurring failure was often that nobody could say what “done” meant. That is my bounded F-pms observation from Marius Manolachi’s locked entity facts, not a measured rate. For opportunity selection, “done” means a baseline, a reviewable output, an owner, and a stop rule.

If you are building a wider backlog, use this page as a spoke to the AI opportunity pillar and compare it with the broader small-business AI prioritization guide.

What should the scorecard measure?

Score the work, not the tool. Use seven dimensions with weights that total 100.

DimensionWeightScore 1Score 3Score 5Evidence to collect
Business impact20No named business resultHelps one team, result is indirectDirectly improves a named objectiveObjective, stakeholder, and baseline measure
Task frequency15Less than weeklySeveral times per weekMultiple times per day or a persistent queueTen-business-day sample or system log
Reversibility15Error is hard to undoHuman review or partial rollback existsRead-only, draft, or easy undoAction boundary and rollback path
Data readiness15Data is unavailable or unauthorizedData exists but needs access or cleanupAuthorized, findable, consistent sample existsSource, permissions, format, and gaps
AI fit15Deterministic rules are plainly betterAI helps one bounded stepLanguage, retrieval, classification, or prediction is central and testableTask decomposition and non-AI alternative
Owner capacity10No accountable ownerOwner exists but review time is unclearNamed owner has time and the needed skillsOwner, review hours, and escalation path
Measurement ease10No credible baselineProxy is availableBaseline, threshold, and cadence are explicitUnit of analysis and success rule

The categories translate Microsoft’s strategic fit and BXT questions into a compact worksheet. Microsoft explicitly covers business value and change timeframe, user personas and change resistance, operational risk and safeguards, and AI/LLM fit. OECD’s SME paper adds a useful constraint: data, skills, connectivity, and finance are adoption enablers, so a promising idea without data or an owner is not ready just because the model can do the task. Frequency is an operational measure grounded in the recent-work evidence a small company can collect, with the San Francisco Fed’s reported use cases and capacity barriers explaining why recurring work deserves attention. Reversibility is an operational form of Microsoft’s change and safeguard questions and NIST’s Manage function, which supports documenting a response and disengaging when outcomes do not match intended use.

How do you apply the worksheet?

Use this sequence for every candidate:

  1. Write the current workflow and the result it should improve.
  2. Collect the evidence named in the table. Do not award a 4 or 5 from enthusiasm alone.
  3. Score each dimension from 1 to 5. Explain every score in one sentence.
  4. Calculate weighted score = sum(score × weight) ÷ 5.
  5. Apply the vetoes below. A veto is not a penalty. It removes the candidate from the first-investigation shortlist.
  6. If two candidates are within three points, choose the more reversible one. If still tied, choose lower risk exposure, then the cheaper and faster investigation.

Risk vetoes

Veto questionIf yes, what happens?
Would a wrong result affect employment, credit, legal rights, health, safety, payments, or a material customer entitlement without qualified human review?Veto until the qualified reviewer and approval point are explicit.
Is the input data unauthorized, sensitive, or subject to unclear privacy or intellectual-property rules?Veto until the data boundary and approved handling are documented.
Can the system send, publish, delete, purchase, promise, or change a record without a reversible approval step?Veto the autonomous action. Narrow the investigation to read-only output or drafts.
Is there no accountable owner, baseline, or stop condition?Veto the pilot. Prepare the missing evidence first.

This follows the shape of NIST’s AI Risk Management Framework: Govern the responsibility and policy, Map the context and harms, Measure performance and risk, then Manage the response and the decision to continue or disengage. NIST’s AI RMF Core is voluntary and not a checklist, but it is a useful guardrail against treating risk as a small deduction from a large score.

How do six representative opportunities rank?

The following is a worked example for a 20-person services company with a shared inbox, recurring operational documents, a modest knowledge base, no dedicated data science team, and an owner who can review a pilot for a few hours per week. These are assumptions, not client data or measured rates.

OpportunityImpactFrequencyReversibleDataAI fitOwnerMeasureScoreVeto result
Internal knowledge retrieval345344374Pass with citations and read-only access
Document or inbox triage454454486Pass if routing and drafts stay human-approved
Customer-service assistance454453484Pass for suggestions; narrow high-impact cases
Sales research345344374Pass with approved sources and no automatic outreach
Marketing production345455382Pass for drafts; review claims and publication
Forecasting522322255Veto in this context because the owner and review path are weak

The winning arithmetic is transparent: (4×20 + 5×15 + 4×15 + 4×15 + 5×15 + 4×10 + 4×10) ÷ 5 = 86.

The candidate set is not arbitrary. The San Francisco Fed reports small-business AI applications across productivity, marketing, written communications, visuals, customer service, analytics and forecasting, while also reporting barriers involving policy, cost, staff time, training, implementation knowledge, accuracy, and intellectual property. Its 2026 brief is qualitative context, not a ranking of what your company should do. A separate Intuit QuickBooks survey of more than 2,200 US businesses with up to 100 employees similarly reports marketing, customer service, administrative work, data processing, and bookkeeping as common uses. That survey is descriptive and commissioned by Intuit, so I use it as a cross-check, not as independent proof of priority.

When should a different company choose forecasting?

Change the context, then recalculate. Consider a 20-person subscription business with several years of stable demand history, a finance or operations owner who reviews weekly forecasts, and a forecast used only to prepare a human planning decision.

OpportunityImpactFrequencyReversibleDataAI fitOwnerMeasureScoreResult
Document or inbox triage344352370Still viable, but owner capacity is weak
Customer-service assistance454453484Viable for reviewed suggestions
Forecasting553544589First investigation candidate

Forecasting wins here because frequency, data readiness, owner capacity, measurement, and reversibility changed. It is still advisory. It does not place an order, change a price, or commit cash. If the team removes that approval boundary, the forecasting candidate hits a veto again.

What should the first investigation brief contain?

For the typical context, keep the first test narrow: classify incoming documents or inbox items, suggest urgency and owner, and draft a next step. Do not let it send, delete, pay, publish, or update a system of record.

Brief fieldFilled investigation choice
HypothesisAI-assisted triage can reduce handling time without increasing critical misroutes.
Baseline metricMedian minutes from receipt to correct queue, measured on 40 recent historical items. Record current routing accuracy and critical misroutes as secondary measures.
Sample40 items, stratified across the four largest categories, with the category and correct queue labeled by the current owner.
ModeShadow mode first. Produce a classification, urgency, owner suggestion, and draft next step without writing to systems or contacting anyone.
Human approval pointThe named operations owner approves the queue, urgency, and any draft before routing, external communication, deletion, payment, or commitment.
Success thresholdAt least 25% lower median handling time and at least 95% correct routing, with zero unreviewed high-consequence actions. This is a proposed threshold, not a measured result.
Cost fieldsModel and tool cost per item, fixed monthly cost, reviewer minutes per item, correction minutes, and total cost per accepted item.
Latency fieldsModel p50 and p95 latency, end-to-end time to a reviewable result, queue wait, and human approval time.
Stop conditionStop immediately for unauthorized data transfer, an unapproved external action, or a high-consequence misroute. Stop and redesign after two consecutive review batches below 95% routing accuracy or when the full review path fails to reach the 25% time reduction.
Decision after sampleContinue to a bounded live shadow test, narrow the categories, or stop. Do not expand because the demo looked good.

The brief turns Microsoft’s success-measurement and accountability questions into an operational test. It also gives NIST’s Measure and Manage functions something concrete to inspect. The sample is deliberately small enough for a 20-person team to label and review, but it is not large enough to prove long-term reliability.

What should you remember before copying the winner?

The winner is not “AI triage.” The winner is the candidate that best fits a stated context after value, recurrence, reversibility, data, AI fit, owner capacity, measurement, and risk have been checked.

This worksheet has important limits. Its scores are illustrative assumptions, not measurements from a real company. It does not estimate ROI, model accuracy, legal compliance, or adoption. The surveys describe reported uses and barriers, not outcomes for your team. Change the weights when your strategy, data, owner, or risk tolerance changes.

Start with the candidate that gives your team the fastest trustworthy learning loop. For the representative company above, that is document or inbox triage in shadow mode. For the alternate company, it is forecasting with a human approval boundary. The first investigation is a decision about evidence, not a permanent commitment to a technology.

Questions people ask next

Is document or inbox triage always the best first AI opportunity?

No. It wins the worked example because it is frequent, bounded, data-ready, and easy to review. A different company can choose forecasting, customer-service assistance, or another candidate when its data, owner, baseline, and risk boundary are stronger.

What should a small company do if two AI opportunities tie?

Choose the more reversible candidate first. If reversibility also ties, choose the one with the lower risk exposure, then the shorter investigation with the cheaper and faster feedback loop.

What is the safest first version of document triage?

Run it in shadow mode on historical or newly arriving items. Let AI classify and draft suggestions, but keep routing confirmation, external communication, deletion, payment, and other consequential actions with a named human owner.