Field note · implementation
AI Implementation Estimates Miss Integration and Change-Management Work
AI implementation quotes often price the model feature first. Use this ten-surface worksheet to expose integration, adoption, rollout, and ownership work.

The first AI estimate is often a feature estimate wearing a delivery costume. It counts the prompt, the model call, and a small interface. Then the team tries to connect it to real work.
When I taught product managers to move from writing specifications to building and shipping, the missing line was rarely the model. It was the boundary around the model: who uses the result, which system it touches, what happens when it is wrong, and who owns it next. (F-pms)
Reproduce the failure with a proposal trace
You can reproduce the estimate failure without a live build: hold the use case constant, trace the quoted feature through data, workflow, release, and ownership, and mark every unanswered handoff as missing work. In the bounded support-triage scenario below, the first quote stops after the model and interface.
| Trace step | What the model-first quote says | Question the trace must answer | Result in this scenario |
|---|---|---|---|
| Feature | Prompting, output shape, interface, and error display | Can the feature produce a plausible draft? | Included in 8 days |
| Boundary | No data, integration, or identity line | Which source is read, with what permission, and how are failures logged? | Missing |
| Human work | No workflow, evaluation, or training line | Who approves, corrects, routes, and escalates the result? | Missing |
| Release | No rollout, support, or ownership line | Who releases it, handles defects, refreshes the evaluation, and retires it? | Missing |
This is a reproduction of the failure in the hypothetical decision artifact, not a project trace or an industry benchmark. It shows how a 10-day feature quote can look complete while leaving the operational outcome undefined.
Diagnose the missing boundary
An AI implementation estimate is complete only when it prices the operational boundary around the model, not just the model behavior.
The practical test is simple. Ask whether the estimate names work for data access, system integration, workflow change, evaluation, governance, adoption, rollout, support, and ongoing ownership. If those lines are missing, the number may be a prototype quote. It is not yet a production estimate.
The primary sources point in the same direction, although none gives a universal cost multiplier. McKinsey's 2025 global survey says nearly two-thirds of respondents had not begun scaling AI across the enterprise and identifies workflow redesign as a key success factor. Gartner's survey names estimating and demonstrating business value as the top adoption barrier for 49% of respondents. (McKinsey's State of AI, Gartner's survey)
The sourceable artifact on this page is a ten-surface estimate-completeness worksheet, filled with one bounded hypothetical implementation. It is original analysis, not industry benchmark data.
| Estimate type | What it includes | What decision it can support |
|---|---|---|
| Model-first quote | Feature behavior, prompt or model calls, and a thin interface | Is the idea technically plausible? |
| Complete bounded release estimate | Model, data, integration, workflow, evaluation, governance, training, rollout, support, and ownership | Can we fund and operate this release? |
| Ongoing operating estimate | Review cadence, monitoring, evaluation refresh, source updates, vendor cost, and retirement | Can someone own it after handover? |

Why integration work disappears from the first quote
Integration disappears because a demo can avoid the constraints that make a live workflow expensive.
A demo can use a pasted sample. A release needs identity, permissions, source freshness, retries, logs, failure handling, and a representative environment. A demo can show a good output in isolation. A release has to place that output at the right point in a real process.
Gartner found that embedding generative AI in existing applications was the primary way respondents fulfilled GenAI use cases, reported by 34% of its survey sample. That is a useful clue about the work boundary: the feature is often entering an existing system, not replacing one. (Gartner's deployment survey)
EIOPA's insurance survey describes a similar progression from proof of concept to controlled internal deployment and then integration into more critical business processes. It also reports that human supervision remains central. The full report is a sector survey, not a small-company delivery plan, but its sequence exposes two lines that a model-only quote hides: integration and review. (EIOPA's market survey, EIOPA report PDF)
For every integration surface, ask four questions:
- What system is read, and what system is written?
- Which identity and permission does the workflow use?
- What happens when the source is stale, unavailable, or contradictory?
- Where does a human see, approve, correct, or reject the result?
If a proposal cannot answer those questions, mark integration as an open dependency. Do not bury it inside “engineering.”
Why change-management work appears after the demo
Change work appears late because the demo proves that a person can see an output. It does not prove that the person will trust it, use it at the right moment, or know what to do when it fails.
The workflow changes when a new result creates a new decision, review step, exception path, or ownership question. That work belongs in the estimate even if no new code is required.
McKinsey's 2026 readiness survey draws a useful distinction between personal readiness and organizational readiness. In its sample, 70% of respondents said they felt personally prepared to use AI, while only 27% of leaders said their organizations were ready to make the required shifts. The article also reports an association between workflow redesign and enterprise value in its enablement sample, and between support and training and value capture in its automation sample. These are survey associations, not causal laws or planning multipliers. (McKinsey's readiness survey)
The EU publication on public administrations makes the same scope visible from another angle. Its report metadata says it examines individual use and the broader organizational processes required for strategic assimilation, including governance, data protection, organizational readiness, and technological sovereignty. That is why “training” should not mean a single demo session. (Publications Office of the EU report)
The change-management line should name the behavior that must change:
- a reviewer must approve or correct an AI result;
- an operator must trust a new source of context;
- a manager must change a target or queue rule;
- an escalation owner must handle the exceptions;
- a team must learn how to report failures and improve the workflow.
TryUncle has made a related constraint concrete for me. An AI agent that watches a screen and annotates it live has to respect timing and human approval as product constraints. They cannot be postponed to a handover document. (F-tryuncle)
Repair the estimate with a ten-surface worksheet
Repair the quote by turning each missing boundary into a deliverable with an owner, dependency, effort, and completion check. The worksheet below is the smallest useful repair because it makes an apparently complete feature estimate auditable before anyone argues about the number.
For each row, require five fields: deliverable, effort, owner, dependency, and completion evidence. “Included” is not a deliverable.
| Surface | Minimum question | Completion evidence |
|---|---|---|
| Model or feature build | What behavior is in scope, and what is excluded? | Acceptance cases pass |
| Data readiness | Which sources, fields, permissions, freshness checks, and cleanup are required? | Approved sample and access check |
| Integration | Which systems, identities, events, writes, retries, and failure paths are included? | End-to-end representative test |
| Workflow redesign | What changes for operators, reviewers, approvers, and escalation owners? | Agreed workflow and decision rights |
| Evaluation | Which cases, rubric, baseline, reviewer process, and release threshold apply? | Evaluation record and release decision |
| Security and governance | What data, vendor, retention, access, audit, and policy review is required? | Signed review and controls |
| Training and change management | Who must work differently, and how will independent use be checked? | Rehearsal and transfer check |
| Rollout | What is the release sequence, rollback, success condition, and expansion gate? | Rollout plan and go/no-go record |
| Support | Who handles questions, defects, exceptions, and incidents? | Named queue and escalation path |
| Ongoing ownership | Who monitors, evaluates, updates, pays for, and retires the workflow? | Owner, cadence, budget, retirement trigger |
The worksheet is deliberately boring. Boring is useful here. It turns “production-ready” from a mood into a list of obligations.
Worked estimate: a support-triage feature
This example is hypothetical. It uses person-days as planning assumptions, not money, rates, a client result, or a market average.
A 12-person B2B SaaS support team receives about 1,000 tickets each month. The proposed feature reads a ticket and linked internal guidance, drafts a summary, and suggests a routing label. A support agent approves or edits the output. The system does not send a reply or change a customer record automatically.
The scenario assumes one help-desk API, one internal knowledge source, read-only access, an existing identity provider, one working language, and a controlled release to one queue.
| Estimate surface | Model-first quote | Complete bounded release | What the days cover |
|---|---|---|---|
| Model or feature build | 8 | 8 | Prompting, output shape, interface, and error display |
| Data readiness | 2 | 4 | Sample selection, access checks, freshness, and knowledge cleanup |
| Integration | 0 | 7 | Help-desk read path, identity, retries, logs, and representative environment |
| Workflow redesign | 0 | 5 | Approval, routing rules, exception path, and operator procedure |
| Evaluation | 0 | 4 | Case set, rubric, baseline, reviewer calibration, and release gate |
| Security and governance | 0 | 3 | Data review, vendor and retention check, access scope, and audit decision |
| Training and change management | 0 | 3 | Rehearsal, guidance, feedback loop, and independent-use check |
| Rollout | 0 | 2 | Shadow mode, one-queue release, rollback, and expansion decision |
| Support | 0 | 2 | First-release triage, defect handling, and escalation setup |
| One-time total | 10 | 38 | The model-first number is not a production estimate |
| Ongoing ownership | 0 | 1.5 days/month | Evaluation refresh, source and prompt updates, monitoring, and owner review |
The complete estimate is 28 person-days above the model-first quote. That is the worked result of this scenario, not a claim about typical implementations. Integration, workflow redesign, evaluation, and the operating boundary create most of the difference.
The explicit exclusions are model training or fine-tuning, help-desk replacement, historical migration, multilingual support, 24/7 service levels, automatic customer replies, regulated decision-making, and cleanup beyond the agreed sample.
Uncertainty is medium-high around API permissions, knowledge-source quality, approval behavior, and security review. No percentage contingency is applied. The assumptions themselves are the uncertainty register, and each one has a check before the release estimate becomes a commitment.
Verify the revised estimate before funding it
Verify the artifact in four passes: recompute the arithmetic, check that the scenario boundary is explicit, confirm that each operational obligation has completion evidence, and test the decision rule against the stated exclusions. This verifies the worksheet and its planning logic. It does not claim that a real support team has completed the release.
| Verification check | Evidence in this artifact | Result |
|---|---|---|
| Arithmetic | The complete rows sum to 38 person-days, and 38 minus the 10-day model-first quote equals 28 additional one-time days. | Pass |
| Bounded scope | The scenario names 12 support staff, about 1,000 monthly tickets, one help-desk API, one knowledge source, read-only access, one language, and one queue. | Pass |
| Operational proof | The worksheet requires an end-to-end integration test, agreed workflow, evaluation record, signed controls, rehearsal, rollout record, named support path, and owner cadence. | Pass as a release checklist; not yet a live result |
| Honest decision | Go, revise, and stop conditions are tied to funding, access, ownership, human approval, evaluation, governance, and rollback. | Pass |
If any pass condition is still an assumption when the proposal is signed, keep the decision at revise. Verification evidence is not a promise that the system will work. It is the record that the estimate has exposed what must be proved next.
Choose go, revise, or stop from the estimate
Use the completed worksheet to make the funding decision, not to defend the original quote.
Go when the buyer can fund the 38-day bounded release, grant read-only access, name an owner for 1.5 days per month, and accept the human-approval boundary.
Revise when the buyer can fund only 10 to 37 days, or when an assumption is unverified. Reduce the scope to a shadow-mode prototype with no operational promise. Re-estimate after access, data, and workflow checks.
Stop when nobody owns the workflow, representative data or integration access is unavailable, or the buyer expects autonomous customer-facing action without evaluation, governance, and rollback work.
This is also where AI workflow delivery fits in the wider implementation plan. After the delivery boundary is clear, use the AI platform total-cost worksheet to separate implementation effort from recurring platform and operating cost.
When a small estimate is honest
A small estimate is honest when the scope is genuinely disposable.
Call it a prototype if it uses sample data, has no production integration, has no live users, has no business-system writes, has no governance approval, and has no handover obligation. State those exclusions in the proposal. The reader can then compare a prototype price with a release price without pretending they are the same purchase.
The failure is not that a 10-day prototype exists. The failure is calling it a 10-day implementation.
Before signing, copy the ten rows into the proposal. If a row has no effort, owner, dependency, and completion evidence, mark it open. If the open row changes the go, revise, or stop decision, the estimate is not ready.
If you want help turning an AI proposal into a capability your team can own, Marius Manolachi's AI consulting and tutoring work follows the same boundary: make existing people capable of building and operating AI products on their own work.
Questions people ask next
Should ongoing ownership be part of an AI implementation estimate?
Yes. Name the post-launch owner, review cadence, monitoring, evaluation refresh, source updates, vendor costs, and retirement trigger. If those are not funded, the estimate ends at handover rather than at an owned operational outcome.
What if I only want an AI prototype?
Keep the small estimate, but label it as a prototype. State that it excludes production integration, live users, governance approval, rollout, support, and handover. Re-estimate before treating the prototype as a delivery commitment.