Field note · implementation

Why Does AI Fail on Training Needs Analysis After the Demo?

A six-case controlled test shows why transcript-only AI TNAs recommend training when the missing evidence points to workflow, policy, or tool repair.

10 minute read
  • AI training
  • Needs analysis
  • AI implementation
Illustration of an AI training needs analysis changing after source evidence is added

The demo is persuasive because it turns a messy conversation into a clean recommendation. That polish is also the risk. A transcript can contain a complaint, a proposed training solution, and a confident manager without containing proof that training is the right response.

I ran the same fixed needs-analysis task on six controlled, non-client cases. The only change was whether the model saw the demo transcript alone or the transcript plus evidence about the role, performance, learner, policy, tool, and organizational constraints.

What did the controlled test find?

The evidence packet changed the intervention in five of six cases. Transcript-only outputs recommended training in all six cases, including cases where the supplied evidence showed a hidden field, conflicting policy versions, stale retrieval, delayed access, or an unresolved metric owner.

ConditionScoreImmediate recommendation
Demo transcript only14/30Training in all six cases
Transcript plus source evidence30/30Workflow, policy, data, or process repair in five cases; targeted training in one

The score is not a benchmark for AI in general. It is the result of this dated, six-case fixture. The sourceable finding is narrower and more useful: a transcript-only TNA can sound complete while skipping the evidence needed to distinguish a performance gap from a broken operating condition.

Illustration of an AI training recommendation being rerouted to workflow repair after evidence is added

The test used one model, one output per condition, and a binary five-part rubric. The full prompts, case inputs, raw outputs, and limitations are archived in the research record for this article. The cases are controlled test material, not client outcomes.

Why does the demo transcript create the wrong training story?

A demo transcript often contains a solution-shaped statement before it contains a problem definition. When the manager says, “We need training,” the model can treat that sentence as a constraint instead of a hypothesis.

That breaks the first part of needs analysis. The CDC needs-analysis sequence starts with the goal and the observable gap, then collects data to find the gap's source. It explicitly lists policy, technology, organization, and standard operating procedures as possible alternatives to training.

The transcript-only outputs in this test did three things that made the failure hard to notice:

  • They restated a plausible symptom as a performance gap.
  • They named missing evidence, then recommended training anyway.
  • They made the next step a training activity instead of an evidence-collection activity.

That is a polished form of premature closure. The answer looks cautious because it contains an uncertainty section. Its recommendation still assumes the conclusion.

I have seen the same shape in teaching. When I taught product managers to move from writing specifications to building and shipping, the recurring failure was usually an undefined “done,” not the model. A TNA has the same dependency. If “done” means “the report contains a gap, cause, and training plan,” the model can satisfy the format while never proving that training would change the work. That observation is part of Marius Manolachi's locked teaching record, not a measured rate.

Which evidence changes the training recommendation?

The smallest useful evidence packet has six parts. Each part closes a different escape route for a training-shaped answer.

EvidenceQuestion it answersFailure it can expose
Role and expected performanceWhat should this person do, and what counts as done?A vague complaint with no observable outcome
Observed performanceWhat actually happened in recent work or a replay?A perception presented as a measured gap
Learner evidenceDoes the person lack knowledge, skill, or behavior?Training for people who already know the rule
Policy and processWhat rule, approval, or handoff governs the work?Conflicting or missing operating instructions
Tool and data contextCan the person access and execute the required step?Hidden fields, stale sources, permissions, or timing failures
Organizational constraintsWho owns the decision, and what can change?Training used to hide an ownership or resource problem

This is consistent with the CDC Quality Training Standards, which say a needs assessment should validate that training is needed and identify learners and delivery barriers. It is also why a learner survey is only one layer of evidence. A 2025 medical-education study surveyed 68 faculty and 506 students, found different familiarity and preferences, and used the results to design two parallel programmes. The authors also state that their learner analysis did not diagnose current AI competencies and that the single-school sample limits generalisability. Read the study.

The distinction matters after an AI demo because a demo usually shows capability, not context. It proves that a model can produce a checklist, explanation, lesson, or guided walkthrough. It does not prove that the learner lacks the skill, that the policy is stable, that the tool exposes the required state, or that the organization has named an owner.

The UK government's Skills for AI programme takes a similarly broad view by combining workshops, survey evidence, and real-world case studies into employer guidance. That is a useful reminder that AI capability is an organizational design problem, not only a prompt-training problem.

How can you reproduce the failure before approving training?

Run the same request twice. Hold the model, prompt, output schema, and rubric constant. Change only the evidence condition.

  1. Freeze the demo transcript. Keep the manager's complaint and proposed solution exactly as shown. Do not add your interpretation to condition A.
  2. Write the evidence packet. Add role, expected performance, observed examples, learner evidence, policy and process, tool and data context, and constraints for condition B.
  3. Force a decision. Require the model to choose TRAIN, REPAIR WORKFLOW/POLICY/TOOL, or COLLECT EVIDENCE. Allow a combined action only when the evidence supports it.
  4. Score five observable criteria. Check the performance gap, cause distinction, uncertainty, missing-evidence request, and defensible next step. Do not score polish or confidence.
  5. Inspect the decision change. The useful output is not the average score. It is the case where the recommendation changes and the evidence that caused the change.

The fixed test prompt in this run told the model to treat the transcript as a report of what was said, not as proof that the proposed solution was correct. That instruction did not eliminate the failure. With transcript-only inputs, the model still produced training-shaped next steps. The evidence packet supplied the facts needed to constrain the decision.

This is also a governance boundary. NIST describes the AI Risk Management Framework as voluntary guidance for incorporating trustworthiness into the design, development, use, and evaluation of AI systems. In practice, that means the TNA output should be treated as decision support with a visible evidence boundary, not as an approval to buy training.

What does the failure matrix say?

The matrix makes the diagnosis visible before anyone argues about prompt quality.

CaseTranscript-only failureEvidence-grounded diagnosisFirst action
Proposal intakeCalls missing fields a prompt-training problemCoordinators know the rule; the cost-center field is hiddenRepair the intake tool
Support escalationBlames weak prompts for escalationsProduct changes are missing from retrieval; a smaller rule gap may remainRepair sources, then test the rule
Procurement exceptionsCalls bypasses a policy-knowledge gapTwo policy versions conflict and the portal has no exception routeRepair policy and portal
Analytics reportingCalls inconsistent metrics a SQL skill gapDefinitions, source views, and ownership disagreeRepair governance and data source
Workspace onboardingCalls failed setup an onboarding-training gapAccess arrives late; setup works once access existsRepair provisioning handoff
Incident severitySuggests guided trainingCurrent policy and tool are available; responders fail a specific scenarioRun targeted training

The fifth case is especially important. Training often becomes a substitute for fixing a dependency that sits outside the learner's control. The model cannot see that dependency from a demo transcript unless someone supplies the access timing, ownership, and observed setup result.

What is the worked decision rule?

Do not approve training from a transcript-only TNA when the output cannot connect the observed gap to a learner-controlled knowledge, skill, or behavior change. Route the case to evidence collection or workflow, policy, tool, data, or ownership repair first.

For the proposal-intake case, the decision record is:

Decision fieldRecord
OutcomeComplete proposals reach finance and legal review with required fields present
Observed gap7 of 10 recent drafts missed the cost center
Learner evidenceCoordinators named the requirement and passed the checklist check
System evidenceThe field was hidden and no missing-field warning existed
DecisionRepair workflow/tool, not broad training
VerificationMake the field visible or add a warning, then replay ten proposals
Training vetoDo not schedule training until the repaired workflow is tested
EscalationIf errors remain, collect a targeted performance sample and train that behavior

This record follows the CDC sequence: define the gap, collect data about people and systems, analyze causes, decide whether training is a solution, and only then conduct a training needs analysis. The training plan is downstream of the decision. The demo reversed that order.

The same rule applies to an AI-generated training proposal, a consultant's workshop recommendation, or an internal team's rollout brief. A polished recommendation is not evidence of a training need.

When is training still the right answer?

Training remains the right first action when the evidence shows a specific performance gap, the current policy and tools make the desired behavior possible, and the gap is plausibly changed by knowledge, skill, or practice.

That is what happened in C6. Responders had the current severity rubric, access to the incident tool, and the escalation contact. They classified ordinary latency incidents correctly but misclassified credential-exposure scenarios in a blind replay. The next step was targeted scenario practice, followed by the same case family after training. The human escalation decision stayed in place.

This is the principal exception to the veto. The rule is not “never train after a demo.” It is “do not let the demo choose training before evidence has isolated a trainable gap.”

What does this test not prove?

It does not prove that AI always fails at needs analysis, that transcript-only inputs always produce training recommendations, or that six cases predict workplace performance. It is one model, one pass, six controlled synthetic cases, and two conditions. The cases were designed to test known distinctions. No temperature or seed was available, and no second model or rerun was used.

The useful claim is bounded: in this fixture, adding source evidence changed five of six decisions and raised the rubric score from 14/30 to 30/30. The full archive lets another reviewer challenge the case design, rerun the prompts, or add counterexamples.

If you are evaluating a training demo, ask for the evidence packet before asking for more polished output. If your team needs help turning that packet into a safe implementation decision, start with the AI workflow implementation guide, then compare the recommendation against the failure clinic on AI evals that pass while users still fail. Marius Manolachi's AI consulting and tutoring work follows the same capability boundary: make the existing team able to build and judge the work on its own.

Questions people ask next

Can a demo transcript ever be enough for a training needs analysis?

It can identify a question or a perceived gap, but it is not enough to choose training when the decision has operational consequences. Add observed performance, expected performance, learner evidence, system and policy context, and constraints first.

What should I do when the AI recommends training but the cause is unclear?

Do not approve the training yet. Ask for the smallest evidence packet that can separate a skill gap from a policy, process, tool, data, or ownership problem, then rerun the analysis with the same rubric.