Field note · implementation
Why Does AI Fail on Training Needs Analysis After the Demo?
A six-case controlled test shows why transcript-only AI TNAs recommend training when the missing evidence points to workflow, policy, or tool repair.

The demo is persuasive because it turns a messy conversation into a clean recommendation. That polish is also the risk. A transcript can contain a complaint, a proposed training solution, and a confident manager without containing proof that training is the right response.
I ran the same fixed needs-analysis task on six controlled, non-client cases. The only change was whether the model saw the demo transcript alone or the transcript plus evidence about the role, performance, learner, policy, tool, and organizational constraints.
What did the controlled test find?
The evidence packet changed the intervention in five of six cases. Transcript-only outputs recommended training in all six cases, including cases where the supplied evidence showed a hidden field, conflicting policy versions, stale retrieval, delayed access, or an unresolved metric owner.
| Condition | Score | Immediate recommendation |
|---|---|---|
| Demo transcript only | 14/30 | Training in all six cases |
| Transcript plus source evidence | 30/30 | Workflow, policy, data, or process repair in five cases; targeted training in one |
The score is not a benchmark for AI in general. It is the result of this dated, six-case fixture. The sourceable finding is narrower and more useful: a transcript-only TNA can sound complete while skipping the evidence needed to distinguish a performance gap from a broken operating condition.

The test used one model, one output per condition, and a binary five-part rubric. The full prompts, case inputs, raw outputs, and limitations are archived in the research record for this article. The cases are controlled test material, not client outcomes.
Why does the demo transcript create the wrong training story?
A demo transcript often contains a solution-shaped statement before it contains a problem definition. When the manager says, “We need training,” the model can treat that sentence as a constraint instead of a hypothesis.
That breaks the first part of needs analysis. The CDC needs-analysis sequence starts with the goal and the observable gap, then collects data to find the gap's source. It explicitly lists policy, technology, organization, and standard operating procedures as possible alternatives to training.
The transcript-only outputs in this test did three things that made the failure hard to notice:
- They restated a plausible symptom as a performance gap.
- They named missing evidence, then recommended training anyway.
- They made the next step a training activity instead of an evidence-collection activity.
That is a polished form of premature closure. The answer looks cautious because it contains an uncertainty section. Its recommendation still assumes the conclusion.
I have seen the same shape in teaching. When I taught product managers to move from writing specifications to building and shipping, the recurring failure was usually an undefined “done,” not the model. A TNA has the same dependency. If “done” means “the report contains a gap, cause, and training plan,” the model can satisfy the format while never proving that training would change the work. That observation is part of Marius Manolachi's locked teaching record, not a measured rate.
Which evidence changes the training recommendation?
The smallest useful evidence packet has six parts. Each part closes a different escape route for a training-shaped answer.
| Evidence | Question it answers | Failure it can expose |
|---|---|---|
| Role and expected performance | What should this person do, and what counts as done? | A vague complaint with no observable outcome |
| Observed performance | What actually happened in recent work or a replay? | A perception presented as a measured gap |
| Learner evidence | Does the person lack knowledge, skill, or behavior? | Training for people who already know the rule |
| Policy and process | What rule, approval, or handoff governs the work? | Conflicting or missing operating instructions |
| Tool and data context | Can the person access and execute the required step? | Hidden fields, stale sources, permissions, or timing failures |
| Organizational constraints | Who owns the decision, and what can change? | Training used to hide an ownership or resource problem |
This is consistent with the CDC Quality Training Standards, which say a needs assessment should validate that training is needed and identify learners and delivery barriers. It is also why a learner survey is only one layer of evidence. A 2025 medical-education study surveyed 68 faculty and 506 students, found different familiarity and preferences, and used the results to design two parallel programmes. The authors also state that their learner analysis did not diagnose current AI competencies and that the single-school sample limits generalisability. Read the study.
The distinction matters after an AI demo because a demo usually shows capability, not context. It proves that a model can produce a checklist, explanation, lesson, or guided walkthrough. It does not prove that the learner lacks the skill, that the policy is stable, that the tool exposes the required state, or that the organization has named an owner.
The UK government's Skills for AI programme takes a similarly broad view by combining workshops, survey evidence, and real-world case studies into employer guidance. That is a useful reminder that AI capability is an organizational design problem, not only a prompt-training problem.
How can you reproduce the failure before approving training?
Run the same request twice. Hold the model, prompt, output schema, and rubric constant. Change only the evidence condition.
- Freeze the demo transcript. Keep the manager's complaint and proposed solution exactly as shown. Do not add your interpretation to condition A.
- Write the evidence packet. Add role, expected performance, observed examples, learner evidence, policy and process, tool and data context, and constraints for condition B.
- Force a decision. Require the model to choose TRAIN, REPAIR WORKFLOW/POLICY/TOOL, or COLLECT EVIDENCE. Allow a combined action only when the evidence supports it.
- Score five observable criteria. Check the performance gap, cause distinction, uncertainty, missing-evidence request, and defensible next step. Do not score polish or confidence.
- Inspect the decision change. The useful output is not the average score. It is the case where the recommendation changes and the evidence that caused the change.
The fixed test prompt in this run told the model to treat the transcript as a report of what was said, not as proof that the proposed solution was correct. That instruction did not eliminate the failure. With transcript-only inputs, the model still produced training-shaped next steps. The evidence packet supplied the facts needed to constrain the decision.
This is also a governance boundary. NIST describes the AI Risk Management Framework as voluntary guidance for incorporating trustworthiness into the design, development, use, and evaluation of AI systems. In practice, that means the TNA output should be treated as decision support with a visible evidence boundary, not as an approval to buy training.
What does the failure matrix say?
The matrix makes the diagnosis visible before anyone argues about prompt quality.
| Case | Transcript-only failure | Evidence-grounded diagnosis | First action |
|---|---|---|---|
| Proposal intake | Calls missing fields a prompt-training problem | Coordinators know the rule; the cost-center field is hidden | Repair the intake tool |
| Support escalation | Blames weak prompts for escalations | Product changes are missing from retrieval; a smaller rule gap may remain | Repair sources, then test the rule |
| Procurement exceptions | Calls bypasses a policy-knowledge gap | Two policy versions conflict and the portal has no exception route | Repair policy and portal |
| Analytics reporting | Calls inconsistent metrics a SQL skill gap | Definitions, source views, and ownership disagree | Repair governance and data source |
| Workspace onboarding | Calls failed setup an onboarding-training gap | Access arrives late; setup works once access exists | Repair provisioning handoff |
| Incident severity | Suggests guided training | Current policy and tool are available; responders fail a specific scenario | Run targeted training |
The fifth case is especially important. Training often becomes a substitute for fixing a dependency that sits outside the learner's control. The model cannot see that dependency from a demo transcript unless someone supplies the access timing, ownership, and observed setup result.
What is the worked decision rule?
Do not approve training from a transcript-only TNA when the output cannot connect the observed gap to a learner-controlled knowledge, skill, or behavior change. Route the case to evidence collection or workflow, policy, tool, data, or ownership repair first.
For the proposal-intake case, the decision record is:
| Decision field | Record |
|---|---|
| Outcome | Complete proposals reach finance and legal review with required fields present |
| Observed gap | 7 of 10 recent drafts missed the cost center |
| Learner evidence | Coordinators named the requirement and passed the checklist check |
| System evidence | The field was hidden and no missing-field warning existed |
| Decision | Repair workflow/tool, not broad training |
| Verification | Make the field visible or add a warning, then replay ten proposals |
| Training veto | Do not schedule training until the repaired workflow is tested |
| Escalation | If errors remain, collect a targeted performance sample and train that behavior |
This record follows the CDC sequence: define the gap, collect data about people and systems, analyze causes, decide whether training is a solution, and only then conduct a training needs analysis. The training plan is downstream of the decision. The demo reversed that order.
The same rule applies to an AI-generated training proposal, a consultant's workshop recommendation, or an internal team's rollout brief. A polished recommendation is not evidence of a training need.
When is training still the right answer?
Training remains the right first action when the evidence shows a specific performance gap, the current policy and tools make the desired behavior possible, and the gap is plausibly changed by knowledge, skill, or practice.
That is what happened in C6. Responders had the current severity rubric, access to the incident tool, and the escalation contact. They classified ordinary latency incidents correctly but misclassified credential-exposure scenarios in a blind replay. The next step was targeted scenario practice, followed by the same case family after training. The human escalation decision stayed in place.
This is the principal exception to the veto. The rule is not “never train after a demo.” It is “do not let the demo choose training before evidence has isolated a trainable gap.”
What does this test not prove?
It does not prove that AI always fails at needs analysis, that transcript-only inputs always produce training recommendations, or that six cases predict workplace performance. It is one model, one pass, six controlled synthetic cases, and two conditions. The cases were designed to test known distinctions. No temperature or seed was available, and no second model or rerun was used.
The useful claim is bounded: in this fixture, adding source evidence changed five of six decisions and raised the rubric score from 14/30 to 30/30. The full archive lets another reviewer challenge the case design, rerun the prompts, or add counterexamples.
If you are evaluating a training demo, ask for the evidence packet before asking for more polished output. If your team needs help turning that packet into a safe implementation decision, start with the AI workflow implementation guide, then compare the recommendation against the failure clinic on AI evals that pass while users still fail. Marius Manolachi's AI consulting and tutoring work follows the same capability boundary: make the existing team able to build and judge the work on its own.
Continue with a related field note
Questions people ask next
Can a demo transcript ever be enough for a training needs analysis?
It can identify a question or a perceived gap, but it is not enough to choose training when the decision has operational consequences. Add observed performance, expected performance, learner evidence, system and policy context, and constraints first.
What should I do when the AI recommends training but the cause is unclear?
Do not approve the training yet. Ask for the smallest evidence packet that can separate a skill gap from a policy, process, tool, data, or ownership problem, then rerun the analysis with the same rubric.