Training Measurement

Confidence Is Not Competence, and Our Own Study Says So

A colleague pointing at a dashboard during a review: the row that gets collected is not the row that counts
Quick answer

Many training evaluation forms ask a version of the same question: how confident do you feel now? In the study behind our method, the answer went up after a slide deck, up after a video and up after a short online rehearsal with a role-play coach, with no significant difference between the three. The blind-rated conversations that followed told a different story. Here are the confidence figures next to the delivery scores, and what we record instead.

The numbers, in full

Fifty-seven participants learned the same feedback model, Situation, Behaviour, Impact, in one of three ways: an instructional slideshow, an instructional video, or an online coaching session with a role-play coach, each format capped at 20 minutes. Each then held a simulated feedback conversation, which two independent assessors scored blind on a 1 to 7 scale. Confidence was self-rated on a 1 to 7 scale at three points: before the intervention, straight after it, and after the simulation. The full method and its limits are on our research page.

Every group climbed at every step. The format had no significant effect on confidence (p = .958), and the interaction between time and format was not significant either (p = .541). Read that carefully: it is an absence of evidence for a difference between the formats, not evidence that they are the same. We will come back to that distinction, because we got it wrong ourselves.

Now the other measure. Blind-rated delivery averaged 5.25 out of 7 for the rehearsal group, 4.26 for slides and 3.82 for video. Rehearsal was significantly better than both passive formats, with Cohen's d of 0.75 against slides and 1.08 against video, and slides and video did not differ from each other. Put the two rows side by side. Final confidence sat between 5.38 and 5.53 across the three groups. Rated delivery sat between 3.82 and 5.25. The people who had watched a video reported almost the same final confidence as the people who had rehearsed, and were rated a long way behind them.

The study also checked whether a participant's final confidence rating predicted their assessed performance. It did not. The paper discusses this gap between confidence and performance through the lens of the Dunning-Kruger effect, but cautions that, without a control condition, the effect does not fully apply here. The study did not rely on the question "how confident are you?" alone: it asked it at three points and had two independent assessors score a recording of each conversation.

Why the form still gets sent

Confidence is cheap to collect and it goes up. That is the whole problem. The TalentLMS 2026 L&D Benchmark Report, a September 2025 survey of 101 HR managers in the United States, finds business impact selected by 37% of them as a measure of L&D success, and satisfaction with training by 28%. The Department for Education's Employer Skills Survey 2024 puts UK employer training spend at £1,700 per employee, £53.0bn in total. A confidence score is a thin receipt for that.

Even the framework most evaluation forms claim to follow does not treat confidence as behaviour. Kirkpatrick Partners define Level 2, Learning, as the degree to which participants acquire "the intended knowledge, skills, attitude, confidence, and commitment" they need, and Level 3, Behaviour, as the degree to which they actually perform the critical behaviours in their own environment. Confidence is filed under Level 2 by the people who maintain the model. A confidence question is a Level 2 instrument at best, and in our study a participant's final confidence rating did not predict how their simulated conversation was scored. Our explainer on the Kirkpatrick model sets out the four levels.

The uncomfortable pair of rows: final confidence ran from 5.38 to 5.53 across three groups. Blind-rated delivery ran from 3.82 to 5.25. Only one of those rows is what a standard evaluation form collects.

What the study cannot carry

This is one study of 57 people, run in 2020 for a taught module at UCL, unpublished and not replicated. It has no control condition, so it cannot say how much anyone improved, and no measurement of ability before the intervention, only of confidence. The practice condition was a short online coaching session, not our programme, and we do not present it as a measurement of what we deliver. The delivery score comes from one simulated conversation held immediately afterwards, which is not the same thing as behaviour at work a month later. What the study can say is narrow and, within that, clear: three formats produced confidence scores with no significant difference between them, and delivery scores that differed significantly.

We have had to correct our own wording on this. Earlier drafts of our research page described the rise in confidence as identical across the groups, and a later draft called it equal. Both overstate it. A non-significant difference is a failure to find a difference, not proof of sameness, and the phrasing now on the page is "with no significant difference between them". If we want buyers to read evaluation data carefully, we have to read our own the same way.

What we record instead

In our new manager programme the fundamentals are watched while people work, one manager at a time, rather than collected on a satisfaction form at the end. Each one-to-one conversation with an actor is played, debriefed, rewound to the moment that mattered and played again, and only the second attempt shows whether the manager can do something other than their default when the pressure comes back. The questions are behavioural: did the manager notice their own reaction and slow down, did they make it safe for the actor to push back, did they leave with an owner and a date. None of those is a feeling.

Being straight about the limit of that too: what someone does in the room under pressure is still not what they do on an ordinary Tuesday. That is Kirkpatrick Level 3, behaviour on the job, and it is a separate piece of measurement, set out in our guide to measuring behaviour change. What we do not do is ask people how confident they feel and call the answer evidence. Our own study says that answer went up in every condition, including the two whose delivery was rated significantly weaker.

Confidence Is Not Competence, and Our Own Study Says So: The Takeaways

In the randomised 2020 UCL study by co-founder Ben Laumann (n = 57), participants rated their own confidence before the intervention, after it and after a simulated feedback conversation. Confidence rose in all three groups, and the format made no significant difference to it (p = .958). Blind-rated delivery of the same conversation did differ significantly, with rehearsal ahead of slides (Cohen's d = 0.75) and video (d = 1.08). A post-training confidence score is therefore not a measure of what someone can do. We watch the behaviour instead.

  • In the randomised 2020 UCL study (n = 57), self-rated confidence rose in all three groups, from 4.46, 4.69 and 4.27 to 5.39, 5.38 and 5.53 on a 1 to 7 scale, with no significant difference between formats (p = .958).
  • Blind-rated delivery did differ: 5.25 after rehearsal, 4.26 after slides, 3.82 after video, with Cohen's d of 0.75 and 1.08 in favour of rehearsal.
  • Final confidence did not predict assessed performance. A confidence score is a Level 2 measure by Kirkpatrick's own definition, not evidence of behaviour.
  • The study is small, unpublished, has no control condition and tested a short online coaching session, not our programme. It shows what happened to confidence and delivery on the day, and no more.
  • We watch the first attempt against the replay in the room. On-the-job behaviour is a separate measurement, set out in our guide, rather than a question about how confident people feel.

Frequently Asked Questions

Does higher confidence after training mean the training worked?

Not on its own. In our 2020 UCL study, confidence rose after every format, including the two whose delivery was rated significantly weaker, and a participant's final confidence did not predict their assessed performance. Confidence is worth knowing, and it can be a goal in its own right, but it is not a measure of what someone can do.

Why publish the confidence figures if they make the study look less tidy?

Because they are the point. The delivery result says rehearsal produced better performance. The confidence result says a satisfaction form would have reported a success in all three conditions. Together they explain why we measure observed behaviour and why we tell buyers to be wary of evaluation data that stops at how people felt.

What should an evaluation form ask instead?

Ask it less and watch more. In the room, score the behaviour someone actually produced, ideally on a second attempt after feedback. On the job, agree in advance which behaviour should change, who will observe it and when, which is Kirkpatrick Level 3. Our guide to measuring behaviour change sets that out without needing a research department.

Sources: Sidestream, The Study Behind the Method · Kirkpatrick Partners, The Kirkpatrick Model · TalentLMS, The 2026 Annual L&D Benchmark Report (September 2025) · Department for Education, Employer Skills Survey 2024

Continue Reading

Related Articles

Training Measurement

The 37% Gap: Most Training Still Isn't Measured for Impact

Measurement & ROI

How to Measure the ROI of Behaviour Change Training

Leadership Development

60% Have the Training. Only 59% Have the Time.

Sidestream

Take Action

Bring Us Your
People Problem

Free 30-minute diagnostic call. No deck, no hard sell, just an honest conversation about whether we can help.