A one-line summary of evidence leaves out the method section. Our line is that a randomised study at UCL found rehearsal beat slides and video, and it is true. The method section is where it gets interesting, because the rehearsal was a twenty-minute Zoom call.
The short version
The study compared three ways of teaching one feedback model, rehearsal with a role-play coach, a slideshow and a video, with the conversations rated blind. Rehearsal came out ahead. But there was no control group, skill was measured once, on the day, and the rehearsal was a twenty-minute Zoom session. So it is evidence that practising with a responsive partner beats watching, not evidence for our programme as it runs now. We build on the first and measure the rest.
What was actually tested
Fifty-seven participants were randomly allocated to one of three conditions. All three received the same content on the SBI feedback model (Situation, Behaviour, Impact), and all three were timed to a maximum of twenty minutes. Two groups watched: a slideshow or an instructional video. The third had, in the paper's words, "a ZOOM-based coaching session between a role play coach and participants". They went through the model up to three times and practised its components with the coach, without being coached on how they spoke or looked.
Then every participant gave feedback to the same actor, over Zoom, in a recorded conversation, and two independent assessors rated each recording on the three components without knowing which group anyone had been in. That blind rating is the part of the design we would defend hardest.
What it shows
The assessors rated the rehearsal group significantly higher than either watching group. Cohen's d was 1.08 against video and 0.75 against the slideshow, a large and a medium-to-large effect. The slideshow and the video did not differ significantly from each other, so the study gives no grounds for preferring one passive format over the other.
The second finding is confidence. Participants rated how confident they felt before the intervention, after it, and after the conversation. On a seven-point scale the slideshow group went from 4.46 to 5.00 to 5.39, the video group from 4.69 to 5.13 to 5.38, and the practice group from 4.27 to 5.33 to 5.53 (Laumann (2020), UCL, unpublished research paper, Figure 4). Confidence rose over time in every group, and the format made no significant difference to it (p = .958 for modality, p = .541 for the interaction). Rated delivery differed significantly between the groups; reported confidence did not.
What it does not show
Most of these limits are named in the paper itself, a supervised piece of university work from 2020 that is neither peer-reviewed nor independently replicated. None of them makes the finding wrong. All of them narrow what it is a finding about.
- One skill. The SBI model was chosen because it breaks into three observable parts. Nothing in the data says rehearsal beats slides for a negotiation, a disciplinary meeting or a restructure.
- One session. Twenty minutes, once, with up to three passes through the model. A programme that runs over weeks is a different intervention, and the study is silent on it.
- No control group, and no baseline. The paper notes the "absence of a control condition" and, because of it, assumes that everyone improved to some degree. What the design cannot show is by how much, or whether the two watching formats added anything beyond the introduction. The only thing measured beforehand was confidence, not skill.
- No transfer measured. Every assessed conversation took place straight after the session. Whether anyone used the model a week later, or with a real colleague, was never observed; the paper names lasting learning impact, alongside its own debrief stage, as the main point for future work.
- A senior sample. Participants averaged nearly twenty years in professional practice, which the paper suggests may have narrowed the gap between conditions. Nothing here is about new managers or frontline staff.
- Not our programme. The practice condition was a coaching session with a role-play coach on Zoom; the professional actor with pre-agreed responses was the counterpart in the assessed conversation, which all three groups had, and nothing was observed weeks later. The study tells us which lever to pull. It does not certify the machine we built around that lever.
Why we say this on our own pages
For a while our own pages described the study with a word that implies a control group. The paper says plainly that there was none, so we corrected how we described our own study; we now say randomised, which is what it was. A buyer who reads the method section and finds a limit we did not mention will assume the rest of the site is written the same way. The limits are what make the claim credible, so they belong on the page next to it.
What follows for programme design
The study answers one question, for one skill, measured on the day: rehearsal led to higher-rated delivery than watching, and the two watching formats did not differ significantly from each other. Three things in how we work come directly from that, and they are set out on our research page: participants perform rather than watch, the counterpart is someone trained to hold their ground, and outcomes are recorded as observed behaviour rather than self-report.
Everything past that point is design, not evidence. The study used a fictitious colleague; we offer two routes into the rehearsal: a scripted scene built around the pressures the organisation has described, or Real Play, where a participant brings a conversation they actually have coming and briefs the actor on it themselves. The study measured on the day; we would build an observation point weeks later, because a result that only exists in the room is a Kirkpatrick Level 2 result, not the Level 3 change in behaviour at work. The study used one short rehearsal; we would build in return visits to the same conversation, since nothing in our own data says one pass is enough.
None of that is proven by the study. It is what we think follows from it, built to be measured rather than asserted. If a provider tells you their evidence covers everything they sell, ask for the method section.
Where Sidestream fits
We are a behaviour change consultancy that combines organisational psychology with immersive theatre. Our role-play training and immersive simulation training rest on the finding above and are measured against the limits above. The full paper is available on request. Book a free 30-minute call and bring the conversation your managers keep putting off.
What a Twenty-Minute Zoom Role Play Does and Does Not Prove: The Takeaways
Our UCL study compared a twenty-minute Zoom rehearsal with a slideshow and a video, and rehearsal came out ahead. Here is what that proves, the limits the paper itself names, and what we build on top that still has to be measured.
- Randomised, 57 participants, three formats, one feedback model, rated blind by two assessors: rehearsal beat video (d = 1.08) and slides (d = 0.75).
- No control condition, no baseline measure of skill, no follow-up: the study ranks formats on the day and cannot say how much anyone improved or whether it lasted.
- Confidence rose in all three groups with no significant difference between them (p = .958 for modality, p = .541 for the interaction), while rated delivery differed significantly.
- The practice condition was a twenty-minute role-play coaching session on Zoom, so the study is not a measurement of Sidestream's programme.
- Everything beyond the finding (professional actors and scripted scenes, Real Play, later observation, repeated passes) is design to be measured, not evidence to be cited.
Frequently Asked Questions
Does the study prove immersive training works?
It shows that rehearsing one feedback model with a role-play coach for twenty minutes produced better blind-rated delivery than a slideshow or a video, with a large effect against video and a medium-to-large effect against slides. It does not test a full programme, a control group, other skills or transfer to the job.
The participants were randomised, so why say there was no control group?
Randomisation means the three groups were assigned by chance, which protects the comparison between them. A control group would have been a fourth group that received no training, which is what you need to say how much anyone improved. The study had the first and not the second.
Can I read the paper?
Yes. It is an unpublished, supervised piece of UCL research from 2020, available on request. Our research page summarises the design, the numbers and the limits, and our guide to the Kirkpatrick model explains why we measure at Level 3 rather than by satisfaction.
Sources: Sidestream, The Study Behind the Method · Sidestream, Role-Play Training · Sidestream, What Is the Kirkpatrick Model?
