Supporting research — roleplay

What rehearsing a conversation does, and does not, do.

Roleplay has a large research literature, most of it about training health and social-care professionals to hold difficult conversations. This page states what that literature supports, what it does not, and where our own product sits outside it. Every effect size below is quoted with the thing it was compared against, because that is where most claims in this field go wrong.

This page covers roleplay. The guided tutor rests on a different literature and has its own page. Neither body of evidence stands in for the other.

Citations are author–year links into the reference list; every entry there has been checked against the published record.

Back to the home page · Questions and answers

01

The claim we make — and the one we don't.

What we claim is that this product gives a learner somewhere to practise a conversation, as often as they need, and gives their teacher the recording afterwards. Both halves are verifiable in the product itself (§12).

What we do not claim is that practising with our characters makes anyone competent, or that it improves outcomes for the people they will later care for. The best evidence for conversational training reaching real practice comes from a randomised trial of oncologists whose consultations with 2,407 real patients were videotaped (Fallowfield et al., 2002) — and that was a three-day residential course with human simulated patients, video review and a trained facilitator. Nothing in that trial licenses a claim about solo practice with software.

The honest summary of this field is that roleplay is as good as the expensive alternatives, not better than them, that the debrief carries much of the measurable benefit (§4), and that the AI-partner literature is emerging rather than established (§3). We would rather say that than be corrected by a curriculum committee.

02

Every effect size has a comparator. Most vendors drop it.

The number quoted on nearly every simulation product's website comes from one meta-analysis of 609 studies and 35,226 trainees: skills improve with an effect of about 1.09 (Cook et al., 2011). That figure is measured against no intervention at all. It says simulation beats doing nothing.

The same team ran the comparison that actually matters — simulation against other teaching — across 92 studies (Cook et al., 2012):

Satisfaction
0.59 (95% CI 0.36–0.81). The largest effect is on how much learners liked it.
Knowledge
0.30 (0.16–0.43).
Process skills
0.38 (0.24–0.52).
Behaviours
0.77 (−0.13 to 1.66) — not statistically significant.
Effects on patients
0.36 (−0.06 to 0.78) — not statistically significant.

Read the column downward. The effect is largest on satisfaction and disappears by the time it reaches the person being cared for. Simulation also cost more than what it was compared against. A separate review of 50 studies covering 16,742 patients found the same shape: a moderate patient effect against no intervention, and a non-significant one against other instruction (Zendejas et al., 2013).

This is not an argument against roleplay. It is an argument for quoting it accurately, and for the claim in §1 being about access rather than superiority.

03

Who plays the character matters less than you would expect.

The expensive option in this field is a trained simulated patient: a person paid to portray a role consistently, to a professional standard (Lewis et al., 2017). It is reasonable to ask whether anything cheaper can substitute.

A randomised trial found health professionals who practised motivational interviewing with each other reached the same measured competence as those who practised with trained simulated patients (Lane, Hood & Rollnick, 2008). A meta-analysis of ten studies found no overall difference between the two methods (Xiao & Fu, 2025, g = 0.073, ns); the one outcome that favoured simulated patients was learner self-confidence. A randomised trial of 103 medical students found peer roleplay outperformed standardised patients on understanding the other person's perspective (Bosse et al., 2012).

On a screen

Twenty-seven randomised trials comparing virtual simulation with mannequins and real people found no significant difference in communication skills (Jiang et al., 2024, SMD −0.02, 95% CI −0.62 to 0.58). The same review found virtual simulation worse for procedural skills in nursing — which is the honest shape of the finding, and the reason this product does not attempt procedural training. A randomised trial of 120 students found virtual and live simulation produced equivalent communication performance, holding at two months (Liaw et al., 2020).

One widely-quoted figure needs correcting wherever it appears. A review of virtual patients reports a large skills effect against traditional teaching (Kononowicz et al., 2019, SMD = 0.90), but those skills were clinical reasoning and procedural skills, not communication, and the figure carries very high heterogeneity. It is not evidence for a conversation product.

When the character is a language model

This is where the evidence thins, and we would rather say so than imply otherwise. The best-established use is language learning, where four independent meta-analyses converge on a moderate effect for practising with a conversational agent ((Bibauw et al., 2022, d = 0.58); (Wang et al., 2025, g = 0.48); (Lyu et al., 2025, g = 0.61); (Hou & Min, 2026, g = 0.61)). Convergence across four teams is real signal.

Everywhere else it is thin. The authoritative review of AI in vocational education found chatbots only "support self-regulation and performance", and named the field's central problem: an overreliance on pre-experimental designs and self-report measures that fuels a general "success narrative" (Deutscher et al., 2026). A review of language-model virtual patients found outcomes concentrated at the level of learner reaction and knowledge, with limited evidence of transfer to real practice, and reported hallucination rates between 0.31% and 5% (Li & Lutfi, 2026).

Two structural cautions apply to everything in this paragraph. The most candid meta-analysis in the field states that nearly every effectiveness evaluation was run by the people who built the system being evaluated, and calls the result "a very acute publication bias" (Bibauw et al., 2022). And when a meta-analysis of 62 chatbot studies corrected for publication bias, the effect fell from large to small-to-moderate (Laun & Wolff, 2025). We are a vendor publishing a research page; treat this page with the same scepticism.

Warning

No peer-reviewed trial has shown an AI roleplay partner equal to a human one

The most favourable evidence we could find is an un-refereed preprint from a team presenting their own framework. Figures circulating in this market — "275% more confident", "four times faster" — come from consultancy and vendor reports, not from peer-reviewed research, and we do not cite them. For business, sales, leadership and VET soft skills specifically, we could find no peer-reviewed controlled field experiment with objective outcomes at all. That is the least evidenced corner of this literature, and it is the corner most products like ours are sold into.

04

The effect is in the debrief, not the roleplay.

This is the most important finding on the page, and it is the reason the product is built the way it is.

A meta-analysis of 46 samples found that debriefing alone improves performance by about a quarter (Tannenbaum & Cerasoli, 2013, d = .67), consistently across simulated and real settings and across medical and non-medical groups, with larger effects when the debrief is structured and facilitated. That is as large as, or larger than, the effect of the simulation itself against other instruction (§2).

The supporting pattern is consistent. Across 109 studies, feedback was the most frequently cited feature of effective simulation, named in 47% — ahead of repetitive practice at 39%, while the validity of the simulator itself was cited in 3% (Issenberg et al., 2005). Across 177 studies and 11,511 learners, simulation with debriefing worked, but video-assisted debriefing was no better than plain debriefing (Cheng et al., 2014, ES = 0.10, ns). Ten randomised trials found improvement regardless of which debriefing method was used (Levett-Jones & Lapkin, 2014). More technology in the debrief does not help.

The professional standard is written as a requirement, not a recommendation: all simulation-based education must include a planned debriefing process, conducted by a person or system competent to provide appropriate feedback, maintaining psychological safety (INACSL Standards Committee, Decker et al., 2021).

So the design follows the evidence rather than the marketing. A student's conversation is recorded and held; a rubric mark, where one exists, is held for a teacher rather than released; and the teacher reads the transcript, listens to the recording and decides what is sent. The product's job is to make the debrief possible and to make it cheap enough to happen. It is not to replace it — and where this product is used with no teacher review at all, it falls outside the standard above, and we would rather say so than design around it.

05

Feedback is the active ingredient. Repetition is not.

The instinct for a product that offers unlimited practice is to claim that unlimited practice is the point. The evidence does not support that, and the honest version is more interesting.

In the largest review of instructional design features inside simulation — 289 studies, 18,971 trainees — repetitive practice has a large point estimate but does not reach significance (Cook et al., 2013, ES 0.68, 7 studies, p = 0.06). Feedback does, on eighty studies (Cook et al., 2013, ES 0.44, p < .001). In a Cochrane review of 76 studies and 10,124 students, personalised feedback against generic or none was the single best-supported design feature in the literature (Gilligan et al., 2021, SMD 0.58, moderate certainty).

Practice without feedback does not hold. Across 21 studies of motivational interviewing measured by observation rather than self-report, skills eroded over six months without feedback or coaching, and were sustained with it (Schwalbe, Oh & Zweben, 2014, d = −0.30 versus d = 0.03); three to four feedback sessions across six months was the amount required. A randomised trial of 140 clinicians reached the same conclusion from audiotaped real practice (Miller et al., 2004).

Warning

Unsupervised practice can be worse than supervised practice

Across 32 studies and 2,482 trainees, unsupervised simulation was associated with poorer immediate outcomes than instructor-supervised simulation, and negligible effects on retention beyond a week (Brydges et al., 2015, ES −0.34). Interventions that explicitly scaffolded self-regulated learning did better than those that did not. Availability at any hour is a real benefit — it is what makes spaced practice possible — but on its own it is not the mechanism, and we do not sell it as one.

On spacing itself the evidence is genuinely split, and we will not pick the flattering half. Within simulation training, distributing practice helps (Cook et al., 2013, ES 0.66, p = 0.03), and a randomised trial of surgical trainees found spaced practice beat massed practice on live-tissue transfer with total practice time held equal (Moulton et al., 2006). But in the general literature the spacing benefit shrinks as tasks get more complex, falling to nearly nothing for tasks combining high mental and physical demand — the class nearest a difficult conversation (Donovan & Radosevich, 1999, d = 0.07). We found no study anywhere establishing how many roleplay repetitions produce how much improvement. That number does not exist.

06

Practising out loud, and why we do not claim voice is better.

This product is built around speaking. The tempting claim is that speaking practice beats typing practice. It is not supported: a meta-analysis of education chatbots found text-based interactions had the largest effects (Laun & Wolff, 2025), a language-learning meta-analysis found modality did not moderate outcomes at all (Bibauw et al., 2022), and a randomised study that varied voice against text found no significant difference (Fang et al., 2025, preprint).

What is supported is modality matching: practice in the form you will be assessed in. Same-modality practice produced substantially larger gains than practice that crossed modalities, where the effect was small and non-significant (Bibauw et al., 2022, d = 0.84 same-modality speaking versus 0.29 cross-modality, ns). For an oral assessment, that means speaking. That is the claim we make, and it is narrower and better evidenced than "voice is better".

Practising with a machine does reliably reduce anxiety. In a counterbalanced crossover study, learners were measurably less anxious speaking to an AI than to an instructor — and their scores did not differ (Susoy, 2026, anxiety d = 0.39; performance d = 0.13, ns). Most participants still preferred the human, citing empathy and non-verbal communication. Lower anxiety is a real benefit, and it is a benefit in its own right; it is not evidence of better learning, and this is the most common place a product like ours would overclaim.

07

Psychological safety, and the cost of a distressing scenario.

Psychological safety is a property of the group and the facilitation, not of the software (Edmondson, 1999). It is established before a scenario starts, through prebriefing (INACSL Standards Committee, McDermott et al., 2021), and threats to it undermine reflective learning and inhibit transfer (Kolbe et al., 2020). The Australian quality framework makes participant safety an explicit statement and warns against "negative learning" pursued through low-cost shortcuts (Simulation Australasia, 2015). We can supply scaffolding. We cannot supply the conditions in the room.

Difficulty is not depth. In a randomised trial, residents given a simulated scenario ending in an unexpected death showed identical skill retention at three months to those whose patient survived, while reporting significantly higher anxiety, appraising the event as a threat rather than a challenge, and reporting more negative emotion (Khanduja et al., 2023). The distress bought nothing.

The cost also falls on whoever plays the part. One in five standardised patients reported vicarious physical symptoms during or after portraying illness, most strongly when they had personal experience of the condition (Rybak et al., 2025). That particular harm is one a synthetic character genuinely removes — which is a real advantage, and one of the few we can state without qualification.

08

What a machine character gets wrong.

These are properties of the technology, measured by other people, and they are not solved by prompting.

It drifts out of character
Across 21 models, personas showed low stability, and that stability fell further as the conversation got longer (Kovač et al., 2024). Long scenarios degrade. We cap conversation length and keep the transcript so a teacher can see where it went.
It agrees with the learner
Sycophancy is a measured, general behaviour of assistants trained on human preference, and both people and preference models sometimes prefer a convincing agreeable answer to a correct one (Sharma et al., 2023). In a difficult conversation this is a direct threat to validity: a character that softens under pressure teaches the learner that pressure works.
It can say unsafe things
Physician-led testing of four public assistants found 5–13% of answers to patient-posed medical questions were unsafe (Draelos et al., 2026). This product is a rehearsal tool, not a source of clinical guidance, and nothing it says should be treated as advice.
It stereotypes the people it invents
A language model asked to generate clinical vignettes "consistently produc[ed] … stereotype[d] demographic presentations" (Zack et al., 2024), and commercial models have been shown to reproduce race-based medical misconceptions (Omiye et al., 2023). The problem predates AI — racial diversity is underrepresented across manikins and standardised patients generally (Foronda et al., 2020) — but a generator will reproduce it at scale. Characters here are written and reviewed by people. We regard this as ongoing, not solved.

One more caution, applied to ourselves. Of 137 peer-reviewed studies of health chatbots, 99.3% did not identify the model version and fewer than a third addressed ethics, regulation or safety (Huo et al., 2025). The literature favourable to tools like this one is reported no better than the literature against.

09

What a single conversation cannot decide.

The product can put a suggested mark in front of a teacher. It cannot make a competence decision, and the reasons are well measured.

One encounter is not reliable
Across 39 studies, the reliability of objective structured clinical examination scores averaged α = 0.66 across stations — below the 0.8 normally required for a decision about an individual — and interpersonal skills were assessed less reliably than clinical ones (Brannick, Erol-Korkmaz & Prewett, 2011). How someone handles one conversation predicts the next poorly.
The assessor is a bigger variable than the candidate
In a large examination dataset, examiner stringency accounted for more score variance (16%) than candidate ability (11.4%), and correcting for it moved the pass rate from 87.2% to 97.1% (Homer, 2022). Human review inherits this problem; it does not escape it.
Validity evidence is thin
A review of 417 studies concluded that validity evidence for simulation-based assessment "is sparse" and concentrated in a few procedural domains (Cook et al., 2013).
Automated scoring is weakest exactly here
Benchmarked against expert raters, a leading model scored r = 0.90 on structured information-gathering but r = 0.47 on the patient-facing communication section (Yao et al., 2026) — the part this product exists for. Which is why a suggested mark is held, not sent.
Skills decay
Performance declined by 11–34.9% within three months of simulation training in four of four studies reviewed (Legoux et al., 2021). One-off use will not hold.

Simulation also does not substitute for supervised placement in either jurisdiction we sell into. Australian registered-nurse accreditation requires 800 hours of professional experience placement that does not count simulation (ANMAC, 2019), and the United Kingdom caps simulated practice learning at 600 of 2,300 hours (NMC, 2024). The one large randomised study supporting substitution up to 50% did so under strict conditions — trained facilitators, high-quality debriefing, best-practice design — and its authors explicitly warned that simulation is not a remedy for placement shortages (Hayden et al., 2014).

10

The numbers, with their comparators attached.

Simulation versus no teaching
1.09 process skills, 0.50 patient effects (Cook et al., 2011).
Simulation versus other teaching
0.38 process skills; patient effects 0.36, not significant (Cook et al., 2012).
Roleplay versus standardised patients
0.073, not significant — no difference (Xiao & Fu, 2025).
Virtual versus in-person, communication
−0.02 — parity (Jiang et al., 2024).
Debriefing
0.67, about a 25% improvement (Tannenbaum & Cerasoli, 2013).
Personalised feedback versus generic
0.58, moderate certainty (Gilligan et al., 2021).
Conversational agents, language proficiency
0.480.61 across four meta-analyses (Bibauw et al., 2022).
High versus low fidelity
1–2%. Realism is not the mechanism (Norman, Dore & Grierson, 2012).
Communication training, effect on patients
No evidence of benefit for patient satisfaction, patients' perception of communication, or clinician burnout (Moore et al., 2018).
Deliberate practice, professional domains
Under 1% of performance variance, not significant (Macnamara, Hambrick & Oswald, 2014). There is no ten-thousand-hour rule here.

11

Limitations. This is the honest one.

  1. We have not shown that practice here changes behaviour with real people. The peer-roleplay literature contains no studies at all measuring changed real-world behaviour or client outcomes (Gelis et al., 2020), and reviews of language-model virtual patients report limited evidence of transfer (Li & Lutfi, 2026). We have not run such a study either.
  2. Roleplay is not better than the alternatives. It is more available than them. It improves communication skills, but no more than other simulation methods (Gelis et al., 2020), and peer practice matches trained simulated patients (Lane, Hood & Rollnick, 2008).
  3. Realism is not the mechanism. High fidelity beats low fidelity by 1–2% (Norman, Dore & Grierson, 2012), and physical fidelity appears independent of educational effectiveness (Hamstra et al., 2014). A more convincing character is a better product, not a better teacher.
  4. Confidence and competence come apart, most of all here. Students overestimate their communication performance with simulated patients more than they overestimate knowledge (Blanch-Hartigan, 2011); in one simulation study self-ratings and evaluator ratings agreed at ICC = 0.17, with the weakest performers overestimating most (Høegh-Larsen et al., 2023). Self-assessment inside this product is not evidence of capability.
  5. Satisfaction is not learning. Learner reactions predict changes in motivation and self-efficacy far better than they predict skill (Sitzmann et al., 2008). Any "learners loved it" figure we ever publish should be read as exactly that.
  6. Without a debrief, the mechanism is missing. The standard requires one (INACSL Standards Committee, Decker et al., 2021), and debriefing carries d = .67 on its own (Tannenbaum & Cerasoli, 2013). Used without teacher review, this product forfeits the best-evidenced part of the method.
  7. The character drifts, and may agree with the learner. Persona stability falls with conversation length (Kovač et al., 2024), and sycophancy is a measured behaviour of these models (Sharma et al., 2023).
  8. Aged care, our core market, is the least evidenced setting. A scoping review found only 20 studies of simulation in residential aged care worldwide, none Australian or British, and called it a large paucity of evidence (Keane, Franklin & Vaughan, 2020). An Australian cluster-randomised trial across 24 nursing homes and 1,304 residents found simulation-based staff training produced no reduction in unplanned hospital transfer or death in hospital (Tropea et al., 2022, OR 1.14, 95% CI 0.82–1.59), with 26% staff participation. We publish that because it is the study a careful buyer will find, and because its lesson is about implementation.
  9. Communication training has repeatedly failed to move patient outcomes. A Cochrane review found no evidence of benefit for patient satisfaction or clinician burnout (Moore et al., 2018). A randomised trial of 472 trainees found no improvement in patient-rated communication, and patients of trained trainees scored higher on depression (Curtis et al., 2013).
  10. There is no professional standard yet for AI simulated participants. We hold ourselves to the human-participant standards in the meantime ((Lewis et al., 2017); (INACSL Standards Committee, McDermott et al., 2021)) and expect them to change.

12

Check it yourself.

Both claims in §1 are verifiable in a trial account, without taking our word for anything.

The character stays in role, or does not
Push it. Ask it to break character, to tell you the answer, to mark you. Then run a long conversation and watch for the drift described in §8. If you find it, that is the documented behaviour, not a surprise.
Nothing is marked without a person
Finish a rubric-marked activity as a student. The result appears in Responses as Awaiting review and the gradebook stays empty until a teacher opens it and commits.
The evidence is kept
Open the conversation as the teacher. The transcript, the audio and the model's working are all there — which is the material a debrief needs.

If a claim on this page does not survive that test, we would rather hear about it than have it repeated. Tell us.

13

References.

Each entry has been checked against the published record (DOI resolution and publisher metadata). In-text citations link here; use your browser's back button to return to the claim. Where a figure could not be verified from a primary source, the claim it would have supported is not made.

  • Australian Nursing and Midwifery Accreditation Council. (2019). Registered Nurse Accreditation Standards 2019. ANMAC. https://anmac.org.au/accreditation-standards/registered-nurse
  • Bibauw, S., Van den Noortgate, W., François, T., & Desmet, P. (2022). Dialogue systems for language learning: A meta-analysis. Language Learning & Technology, 26(1), 1–24. https://doi.org/10.64152/10125/73488
  • Blanch-Hartigan, D. (2011). Medical students' self-assessment of performance: Results from three meta-analyses. Patient Education and Counseling, 84(1), 3–9. https://doi.org/10.1016/j.pec.2010.06.037
  • Bosse, H. M., Schultz, J.-H., Nickel, M., Lutz, T., Möltner, A., Jünger, J., Huwendiek, S., & Nikendei, C. (2012). The effect of using standardized patients or peer role play on ratings of undergraduate communication training: A randomized controlled trial. Patient Education and Counseling, 87(3), 300–306. https://doi.org/10.1016/j.pec.2011.10.007
  • Brannick, M. T., Erol-Korkmaz, H. T., & Prewett, M. (2011). A systematic review of the reliability of objective structured clinical examination scores. Medical Education, 45(12), 1181–1189. https://doi.org/10.1111/j.1365-2923.2011.04075.x
  • Brydges, R., Manzone, J., Shanks, D., Hatala, R., Hamstra, S. J., Zendejas, B., & Cook, D. A. (2015). Self-regulated learning in simulation-based training: A systematic review and meta-analysis. Medical Education, 49(4), 368–378. https://doi.org/10.1111/medu.12649
  • Cheng, A., Eppich, W., Grant, V., Sherbino, J., Zendejas, B., & Cook, D. A. (2014). Debriefing for technology-enhanced simulation: A systematic review and meta-analysis. Medical Education, 48(7), 657–666. https://doi.org/10.1111/medu.12432
  • Cook, D. A., Hatala, R., Brydges, R., Zendejas, B., Szostek, J. H., Wang, A. T., Erwin, P. J., & Hamstra, S. J. (2011). Technology-enhanced simulation for health professions education: A systematic review and meta-analysis. JAMA, 306(9), 978–988. https://doi.org/10.1001/jama.2011.1234
  • Cook, D. A., Brydges, R., Hamstra, S. J., Zendejas, B., Szostek, J. H., Wang, A. T., Erwin, P. J., & Hatala, R. (2012). Comparative effectiveness of technology-enhanced simulation versus other instructional methods: A systematic review and meta-analysis. Simulation in Healthcare, 7(5), 308–320. https://doi.org/10.1097/SIH.0b013e3182614f95
  • Cook, D. A., Brydges, R., Zendejas, B., Hamstra, S. J., & Hatala, R. (2013). Technology-enhanced simulation to assess health professionals: A systematic review of validity evidence, research methods, and reporting quality. Academic Medicine, 88(6), 872–883. https://doi.org/10.1097/ACM.0b013e31828ffdcf
  • Cook, D. A., Hamstra, S. J., Brydges, R., Zendejas, B., Szostek, J. H., Wang, A. T., Erwin, P. J., & Hatala, R. (2013). Comparative effectiveness of instructional design features in simulation-based education: Systematic review and meta-analysis. Medical Teacher, 35(1), e867–e898. https://doi.org/10.3109/0142159X.2012.714886
  • Curtis, J. R., Back, A. L., Ford, D. W., Downey, L., Shannon, S. E., Doorenbos, A. Z., Kross, E. K., Reinke, L. F., Feemster, L. C., Edlund, B., Arnold, R. W., O'Connor, K., & Engelberg, R. A. (2013). Effect of communication skills training for residents and nurse practitioners on quality of communication with patients with serious illness: A randomized trial. JAMA, 310(21), 2271–2281. https://doi.org/10.1001/jama.2013.282081
  • Deutscher, V., Thomann, H., Zlatkin-Troitschanskaia, O., Weyland, U., Abele, S., Danek, A. H., Greiff, S., Rausch, A., Seeber, S., Seifried, J., & Winther, E. (2026). Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness. Computers and Education: Artificial Intelligence, 11, 100628. https://doi.org/10.1016/j.caeai.2026.100628
  • Donovan, J. J., & Radosevich, D. J. (1999). A meta-analytic review of the distribution of practice effect: Now you see it, now you don't. Journal of Applied Psychology, 84(5), 795–805. https://doi.org/10.1037/0021-9010.84.5.795
  • Draelos, R. L., Afreen, S., Blasko, B., Brazile, T. L., Chase, N., Desai, D. P., Evert, J., Gardner, H. L., Herrmann, L., House, A. V., Kass, S., Kavan, M., Khemani, K., Koire, A., McDonald, L. M., Rabeeah, Z., & Shah, A. (2026). Large language models provide unsafe answers to patient-posed medical questions. npj Digital Medicine, 9(1), 241. https://doi.org/10.1038/s41746-026-02428-5
  • Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383. https://doi.org/10.2307/2666999
  • Fallowfield, L., Jenkins, V., Farewell, V., Saul, J., Duffy, A., & Eves, R. (2002). Efficacy of a Cancer Research UK communication skills training model for oncologists: A randomised controlled trial. The Lancet, 359(9307), 650–656. https://doi.org/10.1016/S0140-6736(02)07810-8
  • Fang, C. M., Liu, A. R., Danry, V., Lee, E., Chan, S. W. T., Pataranutaporn, P., Maes, P., Phang, J., Lampe, M., Ahmad, L., & Agarwal, S. (2025). How AI and human behaviors shape psychosocial effects of extended chatbot use: A longitudinal randomized controlled study [Preprint]. arXiv. https://arxiv.org/abs/2503.17473
  • Foronda, C., Prather, S. L., Baptiste, D., Townsend-Chambers, C., Mays, L., & Graham, C. (2020). Underrepresentation of racial diversity in simulation: An international study. Nursing Education Perspectives, 41(3), 152–156. https://doi.org/10.1097/01.NEP.0000000000000511
  • Gelis, A., Cervello, S., Rey, R., Llorca, G., Lambert, P., Franck, N., Dupeyron, A., Delpont, M., & Rolland, B. (2020). Peer role-play for training communication skills in medical students: A systematic review. Simulation in Healthcare, 15(2), 106–111. https://doi.org/10.1097/SIH.0000000000000412
  • Gilligan, C., Powell, M., Lynagh, M. C., Ward, B. M., Lonsdale, C., Harvey, P., James, E. L., Rich, D., Dewi, S. P., Nepal, S., Croft, H. A., & Silverman, J. (2021). Interventions for improving medical students' interpersonal communication in medical consultations. Cochrane Database of Systematic Reviews, 2(2), Article CD012418. https://doi.org/10.1002/14651858.CD012418.pub2
  • Hamstra, S. J., Brydges, R., Hatala, R., Zendejas, B., & Cook, D. A. (2014). Reconsidering fidelity in simulation-based training. Academic Medicine, 89(3), 387–392. https://doi.org/10.1097/ACM.0000000000000130
  • Hayden, J. K., Smiley, R. A., Alexander, M., Kardong-Edgren, S., & Jeffries, P. R. (2014). The NCSBN National Simulation Study: A longitudinal, randomized, controlled study replacing clinical hours with simulation in prelicensure nursing education. Journal of Nursing Regulation, 5(2), S3–S40. https://doi.org/10.1016/S2155-8256(15)30062-4
  • Høegh-Larsen, A. M., Gonzalez, M. T., Reierson, I. Å., Husebø, S. I. E., Hofoss, D., & Ravik, M. (2023). Nursing students' clinical judgment skills in simulation and clinical placement: A comparison of student self-assessment and evaluator assessment. BMC Nursing, 22, 64. https://doi.org/10.1186/s12912-023-01220-0
  • Homer, M. (2022). Pass/fail decisions and standards: The impact of differential examiner stringency on OSCE outcomes. Advances in Health Sciences Education, 27(2), 457–473. https://doi.org/10.1007/s10459-022-10096-9
  • Hou, Z., & Min, S. (2026). Dialogue-based computer-assisted language learning systems for second language speaking development: A three-level meta-analysis. ReCALL, 38(1), 40–56. https://doi.org/10.1017/S0958344025100268
  • Huo, B., Boyle, A., Marfo, N., Tangamornsuksan, W., Steen, J. P., McKechnie, T., Lee, Y., Mayol, J., Antoniou, S. A., Thirunavukarasu, A. J., Sanger, S., Ramji, K., & Guyatt, G. (2025). Large language models for chatbot health advice studies: A systematic review. JAMA Network Open, 8(2), e2457879. https://doi.org/10.1001/jamanetworkopen.2024.57879
  • INACSL Standards Committee, Decker, S., Alinier, G., Crawford, S. B., Gordon, R. M., Jenkins, D., & Wilson, C. (2021). Healthcare Simulation Standards of Best Practice™: The debriefing process. Clinical Simulation in Nursing, 58, 27–32. https://doi.org/10.1016/j.ecns.2021.08.011
  • INACSL Standards Committee, McDermott, D. S., Ludlow, J., Horsley, E., & Meakim, C. (2021). Healthcare Simulation Standards of Best Practice™: Prebriefing — Preparation and briefing. Clinical Simulation in Nursing, 58, 9–13. https://doi.org/10.1016/j.ecns.2021.08.008
  • Issenberg, S. B., McGaghie, W. C., Petrusa, E. R., Lee Gordon, D., & Scalese, R. J. (2005). Features and uses of high-fidelity medical simulations that lead to effective learning: A BEME systematic review. Medical Teacher, 27(1), 10–28. https://doi.org/10.1080/01421590500046924
  • Jiang, N., Zhang, Y., Liang, S., Lu, C., Xin, T., & Wang, Y. (2024). Effectiveness of virtual simulations versus mannequins and real persons in medical and nursing education: Meta-analysis and trial sequential analysis of randomized controlled trials. Journal of Medical Internet Research, 26, e56195. https://doi.org/10.2196/56195
  • Keane, J. M., Franklin, N. F., & Vaughan, B. (2020). Simulation to educate healthcare providers working within residential age care settings: A scoping review. Nurse Education Today, 85, Article 104228. https://doi.org/10.1016/j.nedt.2019.104228
  • Khanduja, K., Bould, M. D., Andrews, M., Alam, F., Hall, J., Sharma, B., & Matava, C. (2023). Impact of unexpected death in a simulation scenario on skill retention, stress, and emotions: A simulation-based randomized controlled trial. Cureus, 15(5), e39715. https://doi.org/10.7759/cureus.39715
  • Kolbe, M., Eppich, W., Rudolph, J., Meguerdichian, M., Catena, H., Cripps, A., Grant, V., & Cheng, A. (2020). Managing psychological safety in debriefings: A dynamic balancing act. BMJ Simulation & Technology Enhanced Learning, 6(3), 164–171. https://doi.org/10.1136/bmjstel-2019-000470
  • Kononowicz, A. A., Woodham, L. A., Edelbring, S., Stathakarou, N., Davies, D., Saxena, N., Tudor Car, L., Carlstedt-Duke, J., Car, J., & Zary, N. (2019). Virtual patient simulations in health professions education: Systematic review and meta-analysis by the Digital Health Education Collaboration. Journal of Medical Internet Research, 21(7), e14676. https://doi.org/10.2196/14676
  • Kovač, G., Portelas, R., Sawayama, M., Dominey, P. F., & Oudeyer, P.-Y. (2024). Stick to your role! Stability of personal values expressed in large language models. PLOS ONE, 19(8), e0309114. https://doi.org/10.1371/journal.pone.0309114
  • Lane, C., Hood, K., & Rollnick, S. (2008). Teaching motivational interviewing: Using role play is as effective as using simulated patients. Medical Education, 42(6), 637–644. https://doi.org/10.1111/j.1365-2923.2007.02990.x
  • Laun, M., & Wolff, F. (2025). Chatbots in education: Hype or help? A meta-analysis. Learning and Individual Differences, 119, 102646. https://doi.org/10.1016/j.lindif.2025.102646
  • Legoux, C., Gerein, R., Boutis, K., Barrowman, N., & Plint, A. (2021). Retention of critical procedural skills after simulation training: A systematic review. AEM Education and Training, 5(3), e10536. https://doi.org/10.1002/aet2.10536
  • Levett-Jones, T., & Lapkin, S. (2014). A systematic review of the effectiveness of simulation debriefing in health professional education. Nurse Education Today, 34(6), e58–e63. https://doi.org/10.1016/j.nedt.2013.09.020
  • Lewis, K. L., Bohnert, C. A., Gammon, W. L., Hölzer, H., Lyman, L., Smith, C., Thompson, T. M., Wallace, A., & Gliva-McConvey, G. (2017). The Association of Standardized Patient Educators (ASPE) Standards of Best Practice (SOBP). Advances in Simulation, 2, 10. https://doi.org/10.1186/s41077-017-0043-4
  • Li, D., & Lutfi, S. L. (2026). Large language model–based virtual patient systems for history-taking in medical education: Comprehensive systematic review. JMIR Medical Informatics, 14(1), e79039. https://doi.org/10.2196/79039
  • Liaw, S. Y., Ooi, S. W., Rusli, K. D. B., Lau, T. C., Tam, W. W. S., & Chua, W. L. (2020). Nurse-physician communication team training in virtual reality versus live simulations: Randomized controlled trial on team communication and teamwork attitudes. Journal of Medical Internet Research, 22(4), e17279. https://doi.org/10.2196/17279
  • Lyu, B., Lai, C., & Guo, J. (2025). Effectiveness of chatbots in improving language learning: A meta-analysis of comparative studies. International Journal of Applied Linguistics, 35(2), 834–851. https://doi.org/10.1111/ijal.12668
  • Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608–1618. https://doi.org/10.1177/0956797614535810
  • Miller, W. R., Yahne, C. E., Moyers, T. B., Martinez, J., & Pirritano, M. (2004). A randomized trial of methods to help clinicians learn motivational interviewing. Journal of Consulting and Clinical Psychology, 72(6), 1050–1062. https://doi.org/10.1037/0022-006X.72.6.1050
  • Moore, P. M., Rivera, S., Bravo-Soto, G. A., Olivares, C., & Lawrie, T. A. (2018). Communication skills training for healthcare professionals working with people who have cancer. Cochrane Database of Systematic Reviews, 7(7), Article CD003751. https://doi.org/10.1002/14651858.CD003751.pub4
  • Moulton, C.-A. E., Dubrowski, A., MacRae, H., Graham, B., Grober, E., & Reznick, R. (2006). Teaching surgical skills: What kind of practice makes perfect? A randomized, controlled trial. Annals of Surgery, 244(3), 400–409. https://doi.org/10.1097/01.sla.0000234808.85789.6a
  • Nursing and Midwifery Council. (2024). Simulated practice learning: Supporting information for our education and training standards. NMC. https://www.nmc.org.uk/standards/guidance/supporting-information-for-our-education-and-training-standards/simulated-practice-learning/
  • Norman, G., Dore, K., & Grierson, L. (2012). The minimal relationship between simulation fidelity and transfer of learning. Medical Education, 46(7), 636–647. https://doi.org/10.1111/j.1365-2923.2012.04243.x
  • Omiye, J. A., Lester, J. C., Spichak, S., Rotemberg, V., & Daneshjou, R. (2023). Large language models propagate race-based medicine. npj Digital Medicine, 6(1), 195. https://doi.org/10.1038/s41746-023-00939-z
  • Huysmans, R., Hathaway, M., & Healthcare Simulation Standards Advisory Group. (2015). Quality framework for simulation programs in Australian health care settings (QualSim). Simulation Australasia. https://www.simnet.org.au/QualSim_Framework_2016.pdf
  • Rybak, A., Decavele, M., Taytard, J., Similowski, T., Fartoukh, M., & Guerin, C. (2025). When simulation becomes physical: Vicarious symptoms in standardized patients during OSCEs. BMC Medical Education, 25, 1509. https://doi.org/10.1186/s12909-025-07950-w
  • Schwalbe, C. S., Oh, H. Y., & Zweben, A. (2014). Sustaining motivational interviewing: A meta-analysis of training studies. Addiction, 109(8), 1287–1294. https://doi.org/10.1111/add.12558
  • Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2023). Towards understanding sycophancy in language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2310.13548
  • Sitzmann, T., Brown, K. G., Casper, W. J., Ely, K., & Zimmerman, R. D. (2008). A review and meta-analysis of the nomological network of trainee reactions. Journal of Applied Psychology, 93(2), 280–295. https://doi.org/10.1037/0021-9010.93.2.280
  • Susoy, Z. (2026). Reducing anxiety and enhancing performance: The impact of AI chatbots versus human facilitation on EFL speaking assessment outcomes. Frontiers in Psychology, 16, 1745942. https://doi.org/10.3389/fpsyg.2025.1745942
  • Tannenbaum, S. I., & Cerasoli, C. P. (2013). Do team and individual debriefs enhance performance? A meta-analysis. Human Factors, 55(1), 231–245. https://doi.org/10.1177/0018720812448394
  • Tropea, J., Nestel, D., Johnson, C., Hayes, B. J., Hutchinson, A. F., Brand, C., Le, B. H., Blackberry, I., Caplan, G. A., Bicknell, R., Hepworth, G., Lim, W. K., & Gutman, R. (2022). Evaluation of IMproving Palliative care Education and Training Using Simulation in Dementia (IMPETUS-D), a staff simulation training intervention to improve palliative care of people with advanced dementia living in nursing homes: A cluster randomised controlled trial. BMC Geriatrics, 22, 127. https://doi.org/10.1186/s12877-022-02809-x
  • Wang, F., Cheung, A. C. K., Neitzel, A. J., & Chai, C. S. (2025). Does chatting with chatbots improve language learning performance? A meta-analysis of chatbot-assisted language learning. Review of Educational Research, 95(4), 623–660. https://doi.org/10.3102/00346543241255621
  • Xiao, J., & Fu, X. (2025). Is the use of standardized patients more effective than role-playing in medical education? A meta-analysis. Frontiers in Medicine, 12, 1601116. https://doi.org/10.3389/fmed.2025.1601116
  • Yao, Z., Zhang, Z., Tang, C., Bian, X., Zhao, Y., Yang, Z., Wang, J., Zhou, H., Jang, W. S., Ouyang, F., & Yu, H. (2026). MedQA-CS: Objective structured clinical examination (OSCE)-style benchmark for evaluating LLM clinical skills. Proceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics, 6183–6257. https://doi.org/10.18653/v1/2026.eacl-long.292
  • Zack, T., Lehman, E., Suzgun, M., Rodriguez, J. A., Celi, L. A., Gichoya, J., Jurafsky, D., Szolovits, P., Bates, D. W., Abdulnour, R.-E. E., Butte, A. J., & Alsentzer, E. (2024). Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: A model evaluation study. The Lancet Digital Health, 6(1), e12–e22. https://doi.org/10.1016/S2589-7500(23)00225-X
  • Zendejas, B., Brydges, R., Wang, A. T., & Cook, D. A. (2013). Patient outcomes in simulation-based medical education: A systematic review. Journal of General Internal Medicine, 28(8), 1078–1089. https://doi.org/10.1007/s11606-012-2264-5

Read the guided tutor's evidence too.

Supporting research — guided tutor