Skip to main content
BMC Medical Education logoLink to BMC Medical Education
. 2026 May 26;26:1185. doi: 10.1186/s12909-026-09518-8

An exploratory study of the use of artificial intelligence-based virtual patients to enhance dentist-patient communication training

Yixuan Xie 1,2,3,4,#, Zhanpeng Ou 1,2,3,4,#, Yuanding Huang 1,2,3,4, Hong Huang 1,2,3,4, Xiongwen Ran 1,2,3,4, Hongwei Dai 1,2,3,4,, Bo Huang 5,6,, Linjing Shu 1,2,3,4,
PMCID: PMC13386813  PMID: 42192521

Abstract

Background

Effective doctor–patient communication is critical in dentistry for diagnostic accuracy and treatment efficacy. Traditional instructional formats afford limited practice opportunities, impeding the transfer of theoretical knowledge to clinical settings. Standardised patients (SPs) provide authentic interaction but are costly and logistically demanding, restricting training scalability. Large language models (LLMs), capable of generating contextually adaptive dialogues, offer innovative opportunities for dental communication training.

Methods

An AI agent was developed using the DeepSeek large language model. Thirty-eight fourth-year dental students were randomly assigned to an experimental group (theoretical instruction plus AI-based virtual patient consultations) or a control group (theoretical instruction plus peer-to-peer role-play practice). Baseline and post-intervention doctor–patient communication skills were assessed using standardised patients consultations scored with the Set Elicit Give Understand End (SEGUE) scale. Post-intervention questionnaires assessed AI agent usability and participant satisfaction.

Results

No statistically significant between-group difference was observed at baseline (p > 0.05). Following the intervention, the experimental group’s SP consultation scores were significantly higher than those of the control group (p < 0.001), with particularly pronounced gains in the preparation stage and consultation closure. Questionnaire data indicated high levels of participant satisfaction and acceptance.

Conclusions

Integration of theoretical instruction with AI agent-based training demonstrates preliminary efficacy in improving dental students’ doctor–patient communication skills and shows promise as a cost-efficient supplement to conventional training. Current limitations in dialogue flexibility and emotional intelligence should be addressed in future iterations.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12909-026-09518-8.

Keywords: Generative artificial intelligence, Large language models, Medical education, History-taking, Doctor–patient communication, Virtual patient

Background

In the clinical practice of dentistry, effective doctor–patient communication is of critical importance for diagnostic accuracy and treatment efficacy. Adequate communication enables clinicians to obtain comprehensive clinical information, fosters trust between clinician and patient, enhances adherence and satisfaction, and reduces the incidence of disputes [1, 2]. Doctor–patient communication is among the core competencies that dental students must develop and constitutes an indispensable element of medical education [3]. Evidence suggests that most communication difficulties arise from clinicians’ insufficient command of communication principles and skills rather than from deficits in clinical knowledge; even practitioners who are familiar with communication frameworks frequently fail to apply them effectively in real clinical encounters [4, 5]. The systematic development of communication skills during dental training therefore warrants greater priority.

Nevertheless, within the contemporary Chinese dental education system, the teaching of doctor–patient communication skills remains comparatively underdeveloped and fails to meet the practical demands of clinical practice. Communication training programmes in dental institutions are currently limited in scope, with disproportionate emphasis on theoretical content over practical skill development; instructional approaches are largely restricted to classroom lectures and standardised patient consultations, each carrying well-documented limitations [68]. Lectures cannot replicate genuine clinical encounters or afford repeated practice, while standardised patient consultations, though more ecologically valid, are constrained by cost, scheduling, and access [9, 10]. Identifying more effective and scalable approaches to communication training has therefore become an urgent priority in dental education.

In recent years, the rapid development of generative artificial intelligence (GAI) has created substantial new opportunities for medical education. GAI refers to a class of computational techniques that produce new content—including text, images, or molecular structures—by modelling complex data distributions through deep learning [11]. GAI demonstrates considerable potential in medical education by virtue of its strong semantic understanding and generation capabilities [12, 13]. This potential is most evident in areas such as personalised learning, simulated case generation, real-time question-answering, and virtual reality integration [1416]. Virtual patient systems built on this technology can simulate the full spectrum of clinical encounters—including history-taking, physical examination, diagnosis, and treatment planning. By offering personalised learning pathways, reusable scenarios, and a risk-free training environment, and by incorporating real-time feedback, such systems have the potential to improve learning effectiveness and strengthen educational outcomes [1719].

Several large language models have been evaluated for medical education applications, with evidence demonstrating their capacity to generate personalised learning scenarios, support clinical reasoning training, and provide real-time formative feedback. For the present study, the DeepSeek large language model was selected as the underlying engine for the AI agent. This choice was guided by pragmatic academic considerations: DeepSeek demonstrates strong natural-language generation performance comparable to leading proprietary models, supports local server deployment without usage fees, and enables data to be processed within institutional infrastructure—a feature that addresses data governance requirements relevant to educational research in healthcare settings. These characteristics render it particularly suitable for resource-constrained academic institutions seeking to develop and evaluate AI-based teaching tools [20, 21].

Although research on GAI in the medical domain is flourishing, no virtual patient dialogue system specifically designed for dental students had previously been reported. The present study therefore developed an AI-driven virtual patient consultation training platform based on GAI and integrated it with theoretical instruction, offering a novel practical teaching model for cultivating doctor–patient communication skills in dental education. As an exploratory, single-center study, this research primarily aims to estimate the effect size and evaluate the operational feasibility of subsequent confirmatory trials, thereby providing a reference for the broader application of GAI in healthcare professional education.

Methods

Research participants

This study recruited 38 fourth-year undergraduate dental students from the School of Stomatology, Chongqing Medical University, enrolled between March and July 2025. All participants had completed the relevant professional coursework in the first semester of their fourth year and had acquired sufficient theoretical grounding to engage in practical training. Written informed consent was obtained from all participants prior to enrolment; voluntary participation was required to limit the influence of extrinsic motivation on learning outcomes.

AI agent development platform and design

Three AI agents simulating patients with distinct clinical presentations were developed using the Coze agent-building platform, powered by the DeepSeek large language model. Each agent was configured with a defined role, personality profile, and response logic, endowing it with disease-specific background information, clinical status, and emotional characteristics. The agent orchestration interface is illustrated in Fig. 1. Corresponding medical history records, oral examination findings, and imaging results were uploaded to the knowledge base, enabling each agent to generate responses consistent with the patient’s clinical profile (Fig. 2). Quick-access commands for oral examination and imaging review were positioned above the dialogue interface, allowing users to retrieve relevant clinical content on demand. Voice and telephone consultation modes were activated to enable real-time interaction. At the end of each consultation, users could activate the ‘Consultation Score’ command, prompting the agent to evaluate the dialogue against predefined criteria and generate a composite score together with scoring rationale and targeted improvement suggestions (Fig. 3). This immediate, individualised feedback enabled students to identify communication gaps and make targeted improvements, thereby enhancing learning efficiency and outcomes.

Fig. 1.

Fig. 1

Screenshot of AI agent design interface

Fig. 2.

Fig. 2

Screenshot of AI agent chat interface

Fig. 3.

Fig. 3

Screenshot of the AI agent comprehensive score interface

Given that all participants were native Chinese speakers, consultation training sessions with the AI agent were conducted in Chinese; data and screenshots were subsequently translated into English.

Randomised controlled trial design

Using research randomizer, an online random number generator, 38 participants were allocated to either the experimental group or the control group (n = 19 per group). Before the intervention commenced, all participants completed an SP consultation test using an identical case to establish baseline doctor–patient communication proficiency and confirm group comparability. The experimental group received theoretical instruction combined with AI agent-simulated patient consultations, while the control group received theoretical instruction combined with peer-to-peer role-play practice. The number of training sessions, the duration of each session, the selected cases, and the content of the accompanying instructional training were kept identical across groups.

To minimise assessment bias, outcome evaluation was conducted under blinded conditions: For this study, a single professionally qualified assessor (with over 10 years’ clinical and teaching experience) was engaged to complete the scoring of all evaluation indicators and the analysis of the data. This assessor did not participate in the delivery of the training intervention and remained unaware of the group allocation throughout the evaluation process. Prior to scoring, the assessor carefully reviewed the SEGUE scale manual and scoring criteria, and discussed each item with the research team to ensure consensus on the scoring criteria. Independent rater scored each pre and post SP consultation solely on the basis of anonymised video recordings and corresponding transcripts, from which all group-allocation information had been removed. The total score awarded by the assessor was used directly as the assessment result for each student. Raters were therefore unaware of whether a given student belonged to the experimental or the control group. Scores were assigned according to the SEGUE scale. The detailed experimental workflow is presented in Fig. 4.

Fig. 4.

Fig. 4

Technical roadmap for randomized controlled trials

Teaching effectiveness evaluation tools

SEGUE scale

The SEGUE scale was developed by Professor Makoul at the Feinberg School of Medicine, Northwestern University, as a standardised framework for quantifying doctor–patient communication competencies. It is among the most widely validated instruments for assessing communication skills in medical education and clinical practice, demonstrating favourable reliability and validity [22, 23]. The five dimensions capture the full arc of a clinical consultation: Set the stage (S), Elicit information (E), Give information (G), Understand the patient's perspective (U), and End the encounter (E). Each phase is rated using a five-point Likert scale (1 = not achieved at all; 5 = fully achieved), and sub-scores are summed to yield a total score out of 100. All 38 participants completed SP consultation tests using an identical case at both baseline and post-intervention; independent assessors, blinded to group allocation, rated the recorded video dialogues and transcripts using the SEGUE scale, and pre–post score differences were compared between and within groups.

Chatbot Usability Questionnaire

The Chatbot Usability Questionnaire (CUQ), developed by an interdisciplinary team at Ulster University, is a 16-item instrument designed to assess chatbot usability. The CUQ was specifically designed for conversational interfaces and is intended to be directly comparable to the System Usability Scale (SUS), which carries a widely accepted average-usability benchmark of 68 points. Holmes et al. conducted a systematic evaluation of the CUQ’s reliability and validity based on 156 valid questionnaires (26 participants × 3 chatbots × 2 rounds of retesting). In terms of construct validity, the CUQ was able to significantly distinguish between chatbots pre-rated as having good, moderate, and poor usability levels (p < 0.05); In terms of test-retest reliability, scores obtained when the same participants completed the CUQ again approximately two weeks later showed good correlation (r > 0.7); exploratory factor analysis further revealed that the CUQ comprises four latent dimensions: personality, user experience, error handling and onboarding. The above evidence provides preliminary support for the CUQ as a reliable and valid tool for assessing chatbot usability [24]. Odd-numbered items address positive usability attributes and even-numbered items address negative aspects [25]. Responses are scored according to the official CUQ calculation guide, yielding a total score out of 100 [26]. The questionnaire was distributed online to the 19 experimental-group participants. Usability was evaluated across four domains: personality, user experience, error handling, and onboarding [27].

AI agent application experience survey questionnaire

A purpose-built questionnaire was developed to capture students' learning experiences with the AI agent. It was distributed anonymously online via Questionnaire Star to the 19 experimental-group participants. A five-point Likert scale (1–5) was used to assess satisfaction and acceptance across four domains: interest and attitude, interaction experience, learning outcomes, and overall evaluation. Internal consistency was evaluated using Cronbach's α prior to analysis (α = 0.83), confirming adequate reliability.

Sample size

Given the exploratory, single-centre nature of this study, no formal a priori power calculation was performed, consistent with established practice for pilot investigations in health professions education, where the primary aim is to estimate effect sizes and assess procedural feasibility for subsequent confirmatory trials. A total of 38 fourth-year dental students were enrolled on a convenience basis (19 per group). Whilst this sample size may limit statistical power and generalisability, it was considered adequate to generate preliminary evidence regarding the feasibility and potential efficacy of AI agent-based virtual patient training. The results should be interpreted with appropriate caution; the effect sizes reported here (r = 0.61 for total SEGUE score) provide an empirical basis for powering future confirmatory trials.

Statistical processing

Data were analysed using IBM SPSS Statistics version 26.0. Normally distributed continuous data are presented as mean ± SD; non-normally distributed data are reported as median (interquartile range, IQR). Between-group differences in post-intervention SP scores were examined with the Mann–Whitney U test; within-group pre–post differences were assessed with the Wilcoxon signed-rank test. Effect sizes were calculated as r = Z / √N, where Z is the standardised test statistic and N is the total sample size, following established conventions for non-parametric effect size reporting [28]. Statistical significance was set at p < 0.05.

Ethical review

The experimental protocol was approved by the Ethics Committee of the Affiliated Stomatological Hospital of Chongqing Medical University (Ethics Approval No. 034 of 2025). All data were stored anonymously, and anonymous codes were retained by a designated member of the research team. All participants joined voluntarily with written informed consent and agreed that their data could be used for research analysis.

Results

Participant characteristics

A total of 38 fourth-year stomatology undergraduates were enrolled, with 19 allocated to each group. Baseline communication skills did not differ significantly between groups (p > 0.05, r = 0.06), confirming group comparability (Table 1). Baseline SP scores were suboptimal in both groups: the experimental group’s median was 50 points (IQR: 48–53) and the control group’s was 51 points (IQR: 48–54), both below the passing threshold of 60 points [29, 30].

Table 1.

Comparison of baseline characteristics between groups

Variable Experimental group Control group U p r
Sex
 Male 7 9
 Female 12 10
Programme
 Dentistry 100% 100%
Degree level
 Bachelor’s degree 100% 100%
Baseline score 50(48–53) 51(48–54) 192.50 0.729 0.06

Data presented as median (IQR)

Teaching effectiveness

To address the primary research question, score improvement (post-test minus pre-test) was adopted as the principal outcome measure. The median improvement in the experimental group was + 20 points (pre-test: 50 [IQR 48–53]; post-test: 70 [IQR 66–78]), compared with + 9 points in the control group (pre-test: 51 [IQR 48–54]; post-test: 60 [IQR 59–61]). Mann–Whitney U testing confirmed that the between-group difference in post-intervention total scores was statistically significant (U = 51.50, p < 0.001, r = 0.61, large effect), indicating a clinically meaningful superiority of AI agent-assisted training over peer-to-peer practice (Table 2).

Table 2.

Comparison of total SEGUE scores between groups before and after the teaching intervention

Assessment point Experimental group Control group U p r
Pre-test total score 50(48–53) 51(48–54) 192.50 0.729 0.06
Post-test total score 70(66–78) 60(59–61) 51.50 < 0.001 0.61

Data presented as median (IQR)

Sub-domain analyses revealed differential effects across the five SEGUE dimensions. In the preparation stage, the experimental group scored significantly higher than the control group (p < 0.05, r = 0.46, medium effect; median: 11 [IQR 7–11] vs. 7 [IQR 6–8]). The largest between-group difference was observed in the consultation closure dimension (p < 0.001, r = 0.72, large effect; median: 16 [IQR 8–16] vs. 6 [IQR 6–7]), suggesting that AI-driven immediate feedback particularly reinforces structured closure behaviours. For the remaining three dimensions, between-group differences did not reach statistical significance (all p > 0.05). Effect sizes were: information elicitation (U = 225.00, r = 0.28, small-to-medium effect), information giving (U = 171.00, r = 0.05, negligible effect), and empathy/understanding (U = 161.00, r = 0.09, negligible effect). Notably, both groups achieved similarly high post-test scores in information elicitation (experimental: 26 [IQR 26–26]; control: 26 [IQR 26–27]), suggesting a ceiling effect in this dimension. The negligible effect sizes for information giving and empathy/understanding indicate that these domains may require more targeted or extended training. Detailed sub-domain results are presented in Fig. 5.

Fig. 5.

Fig. 5

Statistical results of each dimension after the teaching experiments on doctor-patient communication skills in the two groups. a Comparison of the scores in the preparation stage. b Comparison of the scores in information collection. c Comparison of the scores in information transmission. d Comparison of the scores in empathy understanding. e Comparison of the scores in conversation conclusion. f Comparison of the overall evaluation scores

Within-group pre–post comparisons using the Wilcoxon signed-rank test revealed that experimental group participants achieved significantly higher post-intervention scores across all five SEGUE sub-domains and the total score (p < 0.001 for the total and for preparation, information giving, and consultation closure; p < 0.05 for information elicitation and empathy/understanding; Fig. 6). Control group participants likewise showed significant improvements across all five sub-domains and the total score following the intervention (Fig. 7). These findings indicate that both training modalities produced meaningful improvements in doctor–patient communication competency; however, AI agent-based training was associated with a more pronounced overall effect, particularly in the preparation stage and consultation closure.

Fig. 6.

Fig. 6

Statistical results of the scores of each dimension before and after the teaching experiment of doctor-patient communication skills in the experimental group. a Comparison of the scores in the preparation stage. b Comparison of the scores in information collection. c Comparison of the scores in information transmission. d Comparison of the scores in empathy understanding. e Comparison of the scores in conversation conclusion. f Comparison of the overall evaluation scores

Fig. 7.

Fig. 7

Statistical results of the scores of each dimension before and after the teaching experiment of doctor-patient communication skills in the control group. a Comparison of the scores in the preparation stage. b Comparison of the scores in information collection. c Comparison of the scores in information transmission. d Comparison of the scores in empathy understanding. e Comparison of the scores in conversation conclusion. f Comparison of the overall evaluation scores

Chatbot Usability Questionnaire feedback

Questionnaires were completed by all 19 experimental-group participants (response rate: 100%). Responses were analysed across four CUQ domains: personality, user experience, error handling, and onboarding (Fig. 8; Table 3).

Fig. 8.

Fig. 8

Statistical results of the four categories of the chatbot usability questionnaire

Table 3.

Results from the Chatbot Usability Questionnaire (CUQ)

Index Mean ± SD
Personality
 The chatbot understood me well 3.7 ± 0.5
 The chatbot’s personality was realistic and engaging 3.5 ± 0.6
 The chatbot was welcoming during initial setup 3.7 ± 0.6
 The chatbot seemed too robotic 3.3 ± 0.8
 The chatbot seemed very unfriendly 2.1 ± 0.7
User Experience
 The chatbot was easy to navigate 3.8 ± 0.7
 The chatbot was very easy to use 3.8 ± 0.8
 It would be easy to get confused when using the chatbot 2.5 ± 0.8
 The chatbot was very complex 2.2 ± 0.6
 Chatbot responses were useful, appropriate and informative 3.7 ± 0.7
 The chatbot failed to recognise a lot of my inputs 2.4 ± 0.8
Chatbot responses were irrelevant 2.3 ± 0.8
 Error Handling
 The chatbot coped well with any errors or mistakes 3.4 ± 0.8
 The chatbot seemed unable to handle any errors 2.2 ± 0.8
Onboarding
 The chatbot explained its scope and purpose well 3.6 ± 0.8
 The chatbot gave no indication as to its purpose 2.2 ± 0.6
Total score 65.9 ± 10.2

In the personality domain, 57.89% of participants agreed that the chatbot had a realistic and engaging personality (36.84% neutral); 47.37% were neutral regarding whether it appeared too mechanical, while 36.84% agreed. Most participants (73.68%) found the chatbot welcoming during initial setup, and 73.68% reported that it understood their input well; 84.21% disagreed that it seemed unfriendly. In the user experience domain, 73.69% found the chatbot easy to navigate, and 52.63% disagreed that it was confusing to operate. A large majority (84.21%) agreed that responses were useful, appropriate, and informative, and 73.69% disagreed that responses were irrelevant. In the error handling domain, 47.37% were neutral and 42.10% agreed that the chatbot managed errors well; 78.95% rejected the notion that it was unable to handle errors. For onboarding, 63.15% agreed that the chatbot explained its scope and purpose adequately, and 84.21% disagreed that it gave no indication of its purpose.

The overall mean CUQ score was 65.9 ± 10.2 out of 100. The chatbot performed well in ease of use and initial onboarding but showed room for improvement in personalisation, information quality, and error-handling robustness.

AI agent application experience survey feedback

All 19 experimental-group participants returned completed questionnaires (response rate: 100%). The instrument demonstrated good internal consistency (Cronbach’s α = 0.83). Results are reported across four domains: interest and attitude, interactive experience, learning outcomes, and overall evaluation (Tables 4 and 5). Gender-stratified analyses revealed no statistically significant differences between male and female participants across any domain (p > 0.05).

Table 4.

Results from the AI agent implementation experience survey

Index Max Min Mean ± SD
Interest and attitude
 I am interested in AI agent-assisted instruction within stomatology practicum courses 5 3 4.11 ± 0.57
 I like knowledge acquisition through large language model (LLM) or AI agent-based learning methods. 5 3 4.11 ± 0.66
 AI agent-simulated patient training stimulates my learning motivation and engagement 5 3 4.21 ± 0.63
Interactive experience
 During interactions, AI agents provide fluent and contextually appropriate responses to my inquiries 5 2 3.84 ± 0.69
 AI agent-simulated patients realistically reflect actual clinical consultation processes. 5 1 3.58 ± 0.90
 I consciously incorporate humanistic care considerations during AI-patient interactions 5 3 3.79 ± 0.54
Learning outcome
 Training with AI-simulated patients enhances my understanding of key history-taking components. 5 2 4.05 ± 0.71
 AI-simulated patient training builds my confidence in clinical patient encounters. 5 3 3.89 ± 0.66
 My problem-solving and analytical abilities improve through AI-simulated case training. 5 3 4.05 ± 0.52
 AI-assisted teaching significantly facilitates my adaptation to clinical practice 5 3 4.11 ± 0.57
Summary
 AI-assisted teaching improves teaching efficiency 5 3 4.05 ± 0.52
 AI agent-based instruction warrants broader implementation and development 5 3 4.16 ± 0.50

Table 5.

Gender-stratified group comparison of scores from the AI agent implementation experience survey

Index Gender Number Mean ± SD t p
I am interested in AI agent-assisted instruction within stomatology practicum courses Male 7 4.00 ± 0.58 -0.61 0.552
Female 12 4.17 ± 0.58
I like knowledge acquisition through large language model (LLM) or AI agent-based learning methods Male 7 4.14 ± 0.69 0.19 0.855
Female 12 4.08 ± 0.67
AI agent-simulated patient training stimulates my learning motivation and engagement Male 7 4.00 ± 0.82 -1.12 0.279
Female 12 4.33 ± 0.49
During interactions, AI agents provide fluent and contextually appropriate responses to my inquiries Male 7 3.57 ± 0.98 -1.10 0.305
Female 12 4.00 ± 0.43
AI agent-simulated patients realistically reflect actual clinical consultation processes Male 7 3.57 ± 1.40 -0.02 0.983
Female 12 3.58 ± 0.52
I consciously incorporate humanistic care considerations during AI-patient interactions Male 7 4.00 ± 0.58 1.34 0.199
Female 12 3.67 ± 0.49
Training with AI-simulated patients enhances my understanding of key history-taking components Male 7 4.29 ± 0.49 1.11 0.283
Female 12 3.92 ± 0.79
AI-simulated patient training builds my confidence in clinical patient encounters Male 7 3.71 ± 0.95 -0.75 0.475
Female 12 4.00 ± 0.43
My problem-solving and analytical abilities improve through AI-simulated case training. Male 7 4.00 ± 0.58 -0.33 0.749
Female 12 4.08 ± 0.52
AI-assited teaching significantly facilitates my adaptation to clinical practice Male 7 4.29 ± 0.49 1.06 0.303
Female 12 4.00 ± 0.60
AI-assisted teaching improves teaching efficiency Male 7 3.86 ± 0.69 -1.26 0.224
Female 12 4.17 ± 0.39
AI agent-based instruction warrants broader implementation and development Male 7 4.14 ± 0.69 -0.10 0.924
Female 12 4.17 ± 0.39

In the interest and attitude domain, all three item means exceeded 4.1, indicating high overall engagement. Participants attributed this to the interactive and immersive qualities of the AI agent, which they felt effectively stimulated motivation. In the learning outcomes domain, mean scores exceeded 4.0 for three of the four items: mastery of history-taking techniques, analytical and problem-solving ability, and preparedness for clinical practice. The item concerning confidence in clinical consultations scored somewhat lower (mean 3.89), which, taken alongside the lower interactive experience ratings, suggests that the simulation’s authenticity may have been insufficient to fully translate practice gains into self-efficacy. In the overall evaluation domain, both items exceeded 4.0, reflecting high general satisfaction and endorsement of broader deployment.

Interactive experience scores were the lowest of the four domains (all means < 4.0), indicating limitations in dialogue fluency and contextual appropriateness. Clinical authenticity received the lowest sub-domain score, suggesting that further optimisation of response logic, conversational flow, and scenario fidelity is required. Participants’ improvement suggestions centred on three areas: increasing dialogue flexibility; expanding the case library to encompass greater complexity and diversity; and improving system stability to support application in other practical teaching contexts.

Discussion

The necessity of systematic doctor–patient communication teaching in undergraduate dental education

The World Dental Federation (FDI) has affirmed that effective clinician–patient communication is a cornerstone of high-quality dental care, and the dental education literature broadly recognises communication skill development as essential to clinical competence. Evidence further indicates that these competencies are closely linked to clinical practice behaviours, underscoring the need for systematic communication training in dental curricula. In several countries, dedicated courses complement clinical internships; for example, Japan Dental University offers a ‘Doctor–Patient Relationship’ course that develops practical communication skills through supervised patient contact [9]. Comparable structured courses remain largely absent from Chinese dental curricula, and the present data support this concern: baseline SP scores in both groups fell below the passing threshold of 60 points, indicating a substantive communication skills deficit.

Qualitative review of performance revealed that students most commonly omitted preparatory behaviours—self-introduction, identity verification, environmental reassurance, and agenda-setting—critical for establishing patient trust. In the information-giving phase, students tended to rely on technical terminology without verifying patient comprehension. At consultation closure, summarisation and checking for residual concerns were frequently absent. These deficits reflect the longstanding neglect of communication training in traditional curricula, encompassing limited systematic instruction, formative assessment, and authentic practice opportunities. Following the intervention, both groups achieved significant score improvements: the experimental group’s median post-test score (70 [IQR 66–78]) was significantly higher than the control group’s (60 [IQR 59–61]). Specifically, the experimental group’s median score during the preparation phase was 11 points (IQR 7–11) and the median score at the end of the dialogue was 16 points (IQR 8–16), which were significantly higher than the control group’s median preparation phase score of 7 points (IQR 6–8) and median end-of-dialogue score of 6 points (IQR 6–7). While both training modalities enhanced communication skills, AI agent-based training was associated with greater gains, particularly in the mastery of key consultation techniques. Based on these findings, the present study advocates for the systematic integration of doctor–patient communication training into undergraduate dental curricula.

The learning effect of AI agent virtual patients on doctor–patient communication training

Conventional doctor–patient communication teaching relies primarily on didactic lectures and standardised patient consultations. Lectures are inherently abstract and cannot replicate authentic clinical environments, while SP consultations, though more ecologically valid, are costly and logistically demanding. Both formats therefore impose constraints on the frequency and quality of practice. The present study integrated AI virtual patients with theoretical instruction, enabling students to consolidate conceptual knowledge through targeted, repeated practice and shifting the pedagogical dynamic from didactic transmission to active, dialogue-based exploration.

This reorientation is consistent with constructivist learning theory, which holds that authentic, interactive contexts facilitate deeper knowledge construction than passive instruction alone. Supporting text, voice, and telephone interaction modes, the AI platform realistically simulated clinical encounters and satisfied the contextual, interactive, and practical demands of situated learning. Furthermore, structured communication training using role-play and simulated patient interactions has been shown to improve practical communication competence beyond that achieved through supervised patient care alone [31, 32].

Post-intervention scores improved across all SEGUE dimensions in the experimental group, with the most pronounced gains in preparation, information giving, and consultation closure. Compared with the peer-to-peer control group, the experimental group’s median total score was notably higher (70 [IQR 66–78] vs. 60 [IQR 59–61]). Score differentials were particularly large in preparation (11 [IQR 7–11] vs. 7 [IQR 6–8]) and dialogue closure (16 [IQR 8–16] vs. 6 [IQR 6–7]), likely attributable to the immediate, criterion-referenced feedback provided after each AI-agent session. Survey results reinforced these findings, with participants reporting high scores in interest, learning outcomes, and overall evaluation domains.

Contextualising the mean CUQ score of 65.9 ± 10.2 against published benchmarks is essential for interpreting its significance. The CUQ is designed to be directly comparable to the System Usability Scale (SUS), which carries an established average-usability threshold of 68 points; a score approaching or exceeding this threshold represents acceptable-to-good usability. The present score falls marginally below this threshold, indicating that the AI agent meets the lower boundary of acceptable usability for a first-generation prototype. Comparable figures have been reported in the dental and medical education literature: Yuan et al. recorded a pre-optimisation CUQ of 64.2 for a stomatology AI agent built on ChatGLM, and Holderried et al. reported a CUQ of 77 for a more mature GPT-powered simulated patient after iterative refinement [27]. Viewed within this developmental trajectory, the present score is consistent with first-iteration performance and provides a meaningful baseline from which targeted improvements—particularly in dialogue flexibility, emotional realism, and error recovery—can be systematically addressed in future iterations.

The advantages of AI agent virtual patients in communication training

AI agent consultations facilitate multi-round, open-ended practice. By drawing on uploaded case information to simulate disease progression and clinical diagnostic scenarios, the agent provides a personalised, dynamic learning platform [16, 33, 34]. Upon consultation completion, the system generates real-time, criterion-based scores alongside targeted improvement suggestions, enabling students to address specific weaknesses promptly. This combination of dynamic interaction and immediate feedback addresses core limitations of traditional teaching formats [15, 35, 36].

The platform also supports practice in a safe, controlled environment, alleviating training-related psychological pressure. By simulating patient emotional responses such as anxiety and frustration, it helps learners develop communicative resilience under stress. From a resource perspective, once developed, AI virtual patients can be deployed repeatedly across multiple users without the recurring costs of live standardised patients [37, 38]. Learners can practise via mobile devices in their own time, substantially reducing training costs. Through scenario-specific agent design, the platform can address individual learning requirements with personalised learning pathways calibrated to each student’s knowledge level and learning style [39].

The deficiencies and limitations of AI agent virtual patients in communication training

A potential confound meriting acknowledgement is the novelty effect: participants in the experimental group may have experienced heightened motivation attributable to the novelty of the AI modality rather than its pedagogical content per se. However, several considerations mitigate the magnitude of this concern. First, the peer-to-peer format used in the control group also constituted a structured, interactive activity that differed from routine didactic instruction, and therefore carried a comparable degree of experiential novelty. Second, the training protocol spanned multiple sessions over several weeks; empirical evidence suggests that novelty-driven motivation attenuates rapidly with repeated exposure. Future studies employing validated motivational scales administered at multiple time points would enable a more rigorous assessment of this potential bias.

Current AI agent models show limited capacity to maintain semantic coherence in extended dialogues. Abrupt or repetitive topic shifts can produce inconsistent or off-topic responses, and emotional outputs rely on predefined rules rather than adaptive inference, resulting in formulaic expressions that fail to capture the nuanced emotional variability of real patients [40]. Interaction experience was accordingly the lowest-rated survey domain, with participants citing limited contextual flexibility and insufficient clinical authenticity. Future development should prioritise context-sensitive response generation, more naturalistic emotion modelling, and multimodal interaction [41].

The current agents are at an early developmental stage, with case libraries limited in scope and complexity. As participants complete multiple sessions with similar cases, engagement and marginal skill gains may diminish. Several participants recommended expanding the case portfolio to include paediatric and geriatric presentations and challenging scenarios such as emergency consultations and conflict management. Notwithstanding its pedagogical advantages, GAI-based training effectiveness remains contingent on curriculum design and instructional strategy; it is not a substitute for clinical placements but serves as a supplementary instrument that bridges the gap between theoretical instruction and clinical practice [39, 41, 42]. Future research should employ controlled designs with adequate statistical power—based on the effect sizes reported here (r = 0.61)—to confirm these preliminary findings.

Study limitations

This study has several limitations. As a single-centre exploratory investigation, the sample size was small (n = 38) and the primary analyses were descriptive, limiting statistical precision and generalisability. No formal a priori power calculation was performed, consistent with the pilot nature of the study. Furthermore, outcome assessment relied on a single independent assessor; the absence of a second rater precluded the calculation of inter-rater reliability, which represents a methodological limitation. Future studies should employ at least two independent assessors to quantify inter-rater agreement and strengthen the rigour of outcome evaluation. Participants were self-selected volunteers who may have been predisposed to accept AI-based tools, introducing selection bias. The study was conducted exclusively from a student-use perspective; educator perceptions of AI virtual patient feasibility were not captured [43]. Finally, the AI platform is at an early developmental stage with limited case content and an incomplete curriculum design, restricting the ecological validity of the training protocol.

Conclusions

This exploratory study provides preliminary evidence that a GAI-based virtual patient platform can meaningfully enhance dental students’ doctor–patient communication competencies and is well received by learners as a low-cost, accessible supplement to traditional training. The integration of AI agents into undergraduate dental education showed promise in compensating for the limitations of conventional teaching formats, stimulating student motivation, and supporting the transition to clinical practice. Future work should replicate these findings in larger, multi-site samples; refine the curriculum design; expand and diversify the case library; and evaluate the platform’s transferability to other areas of dental training.

Supplementary Information

Supplementary Material 1. (14.6KB, docx)
Supplementary Material 2. (242.4KB, pdf)

Acknowledgements

The authors acknowledge the financial support provided by the Chongqing Higher Education Teaching Reform Research Project (Project No. 254033) and the Chongqing Medical University School of Stomatology Education Teaching Reform Research Project (Project No. KQJ202501).

Abbreviations

CUQ

Chatbot Usability Questionnaire

FDI

World Dental Federation

GAI

Generative artificial intelligence

IQR

Interquartile range

LLMs

Large language models

SEGUE

Set Elicit Give Understand End

SPs

Standardised patients

SPSS

Statistical Package for the Social Sciences

SUS

System Usability Scale

Authors' contributions

Yixuan Xie conceived the study and drafted the manuscript. Zhanpeng Ou performed the statistical analysis and interpreted the data. Yuanding Huang, Hong Huang, and Xiongwen Ran participated in the study design and helped revise the manuscript. Hongwei Dai and Bo Huang contributed to the study design, provided expert guidance on the research methodology, and assisted in manuscript revision. Linjing Shu (primary corresponding author) provided the research concept, secured funding, supervised the overall study, and critically reviewed the manuscript. All authors read and approved the final manuscript.

Funding

This research was funded by the Chongqing Higher Education Teaching Reform Research Project (Project No. 254033) and the Chongqing Medical University School of Stomatology Education Teaching Reform Research Project (Project No. KQJ202501).

Data availability

The datasets used and analysed during the current study are available from the corresponding author on reasonable request.

Declarations

Ethics approval and consent to participate

The experimental protocol was approved by the Ethics Committee of the Affiliated Stomatological Hospital of Chongqing Medical University (Ethics Approval No. 034 of 2025) and complies with the Declaration of Helsinki. All data were stored anonymously; anonymous codes were retained by a designated member of the research team. All participants joined voluntarily with written informed consent and agreed that their data could be used for research analysis.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Yixuan Xie and Zhanpeng Ou contributed equally and should be regarded as co-first authors.

Contributor Information

Hongwei Dai, Email: dai64@hospital.cqmu.edu.cn.

Bo Huang, Email: hbxx8818@126.com.

Linjing Shu, Email: 501128@cqmu.edu.cn.

References

  • 1.Hausberg MC, Hergert A, Kröger C, Bullinger M, Rose M, Andreas S. Enhancing medical students’ communication skills: development and evaluation of an undergraduate training program. BMC Med Educ. 2012;12:16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Chen X, Liu C, Yan P, Wang H, Xu J, Yao K. The impact of doctor-patient communication on patient satisfaction in outpatient settings: implications for medical training and practice. BMC Med Educ. 2025;25:830. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zhou S, Qu Y, Song Y-J, Wang J-C, Zhao L-B, Hu. N-R-S. Education on the importance of doctor-patient communication in orthodontic clinical teaching. J Craniofac Surg. 2025;36:e248–51. [DOI] [PubMed] [Google Scholar]
  • 4.Shiraly R, Mahdaviazad H, Pakdin A. Doctor-patient communication skills: a survey on knowledge and practice of iranian family physicians. Bmc Fam Pract. 2021;22:130. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Moezzi M, Rasekh S, Zare E, Karimi M. Evaluating clinical communication skills of medical students, assistants, and professors. BMC Med Educ. 2024;24:19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Jiang Y, Shi L, Cao J, Zhu L, Sha Y, Li T, et al. Effectiveness of clinical scenario dramas to teach doctor-patient relationship and communication skills. BMC Med Educ. 2020;20:473. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Adnan AI. Effectiveness of communication skills training in medical students using simulated patients or volunteer outpatients. Cureus. 2022;14:e26717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Moore R. Maximizing student clinical communication skills in dental education-a narrative review. Dent J. 2022;10:57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.He Y-Z, Liao P-C, Chang Y-T. Enhancing patient-centred care in taiwan’s dental education system: Exploring the feasibility of doctor-patient communication education and training. J Dent Sci. 2023;18:1831–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Holderried F, Stegemann-Philipps C, Herrmann-Werner A, Festl-Wietek T, Holderried M, Eickhoff C, et al. A language model–powered simulated patient with automated feedback for history taking: prospective study. JMIR Med Educ. 2024;10:e59213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Fahrner LJ, Chen E, Topol E, Rajpurkar P. The generative era of medical AI. Cell. 2025;188:3648–60. [DOI] [PubMed] [Google Scholar]
  • 12.Wang S, Mo C, Chen Y, Dai X, Wang H, Shen X. Exploring the performance of ChatGPT-4 in the taiwan audiologist qualification examination: preliminary observational study highlighting the potential of AI chatbots in hearing care. JMIR med educ. 2024;10:e55595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Aster A, Laupichler MC, Rockwell-Kollmann T, Masala G, Bala E, Raupach T. ChatGPT and other large language models in medical education - scoping literature review. Med Sci Educ. 2025;35:555–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Xu X, Chen Y, Miao J. Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review. J Educ Eval Health Prof. 2024;21:6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Meng X, Yan X, Zhang K, Liu D, Cui X, Yang Y, et al. The application of large language models in medicine: A scoping review. Iscience. 2024;27:109713. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Lu B, Wei Z, Li X, Yin Y, Linghu J, Wang Y, et al. Progress of a novel dentistry teaching model based on the combination of virtual reality and artificial intelligence technologies in optimizing cognitive load: a systematic review. J Dent Educ. 2026;90:725–42. [DOI] [PubMed]
  • 17.Chung K, Cho HY, Park JY. A chatbot for perinatal women’s and partners’ obstetric and mental health care: development and usability evaluation study. JMIR Med Inf. 2021;9:e18607. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Chang C-C, Tseng P-L, Liu C-C, Ming J-L, Fan S-H, Tung C-Y. Development and validation of the physician’s health literacy competence scale: a step towards effective doctor-patient communication. Med (Baltim). 2025;104:e41643. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lucas HC, Upperman JS, Robinson JR. A systematic review of large language models and their implications in medical education. Med Educ. 2024;58:1276–85. [DOI] [PubMed] [Google Scholar]
  • 20.Peng Y, Malin BA, Rousseau JF, Wang Y, Xu Z, Xu X, et al. From GPT to DeepSeek: significant gaps remain in realizing AI in healthcare. J Biomed Inf. 2025;163:104791. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Kayaalp ME, Prill R, Sezgin EA, Cong T, Królikowska A, Hirschmann MT. DeepSeek versus ChatGPT: multimodal artificial intelligence revolutionizing scientific discovery. From language editing to autonomous content generation-redefining innovation in research and practice. Knee Surg Sports Traumatol Arthrosc: Off J ESSKA. 2025;33:1553–6. [DOI] [PubMed] [Google Scholar]
  • 22.Zeng N, Lu H, Li S, Yang Q, Liu F, Pan H, et al. Application of the combination of CBL teaching method and SEGUE framework to improve the doctor-patient communication skills of resident physicians in otolaryngology department. BMC Med Educ. 2024;24:201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Dorrestein L, Ritter C, De Mol Z, Wichtel M, Cary J, Vengrin C, et al. Validity evidence for communication skills assessment in health professions education: A scoping review. BMJ Open. 2025;15:e096799. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Holmes S, Bond R, Moorhead A, Zheng J, Coates V, McTear M. Towards Validating a Chatbot Usability Scale. In: Design, User Experience, and Usability. HCII 2023. Lecture Notes in Computer Science. Springer; 2023;14033:321–39.
  • 25.Holmes S, Moorhead A, Bond R, Zheng H, Coates V, McTear M. Usability testing of a healthcare chatbot: can we use conventional methods to assess conversational user interfaces? In: Proceedings of the 31st European Conference on Cognitive Ergonomics; 2019 Sep 10-13; Belfast, United Kingdom. New York: ACM; 2019. p. 207–14.
  • 26.Holmes S. Chatbot Usability Questionnaire Usage Guide.
  • 27.Holderried F, Stegemann-Philipps C, Herschbach L, Moldt J-A, Nevins A, Griewatz J, et al. A Generative Pretrained Transformer (GPT)-Powered Chatbot as a Simulated Patient to Practice History Taking: Prospective, Mixed Methods Study. JMIR Med Educ. 2024;10:e53961. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Cohen J. A power primer. Psychol Bull. 1992;112:155–9. [DOI] [PubMed] [Google Scholar]
  • 29.Tsikas SA, Afshar K. Clinical experience can compensate for inferior academic achievements in an undergraduate objective structured clinical examination. BMC Med Educ. 2023;23:167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Liu M, Liu K-M. Setting pass scores for clinical skills assessment. Kaohsiung J Med Sci. 2008;24:656–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Huang S, Wen C, Bai X, Li S, Wang S, Wang X, et al. Exploring the application capability of ChatGPT as an instructor in skills education for dental medical students: randomized controlled trial. J Med Internet Res. 2025;27:e68538. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Haak R, Rosenbohm J, Koerfer A, Obliers R, Wicht MJ. The effect of undergraduate education in communication skills: a randomised controlled clinical trial. Eur J Dent Educ: Off J Assoc Dent Educ Eur. 2008;12:213–8. [DOI] [PubMed] [Google Scholar]
  • 33.Wang X, Guo Y, Tian J, Jiang X, Yu S, Xia B. Study on the training effect of SIMROID robot in improving medical students’ patient-clinician communication and behavior management skills. J Dent Educ. 2025. [DOI] [PMC free article] [PubMed]
  • 34.Riedel M, Kaefinger K, Stuehrenberg A, Ritter V, Amann N, Graf A, et al. ChatGPT’s performance in german OB/GYN exams - paving the way for AI-enhanced medical education and clinical practice. Front Med. 2023;10:1296615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Hirosawa T, Yokose M, Sakamoto T, Harada Y, Tokumasu K, Mizuta K, et al. Utility of generative artificial intelligence for japanese medical interview training: Randomized crossover pilot study. JMIR Med Educ. 2025;11:e77332–77332. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Benítez TM, Xu Y, Boudreau JD, Kow AWC, Bello F, Van Phuoc L, et al. Harnessing the potential of large language models in medical education: promise and pitfalls. J Am Med Inf Assoc: JAMIA. 2024;31:776–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Mascarenhas S, Al-Halabi M, Otaki F, Nasaif M, Davis D. Simulation-based education for selected communication skills: exploring the perception of post-graduate dental students. Korean J Med Educ. 2021;33:11–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Rädel-Ablass K, Schliz K, Schlick C, Meindl B, Pahr-Hosbach S, Schwendemann H, et al. Teaching opportunities for anamnesis interviews through AI based teaching role plays: a survey with online learning students from health study programs. BMC Med Educ. 2025;25:259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Eysenbach G. The role of ChatGPT, generative language models, and artificial intelligence in medical education: a conversation with ChatGPT and a call for papers. JMIR Med Educ. 2023;9:e46885. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Sallam M, Salim NA, Barakat M, Al-Tammemi AB. ChatGPT applications in medical, dental, pharmacy, and public health education: a descriptive study highlighting the advantages and limitations. Narra J. 2023;3:e103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Mu Y, He D. The potential applications and challenges of ChatGPT in the medical field. Int J Gen Med. 2024;17:817–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Bianchi J, Zheng M. Leveraging generative artificial intelligence in teaching, scholarship and dental education: use cases and reflections. Orthod Craniofac Res. 2025;28(Suppl 1):S11–7. [DOI] [PubMed] [Google Scholar]
  • 43.Al-Zubaidi SM, Muhammad Shaikh G, Malik A, Zain Ul Abideen M, Tareen J, Alzahrani NSA, et al. Exploring faculty preparedness for artificial intelligence-driven dental education: A multicentre study. Cureus. 2024;16:e64377. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (14.6KB, docx)
Supplementary Material 2. (242.4KB, pdf)

Data Availability Statement

The datasets used and analysed during the current study are available from the corresponding author on reasonable request.


Articles from BMC Medical Education are provided here courtesy of BMC

RESOURCES