Abstract
Background:
The rapid development of artificial intelligence (AI) technology is profoundly reshaping the medical education model. However, to date, there has been no bibliometric study specifically focusing on AI and medical education.
Methods:
We retrieved 918 records from the Web of Science™ Core Collection. Using CiteSpace and VOSviewer, we conducted a scientometric analysis of these records, including temporal and spatial distribution, author distribution, references, journals, and keywords.
Results:
The analysis provides foundational information about this research domain, revealing a remarkable growth in scholarly interest over the past decade. Current research hotspots primarily focus on the application of large language models and virtual reality technologies in medical education.
Conclusion:
This study provides essential information for interested researchers. We hope this work will offer new perspectives for advancing the development of AI in medical education.
Keywords: AI, artificial intelligence, bibliometric, large language models, medical education
1. Introduction
Medical education, as a core pillar of the human health system, is undergoing disruptive changes brought about by artificial intelligence (AI) technology.[1] Traditional medical teaching has long relied on theoretical lectures, standardized patients, and clinical practice. However, with the explosive growth of medical knowledge and the increasing complexity of clinical skill demands, traditional methods face bottlenecks in personalized learning, real-time feedback, and interdisciplinary integration. In this context, AI technology has injected new momentum into medical teaching with its powerful data processing, pattern recognition, and adaptive learning capabilities. From virtual patient simulation to intelligent diagnostic assistance systems, from adaptive learning platforms to interactive teaching tools based on natural language processing, AI is driving the transformation of medical education towards precision, personalization, and intelligence.[2]
In recent years, research on the intersection of AI and medical education has been rapidly increasing globally. However, there is still a significant knowledge gap in current research: firstly, technology driven research dominates, while in-depth analysis of the integration of educational theory and technology is relatively scarce; Secondly, the research topics are scattered and lack structured analysis of the development context, core author groups, and international cooperation networks in the field; Thirdly, there is no systematic consensus on empirical research on long-term educational outcomes.[3] These knowledge gaps highlight the urgency of comprehensively sorting out the development trajectory of this field through bibliometric methods.
Bibliometrics, a branch of informatics, focuses on the quantitative and qualitative analysis of literature systems and bibliometric features. It assesses the contributions and impacts of different authors, countries/regions, institutions, disciplines, and journals, as well as evaluates the current state, trends, and frontiers of research activities.[4–6]
This study is based on the Web of Science™ Core Collection (WoSCC) database and uses visualization tools such as CiteSpace and VOSviewer to quantitatively analyze and draw knowledge graphs of international literature on AI in medical teaching from 2015 to 2025. Systematically analyze the core research forces, national cooperation networks, and the historical evolution path and cutting-edge directions of research hotspots in this field over the past decade. We hope to provide evidence support and strategic references for educational technology developers, medical education policy makers, and interdisciplinary researchers.
2. Methods
2.1. Data collection
Literature was extracted from the Web of Science Core Collection (WoSCC). The retrieved data were collected on March 12, 2025, to avoid any potential deviation due to daily updates. The search terms were as follows: TS= (“AI” OR “AI” OR “machine learning” OR “deep learning”) AND TS= (“medical education” OR “medical training” OR “health education” OR “clinical education” OR “medical student education” OR “health professions education” OR “medical learning” OR “medical instruction” OR “medical simulation” OR “medical residency training”), and the dates of the search were March 12, 2015 to March 12, 2025. We eliminated invalid documents, including Early Access, Editorial Material, Letter, Retracted Publication, Data Paper, Meeting Abstract, Correction, Retraction, Book Chapters, News Item and Meeting. A total of 995 records retrieved were retained and the retrieved literature records were downloaded and saved as a plain text file in the format of “Full Record and Cited References,” and stored in download_.txt format., which was used as the sample of the analysis data in the paper (Fig. 1).
Figure 1.
The flowchart illustrating the search strategy and selection process.
2.2. Data analysis
All valid documents retrieved from Web of Science Core Collection underwent an initial deduplication process utilizing the CiteSpace. A total of 918 records were converted to Microsoft Excel 2019, CiteSpace or VOSviewer to perform visual analysis.
CiteSpace is a Java-based software designed for bibliometric analysis and data visualization, developed by Professor Chen Chaomei at Drexel University.[7] It presents the structure, laws, and distribution of scientific knowledge using data mining, information analysis, and atlas drawing.[8] This software is particularly useful for visualizing research hotspots, tracking the evolution of various fields, and predicting future trends, making it an effective tool for big data analysis.[9]
VOSviewer, co-developed by Nees Jan van Eck and Ludo Waltman at Leiden University, is a specialized bibliometric visualization tool that transforms complex academic relationships into interpretable knowledge maps.[10] By mapping co-occurrence relationships among authors, keywords, and institutions through interactive visualizations, it enables researchers to systematically identify emerging research frontiers, track scholarly collaborations, and evaluate disciplinary landscapes.
We used Microsoft Office Excel 2019 to analyze the articles and used CiteSpace and VOSviewer software to analyze the distribution of countries/regions visually, authors and co-cited authors, journals and co-cited journals, co-cited references, keyword cluster analysis, and timelines.
3. Results
3.1. Temporal distribution map of the literature
The number of publications over a period can reflect the research speed and trend in this field (Fig. 2A). Since 2015, annual publications have shown consistent growth: rising from 5 (2015) to 382 (2024), with 60 already recorded by March 12, 2025. Especially, the 2024 total alone exceeds twice the previous year’s volume, particularly highlighting intensified scholarly focus during this period. This upward trajectory indicates sustained research interest and accelerated development in the domain.
Figure 2.
(A) Trends of the number of published over the past 10 years; cooperation map of countries/regions (A) and institution (B) of AI in medical education. AI = artificial intelligence.
3.2. Distribution of countries/regions and institution
A total of 918 articles have been published by 300 institutions across 89 countries/regions. Table 1 lists the top 10 countries/regions and institutions by the number of publications. The most significant contribution comes from the United States (382 articles, 41.61%), followed by the People’s Republic of China (135 articles, 14.71%), England (78 articles, 8.50%), Canada (66 articles, 7.19%), and Germany (51 articles, 5.56%). Together, the top 5 countries account for more than 70% of the total publications. Figure 2B shows the Global Scientific Collaboration Network, where thicker lines indicate closer cooperation between countries. The institution with the highest number of publications is Harvard University (53 articles, 5.77%), followed by Mayo Clinic in the United States (46 articles, 5.01%), University of California System (37 articles, 4.03%), Harvard University Medical Affiliates (37 articles, 4.03%), and Harvard Medical School (32 articles, 3.49%).
Table 1.
The top 10 countries and institutions by number of publications.
| No. | Country | Centrality | Count (%) | Institution | Centrality | Count (%) |
|---|---|---|---|---|---|---|
| 1 | USA | 0.3 | 382 (41.61%) | Harvard University | 0.09 | 53 (5.77%) |
| 2 | Peoples R China | 0.07 | 135 (14.71%) | Mayo Clinic | 0.1 | 46 (5.01%) |
| 3 | England | 0.21 | 78 (8.50%) | University of California System | 0.06 | 37 (4.03%) |
| 4 | Canada | 0.04 | 66 (7.19%) | Harvard University Medical Affiliates | 0.05 | 37 (4.03%) |
| 5 | Germany | 0.15 | 51 (5.56%) | Harvard Medical School | 0.04 | 32 (3.49%) |
| 6 | Australia | 0.05 | 39 (4.25%) | Stanford University | 0.15 | 29 (3.16%) |
| 7 | India | 0.1 | 38 (4.14%) | University of Toronto | 0.04 | 20 (2.79%) |
| 8 | Italy | 0.12 | 35 (3.81%) | University System of Ohio | 0.01 | 19 (2.07%) |
| 9 | Saudi Arabia | 0.09 | 28 (3.05%) | Emory University | 0.05 | 19 (2.07%) |
| 10 | Netherlands | 0.03 | 26 (2.83%) | University of London | 0.04 | 19 (2.07%) |
In Figure 2C, each node represents an institution, with the size of the node proportional to the number of published articles. The lines between nodes indicate collaboration between them; denser lines correspond to closer cooperation. Stanford University (0.15) and Mayo Clinic (0.10) exhibit high centrality and are highlighted with purple circles in Figure 3. This finding suggests that these institutions likely play a pivotal role in the development of the application of AI in medical education.
Figure 3.
Authors (A) and references (B) co-citation network of AI in medical education. (C) Top 10 references with the strongest citation bursts related to AI in medical education. (D) Journals co-citation network of AI in medical education. AI = artificial intelligence.
3.3. Authors and co-cited authors
In the past decade, a total of 710 authors have published articles on AI in medical education. Among them, Friedman, Paul A, had the highest number of published papers (12). Followed by Noseworthy, Peter A (10), Kapa, Suraj (9), Lopez-jimenez, Francisco (7), Masters, Ken (5) (Table2).
When 2 or more authors are cited simultaneously, they are referred to as co-cited authors. Table 2 lists the top 5 most-cited authors, with Kung TH being the most cited (100 citations), followed by Gilson Aidan (67), Masters K (54), Sallam M (53), and Topol EJ (44). In CiteSpace, nodes with a mediation centrality >0.1 are considered key points and are highlighted with purple circles to indicate the importance of co-cited authors (Fig. 3A). The author with the highest centrality is Masters K (0.17), followed by Gulshan V (0.15), Chan Kaisiang (0.13), Cook DA (0.12), Chan LK (0.11).
Table 2.
The top 5 authors and co-cited authors by number of publications.
| No. | Author | Count | Co-cited author | Citation |
|---|---|---|---|---|
| 1 | Friedman, Paul A | 12 | Kung TH | 100 |
| 2 | Noseworthy, Peter A | 10 | Gilson Aidan | 67 |
| 3 | Kapa, Suraj | 9 | Masters K | 54 |
| 4 | Lopez-jimenez, Francisco | 7 | Sallam M | 53 |
| 5 | Masters, Ken | 5 | Topol EJ | 44 |
3.4. Co-cited references and references bursts
Co-citation analysis indicated that 2 references appeared in the reference list of a third citation article, and then the 2 references formed a co-citation relationship.[4] The references co-citation network of AI in medical education was show in Figure 3B. We listed the 5 most frequently cited references related to research on AI in medical education in Table 3. Among the 645 cited references, The article “Kung TH, 2023, Plos Digit Health, V2, P0, DOI 10.1371/journal.pdig.0000198” was the most cited, highlighting its significance in the field.
Table 3.
Top 5 co-cited references and journals related to AI in medical education.
| No. | Reference | Citation | Co-cited journal | Citation |
|---|---|---|---|---|
| 1 | Kung TH, 2023, Plos Digit Health, V2, P0, DOI 10.1371/journal.pdig.0000198 | 77 | Acad Med | 256 |
| 2 | Gilson Aidan, 2023, JMIR Med Educ, V9, Pe45312, DOI 10.2196/45312 | 67 | JAMA-J Am Med Assoc | 226 |
| 3 | dos Santos DP, 2019, Eur Radiol, V29, P1640, DOI 10.1007/s00330-018-5601-1 | 43 | JMIR Med Educ | 212 |
| 4 | Topol EJ, 2019, Nat Med, V25, P44, DOI 10.1038/s41591-018-0300-7 | 41 | J Med Internet Res | 208 |
| 5 | Paranjape Ketan, 2019, JMIR Med Educ, V5, Pe16048, DOI 10.2196/16048 | 36 | Med Teach | 203 |
Figure 3C presents the top 10 references with the strongest citation bursts. The article “dos Santos DP, 2019, Eur Radiol, V29, P1640, DOI 10.1007/s00330-018-5601-1”exhibits the strongest citation burst, with an intensity of 8.59. And the article “Wartman SA, 2018, Acad Med, V93, P1107, DOI 10.1097/ACM.0000000000002 044 “ follows with a citation burst intensity of 7.89.
3.5. Journals co-citation network
The impact of a journal depends on its co-citation frequency, which reflects the influence of a journal in a specific research field. Among 605 co-cited journals, 5 journals were cited over 200 times (Table 3). Acad Med (256) was the most frequently cited journal, followed by JAMA-J Am Med Assoc (226), JMIR Med Educ (212), J Med Internet Res (208) and Med Teach (203). Figure 3D displays the co-citation network of journals, highlighting the top 5 journals with the highest centrality: Am J Roentgenol (0.09), JMIR Med Educ (0.08), Artif Intell (0.08), J Med Internet Res (0.07) and BMC Med Educ (0.07). However, no journal has a centrality exceeding 0.1.
3.6. Research hotspots and frontier analysis
Keywords summarize research themes, and analyzing them helps us understand the research hotspots in a specific field. Table 4 presents the high-frequency keywords. in addition to “AI”(342), “medical education”(267) and “machine learning”(105), keywords with higher frequency in this study include “education”(52), “deep learning”(42), “performance”(38), “large language models (LLMs)”(37), “virtual reality (VR)”(34),“medical students”(31) and “health” (28)(Table 4).
Table 4.
Top 10 keywords in AI in medical education.
| No. | Keywords | Year | Centrality | Counts |
|---|---|---|---|---|
| 1 | Artificial intelligence | 2018 | 0.25 | 342 |
| 2 | Medical education | 2016 | 0.23 | 267 |
| 3 | Machine learning | 2016 | 0.17 | 105 |
| 4 | Education | 2015 | 0.16 | 52 |
| 5 | Deep learning | 2015 | 0.13 | 42 |
| 6 | Performance | 2019 | 0.11 | 38 |
| 7 | Large language models | 2023 | 0.03 | 37 |
| 8 | Virtual reality | 2020 | 0.08 | 34 |
| 9 | Medical students | 2015 | 0.09 | 31 |
| 10 | Health | 2018 | 0.04 | 28 |
CiteSpace software was used to cluster the keywords. Figure 4A displays 10 clusters, representing different research directions. The cluster 0 mainly includes: case-based learning; critical thinking; skills; natural language processing; machine learning; health professions education; gross anatomy; reflective writing. The cluster 1 mainly includes: VR; augmented reality; mixed reality; extended reality; AI; nursing education; clinical virtual simulation; virtual patient; clinical trial. The cluster 2 mainly includes: classification; impact; noise; dimensions; medical education; explainable AI; learning strategy; nucleus segmentation; artificial neural network. The cluster 3 mainly includes: qualitative research; geriatric medicine; internal medicine; mobile health; AI; postgraduate residence training; augmented reality; technology advances; skill-based training. The cluster 4 mainly includes: patient education; vignette; medical education; mathematical models; health literacy; complier average causal effects; wedge cluster; family branch system; mixed-effects model. The cluster 5 mainly includes: critical thinking; case-based learning; doctor-patient relationship; LLMs; AI; medical education; generative pretrained transformer; doctor robot. The cluster 6 mainly includes: 3D printing; medical device; undergraduate medical education; disparities; graduate medical education; medical education; information technology; finite element; recognition. The cluster 7 mainly includes: digital health; orthopedic surgery; machine learning; task analysis; large language model; generative AI; neural networks; clinical neuroscience; mobile apps. The cluster 8 mainly includes: multicriteria classification; skeletal maturation; needs assessment; surgical skills assessment; AI; medical education; surgical education; reinforcement learning; quantitative data. The cluster 9 mainly includes: clinical decision support systems; biomedical research; precision medicine; AI; abdominal CT image; health education; systematic psychometric review; students.
Figure 4.
(A) The cluster of keywords relate to AI in medical education. (B) Top 10 keywords with the strongest citation bursts. (C) Timeline viewer related to AI in medical education.
Based on the keywords co-citation network, we conducted the keywords emergent word detection. The top 10 keywords with the strongest citation bursts are reported. As revealed in Figure 4B, the blue line denotes the time axis while the red segment on the blue time axis demonstrates the burst detection, indicating the start year, end year, and burst duration. The top ranked item by bursts is “LLMs” in Cluster 5, with bursts intensity of 5.98. The second one is “machine learning” in Cluster 8, with bursts intensity of 3.6. The third is “ChatGPT” in Cluster 5, with bursts intensity of 3.5. The 4th is “medical education & training” in Cluster 3, with bursts intensity of 3.29. And the 5th is “diagnosis” in Cluster 5, with bursts intensity of 2.88. Based on the start time of emergence, “LLMs,” “ChatGPT” are the current research frontiers and they are still in an explosive period.
The timeline viewer can display the dynamic evolution path of research hotspots represented by keywords, exploring the temporal characteristics of research areas reflected in clusters and the rise and fall of hotspot keyword research. Documents within the same cluster are placed on the same horizontal line, with time progressing from left to right, from past to present. The number of documents within the same cluster highlights the richness and significance of research findings in that cluster’s field. Figure 4C visually presents the stage hotspots and development directions of AI and medical education from a temporal perspective. The keyword involved in 2015 and 2016 are mainly include “deep learning,” “machine learning,” “big data,” “medical student” and “diagnosis.” The keywords involved in 2022 to 2024 are mainly include “chatgpt,” “scoping review,” “endobronchial ultrasound” “3D visualization” and “medical diagnostic imaging”
4. Discussion
In this study, we statistically analyzed the literature related to AI and medical education, and conducted a visual analysis using CiteSpace and VOSviewer. Since 2015, The number of papers published annually in this field has demonstrated an overall exponential growth trend. This growth trend is closely associated with iterative breakthroughs in AI architectures. The deployment of GPT-3.5 in 2022 propelled the large-scale application of natural language processing technologies in medical education, while the emergence of multimodal LLMs in 2023 further expanded the construction of cross-modal teaching scenarios integrating images, texts, and voices. Notably, the substantial output from the United States (382papers) and China (135 papers), the 2 leading research entities, reflects the dual influence of the digital transformation needs of healthcare systems and national strategic investments. The clustered output from top institutions such as Harvard University (53 papers) illuminates the internal mechanisms of academic resource centralization and research network formation. In addition, we can see that the cooperation between countries and institutions is relatively close, which is conducive to removing academic barriers and further conducting research on AI and medical education.
In terms of individual authors, Professor Friedman Paul A from the Department of Cardiovascular Medicine at Mayo Clinic emerged as a foundational figure in the field, contributing 12 core publications. His research on arrhythmia prediction models and telemedicine data mining established an empirical chain for translating AI from clinical data to educational applications.[11–13] Professor MASTERS K from Sultan Qaboos University, with a high citation centrality of 0.83, has established a significant academic status in research on technological equity. His practice of promoting the inclusive application of AI in medical education in resource-constrained regions fills a theoretical gap in global health education.[14–17]
Among the landmark studies, Professor Tiffany H Kung 2023 article titled “Performance of ChatGPT on United States Medical Licensing Exam: Potential for AI-assisted medical education using LLMs,” published in the PLOS Digit Health, has the highest number of citations and exhibits the strongest citation burst.[18] This research evaluated the performance of a large language model on the United States Medical Licensing Exam, suggest that LLMs may have the potential to assist with medical education, and potentially, clinical decision-making. Professor D Pinto Dos Santos 2019 article titled “Medical students’ attitude towards AI: a multicentre survey” published in European Radiology, ranks first in citation burst intensity.[19] This article indicates that medical students have a positive attitude towards AI and are aware of the potential applications and impacts of AI in medicine.
Analysis of core journals indicated that journals in the JCR Q1 category constitute the main positions for knowledge dissemination, with JAMA-Journal of the American Medical Association (63.1, Q1) having the highest impact factor. Analyzing the distribution of literature sources helps identify the core journals for publishing articles on the application of AI in medical education, aiding scholars in establishing their scientific contributions. This data will assist future researchers in selecting journals when submitting manuscripts related to the application of AI in medical education.
Keywords summarize research themes and core content. By analyzing keyword co-occurrence, we can understand the distribution and development of various research hotspots in a specific field. Among the top 10 most occurring keywords, “ LLMs” and “VR” appears frequently. This highlights the dual technological hotpots reshaping medical education: the ascendancy of LLMs in democratizing clinical knowledge synthesis and the persistent dominance of VR in delivering immersive procedural training.[20,21]
Based on the keyword co-citation network, further analysis of keyword emergence was conducted. Keywords with citation bursts in the past 5 years represent current research hotspots. LLMs and ChatGPT have become the core frontiers of medical education research with the highest intensity values (5.98 and 3.5), marking the deep reconstruction of medical teaching paradigms by generative AI. Its explosive growth (since the end of 2022) is driven by 3 main factors: Firstly, LLMs possess remarkable cross-modal knowledge integration capabilities, enabling them to dynamically analyze vast amounts of medical literature, extract key clinical information, and generate structured cases that align with real-world diagnostic scenarios.[22,23] This not only saves educators significant time in case development but also ensures that the generated materials cover a wide range of clinical manifestations and rare conditions. Secondly, LLMs can provide immersive clinical reasoning training, allowing medical students to engage in interactive conversations with virtual patients.[24] During these conversations, they can simulate realistic patient responses, ask follow-up questions, and provide real-time feedback on diagnostic decisions. This enables learners to practice clinical reasoning in a risk-free environment, thereby enhancing their abilities to correlate symptoms, interpret test results, and develop treatment plans. Furthermore, LLMs facilitate universal access to educational resources. Their low deployment costs enable low-resource areas to access high-quality medical education materials and training tools, breaking geographical and economic barriers. By adapting to local languages and medical backgrounds, they help address the shortage of educational resources in underserved areas and promote equity in medical education.[25,26] In the future, the multimodal integration of LLMs with VR and augmented reality will give rise to context aware intelligent teaching systems, but it is necessary to simultaneously construct a new framework for ethical review and competency assessment to balance technological innovation and the humanistic core of medical education.[27,28]
AI has unveiled a revolutionary landscape for medical education: virtual simulations and personalized learning have redefined skill-training paradigms, big data integration accelerates knowledge renewal, and intelligent assessment systems coupled with remote education promote resource equity, significantly enhancing teaching efficiency and accessibility. However, a critical oversight in current research lies in the dearth of robust empirical data regarding AI’s actual pedagogical impact on medical education outcomes, which often amplifies AI’s successes while downplaying its limitations in pedagogical contexts. For example, studies celebrating LLMs’ ability to generate medical cases rarely address whether these AI-created scenarios adequately capture the messiness of real clinical practice, such as vague patient histories or conflicting symptoms, that are critical for developing diagnostic intuition. Even when empirical data exist, it is often confined to narrow metrics, such as relying solely on student satisfaction surveys to assess teaching effectiveness, despite their low correlation with actual learning gains. This cherry-picking creates an illusion of evidence, masking the reality that AI’s most hyped applications, like automated clinical reasoning tutor, lack proof of their ability to foster the adaptive thinking required in dynamic healthcare environments. More critically, this oversight impedes scrutiny of potential harms, such as the insidious erosion of critical evaluation skills that occurs when learners become accustomed to AI-generated “correct answers” without understanding their underlying logic.
The deep integration of AI in medical education must be rooted in the dual framework of educational theory and medical cognitive rules. Situated learning theory guides AI to construct a dynamic clinical simulation environment, enabling students to hone their diagnostic decision-making abilities through “legitimate peripheral participation” by utilizing a high-fidelity case library and real-time complication generation.[29,30] In the face of the explosive growth of medical knowledge, cognitive load theory becomes crucial – AI reduces the cognitive burden of anatomical learning through augmented reality visualization technology, while dynamically linking disease mechanisms with clinical protocols to enhance knowledge integration.[31] Schon reflective practice model is materialized by AI technology: natural language processing analyzes blind spots in doctor-patient communication, and motion capture systems label surgical operation deviations, generating reflective reports with evidence chains to facilitate students’ transformation from technical operators to clinical thinkers.[32,33] The implementation of technology requires the simultaneous construction of a medical-specific ethical protection network: ensuring the medical rigor of AI teaching cases through a double-blind verification mechanism, training critical questioning of AI suggestions through “intentional misdiagnosis traps,” and using affective computing technology to inversely strengthen humanistic empathy.[34] Only by deeply coupling educational theory, clinical thinking development rules, and AI technology characteristics can we break through the limitations of instrumental applications and truly achieve “reconstructing the medical cognitive paradigm with AI” – transitioning from knowledge transfer to the cultivation of clinical wisdom, ultimately forming a new era of medical education ecosystem with human-machine collaboration and shared responsibility.
Moving forward, medical education must prioritize human-centered AI, embedding ethical frameworks within technological innovation to establish a dynamic, human-machine collaborative ecosystem – where AI handles standardized, data-driven tasks, educators focus on cultivating clinical reasoning and humanistic competencies, and policymakers refine data regulations and equitable resource allocation. Only by balancing instrumental rationality with the intrinsic values of medical education can technology empower the symbiotic evolution of innovation and healthcare humanism, ultimately fostering sustainable advancements in global healthcare talent development.
5. Conclusion
Through a detailed bibliometric analysis of AI and medical education, this study evaluated literature information across different years, countries, institutions, authors, and journals, and analyzed thematic development and future research hotspots. Our research has observed a significant increase in popularity in this field since 2015. This study provides necessary information for interested researchers and identifies potential collaborators. The current research hotspots mainly focus on the application of LLMs and VR methods in medical education
Author contributions
Conceptualization: Wendan Cheng, Haoran Yu.
Data curation: Wendan Cheng.
Formal analysis: Wendan Cheng, Haoran Yu.
Funding acquisition: Wendan Cheng.
Investigation: Wendan Cheng, Zhongyao Hu.
Methodology: Wendan Cheng, Zhongyao Hu, Haoran Yu.
Project administration: Wendan Cheng, Zhongyao Hu.
Resources: Wendan Cheng.
Software: Wendan Cheng.
Visualization: Wendan Cheng, Zhongyao Hu.
Writing – original draft: Wendan Cheng.
Supervision: Zhongyao Hu, Haoran Yu.
Validation: Haoran Yu.
Writing – review & editing: Haoran Yu.
Abbreviations:
- AI
- artificial intelligence
- LLMs
- large language models
- VR
- virtual reality
- WoSCC
- Web of Science™ Core Collection
WC and HY contributed to this article equally.
This work was supported by the Anhui Science and Technology Department [grant number 2208085MH214].
In this work, ethical review is not necessary. Because it is based on publicly available literature, databases, and citation data. These data have been publicly released and do not involve sensitive personal information or privacy issues.
The authors have no conflicts of interest to disclose.
All data generated or analyzed during this study are included in this published article [and its supplementary information files].
How to cite this article: Cheng W, Hu Z, Yu H. A bibliometric analysis of artificial intelligence in medical education (2015–2025). Medicine 2025;104:46(e45684).
Contributor Information
Wendan Cheng, Email: sunyccc@126.com.
Zhongyao Hu, Email: huzhongyao000@163.com.
References
- [1].Rampton V, Mittelman M, Goldhahn J. Implications of artificial intelligence for medical education. The Lancet Digital Health. 2020;2:e111–2. [DOI] [PubMed] [Google Scholar]
- [2].Jowsey T, Stokes-Parish J, Singleton R, Todorovic M. Medical education empowered by generative artificial intelligence large language models. Trends Mol Med. 2023;29:971–3. [DOI] [PubMed] [Google Scholar]
- [3].Furfaro D, Celi LA, Schwartzstein RM. Artificial intelligence in medical education. Chest. 2024;165:771–4. [DOI] [PubMed] [Google Scholar]
- [4].Ma D, Guan B, Song L, et al. A bibliometric analysis of exosomes in cardiovascular diseases from 2001 to 2021. Front Cardiovasc Med. 2021;8:734514. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Wei N, Xu Y, Li Y, et al. A bibliometric analysis of T cell and atherosclerosis. Front Immunol. 2022;13:948314. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Tan L, Wang X, Yuan K, et al. Structural and temporal dynamics analysis on drug-eluting stents: history, research hotspots and emerging trends. Bioact Mater. 2023;23:170–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Chen C. Searching for intellectual turning points: progressive knowledge domain visualization. Proc Natl Acad Sci USA. 2004;101(suppl_1):5303–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Chen C. CiteSpace II: detecting and visualizing emerging trends and transient patterns in scientific literature. J Am Soc Inf Sci. 2006;57:359–77. [Google Scholar]
- [9].Donnelly JP. A systematic review of concept mapping dissertations. Eval Program Plan. 2017;60:186–93. [DOI] [PubMed] [Google Scholar]
- [10].Van Eck NJ, Waltman L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics. 2010;84:523–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Lüscher TF, Wenzl FA, D’Ascenzo F, Friedman PA, Antoniades C. Artificial intelligence in cardiovascular medicine: clinical applications. Eur Heart J. 2024;45:4291–304. [DOI] [PubMed] [Google Scholar]
- [12].Siontis KC, Noseworthy PA, Attia ZI, Friedman PA. Artificial intelligence-enhanced electrocardiography in cardiovascular disease management. Nat Rev Cardiol. 2021;18:465–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Attia ZI, Kapa S, Lopez-Jimenez F, et al. Screening for cardiac contractile dysfunction using an artificial intelligence–enabled electrocardiogram. Nat Med. 2019;25:70–4. [DOI] [PubMed] [Google Scholar]
- [14].Masters K. Artificial intelligence in medical education. Med Teach. 2019;41:976–80. [DOI] [PubMed] [Google Scholar]
- [15].Masters K, Benjamin J, Agrawal A, MacNeill H, Pillow MT, Mehta N. Twelve tips on creating and using custom GPTs to enhance health professions education. Med Teach. 2024;46:752–6. [DOI] [PubMed] [Google Scholar]
- [16].Mehta N, Agrawal A, Benjamin J, Mehta S, MacNeill H, Masters K. Pedagogy and generative artificial intelligence: applying the PICRAT model to Google NotebookLM. Med Teach. 2024:1–3. [DOI] [PubMed] [Google Scholar]
- [17].Masters K, Salcedo D. A checklist for reporting, reading and evaluating Artificial Intelligence Technology Enhanced Learning (AITEL) research in medical education. Med Teach. 2024;46:1175–9. [DOI] [PubMed] [Google Scholar]
- [18].Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2:e0000198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Pinto Dos Santos D, Giese D, Brodehl S, et al. Medical students’ attitude towards artificial intelligence: a multicentre survey. Eur Radiol. 2019;29:1640–6. [DOI] [PubMed] [Google Scholar]
- [20].Xu X, Chen Y, Miao J. Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review. J Educ Eval Health Prof. 2024;21:6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Fazlollahi AM, Bakhaidar M, Alsayegh A, et al. Effect of artificial intelligence tutoring vs expert instruction on learning simulated surgical skills among medical students: a randomized clinical trial. JAMA Netw Open. 2022;5:e2149008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [22].Lucas HC, Upperman JS, Robinson JR. A systematic review of large language models and their implications in medical education. Med Educ. 2024;58:1276–85. [DOI] [PubMed] [Google Scholar]
- [23].Goh E, Gallo RJ, Strong E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat Med. 2025;31:1233–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Johri S, Jeong J, Tran BA, et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat Med. 2025;31:77–86. [DOI] [PubMed] [Google Scholar]
- [25].Borg A, Georg C, Jobs B, et al. Virtual patient simulations using social robotics combined with large language models for clinical reasoning training in medical education: mixed methods study. J Med Internet Res. 2025;27:e63312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Wang C, Li S, Lin N, et al. Application of large language models in medical training evaluation – using ChatGPT as a standardized patient: multimetric assessment. J Med Internet Res. 2025;27:e59435. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Ong JCL, Chang SY-H, William W, et al. Ethical and regulatory challenges of large language models in medicine. Lancet Digital Health. 2024;6:e428–32. [DOI] [PubMed] [Google Scholar]
- [28].Rahimzadeh V, Kostick-Quenet K, Blumenthal Barby J, McGuire AL. Ethics education for healthcare professionals in the era of ChatGPT and other large language models: do we still need it? Am J Bioeth. 2023;23:17–27. [DOI] [PubMed] [Google Scholar]
- [29].O’Brien BC, Battista A. Situated learning theory in health professions education research: a scoping review. Adv Health Sci Educ Theory Pract. 2020;25:483–509. [DOI] [PubMed] [Google Scholar]
- [30].Haigh J. Expansive learning in the university setting: the case for simulated clinical experience. Nurse Educ Pract. 2007;7:95–102. [DOI] [PubMed] [Google Scholar]
- [31].Gkintoni E, Antonopoulou H, Sortwell A, Halkiopoulos C. Challenging cognitive load theory: the role of educational neuroscience and artificial intelligence in redefining learning efficacy. Brain Sci. 2025;15:203. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Plant J, Li S-TT, Blankenburg R, Bogetz AL, Long M, Butani L. Reflective practice in the clinical setting: a multi-institutional qualitative study of pediatric faculty and residents. Acad Med. 2017;92:S75–83. [DOI] [PubMed] [Google Scholar]
- [33].Hallett CE. Learning through reflection in the community: the relevance of Schon’s theories of coaching to nursing education. Int J Nurs Stud. 1997;34:103–10. [DOI] [PubMed] [Google Scholar]
- [34].Palmer J. How to harness AI’s potential in research – responsibly and ethically. Nature. 2024;632:1181–3. [DOI] [PubMed] [Google Scholar]




