ABSTRACT
Artificial intelligence (AI) is entering orthopaedic education through image interpretation, simulation, technical skill assessment, examination support and content generation. Most evidence, however, comes from undergraduate and postgraduate settings, whereas practising surgeons require continuing medical education and continuing professional development (CME/CPD) that supports safe adoption, maintenance of competence and practice improvement. This matters because AI may improve task performance while it is present without producing durable learning or safer independent practice. This critical narrative review asks: what learner benefit does the current orthopaedic education literature demonstrate, and how should that evidence be translated into AI-enabled CME/CPD for practising orthopaedic clinicians? Orthopaedic-specific reviews and primary studies were synthesised and supplemented by a targeted PubMed and Google Scholar update through 4 August 2026. Across modalities, a consistent pattern emerges: AI is strongest in tightly bounded tasks with objective outputs, whereas evidence for retention, transfer, workplace performance and patient benefit is sparse. Diagnostic support can improve accuracy, speed or confidence while available; machine-learning models can distinguish expertise levels in selected simulated procedures; and large language models perform strongly on some text-based examinations but remain vulnerable to explanation errors, hallucinated reasoning, model and version dependence, and image-rich tasks. One recent historical cohort study suggests that structured AI-assisted peer teaching may improve knowledge, clinical reasoning and three-month retention. We therefore propose an AI-specific outcomes framework that separates model capability and assisted performance from learning, transfer, workplace performance and patient or system outcomes, alongside four design principles for CPD. For CME/CPD, the priority is not simply to use AI more often, but to ensure that AI-enabled education produces demonstrable learning, safe transfer and improved practice.
KEYWORDS: Artificial intelligence, orthopaedics, continuing medical education, continuing professional development, surgical education, simulation
Introduction
Orthopaedic practice depends on visual pattern recognition, biomechanical reasoning, procedural planning and psychomotor skill. These features make the speciality particularly suitable for AI systems that can analyse images, quantify movement, classify performance and generate structured feedback. The same technologies are now being considered not only for medical students and residents, but also for specialists learning new procedures, responding to changing evidence, maintaining competence and integrating AI-enabled clinical tools into practice.
Existing orthopaedic reviews have largely mapped applications in surgical training. They describe fracture-interpretation support, virtual-reality and sensor-based simulation, automated expertise classification, large language model (LLM) examination performance, content generation and emerging personalised learning [1–8]. Across these domains, however, the same pattern recurs: AI performs most convincingly when the task is tightly bounded, the desired output is objective and performance is measured immediately. The technological signal is therefore stronger than the educational signal. A model may answer a board-style question, highlight a fracture or distinguish novice from expert performance without causing a clinician to learn, retain that learning, transfer it to a new setting or improve patient care.
This distinction is especially important for CME/CPD. The purpose of CPD is not simply exposure to information or technology; it is the maintenance and development of professional competence and, ultimately, better clinical performance and health outcomes. AI also creates a new educational need. Orthopaedic clinicians must learn how to appraise model boundaries, recognise automation bias, protect data, verify generated content and preserve accountability when human and machine recommendations interact. Without this training, rapid clinical adoption could outpace the profession’s ability to use AI critically and safely. These are career-long competencies rather than topics confined to undergraduate curricula.
This review therefore asks two linked questions: what does the current orthopaedic education literature demonstrate about learner benefit, and how should that evidence be translated into AI-enabled CME/CPD for practising orthopaedic clinicians? Its contribution is an outcomes-based framework that separates technical model capability and tool-assisted performance from learning, retention, workplace performance and patient or system benefit.
Review Approach
This article is a critical narrative review rather than a systematic or scoping review. The evidence base was anchored to the orthopaedic-specific systematic scoping review by Al-Saadawi et al., which searched four databases through August 2024, and to subsequent orthopaedic reviews published in 2025 and 2026 [1–6]. A targeted update was undertaken using PubMed and Google Scholar, backward citation-chasing and primary-source checking for peer-reviewed literature published through 4 August 2026.
Search concepts combined orthopaedic or orthopaedic surgery with artificial intelligence, machine learning, deep learning, computer vision, large language models, vision-language models or ChatGPT, and with education, training, simulation, resident, fellow, surgeon, continuing medical education, continuing professional development, assessment, examination, peer teaching or curriculum. Priority was given to primary studies with an identifiable orthopaedic educational context and to evidence concerning AI education for healthcare professionals. Purely clinical AI studies without an educational, training or professional-development component were not included in the main synthesis.
Studies were interpreted according to the educational outcome assessed, the objectivity of the target task, the degree of external or workplace validation and whether the study demonstrated learner benefit rather than model performance alone. Because the search was targeted and not independently duplicated across all subscription databases, no PRISMA flow diagram or formal claim of comprehensive coverage is made. Where implications for practising surgeons are extrapolated from student or resident evidence, this is stated explicitly.
The Current Evidence Landscape
The orthopaedic-specific evidence base is recent and heterogeneous. The 2025 scoping review identified 21 primary studies involving 273 participants, with most papers appearing during the preceding five years [1]. Subsequent work has expanded rapidly in LLM benchmarking, multimodal examinations, extended-reality training, AI-assisted peer teaching and surveys of real-world use [2–6,9]. Growth in longitudinal educational trials and transfer to clinical performance has been much slower.
Synthesised across modalities, the evidence is not simply uneven in quantity; it is concentrated at the lower end of the educational-outcome hierarchy. Fracture classification, simulator telemetry and multiple-choice examinations provide immediate, objective endpoints and therefore dominate the literature. Educationally complex outcomes such as judgement, communication, operative independence, retention, transfer and practice change are harder to measure and remain under-represented. Table 1 maps the principal applications, the learner populations studied and the unresolved evidence gap. This distribution is central to the paper’s argument: evidence that an AI system can perform a task, or can improve performance while present, should not be treated as equivalent to evidence that a clinician has learned.
Table 1.
Orthopaedic AI applications across the career continuum.
| Application | Evidence in current literature | Predominant learner population | Potential CME/CPD use | Critical evidence gap |
|---|---|---|---|---|
| Diagnostic support | Improved accuracy, speed or confidence in selected learner studies | Predominantly medical students and trainees; includes interns, residents and fellows | Deliberate practice for image interpretation; safe adoption of decision support | Unaided retention, calibration, boundary recognition and behaviour when AI is wrong |
| Simulation and skill assessment | High discrimination between expertise levels in bounded simulator tasks | Predominantly residents/trainees; small novice-expert or post-residency comparisons | Learning new procedures; refreshing infrequent skills; targeted remediation | Transfer to novel tasks, theatre performance and clinical outcomes |
| LLM knowledge support | Rapid improvement on some text-based examinations; performance remains modality- and model-dependent | Mostly model benchmarking against resident/fellow examination standards rather than learner interventions | Evidence orientation, reflective questioning and individual learning plans | Accuracy of explanations, source verification, multimodal performance and impact on practice |
| Content generation | Faster production of summaries and draft materials, with substantial error burden in generated questions | Fellow-generated comparators and expert reviewers; limited direct learner-outcome evidence | Drafting CPD content and formative questions | Independent review burden, errors, conflicts and version drift |
| AI-assisted peer learning | Encouraging gains in knowledge, reasoning and three-month retention in one historical cohort | Medical students | Case-based discussion and communities of practice | Randomised evidence and relevance to practising clinicians |
| Practice analytics | Conceptually supported by automated performance measurement | No established practising-surgeon CPD population | Identification of professional gaps and targeted feedback | Governance, fairness, acceptability and separation from punitive surveillance |
Current educational evidence, potential relevance to CME/CPD and the principal remaining evidence gap.
Diagnostic Teaching and Image Interpretation
Diagnostic teaching is one of the clearest near-term applications. Cheng et al. evaluated AI-generated heatmaps for hip-fracture interpretation in fifth-year medical students and reported greater improvement in post-learning performance than with standard radiographic teaching alone [10]. Kim et al. reported that a transfer-learning ensemble for foot-fracture diagnosis improved accuracy across fellows, residents, interns and students while reducing decision time [11]. Small resident studies summarised in the orthopaedic reviews have similarly reported improvements in detection, confidence and interpretation speed.
The evidence also illustrates an important boundary problem. The hip-fracture heatmap and foot-fracture ensemble were developed and evaluated for different anatomical and diagnostic tasks and should not be assumed to be interchangeable [10,11]. More broadly, an algorithm validated on adult plain radiographs cannot be assumed to perform similarly on paediatric images, postoperative films containing metalwork, unusual projections or a different modality such as soft-tissue MRI. Larger neural networks are increasingly deployed for clinical imaging, but unless they are evaluated as teaching interventions they demonstrate diagnostic assistance rather than educational effectiveness. A practical educational risk is that users learn to follow the AI report rather than interrogate the image and the conditions under which the model was validated.
These findings remain relevant to CPD because radiographic interpretation continues throughout professional practice and decision-support systems may be introduced after formal training has ended. However, the educational mechanism must be distinguished from the immediate performance effect of having the tool available. A clinician may become faster or more accurate while an AI overlay is present without acquiring transferable knowledge. Conversely, repeated comparison between personal judgement, AI output and expert adjudication could be designed as deliberate practice.
A CPD programme built around diagnostic AI should therefore measure unaided post-activity performance, calibration of confidence, response to deliberately incorrect AI advice and retention after support is removed. It should include cases outside the model’s familiar distribution and require participants to explain why they accept or reject an AI suggestion. In practical terms, this means testing whether the learner can recognise when a model designed for one anatomy, modality or case distribution is being applied beyond that boundary. Immediate assisted accuracy alone is an insufficient CME/CPD outcome.
Simulation and Automated Technical-Skill Assessment
Orthopaedic simulation has produced some of the highest reported classification accuracies in the field. Bissonnette et al. showed that machine-learning algorithms could distinguish levels of training during a virtual-reality spinal task, with the best-performing support vector machine reaching 97.6% accuracy [12]. Alkadri et al. used a multilayer perceptron to assess a simulated annulus-incision procedure and reported 80% testing accuracy while identifying the simulator metrics most influential in classification [13]. Demirel et al. developed quantitative metrics for arthroscopic rotator cuff repair that differentiated novice and expert performance [14].
These studies demonstrate that AI can convert aspects of tacit supervision into reproducible measurements. Force, velocity, tool trajectory, task time, tissue interaction and procedural completeness can be analysed more consistently than by unaided observation alone. Explainable metrics may also support targeted coaching by showing precisely which behaviours differ from an expert reference. Reporting guidance for machine-learning assessment of surgical expertise highlights the importance of transparent data partitioning, validation and interpretation [15].
Taken together, however, the studies demonstrate robust discrimination of expertise in a small number of highly specified tasks rather than a mature CPD infrastructure. Orthopaedics spans many anatomical regions, procedures, implants and operative platforms, while current simulations are comparatively few and often procedure- or platform-specific. Aviation provides a useful comparator: simulation is used for recurrent competence, training on new equipment and rehearsal of uncommon emergencies. Similar use cases are attractive for orthopaedic CPD, including competence maintenance, adoption of new techniques and low-frequency high-risk events. The challenge is greater in surgery because one clinician may need to remain competent across multiple procedures and technologies rather than a single aircraft type. The central question is therefore not whether AI can score one simulator task, but whether AI-guided practice transfers across tasks, anatomy and real clinical environments.
For practising surgeons, the potential CPD applications include learning a new implant or operative approach, refreshing an infrequently performed procedure, responding to an identified performance gap and documenting progression before supervised clinical adoption. Yet the principal limitation remains the distance between simulator classification and clinical competence. A model may identify whether a participant resembles an expert within one simulator without establishing that the participant can adapt to anatomical variation, manage complications or perform safely in theatre. AI-enabled simulation should therefore be embedded in a broader pathway that includes expert review, repeated practice, assessment on an unfamiliar task and supervised workplace transfer.
Examination Support, Knowledge Updating and Large Language Models
The LLM literature has developed faster than any other area of orthopaedic AI education. Early studies were appropriately cautious. Lum reported that ChatGPT was unlikely to pass an American orthopaedic board-style examination, Kung et al. found that ChatGPT-3.5 answered 46 of 102 Orthopaedic In-Training Examination questions correctly and Massey et al. found that residents outperformed both ChatGPT-3.5 and ChatGPT-4 [16–18].
Performance has improved markedly with newer models and different datasets. Khan et al. reported that GPT-4.0 achieved 73.9% on a UK FRCS Trauma and Orthopaedics dataset [19]. Diniz et al. found that reasoning-optimised models achieved substantially higher accuracy than GPT-4 on a large text-only orthopaedic question set [20]. More recent models have also performed strongly on text-based OITE material [21]. These results show that conclusions about LLM capability are model-, version-, dataset- and modality-specific rather than stable properties of “AI”.
Raw examination score remains an incomplete educational outcome. Hayes et al. showed that performance deteriorated when image information was inadequate or generated by AI rather than described by an orthopaedic expert [22]. Jain et al. reported resident-comparable scores on selected examinations, but only 58.7% of responses were rated ideal and more than one third were considered unacceptable [23]. Ko et al. found that larger open-source vision-language models could approach junior-resident performance on a multimodal examination, while smaller models performed below first-year residents and image-rich domains remained difficult [24]. Together, these studies show that benchmark gains do not remove dependence on input quality, modality or model architecture.
For CME/CPD, LLMs may support rapid evidence orientation, case discussion, reflective questioning and the drafting of individual learning plans. They should not be treated as independent repositories of current guidance.
Training and CPD activities should therefore require participants to compare LLM responses with primary sources, identify unsupported claims, assess uncertainty and revise outputs for clinical use. They should also ask whether the prompt itself frames the problem in a biased way and whether the evidence available to the model reflects publication bias or incomplete representation of negative findings. This converts AI from an answer generator into an object of critical appraisal and makes verification part of the learning outcome.
Content Generation, Peer Teaching and Personalised Learning
Generative AI may reduce the time required to draft educational summaries, explanations and assessment material. DeCook et al. reported that ChatGPT-4 produced total knee arthroplasty educational summaries that were rated favourably compared with fellow-generated material and were created much faster [25]. Question generation is more problematic: Dagher et al. found that 70% of ChatGPT-4-generated orthopaedic examination questions contained at least one error, and only half were considered usable after minor or moderate modification [26]. CPD providers can therefore use AI for first drafts, but content review, source verification, conflict-of-interest controls and clear responsibility for the final activity remain essential.
The strongest recent learner-outcome evidence comes from an historical cohort study of AI-assisted peer teaching in orthopaedic clinical education. Yu et al. compared two consecutive cohorts of medical students and reported higher knowledge and OSCE scores, particularly for clinical reasoning, with the difference in knowledge maintained at three months [27]. The study is important because it measured learning and retention rather than model performance alone. Its non-randomised design, temporal cohorts and combined intervention mean that the independent effect of AI cannot be isolated, but the findings support evaluation of AI as part of a structured social-learning intervention rather than as a standalone chatbot.
True adaptive orthopaedic learning remains more aspiration than established evidence. Most current studies evaluate isolated components: a chatbot, simulator score, diagnostic overlay or generated summary. This is another example of AI boundaries limiting claims of true personalisation: successful adaptation within one task or platform does not establish that the system has identified the clinician’s broader professional needs. Personalisation should therefore not be driven primarily by opaque platform analytics or the volume of content a user consumes. For practising clinicians, CME personalisation should begin with a documented professional appraisal or practice gap and should adapt content, cases and feedback to the identified gap.
Why This Evidence Matters for CME/CPD
Most orthopaedic AI education studies involve students or residents, but the underlying learning needs persist after certification. Specialists must incorporate new evidence, devices and operative techniques; maintain skills that may be used infrequently; understand AI systems introduced into clinical workflows; and recognise when technology changes the distribution of responsibility within a team. The limited direct evidence in practising orthopaedic surgeons should therefore be treated as a research gap, not as evidence that CPD is unaffected.
A 2026 multicentre survey of 534 orthopaedic residents found that 71.9% had received no formal AI training, despite substantial use for literature searching, summarisation and drafting [9]. This pattern is likely to create a moving cohort of clinicians who use AI before receiving structured education in verification, governance or safe application. Broader reviews of AI education for healthcare professionals similarly identify fragmented curricula and limited evidence that training changes workplace behaviour [28,29]. A professional CPD position statement has consequently emphasised ethical and responsible AI use as a priority for the field [30]. Recent CPD scholarship further argues that clinicians must be prepared not only to use AI tools but also to interpret, supervise and steward AI-embedded practice environments [31].
The practical implication is that AI-related CPD should have two complementary strands. The first is education with AI: using diagnostic overlays, simulation analytics or LLMs to support learning. The second is education about AI: developing the competence to appraise, supervise and safely use AI in professional practice. Programmes that offer the first without the second risk strengthening automation dependence rather than professional capability.
Structured training in the capabilities, boundaries, verification requirements and governance of each AI modality should therefore precede or accompany reliance on that modality in orthopaedic practice.
A Proposed Outcomes Framework for AI-Enabled Orthopaedic CME/CPD
CME/CPD evaluation frameworks distinguish participation and satisfaction from learning, competence, performance and health outcomes [32]. AI introduces an additional layer because a technology can perform well even when the learner does not improve. Table 2 therefore places technical model validation and assisted performance below progressively stronger educational and professional outcomes, and Figure 1 summarises the distinction.
Table 2.
Outcomes framework for AI-enabled orthopaedic CME/CPD.
| Level | Question | Illustrative measure | Approximate Moore correspondence | Current orthopaedic evidence |
|---|---|---|---|---|
| 0. Model capability | Can the AI perform or score the task? | Accuracy, AUROC, examination score, expertise classification | No direct equivalent: technology outcome, not learner outcome | Common |
| 1. Assisted performance | Does the clinician perform better while AI is present? | Accuracy, speed, confidence, simulator metric | No direct equivalent; deliberately separated from Moore Level 5 performance | Moderate |
| 2. Learning and competence | Does unaided knowledge, reasoning or skill improve? | Post-test, OSCE or unaided skills assessment | Moore Levels 3–4 | Limited |
| 3. Retention and transfer | Is improvement maintained and generalised? | Delayed test, novel case, different simulator or setting | Cross-cuts Moore Levels 3–4 and bridges to Level 5 | Very limited |
| 4. Workplace performance | Does professional behaviour or practice change? | Audit, observed practice, adherence, escalation and calibration | Moore Level 5 | Minimal |
| 5. Patient, team or system outcome | Does implementation improve safety, equity, capacity or health outcomes? | Patient outcomes, teamwork, cost-effectiveness and access | Moore Levels 6–7, extended to team/system outcomes | Not established |
An AI-specific hierarchy that distinguishes technical and assisted performance from progressively stronger educational and professional outcomes, with approximate mapping to Moore Levels 1–7.
Figure 1.

AI-enabled orthopaedic CME/CPD outcomes hierarchy and approximate relationship to Moore’s framework.
The framework is not intended to replace Moore’s outcomes framework. Its primary purpose is as an AI-specific evaluation and research taxonomy, with a secondary use as a curriculum-design tool: investigators and CPD providers can pre-specify the level of outcome they expect an intervention to influence and align assessment accordingly. Moore Levels 1–2 (participation and satisfaction) remain relevant process measures but are not included as evidence of AI educational effectiveness. The proposed Level 0 (model capability) has no Moore equivalent because it evaluates the technology rather than the learner. Level 1 (assisted performance) is also separated deliberately: better performance while AI is present may occur without durable learning and should not automatically be labelled Moore Level 5 performance. Proposed Level 2 broadly corresponds to Moore Levels 3–4 (learning and competence), Level 3 adds explicit retention and transfer across contexts, Level 4 corresponds to Moore Level 5 workplace performance, and Level 5 encompasses Moore Levels 6–7 while extending attention to team, system, equity and cost outcomes.
At Level 0, technical validation asks whether the AI can complete or score the task. Level 1 assesses whether the clinician performs better while the tool is present. Level 2 asks whether unaided knowledge, reasoning or skill improves. Level 3 asks whether that improvement is retained and generalised to a different case, simulator or setting. Level 4 examines actual workplace behaviour, including how clinicians respond when the AI is uncertain or wrong. Level 5 concerns patient, team and system outcomes, including safety, equity, supervision capacity and cost-effectiveness.
This hierarchy changes the design and interpretation of educational studies. An LLM benchmark should not be described as evidence of educational effectiveness. A diagnostic-support trial should include unaided follow-up. A simulator classifier should be evaluated as a feedback intervention rather than only as a classifier. An AI literacy course should assess behaviour using flawed or uncertain outputs, not only confidence or satisfaction (Table 3). The framework therefore links the central evidence gap in this review to a practical research rule: claims should not exceed the outcome level actually measured.
Table 3.
Working definitions of AI terminology used in this review.
| Term | Working definition |
|---|---|
| Hallucinated content or reasoning | An answer, explanation or reasoning chain that is presented plausibly but is unsupported, internally inconsistent or not grounded in the supplied information or reliable evidence. |
| Vision dependence | Variation in model performance according to whether visual information is available and how well the model can interpret it; image-rich tasks may perform differently from text-only tasks. |
| AI boundaries | The task, population, anatomy, modality, data distribution, workflow and model/version conditions within which performance has been developed or evaluated. Performance should not be assumed outside these boundaries. |
| Machine learning | Methods in which computational models learn patterns from data to make predictions, classifications or other outputs without every decision rule being explicitly programmed. |
| Deep learning | A subset of machine learning using multilayer neural networks, commonly applied to imaging, language and complex high-dimensional data. |
| Computer vision | AI methods for analysing visual data such as radiographs, operative video or simulator images. |
| Large language model (LLM) | A generative model trained on large text corpora to predict and generate language; capabilities and limitations vary by model, version, prompting and access to external tools or sources. |
| Vision-language model (VLM) | A model that combines language processing with visual inputs, enabling joint reasoning over text and images. |
| ChatGPT and other LLM interfaces | ChatGPT is one interface/model family and is not synonymous with all LLMs. Orthopaedic evidence increasingly includes other proprietary, open-source and multimodal models. |
| Ambient voice technology (AVT) | Speech-enabled systems that capture clinical conversations and use AI to generate transcription, summaries or structured documentation for clinician review. |
Definitions are intended to clarify how terms are used in the educational argument rather than provide exhaustive technical specifications.
Designing AI-Enabled Orthopaedic CPD
Four design principles emerge when the orthopaedic evidence is read alongside established CME/CPD design. First, education should begin with a documented professional practice gap, an established principle of CME planning rather than a novel inference from AI studies [32,33]. This matters particularly in the present evidence base because the apparent benefits of AI are task- and modality-specific; technology should be selected only when it offers a plausible mechanism for addressing an identified gap. Second, activities should require active judgement. The diagnostic and LLM literature shows why learners should compare, critique, explain and decide rather than passively accept generated output [10,11,22,23,26]. Third, supervised application should precede independent use, especially for procedural or patient-facing systems, because strong simulator classification has not yet established transfer to theatre performance [12–15]. Fourth, outcome evaluation should extend beyond immediate satisfaction or assisted performance and include unaided learning, retention, transfer or workplace behaviour whenever feasible; the scarcity of such outcomes is itself a consistent finding of the review [27,28,31,32].
For diagnostic AI, this may involve case sets containing correct, incorrect and uncertain model outputs, followed by delayed unaided assessment. For procedural simulation, it may involve baseline assessment, AI-guided deliberate practice, expert debriefing and transfer to a novel task. For LLM use, it may involve source verification, prompt documentation, recognition of hallucinated content and explicit escalation when uncertainty remains. For CPD providers using generative AI, quality assurance should include human ownership of learning objectives, independent review of clinical content, documentation of tool and version, and disclosure where AI materially contributed to educational material.
AI also creates opportunities for practice-based CPD. With appropriate governance, routinely collected performance data could identify learning needs and provide targeted feedback. However, surveillance concerns, data quality, fairness and the separation of formative development from punitive performance management require careful attention. Clinicians are unlikely to engage honestly with AI-generated feedback if the purpose and consequences of data use are unclear.
Implementation, Ethics and Governance
Educational implementation raises issues that benchmark studies rarely capture. Images, operative video, simulator logs, prompts and learner analytics may contain identifiable or sensitive information. Commercial AI platforms can create uncertainty about data storage, secondary use and model training. Local governance, data minimisation and clear institutional policies are essential before educational or clinical material is uploaded to external systems.
Ambient voice technology provides a near-term example of why education about AI cannot wait for orthopaedic-specific educational trials. NHS England is supporting scaled adoption of AI-enabled ambient scribing and explicitly advises that staff understand product capabilities and limitations, review outputs for inaccuracies, and address privacy, automation bias and accountability within governance arrangements [34]. These systems are not orthopaedic teaching tools, and no educational benefit should be inferred from their clinical deployment. Nevertheless, orthopaedic clinicians who use them may require CPD in consent, verification, safe documentation and recognition of clinically important omissions or distortions.
Bias can arise from narrow datasets, unrepresentative learners, institution-specific assessment practices or unequal access to advanced hardware and paid models. Version drift creates an additional problem: an activity validated with one model may not behave the same way months later. CME/CPD reports should therefore document the tool, version, access date, input modality, prompts or task structure, human oversight and relevant subgroup performance.
Professional accountability remains human. AI can contribute to an explanation or recommendation, but it cannot assume responsibility for patient care, competency decisions or the independence of accredited education. Relevant reporting guidance includes CLAIM for imaging studies, TRIPOD+AI for prediction models and CONSORT-AI for interventional trials [35–37]. Educational studies should additionally report the learner group, comparator, follow-up, degree of unaided assessment and intended level of CME/CPD outcome.
Research Priorities
The next generation of orthopaedic AI education research should include practising clinicians and measure outcomes that matter to CPD. Priorities are: multicentre randomised or well-designed quasi-experimental evaluations; delayed unaided assessment; transfer from simulation to cadaveric or workplace performance; complete multimodal tasks; behaviour when AI advice is deliberately incorrect; comparison of AI, expert and combined feedback; equity and cost across differently resourced settings; and transparent reporting of model version and human oversight.
A particularly valuable design would recruit residents and practising surgeons learning the same new orthopaedic task, compare expert feedback, AI feedback and a combined approach, and assess unaided performance on a novel case after a delay. Such a study would test whether the intervention supports learning across career stages and whether the effect reaches competence or workplace performance rather than stopping at tool-assisted accuracy.
Limitations
This is a critical narrative review rather than a comprehensive systematic review. It was anchored to published orthopaedic reviews and supplemented by a targeted update, but searching was not independently duplicated across all subscription databases and formal risk-of-bias assessment was not undertaken for every primary study. Direct orthopaedic CME/CPD intervention evidence is sparse, so several proposed applications for practising surgeons are explicit extrapolations from undergraduate, postgraduate and broader healthcare-professional evidence. The LLM evidence is also disproportionately concentrated on ChatGPT/OpenAI-family evaluations, although newer studies include open-source vision-language models and DeepSeek-V3 [24,27]; findings from one model family, version or access mode should therefore not be generalised to all LLMs. The field is changing rapidly, and model-specific results may become outdated quickly. Finally, the proposed outcomes framework is interpretive and has not been prospectively validated.
Conclusion
AI in orthopaedic education is in active use and is no longer speculative. Demonstrated educational value remains limited relative to technological capability. The strongest evidence concerns bounded tasks with objective outputs, including fracture-interpretation support and simulator-based assessment. LLMs can perform strongly on selected text-based examinations, yet explanation quality, hallucinated reasoning, model boundaries and image interpretation remain important concerns. Emerging learner-outcome evidence is encouraging, but direct evidence of workplace practice change and patient benefit is not established.
For CME/CPD, the priority is not simply to use AI more often, but to ensure that AI-enabled education produces demonstrable learning, safe transfer, and improved practice. Orthopaedic clinicians need to learn with AI, appraise AI, and remain accountable when AI influences care. Programmes should begin with a professional practice gap, require active judgement, include governance, and measure outcomes beyond the model’s own performance. Until stronger evidence develops, AI should extend expert-led professional development rather than replace it.
Funding Statement
This work received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.
Disclosure statement
No potential conflict of interestwas reported by the author(s).
OpenAI ChatGPT (GPT-5.6, accessed 4 August and 7 September 2026) was used to assist with manuscript restructuring, language editing and bibliographic checking. The author independently reviewed and revised all output, checked the cited sources and accepts responsibility for the accuracy, originality and integrity of the manuscript.
Ethics statement
Ethics approval and consent were not required because this article reviews published literature and involved no new data collection or identifiable participant information.
References
- [1].Al-Saadawi A, Tehranchi S, Ahmed S, et al. Exploring the current applications of artificial intelligence in orthopaedic surgical training: a systematic scoping review. Cureus. 2025;17(4):e81671. doi: 10.7759/cureus.81671 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [2].Banskota B, Bhusal R, Yadav PK, et al. Artificial intelligence in orthopaedic education, training and research: a systematic review. BMC Med Educ. 2025;25(1):1594. doi: 10.1186/s12909-025-08162-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- [3].Szatkowski JP, Druten E, Soni C, et al. Artificial intelligence in orthopaedic education: a narrative review. Ann Joint. 2025;10:34. doi: 10.21037/aoj-25-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Ren K, Weng Q, Chen Q, et al. The application of large language models in orthopedic postgraduate education: potentials, challenges, and future prospects. J Orthop Surg Res. 2026;21(1):339. doi: 10.1186/s13018-026-06844-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Merino Molina V, Barrera Medina RA, Martinez Mora JS, et al. Extended reality and emerging artificial intelligence in orthopedic surgical training: a scoping review of educational outcomes. J Robot Surg. 2026;20(1):526. doi: 10.1007/s11701-026-03230-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- [6].Leal JA. Artificial intelligence in orthopaedic and trauma surgery education: applications, ethics, and future perspectives. J Am Acad Orthop Surg Glob Res Rev. 2025;9(9):e25.00174. doi: 10.5435/JAAOSGlobal-D-25-00174 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [7].Gordon M, Daniel M, Ajiboye A, et al. A scoping review of artificial intelligence in medical education: beme Guide No. 84. Med Teach. 2024;46(4):446–10. doi: 10.1080/0142159X.2024.2314198 [DOI] [PubMed] [Google Scholar]
- [8].Kirubarajan A, Young D, Khan S, et al. Artificial intelligence and surgical education: a systematic scoping review of interventions. J Surg Educ. 2022;79(2):500–515. doi: 10.1016/j.jsurg.2021.09.012 [DOI] [PubMed] [Google Scholar]
- [9].Oner SK, Ocak B, Demirkiran ND, et al. Artificial intelligence in orthopedic residency training in Turkiye: a multicenter cross-sectional survey. Med (Baltim). 2026;105(20):e48792. doi: 10.1097/MD.0000000000048792 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [10].Cheng CT, Chen CC, Fu CY, et al. Artificial intelligence-based education assists medical students’ interpretation of hip fracture. Insights Imag. 2020;11(1):119. doi: 10.1186/s13244-020-00932-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [11].Kim T, Goh TS, Lee JS, et al. Transfer learning-based ensemble convolutional neural network for accelerated diagnosis of foot fractures. Phys Eng Sci Med. 2023;46(1):265–277. doi: 10.1007/s13246-023-01215-w [DOI] [PubMed] [Google Scholar]
- [12].Bissonnette V, Mirchi N, Ledwos N, et al. Artificial intelligence distinguishes surgical training levels in a virtual reality spinal task. J Bone Joint Surg Am. 2019;101(23):e127. doi: 10.2106/JBJS.18.01197 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [13].Alkadri S, Ledwos N, Mirchi N, et al. Utilizing a multilayer perceptron artificial neural network to assess a virtual reality surgical procedure. Comput Biol Med. 2021;136:104770. doi: 10.1016/j.compbiomed.2021.104770 [DOI] [PubMed] [Google Scholar]
- [14].Demirel D, Palmer B, Sundberg G, et al. Scoring metrics for assessing skills in arthroscopic rotator cuff repair: performance comparison study of novice and expert surgeons. Int J CARS. 2022;17(10):1823–1835. doi: 10.1007/s11548-022-02683-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [15].Winkler-Schwartz A, Bissonnette V, Mirchi N, et al. Artificial intelligence in medical education: best practices using machine learning to assess surgical expertise in virtual reality simulation. J Surg Educ. 2019;76(6):1681–1690. doi: 10.1016/j.jsurg.2019.05.015 [DOI] [PubMed] [Google Scholar]
- [16].Lum ZC. Can artificial intelligence pass the American board of orthopaedic surgery examination? Orthopaedic residents versus ChatGPT. Clin Orthop Relat Res. 2023;481(8):1623–1630. doi: 10.1097/CORR.0000000000002704 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [17].Kung JE, Marshall C, Gauthier C, et al. Evaluating ChatGPT performance on the orthopaedic in-training examination. JBJS Open Access. 2023;8(3):e23.00056. doi: 10.2106/JBJS.OA.23.00056 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Massey PA, Montgomery C, Zhang AS. Comparison of ChatGPT-3.5, ChatGPT-4, and orthopaedic resident performance on orthopaedic assessment examinations. J Am Acad Orthop Surg. 2023;31(23):1173–1179. doi: 10.5435/JAAOS-D-23-00396 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [19].Khan AM, Sarraf KM, Simpson AI. Enhancements in artificial intelligence for medical examinations: a leap from ChatGPT 3.5 to ChatGPT 4.0 in the FRCS trauma and orthopaedics examination. Surgeon. 2025;23(1):13–17. doi: 10.1016/j.surge.2024.11.008 [DOI] [PubMed] [Google Scholar]
- [20].Diniz P, Oeding JF, Dean MC, et al. Reasoning-optimised large language models reach near-expert accuracy on board-style orthopaedic exams: a multi-model comparison on 702 multiple-choice questions. Knee Surg Sports Traumatol Arthrosc. 2026;34(2):752–762. doi: 10.1002/ksa.70222 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Hayes DS, Barre A, Muchow RD. Assessing the third wave of generative AI: performance of advanced models on text-based questions from the 2024 orthopaedic in-training examination. J Am Acad Orthop Surg. 2026;34(12):e1636–e1648. doi: 10.5435/JAAOS-D-25-00441 [DOI] [PubMed] [Google Scholar]
- [22].Hayes DS, Foster BK, Makar GS, et al. Artificial intelligence in orthopaedics: performance of ChatGPT on text and image questions on a complete AAOS orthopaedic in-training examination. J Surg Educ. 2024;81(11):1645–1649. doi: 10.1016/j.jsurg.2024.08.002 [DOI] [PubMed] [Google Scholar]
- [23].Jain N, Gottlich C, Fisher J, et al. ChatGPT-4o is not a reliable study source for orthopaedic surgery residents. JBJS Open Access. 2025;10(3):e25.00112. doi: 10.2106/JBJS.OA.25.00112 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Ko S, Lee J, Ko K, et al. Benchmarking open-source vision language models in orthopedic in-training examination: a comparison with residents, domain-specific evaluation, and parameter scaling. Clin Orthop Surg. 2026;18(1):159–166. doi: 10.4055/cios25183 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25].DeCook R, Muffly BT, Mahmood S, et al. AI-generated graduate medical education content for total joint arthroplasty: comparing ChatGPT against orthopaedic fellows. Arthroplast Today. 2024;27:101412. doi: 10.1016/j.artd.2024.101412 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Dagher T, Haydon R, Wolf JM, et al. Can ChatGPT-4 replace orthopedic question banks? Evaluating AI’s ability to write questions for the orthopedic in-training exam. Curr Probl Surg. 2026;74:101943. doi: 10.1016/j.cpsurg.2025.101943 [DOI] [PubMed] [Google Scholar]
- [27].Yu C, Li F, Zhang N, et al. Effectiveness of artificial intelligence-assisted peer teaching in orthopedic clinical education: historical cohort study. JMIR Med Educ. 2026;12:e87959. doi: 10.2196/87959 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Woods L, Lyons K, Van Der Vegt A, et al. Assessing the effectiveness of artificial intelligence education and training for healthcare workers: a systematic review. BMC Med Educ. 2026;26(1):549. doi: 10.1186/s12909-026-08969-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Charow R, Jeyakumar T, Younus S, et al. Artificial intelligence education programs for health care professionals: scoping review. JMIR Med Educ. 2021;7(4):e31043. doi: 10.2196/31043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [30].Crim A, Skidmore RM, Bowser AD, et al. The Alliance for continuing education in the health professions’ position on artificial intelligence in cpd. J Cme. 2025;14(1):2545647. doi: 10.1080/28338073.2025.2545647 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [31].Samuel A, Cervero RM. Continuing professional development in the artificial intelligence era. J Contin Educ Health Prof. 2026. Online ahead of print. 46(3):160–165. doi: 10.1097/CEH.0000000000000668 [DOI] [PubMed] [Google Scholar]
- [32].De M Jr, Green JS, Gallis HA. Achieving desired results and improved outcomes: integrating planning and assessment throughout learning activities. J Contin Educ Health Prof. 2009;29(1):1–15. doi: 10.1002/chp.20001 [DOI] [PubMed] [Google Scholar]
- [33].Wittich CM, Chutka DS, Mauck KF, et al. Perspective: a practical approach to defining professional practice gaps for continuing medical education. Acad Med. 2012;87(5):582–585. doi: 10.1097/ACM.0b013e31824d4d5f [DOI] [PubMed] [Google Scholar]
- [34].NHS England . Guidance on the use of AI-enabled ambient scribing products in health and care settings. Version 3 Published. 2025; [cited 2026 Sep 7]. Available from: https://www.england.nhs.uk/long-read/guidance-on-the-use-of-ai-enabled-ambient-scribing-products-in-health-and-care-settings/
- [35].Collins GS, Dhiman P, Navarro CLA, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
- [36].Liu X, Cruz Rivera S, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26(9):1364–1374. doi: 10.1038/s41591-020-1034-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for artificial intelligence in medical imaging (claim): 2024 update. Radiol Artif Intell. 2024;6(4):e240300. doi: 10.1148/ryai.240300 [DOI] [PMC free article] [PubMed] [Google Scholar]
