Abstract
Objective
Surgical outcomes vary widely, in part due to variability in surgeon performance. However, current strategies for assessing and improving surgical performance suffer from issues of objectivity and scalability due to reliance on human evaluators. In this narrative review, we review recent advances in the use of artificial intelligence to address these shortcomings by automating surgical skills assessment and feedback.
Methods
We searched PubMed for studies published between 2015 and 2025 pertaining to artificial intelligence for surgical training. Search terms included “artificial intelligence” or “machine learning” or “deep learning” and “surgical feedback” or “surgical training” or “surgical skill.” Articles were identified with special attention given to those published in the last 5 years with a focus on AI for surgical skill assessment or feedback.
Results
AI has been used to successfully automate surgical skill assessment across a variety of surgical disciplines via approaches such as kinematics, sabermetrics, computer vision, and gesture analysis. Many of these studies have developed AI models capable of a binary classification of skill (novice versus expert) which demonstrate concordance when verified against ground truths from human raters. Based on these skills assessments, AI approaches may be further leveraged to generate automatic feedback, which has proven effective in improving surgeon performance metrics, particularly for underperformers. AI has also shown utility in categorizing and analyzing the content and impact of live surgical feedback, enabling more efficient analysis of how feedback can be best delivered to trainees.
Conclusions
AI is a promising tool for augmenting surgical training and improving the objectivity and scalability of surgical skill assessment and feedback. Thus far, AI models are adept at detecting relatively large differences in surgical performance and providing rudimentary feedback. Further work is required to create models capable of doing more fine-tuned skill assessments and generating more detailed, constructive feedback.
Keywords: artificial intelligence, surgical training, surgical education, surgical assessment, surgical skill, surgical feedback, deep learning, machine learning
1. Introduction/Background
Modern day surgical techniques are constantly and rapidly evolving. As such, there is a need for optimizing surgical training, assessment, and feedback, both for surgical trainees as well as experienced surgeons adapting to new surgical tools and techniques. Artificial intelligence presents a novel and promising tool for revolutionizing surgical education. This narrative review explores the evolution and current landscape of AI applications for surgical training, from skills assessment to automated feedback (see key references in Box 1).
Box 1: Key references highlighting AI-based approaches to surgical training.
- 1.Ghodoussipour S, Reddy SS, Ma R, Hwang DH, Nguyen J, Hung AJ. An Objective Assessment of Performance during Robotic Partial Nephrectomy: Validation and Correlation of Automated Performance Metrics with Intraoperative Outcomes. Journal of Urology. 2021. May;205(5):1294–302. [DOI] [PubMed] [Google Scholar]
- 2.Shafiei SB, Shadpour S, Mohler JL, Rashidi P, Toussi MS, Liu Q, et al. Prediction of Robotic Anastomosis Competency Evaluation (RACE) metrics during vesico-urethral anastomosis using electroencephalography, eye-tracking, and machine learning. Sci Rep. 2024. June 25;14(1):14611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lavanchy JL, Zindel J, Kirtac K, Twick I, Hosgor E, Candinas D, et al. Automation of surgical skill assessment using a three-stage machine learning algorithm. Sci Rep [Internet]. 2021. Mar 4 [cited 2025 July 24];11(1). Available from: https://www.nature.com/articles/s41598-021-84295-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ma R, Ramaswamy A, Xu J, Trinh L, Kiyasseh D, Chu TN, et al. Surgical gestures as a method to quantify surgical performance and predict patient outcomes. npj Digit Med. 2022. Dec 22;5(1):187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Ma R, Kiyasseh D, Laca JA, Kocielnik R, Wong EY, Chu TN, et al. Artificial Intelligence-Based Video Feedback to Improve Novice Performance on Robotic Suturing Skills: A Pilot Study. Journal of Endourology. 2024. Aug 1;38(8):884–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
2. Methods
We searched PubMed for studies published between 2015 and 2025 pertaining to artificial intelligence for surgical training. Search terms included “artificial intelligence” or “machine learning” or “deep learning” and “surgical feedback” or “surgical training” or “surgical skill.” Articles were identified with special attention given to those published in the last 5 years with a focus on AI for surgical skill assessment or feedback.
3. Defining Surgical Performance and Current Challenges in Surgical Training, Assessment, & Feedback
Surgical performance is a multifactorial concept and consists of both technical skills (such as precision, dexterity, and tissue handling) as well as non-technical skills (such as decision-making, situational awareness, teamwork and communication) [1]. There is a growing body of literature demonstrating a clear correlation between surgical performance and patient outcomes [2-5]. Therefore, precise and accurate assessment of surgical performance is the necessary foundation for any novel surgical training method. In recent years, a variety of tools have been developed with this aim, which may be divided into 3 broad categories:
Global Rating Scales: score various domains of surgical skill on a Likert scale; these may be either generic, e.g. Global Evaluative Assessment of Robotic Skills (GEARS) [6], End-To-End Assessment of Suturing Expertise (EASE) [7], Global Operative Assessment of Laparoscopic Skills (GOALS) [8], or procedure/task-specific, e.g. Robotic Anastomosis Competency Evaluation (RACE) for the urethrovesical anastomosis step of robotic-assisted radical prostatectomy [9], Global Operative Assessment of Laparoscopic Skills Groin Hernia (GOALS-GH) for laparoscopic inguinal hernia repair [10].
Task specific checklists: procedure-specific, step-wise checklists where each item is scored in a binary fashion (done correctly vs. not done correctly) [11, 12].
Safety metrics/simulation-based assessment: e.g. FLS (Fundamentals of Laparoscopic Surgery) and Fundamentals of Robotic Surgery (FRS)/Fundamental Skills of Robotic Surgery (FSRS), certification exams with corresponding curricula for achieving proficiency in basic laparoscopic or robotic skills. Assessments are based on a pre-determined set of tasks, each focused on a specific skill, which participants must complete within the allocated time limit and without exceeding allowable errors [13-16].
Despite these existing tools, the current landscape of surgical performance assessment continues to face several key challenges:
Objectively quantifying surgical performance: the above methods are reliant on individual raters and therefore retain some human subjectivity and bias; further, many aspects of surgery are difficult to quantify
Defining benchmarks in high stakes assessments: on the spectrum of surgical skill, defining a threshold for adequate vs. inadequate performance remains a high-stakes objective
Inter-assessor variability and intra-rater reliability between human evaluators: the current model of surgical training overwhelmingly relies on assessment and feedback delivered on a case-by-case basis from a more senior, experienced attending surgeon to a less-experienced trainee, resulting in the potential for significant variability between assessments by different assessors, or even if the assessor is consistent, variability between cases
Assessment Burnout / Scalability: current assessment methods that require human raters are not feasible for widespread implementation due to being too labor intensive; as such, trainees may suffer from a paucity of consistent, good quality feedback
Unclear predictive validity of current feedback models
In the following sections, we will describe the latest efforts utilizing AI to address these unmet challenges. By automating surgical skill assessment and feedback, AI has the potential to significantly improve the objectivity of evaluations, mitigate human bias, and enable frequent, regular, and reliable surgical feedback delivery, thereby improving surgical performance and patient outcomes.
4. Artificial Intelligence for Surgical Skills Assessment
Existing efforts to automate surgical skills assessment using AI tend to follow one of the following approaches- kinematics, sabermetrics, computer vision, and gestures- although other AI-based approaches are also being explored.
A. Kinematics
Kinematics refers to characteristics of movement such as instrument travel time, path length, and velocity. As robotic surgery becomes increasingly prevalent, automated recording of instrument actions by robotic surgical systems has enabled the large-scale collection of kinematic data available to be analyzed with artificial intelligence. Such data has previously been applied to automate assessment of detailed suturing technical skills: Hung et al first showed that using robotic instrument kinematic data, an AI model was able to score surgeons completing a suturing exercise on subdomains including needle entry-angle, needle driving, and needle withdrawal with decent performance (AUCs ranging from 0.698 to 0.705 for the above 3 domains) [17]. Compared to using human raters to repeatedly score each subdomain for each stitch, this automated method is significantly less labor-intensive for suturing skill assessment.
Using kinematics, AI is also capable of distinguishing between novice and expert surgeons. In a study by Fard et al, 8 global movement features such as speed, motion smoothness, and path length were collected from surgeons with varying robotic experience while performing knot tying and suturing using the da Vinci robot. Surgeons were then scored using global rating scores to identify novices versus experts. Using the aforementioned kinematic features, the AI model was able to correctly classify surgeon skill level with 82.3% and 89.9% accuracy for knot tying and suturing, respectively [18]. Similarly, in a study of 50 robotic-assisted partial nephrectomies, authors found significant differences in kinematics profiles between expert and trainee surgeons during 5 key steps: colon mobilization, opening of Gerota’s fascia and lifting the ureter, hilar dissection, defatting of the kidney/exposure of the tumor, and intraoperative ultrasound/scoring of the tumor. Generally, experts were more efficient and directed in their movements, indicated by decreased instrument moving times, shorter instrument path lengths, faster instrument velocities, and reduced overall task duration [19].
Machine learning algorithms have further leveraged kinematic data to predict postoperative outcomes as a surrogate marker for surgical performance. Hung et al showed that kinematic data was able not only to help evaluate surgical performance during robotic assisted radical prostatectomy but also predict outcomes such as surgery time, length of stay, foley catheter duration, and urinary continence recovery [20, 21]. While other studies frequently rely on skills assessment scores by human raters for ground truth data, the demonstrated relationship between kinematics metrics with postoperative outcomes lends an additional degree of credibility to kinematics as a predictor of surgical performance.
As the field of AI for surgical skill assessment evolves, kinematics information is being increasingly combined with other data forms to develop more robust AI models incorporating multiple inputs. Cui et al recently proposed a framework for skill assessment considering the interconnected nature of skills required in surgery. In their study, they divided suturing into 6 sub-skills: needle repositioning, needle hold ratio, needle hold angle, driving smoothness, wrist rotation, and wrist rotation needle withdrawal. Knowing that performance in these individual sub-skills is deeply interconnected (e.g. needle withdrawal performance is more likely to be ideal when needle hold angle is ideal), they created an AI system incorporating the known relationships between different suturing sub-skills to simultaneously score each sub-skill domain. The machine learning model described in their study then extracted features of the above 6 sub-skills from input videos and combined such data with kinematic information in order to predict 3 month urinary continence recovery after robotic-assisted radical prostatectomy, with a mean AUC of 0.70 +/− 0.02 across all 6 sub-skills, an improvement compared to 0.67 ± 0.02 with video data alone [22]. Here, AI demonstrates its effectiveness at not only automating the assessment of suturing skills but also interpreting detailed kinematic information to predict objective functional outcomes after surgery, translating complex surgical data into quantifiable and analyzable metrics.
B. Sabermetrics
AI has also been applied to other digital metrics with the goal of automating surgical performance assessment. Inspired by performance monitoring in elite sports, machine-learning has been used to integrate audio-visual capture of the operating room with sensor-based measurements to analyze surgical performance, a practice termed “surgical sabermetrics” [23]. Research in this area thus far has identified surgeon cognitive load as an area of particular interest, as excessive amounts have been shown to have detrimental effects on surgeons’ technical performance. Sensor-based measurements have been applied both to physiologic metrics (such as heart rate variability, eye-tracking, and electroencephalography [EEG]) as well as nonphysiologic ones (for example, movement metrics using proximity sensors). This data may be used to holistically analyze surgical performance during cognitively challenging tasks, track improvement in the cognitive load required to complete a task with increasing proficiency or following training interventions, and help develop effective strategies for mitigating the negative effects of high cognitive load or stress during critical portions of the procedure.
Shafiei et al has similarly investigated the utility of collecting sensor-based information for surgical skills assessment in robotic assisted surgery. In their recent studies, they focus on EEG and eye-tracking as metrics to predict surgical performance. EEG sheds light on the cognitive and motor processes involved in performing a surgical task, as well as on the surgeon’s cognitive load, providing insight into levels of concentration, fatigue, and stress. Eye-tracking enables non-intrusive measurement of visual attention, with experienced surgeons typically demonstrating more efficient and purposeful eye movements, and less experienced surgeons exhibiting more frequent and random eye saccades. In one study, EEG and eye tracking data was recorded from 23 participants of different skill levels (inexperienced, competent, and experienced) performing robotic-assisted vesico-urethral anastomoses on both plastic and animal models. All participants were scored by 3 independent raters using the Robotic Anastomosis Competency Evaluation (RACE) on domains such as needle positioning, needle entry, needle driving and tissue trauma, suture placement, tissue approximation, and knot-tying. Using the collected EEG and eye-tracking data, machine learning models were able to accurately predict RACE scores across all skill levels [24]. In another study Shafiei et al, EEG and eye tracking data was collected from 11 physicians of various skill levels and surgical specialties during robotic surgery on live pigs (11 hysterectomies, 11, cystectomies, and 21 nephrectomies). Three primary subtasks from all operative videos were extracted- blunt, cold sharp, and thermal dissection- and assessed using GEARS by an expert robotic surgeon. Again, machine-learning models were able to use EEG and eye-tracking data to accurately predict GEARS scores, illustrating the utility of AI in empowering novel, objective, and automated approaches to surgical skill assessment [25].
C. Computer Vision
Termed “computer vision,” AI has been applied towards the identification and classification of objects in images and videos, resulting in the generation of algorithms capable of human-level object detection. These models may be used not only to derive kinematics data from the identification surgical instruments, but also to recognize anatomic structures, identify safety hazards, and detect differences in surgeon skill level using metrics such as fluidity of motion, economy of motion, tissue handling, bimanual dexterity, instrument use, and total operative duration. In a 2023 study, a deep learning model was trained to identify and localize instances of surgical instruments from recorded neurosurgery videos. The model was able to compute tool motion and tool handling metrics by tracking the detected instrument locations through time. When compared to those of novice surgeons, instrument metrics of expert surgeons exhibited features associated with more controlled, efficient, coordinated, and confident movements- for example, lower velocity, lower acceleration, lower jerks, reduced path length, increased bi-manual handling, shorter idle time, and smaller inter tool-tip distances [26].
Computer vision algorithms like these have been shown to be successful at assessing degree of surgical skill across multiple surgical approaches (open, laparoscopic, robotic, and endoscopic) and specialty areas. Azari et al employed computer vision-based motion tracking of surgeon hand movements to accurately predict OSATS scores from expert raters in 16 open surgeries, spanning colorectal, complex upper gastrointestinal, hepatobiliary, surgical oncology, transplant, vascular, thoracic, and cardiac operations [27]. In the realm of ophthalmology, computer vision tracking of surgical instruments during cataract surgery was able to distinguish between junior and senior surgeons using metrics such as instrument path length and number of instrument movements [28]. A similar study in otolaryngology found that 3 computer vision derived metrics - total number of movements, camera path length, and surgical time - predicted surgeon skill level during endoscopic dacryocystorhinostomy [29]. In another study, Lavanchy et al developed a machine learning algorithm to automate surgical skill assessment using video footage of laparoscopic cholecystectomies. Good surgical performance was determined by narrow and focused instrument handling within the operating field and poor surgical performance was indicated by frequent changes of instrument direction within a larger field. Their algorithm was able to distinguish between good versus poor surgical performance with an 87% accuracy [30].
In addition to being able to assess surgical skill in the operating room, these AI tools can accurately classify surgical expertise in virtual reality (VR) surgical simulators. The Virtual Operative Assistant described by Mirchi et al is one example of a surgical system incorporating AI to assess proficiency in subpial brain tumor resection using a virtual reality platform. The study showed that the Virtual Operative Assistant accurately identified all skilled participants and 82% of novice participants [31]. These studies collectively illustrate AI’s effectiveness at a binary classification of surgical skill – novice vs. expert – via automated methods that mitigate subjectivity and bias, minimize the need for labor intensive human involvement, and enable scalability.
Yet another application of computer vision is the recognition of standardized surgical field development. The best described instance of standardized surgical field development, the critical view of safety during laparoscopic cholecystectomy, has been studied as an opportunity for artificial intelligence to automatically segment hepatocystic anatomy and enhance surgical safety: Mascagni et al previously created a deep learning model capable of assessing achievement of the critical view of safety with a mean average precision and balanced accuracy of 71.9% and 71.4%, respectively [32]. Similarly, Igaki et al were able to use deep learning to automate assessment of standardized surgical field development in laparoscopic sigmoid colon resection. Their model’s outputs strongly correlated with performance scores from two independent raters (Spearman rank correlation coefficient of 0.81), and were able to detect low-scoring and high-scoring surgeons with AUCs of 0.93 and 0.94, respectively [33].
Phase recognition represents another area of computer vision showing promise in automating surgical skill assessment. In one study of 410 robotic-assisted radical prostatectomies from 18 Japanese medical centers, Zhao et al developed a deep-learning model with the ability to automatically distinguish between various steps of the operation with 89% accuracy [34]. They were able to predict expert surgeons from novices based on the duration of 3 critical steps (Retzius space expansion, dorsal venous complex incision/apex treatment, and urethrovesical anastomosis) with 86.2% accuracy. Similarly, Nakajima et al developed a deep learning model capable of recognizing the 5 surgical phases of laparoscopic sigmoidectomy with an overall 86% accuracy. Based on metrics such as time for mobilization of the colon, time for dissection of the mesorectum plus transection of the rectum plus anastomosis, and phase transition counts, their model was able to reliably distinguish between expert, intermediate, and novice surgeons [35].
D. Surgical Gestures
The ability to characterize and quantify surgical motion to a high degree of detail is fundamental to surgical data science. “Surgical gestures” represent one novel approach to achieving this goal. By segmenting surgical motion into moment-to-moment, granular actions, or “gestures,” long and complex operations may be deconstructed into strings of discrete moves amenable to quantitative analysis.
Various systems for gesture classification have been proposed. Ma et al developed the original gesture classification system consisting of 9 distinct dissection gestures (e.g. peel/push, spread, cold cut) and 4 supporting gestures (e.g. retraction, camera move), and found that based on gesture selection, machine learning models can be constructed to accurately predict erectile function recovery after radical prostatectomy [36].
Similarly, Van Amsterdam et al developed a list of 7 distinct gestures specific to the dorsal venous complex suturing phase of robotic-assisted radical prostatectomy (e.g. positioning the needle tip, pushing the needle through tissue, pulling the needle out of tissue, tying a knot, cutting a suture). Their AI model integrated both kinematic as well as visual data to automate surgical gesture recognition and was largely concordant with ground truth manual annotations [37]. Using the same dataset and gesture classification, Sirajudeen et al further applied deep learning in order to automatically predict errors (e.g. multiple attempts, needle drop, instruments out of view, excessive force) by gesture. They found that the greatest number of errors were made with the ‘pulling the needle out of the tissue’ gesture followed by ‘tying the knot,’ whereas the fewest errors were made with the ‘cutting the suture’ gesture [38].
Sato et al sought to differentiate between novice and expert surgeons performing robot-assisted radical prostatectomy via AI-automated recognition of dissection times (defined as duration of forceps energy activation) versus exposure times (defined as the combined duration of third arm manipulation and camera movement). They found that exposure time was significantly shorter for experts than for novices, but there was no significant difference in the dissection time between novices and experts. Consequently, the ratio of time spent dissecting to the time spent gaining exposure was significantly lower for the novice compared to expert group (0.174 vs. 0.315, p = 0.003) [39].
By reducing complicated operations into a standardized, digestible semantic vocabulary, these AI-based approaches are thus able to quantify and analyze the specific gestures and gesture sequences that determine excellent versus poor surgical skills, enabling clear, objective, and immediate assessment of surgical performance.
5. Artificial Intelligence for Providing and Analyzing Surgical Feedback
Feedback plays a fundamental role in surgical learning, for both immediate performance adjustments and long-term skill acquisition. Ineffective feedback intraoperatively can have grave consequences. Understanding how feedback is delivered, what constitutes optimal feedback, and how to deliver more consistent, useful, and frequent feedback to trainees is essential and requires additional inquiry. In recent years, AI has begun to transform the way feedback is understood and how it is delivered, making it more objective, personalized, and accessible.
One of the clearest applications of AI in this domain is in simulation-based robotic training. A 2024 pilot study by Ma et al demonstrated that AI-based video feedback significantly improved performance of novice surgeons during simulated robotic suturing exercises [40]. Forty-two participants with no robotic surgical experience were randomized to control or feedback groups and video-recorded while completing two rounds of a dry lab suturing task designed to simulate the vesicourethral anastomosis step of radical prostatectomy on the da Vinci surgical robot. Participants were assessed on needle handling and needle driving. Subjects in the feedback group were presented with representative video clips from their first round and informed of their AI-based skill assessment (where each stitch received a binary skill score- low skill or high skill) along with a textual explanation (e.g. “you successfully regrabbed the needle fewer than 3 times” for high skill or “to improve: minimize number of regrabs of the needle (<2 times)” for low skill). Meanwhile, subjects in the control group were only shown randomly selected video clips from the first round. Participants from each group were further labeled as underperformers or innate-performers based on a median split of their baseline technical skill scores from the first round. The feedback group had a significantly larger improvement in needle handling score compared with the control group. Furthermore, while innate-performers exhibited similar improvements across rounds regardless of feedback, underperformers in the feedback group improved significantly more than the control group in needle handling. While surgical trainees are typically reliant on their supervising attending surgeon for feedback, which may vary in quality and frequency due to time constraints, divided attention, patient safety concerns, or other limitations, this study demonstrates a role for AI in augmenting the consistency, quality and frequency of immediate, real-time feedback.
AI-based feedback has also been explored in laparoscopic surgery. A 2025 study introduced SmartCoach, an AI-enhanced coaching system applied to laparoscopic pancreatoduodenectomy (LPD), a technically demanding procedure typically reserved for highly experienced surgeons. While review and self-debriefing of surgical videos is a cornerstone of surgical education in the age of minimally invasive surgery, inefficiency and lack of expert feedback constitute major barriers to optimal learning. Their SmartCoach system provides real-time, automated intraoperative performance assessments, structured post-operative debriefings, and targeted improvement strategies for surgeons [41]. Their work, while still in its early stages, demonstrates that even experienced professionals can benefit from AI-guided feedback, particularly for complex, technically demanding procedures. In addition, AI can serve to enhance and streamline the otherwise extremely time-consuming process of video review and self-debriefing by providing focused feedback in the areas of greatest impact, thus making regular performance review more effective and feasible for busy surgeons.
In real-world surgical practice, AI has been used during robotic urologic procedures to categorize live surgical feedback to trainees. A 2024 study employed a human–AI collaborative approach to automatically cluster and categorize intraoperative feedback, enhancing the understanding of trainees’ educational needs by identifying recurring patterns [42]. Authors developed a semi-automated surgical feedback analysis framework which received raw text transcripts of surgical feedback delivered during live surgical cases and automatically detected themes throughout lines of feedback to create clinically meaningful categories or topics (e.g. instrument positioning, tissue and vein handling, identifying surgical planes). These feedback categories were then evaluated for their interpretability and ability to predict clinical outcomes: certain categories of feedback were found to be more likely to result in behavior change in the trainee (e.g. handling bleeding, sweeping techniques), whereas others were more likely to result in just verbal acknowledgement. By bypassing the need for human effort to interpret and categorize surgical transcripts, AI can therefore enable large scale analysis of real-world surgical feedback. This understanding is fundamental to expanding our ability to provide consistent, effective, and ultimately automated surgical feedback to trainees.
Further work is still required to ensure AI-based feedback is consistently reliable and unbiased for all surgeons. Kiyasesh et al carried out one such self-evaluation of their own AI system, which they previously showed to be capable of assessing surgeon skill from surgical video while simultaneously highlighting aspects of the video relevant to the assessment. However, when comparing AI-based explanations of skill assessment to explanations generated by human experts, they found that while explanations often aligned, they were not equally reliable for different sub-cohorts of surgeons (e.g. novices vs. experts). They were able to improve the reliability of AI-based explanations and mitigate this bias by using human explanations as supervision to explicitly teach their AI system how to highlight important video frames. However their findings demonstrate that caution is warranted when relying on novel AI-based feedback systems showing initial promise, and that further work is needed to vigorously test these systems prior to widespread implementation [43].
6. Conclusion
AI shows significant promise as an effective tool to automate surgical skill assessment and surgical feedback, making surgical training more objective, more accessible, and less laborious, with novel and increasingly sophisticated applications continuing to emerge as the field grows. The included studies collectively illustrate that current AI models for surgical skill assessment are facile at a binary classification of performance: novice versus skilled, inexperienced versus expert. The future of AI in this field, however, must turn towards both more global and granular analyses of surgical skill in-order to define benchmarks of proficiency for high-risk procedures, locate trainees along a learning curve for a particular task or operation, and improve applicability to real-world education. Refining surgical skill assessment will in turn empower AI to provide more detailed and constructive feedback to surgeons on how to improve their skills. Further study is still needed to design AI feedback models capable of not only providing assessments of skill and identifying areas for improvement, but also interpreting and communicating the specific adjustments necessary to achieve ideal performance.
Declaration of Interests Statement
Andrew J. Hung has received consulting fees for Medtronic. No other disclosures.
Research reported in this publication was supported by the National Cancer Institute of the National Institutes of Health under Award Numbers R01CA251579, R01CA273031, and R01CA298988. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
References
- [1].Collins JW, Dell’Oglio P, Hung AJ, Brook NR. The Importance of Technical and Nontechnical Skills in Robotic Surgery Training. European Urology Focus 2018; 4: 674–676. [DOI] [PubMed] [Google Scholar]
- [2].Birkmeyer JD, Finks JF, O’Reilly A, et al. Surgical Skill and Complication Rates after Bariatric Surgery. N Engl J Med 2013; 369: 1434–1442. [DOI] [PubMed] [Google Scholar]
- [3].Heard JR, Ghaffar U, Ma R, et al. Surgical Performance Metrics for 1-Year Patient-Reported Outcomes After Radical Prostatectomy. JAMA Surg 2025; 160: 674. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [4].Chu TN, Wong EY, Ma R, et al. A Multi-institution Study on the Association of Virtual Reality Skills with Continence Recovery after Robot-assisted Radical Prostatectomy. European Urology Focus 2023; 9: 1044–1051. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [5].Goldenberg MG, Goldenberg L, Grantcharov TP. Surgeon Performance Predicts Early Continence After Robot-Assisted Radical Prostatectomy. Journal of Endourology 2017; 31: 858–863. [DOI] [PubMed] [Google Scholar]
- [6].Aghazadeh MA, Jayaratna IS, Hung AJ, et al. External validation of Global Evaluative Assessment of Robotic Skills (GEARS). Surg Endosc 2015; 29: 3261–3266. [DOI] [PubMed] [Google Scholar]
- [7].Haque TF, Hui A, You J, et al. An Assessment Tool to Provide Targeted Feedback to Robotic Surgical Trainees: Development and Validation of the End-to-End Assessment of Suturing Expertise (EASE). Urology Practice 2022; 9: 532–539. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [8].Vassiliou MC, Feldman LS, Andrew CG, et al. A global assessment tool for evaluation of intraoperative laparoscopic skills. The American Journal of Surgery 2005; 190: 107–113. [DOI] [PubMed] [Google Scholar]
- [9].Raza SJ, Field E, Jay C, et al. Surgical Competency for Urethrovesical Anastomosis During Robot-assisted Radical Prostatectomy: Development and Validation of the Robotic Anastomosis Competency Evaluation. Urology 2015; 85: 27–32. [DOI] [PubMed] [Google Scholar]
- [10].Kurashima Y, Feldman LS, Al-Sabah S, Kaneva PA, Fried GM, Vassilou MC. A tool for training and evaluation of laparoscopic inguinal hernia repair: the Global Operative Assessment of Laparoscopic Skills-Groin Hernia (GOALS-GH). The American Journal of Surgery 2011; 201: 54–61. [DOI] [PubMed] [Google Scholar]
- [11].Campos MEC, Oliveira MMRD, Reis AB, Assis LBD, Iremashvili V. Development and validation a task-specific checklist for a microsurgical varicocelectomy simulation model. Int braz j urol 2020; 46: 796–802. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [12].Tang B. Identification and Categorization of Technical Errors by Observational Clinical Human Reliability Assessment (OCHRA) During Laparoscopic Cholecystectomy. Arch Surg 2004; 139: 1215. [DOI] [PubMed] [Google Scholar]
- [13].Higgins RM, Turbati MS, Goldblatt MI. Preparing for and passing the fundamentals of laparoscopic surgery (FLS) exam improves general surgery resident operative performance and autonomy. Surg Endosc 2023; 37: 6438–6444. [DOI] [PubMed] [Google Scholar]
- [14].Raza SJ, Froghi S, Chowriappa A, et al. Construct Validation of the Key Components of Fundamental Skills of Robotic Surgery (FSRS) Curriculum—A Multi-Institution Prospective Study. Journal of Surgical Education 2014; 71: 316–324. [DOI] [PubMed] [Google Scholar]
- [15].Stegemann AP, Ahmed K, Syed JR, et al. Fundamental Skills of Robotic Surgery: A Multi-institutional Randomized Controlled Trial for Validation of a Simulation-based Curriculum. Urology 2013; 81: 767–774. [DOI] [PubMed] [Google Scholar]
- [16].Satava RM, Stefanidis D, Levy JS, et al. Proving the Effectiveness of the Fundamentals of Robotic Surgery (FRS) Skills Curriculum: A Single-blinded, Multispecialty, Multi-institutional Randomized Control Trial. Annals of Surgery 2020; 272: 384–392. [DOI] [PubMed] [Google Scholar]
- [17].Hung AJ, Rambhatla S, Sanford DI, et al. Road to automating robotic suturing skills assessment: Battling mislabeling of the ground truth. Surgery 2022; 171: 915–919. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [18].Fard MJ, Ameri S, Darin Ellis R, Chinnam RB, Pandya AK, Klein MD. Automated robot‐assisted surgical skill evaluation: Predictive analytics approach. Robotics Computer Surgery; 14. Epub ahead of print February 2018. DOI: 10.1002/rcs.1850. [DOI] [PubMed] [Google Scholar]
- [19].Ghodoussipour S, Reddy SS, Ma R, Hwang DH, Nguyen J, Hung AJ. An Objective Assessment of Performance during Robotic Partial Nephrectomy: Validation and Correlation of Automated Performance Metrics with Intraoperative Outcomes. Journal of Urology 2021; 205: 1294–1302. [DOI] [PubMed] [Google Scholar]
- [20].Hung AJ, Chen J, Ghodoussipour S, et al. A deep-learning model using automated performance metrics and clinical features to predict urinary continence recovery after robot-assisted radical prostatectomy. BJU International 2019; 124: 487–495. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [21].Hung AJ, Chen J, Che Z, et al. Utilizing Machine Learning and Automated Performance Metrics to Evaluate Robot-Assisted Radical Prostatectomy Performance and Predict Outcomes. Journal of Endourology 2018; 32: 438–444. [DOI] [PubMed] [Google Scholar]
- [22].Cui Z, Ma R, Yang CH, et al. Capturing relationships between suturing sub-skills to improve automatic suturing assessment. npj Digit Med 2024; 7: 152. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [23].Howie EE, Ambler O, Gunn EGM, et al. Surgical Sabermetrics: A Scoping Review of Technology-enhanced Assessment of Nontechnical Skills in the Operating Room. Annals of Surgery 2024; 279: 973–984. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [24].Shafiei SB, Shadpour S, Mohler JL, et al. Prediction of Robotic Anastomosis Competency Evaluation (RACE) metrics during vesico-urethral anastomosis using electroencephalography, eye-tracking, and machine learning. Sci Rep 2024; 14: 14611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [25].Shafiei SB, Shadpour S, Mohler JL, Kauffman EC, Holden M, Gutierrez C. Classification of subtask types and skill levels in robot-assisted surgery using EEG, eye-tracking, and machine learning. Surg Endosc 2024; 38: 5137–5147. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [26].Deepika P, Deepesh KVV, Vadali PS, Rao M, Vazhayil V, Uppar AM. Computer Assisted Objective Assessment of Micro-Neurosurgical Skills From Intraoperative Videos. IEEE Open J Eng Med Biol 2023; 4: 11–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [27].Azari DP, Frasier LL, Quamme SRP, et al. Modeling Surgical Technical Skill Using Expert Assessment for Automated Computer Rating. Annals of Surgery 2019; 269: 574–581. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [28].Balal S, Smith P, Bader T, et al. Computer analysis of individual cataract surgery segments in the operating room. Eye 2019; 33: 313–319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [29].Wawrzynski JR, Smith P, Tang L, et al. Tracking camera control in endoscopic dacryocystorhinostomy surgery. Clinical Otolaryngology 2015; 40: 646–650. [DOI] [PubMed] [Google Scholar]
- [30].Lavanchy JL, Zindel J, Kirtac K, et al. Automation of surgical skill assessment using a three-stage machine learning algorithm. Sci Rep; 11. Epub ahead of print 4 March 2021. DOI: 10.1038/s41598-021-84295-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [31].Mirchi N, Bissonnette V, Yilmaz R, Ledwos N, Winkler-Schwartz A, Del Maestro RF. The Virtual Operative Assistant: An explainable artificial intelligence tool for simulation-based training in surgery and medicine. PLoS ONE 2020; 15: e0229596. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [32].Mascagni P, Vardazaryan A, Alapatt D, et al. Artificial Intelligence for Surgical Safety: Automatic Assessment of the Critical View of Safety in Laparoscopic Cholecystectomy Using Deep Learning. Annals of Surgery 2022; 275: 955–961. [DOI] [PubMed] [Google Scholar]
- [33].Igaki T, Kitaguchi D, Matsuzaki H, et al. Automatic Surgical Skill Assessment System Based on Concordance of Standardized Surgical Field Development Using Artificial Intelligence. JAMA Surg 2023; 158: e231131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [34].Zhao X, Takenaka S, Iuchi S, et al. Automatic recognition of surgical phase of robot-assisted radical prostatectomy based on artificial intelligence deep-learning model and its application in surgical skill evaluation: a joint study of 18 medical education centers. Surg Endosc. Epub ahead of print 10 July 2025. DOI: 10.1007/s00464-025-11967-z. [DOI] [PubMed] [Google Scholar]
- [35].Nakajima K, Kitaguchi D, Takenaka S, et al. Automated surgical skill assessment in colorectal surgery using a deep learning-based surgical phase recognition model. Surg Endosc 2024; 38: 6347–6355. [DOI] [PubMed] [Google Scholar]
- [36].Ma R, Ramaswamy A, Xu J, et al. Surgical gestures as a method to quantify surgical performance and predict patient outcomes. npj Digit Med 2022; 5: 187. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [37].Van Amsterdam B, Funke I, Edwards E, et al. Gesture Recognition in Robotic Surgery With Multimodal Attention. IEEE Trans Med Imaging 2022; 41: 1677–1687. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [38].Sirajudeen N, Boal M, Anastasiou D, et al. Deep learning prediction of error and skill in robotic prostatectomy suturing. Surg Endosc 2024; 38: 7663–7671. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [39].Sato K, Takenaka S, Kitaguchi D, et al. Objective surgical skill assessment based on automatic recognition of dissection and exposure times in robot-assisted radical prostatectomy. Langenbecks Arch Surg 2025; 410: 39. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [40].Ma R, Kiyasseh D, Laca JA, et al. Artificial Intelligence-Based Video Feedback to Improve Novice Performance on Robotic Suturing Skills: A Pilot Study. Journal of Endourology 2024; 38: 884–891. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [41].Cheng K, Wu S, Peng B, Wang X. An artificial intelligence enhanced coaching mode. International Journal of Surgery. Epub ahead of print 23 June 2025. DOI: 10.1097/JS9.0000000000002713. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [42].Kocielnik R, Yang CH, Ma R, et al. Human AI collaboration for unsupervised categorization of live surgical feedback. npj Digit Med 2024; 7: 372. [DOI] [PMC free article] [PubMed] [Google Scholar]
- [43].Kiyasseh D, Laca J, Haque TF, et al. A multi-institutional study using artificial intelligence to provide reliable and fair feedback to surgeons. Commun Med 2023; 3: 42. [DOI] [PMC free article] [PubMed] [Google Scholar]
