Table 4.
The evaluation results of core AI tools.
| AI tool category | Representative examples | Accuracy | Generalizability | Safety | Clinical integration | Regulatory status | References |
|---|---|---|---|---|---|---|---|
| Segmentation models | nnUNet, U-Net-based rotator cuff model | DSC: 0.86–0.95 (humerus/glenoid/rotator cuff) | Limited by training data (few obese patients) | No direct procedural risk; data privacy dependent on hospital systems | Compatible with ultrasound/MRI; requires radiologist collaboration | Not independently regulated (integrated into software) | Mu et al. (2021), Dai et al. (2024), Medina et al. (2021), and Alipour et al. (2024) |
| Navigation systems | ExactechGPS®, Joint VTS | Injection accuracy: 90 ~ 96.6% | Moderate (validated in RSA/TSA; limited in frozen shoulder) | Complication rate < 1%; real-time vibration alerts | Integrates with ultrasound/robots; 1–2 weeks learning curve | NMPA Class III/FDA 510(k)/CE Class IIb | Xu et al. (2025), Andriollo et al. (2024), and Kuratani et al. (2022) |
| Real-time learning platforms | Federated averaging-based closed-loop systems | Adaptive accuracy improvement: 5 ~ 10% after 100 cases# | High (multi-center data integration) | Edge computing protects privacy; model updates require validation | validationSeamless with intraoperative workflow; no additional operator burden | Regulatory gap (continuous learning not fully standardized) | Xu et al. (2025) and Schweihoff et al. (2021) |
| Specialized tools for subgroups | AI models for BMI > 35 patients | First-pass success rate: 85% (vs. 65% for conventional AI) | High (targets obese/large tear patients) | Reduces soft tissue injury risk by 20% | Requires high-resolution ultrasound; short learning curve | NMPA/FDA pending (pilot stage) | Huang et al. (2015) and Wu et al. (2025) |
The performance of real-time learning platforms (5 ~ 10% accuracy improvement after 100 cases) is inferred from federated learning technical characteristics, not direct clinical trial data.