Skip to main content
NPJ Digital Medicine logoLink to NPJ Digital Medicine
. 2026 Jun 4;9:684. doi: 10.1038/s41746-026-02855-4

Fully automated system predicts osteoporotic vertebral fracture across institutions using lumbar MRI paraspinal muscle signatures

Weicong Zhang 1,#, Yangjie Qin 1,#, Yixiu Hao 2,#, Weiying Liang 3, Junjie Lu 4, Weijia Zhu 5, Xiangwei Yuan 6, Haoyang Zhou 7, Yingnan Zhao 1, Qinghua Xie 6, Yu Liu 7, Didi Hu 8, Zhuodong Liang 9, Bao Feng 7,✉, Wansheng Long 1,✉
PMCID: PMC13554267  PMID: 42243528

Abstract

Paraspinal muscle (PM) degeneration is a crucial yet frequently overlooked risk factor for osteoporotic vertebral fractures (OVF). We developed PM Segmentation and Classification of OVF (PMSAC-OVF), a fully automated, multi-institutional system that segments lumbar PMs on MRI, extracts federated learning (FL) and radiomics features, and integrates them with clinical variables for OVF prediction. Leveraging a vision foundation model framework, the system enables privacy-preserving, cross-institutional training and lightweight local deployment. Data from 2,884 patients across five institutions (2014–2024) were analyzed. The automated segmentation module demonstrated expert-level accuracy (Dice coefficient: 0.952, Intersection over Union: 0.909) while reducing processing time to seconds. For prediction, FL and radiomics models yielded pooled AUCs of 0.827 (range: 0.819-0.861) and 0.803 (0.793-0.892), respectively. Trimodal models integrating radiomics signatures (RS), FL signatures (FLS), and clinical variables achieved a pooled AUC of 0.840 (0.822-0.916), significantly outperforming clinical-only models (AUC: 0.742, 0.641-0.778). SHapley Additive exPlanations identified RS, FLS, and bone mineral density as the top predictors, highlighting the complementary value of image-derived features. PMSAC-OVF provides a robust, interpretable, and scalable solution for OVF prediction in heterogeneous clinical settings, potentially facilitating early identification and personalized intervention for high-risk individuals.

Subject terms: Computational biology and bioinformatics, Diseases, Health care, Medical research

Introduction

Osteoporosis is a systemic skeletal disease characterized by decreased bone mass and the deterioration of bone microarchitecture, leading to compromised bone strength and increased susceptibility to fractures1,2. As the most common complication of osteoporosis, osteoporotic vertebral fractures (OVF) can induce severe consequences, including chronic pain, reduced height, impaired mobility, kyphosis, disability, diminished quality of life, and elevated mortality3. With the global population aging rapidly, the rising prevalence of both osteoporosis and OVF has imposed a substantial socioeconomic burden. Worldwide, approximately 8.9 million osteoporotic fractures occur annually4, with the United States alone reporting an estimated 2 million cases, incurring $17 billion in medical expenses each year2. In China, a nationwide epidemiological survey found OVF prevalence rates of 10.5% in men and 9.7% in women aged 40 years and older. Osteoporotic fracture-related costs in China are projected to reach $18 billion by 2035 and increase by an additional 23% by 20505. These data highlight the urgent need for improved OVF prediction to support early intervention and guide public health strategies.

Dual-energy X-ray absorptiometry (DXA)-derived bone mineral density (BMD) remains the gold standard for osteoporosis diagnosis, as defined by the WHO3,6–8. However, limited access to DXA scanners in primary care settings contributes to underdiagnosis and missed treatment opportunities6,9,10. Moreover, a substantial proportion of fragility fractures occur in individuals without osteoporotic BMD levels.

Consequently, current clinical strategies are shifting toward risk-based fracture prediction rather than relying solely on BMD11. Several studies have demonstrated that BMD-based models for predicting central fracture achieve areas under the receiver operating characteristic curve (AUC) values ranging from 0.60 to 0.86, with reduced predictive validity in younger populations12. These findings suggest that BMD alone is insufficient as the optimal predictor of fractures.

To enhance fracture prediction, several multifactorial tools have been developed, including the Fracture Risk Assessment Tool (FRAX), QFracture, and the Garvan Fracture Risk Calculator. These tools estimate future fracture risk using clinical and demographic factors—such as fall history, parental fracture, glucocorticoid use, and alcohol consumption—with or without BMD input13–15. Despite their widespread use, these tools exhibit only moderate discriminatory performance, with reported AUCs ranging from 0.60 to 0.8012. These limitations underscore the need for more accurate and individualized prediction approaches.

Artificial intelligence (AI)-driven models offer a promising avenue for improving fracture prediction. Advanced machine learning (ML) algorithms can automatically extract and analyze high-dimensional features from medical images, enabling more precise and personalized risk assessments16. Nevertheless, the clinical adoption of AI models remains constrained by critical validation challenges. Most medical AI systems require large and diverse datasets to ensure reliability16,17. Yet, substantial data heterogeneity (DH)—caused by variations in imaging devices, acquisition protocols, and image quality—poses significant obstacles to model robustness and generalizability across institutions18. Although several studies have demonstrated the feasibility of ML-based OVF prediction using spinal radiographs, CT, or DXA scans in small cohorts19, persistent DH-related validation challenges continue to hinder their broader clinical application.

Traditional models for predicting OVF have primarily focused on BMD and vertebral morphology. However, paraspinal muscles (PMs), which are anatomically and functionally integrated with the spine, have been largely overlooked in fracture risk assessment. Growing evidence highlights the biomechanical influence of PM degeneration on increasing fracture susceptibility20. Progressive muscle degeneration diminishes mechanochemical stimulation of bone, ultimately disrupting bone remodeling homeostasis and accelerating osteoporotic progression21,22. Furthermore, impaired muscular contractility leads to maladaptive load redistribution within the vertebrae, thereby intensifying focal stress and elevating fracture risk23. Notably, imaging biomarkers of PM degeneration—such as decreased muscle mass and increased fat infiltration—can be detected earlier than osteoporotic bone loss using routine MRI or CT scans24–26. In addition, a higher prevalence of sarcopenia has been reported in patients with OVF compared to those without27,28. Collectively, these findings support the potential utility of PM-derived features as valuable OVF predictors.

To address the pervasive DH challenges, we previously developed a Vision Foundation Model General Lightweight (VFMGL) framework, which has demonstrated strong robustness and generalizability across diverse medical image analysis tasks29. Building upon this state-of-the-art approach and the established pathophysiological association between PM degeneration and OVF, this study aimed to adapt the VFMGL framework to leverage robust PM-based imaging features for OVF prediction. Specifically, we constructed and validated an advanced multi-institutional system—PM Segmentation And Classification of OVF (PMSAC-OVF)—which fully automated the following computational pipeline: (1) 3D segmentation of PM regions of interest (ROIs); (2) extraction of general deep learning (DL) features via a vision foundation model (VFM), enabling efficient and lightweight deployment; (3) implementation of a privacy-preserving federated learning (FL) framework to facilitate collaborative, decentralized model training across institutions; and (4) multimodal integration of FL, radiomics, and clinical signatures to enhance predictive performance and model interpretability (Fig. 1).

Fig. 1. Overview of the PMSAC-OVF system.

Fig. 1

a Automated 3D segmentation of bilateral paraspinal muscles, including the psoas major and the erector spinae–multifidus complex. Volumetric segmentation was performed via nnU-Net, followed by slice-wise feature extraction and aggregation. b Construction of the VFMGL-based federated learning framework, incorporating HGKT and DDBL. c Radiomics analysis. d Clinical variable collection. e Multimodal feature integration for OVF prediction. BMD bone mineral density, BMI body mass index, DDBL data deduction at batch-level, ELM extreme learning machine, FLS federated learning signature, HGKT heterogeneous-model general knowledge transfer, KD knowledge distillation, MedSAM medical segment anything model, MLR multivariate logistic regression, Lasso least absolute shrinkage and selection operator, OVF osteoporotic vertebral fractures, PMSAC-OVF Paraspinal Muscle Segmentation And Classification of Osteoporotic Vertebral Fractures, ROI region of interest, RS radiomics signature, VFM vision foundation model, VFMGL vision foundation model general lightweight.

Results

Study population and clinical characteristics

Clinical data were collected from five medical institutions, comprising a total of 2,884 eligible patients who had undergone both axial lumbar T2-weighted (T2W) MRI and DXA-derived BMD assessments (Fig. 2). The demographic and clinical characteristics of patients from each institution are summarized in Table 1. OVF and non-OVF cases were relatively evenly distributed across the five institutions. Patients with OVF consistently exhibited significantly lower BMD values than those without. In Institution A, which contributed the largest subset of data, OVF patients were also significantly older and had lower body mass index (BMI) values. Statistically significant sex differences were not consistently observed across all institutional datasets.

Fig. 2. Flowchart of patient selection.

Fig. 2

The diagram illustrates the inclusion and exclusion process across five participating institutions. BMD bone mineral density, DXA dual-energy X-ray absorptiometry, n number of patients, OVF osteoporotic vertebral fractures, T2W, T2-weighted.

Table 1.

Demographic and clinical characteristics of the study population across five institutions

Institution Set Type Age P value Sex P value BMI P value BMD P value
Male Female
A (1880) Train-set (1316) OVF (647) 70.72 ± 9.66 <0.001 145 502 0.059 23.03 ± 3.57 <0.001 0.79 ± 0.16 <0.001
NOVF (669) 66.83 ± 8.34 180 489 24.28 ± 3.76 0.97 ± 0.21
Test-set (564) OVF (271) 69.79 ± 8.97 <0.001 66 205 0.599 22.65 ± 3.78 <0.001 0.79 ± 0.16 <0.001
NOVF (293) 66.62 ± 9.44 77 216 24.00 ± 3.53 0.98 ± 0.21
B (120) Train-set (84) OVF (44) 71.84 ± 10.86 0.729 12 32 0.032 22.18 ± 3.80 0.251 0.77 ± 0.13 <0.001
NOVF (40) 71.05 ± 9.88 20 20 23.53 ± 4.04 0.92 ± 0.18
Test-set (36) OVF (20) 73.45 ± 10.01 0.007 7 13 0.813 22.30 ± 3.21 0.648 0.84 ± 0.22 0.049
NOVF (16) 65.25 ± 6.37 5 11 22.98 ± 3.38 0.96 ± 0.19
C (373) Train-set (261) OVF (151) 71.11 ± 9.90 0.001 26 125 0.006 22.68 ± 3.44 0.006 0.77 ± 0.18 <0.001
NOVF (110) 66.85 ± 9.93 35 75 23.78 ± 3.29 0.91 ± 0.17
Test-set (112) OVF (54) 70.02 ± 11.04 0.232 16 38 0.721 22.81 ± 3.33 0.069 0.81 ± 0.19 0.01
NOVF (58) 67.71 ± 9.30 19 39 24.17 ± 3.69 0.93 ± 0.25
D (256) Train-set (179) OVF (96) 72.40 ± 10.49 <0.001 33 63 0.364 23.61 ± 4.79 0.508 0.85 ± 0.20 <0.001
NOVF (83) 64.98 ± 10.41 34 49 23.88 ± 3.79 1.03 ± 0.20
Test-set (77) OVF (41) 71.78 ± 10.14 0.090 9 32 0.977 23.00 ± 4.01 0.012 0.82 ± 0.19 <0.001
NOVF (36) 67.50 ± 11.75 8 28 24.96 ± 3.59 1.05 ± 0.25
E (255) Train-set (178) OVF (76) 67.05 ± 11.51 0.081 19 57 0.485 23.41 ± 3.28 0.146 0.85 ± 0.16 <0.001
NOVF (102) 64.22 ± 9.99 21 81 24.17 ± 3.10 1.03 ± 0.21
Test-set (77) OVF (35) 65.14 ± 9.26 0.605 8 27 0.880 23.38 ± 3.42 0.951 0.83 ± 0.19 0.017
NOVF (42) 66.28 ± 9.88 9 33 23.41 ± 3.38 0.92 ± 0.18

Age, BMI, and BMD values are shown as mean ± standard deviation. Numbers in parentheses represent the number of patients.

BMD bone mineral density, BMI body mass index, NOVF non-osteoporotic vertebral fracture, OVF osteoporotic vertebral fracture.

Automated segmentation performance and efficiency

To ensure the reliability of the upstream PM segmentation, we quantitatively evaluated the 3D nnU-Net model on an independent test set of 53 participants from Institution A. As detailed in Supplementary Table 1, the model demonstrated high anatomical fidelity relative to the expert reference standard, yielding a Dice similarity coefficient (DSC) of 0.952 and an intersection over union (IoU) of 0.909. Regarding computational efficiency, the fully automated nnU-Net model reduced the segmentation time to mere seconds per case. This represents a substantial improvement over manual delineation and semi-automated interaction (MedSAM), which required approximately 33 minutes and 14 minutes per participant, respectively.

Performance of radiomics, FL, and clinical models

Using the largest data subset from Institution A, we first evaluated the test performance of radiomics and FL models trained on different combinations of PM features. As shown in Supplementary Tables 2 and 3, incorporating all four PMs yielded the highest AUC values (radiomics: 0.809; FL: 0.819). DeLong’s tests confirmed that this comprehensive four-muscle signature significantly outperformed models based on muscle subgroups (radiomics: p = 0.005 vs. bilateral erector spinae–multifidus complexes, p = 0.019 vs. bilateral psoas majors; FL: p < 0.001 and p = 0.003, respectively). These findings supported the selection of the four-muscle combination for all subsequent cross-model comparisons.

Based on this optimal configuration, the radiomics models achieved test AUCs of 0.809, 0.809, 0.892, 0.794, and 0.793 across Institutions A through E, respectively (pooled AUC: 0.803). Under the FL framework, we evaluated the predictive performance of six candidate classifiers (Fig. 3). Local models employing the extreme learning machine (ELM) consistently demonstrated the best performance, with corresponding test AUCs of 0.819, 0.856, 0.845, 0.856, and 0.861 (pooled AUC: 0.827). Notably, these independent test results closely aligned with the internal 5-fold cross-validation benchmarks (mean AUCs: 0.795 ± 0.004, 0.850 ± 0.007, 0.840 ± 0.022, 0.841 ± 0.017, and 0.825 ± 0.013; see Supplementary Table 11 and Supplementary Fig. 2), confirming the robustness and stability of the FL model configuration. By contrast, the clinical baseline models exhibited inferior discrimination, with AUCs ranging from 0.641 to 0.778 across institutions (pooled AUC: 0.742).

Fig. 3. Performance comparison of six different classifiers within the FL framework.

Fig. 3

ROC curves are shown for the independent test set of each institution (A–E). The ELM classifier consistently demonstrated superior performance compared to other algorithms, justifying its selection for the final FL classifier. AUC areas under the ROC curve, DNN deep neural network, ELM extreme learning machine, FL federated learning, FPR false positive rate, LR logistic regression, RF random forest, ROC receiver operating characteristic, SVM support vector machine, TPR true positive rate, XGBoost extreme gradient boosting.

Performance comparisons of multimodal models

To assess the benefit of multimodal feature fusion, we evaluated six distinct modeling strategies (Figs. 4 and 5). Models combining clinical and image-derived features consistently outperformed those relying solely on clinical data. Notably, the trimodal models—integrating radiomics signatures (RS), federated learning signatures (FLS), and clinical variables—demonstrated a superior overall performance balance across all institutions, with AUCs ranging from 0.822 to 0.916. Comprehensive performance metrics comparing the trimodal model against the clinical baseline are detailed in Table 2. Statistical analyses using DeLong’s tests and reclassification metrics—integrated discrimination improvement (IDI) and net reclassification improvement (NRI)—confirmed that the trimodal strategy significantly surpassed all unimodal models and the bimodal clinical+RS model across all datasets (see Table 3 and Supplementary Tables 4–9).

Fig. 4. ROC curves comparing the performance of six modeling approaches across five institutions and the pooled test set.

Fig. 4

AUC areas under the ROC curve, CLM clinical model, CL + FLS model integrating clinical variables and federated learning signature, CL + RS model integrating clinical variables and radiomics signature, CL + RS + FLS model integrating clinical variables, radiomics signature, and federated learning signature, FLM federated learning model, FPR false positive rate, RAM radiomics model, ROC receiver operating characteristic, TPR true positive rate.

Fig. 5. Radar charts illustrating the comprehensive performance of six modeling approaches across five institutions and the pooled test set.

Fig. 5

The trimodal model (CL + RS + FLS) consistently encompassed the largest area, indicating superior overall performance balance compared to other models. AUC areas under the receiver operating characteristic curve, CLM clinical model, CL + FLS model integrating clinical variables and federated learning signature, CL + RS model integrating clinical variables and radiomics signature, CL + RS + FLS model integrating clinical variables, radiomics signature, and federated learning signature, FLM federated learning model, NPV negative predictive value, PPV positive predictive value, RAM radiomics model.

Table 2.

Performance comparison of the trimodal model versus the clinical model across five institutions and the pooled test set

Institution Model AUC Accuracy Sensitivity Specificity PPV NPV F1 P value
A CLM

0.778

(0.739-0.817)

0.699

(0.660-0.736)

0.738

(0.687-0.788)

0.662

(0.607-0.718)

0.669

(0.614-0.723)

0.732

(0.682-0.785)

0.702

(0.660-0.742)

0.048
CL + RS + FLS

0.822

(0.786-0.854)

0.762

(0.725-0.796)

0.753

(0.703-0.800)

0.771

(0.722-0.814)

0.753

(0.701-0.798)

0.771

(0.722-0.819)

0.753

(0.710-0.788)

B CLM

0.694

(0.515-0.855)

0.667

(0.500-0.806)

0.500

(0.292-0.727)

0.875

(0.696-1.000)

0.833

(0.611-1.000)

0.583

(0.375-0.783)

0.625

(0.413-0.813)

0.053
CL + RS + FLS

0.859

(0.721-0.959)

0.750

(0.611-0.889)

0.700

(0.500-0.889)

0.812

(0.615-1.000)

0.824

(0.615-1.000)

0.684

(0.467-0.882)

0.757

(0.579-0.889)

C CLM

0.654

(0.549-0.755)

0.607

(0.518-0.696)

0.407

(0.286-0.539)

0.793

(0.683-0.891)

0.647

(0.472-0.805)

0.590

(0.481-0.700)

0.500

(0.364-0.611)

<0.001
CL + RS + FLS

0.916

(0.859-0.963)

0.804

(0.732-0.875)

0.722

(0.596-0.830)

0.879

(0.789-0.962)

0.848

(0.738-0.950)

0.773

(0.667-0.864)

0.780

(0.681-0.862)

D CLM

0.724

(0.600-0.831)

0.610

(0.506-0.727)

0.756

(0.619-0.875)

0.444

(0.282-0.606)

0.608

(0.471-0.741)

0.615

(0.421-0.800)

0.674

(0.554-0.776)

0.009
CL + RS + FLS

0.874

(0.792-0.945)

0.792

(0.701-0.870)

0.805

(0.676-0.922)

0.778

(0.641-0.906)

0.805

(0.667-0.927)

0.778

(0.625-0.903)

0.805

(0.710-0.889)

E CLM

0.641

(0.509-0.760)

0.519

(0.403-0.636)

0.829

(0.700-0.944)

0.262

(0.133-0.390)

0.483

(0.352-0.607)

0.647

(0.400-0.867)

0.611

(0.483-0.718)

<0.001
CL + RS + FLS

0.892

(0.808-0.956)

0.753

(0.649-0.844)

0.600

(0.425-0.769)

0.881

(0.775-0.974)

0.808

(0.650-0.957)

0.725

(0.604-0.846)

0.689

(0.542-0.812)

Pooled set CLM 0.742 (0.712-0.775) 0.669 (0.636-0.699) 0.715 (0.671-0.756) 0.625 (0.580-0.667) 0.643 (0.600-0.685) 0.698 (0.653-0.746) 0.677 (0.640-0.712) <0.001
CL + RS + FLS 0.840 (0.812-0.867) 0.766 (0.737-0.792) 0.727 (0.681-0.769) 0.802 (0.767-0.838) 0.777 (0.736-0.817) 0.756 (0.718-0.794) 0.751 (0.716-0.783)

Data are presented as value (95% confidence interval). P values indicate the statistical significance of the difference in AUC between the trimodal model and the clinical model (DeLong’s test).

AUC areas under the receiver operating characteristic curve, CLM clinical model, CL + RS + FLS trimodal model integrating clinical variables, radiomics signature, and federated learning signature, NPV negative predictive value, PPV positive predictive value.

Table 3.

Reclassification improvement and discrimination capability of the trimodal model (CL + RS + FLS) across five institutions and the pooled test set

Institution Metric vs. RAM vs. FLM vs. CLM vs. CL + RS vs. CL + FLS
A NRI 0.590 (p < 0.001) 0.457 (p < 0.001) 0.357 (p < 0.001) 0.036 (p = 0.018) 0.008 (p = 0.872)
IDI 0.319 (p < 0.001) 0.367 (p < 0.001) 0.214 (p < 0.001) 0.010 (p < 0.001) 0.002 (p = 0.947)
B NRI 0.913 (p < 0.001) 0.563 (p < 0.001) 0.438 (p = 0.049) 0.238 (p = 0.033) 0.015 (p = 0.468)
IDI 0.370 (p < 0.001) 0.362 (p < 0.001) 0.252 (p = 0.006) 0.079 (p = 0.013) 0.116 (p = 0.214)
C NRI 0.118 (p = 0.053) 0.573 (p < 0.001) 0.823 (p < 0.001) 0.054 (p = 0.408) 0.167 (p = 0.119)
IDI 0.091 (p < 0.001) 0.499 (p < 0.001) 0.451 (p < 0.001) 0.038 (p = 0.073) 0.166 (p = 0.001)
D NRI 0.369 (p = 0.014) 0.354 (p = 0.001) 0.313 (p = 0.045) 0.028 (p = 0.809) 0.045 (p = 0.717)
IDI 0.351 (p < 0.001) 0.381 (p < 0.001) 0.224 (p = 0.002) 0.041 (p = 0.290) 0.058 (p = 0.210)
E NRI 1.186 (p < 0.001) 0.471 (p < 0.001) 0.691 (p < 0.001) 0.233 (p = 0.009) 0.071 (p = 0.591)
IDI 0.469 (p < 0.001) 0.434 (p < 0.001) 0.383 (p < 0.001) 0.104 (p < 0.001) 0.081 (p = 0.198)
Pooled set NRI 0.579 (p < 0.001) 0.470 (p < 0.001) 0.450 (p < 0.001) 0.062 (p = 0.001) 0.048 (p = 0.229)
IDI 0.309 (p < 0.001) 0.393 (p < 0.001) 0.264 (p < 0.001) 0.027 (p < 0.001) 0.042 (p = 0.044)

Statistical significance was determined using the two-tailed Z-test. NRI cutoffs were set at 0.2 and 0.6.

CLM clinical model, CL + FLS model integrating clinical variables and federated learning signature, CL + RS model integrating clinical variables and radiomics signature, CL + RS + FLS model integrating clinical variables, radiomics signature, and federated learning signature, FLM federated learning model, IDI integrated discrimination improvement, NRI net reclassification improvement, RAM radiomics model.

While differences between the trimodal model and the bimodal clinical+FLS model were not statistically significant in terms of AUC (pooled set: 0.840 vs. 0.847), NRI, or IDI, both approaches exhibited robust performance. Decision curve analysis (DCA, Fig. 6) further corroborated these findings, showing that multimodal models consistently generated greater net clinical benefit than unimodal strategies. Specifically, the trimodal model yielded a net benefit comparable to that of the clinical+FLS model across the majority of clinically reasonable threshold probabilities, particularly in the global pooled test set, highlighting the general superiority of the multimodal integration approach.

Fig. 6. Decision curve analysis assessing the clinical utility of six modeling approaches across five institutions and the pooled test set.

Fig. 6

The multimodal model consistently provided greater net benefit across the majority of clinically reasonable threshold probabilities, indicating superior clinical utility compared to “treat-all”, “treat-none”, and other unimodal strategies. CLM clinical model, CL + FLS model integrating clinical variables and federated learning signature, CL + RS model integrating clinical variables and radiomics signature, CL + RS + FLS model integrating clinical variables, radiomics signature, and federated learning signature, FLM federated learning model, RAM radiomics model.

Feature contribution analysis

To elucidate the decision-making process of the trimodal model, the contributions of the RS, FLS, and four clinical variables—age, sex, BMI, and lumbar BMD—to the multivariate logistic regression (MLR) classifier were visualized using SHapley Additive exPlanations (SHAP) beeswarm plots (Fig. 7). Across all participating institutions, RS consistently emerged as the most influential predictor, highlighting the critical role of textural muscle features. In Institutions B, C, and E, RS, FLS, and BMD robustly ranked as the top three contributors in descending order of importance. The SHAP value distributions revealed clear directional associations: higher RS and FLS values were positively correlated with an increased probability of OVF prediction, whereas higher lumbar BMD showed a strong negative association, consistent with its protective physiological role.

Fig. 7. Model interpretability analysis using SHAP beeswarm plots across five institutional test sets.

Fig. 7

The plots visualize the impact of each feature on the trimodal (CL + RS + FLS) model’s output. Features are ranked by their mean absolute SHAP values (global importance). The RS and FLS emerged as top predictors across the majority of institutions, often surpassing traditional clinical markers like BMD, highlighting the incremental value of image-based biomarkers. BMD bone mineral density, BMI body mass index, CL + RS + FLS model integrating clinical variables, radiomics signature, and federated learning signature, FLS federated learning signature, RS radiomics signature, SHAP SHapley Additive exPlanations.

Discussion

AI-based analysis of routine medical imaging presents a transformative strategy for OVF prediction, enabling precise assessment of osteoporosis severity, timely initiation of personalized interventions, and effective reduction of the fragility fracture burden2,3. Nevertheless, reliable OVF prediction remains a significant clinical challenge. Recognizing the established pathophysiological association between PM degeneration and OVF development20, we proposed PMSAC-OVF, a novel multi-institutional system that fully automates OVF prediction based on PM-derived features from lumbar MRI. By synergizing DL-based segmentation with the VFMGL framework, our system addresses critical barriers in DH and manual workload, enhancing predictive efficiency and supporting scalable deployment in high-throughput clinical settings.

A pivotal component of this system is its ability to circumvent the bottleneck of manual annotation. Reliable and efficient PM segmentation is a prerequisite for clinical adoption but has historically been hindered by the prohibitive labor costs of manual delineation—typically requiring approximately 33 minutes per patient. Even semi-automated interactive tools often necessitate substantial user input, limiting their utility in large-scale studies. In contrast, our fully automated 3D nnU-Net model completes segmentation in mere seconds while maintaining anatomical fidelity comparable to human experts (DSC: 0.952, IoU: 0.909). By serving as a robust upstream component for subsequent feature extraction, this module establishes a seamless, end-to-end automated workflow. This capability effectively eliminates the dependency on manual intervention, streamlining large-scale quantitative analyses and facilitating the routine integration of PM assessment into clinical practice.

Beyond automation, data accessibility remains a critical barrier. Stringent data privacy regulations inherently restrict inter-institutional data exchange, often confining models to fragmented, single-institution datasets that fail to capture the heterogeneity of real-world clinical scenarios17,30,31. Such data siloing exacerbates the risk of overfitting and compromises model generalizability during external validation32. In particular, multi-institutional DH—stemming from variability in scanner vendors, acquisition protocols, image quality, and patient demographics—considerably impairs the robustness of AI models deployed across diverse healthcare settings18,33.

To address these challenges, VFMs provide transferable, high-level feature representations learned from large-scale natural image datasets and have demonstrated robust performance across a broad spectrum of visual tasks34,35. However, the significant domain shift between natural and medical images typically necessitates resource-intensive retraining on annotated medical data36–38, which limits the practical implementation of VFMs in clinical workflows. To overcome this limitation, our framework combined the frozen VFM—specifically DINOv235—backbone with lightweight local models, allowing efficient adaptation of general visual knowledge to institution-specific medical imaging distributions without the computational burden of full-model fine-tuning. This hybrid architecture enhances robustness under varied imaging conditions and facilitates real-world deployment in heterogeneous clinical environments. Moreover, the incorporation of this VFM-based feature extraction within an FL framework enables collaborative model development across institutions without sharing raw patient data, thus ensuring strict privacy compliance. Within this framework, low-heterogeneity samples are identified locally to guide model convergence toward shared, scanner-invariant feature representations, thereby improving generalizability in multi-institutional settings.

As essential components of the spinal support system, PMs play a crucial role in maintaining spinal structural integrity. The primary PMs include the ventral psoas major and the dorsal erector spinae–multifidus complex, which together fulfill complementary biomechanical functions. Specifically, the psoas major stabilizes the anterior spinal column, while the posterior muscles maintain segmental stability and postural alignment. Degeneration of either group can disrupt normal spinal load distribution, accelerate osteoporosis progression, and increase fracture risk23,39,40. Therefore, incorporating features from both anterior and posterior PMs provides a more comprehensive and physiologically informed basis for OVF prediction. This was statistically validated by our results: radiomics and FL models integrating all four PMs achieved significantly higher predictive performance compared to those relying on partial muscle subsets, confirming that holistic muscle profiling is superior to region-specific analysis.

PM degeneration—characterized by muscle atrophy and fatty infiltration—has been recognized as a significant predictor of OVF onset and progression20,24,25. However, conventional macroscopic metrics (e.g., cross-sectional area) often fail to capture subtle, early-stage deterioration in muscle quantity and quality. In contrast, our study leveraged advanced image features—specifically, handcrafted radiomics textures and abstract DL representations—to detect these micro-structural changes. In the global pooled test set, the radiomics models achieved robust discrimination (AUC: 0.803, accuracy: 0.727), confirming the value of texture analysis. Notably, the FL models exhibited superior performance (AUC: 0.827, accuracy: 0.742), indicating that data-driven DL features extracted via the VFM may capture more complex, non-linear patterns of PM degeneration than predefined radiomics descriptors. Crucially, both approaches generated personalized probabilistic risk scores. These scores can facilitate the early identification of high-risk individuals—potentially before significant bone mass loss occurs—guiding tailored interventions such as PM-strengthening resistance training, nutritional optimization, or pharmacotherapy to mitigate muscle degeneration and preserve spinal stability.

Multimodal feature fusion—uniting prior medical knowledge with information-rich imaging signatures—has been widely shown to enhance predictive performance across diverse diagnostic tasks41–43. Consistent with these findings, our trimodal models significantly outperformed all unimodal models across all institutions. Although the trimodal strategy yielded performance metrics and net clinical benefits comparable to the bimodal clinical+FLS model, the inclusion of RS offers distinct clinical advantages. Unlike abstract DL representations, radiomics features provide quantifiable, human-readable descriptors of muscle shape and texture, offering a layer of biological interpretability. SHAP analysis corroborated this synergy, identifying both RS and FLS as top-tier predictors alongside lumbar BMD, while age, sex, and BMI contributed minimally. This suggests that RS captures complementary structural information, enhancing model robustness, particularly in ambiguous cases where DL features alone might be insufficient. Furthermore, by employing a late fusion strategy, we mitigated the risk of overfitting high-dimensional imaging features, ensuring that each modality—PM status (RS and FLS), central bone quality (BMD), and demographic factors—contributes independently to a comprehensive OVF prediction.

This study has several limitations. First, the dataset was obtained from only five institutions. Expanding the cohort to encompass broader geographic and demographic diversity may improve model generalizability. Second, certain clinically relevant variables—such as menopausal status, use of bone metabolism-related medications, and bone turnover markers—were unavailable, which could enhance both predictive performance and clinical interpretability. Third, integrating bone structural features from spinal radiographs or CT scans may enrich the feature space and further optimize prediction accuracy. Fourth, a failure analysis based on confusion matrices (Supplementary Fig. 3) suggests that severe PM fatty infiltration in patients with preserved bone integrity may lead to false positives, warranting further study into these discordant cases. Finally, future research will investigate the relationship between PM features and OVF severity, and explore personalization strategies—such as federated meta-learning44—to strengthen model robustness, particularly in data-constrained institutional settings.

In conclusion, this study demonstrates the effectiveness of the PMSAC-OVF system for fully automated OVF prediction by integrating 3D PM segmentation, multimodal feature extraction and fusion, and individualized risk estimation. The system enhances model robustness and generalizability while preserving data privacy, supporting practical deployment in heterogeneous, multi-institutional clinical environments. Its application may enable the early identification of high-risk individuals and inform tailored preventive or therapeutic strategies, thereby mitigating both the clinical impact and societal burden of OVF.

Methods

Data source

This multi-institutional retrospective study received ethical approval (No. 2024-267 A) from the Ethics Review Committees of five participating medical institutions: Jiangmen Central Hospital (Institution A), Maoming People’s Hospital (Institution B), Baishi Orthopedic Hospital (Institution C), Guangzhou First People’s Hospital (Institution D), and the Fifth Affiliated Hospital, Sun Yat-sen University (Institution E). The requirement for informed consent was formally waived due to the retrospective nature. Prior to processing, all patient data were strictly de-identified in compliance with ethical regulations. Direct identifiers (e.g., names, dates) were replaced with unique study codes, and DICOM metadata were stripped of protected health information, retaining only essential acquisition parameters. To ensure methodological transparency and reproducibility, the study design and reporting strictly adhered to the MEthodological TRansparency Index for Cognitive Systems (METRICS)45 and the Checklist for Artificial Intelligence in Medical Imaging (CLAIM) 2024 guidelines46. Completed checklists detailing compliance are provided in the Supplementary material.

Eligible participants met the following inclusion criteria: (1) enrollment between 2014 and 2024, and age over 40 years, targeting the population at risk for age-related OVF; (2) availability of both lumbar MRI and DXA-derived BMD examinations; and (3) a maximum interval of six months between the MRI and DXA scans. Exclusion criteria were as follows: (1) absence of axial lumbar T2W MRI sequence; (2) secondary osteoporosis attributed to endocrine or metabolic disorders; (3) OVF resulting from high-energy trauma, malignancy, or infection; (4) MRI performed more than two weeks after OVF onset or a history of prior OVF; (5) structural spinal abnormalities, such as spondylolisthesis, scoliosis, or kyphosis; (6) previous spinal surgery, including internal fixation or vertebroplasty, to avoid magnetic susceptibility artifacts and confounding iatrogenic muscle changes; (7) incomplete clinical data (e.g., missing weight or height); and (8) inadequate image quality due to significant artifacts (e.g., motion ghosting, susceptibility artifacts from metal, or severe chemical shift) that obscured PM delineation.

To ensure the clinical translatability of the developed models, all MRI and DXA data were acquired as part of routine clinical workflows for lumbar spine assessment. For patients with multiple MRI and DXA records, only the earliest qualifying pair was analyzed. Detailed image acquisition parameters (e.g., scanner vendor, field strength, slice thickness) for all participating institutions are summarized in Supplementary Table 10, highlighting the DH inherent in this multi-institutional cohort. Each institutional dataset was randomly partitioned at the patient level, with 70% allocated for model training and 30% reserved for independent hold-out testing. To emulate a real-world, multi-institutional deployment scenario, models were trained separately within each institution using only local data, without any inter-institutional data exchange.

Image annotation, automated segmentation, and preprocessing

Spine MRI is the clinical reference standard for diagnosing first-episode acute OVF2,3. Two board-certified musculoskeletal radiologists, each with five years of experience in spinal imaging interpretation, independently classified all cases as either OVF or non-OVF based on thoracic and lumbar MRI scans. Any inter-observer discrepancies were resolved to reach consensus after review by a senior radiologist with more than ten years of subspecialty experience.

To automate PM segmentation, we developed and implemented a customized segmentation platform based on the Medical Segment Anything Model (MedSAM) framework34. This platform was applied to axial T2W lumbar MRI scans from 100 participants, stored in Digital Imaging and Communications in Medicine (DICOM) format. The segmentation process comprised four sequential phases: (1) Expert annotation: Experienced radiologists manually delineated bilateral PMs (the psoas major and the erector spinae–multifidus complex) at lumbar levels L1-L5 using constrained bounding boxes. Due to the difficulty in distinguishing the erector spinae from the multifidus on MRI, these muscles were treated as a single functional unit in this study47,48. (2) Computational segmentation: The MedSAM-based algorithm generated preliminary masks for the target muscle regions. (3) ROI refinement: Senior musculoskeletal radiologists reviewed and corrected the segmentation contours to ensure anatomical precision. These final refined ROIs served as the reference standard for model development. (4) Segmentation model training and evaluation: The annotated dataset of 100 cases was randomly partitioned into a training set (n = 47) and an independent test set (n = 53). We developed a 3D nnU-Net model on the training set for automated segmentation of the four PMs49, implementing a 5-fold cross-validation strategy to evaluate internal stability. Final segmentation performance was assessed on the independent test set. Crucially, the trained nnU-Net model inherently localizes the target L1-L5 PMs based on learned anatomical context. Structures outside this specific range are automatically excluded, thereby robustly accommodating variable MRI scanning fields of view without requiring manual cropping or additional localization algorithms. The complete segmentation workflow is illustrated in Fig. 1a.

Following fully automated 3D PM segmentation, feature extraction was performed. To leverage both volumetric context and high-resolution 2D features, all valid axial slices within the segmented volume were utilized. Each 2D axial slice was resampled to a standardized in-plane resolution of 224×224 pixels. The extracted slice-level deep features were subsequently aggregated to construct a comprehensive case-level representation for each participant.

Federated learning framework construction

To address cross-institutional DH in OVF prediction, we developed an FL framework that integrated lightweight feature extraction with privacy-preserving decentralized model training29. The backbone architecture was based on the VFMGL framework, which combined a pretrained large-scale VFM (DINOv2) with a deployable lightweight network (ResNet18). DINOv2, a self-supervised vision transformer trained on diverse image collections, was selected for its superior transferability across downstream medical tasks35.

The FL framework was underpinned by two synergistic modules: (1) heterogeneous-model general knowledge transfer (HGKT), and (2) data deduction at batch level (DDBL). To mitigate the risk of negative transfer caused by domain discrepancies between natural and medical images, HGKT dynamically selected and transferred only task-relevant features from the frozen VFM to the local ResNet18 models via adaptive layer matching and weighted knowledge transfer. This ensured local networks inherited robust visual representations without the computational burden of full VFM deployment. DDBL enhanced generalizability by identifying low-heterogeneity samples within each institution that aligned with global inference outputs. These samples guided local models to learn shared, institution-invariant feature representations through knowledge distillation, improving cross-institutional consistency while preserving data privacy. (Fig. 1b)

The FL pipeline leveraged transfer learning for robust initialization. The local ResNet18 backbones were initialized with ImageNet-pretrained weights, while the DINOv2 (ViT-B/14) encoder utilized its official pretrained checkpoint. Any newly appended task-specific layers were randomly initialized using the default PyTorch scheme. The entire network was then fine-tuned end-to-end on the institution-specific training data.

Cross-institutional training was coordinated via the federated averaging (FedAvg) algorithm50. In each of the ten global communication rounds, participating institutions independently trained their local models on private datasets. Locally updated weights were transmitted to a central server, aggregated using sample-size-weighted averaging, and redistributed. Local training employed stochastic gradient descent (SGD) with a cross-entropy loss function. Hyperparameters were set to a learning rate of 0.0001, momentum of 0.9, weight decay of 0.0001, and a batch size of 25. To ensure configuration stability and rule out overfitting, a 5-fold cross-validation was performed exclusively within the training cohort. This internal validation step rigorously verified hyperparameter stability prior to final model training on the complete training set. The generalizability of each local model was evaluated on its respective independent hold-out test set.

To facilitate consistent multimodal integration, we employed ELM classifiers to generate institution-specific probabilistic feature representations. The output class-wise probabilities were defined as the FLS. These compact signatures mitigated issues such as gradient instability and local optima entrapment, providing an optimal, privacy-preserving input for downstream multimodal fusion.

Radiomics model construction

High-throughput quantitative radiomics features were extracted from lumbar MRI scans using the open-source PyRadiomics package in Python (via the SimpleITK library). For each patient, features were independently derived from four distinct volumes of interest (VOIs): the bilateral psoas major and the bilateral erector spinae–multifidus complexes.

To ensure inter-dataset consistency and reproducibility, multiple preprocessing steps were implemented. First, voxel spacing was standardized to 0.5×0.5×8.0 mm using B-spline interpolation, thereby preserving anatomical fidelity along the slice axis. Second, image intensities were normalized via Z-score transformation to standardize the dynamic range. A fixed gray-level bin width of 25 was applied for discretization, a setting robust against intensity variations.

A total of 1,223 radiomics features were extracted per muscle VOI from both the original and filtered images. Image filters included wavelet transformations and Laplacian of Gaussian (LoG) with sigma values of 1.0, 2.0, 3.0, and 5.0. The extracted feature classes encompassed first-order statistics, 2D/3D shape descriptors, and higher-order texture features derived from the gray-level co-occurrence matrix (GLCM), gray-level run-length matrix (GLRLM), gray-level size-zone matrix (GLSZM), and gray-level dependence matrix (GLDM).

Feature selection was conducted in three sequential steps to mitigate overfitting and identify the most discriminative predictors: (1) a Mann-Whitney U test was initially performed to eliminate non-discriminative features (p ≥ 0.05);(2) Pearson correlation analysis was conducted to remove redundant features with correlation coefficients |r | > 0.9; and (3) the least absolute shrinkage and selection operator (LASSO) regression was employed to identify the optimal feature subset, with the regularization parameter (λ) optimized via 10-fold cross-validation. The specific radiomics features retained for each muscle-specific model, along with their respective coefficients, are comprehensively documented in Supplementary Tables 12, 14, 16, 18, and 20.

Following feature selection, a two-stage stacking ensemble modeling strategy was employed to construct the final predictive signature (Supplementary Fig. 1). First, individual XGBoost classifiers were trained independently for each of the four muscles using their respective selected features. Model hyperparameters were tuned using the Tree-structured Parzen Estimator (TPE) algorithm within the Optuna optimization framework. Second, the probabilistic risk scores generated by the four muscle-specific XGBoost classifiers were combined using MLR to produce the final ensemble prediction. The optimal ensemble weights assigned to each muscle group are detailed in Supplementary Tables 13, 15, 17, 19, and 21. The resulting integrated probabilistic output was defined as the RS.

Clinical modeling, multimodal integration, and interpretability analysis

The clinical models were developed in two sequential steps. First, variable selection was performed. Univariate analyses—including the independent samples t‑test, Mann‑Whitney U test, and Pearson chi‑square test—were applied to four clinical variables (age, sex, BMI, and lumbar BMD; Fig. 1d) to identify statistically significant predictors (p < 0.05). Variables meeting this criterion were subsequently evaluated for multicollinearity using the variance inflation factor (VIF), and those exhibiting high collinearity were excluded. Second, an MLR model was constructed using the remaining variables, with stepwise backward elimination employed to optimize model performance.

For multimodal integration, the RS, FLS, and clinical variables were regarded as distinct yet complementary feature domains. These cross-modal signatures were combined using a late-fusion MLR classifier to assess their joint predictive capability. To provide transparency and interpret the model’s decision-making process, SHAP values were computed, quantifying the relative contributions of RS, FLS, and each clinical variable to the final OVF prediction51. The modeling workflow is summarized in Fig. 1e.

Statistical analysis

The sample size was determined based on the maximal availability of eligible retrospective data within the study period. Demographic characteristics were described as mean ± standard deviation for continuous variables and as frequencies with corresponding percentages for categorical variables. Intergroup comparisons were conducted using the independent two-sample t-tests for continuous variables and the Pearson χ² test for categorical variables. The automated segmentation performance was quantified using the DSC and IoU. Classification performance was assessed through AUC, accuracy, sensitivity, specificity, positive and negative predictive values (PPV and NPV), and F1-score, each reported with corresponding 95% confidence intervals (CIs). The statistical significance of AUC differences between models was evaluated using DeLong’s test. Improvements in model reclassification and discrimination were quantified using the NRI and IDI metrics. Clinical utility was evaluated through DCA. A two-tailed p value < 0.05 was considered statistically significant. All analyses were performed using Python (v3.9.18) and relevant statistical libraries.

Hardware and software environment

FL training was performed on a high-performance computing cluster equipped with an Intel Xeon Platinum 8358 P CPU (2.60 GHz) and an NVIDIA RTX A6000 GPU (48 GB GDDR6 memory, CUDA v12.2). The FL framework and DL algorithms were implemented using PyTorch v2.0.0+cu117 within a 64-bit Python v3.9.18 environment. MATLAB R2021b was utilized for auxiliary data processing tasks. All statistical analyses were conducted in Python using the SciPy v1.11.0 and Statsmodels v0.14.1 libraries.

Supplementary information

Supplementary material (2.7MB, pdf)

Acknowledgements

This study was supported by the National Natural Science Foundation of China (82460361 to B.F., 12261027 to Y.L.), Chongqing Big Data Collaborative Innovation Center (CQBDCIC202304 to Y.L.), GUAT Special Research Projecton the Strategic Development of Distinctive Interdisciplinary Fields (TS2024231 to B.F.), and the Bagui Youth Top Talent training Program (to B.F.).

Author contributions

W.C.Z. conceived and designed the study and drafted the initial manuscript. Y.J.Q. conducted the federated learning experiments and contributed to manuscript writing. Y.X.H. performed data analysis and interpretation. W.Y.L., J.J.L., Y.X.H. and W.J.Z. acquired multi-institutional imaging and clinical data. X.W.Y. and Q.H.X. conducted the literature review and provided clinical and theoretical guidance. H.Y.Z. developed the paraspinal muscle segmentation models. Y.N.Z. performed the radiomics analyses. Y.L. carried out the statistical analyses. D.D.H. and Z.D.L. curated and managed the dataset. B.F. contributed to the methodological framework. W.S.L. supervised the overall project and critically revised the manuscript. All authors reviewed and approved the final manuscript for submission.

Data availability

The datasets generated and analyzed during the current study are not publicly available due to strict patient privacy regulations and institutional data governance policies, but are available from the corresponding author on reasonable request.

Code availability

To facilitate reproducibility and open science, the complete source code for the PMSAC-OVF system—including the VFMGL framework, radiomics analysis, and model inference scripts—is publicly available on GitHub at https://github.com/XmySz/PMSAC-OVF.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Weicong Zhang, Yangjie Qin, Yixiu Hao.

Contributor Information

Bao Feng, Email: fengbao1986.love@163.com.

Wansheng Long, Email: jmlws2@163.com.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s41746-026-02855-4.

References

  • 1.Compston, J. E., McClung, M. R. & Leslie, W. D. Osteoporosis. Lancet393, 364–376 (2019). [DOI] [PubMed] [Google Scholar]
  • 2.Ensrud, K. E. & Crandall, C. J. Osteoporosis. Ann. Intern. Med.177, ITC1–ITC16 (2024). [DOI] [PubMed] [Google Scholar]
  • 3.Preventive Services Task Force, U. S. et al. Screening for osteoporosis to prevent fractures: US preventive services task force recommendation statement. JAMA333, 498–508 (2025). [DOI] [PubMed] [Google Scholar]
  • 4.The Lancet Diabetes Endocrinology. Osteoporosis: overlooked in men for too long. Lancet Diabetes Endocrinol. 9, 1 (2021). [DOI] [PubMed]
  • 5.Wang, L. et al. Prevalence of osteoporosis and fracture in China: the China osteoporosis prevalence study. JAMA Netw. Open4, e2121106 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.LeBoff, M. S. et al. The clinician’s guide to prevention and treatment of osteoporosis. Osteoporos. Int.33, 2049–2102 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Huang, C. F. et al. Asia-pacific consensus on osteoporotic fracture prevention in postmenopausal women with low bone mass or osteoporosis but no fragility fractures. J. Formos. Med Assoc.122, S14–S20 (2023). [DOI] [PubMed] [Google Scholar]
  • 8.Gregson, C. L. et al. UK clinical guideline for the prevention and treatment of osteoporosis. Arch. Osteoporos.17, 80 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zeng, Q. et al. The prevalence of osteoporosis in China, a nationwide, multicenter DXA survey. J. Bone Min. Res.34, 1789–1797 (2019). [DOI] [PubMed] [Google Scholar]
  • 10.Cheng, X. et al. Opportunistic screening using low-dose CT and the prevalence of osteoporosis in China: a Nationwide, multicenter study. J. Bone Min. Res.36, 427–435 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Fuggle, N. R. et al. Fracture prediction, imaging and screening in osteoporosis. Nat. Rev. Endocrinol.15, 535–547 (2019). [DOI] [PubMed] [Google Scholar]
  • 12.Kahwati, L. C. et al. Screening for osteoporosis to prevent fractures: a systematic evidence review for the US Preventive Services Task Force. JAMA333, 509–531 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Agarwal, A. et al. Predictive performance of the Garvan Fracture Risk Calculator: a registry-based cohort study. Osteoporos. Int.33, 541–548 (2022). [DOI] [PubMed] [Google Scholar]
  • 14.Fink, H. A. et al. Performance of fracture risk assessment tools by race and ethnicity: a systematic review for the ASBMR Task Force on clinical algorithms for fracture risk. J. Bone Min. Res.38, 1731–1741 (2023). [DOI] [PubMed] [Google Scholar]
  • 15.Cozadd, A. J., Schroder, L. K. & Switzer, J. A. Fracture risk assessment: an update. J. Bone Jt. Surg. Am.103, 1238–1246 (2021). [DOI] [PubMed] [Google Scholar]
  • 16.Haug, C. J. & Drazen, J. M. Artificial intelligence and machine learning in clinical medicine, 2023. N. Engl. J. Med.388, 1201–1208 (2023). [DOI] [PubMed] [Google Scholar]
  • 17.Price, W. N. 2nd & Cohen, I. G. Privacy in the age of medical big data. Nat. Med.25, 37–43 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Jiang, M. R., Wang, Z. R. & Dou, Q. HarmoFL: harmonizing local and global drifts in federated learning on heterogeneous medical images. Proc. 36th AAAI Conf. Artif. Intell.36, 914–922 (2022).
  • 19.Allam, A. K., Anand, A., Flores, A. R. & Ropper, A. E. Computer vision in osteoporotic vertebral fracture risk prediction: a systematic review. Neurospine20, 1112–1123 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chen, Z. et al. Role of paraspinal muscle degeneration in the occurrence and recurrence of osteoporotic vertebral fracture: a meta-analysis. Front Endocrinol (Lausanne). 13, 1073013 (2023). [DOI] [PMC free article] [PubMed]
  • 21.Sheng, R. et al. Muscle-bone crosstalk via endocrine signals and potential targets for osteosarcopenia-related fracture. J. Orthop. Transl.43, 36–46 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Dong, Y., Yuan, H., Ma, G. & Cao, H. Bone-muscle crosstalk under physiological and pathological conditions. Cell Mol. Life Sci.81, 310 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Auger, J. D., Frings, N., Wu, Y., Marty, A. G. & Morgan, E. F. Trabecular architecture and mechanical heterogeneity effects on vertebral body strength. Curr. Osteoporos. Rep.18, 716–726 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Jeon, I., Kim, S. W. & Yu, D. Paraspinal muscle fatty degeneration as a predictor of progressive vertebral collapse in osteoporotic vertebral compression fractures. Spine J.22, 313–320 (2022). [DOI] [PubMed] [Google Scholar]
  • 25.Huang, W. et al. The association between paraspinal muscle degeneration and osteoporotic vertebral compression fracture severity in postmenopausal women. J. Back Musculoskelet. Rehabil.36, 323–329 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Bergh, J. P. et al. The clinical application of high-resolution peripheral computed tomography (HR-pQCT) in adults: state of the art and future directions. Osteoporos. Int.32, 1465–1485 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Wang, W. F. et al. The association between sarcopenia and osteoporotic vertebral compression refractures. Osteoporos. Int.30, 2459–2467 (2019). [DOI] [PubMed] [Google Scholar]
  • 28.Eguchi, Y. et al. Reduced leg muscle mass and lower grip strength in women are associated with osteoporotic vertebral compression fractures. Arch. Osteoporos.14, 112 (2019). [DOI] [PubMed] [Google Scholar]
  • 29.Lu, S. et al. General lightweight framework for vision foundation model supporting multi-task and multi-center medical image analysis. Nat. Commun.16, 2097 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Pyrrho, M., Cambraia, L. & de Vasconcelos, V. F. Privacy and health practices in the digital age. Am. J. Bioeth.22, 50–59 (2022). [DOI] [PubMed] [Google Scholar]
  • 31.Fang, C. et al. Decentralised, collaborative, and privacy-preserving machine learning for multi-hospital data. EBioMedicine101, 105006 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kalra, S., Wen, J., Cresswell, J. C., Volkovs, M. & Tizhoosh, H. R. Decentralized federated learning through proxy model sharing. Nat. Commun.14, 2899 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Shen, T. et al. Federated mutual learning: a collaborative machine learning method for heterogeneous data, models, and objectives. Front. Inf. Technol. Electron Eng.24, 1390–1402 (2023). [Google Scholar]
  • 34.Ma, J. et al. Segment anything in medical images. Nat. Commun.15, 654 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Maxime, O. et al. DINOv2: learning robust visual features without supervision. Transact. Mach. Learn. Res.1 (2024).
  • 36.Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature616, 259–265 (2023). [DOI] [PubMed] [Google Scholar]
  • 37.Sun, Y. et al. A data-efficient strategy for building high-performing medical foundation models. Nat. Biomed. Eng.9, 539–551 (2025). [DOI] [PubMed] [Google Scholar]
  • 38.Zhang, S. & Metaxas, D. On the challenges and perspectives of foundation models for medical image analysis. Med. Image Anal.91, 102996 (2024). [DOI] [PubMed] [Google Scholar]
  • 39.Gurusamy, P. et al. Density and fat fraction of the psoas, paraspinal, and oblique muscle groups are associated with lumbar vertebral bone mineral density in a multi-ethnic community-living population: the multi-ethnic study of atherosclerosis. J. Bone Min. Res.37, 1537–1544 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Wang, M., Liu, J., Chen, X., Wang, X. & Chen, X. Muscle quality and spine fractures. J. Cachexia Sarcopenia Muscle13, 1426–1428 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Liu, Y. et al. Development of a CT-Based comprehensive model combining clinical, radiomics with deep learning for differentiating pulmonary metastases from noncalcified pulmonary hamartomas: a retrospective cohort study. Int J. Surg.110, 4900–4910 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Saad, M. B. et al. Predicting benefit from immune checkpoint inhibitors in patients with non-small-cell lung cancer by CT-based ensemble deep learning: a retrospective study. Lancet Digit Health5, e404–e420 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Huang, Y. Q. et al. Enhanced risk stratification for stage II colorectal cancer using deep learning-based CT classifier and pathological markers to optimize adjuvant therapy decision. Ann Oncol. 36, 1178-1189 (2025). [DOI] [PubMed]
  • 44.Vettoruzzo, A., Bouguelia, M. R., Vanschoren, J., Rognvaldsson, T. & Santosh, K. C. Advances and challenges in meta-learning: a technical review. IEEE Trans. Pattern Anal. Mach. Intell.46, 4763–4779 (2024). [DOI] [PubMed] [Google Scholar]
  • 45.Kocak, B. et al. METhodological RadiomICs Score (METRICS): a quality scoring tool for radiomics research endorsed by EuSoMII. Insights Imaging15, 8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Tejani, A. S. et al. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radio. Artif. Intell.6, e240300 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Dourthe, B. et al. Automated segmentation of spinal muscles from upright open MRI using a multiscale pyramid 2D convolutional neural network. Spine (Phila Pa 1976). 47, 1179-1186 (2022). [DOI] [PubMed]
  • 48.Baur, D. et al. Analysis of the paraspinal muscle morphology of the lumbar spine using a convolutional neural network (CNN). Eur. Spine J.31, 774–782 (2022). [DOI] [PubMed] [Google Scholar]
  • 49.Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. methods18, 203–211 (2021). [DOI] [PubMed] [Google Scholar]
  • 50.Mcmahan, H. B., Moore, E., Ramage, D., Hampson, S. & Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. 20th Int. Conf. Artif. Intell. Stat. (AISTATS)54, 1273–1282 (2016). [Google Scholar]
  • 51.Lundberg, S. M. & Lee, S. I. A unified approach to interpreting model predictions. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS). 4765–4774 (2017).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary material (2.7MB, pdf)

Data Availability Statement

The datasets generated and analyzed during the current study are not publicly available due to strict patient privacy regulations and institutional data governance policies, but are available from the corresponding author on reasonable request.

To facilitate reproducibility and open science, the complete source code for the PMSAC-OVF system—including the VFMGL framework, radiomics analysis, and model inference scripts—is publicly available on GitHub at https://github.com/XmySz/PMSAC-OVF.


Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group

RESOURCES