Skip to main content
BMC Medical Imaging logoLink to BMC Medical Imaging
. 2026 Apr 21;26:338. doi: 10.1186/s12880-026-02363-7

Explainable machine learning-based mortality prediction in critically ill patients with rheumatoid arthritis-associated lung disease: a radiomics and clinical data integration study

Yubing He 1,#, Rui Liang 1,#, Yawen Zhang 1,#, Xinghui Li 1,✉
PMCID: PMC13348790  PMID: 42015074

Abstract

Background

Intensive Care Unit (ICU) mortality risk is high in rheumatoid arthritis (RA) patients, yet effective prognostic tools remain scarce.

Objectives

To develop and evaluate a machine learning (ML)-based prognostic model for predicting hospital mortality in severe RA patients.

Methods

This retrospective cohort study utilized data from the Medical Information Mart for Intensive Care IV (MIMIC-IV) and the Medical Information Mart for Intensive Care Chest X-ray (MIMIC-CXR) databases, including 1,951 chest X-rays from 984 patients with RA. The primary outcome was all-cause in-hospital mortality. Radiomics features were extracted using PyRadiomics, with 74 features retained after quality control. Key features were selected using the Boruta algorithm and integrated with clinical variables to develop three modeling strategies: clinical-only, radiomics-only, and combined models. Nine ML algorithms were applied using a 60/40 training-test split with 10-fold cross-validation. To address class imbalance (mortality rate: 7.7%), the Synthetic Minority Over-sampling Technique (SMOTE) was applied. Net Reclassification Improvement (NRI) and Integrated Discrimination Improvement (IDI) evaluated the predictive incremental value of imaging omics features combined with clinical data compared to single-modality data. SHapley Additive exPlanations (SHAP) was used to interpret model predictions.

Results

A total of 984 RA patients were included, of whom 76 (7.7%) experienced in-hospital mortality. The neural network model demonstrated superior performance, with an area under the curve (AUC) of 0.887 (95% CI: 0.752–0.934) in the training set and 0.800 (95% CI: 0.714–0.856) in the test set. The combined clinical-radiomics model showed significant incremental value compared to the clinical-only model (NRI: 0.3773, 95% CI: 0.3499–0.4047, P < 0.001; IDI: 0.2303, 95% CI: 0.2197–0.2409, P < 0.001) and the radiomics-only model (NRI: 0.2249, 95% CI: 0.1510–0.2988, P < 0.001; IDI: 0.1939, 95% CI: 0.1655–0.2222, P < 0.001). Key predictive features included blood urea nitrogen (BUN), Wavelet-LL 10th Percentile, and Wavelet-Haar HH Entropy.

Conclusion

The integrated ML model combining chest radiomics and clinical data effectively predicts mortality risk in critically ill RA patients, offering good generalizability and interpretability. It provides a practical, interpretable framework for clinical risk stratification and lays the foundation for the development of intelligent prognostic systems for patients with severe RA and resource allocation.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12880-026-02363-7.

Keywords: Rheumatoid arthritis, Machine learning, Radiomics, Mortality prediction, Prognostic model

Introduction

Rheumatoid arthritis (RA) is a systemic autoimmune disorder affecting approximately 0.5%-1% of the global population [1]. It leads to substantial incidence and disability not only from joint damage but also from extra-articular complications such as cardiovascular diseases, pulmonary disorders, and infections [2, 3], significantly increasing the risk of intensive care unit (ICU) admission [4]. Studies show that the one-year and five-year mortality rates for rheumatoid arthritis-associated interstitial lung disease (RA-ILD) are 13.9% and 39.0%, respectively, compared to 3.8% and 18.2% for patients with RA without interstitial lung disease (ILD). Mortality rates are 2 to 10 times higher in patients with RA-ILD than in those with RA without ILD [5]. However, conventional severity scores such as the Simplified Acute Physiology Score II (SAPS II) and Sequential Organ Failure Assessment (SOFA) perform poorly in RA patients, with reported areas under the curve (AUC) of only 0.60–0.70 [6], This poor performance stems from the fact that these scores do not incorporate RA-specific pathophysiological features, including immune dysregulation, chronic systemic inflammation, and progressive pulmonary involvement [7]. Consequently, there is an urgent need for prognostic tools tailored to this high-risk population, with a particular emphasis on quantifying pulmonary pathology.

Pulmonary complications, especially immune-mediated interstitial lung disease (ILD), are a frequent and grave manifestation of RA, significantly increasing mortality risk and representing an essential factor in critical care prognosis assessment [8–10]. Chest imaging in RA-ILD typically reveals reticular opacities, honeycombing, traction bronchiectasis, and ground-glass opacities [9].Unlike non-RA ILD (e.g., idiopathic pulmonary fibrosis), RA-ILD is characterized by RA-specific immune dysregulation-such as elevated levels of Tumor Necrosis Factor-alpha (TNF-α), Interleukin−6 (IL−6), and autoantibody-mediated lung tissue damage-which leads to distinct patterns of disease progression and radiomic signatures [8, 11]. Nevertheless, visual assessment by clinicians is often insensitive to early or subtle pathological changes, creating a diagnostic and prognostic gap [12, 13].

Radiomics enables high-throughput extraction of quantitative features (e.g., texture, gray level, spatial heterogeneity) from medical images, which is more sensitive for detecting subtle changes in lung parenchyma and has been shown to be significantly correlated with disease activity and lung function impairment. Previous prognostic models for critically ill RA patients have been largely unimodal, but each has inherent limitations. Clinical-only models achieve limited predictive performance (AUC ≤ 0.70) because they lack quantitative information on pulmonary pathology [13]. Conversely, imaging‑only models fail to capture systemic immune dysregulation and multi-organ involvement, making them less clinically actionable [14]. The limitations of unimodal methods indicate that a multimodal model integrating pulmonary radiomic features with clinical data is a key strategy to improve the prognostic predictive efficacy in critically ill RA patients.

However, radiomic features often lack clear pathophysiological explanations and are prone to technical variations, leading to insufficient stability when applied independently [15, 16]. Additionally, relying solely on imaging fails to comprehensively reflect systemic clinical conditions [17, 18]. To address it, we integrated radiomic features from chest X-rays with clinical parameters into a multimodal approach to construct a prognostic model with superior predictive efficacy and stronger interpretability compared with unimodal clinical or radiomic models for predicting ICU mortality in RA patients. By synergizing localized, quantitative imaging signatures with holistic clinical data, we seek to create a robust, interpretable tool for early risk stratification, thereby facilitating timely interventions and optimized resource allocation for this high-risk population.

Materials and methods

Study design and data source

This retrospective cohort study utilized the Medical Information Mart for Intensive Care IV (MIMIC-IV) database version 3.1 (https://mimic.mit.edu/) [19] and the Medical Information Mart for Intensive Care Chest X-ray (MIMIC-CXR) database version 2.0.0 (https://physionet.org/content/mimic-cxr/2.0.0/). The MIMIC-IV database contains comprehensive clinical data from patients admitted to the Beth Israel Deaconess Medical Center between 2008 and 2019, while MIMIC-CXR provides corresponding chest radiograph images. The study was conducted in accordance with the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) guidelines. An overview of the study design, including patient selection, feature extraction, model development, and evaluation workflow, is presented in Fig. 1.

Fig. 1.

Fig. 1

Study design and workflow flowchart. A schematic diagram depicting the entire study process, from data acquisition to model interpretation. The flowchart is divided into four main phases: (1) Data source and patient selection from the MIMIC-IV and MIMIC-CXR databases, (2) Feature engineering including clinical data preprocessing and radiomics feature extraction/selection, (3) Development and internal validation of nine machine learning models with the optimal model selected based on cross-validation performance, and (4) Model interpretation using SHAP and incremental value analysis (NRI/IDI)

Patient selection

Patients with RA were identified using International Classification of Diseases (ICD) codes (ICD-9: 714.x; ICD-10: M05.x, M06.x, M08.x). Inclusion criteria were: (1) age > 18 years at ICU admission, (2) first ICU admission during hospitalization, and (3) ICU length of stay ≥ 24 h. Patients were excluded if core clinical variables had > 20% missing values or if no chest X-ray was available within 24 h of ICU admission. The detailed selection process and final cohort size (n = 984) are illustrated in Fig. 2.

Fig. 2.

Fig. 2

Study population inclusion and exclusion flowchart. Flowchart depicting the selection process of rheumatoid arthritis patients from the MIMIC-IV database, showing the number of patients excluded at each step and the final cohort size

Clinical data extraction

Clinical variables were extracted within the first 24 h of ICU admission, including demographics (age, sex, race, body mass index [BMI]), severity scores SAPS II, SOFA, Acute Physiology Score III [APS III], Oxford Acute Severity of Illness Score [OASIS], Charlson Comorbidity Index), vital signs, and laboratory values (white blood cell [WBC] count, hemoglobin, platelet count, creatinine, blood urea nitrogen [BUN], albumin, lactate, C-reactive protein [CRP]). Comorbidities were identified using ICD codes. The primary outcome was hospital mortality. ICU length of stay was recorded for cohort description only and was not used as a predictor variable, as it contains information unavailable at the time of admission. RA-specific disease activity measures (e.g., Disease Activity Score 28 [DAS28], rheumatoid factor, anti-citrullinated protein antibody status) were not available in MIMIC-IV and are acknowledged as a limitation.

Radiomics feature extraction

Only posteroanterior (PA) and anteroposterior (AP) chest X-ray images acquired within 24 h of ICU admission were included. A total of 1,951 images from 984 patients were processed using PyRadiomics (version 3.0.1) in Python 3.7, yielding 483 initial features per image. A binary mask was generated using the 20th percentile of non-zero pixel intensity as the threshold to define the region of interest (ROI). This threshold-based approach was adopted as a computationally efficient method given the heterogeneous imaging conditions of ICU radiographs; however, it may include non-pulmonary structures such as the mediastinum and thoracic wall. Sensitivity analyses using the 10th and 30th percentile thresholds were conducted to evaluate the robustness of this choice. Features were extracted from both original and wavelet-transformed images using Daubechies-1, Daubechies-4, and Haar filters, encompassing first-order statistics, Gray Level Co-occurrence Matrix (GLCM), Gray Level Run Length Matrix (GLRLM), Gray Level Size Zone Matrix (GLSZM), Neighboring Gray Tone Difference Matrix (NGTDM), and Gray Level Dependence Matrix (GLDM). When both AP and PA views were available for a patient within the same time window, features were averaged to obtain a single patient-level profile.

Data preprocessing

Variables with > 30% missing values were excluded. Remaining missing values were imputed using random forest-based multiple imputation (missForest, version 1.5, 5 iterations). Patients excluded for excessive missing data may differ systematically from included patients, representing a potential source of selection bias. Radiomics features underwent sequential quality control: removal of features with > 20% missing values (n = 240), zero or near-zero variance (n = 13), and high intercorrelation (|r| > 0.95, n = 152), yielding 74 features for analysis.

Feature selection

The Boruta algorithm (version 8.0.0) with random forest was applied on the training set. Both confirmed and tentative features were retained; the potential noise from tentative features was mitigated by downstream regularization and cross-validation.

Model development

Nine algorithms were evaluated: logistic regression, least absolute shrinkage and selection operator (LASSO), Ridge regression, random forest, extreme gradient boosting (XGBoost), linear and radial basis function support vector machine (SVM), neural network, and gradient boosting machine. Data were split 60/40 (training/testing) via stratified sampling. Hyperparameters were optimized using 10-fold cross-validation repeated 3 times within the training set, with AUC as the primary metric. The Synthetic Minority Over-sampling Technique (SMOTE) was applied within cross-validation folds only. The testing set was entirely held out from model training and selection.

Temporal validation

An additional temporal split validation was performed (training: 2008–2016; testing: 2017–2019) to assess generalizability under potential temporal drift in clinical practice and imaging protocols.

Benchmark comparison

The best-performing model was compared against single-predictor logistic regression models constructed using SAPS II, APS III, OASIS, SOFA, Charlson Comorbidity Index, and age, with DeLong tests for paired AUC comparison. The Acute Physiology and Chronic Health Evaluation II (APACHE II) score was not available in MIMIC-IV; APS III, which shares a similar conceptual framework, served as a surrogate.

Model interpretation

SHapley Additive exPlanations (SHAP) analysis (kernelshap, version 0.4.1) was used to quantify feature contributions. Visualizations included beeswarm plots, force plots, waterfall plots for representative patients, and SHAP interaction plots. Individual patient-level radiomics visualizations (original image, intensity heatmap, ROI mask, and feature heatmap) were paired with SHAP outputs. To assess the biological plausibility of selected radiomics features, feature values were compared across pulmonary imaging patterns (normal, ground-glass opacity, consolidation, reticular, and mixed) using Kruskal-Wallis tests.

Incremental value and calibration

Three models (clinical-only, radiomics-only, and combined) were compared using Net Reclassification Improvement (NRI) with a 10% mortality risk threshold and Integrated Discrimination Improvement (IDI), calculated via PredictABEL (version 1.2-4) with 1,000 bootstrap iterations. Model calibration was evaluated using calibration plots, Brier scores, and the Hosmer-Lemeshow goodness-of-fit test.

Survival analysis

Kaplan-Meier analysis was performed to stratify patients into high- and low-risk groups based on the median predicted probability. Survival curves were compared using the log-rank test, and hazard ratios (HR) were estimated via Cox proportional hazards regression.

Post-hoc Power analysis

A post-hoc power analysis was conducted based on the observed AUC, event rate, and sample size to evaluate the statistical adequacy of the study.

Statistical analysis

Continuous variables were presented as mean ± standard deviation or median [interquartile range]; categorical variables as frequencies (%). Group comparisons used Student’s t-test or the Mann-Whitney U test for continuous variables, and the chi-square or Fisher’s exact test for categorical variables. Significance was set at two-sided P < 0.05. All analyses were performed using R version 4.3.0.

Results

Patient selection and baseline characteristics

Among 1,808 patients with RA from the MIMIC-IV 3.1 database, those with non-first admissions, age ≤ 18 years, ICU stay < 24 h, or missing chest X‑ray radiomic data were sequentially excluded, resulting in a final cohort of 984 patients (Fig. 2).

Baseline characteristics stratified by survival status showed significant differences in core features and illness severity indicators between the survival group (n = 908, 92.3%) and the non‑survival group (n = 76, 7.7%) (Table 1). The non‑survival group was older, had higher illness severity scores (SAPS II, SOFA, and Charlson Comorbidity Index), significantly elevated renal function markers (creatinine, BUN), and a longer median ICU length of stay(all P < 0.05). No significant differences were observed between the two groups in sex, body weight, mean arterial pressure, or heart rate. These findings suggest that mortality risk in critically ill RA patients is associated with older age, greater illness severity, renal dysfunction, and prolonged ICU stay.

Table 1.

Baseline characteristics of patients with rheumatoid arthritis

Characteristics Level Total (n = 984) Survival (n = 908) Death (n = 76) P-value SMD
Age (years) (mean ± SD) 70.16 (12.94) 69.42 (12.84) 78.95 (10.77) < 0.001 0.804
Gender (%) Male 324 (32.9) 296 (32.6) 28 (36.8) 0.529 0.089
Female 660 (67.1) 612 (67.4) 48 (63.2)
Weight (kg) (mean ± SD) 79.15 (24.24) 79.10 (24.14) 79.71 (25.57) 0.834 0.024
SAPS II score (median [Q1-Q3]) 34.50 [28.00, 43.00] 33.00 [27.00, 41.00] 48.00 [37.00, 60.00] < 0.001 0.928
SOFA score (median [Q1-Q3]) 3.00 [2.00, 6.00] 3.00 [2.00, 5.00] 7.00 [5.00, 8.00] < 0.001 0.851
Charlson comorbidity index (median [Q1-Q3]) 6.00 [4.00, 8.00] 6.00 [4.00, 8.00] 8.00 [7.00, 10.00] < 0.001 0.925
Mean arterial pressure (mmHg) (mean ± SD) 83.74 (20.70) 84.00 (20.43) 80.53 (23.59) 0.159 0.158
Heart rate (bpm) (mean ± SD) 91.86 (22.34) 91.58 (22.19) 95.16 (24.04) 0.180 0.155
Temperature (°C) (mean ± SD) 36.68 (0.85) 36.72 (0.76) 36.22 (1.50) < 0.001 0.415
Creatinine (mg/dL) (mean ± SD) 1.33 (1.12) 1.27 (1.08) 2.03 (1.33) < 0.001 0.628
BUN (mg/dL) (mean ± SD) 28.68 (26.12) 26.61 (24.37) 53.47 (32.96) < 0.001 0.927
WBC count (×10⁹/L) (mean ± SD) 11.77 (6.59) 11.65 (6.40) 13.18 (8.45) 0.051 0.205
Hemoglobin (g/dL) (mean ± SD) 10.19 (2.00) 10.21 (2.04) 10.03 (1.47) 0.464 0.098
ICU length of stay (hours) (median [Q1-Q3]) 54.45 [36.81, 102.36] 49.82 [35.66, 97.49] 93.76 [61.43, 219.26] < 0.001 0.470

Data are presented as mean ± SD, median [Q1-Q3], or n (%)

SD, standard deviation; IQR, interquartile range; SAPS, Simplified Acute Physiology Score; SOFA, Sequential Organ Failure Assessment; BUN, blood urea nitrogen; WBC, white blood cell; ICU, intensive care unit

Feature selection for risk prediction in critically Ill RA patients

To illustrate the feature extraction process, a representative ICU RA patient case is presented. The original chest X-ray (Fig. 3A) clearly displays the bilateral lung fields, cardiac silhouette, and bony thoracic structures. The pseudocolor mapping (Fig. 3B) visually highlights the grayscale distribution differences, revealing potential pathological areas. The ROI segmentation (Fig. 3C) automatically extracts the lung fields using the 20th percentile threshold method, effectively removing background and non-lung tissue interference. The feature heatmap (Fig. 3D) showcases the spatial heterogeneity of multi-dimensional texture features, providing a basis for the quantitative analysis of lung pathology imaging phenotypes. Through this process, 483 initial features were extracted, and after quality control, 74 features were retained.

Fig. 3.

Fig. 3

Representative examples of radiomics feature extraction from chest X-rays. (A) Original chest X-ray image. (B) Intensity heatmap visualization using JET colormap. (C) Radiomics region of interest (ROI) identified using threshold > 20th percentile. (D) Radiomics feature heatmap showing the spatial distribution of extracted features overlaid on the original image

Combining clinical data with the 74 radiomic features, we performed feature selection using the Boruta algorithm and a random forest classifier, ultimately identifying the top 30 key features for risk prediction in critically ill RA patients (Fig. 4). Among these, BUN, Wavelet-LL 10th Percentile, Wavelet-Haar HH Entropy, Wavelet-HH Kurtosis, Texture Busyness, Mean Gray Level, Temperature, and Wavelet-HL Maximum were consistently ranked as the most important. The distribution of key radiomics features across different pulmonary imaging patterns (normal, ground-glass opacity, consolidation, reticular, mixed) is shown in Supplementary Figure S1, supporting their biological plausibility.

Fig. 4.

Fig. 4

Ridge plot showing the distribution of feature importance for the top 30 features. Features are ranked by mean importance across multiple runs of the Boruta algorithm. Colors indicate feature decision status: Confirmed (dark blue), Tentative (light blue), or Rejected (gray). The x-axis represents feature importance (Z-score), and the y-axis lists feature names

Construction and performance comparison of the risk prediction model for critically Ill RA patients

Based on the identified key features, we developed RA mortality risk prediction models using nine machine learning (ML) algorithms and systematically evaluated their performance. The results (Fig. 5) showed that logistic regression, LASSO, and random forest methods achieved high AUC values on the training set but exhibited significant overfitting on the test set. Gradient boosting and XGBoost displayed insufficient generalization, with a notable decline in AUC on the test set. In contrast, Ridge regression and SVM with RBF performed reasonably well on the test set, but their overall performance was slightly inferior to that of neural networks. Comprehensive analysis revealed that the neural network model maintained a good balance between the training set (AUC = 0.887) and test set (AUC = 0.800), demonstrating the best generalization ability and, therefore, was selected as the final model for this study.

Fig. 5.

Fig. 5

Receiver operating characteristic (ROC) curves comparing model performance on training and testing datasets. Nine machine learning models are shown with their respective areas under the curve (AUC). Solid lines represent training set performance, and dashed lines represent testing set performance. The diagonal gray line indicates random performance (AUC = 0.5)

To quantitatively evaluate the prognostic improvement offered by integrating radiomics with clinical data, we compared the combined model against models using clinical features or radiomic features alone (Table 2). The combined clinical‑radiomics model showed significant incremental value over both the clinical‑only model and the radiomics‑only model, as assessed by NRI and IDI (all P < 0.001). Categorical NRI using a 10% mortality risk threshold further confirmed these findings (Supplementary Table S1).

Table 2.

Comparison of prognostic performance between combined and single models

Comparison NRI 95% CI P value IDI 95% CI P value
Clinical + Radiomics vs. Clinical 0.3773 0.3499 to 0.4047 < 0.001 0.2303 0.2197 to 0.2409 < 0.001
Clinical + Radiomics vs. Radiomics 0.2249 0.1510 to 0.2988 < 0.001 0.1939 0.1655 to 0.2222 < 0.001

Note: NRI = Net Reclassification Improvement; IDI = Integrated Discrimination Improvement; CI = Confidence Interval

Comprehensive evaluation of the neural network model performance

To further assess the discriminative ability and calibration of the optimal neural network model, we compared it with established clinical scoring systems and evaluated its predictive accuracy (Fig. 6). In the training set, the neural network model significantly outperformed the conventional severity scores APS III, SAPS II, OASIS, SOFA, Charlson Comorbidity Index, and age (all P < 0.05; Fig. 6A). In the testing set, the neural network achieved the highest AUC among all comparators and significantly outperformed SOFA, Charlson Comorbidity Index, and age (all P < 0.05; Fig. 6B). Calibration was excellent in both the training set (Brier score and Hosmer-Lemeshow test P > 0.05; Fig. 6C) and the testing set (Fig. 6D). Sensitivity analysis using different ROI thresholds (10th and 30th percentiles) and temporal split validation (2008–2016 vs. 2017–2019) confirmed the robustness and generalizability of the model (Supplementary Figures S2, S3). Post-hoc power analysis indicated adequate statistical power (99.87%) for the primary outcome.

Fig. 6.

Fig. 6

Comprehensive performance evaluation of the Neural Network model. (A) Training set ROC curves comparing the Neural Network model against APSIII, SAPSII, OASIS, SOFA, Charlson comorbidity index, and Age as single-predictor benchmark models, with DeLong test p-values annotated. (B) Testing set ROC curves with the same comparisons. (C) Calibration curve for the training set showing the relationship between predicted probabilities and observed outcomes across decile groups. (D) Calibration curve for the testing set. In panels A–B, AUC values with 95% CIs and DeLong test p-values are shown in the legend. In panels C–D, red dots and lines represent observed calibration points, the blue line represents LOESS smoothed calibration, and the dashed gray line indicates perfect calibration; Brier scores and Hosmer-Lemeshow test p-values are annotated

Global and individual interpretability analysis of the risk prediction model

To elucidate the decision-making mechanism and feature contributions of the aforementioned model, we conducted interpretability analysis using SHAP values. SHAP analysis revealed consistent feature contributions across training and test sets. BUN, Wavelet-HL Maximum, and Texture Busyness were positively associated with mortality risk, whereas Wavelet-LL 10th Percentile, Mean Gray Level, and Temperature showed protective effects (Fig. 7A and B). Force plots illustrated individual-level predictions (Fig. 7C and D). SHAP interaction analysis identified positive synergies between BUN and Wavelet‑LL 10th Percentile, as well as between Texture Busyness and Mean Gray Level (Supplementary Figure S4).

Fig. 7.

Fig. 7

SHAP analysis visualizations for model interpretation. (A) Beeswarm plot for training set showing SHAP value distributions for all features. (B) Beeswarm plot for testing set. (C) Force plot for training set displaying cumulative SHAP contributions across all patients. (D) Force plot for testing set. Features are ordered by importance, with colors representing feature values (red for high, blue for low)

Prediction contribution and imaging validation for typical cases in the risk prediction model for critically Ill RA patients.

To further empirically assess the model’s predictive value in multimodal feature fusion, we selected one high-risk and one low-risk RA patient for in-depth case analysis. In the high-risk patient, elevated BUN, reduced Wavelet-LL 10th Percentile, and increased Texture Busyness collectively drove a high predicted mortality risk(Fig. 8A). Furthermore, the chest X-ray and its radiomic heatmap revealed that the textured areas in the patient’s chest X-ray closely matched the predicted features, further validating the crucial role of radiomic features in high-risk patients (Fig. 8C). In the low-risk patient, opposite patterns were observed. Radiomics heatmaps visually aligned with the model’s predictions, supporting the clinical plausibility of the imaging features (Fig. 8B and D). These findings demonstrate that radiomics features, particularly texture features and gray-level distribution features, in synergy with clinical indicators such as BUN, can effectively characterize mortality risk in patients with RA. By integrating SHAP attribution analysis with radiomics visualization, the predictive logic of the model was intuitively validated at the individual level, further enhancing its clinical credibility.

Fig. 8.

Fig. 8

SHAP waterfall plots and radiomics visualizations for representative patients. (A) Waterfall plot for a high-risk mortality patient showing individual feature contributions to the prediction. (B) Waterfall plot for a low-risk mortality patient. (C) Radiomics visualizations for the high-risk patient including original X-ray, intensity heatmap, ROI mask, and feature heatmap. (D) Corresponding radiomics visualizations for the low-risk patient

Survival analysis based on model risk stratification

Patients were divided into high-risk (n = 492) and low-risk (n = 492) groups based on the median predicted probability derived from the neural network model. Kaplan–Meier survival curve analysis demonstrated that patients in the high-risk group had significantly lower survival rates compared with those in the low-risk group (log-rank, P < 0.001), with a HR of 3.38 (95% CI: 1.99–5.73) from Cox regression (Fig. 9).

Fig. 9.

Fig. 9

Kaplan-Meier survival curves stratified by Neural Network model risk prediction. Patients were divided into high-risk (n = 492) and low-risk (n = 492) groups based on the median predicted probability. The shaded areas represent 95% confidence intervals. The dashed vertical line indicates Day 28. The number at risk table is shown below the survival curves. Log-rank test p-value and HR with 95% confidence interval from Cox proportional hazards regression are annotated

Discussion

The persistently high mortality rate among critically ill patients with RA underscores the limitation of conventional prognostic tools, which fail to capture the unique pathophysiology of this population, particularly the contribution of subclinical pulmonary involvement [20–22]. In response, this study demonstrated that a multimodal model integrating routine clinical data with radiomic features derived from chest X-rays can effectively stratify the mortality risk of critically ill RA patients. Our neural network model achieved robust discriminative performance and good calibration, with substantial incremental value over clinical‑only or radiomics‑only models. SHapley Additive exPlanations (SHAP) analysis identified blood urea nitrogen (BUN) and wavelet‑based texture features—such as Wavelet‑HH Entropy and Wavelet‑LL 10th Percentile—as key predictors. These findings highlight the synergistic value of multimodal data and the potential of radiomics to capture subclinical pulmonary pathology that is not readily discernible by visual inspection.

This study is the first to apply wavelet transform to the detection and analysis of RA-associated pulmonary involvement. In contrast to traditional imaging analysis methods that mainly rely on grayscale and morphological indicators, wavelet transform can reveal the complex pathological processes of coexistent diffuse interstitial involvement and focal inflammation/fibrosis at multiple scales [23–25], thus providing a rich information foundation for risk prediction in RA-related pulmonary lesions. However, their full clinical value still relies on their integration with robust predictive models [26]. We systematically evaluated a variety of ML algorithms and confirmed that the neural network model had the optimal generalization performance, which could effectively adapt to the complex nonlinear interactions in the high-dimensional multimodal feature space [27]. More importantly, we decoded the “black box” of the ML model using the SHAP framework, converting model predictions into actionable insights for each variable [28, 29], which overcomes a key barrier to the clinical application of ML.

The good clinical interpretability of our model conferred by SHAP analysis lays a solid foundation for its clinical translation. BUN being identified as the foremost predictive factor is highly consistent with the conclusions of existing studies, reinforcing the critical role of renal dysfunction and metabolic stress in the outcomes of critical illnesses [30]. Meanwhile, the wavelet transform texture features emerging as core predictive factors suggest that they can reflect pathological features that are not discernible by physicians’ visual interpretation [31]. Although the precise biological correlations of these features require further investigation, their strong association with mortality risk implies that they may reflect lung parenchymal heterogeneity, interstitial changes, or altered tissue density resulting from the systemic inflammatory state of RA [32–34]. SHAP interaction analysis confirmed a positive synergy between BUN and Wavelet-LL 10th Percentile, which further validates the biological plausibility of the multimodal approach. Abnormalities in these features can collectively indicate a high-risk status in patients. In summary, the integration of classic clinical markers such as BUN with radiomic features enables the comprehensive capture of multidimensional information on disease progression, providing a more reliable theoretical basis for early risk warning and personalized treatment of RA patients.

We compared our neural network model with conventional ICU mortality scores including SAPS II, SOFA, APS III, OASIS, and the Charlson Comorbidity Index. The model significantly outperformed all these scores in the training set. In the test set, the neural network model achieved the highest AUC (0.800), significantly outperforming SOFA, the Charlson Comorbidity Index, and age, and exhibiting comparable performance to APS III, SAPS II, and OASIS. Although the APACHE II score was unavailable in the MIMIC-IV database, its reported AUC in the general ICU population ranges from 0.70 to 0.80 [35], and the performance of our model is at least comparable to it. Furthermore, our model demonstrated excellent calibration, whereas conventional scores often show poor calibration in the RA subgroup [7]. These findings suggest that the integration of radiomic features can improve the sensitivity and prognostic reliability of mortality prediction in critically ill RA patients.

Although this study has made progress in methodology and model development, several limitations remain to be addressed in future research. First, the retrospective single-center design may limit generalizability; prospective multicenter external validation is needed. Second, we used a percentile-based threshold for lung ROI definition rather than anatomical segmentation, which may include non-lung tissue (e.g., mediastinum, chest wall). Although sensitivity analysis with different thresholds (10th, 20th, 30th percentiles) showed stable model performance (test AUC range 0.781-0.800), future studies should adopt deep learning-based lung segmentation to improve anatomical precision. Additionally, the chest X-ray images were in JPEG format, which may lose grayscale dynamic range and textural details compared to DICOM; this could affect radiomics feature accuracy. Moreover, RA-specific disease activity indices (e.g., DAS28, rheumatoid factor, anti-CCP) were not available in the MIMIC-IV database, representing a significant gap because these factors influence prognosis. Furthermore, the study lacked an external validation cohort; temporal split validation (2008–2016 vs. 2017–2019) served as an internal check but does not replace independent multicenter data. Another concern is the potential overfitting risk: with 74 radiomics features and 76 mortality events, there remains a risk despite regularization and cross-validation. We also did not systematically evaluate the robustness of radiomics features against variations in imaging acquisition parameters (e.g., tube voltage, exposure dose). Finally, the computational requirements for real‑time bedside deployment have not been assessed.

Conclusion

This study developed and validated a multimodal model integrating chest X‑ray radiomics with clinical data that effectively stratifies mortality risk in critically ill RA patients. The neural network model demonstrated strong discriminative ability and good calibration, with BUN and wavelet‑based texture features as key predictors. Significant improvements in NRI and IDI quantitatively confirmed the added value of combining radiomic and clinical data over unimodal approaches. By using SHAP to provide global and individual‑level interpretability, we offer a practical, explainable framework for risk stratification using routinely available data. While further external validation and implementation studies are needed, this work lays a foundation for future intelligent prognostic systems in RA critical care.

Supplementary Information

Below is the link to the electronic supplementary material.

12880_2026_2363_MOESM1_ESM.docx (9.6KB, docx)

Supplementary Material 1: Reclassification tables for NRI analysis using a 10% mortality risk threshold. Panel A compares the Clinical + Radiomics combined model versus the Clinical model alone; Panel B compares the Clinical + Radiomics combined model versus the Radiomics model alone. For each comparison, cross-tabulation tables are presented separately for events (deaths, n = 76) and non-events (survivors, n = 908), showing the number of patients reclassified into high-risk (≥10%) or low-risk (<10%) categories. Event NRI, non-event NRI, and overall categorical NRI are reported. NRI = Net Reclassification Improvement.

12880_2026_2363_MOESM2_ESM.pdf (24KB, pdf)

Supplementary Material 2: Radiomics feature values across pulmonary imaging patterns for biological plausibility validation. Boxplots showing the distribution of six key radiomics features (Wavelet LL 10th Percentile, Texture Busyness, Mean Gray Level, Wavelet HL Maximum, Wavelet Haar HH Entropy, and Wavelet HH Kurtosis) across five imaging patterns: Normal, ground-glass opacity (GGO), Consolidation, Reticular, and Mixed. Individual data points are overlaid. Kruskal-Wallis test p-values are shown for each feature.

12880_2026_2363_MOESM3_ESM.pdf (16.2KB, pdf)

Supplementary Material 3: Sensitivity analysis of ROI threshold selection on Neural Network model performance. ROC curves comparing training set (red) and testing set (blue) performance using the 10th and 30th percentile thresholds for ROI definition. AUC values with 95% confidence intervals are shown in the legend.

12880_2026_2363_MOESM4_ESM.pdf (12.4KB, pdf)

Supplementary Material 4: Temporal validation of the Neural Network model. ROC curves comparing random split (red solid line) and temporal split (blue dashed line; training: 2008–2016, testing: 2017–2019) validation strategies for the 20th percentile threshold. (A) Training set comparison. (B) Testing set comparison. AUC values with 95% confidence intervals are shown in the legend.

12880_2026_2363_MOESM5_ESM.pdf (9.8MB, pdf)

Supplementary Material 5: SHAP interaction matrix scatter plots for the Neural Network model. (A) Training set SHAP interaction effects among top features. (B) Testing set SHAP interaction effects. Colors represent feature values, and the y-axis shows SHAP interaction values indicating the combined effect of feature pairs on model predictions.

Acknowledgements

The authors would like to acknowledge the developers and curators of the MIMIC-IV and MIMIC-CXR databases for making these valuable resources publicly available. We also thank the research team members for their contributions to data processing and analysis.

Abbreviations

RA

Rheumatoid Arthritis

ICU

Intensive Care Unit

SMOTE

Synthetic Minority Over-sampling Technique

MIMIC-IV

Medical Information Mart for Intensive Care IV

MIMIC-CXR

Medical Information Mart for Intensive Care Chest X-ray

ILD

interstitial lung disease

RA-ILD

arthritis-associated interstitial lung disease

ICD

International Classification of Diseases

BMI

Body Mass Index

APS III

Acute Physiology Score III

OASIS

Oxford Acute Severity of Illness Score

WBC

White Blood Cell

CRP

C-reactive protein

ML

Machine Learning

SVM

Support Vector Machine

XGBoost

eXtreme Gradient Boosting

LASSO

Least Absolute Shrinkage and Selection Operator

RBF

Radial Basis Function

AUC

Area Under the Curve

HR

Hazard Ratio

CI

Confidence Interval

AP

Anteroposterior

PA

Posteroanterior

ROI

Region of Interest

BUN

Blood Urea Nitrogen

GLCM

Gray Level Co-occurrence Matrix

GLDM

Gray Level Dependence Matrix

GLRLM

Gray Level Run Length Matrix

GLSZM

Gray Level Size Zone Matrix

IDI

Integrated Discrimination Improvement

NGTDM

Neighboring Gray Tone Difference Matrix

NRI

Net Reclassification Improvement

SAPS II

Simplified Acute Physiology Score II

SHAP

SHapley Additive exPlanations

SOFA

Sequential Organ Failure Assessment

APACHE II

Acute Physiology and Chronic Health Evaluation II

TRIPOD

Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis

DAS28

Disease Activity Score 28

TNF-α

Tumor Necrosis Factor-alpha

IL-6

Interleukin-6

Author contributions

Yubing He: Conceptualization, Methodology, Software, Formal analysis, Data Curation, Writing – Original Draft, Visualization. Rui Liang: Validation, Investigation, Resources, Writing – Review & Editing. Yawen Zhang: Validation, Investigation, Resources, Writing – Review & Editing. Xinghui Li: Supervision, Project administration, Writing – Review & Editing. All authors read and approved the final manuscript.

Funding

This study was supported by·the Scientific Research Project of Sichuan·Health and Family Planning. Commission (21PJ102); and the Doctoral Startup Fund of the Affiliated Hospital of North Sichuan Medical College (BS20211116) to Xinghui Li.

Data availability

The clinical datasets (MIMIC-IV v3.1) analyzed during the current study are available in the PhysioNet repository, https://mimic.mit.edu/. The chest X-ray image dataset (MIMIC-CXR v2.0.0) is available in the PhysioNet repository, https://physionet.org/content/mimic-cxr/2.0.0/. The code used for data extraction, radiomics analysis, and model development is available from the corresponding author upon reasonable request.

Declarations

Ethics approval and consent to participate

This study used secondary analysis of publicly available anonymized datasets, and was exempted from institutional ethics approval requirements.

Consent for publication

Not applicable, as this study uses de-identified, publicly available data without any individual patient identifiers. No personal information or images that could identify patients are included in this manuscript.

Competing interests

The authors declare no competing interests.

Declaration of AI and AI-assisted Technologies.

During the preparation of this work, the authors used DeepSeek-V3.2 for the purpose of language polishing, grammar checking, and improving the readability of the manuscript. Subsequently, all authors thoroughly reviewed and edited the manuscript and bear full responsibility for the final version.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Yubing He, Rui Liang and Yawen Zhang contributed equally to this work.

References

  • 1.Zhang Z, Gao X, Liu S, et al. Global, regional, and national epidemiology of rheumatoid arthritis among people aged 20–54 years from 1990 to 2021[J]. Sci Rep. 2025;15(1):10736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Li KZ, Tan CS. Corneal, scleral, choroidal, and foveal thickness in patients with rheumatoid arthritis[J]. Turk J Ophthalmol. 2018;48(6):326–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Mousavi SE, Nejadghaderi SA, Khabbazi A, et al. The burden of rheumatoid arthritis in the middle east and north africa region, 1990–2019[J]. Sci Rep. 2022;12(1):19297. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ridha A, Hussein S, AlJabban A, et al. The clinical impact of seropositivity on treatment response in patients with rheumatoid arthritis treated with etanercept: a real-world iraqi experience[J]. Open Access Rheumatol. 2022;14:113–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Hyldgaard C, Hilberg O, Pedersen AB, et al. A population-based cohort study of rheumatoid arthritis-associated interstitial lung disease: comorbidity and mortality[J]. Ann Rheum Dis. 2017;76(10):1700–6. [DOI] [PubMed] [Google Scholar]
  • 6.Fujiwara T, Tokuda K, Momii K, et al. Prognostic factors for the short-term mortality of patients with rheumatoid arthritis admitted to intensive care units[J]. BMC Rheumatol. 2020;4(1):64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Huang H, Chen R, Shao C, et al. Diffuse lung involvement in rheumatoid arthritis: a respiratory physician’s perspective[J]. Chin Med J (Engl). 2023;136(3):280–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Kadura S, Raghu G. Rheumatoid arthritis-interstitial lung disease: manifestations and current concepts in pathogenesis and management[J]. Eur Respir Rev, 2021, 30(160). [DOI] [PMC free article] [PubMed]
  • 9.Shi Y, Lin S, Shi Y, et al. From gut to lung: the role of bile acids in rheumatoid arthritis-associated interstitial lung disease (ra-ild)[J]. J Inflamm Res. 2025;18:10331–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Drakopanagiotakis F, Stavropoulou E, Tsigalou C et al. The role of the microbiome in connective-tissue-associated interstitial lung disease and pulmonary vasculitis[J]. Biomedicines, 2022, 10(12). [DOI] [PMC free article] [PubMed]
  • 11.Rea G, Sverzellati N, Bocchino M et al. Beyond visual interpretation: quantitative analysis and artificial intelligence in interstitial lung disease diagnosis expanding horizons in radiology[J]. Diagnostics (Basel), 2023, 13(14). [DOI] [PMC free article] [PubMed]
  • 12.Wang Y, Chen S, Zheng S, et al. The role of lung ultrasound b-lines and serum kl-6 in the screening and follow-up of rheumatoid arthritis patients for an identification of interstitial lung disease: review of the literature, proposal for a preliminary algorithm, and clinical application to cases[J]. Arthritis Res Ther. 2021;23(1):212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Magnin CY, Lauer D, Ammeter M, et al. From images to clinical insights: an educational review on radiomics in lung diseases[J]. Breathe (Sheff). 2025;21(1):230225. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Lafata KJ, Wang Y, Konkel B, et al. Radiomics: a primer on high-throughput image phenotyping[J]. Abdom Radiol (NY). 2022;47(9):2986–3002. [DOI] [PubMed] [Google Scholar]
  • 15.Ferrari R, Trinci M, Casinelli A, et al. Radiomics in radiology: what the radiologist needs to know about technical aspects and clinical impact[J]. Radiol Med. 2024;129(12):1751–65. [DOI] [PubMed] [Google Scholar]
  • 16.Joye AA, Bogowicz M, Gote-Schniering J, et al. Radiomics on slice-reduced versus full-chest computed tomography for diagnosis and staging of interstitial lung disease in systemic sclerosis: a comparative analysis[J]. Eur J Radiol Open. 2024;13:100596. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Salmanpour MR, Amiri S, Gharibi S et al. Radiological and biological dictionary of radiomics features: addressing understandable ai issues in personalized prostate cancer, dictionary version pm1.0[J]. J Imaging Inf Med, 2025. [DOI] [PMC free article] [PubMed]
  • 18.Malyszko J, Basak G, Batko K, et al. Erratum to: hematological disorders following kidney transplantation[J]. Nephrol Dial Transplant; 2020. [DOI] [PubMed]
  • 19.Johnson AEW, Bulgarelli L, Shen L, et al. Mimic-iv, a freely accessible electronic health record dataset[J]. Sci Data. 2023;10(1):1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Lezcano-Valverde JM, Salazar F, Leon L, et al. Development and validation of a multivariate predictive model for rheumatoid arthritis mortality using a machine learning approach[J]. Sci Rep. 2017;7(1):10189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Robles-Perez A, Luburich P, Bolivar S, et al. A prospective study of lung disease in a cohort of early rheumatoid arthritis patients[J]. Sci Rep. 2020;10(1):15640. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Shao Y, Zhang H, Shi Q, et al. Clinical prediction models of rheumatoid arthritis and its complications: focus on cardiovascular disease and interstitial lung disease[J]. Arthritis Res Ther. 2023;25(1):159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Nowak S, Creuzberg D, Theis M, et al. Comparing multi-texture fibrosis analysis versus binary opacity-based abnormality detection for quantitative assessment of idiopathic pulmonary fibrosis[J]. Sci Rep. 2025;15(1):1479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Acharya UR, Faust O, Sree SV, et al. Thyroscreen system: high resolution ultrasound thyroid image characterization into benign and malignant classes using novel combination of texture and discrete wavelet transform[J]. Comput Methods Programs Biomed. 2012;107(2):233–41. [DOI] [PubMed] [Google Scholar]
  • 25.Sharifi Kia D, Kim K, Simon MA. Current understanding of the right ventricle structure and function in pulmonary arterial hypertension[J]. Front Physiol. 2021;12:641310. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Lee T, Liu Y, Chang C, et al. Development of a risk prediction model for radiation dermatitis following proton radiotherapy in head and neck cancer using ensemble machine learning[J]. Radiat Oncol. 2024;19(1):78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Triantafyllidis A, Kondylakis H, Katehakis D, et al. Deep learning in mhealth for cardiovascular disease, diabetes, and cancer: systematic review[J]. JMIR Mhealth Uhealth. 2022;10(4):e32344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Lu S, Swisher CL, Chung C, et al. On the importance of interpretable machine learning predictions to inform clinical decision making in oncology[J]. Front Oncol. 2023;13:1129380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Oei SP, Bakkes T H G F, Mischi M, et al. Artificial intelligence in clinical decision support and the prediction of adverse events[J]. Front Digit Health. 2025;7:1403047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Seki M, Nakayama M, Sakoh T, et al. Blood urea nitrogen is independently associated with renal outcomes in japanese patients with stage 3–5 chronic kidney disease: a prospective observational study[J]. BMC Nephrol. 2019;20(1):115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Aerts HJWL, Velazquez ER, Leijenaar RTH, et al. Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach[J]. Nat Commun. 2014;5:4006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Phillips I, Ajaz M, Ezhil V, et al. Clinical applications of textural analysis in non-small cell lung cancer[J]. Br J Radiol. 2018;91(1081):20170267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Espinasse M, Pitre-Champagnat S, Charmettant B et al. Ct texture analysis challenges: influence of acquisition and reconstruction parameters: a comprehensive review[J]. Diagnostics (Basel), 2020, 10(5). [DOI] [PMC free article] [PubMed]
  • 34.Jing R, Wang J, Li J, et al. A wavelet features derived radiomics nomogram for prediction of malignant and benign early-stage lung nodules[J]. Sci Rep. 2021;11(1):22330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Knaus WA, Draper EA, Wagner DP, et al. Apache ii: a severity of disease classification system[J]. Crit Care Med. 1985;13(10):818–29. [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

12880_2026_2363_MOESM1_ESM.docx (9.6KB, docx)

Supplementary Material 1: Reclassification tables for NRI analysis using a 10% mortality risk threshold. Panel A compares the Clinical + Radiomics combined model versus the Clinical model alone; Panel B compares the Clinical + Radiomics combined model versus the Radiomics model alone. For each comparison, cross-tabulation tables are presented separately for events (deaths, n = 76) and non-events (survivors, n = 908), showing the number of patients reclassified into high-risk (≥10%) or low-risk (<10%) categories. Event NRI, non-event NRI, and overall categorical NRI are reported. NRI = Net Reclassification Improvement.

12880_2026_2363_MOESM2_ESM.pdf (24KB, pdf)

Supplementary Material 2: Radiomics feature values across pulmonary imaging patterns for biological plausibility validation. Boxplots showing the distribution of six key radiomics features (Wavelet LL 10th Percentile, Texture Busyness, Mean Gray Level, Wavelet HL Maximum, Wavelet Haar HH Entropy, and Wavelet HH Kurtosis) across five imaging patterns: Normal, ground-glass opacity (GGO), Consolidation, Reticular, and Mixed. Individual data points are overlaid. Kruskal-Wallis test p-values are shown for each feature.

12880_2026_2363_MOESM3_ESM.pdf (16.2KB, pdf)

Supplementary Material 3: Sensitivity analysis of ROI threshold selection on Neural Network model performance. ROC curves comparing training set (red) and testing set (blue) performance using the 10th and 30th percentile thresholds for ROI definition. AUC values with 95% confidence intervals are shown in the legend.

12880_2026_2363_MOESM4_ESM.pdf (12.4KB, pdf)

Supplementary Material 4: Temporal validation of the Neural Network model. ROC curves comparing random split (red solid line) and temporal split (blue dashed line; training: 2008–2016, testing: 2017–2019) validation strategies for the 20th percentile threshold. (A) Training set comparison. (B) Testing set comparison. AUC values with 95% confidence intervals are shown in the legend.

12880_2026_2363_MOESM5_ESM.pdf (9.8MB, pdf)

Supplementary Material 5: SHAP interaction matrix scatter plots for the Neural Network model. (A) Training set SHAP interaction effects among top features. (B) Testing set SHAP interaction effects. Colors represent feature values, and the y-axis shows SHAP interaction values indicating the combined effect of feature pairs on model predictions.

Data Availability Statement

The clinical datasets (MIMIC-IV v3.1) analyzed during the current study are available in the PhysioNet repository, https://mimic.mit.edu/. The chest X-ray image dataset (MIMIC-CXR v2.0.0) is available in the PhysioNet repository, https://physionet.org/content/mimic-cxr/2.0.0/. The code used for data extraction, radiomics analysis, and model development is available from the corresponding author upon reasonable request.


Articles from BMC Medical Imaging are provided here courtesy of BMC

RESOURCES