Summary
Background
Acute respiratory distress syndrome (ARDS) is a life-threatening condition with a high incidence and mortality rate in intensive care unit (ICU) admissions. Early identification of patients at high risk for developing ARDS is crucial for timely intervention and improved clinical outcomes. However, the complex pathophysiology of ARDS makes early prediction challenging. This study aimed to develop an artificial intelligence (AI) model for automated lung lesion segmentation and early prediction of ARDS to facilitate timely intervention in the intensive care unit.
Methods
A total of 928 ICU patients with chest computed tomography (CT) scans were included from November 2018 to November 2021 at three centers in China. Patients were divided into a retrospective cohort for model development and internal validation, and three independent cohorts for external validation. A deep learning-based framework using the UNet Transformer (UNETR) model was developed to perform the segmentation of lung lesions and early prediction of ARDS. We employed various data augmentation techniques using the Medical Open Network for AI (MONAI) framework, enhancing the training sample diversity and improving the model's generalization capabilities. The performance of the deep learning-based framework was compared with a Densenet-based image classification network and evaluated in external and prospective validation cohorts. The segmentation performance was assessed using the Dice coefficient (DC), and the prediction performance was assessed using area under the receiver operating characteristic curve (AUC), sensitivity, and specificity. The contributions of different features to ARDS prediction were visualized using Shapley Explanation Plots. This study was registered with the China Clinical Trial Registration Centre (ChiCTR2200058700).
Findings
The segmentation task using the deep learning framework achieved a DC of 0.734 ± 0.137 in the validation set. For the prediction task, the deep learning-based framework achieved AUCs of 0.916 [0.858–0.961], 0.865 [0.774–0.945], 0.901 [0.835–0.955], and 0.876 [0.804–0.936] in the internal validation cohort, external validation cohort I, external validation cohort II, and prospective validation cohort, respectively. It outperformed the Densenet-based image classification network in terms of prediction accuracy. Moreover, the ARDS prediction model identified lung lesion features and clinical parameters such as C-reactive protein, albumin, bilirubin, platelet count, and age as significant contributors to ARDS prediction.
Interpretation
The deep learning-based framework using the UNETR model demonstrated high accuracy and robustness in lung lesion segmentation and early ARDS prediction, and had good generalization ability and clinical applicability.
Funding
This study was supported by grants from the Shanghai Renji Hospital Clinical Research Innovation and Cultivation Fund (RJPY-DZX-008) and Shanghai Science and Technology Development Funds (22YF1423300).
Keywords: Deep learning, Automated lung CT segmentation, UNet Transformer model, Acute respiratory distress syndrome, Prediction
Research in context.
Evidence before this study
We searched PubMed with the terms “(deep learning OR artificial intelligence)” AND “acute respiratory distress syndrome” AND “prediction” published from database inception up to March 24, 2024, with no language restrictions. Several studies have investigated the use of machine learning algorithms for early prediction of ARDS using clinical variables, biomarkers, and physiological data. However, these studies have limitations such as small sample sizes, single-center designs, and lack of external validation. Moreover, the integration of quantitative lung lesion parameters derived from CT images with clinical data for ARDS prediction remains underexplored.
Added value of this study
Our study developed a deep learning-based framework using the UNETR model for automated lung CT segmentation and early prediction of ARDS. The framework was trained on a large dataset and validated in both retrospective and prospective cohorts. The integration of quantified lung lesion parameters and clinical data significantly improved the prediction performance. The deep learning-based framework also outperformed a Densenet-based image classification network, highlighting the importance of incorporating clinical expert knowledge in model development.
Implications of all the available evidence
Our findings demonstrate that a deep learning-based framework integrating quantitative lung lesion parameters and clinical data can accurately predict ARDS in the early stages, enabling the identification of high-risk patients for timely intervention. The framework exhibits good generalization ability and clinical applicability, as evidenced by its robust performance across multiple validation cohorts. Future research should focus on integrating multi-center data through federated learning and deploying the model on edge computing devices to enhance its real-world applicability and improve patient outcomes in the ICU.
Introduction
Acute respiratory distress syndrome (ARDS) is a life-threatening condition with an estimated global incidence of 10.4% of all intensive care unit (ICU) admissions and a pooled mortality rate of 43%.1,2 Despite advances in critical care medicine, ARDS remains a significant healthcare burden. The early identification of patients at high risk for developing ARDS is paramount for timely intervention and improved clinical outcomes. However, the complex pathophysiology of ARDS coupled with the heterogeneity of its clinical presentation makes it difficult to identify patients at high risk for developing the condition in its early stages.3, 4, 5, 6, 7
Computed tomography (CT) imaging provides a more detailed evaluation of lung pathology and is recognized for its potential in early ARDS detection.8 But the interpretation of CT images is highly subjective and necessitates considerable expertise resulting in inter-observer variability among radiologists and can lead to inconsistent diagnoses and potentially delaying treatment initiation.9,10
The advent of artificial intelligence holds promise for overcoming these hurdles and revolutionizing the early prediction and diagnosis of ARDS. Deep learning algorithms have the capability to automate the analysis of medical images, offering rapid and consistent quantification of lung abnormalities.11 Though artificial intelligence (AI) approaches have shown potential for early detection of ARDS, several challenges continue to exist. The development of robust AI models requires large, diverse and well-annotated datasets.12 The acquisition of such datasets is particularly challenging as the condition is relatively rare and the annotation of lung abnormalities on CT images is time-consuming and requires expert knowledge.13 Moreover, the performance of AI models may vary significantly across different patient populations and clinical settings.14
This study aims to develop a deep learning-based framework using the UNet Transformer (UNETR) model for automated lung lesion segmentation and early ARDS prediction. By integrating quantitative evaluations of lung lesions with clinical information, our approach seeks to enhance ARDS prediction accuracy. The model's reliability and generalizability were rigorously validated across three medical centers, highlighting its potential for improving early ARDS prediction.
Methods
Ethics
The study was approved by the Institutional Ethics Committee of Shanghai Jiao Tong University School of Medicine Affiliated Renji Hospital (approval number: KY2021-247-B). Informed consent was waived for the retrospective study, and written informed consent was obtained from all participants in the prospective study (ChiCTR2200058700). This study followed these guidelines in the TRIPOD-AI checklist.15
Study design and data sources
This multicenter cohort study included 928 patient from the ICUs of three hospitals. Patients admitted to the ICUs of participating centers were considered for inclusion in the study if they met any of the following criteria: 1) diagnosed with sepsis or septic shock; 2) suffered from severe trauma; 3) experienced aspiration; 4) diagnosed with severe acute pancreatitis; 5) required massive transfusion, defined as a single transfusion volume exceeding 1 to 1.5 times the patient's own blood volume, or transfusion volume exceeding half of the patient's own blood volume within 1 h, or transfusion rate exceeding 1.5 ml/(kg·min); 6) diagnosed with severe pneumonia; or 7) inhaled toxic gases. Patients were excluded from the study if they met any of the following criteria: 1) age under 18 years; 2) ICU length of stay less than 48 h; or 3) without CT images or CT images with low quality. The sample size for this study was determined based on two considerations to ensure the robustness and generalizability of our deep learning model. Given the complexity of the UNETR model, a substantial amount of data was necessary to effectively train the model and avoid overfitting. Empirical guidelines suggest having at least 10 times the number of model parameters in terms of training samples, providing a baseline for the minimum data required. For the retrospective cohort, data were obtained from Shanghai Jiao Tong University School of Medicine Affiliated Renji Hospital (Medical Center 1), Shanghai Public Health Clinical Center (Medical Center 2), and Shanghai Jiao Tong University Affiliated Sixth People's Hospital (Medical Center 3) covering the period from November 2018 to November 2021. Medical Center 1 provided data on 500 patients and 149,508 CT slices, with data from 350 patients used for training and data from 150 patients used for internal validation. A total of 8729 CT scans from the training group were manually segmented for subsequent training of the segmentation network. Medical Center 2 contributed data on 155 patients and 46,211 CT slices, serving as external validation group Ⅰ. Medical Center 3 contributed data on 92 patients and 26,785 CT slices, serving as external validation group Ⅱ. For the prospective cohort, patient information was collected from Medical Center 1 including 181 patients and 54,119 CT slices, which were used for prospective validation (Fig. 1). The initial laboratory results upon ICU admission and CT images obtained within 24 h before or after ICU admission for all included patients were utilized for analysis in this study. A schematic representation of the study design is provided in Fig. 2.
Fig. 1.
Overview of the entire dataset for our study. In our study, we included a total of 928 patients and 276,623 CT slices. The model underwent internal validation, and was further validated across two independent external cohorts as well as a prospective cohort.
Fig. 2.
The automated AI analysis framework for ARDS prediction in our study. The pipeline of our AI-based ARDS prediction framework is as follows: 1. Collection of patient CT images and clinical metadata; 2. Manual segmentation of CT images; 3. Development of an auto lung CT segmentation network based on UNETR; 4. 3D reconstruction of CT images and extraction of quantified parameters of different lung lesions; 5. The development of machine learning models based on the combination of lung lesion parameters and clinical metadata. 6. Validation of established ARDS prediction model in different cohorts and performance comparison with a CT classification network trained using a series of non-segmented CT images. Norm, Normalization, MLP; Multi-Layer Perceptron; Deconv, Deconvolution; Conv, Convolution; BN, Batch Normalization; ReLU, Rectified Linear Unit; ROC, Receiver Operating Characteristic.
Diagnosis of ARDS
The diagnosis of ARDS in this study was performed in accordance with the Berlin Definition criteria which include the following: acute onset of symptoms, bilateral infiltrates on chest imaging not fully explained by effusions, lobar/lung collapse, or nodules, respiratory failure not fully attributable to cardiac failure or fluid overload, and a PaO2/FiO2 ratio ≤300 mmHg with a minimum PEEP of 5 cm H2O.16 Diagnosis was independently conducted by three ICU clinicians (YZ, SM and JW) who reviewed the clinical data, laboratory results, and imaging findings of the patients. In instances where there was a lack of consensus among the three clinicians, the case was escalated to two senior ICU clinicians (ZH and YG) with substantial experience for further consultation and discussion to reach a conclusive diagnosis.
Imaging preprocessing and CT images manual segmentations
A total of 8729 CT scan slices were manually segmented for the training of the automated lung lesion segmentation network. Manual segmentation was performed using 3D Slicer, a comprehensive open-source platform for medical image informatics, image processing, and three-dimensional visualization. The segmentation labels were selected to represent common pulmonary CT lesion images in critically ill patients, including ground-glass opacity, consolidation, pulmonary fibrosis, and pleural effusion. The segmentations were annotated and reviewed by senior critical care physicians with years of clinical experience. To address the need for a large volume of training data, the dataset was significantly expanded through data augmentation using the Medical Open Network for AI (MONAI) framework developed by NVIDIA, thereby enhancing the diversity of training samples for the deep learning model. The augmentation process included techniques such as scale intensity range, random cropping, random affine transformations, random flip, random rotation, and random intensity shift. This approach improved the model's ability to generalize and substantially reduced the risk of overfitting. Furthermore, we employed the Random Crop by Positive and Negative Label (RCPNL) data augmentation transform within the MONAI framework, resizing CT images to a dimension of 96 × 96 × 128 to facilitate training. More details about the RCPNL data augmentation transform method are provided in Supplementary Material 1.
Automated lung CT segmentation framework
We employed the UNETR architecture as the backbone of our automated lung CT segmentation framework.17 During the training and evaluation of the automated lung CT segmentation framework, we utilized the Sliding Window Inference (SWI) method to reassemble the segmented outputs produced by the model and compared them with the ground truth labels. Additionally, we used a composite loss function combining dice loss with weighted cross-entropy at the pixel level, which captures various aspects of the segmentation task more effectively, aiding in the model's convergence and improving its generalization capabilities.18 The dice coefficient (DC) was calculated as follows:
where A is the set of predicted pixels and B is the set of ground truth pixels. The DC for the entire dataset was computed as a weighted average of the DC for each label, where the weights were determined by the proportion of each label in the dataset. The standard deviation (SD) was calculated to measure the variability of the DCs across the dataset.
We employed the AdamW optimizer with an initial learning rate of 1e-4 and a weight decay of 1e-5. Due to the high resolution of medical images and constraints on memory, we set the training batch size to 1. All models were developed from scratch using the PyTorch deep learning framework (version 2.2) in Python. The implementation of the deep learning framework is available at: https://github.com/ZhouYang30124/DL_ARDS. More details about the SWI method are provided in Supplementary Material 2.
Lung CT classification network
To compare the performance of a fully deep learning-based image classification network and a segmentation network trained with human knowledge for ARDS prediction, we developed a 3D image classification network based on the DenseNet architecture.19 This network uses raw, unsegmented CT images as input and predicts the risk of ARDS development. CT images of patients who developed ARDS during the course of the study were labeled as positive samples, while those from patients who did not develop ARDS were labeled as negative samples. These labeled CT images were used to train and validate the classification model. The network was implemented using the MONAI framework, which is tailored for efficient medical imaging tasks. Image preprocessing included intensity scaling, resizing, and random rotation to enhance data variability. The model utilizes a cross-entropy loss function for optimization and employs the Adam optimizer with an initial learning rate of 1e-4 to ensure effective learning progress. The implementation of the deep learning framework is available at: https://github.com/ZhouYang30124/DL_ARDS.
ARDS prediction model with integrated data and interpretation
We developed and compared seven models, including Logistic Regression (LR), K-Nearest Neighbors (KNN), Gaussian Naive Bayes (GNB), Random Forest (RF), Extreme Gradient Boosting (XGB), AdaBoost (ADB), and Gradient Boosting Decision Tree (GBDT) to predict ARDS. Following the development of our lung lesion segmentation models, we conducted 3D reconstructions of the segmented CT images to quantify the different lung lesions. We calculated parameters of lesions and integrated the data with patient clinical information and laboratory test results. Specifically, we were able to compute the percentage of lung volume affected by ground-glass opacity, consolidation, pulmonary fibrosis, and pleural effusion, as well as the average CT density of these lesion regions. By assessing the performance of seven different machine learning models, we identified the model that demonstrated the best predictive accuracy for ARDS. More details about hyperparameter tuning for the machine learning models and the final parameters are provided in Supplementary Material 3. To enhance the interpretability of our findings, we visualized the contributions of different features using Shapley Explanation Plots. This method allowed us to depict the impact of individual features on the model's predictions, providing insights into which variables were most influential in predicting ARDS outcomes, thus offering a clear visual explanation of the model's decision-making process.
Statistical analysis
Descriptive statistics were calculated for baseline characteristics across all participants and within each cohort. To address with the problem of missing baseline data, which were assumed to be missing at random, we employed multiple imputation using chained equations and verified imputation consistency by comparing distributions between imputed and complete datasets.20 Hyperparameter tuning for the developed machine learning models was performed via a grid search, incorporating 5-fold cross-validation to optimize settings. Shapley Additive Explanations plots were used to assess feature importance in our prediction models. The machine learning algorithms were implemented using Python 3.8.13 with libraries from Scikit-learn, ensuring robust and reproducible analysis. The predictive performance of all models was evaluated using ROC curves, calibration curves, and confusion matrices to compare their effectiveness. The optimum cut point for the confusion matrix was defined using Youden's Index (J), which is calculated as: J = Sensitivity + Specificity − 1. A 2-sided P < 0.05 was considered statistically significant.
Role of the funding source
The funder of the study had no role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the manuscript; and decision to submit the manuscript for publication.
Results
Dataset overview and patient characteristics for the study
Patient data were retrospectively collected from three medical centers between November 2018 and November 2021 (Fig. 1). In the overall patient population, the incidence of ARDS was 29.5% (274/928). The incidence in the retrospective cohort from Medical Center 1 was 29.0% (145/500), while in the prospective cohort, it was 27.6% (50/181). The incidence in Medical Center 2 was 30.3% (47/155), and in Medical Center 3, it was 34.8% (32/92). The median age of patients in the training and internal validation cohorts from Medical Center 1 was 67 years (IQR: 59–75). The proportion of male patients was 68.3% in the training cohort and 66.7% in the internal validation cohort. There were no significant differences in baseline clinical characteristics between the two cohorts (Table 1). Baseline clinical characteristics for all validation cohorts are provided in Supplementary Table S1. The distribution of ARDS causes across different cohorts is provided in Supplementary Table S2.
Table 1.
Baseline characteristics of patients in the training cohort and internal validation cohort.
| Characteristics | No. (%) |
P | |
|---|---|---|---|
| Training cohort (n = 350) | Validation cohort (n = 150) | ||
| Basic information | |||
| Age, median (IQR), years | 67 (57, 75) | 67 (59, 75) | 0.774 |
| Male sex | 239 (68.3%) | 100 (66.7%) | 0.802 |
| BMI, median (IQR), kg/m2 | 26.1 (22.6, 30.5) | 26.0 (22.6, 30.3) | 0.671 |
| Hypertension | 80 (22.9%) | 33 (22.0%) | 0.926 |
| Diabetes | 38 (10.9%) | 16 (10.7%) | 0.992 |
| Vital signs at admission | |||
| SBP, median (IQR), mmHg | 126 (113, 140) | 126 (111, 141) | 0.855 |
| DBP, median (IQR), mmHg | 68 (57, 78) | 69 (59, 78) | 0.378 |
| Heart rate, median (IQR),/min | 94 (80, 110) | 92 (79, 110) | 0.305 |
| RR, median (IQR),/min | 20 (18, 23) | 20 (18, 24) | 0.378 |
| Complete blood count tests | |||
| WBC, median (IQR), 109/L | 10.86 (7.11, 15.36) | 10.55 (7.04, 13.03) | 0.116 |
| Neutrophils, median (IQR), 109/L | 9.38 (5.55, 13.18) | 9.40 (5.35, 11.15) | 0.102 |
| Lymphocytes, median (IQR), 109/L | 0.79 (0.49, 1.17) | 0.67 (0.43, 1.11) | 0.106 |
| RBC, median (IQR), 1012/L | 3.34 (2.64, 4.05) | 3.33 (2.64, 4.06) | 0.871 |
| Platelets, median (IQR), 109/L | 188.5 (124.3, 246.5) | 172.5 (105.3, 226.3) | 0.073 |
| Hemoglobin, median (IQR), g/L | 104.0 (79.0, 122.8) | 98.0 (81.3, 119.0) | 0.683 |
| Arterial blood gas tests | |||
| Ph, median (IQR) | 7.39 (7.32, 7.45) | 7.40 (7.32, 7.44) | 0.993 |
| Sodium, median (IQR), mmol/L | 139.0 (135.8, 141.3) | 139.1 (136.0, 143.0) | 0.274 |
| Potassium, median (IQR), mmol/L | 3.7 (3.5, 4.1) | 3.8 (3.525, 4.2) | 0.275 |
| Calcium, median (IQR), mmol/L | 1.10 (1.06, 1.13) | 1.09 (1.06, 1.13) | 0.820 |
| HCO3-, median (IQR), mmol/L | 24.4 (22.0, 26.10) | 24.0 (21.5, 26.85) | 0.705 |
| Lactate, median (IQR), mmol/L | 1.8 (1.3, 3.075) | 2.0 (1.30, 2.98) | 0.529 |
| Biochemical tests | |||
| ALT, median (IQR), U/L | 33.0 (17.0, 72.4) | 32.6 (17.0, 74.1) | 0.873 |
| AST, median (IQR), U/L | 41.0 (24.0, 88.0) | 40.5 (27.0, 88.0) | 0.870 |
| Direct Bilirubin, median (IQR), μmol/L | 5.7 (3.2, 11.3) | 5.6 (3.8, 13.0) | 0.767 |
| Total Bilirubin, median (IQR), μmol/L | 14.2 (8.9, 25.5) | 14.0 (9.9, 24.5) | 0.654 |
| Creatinine, median (IQR), μmol/L | 76 (55, 106) | 69 (53, 103) | 0.093 |
| Albumin, median (IQR), g/L | 31.0 (26.6, 35.6) | 30.5 (26.6, 34.9) | 0.614 |
| CRP, median (IQR),mg/L | 82.8 (42.2, 131.8) | 84.1 (43.7, 134.7) | 0.905 |
| Coagulation tests & Cardiac biomarkers tests | |||
| Thrombin Time, median (IQR), seconds | 15.7 (14.5, 17.3) | 15.7 (14.5, 17.3) | 0.455 |
| APTT, median (IQR), seconds | 30.1 (27.4, 34.7) | 31.0 (28.1, 36.0) | 0.106 |
| Troponin I, median (IQR), ng/mL | 0.03 (0.01, 0.08) | 0.02 (0.01, 0.06) | 0.273 |
| CKMB, median (IQR), ng/mL | 3.4 (1.8, 7.975) | 3.55 (1.92, 6.35) | 0.642 |
| SOFA, median (IQR) | 2 (1,5) | 2 (1,5) | 0.915 |
| Ventilation | 0.533 | ||
| Mechanical ventilation | 74 (21.1%) | 28 (18.7%) | |
| Non-mechanical ventilation | 276 (78.9%) | 122 (81.3%) | |
Abbreviations: IQR, the interquartile range; BMI, Body Mass Index; SBP, Systolic Blood Pressure; DBP, Diastolic Blood Pressure; RR, Respiratory Rate; WBC, White Blood Cells; RBC, Red Blood Cells; ALT, Alanine Aminotransferase; AST, Aspartate Aminotransferase; CRP, C-Reactive Protein; APTT, Activated Partial Thromboplastin Time; CKMB, Creatine Kinase-MB; SOFA, Sequential Organ Failure Assessment.
Evaluation of the automated lung lesion segmentation model
To evaluate the performance of our AI system on CT slice segmentation, we calculated the DC values to assess the overlap between the predicted outputs and ground truth labels. After post-processing the predictions and labels, the results were aggregated across all validation batches to determine an overall average DC. The automated lung CT segmentation model demonstrated a decrease in training loss over 25,000 iterations, indicating a progressive learning process and convergence. In the test set, the automated segmentation model achieved a DC of 0.734, indicating high accuracy in lesion boundary delineation compared to human experts (Supplementary Fig. S1). Specifically, the DCs for the lung field, ground-glass opacity, consolidation, pulmonary fibrosis, and pleural effusion were as follows: 0.967 ± 0.098, 0.741 ± 0.131, 0.756 ± 0.157, 0.629 ± 0.018, and 0.701 ± 0.201, respectively. The training loss and validation DC for the UNETR-based automated segmentation model during the training process are shown in Supplementary Fig. S2.
Predicting ARDS with quantified lung lesion parameters and clinical data
In this study, we established ARDS prediction models using seven different machine learning algorithms. These models were based on two types of datasets: one derived solely from lung lesion parameters quantified by an automated CT segmentation network, and another from comprehensive data integrating general clinical information and laboratory test results of patients with the quantified lung lesion parameters. The results indicated that among the models constructed using lung lesion parameters, the XGB model exhibited superior performance, with an AUC of 0.860 and a 95% Confidence Interval (CI) of 0.783–0.930. For models built on comprehensive data, the XGB model again outperformed the others, achieving an AUC of 0.916 with a 95% CI of 0.858–0.961 (Fig. 3). Our findings demonstrate an improvement in the overall predictive performance of ARDS models when general clinical information and laboratory test results are incorporated into the machine learning training dataset.
Fig. 3.
The ROC curves of the machine learning models for ARDS prediction based on quantified lesion features and the combination of quantified lesion features and clinical metadata. A. The ROC curves of the machine learning model for ARDS prediction based on quantified lesion features; B. The ROC curves of the machine learning model for ARDS prediction based on the combination of quantified lesion features and clinical metadata. The XGBoost shows the best predictive performance for ARDS prediction. LR, Logistic Regression; KNN, K-Nearest Neighbors; GNB, Gaussian Naive Bayes; RF, Random Forest; XGB, eXtreme Gradient Boosting; ADB, AdaBoost; GBDT, Gradient Boosting Decision Tree.
External and prospective validation of the ARDS prediction model
In this study, the XGB model based on comprehensive data that integrated general clinical information, laboratory test results, and quantified lung lesion parameters was evaluated in external and prospective validation cohorts. The results demonstrated that the model achieved robust performance across all validation cohorts. Specifically, the model achieved an AUC of 0.916 with a 95% CI of 0.858–0.961 in the internal validation cohort. The AUC was 0.865 with a 95% CI of 0.774–0.945 in external validation cohort I. The AUC reached 0.901 with a 95% CI of 0.835–0.955 in external validation cohort II. The prospective validation cohort demonstrated an AUC of 0.876 with a 95% CI of 0.804–0.936 (Fig. 4). The calibration curve and confusion matrix for each validation cohort indicated that the model consistently performed well in predicting ARDS (Supplementary Figs. S3 and S4). Detailed information on the model's predictive performance in all validation cohorts is provided in Supplementary Table S3.
Fig. 4.
The ROC curves for ARDS prediction in all validation cohorts. A. Internal validation cohort; B. External validation cohort I; C. External validation cohort II; D. Prospective validation cohort.
Graphical interpretation of the ARDS prediction model
In our study, the clinical and radiological features that contributed to the onset of ARDS were further analyzed to develop an AI-assisted model for clinical prediction. To interpret the effects and relative contributions of the lung lesion features and clinical parameters on ARDS prediction, we implemented a Shapley Additive Explanation plot. As expected, the lesion features were identified as significant contributors to ARDS prediction. Several clinical parameters including C-reactive protein (CRP), albumin, bilirubin, platelet count, and aspartate aminotransferase (AST) levels, as well as general clinical characteristics such as age, also played important roles in predicting ARDS (Fig. 5).
Fig. 5.
The Shapley Explanation Plot for the ARDS prediction model based on the combination of quantified lesion features and clinical metadata. A. The relative contributions of CT and clinical parameters for ARDS prediction. Features on the right of the risk explanation bar pushed the risk higher, and features on the left pushed the risk lower. B. The relative contribution of each of the CT or clinical parameters to predict the risk of ARDS. C. An example of ARDS prediction estimation. We selected a patient from the ARDS group to show interpretability of the effects of lung-lesion features and clinical parameters as the input risk factors for ARDS prediction. The effects of input from lung-lesion features and clinical parameters for risk prediction. Pink features pushed the risk higher (to the right) and blue features pushed the risk lower (to the left).
Performance of densenet-based CT image classification network
In this study, we established a classification network designed to predict ARDS from unsegmented, raw pulmonary CT images. In the internal validation set, the network achieved an AUC of 0.796 with a 95% CI of 0.712–0.879 (Fig. 6). According to the confusion matrix, the network demonstrated a commendable predictive capacity. However, its performance was significantly inferior to the previously mentioned ARDS prediction models that employed a comprehensive integration of patient clinical information and laboratory test results. Using DeLong's test, we compared the AUC values of the Densenet model and the XGBoost model. The P-value for this comparison was 0.002, indicating there is a statistically significant difference between the two models' performance. The training process and validation of the Densenet classification network are shown in Supplementary Fig. S5.
Fig. 6.
The training process and validation of the Densenet classification network. A. The ROC curve of the Densenet classification network in the internal validation cohort; B. Confusion matrix of the Densenet classification network in the internal validation cohort.
Discussion
In this study, we successfully developed and validated an end-to-end deep learning pipeline utilizing the UNETR model for efficient and accurate segmentation of lung lesions and early ARDS prediction. This model was built upon on 276,623 CT slices from 928 patients and demonstrated high accuracy with an AUC of 0.916 and exhibited robustness across multiple cohorts, emphasizing the advantages of integrating clinical data. Recently, a new global definition for ARDS has been proposed, aiming to improve diagnostic consistency by better capturing the syndrome's complexity compared to the Berlin Definition. This new definition includes the use of pulse oximetry (SpO₂/FIO₂) as an alternative to arterial blood gas measurements, high-flow nasal oxygen (HFNO) for managing severe hypoxemic respiratory failure, and chest ultrasound as an imaging modality, especially in resource-limited settings.21 However, the refined definition still faces challenges in early prediction of ARDS, highlighting the critical need for research into early ARDS markers and the development of integrated, multimodal clinical prediction models to identify patients at high risk for ARDS.
CT imaging with its detailed insights into lung pathology, holds significant promise for enhancing early ARDS prediction.22 However, CT image interpretation is notably subjective, necessitating extensive radiological expertise. This variability in interpretation among radiologists can result in varied diagnoses, potentially delaying the commencement of necessary treatments.9,10 Moreover, the process of manual segmentation and quantifying lung abnormalities in CT scans is not only time-consuming but also requires substantial effort, which may not be feasible in the fast-paced setting of an intensive care unit.23 AI can expedite the analysis of CT scans by automating the identification of lung lesions, offering a swift and uniform approach to prediction and diagnosis.24 This technological advancement has the potential to significantly reduce subjectivity and variability in interpretations, thereby enhancing the efficiency of the clinical process in critical care settings.
The application of image segmentation techniques has become an integral part of medical imaging, particularly in the field of oncology.25,26 These methods enable precise delineation of lesions within medical images, which is crucial for accurate diagnosis, cancer staging, treatment planning, and monitoring treatment response. An example of such advancements is the implementation of the U-Net architecture for the segmentation of brain tumors from magnetic resonance imaging (MRI) scans.27 U-Net's application has significantly enhanced the precision in delineating gliomas, facilitating accurate tumor localization and volume calculations essential for effective treatment and prognosis assessment. Drawing inspiration from the substantial progress made in medical image segmentation, we recognize a significant potential for these technological advancements to be applied to the domain of pulmonary medicine. The implementation of sophisticated models, such as the deep learning-based UNETR model investigated in our study, may provide valuable insights and improvements in tackling the longstanding challenges associated with the early detection of ARDS. Compared to the U-Net architecture, the UNETR model used in this study benefits from its transformer-based design, which effectively captures long-range dependencies and global context, potentially enhancing segmentation accuracy and robustness in lung lesion delineation and early ARDS prediction.
Research in early ARDS prediction has increasingly incorporated machine learning methodologies to leverage both clinical and medical imaging data.28 Several studies have demonstrated the use of various machine learning algorithms including logistic regression, support vector machines, and neural networks to improve the prediction of ARDS onset using a combination of clinical variables, biomarkers, and physiological data. For instance, our previous study focused on predicting COVID-19-induced ARDS by integrating two prediction models based on CT images and clinical data using penalized logistic regression.29 A comprehensive review of early ARDS prediction studies highlighted the heterogeneity in study outcomes and the significant potential of machine learning methods. This review emphasized the lack of standardization in study design, data collection, and outcome measures, which hinders the comparability and generalizability of the findings.30 To address these gaps, the development of more sophisticated models that can efficiently handle and integrate large-scale, heterogeneous data from multiple sources is necessary. This includes the use of advanced machine learning techniques such as deep learning and multi-modal data fusion.
In our study, we developed a classification network based on unsegmented raw CT images to predict the onset of ARDS, which showed good predictive performance. However, its accuracy was significantly lower compared to the model that integrated quantified lung lesion features with clinical data. This underscores the importance of developing automated lung lesion segmentation models that effectively incorporate and learn from human expert knowledge. By integrating expert insights into segmentation, the model's performance exceeded that of standard convolutional neural networks relying solely on raw CT images. This synergy between human expertise and AI highlights the need to combine clinical insights with advanced machine learning techniques to create more effective and accurate predictive models for ARDS. Our findings emphasize that leveraging expert clinical knowledge in developing automated segmentation models can greatly enhance the accuracy and applicability of ARDS prediction, enabling more nuanced and clinically relevant models that reflect the complexity of ARDS pathophysiology.
We acknowledge several limitations in our study. First, the data were collected from three centers in China, which may limit the generalizability of our findings to other regions with different patient populations and healthcare practices. Second, while our model demonstrated high accuracy and robustness, it relies on the quality and consistency of the CT images and clinical data, which can vary between institutions. Third, the model's performance in real-time clinical settings has not been evaluated, and future studies should focus on integrating the model into clinical workflows and assessing its impact on patient outcomes.
In future research, the integration of data from multiple centers using techniques such as federated learning could be explored to further improve the performance and generalizability of the ARDS prediction model. Federated learning allows for the collaborative training of models across different institutions without the need for direct data sharing, addressing potential privacy and data ownership concerns.31, 32, 33 By leveraging a diverse range of patient populations and clinical settings, the model's robustness and applicability to real-world scenarios can be enhanced. Furthermore, the deployment of the developed model on edge computing devices such as the NVIDIA Jetson Nano could facilitate the integration of AI-driven decision support systems directly into the ICU environment. This approach would enable real-time predictions and timely interventions, empowering clinicians to make informed decisions at the point of care.34 The combination of advanced machine learning techniques, multi-center collaborations, and edge computing implementations holds great promise for transforming the early detection and management of ARDS in critical care settings, ultimately improving patient outcomes and optimizing resource allocation in the ICU.
This study advances ARDS prediction with a UNETR model but still faces challenges in clinical application, especially in developing user-friendly tools for real-time use by clinicians. Effective integration with hospital systems and creating intuitive interfaces for displaying ARDS risk based on lung imaging are crucial next steps. These require a comprehensive data integration framework to ensure compatibility with diverse hospital IT infrastructures.
Leveraging deep learning and over 276,623 CT slices, our study pioneers an AI-based UNETR model for early ARDS prediction, achieving expert-level accuracy and robust validation across multiple cohorts. Our study highlights the effective collaboration of AI and human expertise in advancing clinical practices. In future research, integrating multi-center data through federated learning and deploying the model on edge computing devices could further enhance its performance and applicability in real-world ICU settings, ultimately improving patient outcomes and resource allocation.
Contributors
YZ, QZ, ZH and YG were responsible for concept and design. YZ, SM, JW, QX, ZZ and SQ provided statistical analysis. All authors were involved in drafting and technical support in deep learning methods. WW, JF, CL, SX, XZ and FL were responsible for acquisition, analysis, or interpretation of data. YZ, SM and JW were involved in drafting the manuscript. YZ, QZ, ZH and YG had full access to all the data and verified the underlying data. All authors were involved in reviewing the manuscript and approved the final manuscript for submission.
Data sharing statement
The network and source code are fully available (https://github.com/ZhouYang30124/DL_ARDS). The data that supporting the findings of this study is available from the corresponding author upon request.
Declaration of interests
YG received funding from Shanghai Renji Hospital Clinical Research Innovation and Cultivation Fund (RJPY-DZX-008). SM received funding from Shanghai Science and Technology Development Funds (22YF1423300). All other authors declare no competing interests. All other authors declare no competing interests.
Acknowledgements
We thank all study participants. The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as potential conflicts of interest.
Footnotes
Supplementary data related to this article can be found at https://doi.org/10.1016/j.eclinm.2024.102772.
Contributor Information
Quanhong Zhou, Email: zhouanny@hotmail.com.
Zhengyu He, Email: zhengyuheshsmu@163.com.
Yuan Gao, Email: rj_gaoyuan@163.com.
Appendix A. Supplementary data
References
- 1.Bellani G., Laffey J.G., Pham T., et al. Epidemiology, patterns of care, and mortality for patients with acute respiratory distress syndrome in intensive care units in 50 countries. JAMA. 2016;315(8):788–800. doi: 10.1001/jama.2016.0291. [DOI] [PubMed] [Google Scholar]
- 2.Matthay M.A., Zemans R.L., Zimmerman G.A., et al. Acute respiratory distress syndrome. Nat Rev Dis Prim. 2019;5(1):18. doi: 10.1038/s41572-019-0069-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Rezoagli E., Fumagalli R., Bellani G. Definition and epidemiology of acute respiratory distress syndrome. Ann Transl Med. 2017;5(14):282. doi: 10.21037/atm.2017.06.62. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Gajic O., Dabbagh O., Park P.K., et al. Early identification of patients at risk of acute lung injury: evaluation of lung injury prediction score in a multicenter cohort study. Am J Respir Crit Care Med. 2011;183(4):462–470. doi: 10.1164/rccm.201004-0549OC. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Levitt J.E., Calfee C.S., Goldstein B.A., Vojnik R., Matthay M.A. Early acute lung injury: criteria for identifying lung injury prior to the need for positive pressure ventilation∗. Crit Care Med. 2013;41(8):1929–1937. doi: 10.1097/CCM.0b013e31828a3d99. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Zhai R., Sheu C.C., Su L., et al. Serum bilirubin levels on ICU admission are associated with ARDS development and mortality in sepsis. Thorax. 2009;64(9):784–790. doi: 10.1136/thx.2009.113464. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Soto G.J., Kor D.J., Park P.K., et al. Lung injury prediction score in hospitalized patients at risk of acute respiratory distress syndrome. Crit Care Med. 2016;44(12):2182–2191. doi: 10.1097/CCM.0000000000002001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Chung M., Bernheim A., Mei X., et al. CT imaging features of 2019 novel coronavirus (2019-nCoV) Radiology. 2020;295(1):202–207. doi: 10.1148/radiol.2020200230. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Rubenfeld G.D., Caldwell E., Granton J., Hudson L.D., Matthay M.A. Interobserver variability in applying a radiographic definition for ARDS. Chest. 1999;116(5):1347–1353. doi: 10.1378/chest.116.5.1347. [DOI] [PubMed] [Google Scholar]
- 10.Meade M.O., Cook R.J., Guyatt G.H., et al. Interobserver variation in interpreting chest radiographs for the diagnosis of acute respiratory distress syndrome. Am J Respir Crit Care Med. 2000;161(1):85–90. doi: 10.1164/ajrccm.161.1.9809003. [DOI] [PubMed] [Google Scholar]
- 11.Nam J.G., Park S., Hwang E.J., et al. Development and validation of deep learning-based automated detection algorithm for malignant pulmonary nodules on chest radiographs. Radiology. 2019;290(1):218–228. doi: 10.1148/radiol.2018180237. [DOI] [PubMed] [Google Scholar]
- 12.Song Y., Zheng S., Li L., et al. Deep learning enables accurate diagnosis of novel coronavirus (COVID-19) with CT images. IEEE ACM Trans Comput Biol Bioinf. 2021;18(6):2775–2780. doi: 10.1109/TCBB.2021.3065361. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Gerard S.E., Herrmann J., Kaczka D.W., Musch G., Fernandez-Bustamante A., Reinhardt J.M. Multi-resolution convolutional neural networks for fully automated segmentation of acutely injured lungs in multiple species. Med Image Anal. 2020;60 doi: 10.1016/j.media.2019.101592. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Albahri O.S., Zaidan A.A., Albahri A.S., et al. Systematic review of artificial intelligence techniques in the detection and classification of COVID-19 medical images in terms of evaluation and benchmarking: taxonomy analysis, challenges, future solutions and methodological aspects. J Infect Public Health. 2020;13(10):1381–1396. doi: 10.1016/j.jiph.2020.06.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Collins G.S., Moons K.G.M., Dhiman P., et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385 doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.ARDS Definition Task Force. Ranieri V.M., Rubenfeld G.D., et al. Acute respiratory distress syndrome: the Berlin Definition. JAMA. 2012;307(23):2526–2533. doi: 10.1001/jama.2012.5669. [DOI] [PubMed] [Google Scholar]
- 17.Hatamizadeh A., Yang D., Roth H., Xu D. 2022 IEEE/CVF winter conference on applications of computer vision (WACV) 2022. UNETR: transformers for 3D medical image segmentation; pp. 574–584. [DOI] [Google Scholar]
- 18.Milletari F., Navab N., Ahmadi S.-A. 2016 fourth international conference on 3D vision (3DV) IEEE; 2016. V-net: fully convolutional neural networks for volumetric medical image segmentation; pp. 565–571. [Google Scholar]
- 19.Huang G., Liu Z., Van Der Maaten L., Weinberger K.Q. 2017 IEEE conference on computer vision and pattern recognition (CVPR) 2017. Densely connected convolutional networks; pp. 2261–2269. [DOI] [Google Scholar]
- 20.van Buuren S., Groothuis-Oudshoorn K. Mice: multivariate imputation by chained equations in R. J Stat Software. 2011;45(3):1–67. [Google Scholar]
- 21.Matthay M.A., Arabi Y., Arroliga A.C., et al. A new global definition of acute respiratory distress syndrome. Am J Respir Crit Care Med. 2024;209(1):37–47. doi: 10.1164/rccm.202303-0558WS. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Pontone G., Scafuri S., Mancini M.E., et al. Role of computed tomography in COVID-19. J Cardiovasc Comput Tomogr. 2021;15(1):27–36. doi: 10.1016/j.jcct.2020.08.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Hofmanninger J., Prayer F., Pan J., Röhrich S., Prosch H., Langs G. Automated lung segmentation in routine imaging is primarily a data diversity problem, not a methodology problem. Eur Radiol Exp. 2020;4(1):50. doi: 10.1186/s41747-020-00173-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Müller D., Soto-Rey I., Kramer F. Robust chest CT image segmentation of COVID-19 lung infection based on limited data. Inform Med Unlocked. 2021;25 doi: 10.1016/j.imu.2021.100681. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Sinha A., Dolz J. Multi-scale self-guided attention for medical image segmentation. IEEE J Biomed Health Inform. 2021;25(1):121–130. doi: 10.1109/JBHI.2020.2986926. [DOI] [PubMed] [Google Scholar]
- 26.Pang S., Du A., Orgun M.A., et al. CTumorGAN: a unified framework for automated computed tomography tumor segmentation. Eur J Nucl Med Mol Imaging. 2020;47(10):2248–2268. doi: 10.1007/s00259-020-04781-3. [DOI] [PubMed] [Google Scholar]
- 27.Çiçek Ö., Abdulkadir A., Lienkamp S.S., Brox T., Ronneberger O. In: Ourselin S., Joskowicz L., Sabuncu M., Unal G., Wells W., editors. vol. 9901. Springer; Cham: 2016. 3D U-Net: learning dense volumetric segmentation from sparse annotation; pp. 424–432. (Medical image computing and computer-assisted intervention – MICCAI 2016. Lecture notes in computer science). [Google Scholar]
- 28.Zeiberg D., Prahlad T., Nallamothu B.K., Iwashyna T.J., Wiens J., Sjoding M.W. Machine learning for patient risk stratification for acute respiratory distress syndrome. PLoS One. 2019;14(3) doi: 10.1371/journal.pone.0214465. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Zhou Y., Feng J., Mei S., et al. A deep learning model for predicting COVID-19 ARDS in critically ill patients. Front Med. 2023;10 doi: 10.3389/fmed.2023.1221711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Parsons P.E. Biomarkers in the acute respiratory distress syndrome: past, present, and future. Trans Am Clin Climatol Assoc. 2022;132:107–116. [PMC free article] [PubMed] [Google Scholar]
- 31.Sheller M.J., Edwards B., Reina G.A., et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci Rep. 2020;10(1) doi: 10.1038/s41598-020-69250-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Rieke N., Hancox J., Li W., et al. The future of digital health with federated learning. NPJ Digit Med. 2020;3:119. doi: 10.1038/s41746-020-00323-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sinaci A.A., Gencturk M., Alvarez-Romero C., et al. Privacy-preserving federated machine learning on FAIR health data: a real-world application. Comput Struct Biotechnol J. 2024;24:136–145. doi: 10.1016/j.csbj.2024.02.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Wang R., Lai J., Zhang Z., Li X., Vijayakumar P., Karuppiah M. Privacy-Preserving federated learning for internet of medical things under edge computing. IEEE J Biomed Health Inform. 2023;27(2):854–865. doi: 10.1109/JBHI.2022.3157725. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.






