ABSTRACT
This systematic review and meta‐analysis evaluated the performance of machine learning (ML) models in predicting mortality among pulmonary embolism (PE) patients, synthesizing data from 17 studies encompassing 844,071 cases. Logistic Regression was the most commonly used algorithm, followed by advanced models like Random Forests, Support Vector Machines, XGBoost, and Neural Networks. Pooled performance metrics from 12 studies demonstrated a sensitivity of 0.88 (95% CI: 0.78–0.94, I 2 = 90.43%), specificity of 0.79 (95% CI: 0.62–0.89, I 2 = 99.53%), positive likelihood ratio of 4.1 (95% CI: 2.2–7.7), negative likelihood ratio of 0.16 (95% CI: 0.08–0.29), diagnostic odds ratio of 26 (95% CI: 10–71), and an AUROC of 0.91 (95% CI: 0.88–0.93), indicating excellent discriminative ability. Subgroup analyses revealed higher sensitivity in advanced ML models (89.7%) and non‐USA studies (97.2%), with advanced ML showing lower specificity heterogeneity (I 2 = 0%). Significant heterogeneity was observed, particularly in specificity (I 2 = 99%), driven by traditional ML and USA‐based studies. Minimal publication bias was noted for sensitivity (Egger's p = 0.942), but specificity showed potential bias (Egger's p = 0.038 after outlier exclusion). These findings suggest that ML models outperform traditional risk stratification tools in predicting PE mortality, offering robust potential for clinical decision‐making, though heterogeneity and retrospective study designs warrant cautious interpretation.
Trial Registration: PROSPERO: CRD420251026696
Keywords: embolism, machine learning, mortality, pulmonary
1. Introduction
Pulmonary embolism (PE), a manifestation of venous thromboembolism (VTE), ranks as the third most common cause of cardiovascular death worldwide, following stroke and myocardial infarction [1]. In the United States, PE contributes to ~300,000 deaths annually, with a 30‐day mortality rate of 2%–7% despite advancements in treatment [1, 2]. PE often originates from deep vein thrombosis (DVT) in the lower extremities, presenting with symptoms ranging from asymptomatic to fatal [3]. Risk factors for adverse outcomes include elevated heart rate, right ventricular dysfunction, cancer, previous surgery, immobilization, and markers of thromboinflammation such as D‐dimer, C‐reactive protein, and white blood cell count [4, 5, 6]. Comorbidities like heart failure (HF) significantly increase mortality risk, with a 9.1% incidence of PE in severe chronic HF patients and a 12.2% in‐hospital mortality rate in PE patients with HF, reduced to 9.7% with interventions like inferior vena cava filters [7]. Other associated conditions include acute kidney injury (AKI), malignancy, and hemodynamic instability, particularly in critically ill PE patients requiring intensive care unit (ICU) management [8, 9, 10]. These comorbidities and risk factors complicate risk stratification, as traditional scores like the simplified Pulmonary Embolism Severity Index (sPESI) and European Society of Cardiology (ESC) guidelines yield moderate predictive accuracy (c‐statistics 0.75–0.85) but low positive predictive values (21%–26%) for complications [11, 12]. This underscores the need for more precise tools to identify high‐risk patients and guide clinical decisions.
Machine learning (ML) models have emerged as powerful tools for improving outcome prediction in hematological and cardiopulmonary conditions, including VTE and PE [13, 14]. Unlike traditional regression‐based models, ML methods analyze complex data distributions, capturing probabilistic relationships and conditional dependencies among variables, and validate results through robust techniques like cross‐validation [15]. In hematological complications like VTE, ML models enhance risk stratification by integrating clinical, laboratory, and imaging data, outperforming conventional models in scenarios such as anticoagulation discontinuation, where 3.91% of 34,447 PE patients faced increased VTE‐related risks due to early therapy cessation [13]. In cardiopulmonology, ML models improve predictions for conditions like heart failure (HF) by incorporating biomarkers (e.g., N‐terminal pro‐B‐type natriuretic peptide, lymphocyte‐to‐white blood cell ratio) and imaging features (e.g., right ventricular strain), addressing the interplay of cardiac dysfunction and PE‐related complications. These advancements suggest ML's potential to refine risk assessment in complex clinical settings.
This systematic review and meta‐analysis evaluates the performance of ML models for predicting 30‐day mortality in PE patients, synthesizing evidence from retrospective cohort studies across multiple countries. By analyzing models like XGBoost, Random Forest, and Neural Networks, which incorporate diverse features such as sPESI, radiomics, and vital signs, we aim to assess their predictive accuracy compared to traditional risk scores. The study seeks to identify optimal ML approaches for risk stratification, addressing gaps in current tools and supporting personalized management of PE patients to reduce mortality and improve outcomes in clinical practice, particularly in high‐risk settings like the ICU.
2. Methods
2.1. Study Protocol and Registration
This systematic review and meta‐analysis was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta‐Analyses (PRISMA) guidelines [16]. The study protocol was registered with PROSPERO (Registration Number: CRD420251026696). PRISMA checklist is available in Supporting File S1.
2.2. Search Strategy
A comprehensive literature search was conducted across PubMed, Scopus, Web of Science, Embase, EBSCO, and Google Scholar to identify studies evaluating ML models for predicting mortality related to PE, published up to May 2025. The search strategy used keywords including “machine learning,” “pulmonary embolism,” and “mortality,” supplemented by controlled vocabulary terms from Emtree and MeSH to capture relevant synonyms and related concepts. The full search syntax is provided in Supporting File S1. No language or publication date restrictions were applied. Reference lists of included studies were manually reviewed to identify additional relevant articles. Detailed search strategy for different databases is available in Table S1.
2.3. Eligibility Criteria
Studies were included if they: (1) evaluated ML or statistical models for predicting mortality in adult populations with PE, and (2) reported data on diagnostic performance, such as true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN), or other performance metrics (e.g., accuracy, sensitivity, specificity). Studies lacking any performance metrics were initially included but excluded from the meta‐analysis if TP, FP, FN, and TN could not be derived. Exclusions also applied to case reports, book chapters, reviews, conference abstracts, or studies not involving ML or statistical models.
2.4. Study Selection and Data Extraction
Two reviewers independently screened titles, abstracts, and full texts for eligibility, resolving discrepancies through consensus. Data were extracted into a standardized Excel spreadsheet, capturing the following variables: first author and year, type of study, country, validation method, selected features, sample size, number of cases with mortality and controls (no mortality) in test and train sets, follow‐up duration, mean age, male percentage, ML algorithm (e.g., Logistic Regression, Random Forest, SVM, XGBoost, Neural Networks), and best predictor algorithm. Figure 1 shows the study selection flowchart.
Figure 1.

Study selection flowchart.
2.5. Quality Assessment
The quality of included studies was assessed using the Prediction Model Risk of Bias Assessment Tool for Artificial Intelligence (PROBAST + AI) [17] and the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) framework [18]. The assessment, conducted across the PROBAST + AI domains (Participants, Predictors, Outcomes, Analyses, and Overall for Development [Dev], Evaluation [Eval], and Applicability [App]), provides insights into the risk of bias (ROB) and applicability of the included studies. The participant domain assessed cohort selection and representativeness, focusing on geographic and demographic diversity. The predictor domain evaluated the consistency and relevance of features, such as clinical, radiologic, or radiomic data. The outcome domain examined the standardization of mortality definitions and follow‐up duration. The analysis domain scrutinized validation methods (e.g., cross‐validation, internal/external validation) and model complexity to identify risks of overfitting or inadequate reporting. Each domain was rated as low, moderate, or high risk of bias, with applicability concerns noted for generalizability. GRADE was used to determine the overall certainty of evidence, considering risk of bias, inconsistency, indirectness, imprecision, and publication bias, with evidence graded as high, moderate, low, or very low. Assessments were conducted independently by two reviewers, with disagreements resolved through consensus, ensuring a systematic evaluation of study quality.
2.6. Data Synthesis and Statistical Analysis
Diagnostic performance data were extracted to calculate performance metrics for each study. A bivariate random‐effects model was employed to pool data, accounting for variability in model performance across diverse settings, features, and algorithms. Subgroup analyses were conducted based on ML model type (traditional vs. advanced) and country to explore potential sources of variation. Heterogeneity was assessed using the likelihood ratio test (LRT) for χ 2 statistics and the I 2 statistic to evaluate the proportion of variation due to between‐study differences. Analyses were performed using Stata (Version 18) with the midas command. Validation methods and algorithm types were summarized to contextualize model robustness, with no additional bias assessments conducted in this primary analysis.
3. Results
3.1. Study Characteristics
This review a total of 844,071 cases across the 17 included retrospective cohort studies, based on the summed sample sizes derived from the provided data. Studies were conducted across multiple countries, including the USA (7 studies), Australia (1), China (1), Iran (1), Turkey (1), Germany (1), Romania (1), Israel (1), Spain/UK (1), and a multinational cohort spanning 27 countries (1), with most studies (11/17) focusing on a 30‐day follow‐up period, except for 1 study with a 215‐day follow‐up and 5 studies with unspecified follow‐up. Mean age, reported in seven studies, ranged from 56.4 to 67.3 years, and the percentage of male participants, reported in ten studies, ranged from 40% to 53.56%. Validation methods varied: six studies used 5‐fold cross‐validation, two used 10‐fold cross‐validation, one used 3‐fold cross‐validation, one used bootstrap‐corrected cross‐validation, three used external validation, one used both internal and external validation, one used 10‐fold cross‐validation with bootstrap resampling, and four did not specify validation methods. Selected features included clinical data (e.g., age, sex, vital signs, sPESI, PESI), laboratory tests (e.g., cTnI, NT‐proBNP, hemoglobin), imaging data (e.g., 3D‐CTPA, radiomic features, lung/cardiac ROI), and comorbidities (e.g., cancer, HF, COPD). ML algorithms employed included Logistic Regression (7 studies), Random Forest (4), XGBoost (4), Support Vector Machines (SVM, 3), Neural Networks (2), Decision Trees (2), Adaptive Boosting (2), and others (e.g., AutoML, CNN‐based SANet, PEP‐Net). The best‐performing algorithms varied, with Logistic Regression (5 studies), XGBoost (2), Neural Networks (1), Random Forest (1), SVM (1), Adaptive Boosting (1), PEP‐Net (1), AutoML (1), Adaptive LASSO (1), Classification Tree Analysis (1), CART Decision Tree (1), and Multimodal CNN‐based fusion (1) identified as top predictors. Table 1 shows detailed studies characteristics.
Table 1.
Studies characteristics.
| References | Type of study | Country | Validation | Selected features | Sample size | Follow‐up | Mean age/male% | Machine learning algorithms | Best predictor algorithm |
|---|---|---|---|---|---|---|---|---|---|
| Aujesky et al. [19] | Retrospective cohort | USA | External validation | Age, sex, cancer, heart failure, chronic lung disease, temperature < 36°C, pulse ≥ 110/min, systolic BP < 100 mmHg, respiratory rate ≥ 30/min, altered mental status, arterial oxygen saturation < 90% (11 clinical data) | 15,625 | 30 days | —/43% | Logistic Regression | Logistic Regression |
| Aujesky et al. [20] | Retrospective cohort | USA | External validation | Age, sex, cancer, heart failure, chronic lung disease, temperature < 36°C, pulse ≥ 110/min, systolic BP < 100 mmHg, respiratory rate ≥ 30/min, altered mental status, arterial oxygen saturation < 90% (11 clinical data) | 15,625 | 30 days | —/43% | Classification Tree Analysis | Classification Tree Analysis |
| Aujesky et al. [21] | Retrospective cohort | USA | — | Age, sex, cancer, heart failure, chronic lung disease, temperature < 36°C, pulse ≥ 110/min, systolic BP < 100 mmHg, respiratory rate ≥ 30/min, altered mental status, arterial oxygen saturation < 90% (11 clinical data) | 10,721 | 30 days | —/40% | Logistic Regression | Logistic Regression |
| Parast et al. [22] | Retrospective cohort | USA | 5‐fold cross‐validation | 10 | 1596 | 30 days | 57.8 ± 16.7 years/44.7% male | Adaptive LASSO Logistic Regression | Adaptive LASSO (aLASSO) |
| Lau et al. [23] | Retrospective cohort | Australia | — | sPESI (age > 80 years, malignancy, chronic cardiopulmonary disease, heart rate ≥ 110 beats/min, systolic BP ≤ 100 mmHg, oxyhemoglobin saturation < 90%), serum sodium, serum bicarbonate | 1426 | — | 67.3 ± 16.5 years/44.7% male | Logistic Regression | Logistic Regression |
| Liu et al. [24] | Retrospective cohort | China | — | cTnI, NT‐proBNP, RVa/LVa ratio, RAd/LAd ratio, Mastora score, septal angle, Qanadli score, reflux into azygos vein | 88 | 30 days | — | Logistic Regression | Logistic Regression |
| Danilatou et al. [25] | Retrospective cohort | USA | Bootstrap corrected cross‐validation | Clinical and radiomics features | 2468 | — | — | AutoML (Random Forest) | AutoML (Random Forest) |
| Danilatou et al. [26] | Retrospective cohort | USA | 5‐fold cross‐validation | 35 features (demographics, clinicolaboratory, medications, procedures, clinical scores, vital signs) | 6853 | 215 days (7.5 months) | 61.5/52% | Random Forest, SVM, XGBoost | Random Forest |
| Mora et al. [27] | Retrospective cohort | Multinational (27 countries) | 10‐fold cross‐validation and bootstrap resampling (50 times) | 70 variables (full model), 23 variables (reduced model) | 1348 | 30 days | —/— | Decision Tree, k‐Nearest Neighbors, SVM, RUSBoost, Neural Network | Neural Network |
| Cahan et al. [28] | Retrospective cohort | Israel | 5‐fold cross‐validation | 3D‐CTPA imaging, EHR tabular data (vital signs, blood gases, lactate, WBC, age, comorbidities, medications) | — | 30 days | —/— | CNN‐based SANet, Swin UNETR, TabNet, XGBoost, 1DCNN, FC layer | Multimodal (CNN‐based intermediate fusion) |
| Sadegh‐Zadeh et al. [29] | Retrospective cohort | Iran | 5‐fold cross‐validation | 4 laboratory tests, 12 vital signs, heart failure, COPD, syncope, RV enlargement, S1T3Q3, fibrinolytic therapy, embolectomy, age, gender | 604 | — | 61.90 ± 17.18/— | Decision Trees, Random Forests, Adaptive Boosting, SVM, Neural Networks (MLP) | Adaptive Boosting |
| Wang et al. [30] | Retrospective cohort | USA | 10‐fold cross‐validation | General info, lab tests (INR, creatinine, etc.), vital signs, complications, treatment (vasopressor, ventilation, etc.), sPESI, ESC, SAPS II, SOFA | 1229 | 30 days | —/51% | XGBoost, Logistic Regression | XGBoost |
| Tur et al. [31] | Retrospective cohort | Turkey | 5‐fold cross‐validation | Lung‐ROI, Cardiac‐ROI from CT scans | 193 | 30 days | —/— | ResNet18, ResNet50, EfficientNetB0, PEP‐Net (3DResNet + XGBoost) | PEP‐Net (3DResNet + XGBoost) |
| Liu et al. [32] | Retrospective cohort | USA | Internal and external validation (eICU‐CRD) | Hemoglobin_min, APS III score, WBC_max, SOFA score, CCI, race, ventilation, peripheral vascular disease, malignant cancer, metastasis of solid tumor, warfarin, NOAC, urine output < 400 mL | 472 | 30 days | —/— | XGBoost, Logistic Regression, Random Forest, CatBoost, Light GBM, SVM | SVM |
| Mesinovic et al. [33] | Retrospective cohort | Spain, UK | 5‐fold cross‐validation | Demographics (age, sex, country), comorbidities (hypertension, diabetes, smoking), symptoms (cough, fever, fatigue), alpha variant | 800,459 | — | 56.4 ± 20.9/48.6% | Logistic Regression, LDA, Naive Bayes, Random Forests, ADABoosting, XGBoost | XGBoost |
| Shahzadi et al. [34] | Retrospective cohort | Germany | 3‐fold cross‐validation | Radiomic features + IMAT, sPESI | 829 | 30 days | 65/53.56% | Logistic Regression | Logistic Regression |
| Teodoru et al. [35] | Retrospective cohort | Romania | None | COVID‐19 infection, NLR, arterial oxyhemoglobin saturation < 90%, altered mental status, PESI, sPESI, WBC count | 160 | — | 65/52.5% | CART Decision Tree | CART Decision Tree |
3.2. Pooled Performance Metrics
Meta‐analysis of 12 best performance models showed pooled sensitivity of 88% (95% CI: 78%–94%, I 2 = 90.4%), indicating that ML models correctly identified 88% of patients at risk of mortality, making them reliable for detecting high‐risk cases in emergency settings. However, high interstudy variability (intraclass correlation coefficient = 0.24, 95% CI: 0.04–0.43) suggests caution when generalizing results across diverse algorithms and populations. Advanced ML models (e.g., Neural Networks, XGBoost) and non‐USA studies achieved higher sensitivity (89.7% and 97.2%, respectively), offering clinicians greater confidence in identifying at‐risk patients in these contexts. The pooled specificity was 79% (95% CI: 62%–89%, I 2 = 99.5%), reflecting moderate accuracy in ruling out mortality risk, but marked heterogeneity (intraclass correlation coefficient = 0.37, 95% CI: 0.18–0.57) driven by outliers (e.g., Aujesky et al. [19], specificity = 34.9%; Lau et al. [23], specificity = 27.1%) (Figure 2). Clinicians should thus verify negative predictions with clinical judgment, especially when using traditional ML models. The positive likelihood ratio (PLR) of 4.1 (95% CI: 2.2–7.7) suggests a positive ML prediction moderately increases the likelihood of mortality, aiding prioritization of high‐risk patients, while the negative likelihood ratio (NLR) of 0.16 (95% CI: 0.08–0.29) indicates a negative prediction strongly reduces this risk, supporting decisions to de‐escalate care. The diagnostic odds ratio (DOR) of 26 (95% CI: 10–71) and area under the receiver operating characteristic curve (AUROC) of 0.91 (95% CI: 0.88–0.93) (Figure 2) confirm excellent discriminative ability, outperforming traditional tools like PESI for risk stratification. Significant heterogeneity (LRT_Q = 246.5, df = 2, p < 0.001; overall I 2 = 99%) was observed, with minimal contribution from threshold effects (0.04), indicating that algorithm type and geographic region drive variability. Advanced ML models showed lower specificity heterogeneity (I 2 = 0%), making them more reliable for consistent performance across settings. Publication bias was minimal for sensitivity (Egger's p = 0.942), but significant for specificity (Egger's p = 0.038 after outlier exclusion), suggesting potential underreporting of low‐specificity studies, particularly in traditional ML and USA subgroups. For physicians, this means ML models, especially advanced ones, can enhance clinical decision‐making by accurately identifying high‐risk PE patients and guiding resource allocation, but negative predictions should be cross‐referenced with clinical scores in settings with high variability. Table 2 shows performance metrics of different ML models.
Figure 2.

Forest plot for sensitivity and specificity based on best‐performing models (left). SROC plot for AUC based on best‐performing models (right).
Table 2.
Performance metrics of different ML models.
| References | Machine learning algorithms | Accuracy | Sensitivity | Specificity | Precision | F1 score | AUROC | AUROC (upper CI) | AUROC (lower CI) |
|---|---|---|---|---|---|---|---|---|---|
| Aujesky et al. [19] | LR | — | — | — | — | — | 0.78 | 0.8 | 0.76 |
| Aujesky et al. [20] | Classification Tree | — | 0.97 | 0.35 | 0.12 | — | — | — | — |
| Aujesky et al. [21] | LR | — | 0.96 | 0.47 | 0.14 | — | 0.78 | — | — |
| Parast et al. [22] | Adaptive LASSO (aLASSO) | — | 0.49 | 0.9 | 0.397 | — | 0.77 | — | — |
| Lau et al. [23] | LR | — | 1 | 0.27 | 0.05 | — | 0.86 | 0.93 | 0.79 |
| Liu et al. [24] | LR | — | 0.8 | 0.8 | — | — | 0.889 | 0.93 | 0.78 |
| Danilatou et al. [25] | RF | — | — | — | — | — | 0.81 | — | — |
| Danilatou et al. [26] | RF | 0.77 | 0.74 | 0.93 | — | 0.74 | 0.93 | 0.95 | 0.91 |
| SVM | 0.86 | 0.7 | 0.89 | — | 0.56 | 0.89 | 0.91 | 0.87 | |
| XGBOOST | 0.87 | 0.29 | 0.95 | — | 0.63 | 0.84 | 0.85 | 0.83 | |
| Mora et al. [27] | Neural Network | 0.965 | 0.961 | 0.96 | 0.963 | — | 0.96 | 0.98 | 0.95 |
| Decision Tree | 80 | 79.8 | 80.2 | 80.4 | — | 0.8 | 0.81 | 0.79 | |
| KNN | 87.2 | 84.8 | 84.8 | 85 | — | 0.85 | 0.86 | 0.84 | |
| SVM | 88.3 | 87.1 | 86.9 | 87.2 | — | 0.88 | 0.89 | 0.86 | |
| RUSBoost | 91.5 | 91.1 | 90.7 | 91.4 | — | 0.91 | 0.92 | 0.9 | |
| Cahan et al. [28] | Multimodal (CNN‐based intermediate fusion) | 0.93 | 0.9 | 0.94 | 0.69 | — | 0.96 | 1 | 0.93 |
| Sadegh‐Zadeh et al. [29] | Adaptive Boosting | — | — | — | — | — | 0.872 | — | — |
| Decision Tree | — | — | — | — | — | 0.771 | — | — | |
| RF | — | — | — | — | 0.87 | — | — | — | |
| SVM | — | — | — | — | — | 0.846 | — | — | |
| Artificial Neural Network | — | — | — | — | — | 0.794 | — | — | |
| Wang et al. [30] | XGBoost | — | — | — | 0.904 | 0.786 | 0.82 | 0.87 | 0.78 |
| LR | — | — | — | 0.885 | 0.715 | 0.746 | 0.7 | 0.8 | |
| Tur et al. [31] | PEP‐Net (3DResNet + XGBoost) | 0.945 | 0.977 | 0.874 | — | — | 0.917 | — | — |
| ResNet18 | 0.783 | 0.782 | 0.535 | — | — | 0.581 | — | — | |
| ResNet50 | 0.803 | 0.525 | 0.703 | — | — | 0.583 | — | — | |
| EfficientNetB0 | 0.716 | 0.661 | 0.774 | — | — | 0.67 | — | — | |
| Liu et al. [32] | SVM | 0.864 | 0.902 | 0.64 | 0.872 | — | 0.963 | — | — |
| RF | — | — | — | — | — | 0.938 | — | — | |
| XGBoost | — | — | — | — | — | 0.927 | — | — | |
| Catboost | — | — | — | — | — | 0.923 | — | — | |
| Light GBM | — | — | — | — | — | 0.895 | — | — | |
| LR | — | — | — | — | — | 0.874 | — | — | |
| Mesinovic et al. [33] | XGBoost | 0.653 | 0.74 | — | — | — | 0.742 | — | — |
| LR | 0.662 | 0.684 | — | — | — | 0.73 | — | — | |
| LDA | 0.78 | 0.78 | — | — | — | 0.73 | — | — | |
| Naive Bayes | 0.749 | 0.23 | — | — | — | 0.71 | — | — | |
| RF | 0.655 | 0.71 | — | — | — | 0.739 | — | — | |
| ADABoosting | 0.655 | 0.715 | — | — | — | 0.731 | — | — | |
| Shahzadi et al. [34] | LR | — | 0.74 | 0.54 | — | — | 0.7 | 0.79 | 0.6 |
| Teodoru et al. [35] | CART Decision Tree | 0.94 | 0.882 | 0.955 | — | — | — | — | — |
3.3. Subgroup Analysis
Traditional ML models (Logistic Regression, CART, Classification Tree Analysis) had a pooled sensitivity of 68.0% (95% CI: 50.6%–85.4%, I 2 = 95.6%), ranging from 48.9% [22] to 97.6% [23]. High heterogeneity, driven by diverse methodologies, suggests unreliable pooling, and clinicians should interpret these models cautiously. Specificity for traditional ML was 60.6% (95% CI: 40.9%–80.3%, I 2 = 99.6%), varying widely from 27.1% [23] to 95.0% [35], reflecting inconsistent performance due to differences in sample sizes and algorithms. This variability advises physicians to combine traditional ML predictions with clinical tools like PESI for low‐risk patients. In contrast, advanced ML models (Random Forest, KNN, SVM, Neural Networks, Boosting models) showed superior sensitivity of 89.7% (95% CI: 80.6%–98.9%, I 2 = 91.2%), ranging from 74.1% [26] to 97.4% [31], with acceptable heterogeneity for pooling. Specificity was robust at 95.1% (95% CI: 94.3%–95.9%, I 2 = 0%), ranging from 34.9% [19] to 96.0% [27], with zero heterogeneity, though Aujesky et al. [19] was an outlier. These findings suggest advanced ML models are highly reliable for identifying both high‐ and low‐risk PE patients, offering a practical advantage in critical care settings. USA‐based studies (five studies) had a sensitivity of 78.5% (95% CI: 64.2%–92.8%, I 2 = 95.3%), ranging from 48.9% [22] to 95.7% [21], with high heterogeneity indicating mixed performance across traditional and advanced ML models. Specificity was 72.5% (95% CI: 61.5%–83.5%, I 2 = 99.3%), ranging from 34.9% [19] to 93.0% [26], with extreme heterogeneity driven by outliers. Physicians in USA settings should thus use ML predictions cautiously, integrating them with clinical assessments. Non‐USA studies (seven studies) showed a high sensitivity of 97.2% (95% CI: 92.8%–100.0%, I 2 = 43.5%), ranging from 75.0% [34] to 97.6% [23], with moderate heterogeneity supporting reliable pooling. Specificity was 62.0% (95% CI: 36.8%–87.2%, I 2 = 99.9%), ranging from 27.1% [23] to 96.0% [27], with extreme heterogeneity precluding pooling. Non‐USA studies offer robust sensitivity for detecting high‐risk patients, but variable specificity suggests clinicians verify negative predictions. Advanced ML and non‐USA sensitivity provide the most reliable estimates, while traditional ML, USA, and non‐USA specificity require cautious interpretation due to outliers (e.g., Aujesky et al. [19]; Lau et al. [23]). For physicians, advanced ML models are preferable for consistent risk stratification across diverse populations, while traditional ML and USA‐based predictions benefit from supplementary clinical evaluation to ensure reliability.
3.4. Publication Bias
Funnel plot (Figure 3) analysis for sensitivity showed slight asymmetry, with most studies clustering at high sensitivity (74%–97.6%) and Parast et al. [22] (sensitivity = 48.9%) as an outlier. The scarcity of low‐sensitivity studies at high standard error suggests potential bias toward publishing higher‐sensitivity results, but Egger's test was nonsignificant (p = 0.942), indicating no substantial publication bias. This supports confidence in ML models' ability to detect high‐risk PE patients, allowing physicians to rely on these tools for identifying patients needing urgent intervention. For specificity, the funnel plot revealed marked asymmetry, with values ranging from 27.1% [23] to 95.9%, and outliers [19, 23] skewing the distribution. Egger's test (p = 0.053) approached significance, suggesting possible underreporting of low‐specificity studies, particularly in traditional ML and USA subgroups, aligning with their high heterogeneity (I 2 = 99.6% and 99.3%, respectively). Physicians should exercise caution when interpreting negative ML predictions, especially from traditional ML models or USA studies, as low‐specificity results may be underrepresented. Cross‐referencing with clinical scores like PESI can enhance decision‐making reliability in these contexts.
Figure 3.

Funnel and Galbraith plot for sensitivity and specificity based on the best‐performing models.
3.5. Quality Assessment
The PROBAST + AI assessment revealed that 12 of the 17 studies had Low Overall ROB for both Development and Evaluation, indicating robust methodological quality in the majority of the included studies. These studies demonstrated strong participant selection, well‐defined predictors, clear outcome measurements, and comprehensive analyses, particularly in reporting performance metrics and validation methods. Four studies—Aujesky et al. [19, 20, 21] and Sadegh‐Zadeh et al. [29]—exhibited High Overall ROB due to incomplete reporting of performance metrics (“not good” rating) and inadequate validation methods (e.g., derivation‐based or unclear validation), potentially inflating performance estimates. Liu et al. [24] also had High ROB, primarily due to a small sample size (n = 88), increasing selection bias risk. Cahan et al. [28] was rated Unclear for Overall ROB due to a missing sample size, though its analyses were robust. All 17 studies were rated Low for Applicability across all domains, confirming their relevance to the review's focus on ML models for PE mortality prediction (Figure 4).
Figure 4.

Quality assessment based on PROBAST + AI tool.
3.6. GRADE Assessment
The evidence, derived from 17 retrospective cohort studies, started with a Low certainty rating due to the observational nature of the data, which relied on electronic health records, clinical variables, laboratory tests, and imaging features (e.g., sPESI, CT‐based radiomics, vital signs). A serious risk of bias was identified across most studies, as many lacked robust external validation (e.g., only three studies explicitly reported external validation), and issues such as overfitting, inadequate reporting of model calibration, or selection bias in retrospective cohorts were common, warranting a downgrade by one level. Inconsistency was not serious, as AUROC values were relatively consistent, ranging from 0.70 to 0.963 across models (e.g., XGBoost, Random Forest, SVM, Neural Networks), though variability in model performance (e.g., sensitivity 0.49–1.0, specificity 0.27–0.955) suggested some heterogeneity in predictive ability. The evidence was direct, addressing the target population (PE patients), intervention (ML‐based prediction), and outcome (mortality). However, imprecision was a serious concern, as several studies reported wide or unreported confidence intervals for AUROC, and small sample sizes or low event rates limited precision, leading to a further downgrade by one level. No upgrading factors were applied, as large effect sizes or dose–response gradients are not typically relevant for predictive model outcomes, and no studies demonstrated consistent confounding adjustments that would increase certainty. Consequently, the final certainty of evidence for the accuracy of ML models in predicting mortality in PE was rated as Low, indicating limited confidence in the reported AUROC estimates. The true predictive performance may differ substantially from reported results, particularly given the lack of external validation and variable model generalizability.
4. Discussion
This systematic review and meta‐analysis of 17 studies, encompassing 844071 patients, demonstrated that ML models offer robust performance in predicting mortality, with a pooled sensitivity of 88% (95% CI: 78%–94%), specificity of 79% (95% CI: 62%–89%), DOR of 26 (95% CI: 10–71), and AUROC of 0.91 (95% CI: 0.88–0.93). Advanced ML models (Neural Networks, XGBoost) outperformed traditional ML models (Logistic Regression, CART), achieving higher sensitivity (89.7% vs. 68.0%) and consistent specificity (95.1%, I 2 = 0%). Non‐USA studies showed superior sensitivity (97.2%) compared to USA studies (78.5%), though specificity varied widely due to heterogeneity (I 2 = 99%). These findings suggest ML models surpass traditional risk scores like the Pulmonary Embolism Severity Index (PESI) in discriminative ability, offering clinicians a powerful tool for risk stratification in PE management. However, significant heterogeneity, particularly in specificity, and potential publication bias for low‐specificity studies warrant cautious interpretation.
The included studies demonstrated diverse ML approaches, with Logistic Regression being the most common, followed by advanced models like Random Forests, SVM, XGBoost, and Neural Networks. Aujesky and colleagues reported high sensitivity (97%–100%) but low specificity (23%–35%) using a clinical prediction rule with 10 nonlaboratory criteria, identifying low‐risk patients with 0%–1.5% 30‐day mortality [20]. In another study, Aujesky and colleagues achieved a high negative predictive value (98%–99%) and AUC of 0.87 in external validation, but specificity remained limited, aligning with our finding of lower specificity in traditional models (60.6%, I 2 = 99.6%) [21]. In contrast, advanced ML models showed superior performance. For instance, Mora and colleagues reported a Neural Network model with an AUC of 0.96, sensitivity of 96%, and specificity of 96%, outperforming Logistic Regression (AUC 0.76) [27]. Likewise, Cahan and colleagues developed a multimodal CNN‐based model achieving an AUC of 0.964, with 90% sensitivity and 94% specificity, significantly surpassing PESI and sPESI [28]. These findings align with our subgroup analysis, where advanced ML models exhibited consistent specificity (I 2 = 0%), suggesting greater reliability for clinical use compared to traditional models.
Comparing specific studies, Parast and colleagues used adaptive LASSO regression, achieving an AUC of 0.741 for 30‐day mortality, which was modest compared to advanced ML models [22] like Danilatou and colleagues, who reported an AUC of 0.93 for early mortality using a Random Forest classifier in ICU patients [26]. Danilatou and colleagues' model outperformed traditional scores like SAPS II (AUC 0.85), consistent with our meta‐analysis's AUROC of 0.91 for ML models [25]. Lau and colleagues enhanced the sPESI by incorporating sodium and bicarbonate levels, improving the AUC from 0.71 to 0.86, highlighting the value of integrating biomarkers with clinical scores [23]. Similarly, Liu and colleagues reported an SVM model with an AUC of 0.886, sensitivity of 90.2%, and specificity of 64.0%, outperforming PESI and sPESI, with warfarin use identified as a key predictor via SHAP analysis [32]. These studies underscore the ability of ML models to leverage diverse data (biomarkers, imaging, EHR features) to enhance predictive accuracy over traditional scores, supporting our finding of superior discriminative ability (AUROC 0.91).
Related studies in the field reinforce these findings. For example, Wang and colleagues used a deep learning model integrating CTPA imaging and clinical data, reporting an AUC of 0.94 and sensitivity of 91% [30], aligning with Cahan et al. [28] and our advanced ML subgroup results. However, Liu and colleagues focused on saddle PE and found that imaging biomarkers (RVa/LVa ratio) and cardiac troponin I predicted mortality with AUCs of 0.86 and 0.78, respectively, suggesting that combining such biomarkers with ML could further enhance performance [24].
The high heterogeneity in our meta‐analysis, particularly for specificity (I 2 = 99%), reflects differences in algorithms, study populations, and validation methods. Traditional ML models and USA‐based studies showed greater variability (I 2 = 99.6% and 99.3%, respectively), likely due to diverse patient characteristics (higher cancer prevalence in Parast et al. [22]) and inconsistent validation approaches (cross‐validation vs. derivation‐based). Advanced ML models and non‐USA studies, with lower heterogeneity (I 2 = 0% for specificity in advanced ML), offer more consistent performance, likely due to standardized data sets and robust algorithms. The retrospective nature of most studies introduces selection bias, limiting generalizability to prospective settings. Publication bias for specificity (Egger's p = 0.038) suggests underreporting of low‐specificity studies, particularly in traditional ML, which may inflate performance estimates.
For clinicians, these findings highlight the potential of ML models, especially advanced ones, to improve risk stratification in PE patients. Models like those by Mora and colleagues and Cahan and colleagues can prioritize high‐risk patients for aggressive interventions and identify low‐risk patients for outpatient management, reducing unnecessary hospitalizations. However, the variability in traditional ML models and USA studies suggests clinicians should integrate ML predictions with clinical scores like PESI, particularly when specificity is low. Future research should focus on prospective studies, standardized validation protocols, and integration of multimodal data to reduce heterogeneity and enhance generalizability.
Compared to traditional risk stratification models, such as PESI, sPESI, SAPS II, and others evaluated in recent studies, ML models demonstrate notable advancements in predictive accuracy, generalizability, and the ability to handle complex data patterns. Traditional models, like those reported by Liu and Ren, rely on fixed clinical variables and achieved moderate AUCs (SAPS II: 0.835; PESI: 0.702) in predicting 28‐day mortality, with SAPS II showing a nonlinear relationship and a critical threshold at 33 [36]. While effective, these models are limited by their reliance on predefined variables and linear assumptions, which may miss intricate interactions in heterogeneous PE populations. The novelty of ML models lies in their capacity to adaptively learn from complex, high‐dimensional data sets, unlike traditional scores that rely on static thresholds. For example, Magdy and colleagues reported a PESI AUC of 0.923 with 85.2% sensitivity and 100% specificity, but this was in a small cohort (n = 60), limiting generalizability [37]. Similarly, the PESI‐Echo score by the CONAREC XX registry improved upon PESI (AUC 0.75) by adding echocardiographic variables (AUC 0.82), yet it remained constrained by manual variable selection [38]. In contrast, ML models automatically identify key predictors using techniques like SHAP analysis, reducing bias and enhancing predictive power. Traditional models also struggle with handling complex patterns in PE cohorts with comorbidities, such as cancer or HF. Cantu‐Martinez and colleagues identified elevated creatinine, troponin, and sPESI as predictors of adverse outcomes, but their LASSO model did not significantly outperform standard regression (p = 0.11), suggesting limited improvement over traditional methods [39]. Similarly, Avci and colleagues found ECG and imaging markers (AUC 0.59) had poor prognostic value, highlighting the limitations of single‐modality predictors [40]. In contrast, ML models excel at integrating multimodal data, such as CTPA findings, laboratory values, and clinical scores, to uncover patterns unaddressed by traditional models. Generalizability is a key advantage of ML models, particularly advanced ones, as demonstrated by their performance across diverse geographic regions in our analysis. Traditional models, like the AMAPI by Zuin and colleagues (AUC 0.83 in stable patients), showed high sensitivity (91.6%) but low specificity (35.7%), limiting their utility in ruling out low‐risk patients. The EPHIPANY Index for cancer patients with PE achieved an AUC of 0.779 but relied on a decision tree with fixed criteria, less adaptable than ML models which maintained good performance in external validation [41]. Akhoundi and colleagues highlighted the predictive value of RV/LV ratio (AUC 0.744) for 30‐day mortality, but its utility was confined to imaging‐based cohorts [42]. ML models integrate such biomarkers with clinical data, enhancing generalizability across heterogeneous populations, including those with cancer or hemodynamic instability. The high heterogeneity in our meta‐analysis (I 2 = 99% for specificity) reflects challenges in traditional ML and USA‐based studies, driven by diverse algorithms and patient characteristics. Traditional models also face variability due to reliance on specific cohorts. ML models mitigate this through adaptive learning, as seen in Danilatou et al. [26], where feature selection reduced 2300 variables to 25 while maintaining high performance (AUC 0.93). This adaptability, coupled with tools like Liu and colleagues' web app [32], enhances clinical applicability, enabling real‐time risk stratification in diverse settings. However, retrospective study designs and potential specificity bias (Egger's p = 0.038) in our analysis suggest the need for prospective validation to confirm ML's advantages.
5. Limitations
This systematic review and meta‐analysis are constrained by several limitations. First, significant heterogeneity was observed, particularly in specificity (I 2 = 99%) and to a lesser extent in sensitivity (I 2 = 90.43%), driven by differences in algorithm types, study populations, and validation methods, especially in traditional ML and USA‐based studies. This heterogeneity complicates the reliability of pooled estimates, particularly for specificity, necessitating narrative synthesis in some subgroups. Second, the retrospective nature of most included studies introduces potential biases, such as selection bias and incomplete data, which may limit the generalizability of findings to prospective clinical settings. Additionally, the variability in validation approaches (e.g., cross‐validation, bootstrap, or unspecified methods) and the lack of uniform reporting on key metrics across studies further challenge the consistency of results. Potential publication bias, particularly for specificity (Egger's p = 0.038 after outlier exclusion), suggests that studies with lower specificity may be underrepresented, potentially inflating performance estimates. These limitations highlight the need for standardized prospective studies to validate ML models and reduce heterogeneity in future research.
6. Conclusion
This systematic review and meta‐analysis show that ML models outperform traditional scoring systems like PESI and sPESI in predicting 30‐day mortality in PE patients. With pooled sensitivity of 0.88, specificity of 0.79, and AUROC of 0.91, ML algorithms, including Random Forest, XGBoost, and Neural Networks, demonstrate strong discriminative ability. Advanced ML models achieve higher sensitivity (0.897) and consistent specificity (I 2 = 0%), leveraging complex data patterns for reliable risk stratification across diverse populations. These findings support integrating ML tools into clinical practice to enhance PE risk assessment and improve outcomes. Future research should focus on prospective validation of ML models, external testing in diverse ICU settings, and integration with real‐time clinical decision support systems to optimize personalized management and reduce mortality.
Author Contributions
Pooya Eini was a main contributor in the design, implementation, and writing of the manuscript. Mohammad Rezayee, Homa Serpoush, and Peyman Eini independently assessed articles and extracted data independently. All authors read and approved the final manuscript. Mohammad Rezayee and Jason Tremblay performed statistical analysis.
Ethics Statement
The authors have nothing to report.
Consent
The authors have nothing to report.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Supplementary Table 1: Search strategy for different databases.
Acknowledgments
In the preparation of this article, the authors utilized the Grammarly application to enhance linguistic accuracy and clarity. The manuscript underwent meticulous double‐checking to ensure precision, and the authors assume full responsibility for the integrity and originality of the content presented herein. The authors received no specific funding for this work.
Data Availability Statement
Data sharing is not applicable to this article as no new data were created or analyzed in this study.
References
- 1. Zhu J. L., Yuan S. Q., Wei X. Y., et al., “Development and Validation of a Simple Nomogram for Predicting the Short‐Term Prognosis of Patients With Pulmonary Embolism,” Heart & Lung 57 (2023): 144–151, 10.1016/j.hrtlng.2022.09.010. [DOI] [PubMed] [Google Scholar]
- 2. Wendelboe A. M. and Raskob G. E., “Global Burden of Thrombosis: Epidemiologic Aspects,” Circulation Research 118, no. 9 (2016): 1340–1347, 10.1161/circresaha.115.306841. [DOI] [PubMed] [Google Scholar]
- 3. Schulman S., Lindmarker P., Holmström M., et al., “Post‐Thrombotic Syndrome, Recurrence, and Death 10 Years After the First Episode of Venous Thromboembolism Treated With Warfarin for 6 Weeks or 6 Months,” Journal of Thrombosis and Haemostasis 4, no. 4 (2006): 734–742, 10.1111/j.1538-7836.2006.01795.x. [DOI] [PubMed] [Google Scholar]
- 4. Wang G., Liu T., Ji W., et al., “Prolonged Elevated Heart Rate Is Association With Adverse Outcome in Severe Pulmonary Embolism: A Retrospective Study,” International Journal of Cardiology 417 (2024): 132581, 10.1016/j.ijcard.2024.132581. [DOI] [PubMed] [Google Scholar]
- 5. Tzourtzos I., Lakkas L., and Katsouras C. S., “Right Ventricular Longitudinal Strain‐Related Indices in Acute Pulmonary Embolism,” Medicina 60, no. 10 (2024): 1586, 10.3390/medicina60101586. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Zhan Y. and Che X., “A Prognostic Prediction Model for Acute Pulmonary Embolism,” Journal of Investigative Medicine 72, no. 8 (2024): 930–937, 10.1177/10815589241283739. [DOI] [PubMed] [Google Scholar]
- 7. Wadhwa V., Gutta N. B., Trivedi P. S., et al., “In‐Hospital Mortality Benefit of Inferior Vena Cava Filters in Patients With Pulmonary Embolism and Congestive Heart Failure,” American Journal of Roentgenology 211, no. 3 (2018): 672–676, 10.2214/ajr.17.19332. [DOI] [PubMed] [Google Scholar]
- 8. Murgier M., Bertoletti L., Darmon M., et al., “Frequency and Prognostic Impact of Acute Kidney Injury in Patients With Acute Pulmonary Embolism. Data From the RIETE Registry,” International Journal of Cardiology 291 (2019): 121–126, 10.1016/j.ijcard.2019.04.083. [DOI] [PubMed] [Google Scholar]
- 9. Winterton D., Bailey M., Pilcher D., Landoni G., and Bellomo R., “Characteristics, Incidence and Outcome of Patients Admitted to Intensive Care Because of Pulmonary Embolism,” Respirology 22, no. 2 (2017): 329–337, 10.1111/resp.12881. [DOI] [PubMed] [Google Scholar]
- 10. Khemasuwan D., Yingchoncharoen T., Tunsupon P., et al., “Right Ventricular Echocardiographic Parameters Are Associated With Mortality After Acute Pulmonary Embolism,” Journal of the American Society of Echocardiography 28, no. 3 (2015): 355–362, 10.1016/j.echo.2014.11.012. [DOI] [PubMed] [Google Scholar]
- 11. Jiménez D., Aujesky D., Díaz G., et al., “Prognostic Significance of Deep Vein Thrombosis in Patients Presenting With Acute Symptomatic Pulmonary Embolism,” American Journal of Respiratory and Critical Care Medicine 181, no. 9 (2010): 983–991, 10.1164/rccm.200908-1204OC. [DOI] [PubMed] [Google Scholar]
- 12. Jiménez D., Kopecna D., Tapson V., et al., “Derivation and Validation of Multimarker Prognostication for Normotensive Patients With Acute Symptomatic Pulmonary Embolism,” American Journal of Respiratory and Critical Care Medicine 189, no. 6 (2014): 718–726, 10.1164/rccm.201311-2040OC. [DOI] [PubMed] [Google Scholar]
- 13. Gulshan V., Peng L., Coram M., et al., “Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs,” Journal of the American Medical Association 316, no. 22 (2016): 2402–2410, 10.1001/jama.2016.17216. [DOI] [PubMed] [Google Scholar]
- 14. Willan J., Katz H., and Keeling D., “The Use of Artificial Neural Network Analysis Can Improve the Risk‐Stratification of Patients Presenting With Suspected Deep Vein Thrombosis,” British Journal of Haematology 185, no. 2 (2019): 289–296, 10.1111/bjh.15780. [DOI] [PubMed] [Google Scholar]
- 15. Sajjadi S. M., Mohebbi A., Ehsani A., et al., “Identifying Abdominal Aortic Aneurysm Size and Presence Using Natural Language Processing of Radiology Reports: A Systematic Review and Meta‐Analysis,” Abdominal Radiology 50 (2025): 3885–3899, 10.1007/s00261-025-04810-5. [DOI] [PubMed] [Google Scholar]
- 16. Page M. J., McKenzie J. E., Bossuyt P. M., et al., “The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,” BMJ 372 (2021): n71, 10.1136/bmj.n71. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Moons K. G. M., Damen J. A. A., Kaul T., et al., “PROBAST + AI: An Updated Quality, Risk of Bias, and Applicability Assessment Tool for Prediction Models Using Regression or Artificial Intelligence Methods,” BMJ 388 (2025): e082505, 10.1136/bmj-2024-082505. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Guyatt G. H., Oxman A. D., Vist G. E., et al., “GRADE: An Emerging Consensus on Rating Quality of Evidence and Strength of Recommendations,” BMJ 336, no. 7650 (2008): 924–926, 10.1136/bmj.39489.470347.AD. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Aujesky D., Obrosky D. S., Stone R. A., et al., “Derivation and Validation of a Prognostic Model for Pulmonary Embolism,” American Journal of Respiratory and Critical Care Medicine 172, no. 8 (2005): 1041–1046, 10.1164/rccm.200506-862OC. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Aujesky D., Obrosky D. S., Stone R. A., et al., “A Prediction Rule to Identify Low‐Risk Patients With Pulmonary Embolism,” Archives of Internal Medicine 166, no. 2 (2006): 169–175, 10.1001/archinte.166.2.169. [DOI] [PubMed] [Google Scholar]
- 21. Aujesky D., Roy P. M., Le Manach C. P., et al., “Validation of a Model to Predict Adverse Outcomes in Patients With Pulmonary Embolism,” European Heart Journal 27, no. 4 (2006): 476–481, 10.1093/eurheartj/ehi588. [DOI] [PubMed] [Google Scholar]
- 22. Parast L., Cai B., Bedayat A., et al., “Statistical Methods for Predicting Mortality in Patients Diagnosed With Acute Pulmonary Embolism,” Academic Radiology 19, no. 12 (2012): 1465–1473, 10.1016/j.acra.2012.09.008. [DOI] [PubMed] [Google Scholar]
- 23. Lau J. K. E., Chow V., Brown A., Kritharides L., and Ng A. C. C., “A Validated Improved Risk Prediction Model for In‐Hospital Death During Acute Presentation With Pulmonary Embolism,” European Heart Journal 37 (2016): 951–952. [Google Scholar]
- 24. Liu M., Miao R., Guo X., et al., “Saddle Pulmonary Embolism: Laboratory and Computed Tomographic Pulmonary Angiographic Findings to Predict Short‐Term Mortality,” Heart, Lung and Circulation 26, no. 2 (2017): 134–142, 10.1016/j.hlc.2016.02.019. [DOI] [PubMed] [Google Scholar]
- 25. Danilatou V., Antonakaki D., Tzagkarakis C., et al., “Automated Mortality Prediction in Critically‐Ill Patients With Thrombosis Using Machine Learning,” in 2020 IEEE 20th International Conference on Bioinformatics and Bioengineering (BIBE) (Bournemouth Univ, Fac Sci & Technol, Bournemouth, Dorset, England Venizeleio Hosp Heraklion, Iraklion, Greece Fdn Res & Technol Hellas FORTH, Inst Comp Sci, Iraklion, Greece, 2020): 247–254.
- 26. Danilatou V., Nikolakakis S., Antonakaki D., et al., “Outcome Prediction in Critically‐Ill Patients With Venous Thromboembolism and/or Cancer Using Machine Learning Algorithms: External Validation and Comparison With Scoring Systems,” International Journal of Molecular Sciences 23, no. 13 (2022): 7132, 10.3390/ijms23137132. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Mora D., Nieto J. A., Mateo J., et al., “Machine Learning to Predict Outcomes in Patients With Acute Pulmonary Embolism Who Prematurely Discontinued Anticoagulant Therapy,” Thrombosis and Haemostasis 122, no. 4 (2022): 570–577, 10.1055/a-1525-7220. [DOI] [PubMed] [Google Scholar]
- 28. Cahan N., Klang E., Marom E. M., et al., “Multimodal Fusion Models for Pulmonary Embolism Mortality Prediction,” Scientific Reports 13, no. 1 (2023): 7544, 10.1038/s41598-023-34303-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Sadegh‐Zadeh S.‐A., Sakha H., Movahedi S., et al., “Advancing Prognostic Precision in Pulmonary Embolism: A Clinical and Laboratory‐Based Artificial Intelligence Approach for Enhanced Early Mortality Risk Stratification,” Computers in Biology and Medicine 167 (2023): 107696, 10.1016/j.compbiomed.2023.107696. [DOI] [PubMed] [Google Scholar]
- 30. Wang G., Xu J., Lin X., et al., “Machine Learning‐Based Models for Predicting Mortality and Acute Kidney Injury in Critical Pulmonary Embolism,” BMC Cardiovascular Disorders 23, no. 1 (2023): 385, 10.1186/s12872-023-03363-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Tur Y., Cicek V., Cinar T., et al., “Mortality Prediction of Pulmonary Embolism Patients With Deep Learning and XGBoost,” in 2024 4th International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), (Institute of Electrical and Electronics Engineers Inc., 2024), 1–6.
- 32. Liu J., Li R., Yao T., et al., “Interpretable Machine Learning Approach for Predicting 30‐Day Mortality of Critical Ill Patients With Pulmonary Embolism and Heart Failure: A Retrospective Study,” Clinical and Applied Thrombosis/Hemostasis 30 (2024), 10.1177/10760296241304764. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Mesinovic M., Wong X. C., Rajahram G. S., et al., “At‐Admission Prediction of Mortality and Pulmonary Embolism in an International Cohort of Hospitalised Patients With COVID‐19 Using Statistical and Machine Learning Methods,” Scientific Reports 14, no. 1 (2024): 16387, 10.1038/s41598-024-63212-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Shahzadi I., Zwanenburg A., Frohwein L. J., et al., “Short‐Term Mortality Prediction in Acute Pulmonary Embolism: Radiomics Values of Skeletal Muscle and Intramuscular Adipose Tissue,” Journal of Cachexia, Sarcopenia and Muscle 15, no. 4 (2024): 1430–1440, 10.1002/jcsm.13488. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Teodoru M., Negrea M. O., Cozgarea A., and Cozma D., “Enhancing Pulmonary Embolism Mortality Risk Stratification Using Machine Learning: The Role of the Neutrophil‐to‐Lymphocyte Ratio,” Journal of Clinical Medicine 13, no. 5 (2024): 1191, 10.3390/jcm13051191. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Liu P. and Ren Y., “Prognostic Value of SAPS II Score for 28‐Day Mortality in ICU Patients With Acute Pulmonary Embolism,” International Journal of Cardiology 430 (2025): 133201, 10.1016/j.ijcard.2025.133201. [DOI] [PubMed] [Google Scholar]
- 37. Magdy D. M., Salama S., Abdelraheem N. S., and Mahmoud S. R., “Performance of Pulmonary Embolism Risk Scores in Predicting Mortality in Patients With Acute Pulmonary Embolism,” Egyptian Journal of Chest Diseases and Tuberculosis 74, no. 1 (2025): 77–84, 10.4103/ecdt.ecdt_69_24. [DOI] [Google Scholar]
- 38. Burgos L. M., Scatularo C. E., Cigalini I. M., et al., “The Addition of Echocardiographic Parameters to PESI Risk Score Improves Mortality Prediction in Patients With Acute Pulmonary Embolism: PESI‐Echo Score,” European Heart Journal: Acute Cardiovascular Care 10, no. 3 (2020): 250–257, 10.1093/ehjacc/zuaa007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Cantu‐Martinez O., Martinez Manzano J. M., Tito S., et al., “Clinical Features and Risk Factors of Adverse Clinical Outcomes in Central Pulmonary Embolism Using Machine Learning Analysis,” Respiratory Medicine 215 (2023): 107295, 10.1016/j.rmed.2023.107295. [DOI] [PubMed] [Google Scholar]
- 40. Avci S., Perincek G., and Karakayali M., “Prediction of Mortality Associated With Cardiac and Radiological Findings in Patients With Pulmonary Embolism,” Journal of Cardiovascular Emergencies 6, no. 4 (2020): 84–90, 10.2478/jce-2020-0020. [DOI] [Google Scholar]
- 41. Zuin M., Rigatelli G., Picariello C., Carraro M., Zonzin P., and Roncon L., “Prognostic Role of a New Risk Index for the Prediction of 30‐Day Cardiovascular Mortality in Patients With Acute Pulmonary Embolism: The Age‐Mean Arterial Pressure Index (AMAPI),” Heart and Vessels 32, no. 12 (2017): 1478–1487, 10.1007/s00380-017-1012-5. [DOI] [PubMed] [Google Scholar]
- 42. Akhoundi N., Langroudi T. F., Rajebi H., et al., “Computed Tomography Pulmonary Angiography for Acute Pulmonary Embolism: Prediction of Adverse Outcomes and 90‐Day Mortality in a Single Test,” Polish Journal of Radiology 84 (2019): 436–446, 10.5114/pjr.2019.89896. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Table 1: Search strategy for different databases.
Data Availability Statement
Data sharing is not applicable to this article as no new data were created or analyzed in this study.
