Abstract
Background
Accurate prediction of survival in oncology can guide targeted interventions. The traditional regression-based Cox proportional hazards (CPH) model has statistical assumptions and may have limited predictive accuracy. With the capability to model large datasets, machine learning (ML) holds the potential to improve the prediction of time-to-event outcomes, such as cancer survival outcomes. The present study aimed to systematically summarize the use of ML models for cancer survival outcomes in observational studies and to compare the performance of ML models with CPH models.
Methods
We systematically searched PubMed, MEDLINE (via EBSCO), and Embase for studies that evaluated ML models vs. CPH models for cancer survival outcomes. The use of ML algorithms was summarized, and either the area under the curve (AUC) or the concordance index (C-index) for the ML and CPH models were presented descriptively. Only studies that provided a measure of discrimination, i.e., AUC or C-index, and 95% confidence interval (CI) were included in the final meta-analysis. A random-effects model was used to compare the predictive performance in the pooled AUC or C-index estimates between ML and CPH models using R. The quality of the studies was evaluated using available checklists. Multiple sensitivity analyses were performed.
Results
A total of 21 studies were included for systematic review and 7 for meta-analysis. Across the 21 articles, diverse ML models were used, including random survival forest (N=16, 76.19%), gradient boosting (N=5, 23.81%), and deep learning (N=8, 38.09%). In predicting cancer survival outcomes, ML models showed no superior performance over CPH regression. The standardized mean difference in AUC or C-index was 0.01 (95% CI: -0.01 to 0.03). Results from the sensitivity analyses confirmed the robustness of the main findings.
Conclusions
ML models had similar performance compared with CPH models in predicting cancer survival outcomes. Although this systematic review highlights the promising use of ML to improve the quality of care in oncology, findings from this review also suggest opportunities to improve ML reporting transparency. Future systematic reviews should focus on the comparative performance between specific ML models and CPH regression in time-to-event outcomes in specific type of cancer or other disease areas.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12874-025-02694-z.
Keywords: Machine learning, Cox proportional hazards model, Survival analysis, Cancer, Real-world data
Background
Cancer poses a significant burden to society in the US and worldwide. Cancer remains as the leading cause of death globally, accounting for about 10 million deaths in 2019 [1, 2]. In the US, an estimated 1.8 million new cancer cases and 0.6 million deaths due to cancer were reported in 2019 [3, 4]. It is projected that globally, cases of newly diagnosed cancer could reach up to 29.5 million by 2040, with the number of cancer-attributed deaths increasing to 16.4 million [4]. Recent economic modeling data show that the projected global economic burden of cancer in 2050 will reach as high as $25.2 trillion [5]. Furthermore, 2019 data from the National Cancer Institute indicate that the economic burden for oncology care in the US could reach as high as $21.1 billion, considering the costs of patient out-of-pocket and time spent for cancer care [6].
Overall survival, defined as the time from treatment initiation to death, is often regarded as the gold-standard patient-centered outcome in oncology studies [7]. Accurately predicting survival outcomes in oncology not only serves as a tool to evaluate patients’ prognostic factors development but also improves physicians’ decision-making by informing the effectiveness of cancer treatment strategies [8, 9]. Therefore, prediction of survival outcomes for patients with cancer is critically important for evaluating improvement in cancer care.
In survival analyses, the most common analytical approach is the Cox proportional hazards (CPH) regression model [10]. However, the CPH model was designed for smaller datasets and its major limitation is that it does not adaptively handle high-dimensional large datasets well [11]. In addition, if the assumption of proportional hazards fails, the CPH regression model could lead to biased estimation [11]. The proportional hazards assumption pertains to the constancy of the influence of different variables on survival over time, with their cumulative effect being proportional within a defined scale [12].
Machine learning (ML) is a branch of artificial intelligence that relies on algorithms to automatically learn patterns of data, especially when using complex, high-dimensional, and heterogeneous datasets, and has wide applications in health services research [13–15]. Over time, ML methods have been gradually adapted to predict time-to-event outcomes, such as survival outcomes in oncology [16–21]. For example, using the Netherlands Cancer Registry that included 36,658 breast cancer patients, Moncada-Torres et al. used several ML-based methods to predict long-term survival outcomes [17]. Their research showed that although some ML-based models (i.e., random survival forest (RSF), survival support vector machines (SVM)) had similar performance compared with classical CPH models, extreme gradient boosting (GB) performed better in predicting survival [17]. To date, systemic reviews across different clinical domains, such as heart failure, hypertension, atherosclerotic cardiovascular disease, and pregnancy care, have compared the use of ML algorithms with traditional statistical models in predicting disease development [22–25]. However, most of these reviews compared ML against traditional logistic regression by focusing on the prediction of binary outcomes (e.g., deceased vs. survived; disease vs. no disease) [22–25]. For example, a systematic review and meta-analysis by Liu et al., showed superior performance of ML vs. traditional approaches in predicting the atherosclerotic cardiovascular risk prognostication in primary care settings by analyzing 15 observational studies [23]. In another meta-analysis of 32 studies, Chowdhury et al. found similar performance between ML models vs. traditional regression models for predicting hypertension risk in practice [22]. Focusing on cancer outcomes, several reviews have summarized the use of ML in cancer diagnosis, prognosis, or treatment selection; however, none of them specifically focused on synthesizing models to predict mortality risk, a common time-to-event outcome in cancer [26–31]. In addition, these reviews have certain limitations, such as focusing only on specific types of cancer [28, 29] or not performing a quantitative comparison of ML models vs. traditional CPH regression [26, 27, 30, 31]. To summarize, while most available reviews compared ML against traditional regression-based approaches, such as logistic regression (by focusing on the prediction of binary outcomes), none of these prior reviews have specifically focused on comparing models to predict time-to-event outcomes such as mortality risk in cancer patients. Clearly, a knowledge gap still exists about predictive models using survival analysis that leverage ML techniques in predicting survival outcomes in cancer patients.
This systematic review and meta-analysis aimed to (1) systematically identify ML-based risk prediction models for survival analysis in cancer patients and (2) evaluate the performance of ML-based risk prediction models compared with the traditional CPH model for predicting survival outcomes.
Methods
This systematic review and meta-analysis was conducted in alignment with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement as well as the Meta-analysis Of Observational Studies in Epidemiology (MOOSE) reporting guidelines [32, 33]. eTable 1 of the Supplementary Materials (SM) presents the PRISMA checklist. Also, this study followed a published protocol by Debray et al. that provides a flowchart for systematically reviewing validation studies of prediction models [34]. Institutional Review Board approval was not required for this study because summary data from published articles were used (i.e., the authors did not have access to patient-level data or identifiers).
Table 1.
Characteristics of studies assessing cancer survival risk using ML and CPH models
| Author and year | Data source and study year | Sample range | Type of cancer conditions | Inclusion criteria | Exclusion criteria | % Female | Age, mean | Definition of outcome predicted |
|---|---|---|---|---|---|---|---|---|
| Cai et al. 2022 [59] | SEER database from 2004 to 2016 | 59,046 | Non-small cell lung cancer (NSCLC) | (1) Pathologically diagnosed as NSCLC; (2) Without lymph node or other organ metastasis; (3) Diagnosed as stage Tx-4 (the 8th TNM staging system) | (1) Age < 18 years old; (2) Previous or concurrent other cancers; (3) Received neoadjuvant radiotherapy; (4) Survival time < 1 month; (5) No active follow-up; (6) Grade unknown; (7) Location unknown. | Occult NSCLC 45.9%; Other NSCLC 51.4% | Occult NSCLC 71.9 (10.5); Other NSCLC 68.7 (10.1) | Overall survival and cancer specific survival |
| Dong et al. 2022 [54] | SEER database from 1973 to 2014 | 13,952 | Cervical cancer | (1) Patients diagnosed with stage IA to stage IIB cervical cancer confirmed by histopathology according to the FIGO staging system 2019. | (1) Patients with multiple tumors; (2) Patients not having surgery; (3) Patients not having follow-up data. | Not reported | 43 (13) | Overall survival |
| Fan et al. 2022 [55] | SEER database from 1975 to 2016 | 741 | Spinal Ewing’s sarcoma (EWS) | (1) Patients with histologically confirmed EWS; (2) Having active follow-up | (1) patients with “missing/unknown COD” or “N/A not first tumor” SEER cause-specific death classification | 36.60% | 19.85 (12.15) | Overall survival and cancer specific survival |
| Huang et al. 2019 [56] | SEER database from 2005 to 2014 | 1000 | Osteosarcoma | (1) Patients histologically diagnosed with osteosarcoma | (1) Patients with tumor size code 000/888/989–999 (unknown or inaccurate) and extent codes 00 (in situ) or 99 (unknown extent); (2) Patients not diagnosed with osteosarcoma by biopsy; (3) Patients at N0 or M0 stage; (4) Patients without primary site surgery; (5) Patients with non-primary osteosarcoma; (6) patients with missing data | 47% | 25.3 | Overall survival and cause-specific survival |
| Jin et al. 2023 [57] | SEER database from 2012 to 2017 | 4617 | Stage-al3 non-small cell lung cancer | (1), patients pathologically diagnosed with primary stage-iii NSCLC; (2) The existence of one malignant lesion. | (1) Unknown or missing regional lymph nodes | Not reported | Not reported | Overall survival |
| Kim et al. 2022 [58] | SEER database and Korea Tumor Registry System-Biliary Pancreas (KOTUS-BP) from 2004 to 2016 | 9624 from SEER and 3281 from KOTUS-BP | Resected non-metastatic pancreatic ductal adenocarcinoma | Not reported | Not reported | SEER cohort 49.4%; KOTUS cohort 42.1% | SEER cohort 65.6 (10.4); KOTUS cohort 63.8 (10.1) | Overall survival |
| Li et al. 2021 [28] | VA Cancer Registry System and pharmacy dispensation records from the VA Corporate Data Warehouse from 2006 to 2014 | 523 | Follicular lymphoma | (1) Patients with patients with grade 1–3a, stage II–IV FL receiving rituximab combined with cyclophosphamide, vincristine, and prednisone ± doxorubicin or bendamustine combined with rituximab | (1) Patients who received maintenance therapy; (2) Patients without a hematology/oncology visit within 6 months of diagnosis; (3) Patients with another malignancy prior to FL diagnosis | Separated based on cohort | Separated based on cohort | Overall survival |
| Li et al. 2022 [59] | SEER database from 2010 to 2019 | 15,129 | Bone metastatic breast cancer | (1) Patients with cancer evidence with morphological and histopathology diagnosis; (3) Patients with bone metastases at initial diagnosis. | (1) Patients with more than one primary cancer | 100% | 61.39 | Overall survival |
| Lin et al. 2022 [59] | SEER Researcher Plus Database (Nov 2020 Sub) | 3988 | Pancreatic cancer | (1) Patients with the variable “site and morphology code” = ‘Pancreas’; (2) Patients with the behavior recode for malignancy | (1) Age < 18 years old; (2) Patients with unavailable data on cause of death or follow-up survival months; (3) Patients lacking clinical information | 49% in training and 50% in testing data | 65 | Cancer-specific survival |
| Liu et al. 2022 [59] | SEER database from 2004–2018 | 156,154 | Melanoma | (1) Patients with “Site and Morphology.Site recode” = “Melanoma of the Skin” | (1) patients < 18; (2) Patients without a positive histology | 42.73% | 59.23 (15.68) | Overall survival |
| Loureiro et al. 2021 [71] | (1) the nationwide Flatiron Health HER; (2) OAK clinical trial database | 136,719 | 18 different primary cancers | Not reported | Not reported | 50% in training, 50.5% in internal testing, and 37.9% in OAK | 66.47 (10.98) in training, 66.47 (11.05) in internal testing, and 62.79 (9.57) in external testing | Overall survival |
| Ouyang et el. 2022 [62] | SEER database from 2004–2016 | 4636 | Cervical cancer | (1) Patients with hysterectomy and radical hysterectomy; (2) Patients without node or distant metastasis; (3) Patients with cervical adenocarcinoma as the first primary tumor; (4) Patients having histology codes for adenomas and adenocarcinomas; (5) Patients with complete follow-up data | (1) Age < 20 years old; (2) Patients who died within 2 months; (3) Patients undergoing local excision and/or conization, excisional biopsy, amputation of the cervix or those undergoing laser therapy; (4) Patients not having cervical cancer as primary cancer site; (5) Patients with a confirmed diagnosis by autopsy or death | Not reported | Not reported | Overall survival |
| Park et al. 2022 [58] | SEER database from 2000–2017 | 322,348 | Breast cancer | Not reported | (1) patients with missing information | 100.00% | 59.1 (13.0) | Overall survival |
| Senders et al. 2020 [64] | SEER database from 2005 to 2015 | 20,821 | Glioblastoma multiforme | (1) Patients with confirmed diagnosis | (1) Patients who died in the direct postoperative period | Not reported | Not reported | Overall survival |
| She et al. 2020 [42] | SEER database from 2010 to 2015 and Shanghai Pulmonary Hospital data from 2009 to 2013 | 17,322 | Non-small cell lung cancer | (1) Patients with pathologically confirmed primary stage I to IV NSCLC | (1) Patients with missing information | 48.4% in training, 49.2% in internal testing, and 45.7% in Shanghai data | 68 in training, 68 in internal testing, and 60 in external testing | Cancer-specific survival |
| Sun et al. 2023 [65] | SEER database from 2004 to 2015 | 8677 | Laryngeal squamous cell carcinoma | (1) Confirmed diagnosis; (2) Patients with surgical treatment with or without radiotherapy or chemotherapy | (1) Patients with not first primary tumor in the larynx; (2) Patients with metastasis by the time of diagnosis; (3) Patients with missing survival data or survival time less than 3 months | Not reported | Not reported | Overall survival |
| Teng et al. 2022 [66] | SEER database from 2010 to 2015 | 12,840 | Breast cancer | (1) Patients with NOS histology codes 8522/3 (Infiltrating duct and lobular carcinoma breast cancer); Female patients | (1) Patients with IV grade or UNK Stage; (2) Diagnosis age < 80; (3) Survival time < 1 month | 100% | 59.62 (12.1) | Overall survival |
| Tian et al. 2018 [67] | SEER database from 2004 to 2013 | 128,061 | Nonmetastatic colorectal cancer | (1) Patients diagnosed with primary colorectal cancer; (2) Patients with active followed-up with case reports from hospital inpatient departments, radiation treatment centers, laboratories, and physicians’ offices; (3) Patients with survival times longer than 1 month | (1) Patients with undetermined cell types; (2) Not stable or not treated patients; (3) Patients with metastases; (4) patients with unknown tumor size and number of positive regional nodes | 50.32% | 67 | Overall survival |
| Wang et al. 2021 [68] | West China Hospital from 2005 to 2018; SEER database | 1137 | IB-IIA stage non-small lung cancer | Not reported | (1) Age < 18 years old; (2) Patients with unknown survival time | 59% | Not reported | Overall survival |
| Yan et al. 2022 [38] | SEER database from 2000 to 2018 | 3145 | Chondrosarcoma | (1) Patients with confirmed diagnosis of chondrosarcoma; (2) Patients not having bones and joints as primary cancer site | 2) Patients with unknown survival time less than one month; (2) Patients not having chondrosarcoma as primary cancer site | 47% | 52 (18) | Overall survival |
| Yu et al. 2022 [60] | SEER database from 2004 to 2015 | 49,275 | Rectal adenocarcinoma | (1) Age > 20 years old | (1) Patients not having rectal adenocarcinoma as primary cancer site; (2) Patients with unknown tumor grade, survival time, race, marital status, or surgery status. | 40.10% | 62.6 (13.5) | Overall survival |
SD Standard deviation, SEER Surveillance, Epidemiology, and End Results, NSCLC Non-small cell lung cancer, EWS Spinal Ewing’s sarcoma, VA Veterans Affairs
Data sources and search strategy
Electronic databases - PubMed, Medline (via EBSCO), and Embase were systematically searched from database inception until May 3, 2023. Relevant articles included US-based studies that evaluated ML methods vs. the CPH model to predict cancer survival outcomes using large real-world data (RWD) and reported model performance using either the area under the curve (AUC) or the concordance index (C-index), which is also called C-statistic. Relevant databases were searched using keywords including ‘machine learning’, ‘Cox proportional hazards model’, ‘real-world dataset’, ‘cancer’, and ‘mortality’. For these searches, Medical Subject Heading (MeSH) terms and keywords were used together to obtain eligible studies. Details of the search strategies are available in eTable 2 of the SM.
Study selection
We included studies that evaluated ML methods and the CPH model for predicting cancer survival outcomes using observational data, and reported AUC or C-index values comparing ML models with the CPH model. All original studies were included if they: (1) applied at least one ML technique for time-to-event survival prediction; (2) reported model performance using AUC or C-index; (3) evaluated the traditional regression-based CPH model; and (4) analyzed patient-level US real-world databases, e.g., administrative claims, electronic medical records (EMR), or cancer registry data. We focused on original studies that compared ML with the CPH model using RWD for cancer survival outcomes and excluded studies that evaluated binary outcomes, used unsupervised learning, or primarily evaluated feature selection. We also excluded randomized controlled trials (RCTs), simulation studies, and studies using imaging data. Systematic reviews and conference abstracts were also excluded. Initial citations searched from the above-mentioned electronic databases were imported into the Covidence platform [35]. Titles and abstracts were initially screened by several authors of this team (SB, JL, ML, YH). Then, for full article eligibility screening, articles available in the complete paper were retrieved. Three members of this team (SB, JL, ML) were involved in the full-text eligibility review, and other members (YH, KB) were involved in solving any disagreements. Details of the review protocol, including the PICOS/PICOTS table, and complete inclusion or exclusion criteria for article screening, are available in eTable 3 of the SM.
Data extraction
Data extraction was performed using a standardized extraction form in Microsoft Excel spreadsheets by three authors (JL, ML, SB). First, we extracted information on the study characteristics: (1) author (year of publication), (2) real-world datasets used, (3) study design and follow-up duration, (4) participants’ ages and gender, (5) predictor variables, and (6) definition of the survival outcome predicted. Second, we extracted items related to characteristics of the ML model and CPH model: (1) events/sample size, (2) type of ML algorithms utilized, (3) ML model description, (4) ML model performance, (5) ML model validation methods, and (6) traditional CPH model performance. Specifically, regarding model performance, either AUC or C-index was extracted as both of them are popular evaluation metrics for evaluating model discrimination in the context of survival prediction models [36]. In survival models, time-dependent AUC evaluates the model’s discriminative ability at multiple time points, and in addition to the dynamic AUC, an integrated AUC (iAUC), is considered a summary measure of various time-dependent AUC values over the follow-up period. The C-index measures how well a survival model ranks survival times of all possible pairs of subjects. As per Park et al., the iAUC can be considered similar to the C-index [36]. In this study, we extracted either the iAUC or the C-index as the performance measure. Of note, if a study reported multiple AUCs or C-indices for training and testing datasets for a model, model performance measures in the external datasets were considered as the first level of hierarchy for extraction with reasons as follows. In an ideal situation for ML modeling, the dataset is partitioned into training, internal validation, and then independent datasets are used for external validation (if they exist). While training refers to fitting the dataset to obtain a model specification, validation is used to evaluate a model’s performance based on another randomly split dataset; external validation, of note, is performed to evaluate the model performance using an independent dataset, sourcing from an independent population other than the dataset used for the original training [37].
Data synthesis and analysis
We summarized the types of ML algorithms used in the included studies. Note that detailed summaries for some of the commonly used ML models as well as their pros and cons are presented in eTable 4 of the SM. We also illustrate the performance of each ML algorithm by plotting the AUC or C-index for each ML algorithm. When a study utilized multiple ML models based on the same algorithm, we selected the best-performing model to include in our summary. For example, in the scenario of the same ML algorithm being used to produce multiple models, which differed by a set of input features, only the model that generated the highest predictive performance with respect to its C-index or AUC was selected for further analysis. Furthermore, some studies applied multiple ML algorithms for survival prediction. In comparing the performance of ML models with the traditional CPH model for the meta-analysis, the best-performing ML model was selected, regardless of the type of algorithm used. For example, for the Yan (2022) study, we selected the DeepSurv model —a type of neural network (NN)— because it was the best performing ML model with a C-index of 0.832 as compared with the NMLTR model (another NN type; C-index: 0.821) and the RSF model (C-index: 0.803) [38]. This practice is consistent with a published meta-analysis that focused on comparing the traditional regression-based approach against ML models in terms of model performance [39]. The R packages “beanplot” and “beeswarm” were used to draw the box plot and beeswarm plot [40, 41].
Meta-analysis
Next, we performed a meta-analysis on the predictive performance of ML models compared with the traditional CPH model in estimating cancer survival outcomes. Studies meeting the following three criteria were eligible for the meta-analysis: (1) assessing overall survival outcome; (2) reporting AUC or C-index and 95% confidence interval (CI) for the ML method; (3) reporting AUC or C-index and 95% CI for the CPH model. For example, She et al. was ineligible for the meta-analysis, since cancer-specific survival was assessed [42]. For each eligible study, we used the best-performing ML model for inclusion in the meta-analysis. Using the same approach, for each study, the best-performing CPH model was included in the meta-analysis. A meta-analysis of model performance comparing the ML algorithms against that of traditional CPH was performed. As described earlier, either AUC or C-index was used to evaluate the model’s predictive performance.
We generated forest plots to present a summary AUC or C-index along with their 95% CI prediction interval for ML-based models and CPH models. A random-effects model was used to compare the pooled AUCs or C-index values and 95% CI between the ML algorithms and the CPH model. The heterogeneity of these studies (e.g., differences in setting, study population, and methodology) was evaluated using Cochran’s Q test, the τ2 statistic, and the Higgins I2 statistic [43, 44]. Cochran’s Q tests the null hypothesis of the equality of true estimates of the effects estimated from all studies included. A Cochran’s Q test with a p value < 0.05 suggests the presence of statistically significant heterogeneity. The heterogeneity statistic τ2 quantifies the between-study variance of all existing studies, and Higgins I2 represents the fraction of total variance in the estimated effects due to between-study variability. Then, the source of statistically significant heterogeneity was categorized as low, moderate, and high, corresponding to the I2 values ≤ 25%, between 25 and 75% and ≥ 75%, respectively. We employed the R programming language for our analysis, installing and loading the necessary packages: “meta”, “metasens”, and “readxl” [45–47]. Data were imported from an Excel file containing studies and their results. Specifically using the “meta” package, we conducted a meta-analysis using the “metacont” function to compare the AUC or C-index between ML methods and the CPH model. Results were visualized with a forest plot, sorted by the study year and model performance.
We also conducted several sets of sensitivity analyses. First, we performed the leave-one-out method, conducting the meta-analysis sequentially by removing one study at a time to test the impact of each study on the overall effect size [48]. In addition, to explore the impact of heterogeneity due to a study and inform the robustness and stability of the results, another sensitivity analysis was performed using the cumulative sensitivity method over the 7 included studies [39]. Thirdly, given that specific ML techniques influence the variation of ML model performance, another sensitivity analysis was conducted by considering multiple ML models developed as per each unique ML technique in one study. In this regard, a total of 17 models developed across 7 included articles were included for meta-analysis. For example, if a study developed more than 1 type of ML technique, all ML methods as per the unique ML algorithm classifier were included. The following four types of sub-group analyses were conducted: tree-based, kernel-based, deep learning (DL), and NN algorithms by categorizing the ML techniques as per the algorithm classification. Overall, this sensitivity analysis helps in testing the stability of the findings and ensuring the findings are not influenced by the variability in survival prediction caused by a specific ML performance. Furthermore, we also performed additional leave-one-out and cumulative sensitivity methods over these 17 ML models, respectively. All sensitivity analyses of findings were evaluated via a leave-one-out method as well as a cumulative method using “metainf” and “metacum” packages in R respectively [45, 46]. Finally, we performed a sensitivity analysis by conducting the analysis with only studies that reported the C-index (i.e., not AUC) to test the robustness of the findings. In this sensitivity analysis, we first considered only the best-performing ML model from each study. Subsequently, we expanded the analysis to include all ML models developed within each study, similar to the third sensitivity analysis.
Quality evaluation
In this review, the quality assessment of these studies focused on a critical evaluation and the reporting of ML models. A number of reporting guidelines for prediction models in health care are available, including those employing ML. To operationalize the checklist we used for the evaluation of study quality, we adapted items from the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) statement [49] and the Luo checklist [50]. The TRIPOD statement originally contained 22 items (with sub-items), relevant for evaluating the transparent reporting of a prediction model-related study. The Luo checklist was designed to assess the quality of reporting of predictive models in medical research that use machine learning methods and includes 5 broad domains containing 12 items (with sub-items). The Luo checklist has been used by prior ML-related systematic reviews [51, 52].
We included 23 items across 12 domains in our assessment tool considering the TRIPOD statement and the Luo checklist. Each item was marked as “yes”, “partly”, “no”, “unclear”. The overall quality of an individual study was defined as “high”, “moderate”, “low” based on the overall results of domains of individual items. We considered a study as high quality if 16 or more individual items (corresponding to total number of items for 8 domains out of 12 domains) were marked “yes” or “partly”. The study was considered moderate quality if 12–15 items were marked “yes” or “partly” (corresponding to total number of items for 6 domains out of 12 domains); otherwise, the study was marked as low quality if only 11 or fewer items were marked “yes” or “partly”. The checklist items were evaluated for each study by three authors (JL, ML, SB). In cases of any discrepancy in the assessment of study quality, a fourth author (YH) ensured the comprehensiveness of data and quality.
Results
Study identification and selection
Figure 1 shows the PRISMA flowchart of the included studies. As shown in Fig. 1, we identified a total of 324 citations by searching relevant databases (160 MEDLINE, 107 Embase, 57 PubMed). After removing 54 duplicates, there remained 270 studies for title/abstract screening. Further, after title and abstract screening, 208 irrelevant studies were excluded, and 62 full-text articles remained for review. Three authors read the full text of these publications, which led to the exclusion of 41 studies for the reasons shown in Fig. 1. Thus, a total of 21 studies were included in this systematic review.
Fig. 1.
PRISMA diagram
Study characteristics
Table 1 shows the characteristics of all included studies. These US-based studies were carried out using several different data sources to develop ML-based survival models. Surveillance, Epidemiology, and End Results (SEER)-Medicare registry datasets (N = 19, 90.48%) [38, 42, 53–69] were the most commonly used data source, followed by other cancer registry databases including Veterans Affairs (N = 1, 4.76%) [70], and EMR (N = 1, 4.76%) [71]. The median sample size from these included studies was 12,905, ranging from 523 to 322,348.
Different types of cancer diagnoses were evaluated in the included studies. Of included studies, lung cancer (N = 4, 19.05%) [42, 53, 57, 68] was the most common cancer type, followed by breast cancer (N = 3, 14.28%) [59, 63, 66], primary malignant bone neoplasm (N = 3, 14.28%) [38, 55, 56], cervical cancer (N = 2, 9.52%) [54, 62], pancreatic cancer (N = 2, 9.52%) [58, 60], gastrointestinal cancer (N = 2, 9.52%) [67, 69], and other cancer diagnoses, such as central nervous system, lymphoma, melanoma, laryngeal cancer (N = 5, 23.81%) [61, 64, 65, 70, 71]. Of these time-to-event outcomes, most studies evaluated an overall survival outcome (N = 20, 95.24%) [38, 53–59, 61–72], while other studies also evaluated cancer-specific survival outcomes (N = 5, 23.81%) [42, 53, 55, 56, 60].
Use of ML algorithms
Table 2 summarizes the use of ML models across these studies. The most popular ML algorithms used were RSF (N = 16, 76.19%) [38, 54–62, 64–67, 70, 71], GB (N = 5, 23.81%) [59, 63, 64, 66, 71], NN (N = 8, 38.09%) [38, 42, 54, 57, 60, 68, 69, 71], SVM (N = 3, 14.28%) [58, 59, 63], and others including least absolute shrinkage and selection operator (LASSO), CoxBoost, ID3, K-nearest neighbor (KNN), autoencoder, super learner, conditional survival forest, survival tree, recursive partitioning, and time-varying transformation (N = 8, 38.10%) [53, 55, 59, 62–64, 66, 71]. Across these studies, only a few studies applied external validation methods (N = 6, 28.57%) [42, 57, 58, 66, 68, 71]. Details regarding the ML and CPH models used in the included studies are shown in eTable 5 of the SM.
Table 2.
ML algorithms used in the studies and featuring studies (n = 21 studies)
| Type of ML Algorithms | Number of Studies [a] | Featuring studies |
|---|---|---|
| Tree-based Methods | ||
| Random survival forests | 16 | [38, 54–62, 64–67, 70, 71] |
| Gradient boosting | 5 | [59, 63, 64, 66, 71] |
| Neural Networks | 8 | [38, 42, 54, 57, 60, 68, 69, 71] |
| Support vector machine | 3 | [58, 59, 63] |
| Others [b] | 8 | [53, 55, 59, 62–64, 66, 71] |
[a] Most studies have applied more than 1 machine learning algorithms, therefore the sum of the number of studies by machine learning method is greater than the number of included studies
[b] includes least absolute shrinkage and selection operator (LASSO), CoxBoost, ID3, K-nearest neighbor (KNN), autoencoder, super learner, conditional survival forest, survival tree, recursive partitioning, and time-varying transformation
Overall, 42 ML models were evaluated over 21 studies, with an average AUC or C-index of 0.759 and a median AUC or C-index of 0.751 (interquartile range (IQR): 0.121; range: 0.337). A large proportion of these studies (N = 17, 80.95%) had ML models with AUCs equal to 0.7 or above. The mean (standard deviation) [median] AUC or C-index for RSF, GB, NN, SVM, and other ML algorithms was 0.766 (0.085) [0.753], 0.760 (0.056) [0.757], 0.770 (0.097) [0.754], 0.732 (0.088) [0.684], and 0.746 (0.072) [0.723], respectively. Figure 2 shows the performance (AUC or C-index) for each ML algorithm.
Fig. 2.
Performance of different machine learning models used for cancer survival analysis
Among all articles, only 7 articles that reported either AUC or C-index and 95% CI for both ML and CPH methods for overall survival prediction were included for meta-analysis [38, 64–67, 70, 71]. The performance of these 7 ML models was evaluated relative to the CPH model. Figure 3 shows the forest plot of ML vs. CPH of the 7 included studies along with the 95% prediction interval. As depicted in Fig. 3, the meta-analysis revealed a similar predictive accuracy between the CPH model and ML techniques. The standardized mean difference (SMD) in the AUC or C-index was 0.01, with a 95% CI ranging from − 0.01 to 0.03, highlighting the similarities in performance between the two approaches.
Fig. 3.
Forest plot comparing ML and CPH model with 95% prediction interval
Six sets of sensitivity analyses were undertaken to evaluate the robustness of our primary findings. The results of these sensitivity analyses are shown in eTable 6 A-6E of the SM and confirmed the robustness of the main finding, that is, similar model performance of ML models and CPH regression for prediction of overall survival in patients with cancer. In the first and second sensitivity analyses, the leave-one-out method and cumulative method were performed on 7 included articles, both of them showed similar performance of ML and CPH regression. In the third sensitivity analysis, where 17 ML methods over 7 studies were included, the result also did not display a meaningful difference in the AUC or C-index between ML and CPH regression, yielding an SMD in AUC or C-index of 0.02 (95% CI: −0.01 to 0.05). Notably, a pronounced source of heterogeneity was discerned in the predictive efficacy of these models, as quantified by an I2 value of 97%, meaning 97% of heterogeneity is due to between-study variability. The τ2 value of 0.0025 further substantiates this heterogeneity, which was statistically significant (p < 0.01). The fourth and fifth sensitivity analyses using the leave-one-out and cumulative sensitivity analyses on 17 ML models of 7 studies, further corroborated with our primary results and other sensitivity analyses. Lastly, based on the sensitivity analysis only including studies that reported the C-index, we found that the results observed in the primary analysis hold (i.e., no difference between ML and CPH models judged by the C-index).
Quality assessment of included studies
Almost all studies (95.24%) were of high quality with a low risk of bias based on the modified Luo checklist and TRIPOD statement. A few studies failed to report how to handle missing data. Several studies did not provide CIs for AUC or C-index, making them ineligible for meta-analysis. Descriptions of the quality assessment items are provided in eTable 7 of the SM. Our quality assessments for all included studies are available in eTable 8 of the SM.
Discussion
This systematic review and meta-analysis evaluated published studies to provide a comprehensive and rigorous appraisal of the ML literature on the prediction of survival for cancer patients. Although there is a growing interest in the application of ML, a systematic understanding of ML in predicting survival in cancer patients remains understudied. To the best of our knowledge, this is the first systematic review and meta-analysis to evaluate the performance of ML algorithms compared with the traditional regression-based CPH model based on structured data in the prediction of survival outcomes among US patients with cancer. We found that RSF was the most commonly applied ML model, reflecting its preference as the ML algorithm for survival analysis. The ML algorithms under investigation for survival analysis have good predictive capabilities, with a median AUC or C-index of 0.751. Specifically, 17 of these studies had ML models with AUCs or C-index values of 0.7 or higher. However, the meta-analysis showed that these ML models did not outperform the traditional regression-based CPH model.
This systematic review and meta-analysis adds to a growing body of literature on the use of ML for cancer survival prediction [26–31]. A previous systematic review of 13 studies summarized the common ML algorithms, including random forests, SVM, ensemble learning, and DL in predicting cervical cancer survival [29]. Another review focusing on breast cancer 5-year survival also concluded that the most common ML algorithms in this area included tree-based approaches, NN-based DL, SVM, and ensemble learning [28]. However, no other studies have evaluated the pooled performance of ML algorithms compared with the traditional CPH model. In the present pooled analysis of 7 studies, we could not demonstrate an increased improvement with ML algorithms compared with CPH regression in risk prediction for survival outcomes after cancer diagnosis.
In addition, the most popular ML algorithms for survival analysis were RSF, GB, and NN. The results of this review regarding the popularity of RSF and NN in survival prediction in cancer are in alignment with two prior reviews [28, 29]. Rahimi et al. were limited to systematically evaluating studies on cervical cancer survival and concluded that RSF, survival support vector, and deep learning methods as the popular ML algorithms in this field [29]. Specifically for breast cancer survival, Li et al. systematically identified 31 studies and found that the most popular ML methods were tree-based methods and artificial neural networks for breast cancer survival prediction [28]. RSF, as a nonparametric tree-based ensemble method, is built to evaluate censored time-to-event data [73]. The RSF method is advantageous at dealing with noisy survival data with multiple covariates and complex, nonlinear relationships; therefore, it fills a role in being used as an alternative to CPH regression when the proportional hazards assumption is in question [74]. To date, RSF has drawn more attention to its application to time-to-event data. Recent research has explored the use of RSF for prediction of survival or risk of complications in the context of liver transplantation and other surgeries [75, 76]. In addition, the advantages of the NN methods for survival prediction are its adaptability to intricate and nonlinear associations through multiple neuron-like processing layers and its increased flexibility in handling complex clinical variables [77, 78]. There are also examples of published research using NN for survival prediction in other clinical contexts. Katzman et al. developed a NN-based ML models for survival prediction, demonstrating that applying NN techniques can enhance treatment recommendations and subsequently optimize survival outcomes [79].
Furthermore, our findings suggest methodological aspects of ML that need improvement in order to help ensure the applicability of ML models for overall survival risk prediction to better inform clinical decision-making and early palliative care planning in oncology. In a prior review for stroke outcomes, Wang et al. noted that despite a surge in the use of ML for stroke outcomes, only a few studies satisfied the required reporting standards, such as inclusion of diverse rich datasets for model development, detailed reporting of development process, and external validation [80]. In another review for vascular surgery, Li et al. noted that a majority of those studies had high risk of bias with failure to follow standard reporting [81]. According to our systematic review and meta-analysis, the majority of the included studies lacked external validation and only used the same datasets for model validation. This potentially means that these models may have limited use when applied in other real-life settings as they may demonstrate limited predictive accuracy in a different dataset. In three ML-specific systematic reviews, the lack of external validation across many studies using ML models was also noted [80–82]. To guarantee the application of an ML model in clinical practice, enabling the ML model to be replicated in other datasets beyond the sample cohort, future ML researchers should consider the external validation process, which can increase the generalizability of the model, and facilitate its use in clinical decision-making. Future studies examining ML algorithms using RWD should consider following standardized reporting guidelines, such as the TRIPOD statement extended for artificial intelligence and machine learning (TRIPOD-AI), for improving the quality and utility of ML algorithms in clinical practice [83].
Importantly, the quality of the underlying RWD, regardless of ML algorithms used, influences model performance, and researchers need to evaluate data sources fully to understand the variables included as well as their prognostic or predictive value [84–86]. Of note, in the present review, the RWD underlying most of the ML studies included is SEER-Medicare data. Although not the focus of the ML quality evaluation by our review, researchers should be aware of the limitations of the SEER-Medicare data, e.g., laboratory and imaging test results are not captured. Additionally, while the SEER registry data provides reliable information on cancer stage at diagnosis, the use of SEER-Medicare data limits the scope of the analysis to older adults with cancer, limiting generalizability to other populations, such as those who are younger or commercially insured. Future researchers may consider other augmented linked data sources, such as administrative claims linked with EMR data, as such datasets contain EMR-based prognostic variables from patients’ medical visits. The use of EMR-claims linked data, which contain more granular clinical predictors and may improve prediction accuracy and increase clinical utility, is recommended.
Overall, this meta-analysis answered an important question regarding the comparative performance of ML algorithms vs. CPH regression in the context of predicting time-to-event survival outcomes for cancer patients. This meta-analysis did not compare model performance of ML algorithms vs. CPH regression across the specific type of ML algorithms due to the small number of studies across each specific ML algorithm. Thus, future systematic reviews could evaluate the comparative performance between the specific type of ML and the CPH model in time-to-event outcomes in cancer. Also, due to a limited number of studies, this review was unable to conduct a comparison of ML vs. CPH regression in subgroups of specific cancer types. Such model comparisons in time-to-event outcomes prediction in specific cancer areas would be helpful. Also, future research should address the heterogeneity in the time horizon (i.e., follow-up period) among included studies that can impact the results of the meta-analysis. Comparing the performance of ML methods and CPH regression in other disease areas outside of cancer is another promising path to follow.
The findings only focused on the comparison of ML and CPH regression for predicting cancer-specific survival in terms of AUC/C-index, but this does not necessarily fully address the clinical effectiveness of these models. For example, the CPH model, a commonly used statistical approach in cancer survival analysis, is particularly useful for its interpretability and ability to provide a clear explanation of the risk associated with each predictor variable (e.g., age, disease, performance status, cancer type, severity of side effects) on cancer survival time. From the perspective of cancer care in clinical settings, this information is of critical importance, as clinicians can rely on this information to make informed decisions about personalized treatment strategies. Findings from the present meta-analysis only reflect comparisons of predictive accuracy between ML and CPH regression from a mathematical perspective using the AUC/C-index. Accordingly, considerations for choosing between ML and traditional CPH models should extend beyond their comparative predictive accuracy to include the clinical utility of these models and their potential for generating actionable insights to enhance personalized patient care. Ultimately, the goal in cancer therapy is not only to build a tool with the best predictive performance for survival outcomes but, more importantly, to assist clinicians in improving medical decisions and informing suitable treatment strategies.
Furthermore, it is evident that the similar performance of ML and CPH models identified by this review can only reflect their comparable performance in terms of AUC/C-index (i.e., model discrimination), and this finding has several implications for future research. First, assessing the calibration of both ML and CPH models is needed to inform the clinical utility of these models. As such, future research can contribute new evidence on whether ML models outperform the CPH model with regard to generating risk estimates across different clinical risk strata. A poorly calibrated model, despite its high discrimination power, is still unreliable and has limited use for clinicians. Second, the performance of ML models is sensitive to the feature selection methods. Prediction with ML models, even when using the same algorithms and datasets, may vary when different feature selection methods are applied [87, 88]. Therefore, future research should also consider the reporting of the feature selection approach utilized in ML models and, additionally, continue to investigate feature importance captured by these ML models, as compared with the results derived from the established CPH model. This may help provide valuable information on whether ML models are able to offer additional insights beyond the established CPH model in identifying novel prognostic factors.
Given our finding that ML models may not necessarily provide better prediction than the CPH model for cancer survival outcomes, it is important to note some implications for clinicians and researchers in this area. With the growing use of ML models in disease prognosis and clinical care, clinicians, such as oncologists and palliative care physicians, are becoming increasingly aware of the value of these tools in providing new insights into the complex relationships between prognostic factors and clinical outcomes. However, insights generated from these tools may lack external validity. To better inform clinical decision-making, future research should focus on improving ML models using hyperparameter tuning and optimizing feature selection [89, 90] based on expert clinical opinion.
Limitations
This study has some limitations. First, this systematic review and meta-analysis only searched English language publications in major electronic databases; therefore, ML studies that are not included in these electronic databases or not in the English-language are not captured. Moreover, this study focuses on the comparison of ML vs. CPH models in predicting survival for US patients with cancer. We acknowledge that many non-US studies have applied ML for the prediction of cancer survival and may potentially be included, but there are some inherent differences between US studies and non-US studies with respect to different underlying patient characteristics and healthcare systems, among other considerations. Nevertheless, future studies should consider comparing the performance of ML and traditional CPH models for predicting mortality in non-US settings. Second, we considered only ML publications for cancer survival using structured datasets. There may be studies that applied survival-based ML models for survival prediction using unstructured datasets, such as genetic, text, or imaging datasets, which are not included. Future studies are required to evaluate model performance in the prediction of cancer outcomes using ML methods (e.g., natural language processing) that are capable of data mining from unstructured data, such as text-based physicians’ notes. Third, as mentioned in the inclusion criteria, we focused solely on survival outcomes for cancer. We acknowledge that there are other ML studies that have evaluated time-to-event survival outcomes for non-cancer diseases and findings from the present meta-analysis only apply to comparisons of ML and CPH models for cancer patients. Fourth, as AUC and C-index are the commonly reported performance measures for these ML models, we chose to evaluate the comparative performance between ML vs. CPH models using AUC or C-index. We did not evaluate other performance metrics, such as F1 score, specificity, or sensitivity. It might be that judging on other performance measures, ML models might perform the same, better, or worse than the traditional CPH model in cancer survival prediction. This needs further investigation. ML studies may therefore include other methods as baseline/comparison models. Fifth, in the sensitivity analysis in which we included all developed unique ML algorithm models as per each study vs. CPH models for meta-analysis, our results might be biased in terms of standard errors, since the observations are not independent as assumed. Lastly, this review combines results of ML and CPH models across studies without restrictions as to the quality of the studies. We acknowledge that if there is a violation of the proportional hazards assumption in the Cox model, the validity of AUC or C-index values reported for such a model may be questionable.
Conclusions
We conducted a systematic review and meta-analysis to evaluate the performance of ML models vs. the traditional regression-based CPH model that predicts survival in cancer patients using RWD. We found that ML algorithms did not exhibit better performance as judged by AUC or C-index in survival prediction for cancer patients. Future systematic reviews and meta-analyses could evaluate the comparative performance between specific ML methods and the CPH model in time-to-event outcomes in specific cancer areas or other diseases.
Supplementary Information
Acknowledgements
Not applicable.
Abbreviations
- AUC
Area under the curve
- C-index
Concordance index
- CPH
Cox proportional hazards
- DL
Deep learning
- EMR
Electronic medical record
- GB
Gradient boosting
- iAUC
Integrated area under the curve
- IQR
Interquartile range
- KNN
K-nearest neighbor
- LASSO
Least absolute shrinkage and selection operator
- MeSH
Medical Subject Heading
- ML
Machine learning
- MOOSE
Meta-analysis Of Observational Studies in Epidemiology
- NN
Neural networks
- PRISMA
Preferred Reporting Items for Systematic Reviews and Meta-Analysis
- RCTs
Randomized controlled trials
- RSF
Random survival forest
- RWD
Real-world data
- SMD
Standardized mean difference
- SVM
Survival support vector machine
- TRIPOD
Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis
Authors’ contributions
YH, JL, ML, and KB contributed to the development and conceptualization of this review. YH performed the study search. YH, SB, JL, ML, and AA contributed to article screening, data extraction, and quality assessment. YH and AA conducted the meta-analysis with input from KB and JB. YH drafted the initial manuscript. All authors read, edited, and approved the final manuscript.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
The datasets supporting the conclusions of this article are available from the corresponding author upon reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.GBD 2017 Causes of Death Collaborators. Global, regional, and national age-sex-specific mortality for 282 causes of death in 195 countries and territories, 1980–2017: a systematic analysis for the global burden of disease study 2017. Lancet. 2018;392:1736–88. 10.1016/S0140-6736(18)32203-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.GBD 2019 Diseases and Injuries Collaborators. Global burden of 369 diseases and injuries in 204 countries and territories, 1990–2019: A systematic analysis for the global burden of disease study 2019. Lancet. 2020;396:1204–22. 10.1016/S0140-6736(20)30925-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.CDC. Cancer data and statistics. https://www.cdc.gov/cancer/dcpc/data/index.htm. Accessed 20 Aug 2023.
- 4.National Cancer Institute. Cancer statistics. https://www.cancer.gov/about-cancer/understanding/statistics. Accessed 20 Aug 2023.
- 5.Chen S, Cao Z, Prettner K, Kuhn M, Yang J, Jiao L, et al. Estimates and projections of the global economic cost of 29 cancers in 204 countries and territories from 2020 to 2050. JAMA Oncol. 2023;9:465–72. 10.1001/jamaoncol.2022.7826. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.National Cancer Institute. Annual report to the nation part 2: Patient economic burden of cancer care more than $21 billion in the United States in 2019. https://www.cancer.gov. Accessed 23 Sept 2023.
- 7.Delgado A, Guddati AK. Clinical endpoints in oncology - a primer. Am J Cancer Res. 2021;11:1121–31. [PMC free article] [PubMed] [Google Scholar]
- 8.de Freitas R, Nunes RD, Martins E, Curado MP, Freitas NMA, Soares LR, et al. Prognostic factors and overall survival of breast cancer in the City of Goiania, brazil: A population-based study. Rev Col Bras Cir. 2017;44:435–43. 10.1590/0100-69912017005003. [DOI] [PubMed] [Google Scholar]
- 9.Stone PC, Lund S. Predicting prognosis in patients with advanced cancer. Ann Oncol. 2007;18:971–6. 10.1093/annonc/mdl343. [DOI] [PubMed] [Google Scholar]
- 10.Cox DR. Regression models and life-tables. J R Stat Soc Series B Stat Methodol. 1972;34:187–202. 10.1111/j.2517-6161.1972.tb00899.x. [Google Scholar]
- 11.Klein JP, van Houwelingen HC, Ibrahim JG, Scheike TH, editors. Handbook of survival analysis. New York: Chapman and Hall/CRC; 2016. 10.1201/b16248. [Google Scholar]
- 12.Abd ElHafeez S, D’Arrigo G, Leonardis D, Fusaro M, Tripepi G, Roumeliotis S. Methods to analyze time-to-event data: the Cox regression analysis. Oxid Med Cell Longev. 2021;2021:1302811. 10.1155/2021/1302811. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Baştanlar Y, Ozuysal M. Introduction to machine learning. Methods Mol Biol. 2014;1107:105–28. 10.1007/978-1-62703-748-8_7. [DOI] [PubMed] [Google Scholar]
- 14.Doupe P, Faghmous J, Basu S. Machine learning for health services researchers. Value Health. 2019;22:808–15. 10.1016/j.jval.2019.02.012. [DOI] [PubMed] [Google Scholar]
- 15.Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380:1347–58. 10.1056/NEJMra1814259. [DOI] [PubMed] [Google Scholar]
- 16.Hazewinkel A-D, Gelderblom H, Fiocco M. Prediction models with survival data: A comparison between machine learning and the Cox proportional hazards model. 10.1101/2022.03.29.22273112
- 17.Moncada-Torres A, van Maaren MC, Hendriks MP, Siesling S, Geleijnse G. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci Rep. 2021;11:6968. 10.1038/s41598-021-86327-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Nunez J-J, Leung B, Ho C, Bates AT, Ng RT. Predicting the survival of patients with cancer from their initial oncology consultation document using natural language processing. JAMA Netw Open. 2023;6:e230813. 10.1001/jamanetworkopen.2023.0813. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Spooner A, Chen E, Sowmya A, Sachdev P, Kochan NA, Trollor J, et al. A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction. Sci Rep. 2020;10:20410. 10.1038/s41598-020-77220-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Xiao J, Mo M, Wang Z, Zhou C, Shen J, Yuan J, et al. The application and comparison of machine learning models for the prediction of breast cancer prognosis: retrospective cohort study. JMIR Med Inform. 2022;10:e33440. 10.2196/33440. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Gupta S, Tran T, Luo W, Phung D, Kennedy RL, Broad A, et al. Machine-learning prediction of cancer survival: a retrospective study using electronic administrative records and a cancer registry. BMJ Open. 2014;4:e004007. 10.1136/bmjopen-2013-004007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Chowdhury MZI, Naeem I, Quan H, Leung AA, Sikdar KC, O’Beirne M, et al. Prediction of hypertension using traditional regression and machine learning models: a systematic review and meta-analysis. PLoS ONE. 2022;17:e0266334. 10.1371/journal.pone.0266334. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Liu W, Laranjo L, Klimis H, Chiang J, Yue J, Marschner S, et al. Machine-learning versus traditional approaches for atherosclerotic cardiovascular risk prognostication in primary prevention cohorts: A systematic review and meta-analysis. Eur Heart J Qual Care Clin Outcomes. 2023;9:310–22. 10.1093/ehjqcco/qcad017. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Sufriyana H, Husnayain A, Chen Y-L, Kuo C-Y, Singh O, Yeh T-Y, et al. Comparison of multivariable logistic regression and other machine learning algorithms for prognostic prediction studies in pregnancy care: systematic review and meta-analysis. JMIR Med Inform. 2020;8:e16503. 10.2196/16503. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Sun Z, Dong W, Shi H, Ma H, Cheng L, Huang Z. Comparing machine learning models and statistical models for predicting heart failure events: a systematic review and meta-analysis. Front Cardiovasc Med. 2022;9:812276. 10.3389/fcvm.2022.812276. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Deepa p, Gunavathi C. A systematic review on machine learning and deep learning techniques in cancer survival prediction. Prog Biophys Mol Biol. 2022;174:62–71. 10.1016/j.pbiomolbio.2022.07.004. [DOI] [PubMed] [Google Scholar]
- 27.Kourou K, Exarchos TP, Exarchos KP, Karamouzis MV, Fotiadis DI. Machine learning applications in cancer prognosis and prediction. Comput Struct Biotechnol J. 2015;13:8–17. 10.1016/j.csbj.2014.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Li J, Zhou Z, Dong J, Fu Y, Li Y, Luan Z, et al. Predicting breast cancer 5-year survival using machine learning: a systematic review. PLoS ONE. 2021;16:e0250370. 10.1371/journal.pone.0250370. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Rahimi M, Akbari A, Asadi F, Emami H. Cervical cancer survival prediction by machine learning algorithms: a systematic review. BMC Cancer. 2023;23(1):341. 10.1186/s12885-023-10808-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Tran KA, Kondrashova O, Bradley A, Williams ED, Pearson JV, Waddell N. Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med. 2021;13:152. 10.1186/s13073-021-00968-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zhu W, Xie L, Han J, Guo X. The application of deep learning in cancer prognosis prediction. Cancers. 2020;12:603. 10.3390/cancers12030603. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Brooke BS, Schwartz TA, Pawlik TM. MOOSE reporting guidelines for Meta-analyses of observational studies. JAMA Surg. 2021;156:787–8. 10.1001/jamasurg.2021.0522. [DOI] [PubMed] [Google Scholar]
- 33.Moher D, Liberati A, Tetzlaff J, Altman DG, PRISMA Group. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Med. 2009;6:e1000097. 10.1371/journal.pmed.1000097. [PMC free article] [PubMed] [Google Scholar]
- 34.Debray TPA, Damen JAAG, Snell KIE, Ensor J, Hooft L, Reitsma JB, et al. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460. 10.1136/bmj.i6460. [DOI] [PubMed] [Google Scholar]
- 35.Covidence. May : Systematic review software, Veritas health innovation, Melbourne, Australia. https://www.covidence.org/. Accessed 3 2023.
- 36.Park SY, Park JE, Kim H, Park SH. Review of statistical methods for evaluating the performance of survival or other time-to-event prediction models (from conventional to deep learning approaches). Korean J Radiol. 2021;22:1697–707. 10.3348/kjr.2021.0223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Cabitza F, Campagner A, Soares F, García de Guadiana-Romualdo L, Challa F, Sulejmani A, et al. The importance of being external. Methodological insights for the external validation of machine learning models in medicine. Comput Methods Programs Biomed. 2021;208:106288. 10.1016/j.cmpb.2021.106288. [DOI] [PubMed] [Google Scholar]
- 38.Yan L, Gao N, Ai F, Zhao Y, Kang Y, Chen J, et al. Deep learning models for predicting the survival of patients with chondrosarcoma based on a surveillance, epidemiology, and end results analysis. Front Oncol. 2022;12:967758. 10.3389/fonc.2022.967758. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Talwar A, Lopez-Olivo MA, Huang Y, Ying L, Aparasu RR. Performance of advanced machine learning algorithms overlogistic regression in predicting hospital readmissions: A meta-analysis. Explor Res Clin Soc Pharm. 2023;11:100317. 10.1016/j.rcsop.2023.100317. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Kampstra P. Beanplot. A boxplot alternative for visual comparison of distributions. J Stat Softw. 2008;28:1–9. 10.18637/jss.v028.c01.27774042 [Google Scholar]
- 41.Eklund A, Beeswarm. The bee swarm plot, an alternative to stripchart. R package version 0.2. 2016;3. 10.32614/CRAN.package.beeswarm
- 42.She Y, Jin Z, Wu J, Deng J, Zhang L, Su H, et al. Development and validation of a deep learning model for non-small cell lung cancer survival. JAMA Netw Open. 2020;3:e205842. 10.1001/jamanetworkopen.2020.5842. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327:557–60. 10.1136/bmj.327.7414.557. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Deeks J, Higgins J, Altman D, Cochrane Statistical Methods Group. Analysing data and undertaking meta-analyses - Cochrane Handbook for Systematic Reviews of Interventions. 2019;10:241–84. 10.1002/9781119536604.ch10
- 45.Balduzzi S, Rücker G, Schwarzer G. How to perform a meta-analysis with R: a practical tutorial. Evid Based Ment Health. 2019;22:153–60. 10.1136/ebmental-2019-300117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Chen D-G (Din), Peace KE, editors. Applied meta-analysis with R and stata. 2nd edition. New York: Chapman and Hall/CRC; 2021. 10.1201/9780429061240
- 47.Schwarzer G. Meta: an R package for meta-analysis. R News. 2007;7:40–5. [Google Scholar]
- 48.Rappazzo KM, Nichols JL, Rice RB, Luben TJ. Ozone exposure during early pregnancy and preterm birth: a systematic review and meta-analysis. Environ Res. 2021;198:111317. 10.1016/j.envres.2021.111317. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMC Med. 2015;13:1. 10.1186/s12916-014-0241-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Luo W, Phung D, Tran T, Gupta S, Rana S, Karmakar C, et al. Guidelines for developing and reporting machine learning predictive models in biomedical research: a multidisciplinary view. J Med Internet Res. 2016;18:e323. 10.2196/jmir.5870. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Brnabic A, Hess LM. Systematic literature review of machine learning methods used in the analysis of real-world data for patient-provider decision making. BMC Med Inform Decis Mak. 2021;21:54. 10.1186/s12911-021-01403-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Volpe S, Pepa M, Zaffaroni M, Bellerba F, Santamaria R, Marvaso G, et al. Machine learning for head and neck cancer: a safe bet?-A clinically oriented systematic review for the radiation oncologist. Front Oncol. 2021;11:772663. 10.3389/fonc.2021.772663. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Cai J, Yang F, Wang X. Occult non-small cell lung cancer: an underappreciated disease. J Clin Med. 2022;11:1399. 10.3390/jcm11051399. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Dong T, Wang L, Li R, Liu Q, Xu Y, Wei Y, et al. Development of a novel deep learning-based prediction model for the prognosis of operable cervical cancer. Comput Math Methods Med. 2022;2022:4364663. 10.1155/2022/4364663. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Fan G, Yang S, Qin J, Huang L, Li Y, Liu H, et al. Machine learning predict survivals of spinal and pelvic Ewing’s sarcoma with the SEER database. Glob Spine J. 2022;14:1125–36. 10.1177/21925682221134049 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Huang R, Xian S, Shi T, Yan P, Hu P, Yin H, et al. Evaluating and predicting the probability of death in patients with non-metastatic osteosarcoma: a population-based study. Med Sci Monit. 2019;25:4675–90. 10.12659/MSM.915418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Jin L, Zhao Q, Fu S, Cao F, Hou B, Ma J. Development and validation of machine learning models to predict survival of patients with resected stage-III NSCLC. Front Oncol. 2023;13:1092478. 10.3389/fonc.2023.1092478. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Kim H, Park T, Jang J, Lee S. Comparison of survival prediction models for pancreatic cancer: Cox model versus machine learning models. Genomics Inform. 2022;20:e23. 10.5808/gi.22036. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Li C, Liu M, Li J, Wang W, Feng C, Cai Y, et al. Machine learning predicts the prognosis of breast cancer patients with initial bone metastases. Front Public Health. 2022;10:1003976. 10.3389/fpubh.2022.1003976. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Lin J, Yin M, Liu L, Gao J, Yu C, Liu X, et al. The development of a prediction model based on random survival forest for the postoperative prognosis of pancreatic cancer: A SEER-based study. Cancers. 2022;14:4667. 10.3390/cancers14194667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Liu W, Zhu Y, Lin C, Liu L, Li G. An online prognostic application for melanoma based on machine learning and statistics. J Plast Reconstr Aesthet Surg. 2022;75:3853–8. 10.1016/j.bjps.2022.06.069. [DOI] [PubMed] [Google Scholar]
- 62.Ouyang D, Shi M, Wang Y, Luo L, Huang L. Prognostic analysis of pT1-T2aN0M0 cervical adenocarcinoma based on random survival forest analysis and the generation of a predictive nomogram. Front Oncol. 2022. 10.3389/fonc.2022.1049097. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Park JI, Bozkurt S, Park JW, Lee S. Evaluation of race/ethnicity-specific survival machine learning models for Hispanic and black patients with breast cancer. BMJ Health Care Inform. 2023;30:e100666. 10.1136/bmjhci-2022-100666. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Senders JT, Staples P, Mehrtash A, Cote DJ, Taphoorn MJB, Reardon DA, et al. An online calculator for the prediction of survival in glioblastoma patients using classical statistics and machine learning. Neurosurgery. 2020;86:E184–92. 10.1093/neuros/nyz403. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Sun H, Wu S, Li S, Jiang X. Which model is better in predicting the survival of laryngeal squamous cell carcinoma? Comparison of the random survival forest based on machine learning algorithms to Cox regression: analyses based on SEER database. Medicine. 2023;102:e33144. 10.1097/MD.0000000000033144. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Teng J, Zhang H, Liu W, Shu X-O, Ye F. A dynamic bayesian model for breast cancer survival prediction. IEEE J Biomed Health Inform. 2022;26:5716–27. 10.1109/JBHI.2022.3202937. [DOI] [PubMed] [Google Scholar]
- 67.Tian Y, Li J, Zhou T, Tong D, Chi S, Kong X, et al. Spatially varying effects of predictors for the survival prediction of nonmetastatic colorectal cancer. BMC Cancer. 2018;18:1084. 10.1186/s12885-018-4985-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Wang J, Chen N, Guo J, Xu X, Liu L, Yi Z. SurvNet: A novel deep neural network for lung cancer survival analysis with missing values. Front Oncol. 2021;10:588990. 10.3389/fonc.2020.588990 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Yu H, Huang T, Feng B, Lyu J. Deep-learning model for predicting the survival of rectal adenocarcinoma patients based on a surveillance, epidemiology, and end results analysis. BMC Cancer. 2022;22:210. 10.1186/s12885-022-09217-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Li C, Patil V, Rasmussen KM, Yong C, Chien H-C, Morreall D, et al. Predicting survival in veterans with follicular lymphoma using structured electronic health record information and machine learning. Int J Environ Res Public Health. 2021;18:2679. 10.3390/ijerph18052679. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Loureiro H, Becker T, Bauer-Mehren A, Ahmidi N, Weberpals J. Artificial intelligence for prognostic scores in oncology: a benchmarking study. Front Artif Intell. 2021. 10.3389/frai.2021.625573. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Ryu SM, Lee S-H, Kim E-S, Eoh W. Predicting survival of patients with spinal ependymoma using machine learning algorithms with the SEER database. World Neurosurg. 2018. 10.1016/j.wneu.2018.12.091. :S1878-8750(18)32914-0. [DOI] [PubMed] [Google Scholar]
- 73.Ishwaran H, Kogalur UB, Blackstone EH, Lauer MS. Random survival forests. Ann Appl Stat. 2008;2:841–60. 10.1214/08-AOAS169. [Google Scholar]
- 74.Omurlu I, Ture M, Tokatli F. The comparisons of random survival forests and Cox regression analysis with simulation and an application related to breast cancer. Expert Syst Appl. 2009;36:8582–8. 10.1016/j.eswa.2008.10.023. [Google Scholar]
- 75.Kantidakis G, Putter H, Lancia C, de Boer J, Braat AE, Fiocco M. Survival prediction models since liver transplantation - comparisons between Cox models and machine learning techniques. BMC Med Res Methodol. 2020;20:277. 10.1186/s12874-020-01153-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.O’Brien RC, Ishwaran H, Szczotka-Flynn LB, Lass JH, Cornea Preservation Time Study (CPTS) Group. Random survival forests analysis of intraoperative complications as predictors of Descemet stripping automated endothelial keratoplasty graft failure in the cornea preservation time study. JAMA Ophthalmol. 2021;139:191–7. 10.1001/jamaophthalmol.2020.5743. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Smith CC, Chai S, Washington AR, Lee SJ, Landoni E, Field K, et al. Machine-learning prediction of tumor antigen immunogenicity in the selection of therapeutic epitopes. Cancer Immunol Res. 2019;7:1591–604. 10.1158/2326-6066.CIR-19-0155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Yasaka K, Abe O. Deep learning and artificial intelligence in radiology: current applications and future directions. PLoS Med. 2018;15:e1002707. 10.1371/journal.pmed.1002707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y, DeepSurv. Personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. 2018;18:24. 10.1186/s12874-018-0482-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Wang W, Kiik M, Peek N, Curcin V, Marshall IJ, Rudd AG, et al. A systematic review of machine learning models for predicting outcomes of stroke with structured data. PLoS ONE. 2020;15:e0234722. 10.1371/journal.pone.0234722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Li B, Feridooni T, Cuen-Ojeda C, Kishibe T, de Mestral C, Mamdani M, et al. Machine learning in vascular surgery: a systematic review and critical appraisal. NPJ Digit Med. 2022;5:7. 10.1038/s41746-021-00552-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Collins GS, de Groot JA, Dutton S, Omar O, Shanyinde M, Tajar A, et al. External validation of multivariable prediction models: a systematic review of methodological conduct and reporting. BMC Med Res Methodol. 2014;14:40. 10.1186/1471-2288-14-40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Collins GS, Dhiman P, Andaur Navarro CL, Ma J, Hooft L, Reitsma JB, et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open. 2021;11:e048008. 10.1136/bmjopen-2020-048008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Cortes C, Jackel LD, Chiang WP. Limits on learning machine accuracy imposed by data quality. Adv Neural Inf Process Syst. 1994;7:57–62. [Google Scholar]
- 85.Gudivada V, Apon A, Ding J. Data quality considerations for big data and machine learning: going beyond data cleaning and transformations. Int J Adv Softw. 2017;10:1–20. [Google Scholar]
- 86.Ho LV, Ledbetter D, Aczon M, Wetzel R. The dependence of machine learning on electronic medical record quality. AMIA Annu Symp Proc AMIA Symp. 2018;2017 883–91. [PMC free article] [PubMed]
- 87.Sun W, Jiang M, Dang J, Chang P, Yin F-F. Effect of machine learning methods on predicting NSCLC overall survival time based on radiomics analysis. Radiat Oncol. 2018;13:197. 10.1186/s13014-018-1140-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Noroozi Z, Orooji A, Erfannia L. Analyzing the impact of feature selection methods on machine learning algorithms for heart disease prediction. Sci Rep. 2023;13:22588. 10.1038/s41598-023-49962-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Pfob A, Lu S-C, Sidey-Gibbons C. Machine learning in medicine: a practical introduction to techniques for data pre-processing, hyperparameter tuning, and model comparison. BMC Med Res Methodol. 2022;22:282. 10.1186/s12874-022-01758-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Afrash MR, Mirbagheri E, Mashoufi M, Kazemi-Arpanahi H. Optimizing prognostic factors of five-year survival in gastric cancer patients using feature selection techniques with machine learning algorithms: a comparative study. BMC Med Inform Decis Mak. 2023;23:54. 10.1186/s12911-023-02154-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets supporting the conclusions of this article are available from the corresponding author upon reasonable request.



