Abstract
Background
Systemic anticancer therapy (SACT) near the end of life (EOL) reduces the quality of the patient’s remaining life without clinical benefit. Studies investigating machine learning models for predicting cancer mortality to guide treatment decisions have primarily focused on specific types of cancer. This study aimed to evaluate the ability of a pan-cancer model to generalize across 10 cancer types when predicting short-term mortality.
Patients and methods
This study included patients with advanced cancer who were referred to the Department of Oncology at Aalborg University Hospital and died between January 2008 and December 2021 (N = 8690). Clinical data were used to train, validate, and test a pan-cancer model and 10 single-cancer models based on the eXtreme Gradient Boosting (XGBoost) algorithm. The average precision (AP) of the pan-cancer and single-cancer models was assessed and compared. Furthermore, explainable AI with Shapley additive explanations (SHAP) was used to evaluate shared prognostic information across the cancer types.
Results
The mean AP increased from 0.51 using the single-cancer models to 0.56 using the pan-cancer model to predict short-term mortality (random baseline 0.12). Important features identified by SHAP were shared across cancer types, indicating shared predictors of 30-day mortality. The most important features for predicting 30-day mortality were plasma albumin level, white blood cell count, and lactate dehydrogenase levels.
Conclusion
A pan-cancer model enhanced the performance of short-term mortality estimates compared with models based on single cancer types. Thus, including multiple cancer types in predictive modeling in oncology could be advantageous when shared predictors are expected across cancer types.
Key words: pan-cancer models, short-term mortality, counterproductive treatment, machine learning
Highlights
-
•
Comparing a pan-cancer model based on multiple cancers with single-cancer models.
-
•
The predictive performance of 30-day mortality was enhanced with the pan-cancer model.
-
•
We found shared prognostic information across different cancer types.
-
•
Prediction of 30-day mortality could support clinicians in the administration of SACT.
Introduction
Systemic anticancer therapy (SACT) near the end of life (EOL) can be associated with severe short-term side-effects, compromising the quality of life without beneficial effects for patients.1, 2, 3, 4 However, accurate identification of patients with cancer near EOL is difficult, and clinicians tend to overestimate life expectancy in patients with cancer. A nationwide registry-based study found that 16% of patients with cancer in Denmark received SACT 30 days before EOL.5 The correct identification of patients near EOL enables timely cessation of SACT and minimizes unnecessary side-effects.2
In recent years, various machine learning methods for predicting short-term mortality have been investigated using electronic health record (EHR) data. Previous studies primarily investigated single-cancer models.4,6, 7, 8 For single-cancer models it can be challenging to attain large training sets for prognostic machine learning models, particularly in rare cancer types, restricting the performance and generalizability of a machine learning model. Pan-cancer models that include data from multiple cancer types may mitigate these challenges and improve the performance of short-term mortality estimates compared with single-cancer models. In addition, cancer is known to share pathogenic mechanisms across cancer types, and there are shared physiological symptoms related to mortality, including fatigue, pain, and loss of appetite.9,10 The shared pathogenic and physiological factors related to mortality across cancer types give a rationale to explore pan-cancer models for short-term mortality estimates. However, pan-cancer models are seldom explored and are primarily based on cellular and molecular genetic data and pathological imaging.11, 12, 13 This study aimed to compare the performance of a pan-cancer model with 10 single-cancer models using EHR data to predict 30-day mortality in patients with advanced cancer. To the best of our knowledge, this is the first study to compare a pan-cancer model with single-cancer models for short-term mortality estimates using comprehensive time-series clinical data, including diagnosis history, treatment protocols, laboratory tests, and pathological data.
For this study, eXtreme Gradient Boosting (XGBoost) was chosen, as previous research has consistently demonstrated the superior performance of gradient boosting decision trees over alternative methods, including neural network approaches, on tabular data.8,14,15 This study focuses on a single, state-of-the-art approach, as the primary focus was to compare the performance between pan-cancer and single-cancer models rather than to compare different modeling techniques.
Patients and methods
Study population
The study population consisted of patients with cancer from the North Denmark Region referred to the Department of Oncology at Aalborg University Hospital between 1 January 2008 and 31 December 2021 (N = 36 358). The cohort was restricted to patients with advanced cancer as they were at risk of receiving unnecessary SACT near EOL (n = 15 025). Advanced cancer was specified as patients receiving SACT with palliative intent or having a metastatic diagnosis based on the International Statistical Classification of Diseases and Related Health Problems, 10th revision (ICD-10). The MedOnc prescription system distinguishes between treatment contexts, enabling the identification of treatments given with palliative intent. As not all patients with advanced cancer receive SACT with palliative intent, including patients with metastatic disease allowed us to capture patients with advanced cancer who may have never been treated. Rare and ill-defined cancer types were excluded due to potentially high data heterogeneity and insufficient data quality, as well as those rarely treated with SACT (n = 5005) (Supplementary Material Section 1, available at https://doi.org/10.1016/j.esmogo.2025.100172). To mitigate the imbalance in outcomes for the classification of 30-day mortality, only patients who died between 2010 and 2021 were kept (N = 8690). Keeping only patients who died after 2010 ensured a minimum of 2 years of lead time for all patients. The lead time ensured that patients had a sufficient treatment history before inclusion in the study to avoid bias introduced by patients who only had a very short medical history in the dataset. The cohort included 10 cancer types: urinary (n = 555), prostate (n = 900), uterine (n = 232), ovarian (n = 393), colorectal (n = 1558), pancreatic (n = 710), gastroesophageal (n = 606), breast (n = 1285), lung (n = 3044), and brain (n = 400) cancers. Some patients were diagnosed with more than one cancer type and were therefore included in multiple individual cancer types.
Data sources and variables
Data sources included the North Denmark Region’s patient administrative system, the Civil Registration System, the electronic prescribing system for medical cancer treatment used at the Department of Oncology (MedOnc), the pathology database (Patobank), and the Laboratory Information System (LABKA; Supplementary Material Section 2, available at https://doi.org/10.1016/j.esmogo.2025.100172). These data sources included structured tabular data on demographics, treatments, drug doses, procedures, admissions, blood tests, diagnoses, pathology, histology, procedures, side-effects, and comorbidities. The Charlson Comorbidity Index was used to identify comorbidities associated with survival. Each comorbidity from the Charlson Comorbidity Index was used as a separate binary feature (Supplementary Material Section 3, available at https://doi.org/10.1016/j.esmogo.2025.100172).
Data management was carried out using SAS Enterprise Guide 8.3 (SAS Institute Inc, Cary, NC) and Python 3.11.0 (Python Foundation, Wilmington, DE) with pandas 1.5.3 and NumPy 1.26.4.
Data acquisition
The raw data consisted of structured medical records. A data entry was generated for each contact/visit at the Department of Oncology based on available patient data at the time. Only data entries corresponding to contacts posterior to the start of SACT with palliative intent or the first registration of a metastatic diagnosis were included in the following. See Supplementary Material Section 4, available at https://doi.org/10.1016/j.esmogo.2025.100172 for more details.
All cancer cohorts were split into three datasets: training, validation, and test sets. The training dataset included patients who died from 2010 to 2017. The validation dataset included patients who died in 2018 and 2019. The test dataset included patients who died in 2020 and 2021. Patient-based data splitting ensured no data leakage between datasets, as the data retained the natural temporal relationship. The validation and test sets were based on the most recent data, allowing assessment of how well the model generalizes to current hospital populations. This approach ensures that while the model benefits from extensive historical data, its predictive power is evaluated on cases that best reflect present-day clinical practice. Different methods were used for the data generation, including dimensionality reduction, removal of sparse variables, highly correlated features, and multicollinear variables to facilitate the training of the models and enhance model performance.16 The selection process was based exclusively on the training dataset.
Models
The XGBoost algorithm was used as a supervised machine learning algorithm for binary classification to predict the probability of 30-day mortality. The model made a dynamic risk prediction, where predictions for each patient were updated at each visit to the Department of Oncology, assessing the risk based on the latest information available. The XGBoost model was trained and tested using Python 3.11.0 and was based on the ‘xgboost’ algorithm from Scikit-learn. For the XGBoost algorithm, the maximum depth, number of estimators, and learning rate were determined using hyperparameter optimization. The hyperparameters were optimized by Bayesian optimization using scikit-optimize 0.8.1. The hyperparameters were tuned using the training and validation datasets. The area under the precision and recall curve, commonly referred to as the average precision (AP), was used as the target loss during hyperparameter optimization and final training. The AP score was chosen as it is considered a reliable performance measure that considers the imbalances of the datasets and the precision of the model.17 In addition, six metrics were used to assess the performance: specificity, sensitivity, negative predictive value (NPV), positive predictive value (PPV), F1-score, and receiver operating characteristic curve–area under the curve (ROC–AUC), as recommended by Lu et al.7 These performance measures provide a balanced assessment of overall discrimination, class imbalance handling, and clinical relevance.
A single-cancer model was optimized, trained, and tested for each of the 10 cancer types. Furthermore, a pan-cancer model, including data from all cancer types, was optimized, trained, and tested. The variability of the performance for the pan-cancer model and the single-cancer models was assessed by a bootstrapping approach, resampling nine additional training sets with replacement from the original training data. Each bootstrap sample was of the same size as the original dataset, allowing models to be trained on slightly different distributions of the data. The test performance achieved for each single-cancer model was compared with that of the pan-cancer model tested on the corresponding cancer type. Furthermore, to investigate whether each of the single-cancer models learned any shared biological information related to the prognosis of 30-day mortality, all single-cancer models were tested on the test datasets for each cancer type. For example, the performance of the lung cancer model was found using the test dataset from brain, breast, ovarian, etc.
Evaluation of feature importance
Shapley additive explanations (SHAP) were used to investigate feature importance in both the pan-cancer and single-cancer models to evaluate the contributions of each feature. SHAP values were found and evaluated using the shap Python library version 0.46.0. Feature importance was used to highlight the important risk factors related to the 30-day mortality predictions of the models. Each feature impacts the prediction with either a positive or negative contribution to the predicted outcome of the model. To understand the overall impact of each feature’s contribution on the prediction, SHAP importance was considered, which estimates an overall importance measure based on the mean of the absolute SHAP coefficients. A bar chart was generated, presenting the 20 most important features of the pan-cancer model. Furthermore, a heatmap presenting the SHAP importance across all models was assessed to explore potential patterns across cancer types that exhibit similar prognostic characteristics.
Ethical approval and registration
The Danish Patient Safety Authority approved the exemption from informed consent for the project (case number 31-1521-334), and according to the General Data Protection Regulation, the study was registered at the North Denmark Region’s research project inventory (registration number 2019-41).
Results
Study population
A total of 8690 patients with advanced cancer across 10 cancer types were included. Cohort characteristics are shown in Table 1. The dataset included 372 427 contacts, of which 50 939 (13.7%) were within 30 days of death. While the overall distribution of the study population across training, validation, and test sets followed expected patterns, some trends emerged that warrant further attention. The proportion of included patients with a metastatic diagnosis increased from 44.3% in the training dataset to 74.0% in the test dataset. In addition, the proportion of included patients treated with SACT decreased from 82.7% in the training dataset to 70.7% in the test dataset. Furthermore, the proportion of patients aged ≥80 years increased from 12.4% in the training dataset to 19.7% in the validation dataset and to 20.4%. (Supplementary Material Section 5, available at https://doi.org/10.1016/j.esmogo.2025.100172).
Table 1.
Baseline characteristics of the training, validation, and test datasets
| Category | Training | Validation | Test | Total |
|---|---|---|---|---|
| Patients, n | 5308 | 1698 | 1684 | 8690 |
| Sex, n (%) | ||||
| Male | 2620 (49.4) | 882 (51.9) | 859 (51.0) | 4361 (50.2) |
| Female | 2688 (50.6) | 816 (48.1) | 825 (49.0) | 4329 (49.8) |
| Age at death, years, n (%) | ||||
| 18-59 | 993 (18.7) | 236 (13.9) | 201 (11.9) | 1430 (16.5) |
| 60-69 | 1707 (32.2) | 474 (27.9) | 424 (25.2) | 2605 (30.0) |
| 70-79 | 1949 (36.7) | 654 (38.5) | 715 (42.5) | 3318 (38.2) |
| 80+ | 659 (12.4) | 334 (19.7) | 344 (20.4) | 1337 (15.4) |
| Palliative treatment, n (%) | ||||
| Yes | 4388 (82.7) | 1183 (69.7) | 1190 (70.7) | 6761 (77.8) |
| No | 920 (17.3) | 515 (30.3) | 494 (29.3) | 1929 (22.2) |
| Metastasis, n (%) | ||||
| Yes | 2354 (44.3) | 1194 (70.3) | 1247 (74.0) | 4795 (55.2) |
| No | 2954 (55.7) | 504 (29.7) | 437 (26.0) | 3895 (44.8) |
| 1-6 months | 1192 (27.2) | 276 (16.3) | 240 (20) | 1708 (25.3) |
| 6-12 months | 1131 (25.8) | 286 (24.2) | 261 (21.9) | 1678 (24.8) |
| 12+ months | 1865 (42.5) | 588 (49.7) | 644 (54.1) | 3097 (45.8) |
| Contacts with the Department of Oncology | ||||
| All contacts, n | 211 690 | 79 131 | 81 606 | 372 427 |
| Within 30 days of death, n (%) | 29 662 (14.0) | 11 165 (14.1) | 10 112 (12.4) | 50 939 (13.7) |
| Cancers, n (%) | ||||
| Lung | 1877 (35.4) | 594 (35.0) | 573 (34.0) | 3044 (35.0) |
| Colorectal | 1010 (19.0) | 278 (16.4) | 270 (16.0) | 1558 (17.9) |
| Breast | 809 (15.2) | 224 (13.2) | 252 (15.0) | 1285 (14.8) |
| Prostate | 461 (8.7) | 208 (12.2) | 231 (13.7) | 900 (10.4) |
| Pancreatic | 393 (7.4) | 158 (9.3) | 159 (9.4) | 710 (8.2) |
| Gastroesophageal | 355 (6.7) | 134 (7.9) | 117 (6.9) | 606 (7.0) |
| Ovarian | 258 (4.9) | 66 (3.9) | 69 (4.1) | 393 (4.5) |
| Urinary | 317 (6) | 122 (7.2) | 116 (6.9) | 555 (6.4) |
| Brain | 247 (4.7) | 76 (4.5) | 77 (4.6) | 400 (4.6) |
| Uterine | 133 (2.5) | 48 (2.8) | 51 (3.0) | 232 (2.7) |
Performances—comparing models
The hyperparameters were individually tuned for the pan-cancer and single-cancer models. Results from hyperparameter optimization can be found in Supplementary Material Section 6, available at https://doi.org/10.1016/j.esmogo.2025.100172. The results of testing the trained models on their corresponding test datasets are presented in Table 2 with the following performance metrics: specificity, sensitivity, PPV, NPV, F1-score, ROC–AUC, and AP. The results represent a mean value from testing on models trained on the data set and nine additional bootstrapped training sets of the same length. The mean AP for the single-cancer models was 0.51 and the AP for the pan-cancer model was 0.56. The single-cancer models for gastroesophageal and urinary cancer obtained the best AP scores (0.64), whereas the brain cancer model had the lowest AP score (0.34). In general, the models achieved high specificity measures (mean 0.96) and low sensitivity (mean 0.35). The random baseline for AP is 0.12. The random baseline refers to the performance of a model that makes predictions with equal probability for each class, without considering the input features.
Table 2.
Performance metrics: specificity, sensitivity, PPV, NPV, F1-score, ROC–AUC, and AP for the single-cancer models tested on the respective cancer type and pan-cancer model tested on all cancer types
| Model and dataset | Specificity | Sensitivity | PPV | NPV | F1-score | ROC–AUC | AP |
|---|---|---|---|---|---|---|---|
| Lung | 0.95 | 0.40 | 0.58 | 0.90 | 0.48 | 0.85 | 0.53 |
| Colorectal | 0.97 | 0.38 | 0.56 | 0.93 | 0.45 | 0.87 | 0.46 |
| Breast | 0.99 | 0.30 | 0.66 | 0.94 | 0.42 | 0.90 | 0.51 |
| Prostate | 0.97 | 0.26 | 0.61 | 0.89 | 0.36 | 0.86 | 0.48 |
| Pancreatic | 0.94 | 0.46 | 0.58 | 0.91 | 0.51 | 0.87 | 0.56 |
| Gastroesophageal | 0.90 | 0.60 | 0.60 | 0.90 | 0.60 | 0.88 | 0.64 |
| Urinary | 0.98 | 0.37 | 0.74 | 0.89 | 0.49 | 0.87 | 0.64 |
| Brain | 1.00 | 0.00 | 1.00 | 0.86 | 0.01 | 0.74 | 0.34 |
| Ovarian | 0.97 | 0.42 | 0.55 | 0.95 | 0.48 | 0.90 | 0.49 |
| Uterine | 0.91 | 0.32 | 0.45 | 0.86 | 0.37 | 0.81 | 0.46 |
| Pan-cancer | 0.98 | 0.30 | 0.67 | 0.91 | 0.41 | 0.88 | 0.56 |
All models were optimized using AP as the target loss. The results show the mean value obtained from training on both the original training dataset and the nine additional bootstrapped training datasets.
AP, average precision; NPV, negative predictive value; PPV, positive predictive value; ROC–AUC, receiver operating characteristic curve–area under the curve.
Comparing pan-cancer models with general models
The pan-cancer model was applied to the test dataset for each of the 10 cancer types and compared with the performances of the respective single-cancer models. The results are shown in Figure 1. The pan-cancer model achieved improved or similar AP in all cancer types compared with the single-cancer models, except for brain cancer. The mean AP of the single-cancer models was 0.51, whereas the mean AP of the pan-cancer model was 0.56. Urinary cancer achieved the highest AP of all cancer types using the pan-cancer model (0.70). The brain cancer test data obtained a considerably worse performance compared with other cancer types. The largest performance improvement with the pan-cancer model was observed for ovarian cancer (from 0.50 to 0.63). By contrast, the smallest performance improvement was observed for pancreatic cancer (from 0.55 to 0.56).
Figure 1.
The average precision (AP) for the test data from single cancers when using the single-cancer models and the pan-cancer model, respectively. The sizes of the points indicate the relative sample size. The horizontal axis represents the AP of the cancer-specific model using the cancer-specific test data. The vertical axis represents the AP of the pan-cancer model on the cancer-specific test data.
A comparison between the performance of the XGBoost model and a baseline model, logistic regression with elastic net regularization, is presented in Supplementary Material Section 7, available at https://doi.org/10.1016/j.esmogo.2025.100172. The APs of the pan-cancer model and the single-cancer models, when tested for each cancer type, are presented in the heatmap in Figure 2. The heatmap demonstrates that the AP of each model is affected by the dataset on which it is tested. The test dataset for urinary and gastroesophageal cancer generally achieved a high AP across all models, whereas brain cancer generally achieved a low AP across all models. Furthermore, most single-cancer models attained the highest AP on another cancer type than the cancer type that the model originally was trained on. The pan-cancer model outperformed or matched AP in lung, colorectal, breast, prostate, pancreatic, and ovarian cancers.
Figure 2.
Heatmap showing average precision (AP) values for each single-cancer model and pan-cancer model when tested on each of the cancer test datasets. The horizontal axis represents each model, and the vertical axis represents each cancer type.
Feature importance
Important risk factors related to 30-day mortality were analyzed using SHAP values to interpret the contribution of each feature to the prediction of the models. Figure 3 shows a bar chart of the SHAP importance and a beeswarm plot of the distribution of SHAP values. The figure includes the 20 features with the highest feature importance for the pan-cancer model. The beeswarm plot was constructed based on 100 000 predictions.
Figure 3.
A bar chart representing the top 20 features of the pan-cancer model. To the right, a beeswarm plot shows the distribution of Shapley additive explanations (SHAP) values based on 100 000 predictions for the top 20 features. BMI, body mass index.
From the SHAP analysis, plasma albumin level had the highest impact on 30-day mortality predictions, with a SHAP importance of 1.05. A low plasma albumin level was associated with a positive impact on the prediction of 30-day mortality, whereas a high plasma albumin level was related to a negative impact on the prediction, though a less pronounced effect. Thus, low albumin level was associated with an increased risk of 30-day mortality. Leukocytes and lactate dehydrogenase also exhibited high importance in predicting 30-day mortality, with SHAP importance values of 0.42 and 0.24, respectively. The top 20 most important features included laboratory test results, age, body mass index, and diagnosis. To validate the SHAP values and the feature importance, we assessed the effect of permuting individual features on test performance. The largest performance drop was observed for albumin, which is the highest-ranked feature according to SHAP values (Supplementary Material Section 8, available at https://doi.org/10.1016/j.esmogo.2025.100172).
Additionally, we evaluated the SHAP importance of each single-cancer model for the 20 most important features found in the pan-cancer model. Figure 4 presents a heatmap of the SHAP importance for each model. The SHAP importance in this plot is converted to a percentage importance. The three most important features of the pan-cancer model (plasma albumin, leukocytes, and lactate dehydrogenase levels) were selected by a minimum of nine single-cancer models. Furthermore, plasma albumin was the most important feature of all models. The heatmap shows that some of the important features are specific to the pan-cancer model including having lung cancer as diagnosis. In addition, some important features are specific to the single-cancer models, for example, for brain and uterine cancer, creatinine and difference in leukocytes are weighted as more important than for other cancer types.
Figure 4.
Heatmap of the 20 most important features of the pan-cancer model. On the x-axis, the top 20 features are listed from the most important feature on the left to the least important feature on the right. Empty spaces occur in case of features excluded during the data preprocessing. The cancer types are listed on the y-axis. BMI, body mass index; SHAP, Shapley additive explanations.
Discussion
This predictive modeling study aimed to compare the performance of a pan-cancer model with 10 single-cancer models using EHR data to predict 30-day mortality in patients with advanced cancer. We found that the pan-cancer model achieved better performance in 8 out of 10 cancer types when compared with the performance of the single-cancer models (mean single-cancer models 0.51 mean pan-cancer model 0.56). The pan-cancer model is theorized to improve the predictive performance due to additional cancer types containing shared prognostic factors with the target cancer. Studies investigating the molecular interactions shared across cancers and their association with therapy strategies and prognosis have also demonstrated that prognostic features are shared between cancer types.18,19
The analysis of the SHAP importance from the pan-cancer model and each single-cancer model demonstrated that laboratory test results were of great importance. This is supported by Vesteghem et al.8 investigating machine learning for 30-day mortality prediction in lung cancer using the same data sources as in this study. For all the models of this study, plasma albumin level was identified as the most important factor for short-term mortality. Low plasma albumin levels positively influenced the model’s prediction, thus significantly contributing to the likelihood of 30-day mortality. According to a study by Tang et al.,20 an albumin level <4.2 g/dl is closely related to cancer mortality. Albumin can potentially affect the outcome of cancer treatment as it is essential for maintaining the oncotic pressure of the blood and for the transportation of some cancer-targeting drugs.21 Leukocyte count, which had the second-highest SHAP coefficient for the pan-cancer model, is a known predictor of cancer-related mortality. Several studies have found that leukocytosis is a predictor of cancer-related mortality.22,23 Furthermore, the lactate-to-albumin ratio is correlated with the Sequential Organ Failure Assessment (SOFA) score, which may explain why lactate dehydrogenase obtained the third-highest SHAP coefficient.24 In addition, lactate dehydrogenase levels, white blood cell counts, and decreased albumin levels were found to be risk factors for short-term mortality by a study investigating 14-day mortality in patients with advanced or metastatic cancer using laboratory test results.25
There are only a few articles that directly compare the performance of pan-cancer models and single-cancer models. Fan et al.26 predicted overall survival and compared the C-index of a pan-cancer model with 20 single-cancer models, including 11 160 patients. Results cannot be directly compared, as the C-index in their study is a regression-based metric for survival analysis, while our study employs a binary classification of 30-day mortality. However, their pan-cancer model outperformed single-cancer models in 16 of 20 cancer types, thus supporting the findings of our study, that a pan-cancer model can generalize across different cancer types, for survival predictions. The pan-cancer model of their research was trained using a deep learning architecture on multimodal data encompassing clinical data (cancer type, sex, race, histological type, and age) and genetic data including gene expressions. They found that the performance improvement was most likely related to the large sample size attained for the pan-cancer data. Furthermore, they hypothesized that cancer types that did not obtain improvement in performance using the pan-cancer model (kidney renal papillary cell carcinoma, low-grade gliomas, kidney chromophobe, stomach adenocarcinoma) were caused by those types having unique characteristics.26 This explanation could also apply to brain cancer in our study.
Most single-cancer models achieved the highest AP on cancer types other than the one they were trained on. This could be due to several factors. Feature similarity between different cancer types may allow better generalization to other cancer types. Another important factor could be differences in data quality. If the ovarian cancer dataset is smaller or noisier, a model trained on lung cancer may perform better due to more consistent data. Furthermore, class imbalance in specific datasets might influence performance, where a model trained on a more balanced dataset generalizes better.
The AP for each model was highly affected by the dataset, for example, urinary cancer achieved a high AP across all models, whereas brain cancer achieved a low AP. This could indicate that the mortality of some cancer types is more predictable than others. It could also be caused by limitations related to data quality and reliability. A study assessing the validity of SACT registration within the Danish National Patient Registry revealed significant underreporting issues, particularly for complex treatment regimens such as those used in glioblastoma management.21 These reporting challenges may partially explain the significant drop in performance observed using the brain cancer dataset.
Lung cancer, which had the largest sample size of the included cancer types, only achieved little improvement in performance using the pan-cancer model (from 0.53 to 0.55), whereas the second smallest cohort, ovarian cancer, attained the greatest improvement in performance (from 0.50 to 0.63). However, for our study, no overall trend was observed that could indicate that smaller cohorts attained bigger improvements in performance compared with larger cohorts.
The clinical application of our model depends on the intended use, as the balance between false positives and false negatives varies based on the decision-making context. Our results show a high NPV, meaning the model is particularly reliable in identifying patients unlikely to experience early mortality. This can help ensure that patients at low risk can continue receiving treatment without unnecessary concerns. Conversely, with only a moderate PPV of the model, a high-risk prediction should not automatically lead to the discontinuation of SACT but should instead prompt clinicians to consider the treatment’s benefits in the context of the patient’s overall condition. As the model operates as a dynamic risk prediction tool rather than as a one-time classifier, its predictions can change over time. A patient flagged as high-risk at one point, perhaps due to an acute infection, may later have an improved prognosis, allowing for reassessment of treatment. This continuous evaluation ensures that decisions remain flexible to the patient’s evolving clinical status.
The ROC–AUC values (mean 0.86) are relatively high compared with the AP values (mean 0.52). This is explained by the imbalance of the data also reflected by the high specificity measures (mean 0.96) and low sensitivity measures (mean 0.35). High specificity measures indicate a strong ability to correctly identify patients with an expected survival of >30 days. However, the low sensitivity suggests that the models are challenged to capture a significant proportion of patients with 30-day mortality.
We acknowledge that the sensitivity of the model is suboptimal, and achieving higher sensitivity should be a priority before clinical implementation, where the accurate identification of high-risk patients is crucial. However, the current model was primarily designed to demonstrate the feasibility of pan-cancer predictions rather than to optimize for sensitivity alone. In future iterations, we anticipate that the model’s performance can be improved by incorporating additional data sources. Specifically, the inclusion of unstructured data such as patient performance status and treatment toxicity could enable improved sensitivity without compromising other performance metrics.
The models in our study were developed based on single-center data from the North Denmark Region, resulting in highly consistent clinical data. However, this can restrict the generalizability and robustness of the models to broader populations. As the Danish healthcare system is uniform across regions, the impact of this limitation is expected to be minimal within Denmark.
Conclusion
This study evaluated and compared a pan-cancer model trained on 10 different cancer types with single-cancer models using the XGBoost algorithm. The models were trained to predict 30-day mortality in patients with advanced cancer. The pan-cancer model demonstrated successful generalization across multiple cancer types, improving predictive performance compared with single-cancer models in 8 out of 10 cancer types. Important prognostic features included plasma albumin, leukocytes, and lactate dehydrogenase. Prognostic features for short-term mortality were shared across cancer types, emphasizing the potential of pan-cancer models. The feature importance estimates support the identification of critical factors influencing the outcome of the model, allowing for more informed decisions regarding the administration of SACT near EOL.
Acknowledgements
The authors thank Special Consultant Thomas Mulvad Larsen, Business Intelligence Unit, North Denmark Region, for the support provided in relation to the data.
Funding
This work was supported by the Research Hive REPAIR at the Clinical Cancer Research Center, Aalborg University Hospital, Aalborg, Denmark (no grant number).
Disclosure
The authors have declared no conflicts of interest.
Supplementary data
References
- 1.Canavan M.E., Wang X., Ascha M.S., et al. Systemic anticancer therapy and overall survival in patients with very advanced solid tumors. JAMA Oncol. 2024;10:887–895. doi: 10.1001/jamaoncol.2024.1129. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Ma Z., Li H., Zhang Y., et al. Prevalence of aggressive care among patients with cancer near the end of life: a systematic review and meta-analysis. EClinicalMedicine. 2024;71 doi: 10.1016/j.eclinm.2024.102561. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Christakis N.A., Lamont E.B. Extent and determinants of error in doctors’ prognoses in terminally ill patients: prospective cohort study. Br Med J. 2000;320(7233):469–472. doi: 10.1136/bmj.320.7233.469. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Gajra A., Zettler M.E., Miller K.A., et al. Augmented intelligence to predict 30-day mortality in patients with cancer. Future Oncol. 2021;17(29):3797–3807. doi: 10.2217/fon-2021-0302. [DOI] [PubMed] [Google Scholar]
- 5.Mattsson T.O., Pottegård A., Jørgensen T.L., Green A., Bliddal M. End-of-life anticancer treatment—a nationwide registry-based study of trends in the use of chemo-, endocrine, immune-, and targeted therapies. Acta Oncol. (Madr) 2021;60(8):961–967. doi: 10.1080/0284186X.2021.1890332. [DOI] [PubMed] [Google Scholar]
- 6.Reddy V., Nafees A., Raman S. Recent advances in artificial intelligence applications for supportive and palliative care in cancer patients. Curr Opin Support Palliat Care. 2023;17(2):125–134. doi: 10.1097/SPC.0000000000000645. [DOI] [PubMed] [Google Scholar]
- 7.Lu S.C., Xu C., Nguyen C.H., Geng Y., Pfob A., Sidey-Gibbons C. Machine learning-based short-term mortality prediction models for patients with cancer using electronic health record data: systematic review and critical appraisal. JMIR Med Inform. 2022;10(3) doi: 10.2196/33182. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Vesteghem C., Szejniuk W.M., Brøndum R.F., Falkmer U.G., Azencott C.-A., Bøgsted M. Dynamic risk prediction of 30-day mortality in patients with advanced lung cancer: comparing five machine learning approaches. JCO Clin Cancer Inform. 2022;6 doi: 10.1200/CCI.22.00054. [DOI] [PubMed] [Google Scholar]
- 9.Tateo V., Marchese P.V., Mollica V., Massari F., Kurzrock R., Adashek J.J. Agnostic approvals in oncology: getting the right drug to the right patient with the right genomics. Pharmaceuticals. 2023;16(4):614. doi: 10.3390/ph16040614. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Ijaopo E.O., Zaw K.M., Ijaopo R.O., Khawand-azoulai M. A review of clinical signs and symptoms of imminent end-of-life in individuals with advanced illness. 2023;9 doi: 10.1177/23337214231183243. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Halabi S. Pan-cancer prognostic models of clinical outcomes: statistical exercise or clinical tools? Ann Oncol. 2020;31(11):1427–1429. doi: 10.1016/j.annonc.2020.08.2233. [DOI] [PubMed] [Google Scholar]
- 12.Li C., Huang Y., Yi X., Tang Y., Okita R., He J. Pan-cancer prognostic model and immune microenvironment analysis of natural killer cell-related genes. Transl Cancer Res. 2024;13(4):1936–1953. doi: 10.21037/tcr-24-434. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Arslan S., Schmidt J., Bass C., et al. A systematic pan-cancer study on deep learning-based prediction of multi-omic biomarkers from routine pathology images. Commun Med. 2024;4(1):48. doi: 10.1038/s43856-024-00471-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Shwartz-Ziv R., Armon A. Tabular data: deep learning is not all you need. Inf Fusion. 2022;81:84–90. [Google Scholar]
- 15.Borisov V., Leemann T., Sessler K., Haug J., Pawelczyk M., Kasneci G. Deep neural networks and tabular data: a survey. IEEE Trans Neural Netw Learn Syst. 2022;35(6):7499–7519. doi: 10.1109/TNNLS.2022.3229161. [DOI] [PubMed] [Google Scholar]
- 16.Jia W., Sun M., Lian J., Hou S. Feature dimensionality reduction: a review. Complex Intell Syst. 2022;8(3):2663–2693. [Google Scholar]
- 17.Saito T., Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3) doi: 10.1371/journal.pone.0118432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Zhao Y., Li X., Loscalzo J., et al. Transcript and protein signatures derived from shared molecular interactions across cancers are associated with mortality. J Transl Med. 2024;22(1):444. doi: 10.1186/s12967-024-05268-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Ghoshdastider U., Rohatgi N., Mojtabavi Naeini M., et al. Pan-cancer analysis of ligand-receptor cross-talk in the tumor microenvironment. Cancer Res. 2021;81(7):1802–1812. doi: 10.1158/0008-5472.CAN-20-2352. [DOI] [PubMed] [Google Scholar]
- 20.Tang Q., Li X., Sun C.R. Predictive value of serum albumin levels on cancer survival: a prospective cohort study. Front Oncol. 2024;14 doi: 10.3389/fonc.2024.1323192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Akirov A., Masri-Iraqi H., Atamna A., Shimon I. Low albumin levels are associated with mortality risk in hospitalized patients. Am J Med. 2017;130(12):1465.e11–1465.e19. doi: 10.1016/j.amjmed.2017.07.020. [DOI] [PubMed] [Google Scholar]
- 22.Gripp S., Moeller S., Bölke E., et al. Survival prediction in terminally ill cancer patients by clinical estimates, laboratory tests, and self-rated anxiety and depression. J Clin Oncol. 2007;25(22):3313–3320. doi: 10.1200/JCO.2006.10.5411. [DOI] [PubMed] [Google Scholar]
- 23.Maltoni M., Caraceni A., Brunelli C., et al. Prognostic factors in advanced cancer patients: evidence-based clinical recommendations – a study by the Steering Committee of the European Association for Palliative Care. J Clin Oncol. 2005;23(25):6240–6248. doi: 10.1200/JCO.2005.06.866. [DOI] [PubMed] [Google Scholar]
- 24.Chae B.R., Kim Y.J., Lee Y.S. Prognostic accuracy of the Sequential Organ Failure Assessment (SOFA) and quick SOFA for mortality in cancer patients with sepsis defined by systemic inflammatory response syndrome (SIRS) Support Care Cancer. 2020;28(2):653–659. doi: 10.1007/s00520-019-04869-z. [DOI] [PubMed] [Google Scholar]
- 25.Cheng L., DeJesus A.Y., Rodriguez M.A. Using laboratory test results at hospital admission to predict short-term survival in critically ill patients with metastatic or advanced cancer. J Pain Symptom Manage. 2017;53(4):720–727. doi: 10.1016/j.jpainsymman.2016.11.008. [DOI] [PubMed] [Google Scholar]
- 26.Fan Z., Jiang Z., Liang H., Han C. Pancancer survival prediction using a deep learning architecture with multimodal representation and integration. Bioinform Adv. 2023;3(1) doi: 10.1093/bioadv/vbad006. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




