Abstract
Elderly patients with non-small cell lung cancer (NSCLC) and bone metastases face a dire prognosis, creating an urgent need for accurate short-term mortality prediction to guide care. To develop and validate machine learning (ML) models for predicting 3-month cancer-specific mortality in patients aged 70 years or older with NSCLC and bone metastases. We analyzed data from 1,773 patients (aged ≥ 70) from the SEER database (2010–2020). The cohort was randomly split into training (70%) and validation (30%) sets. Seven ML algorithms were trained and evaluated using a comprehensive set of performance metrics, including area under the curve (AUC), calibration, and decision curve analysis for clinical utility. The overall 3-month mortality rate was 48.5%. Among the seven models, the logistic regression model demonstrated superior and stable overall performance. It achieved the highest AUC of 0.79 on the validation set and maintained the highest average accuracy (0.78 ± 0.02) across 10-fold cross-validation. Decision curve analysis confirmed its superior net clinical benefit across most threshold probabilities. Key influential predictors identified included the absence of chemotherapy, presence of liver or brain metastases, and a shorter time from diagnosis to treatment initiation. This study developed and validated an interpretable ML-based prediction model that accurately identifies elderly NSCLC patients with bone metastases who are at high risk of early death. The logistic regression model, selected as the optimal tool, can assist clinicians in making individualized decisions, potentially guiding more aggressive supportive care for high-risk patients and definitive treatments for those with a better prognosis. However, this model was developed and internally validated using a single SEER cohort; external validation in independent datasets is required before clinical application.
Keywords: Non-small cell lung carcinoma, Bone neoplasms, Machine learning, Prognosis, Mortality
Subject terms: Cancer, Computational biology and bioinformatics, Medical research, Oncology, Risk factors
Introduction
Lung cancer constitutes a formidable and persistent global health crisis, maintaining its position as a leading cause of both cancer-related incidence and mortality. Contemporary epidemiological data, as consolidated by Global Cancer Statistics, revealed an estimated 2.2 million new diagnoses worldwide in 2020, representing over 11% of all new cancer cases, and was responsible for approximately 1.8 million deaths, underscoring its profound impact on public health systems1,2. Within this pathological landscape, non-small cell lung cancer (NSCLC) emerges as the predominant histological subtype, accounting for roughly 85% of all cases. A critical determinant of the adverse clinical trajectory in NSCLC is the development of distant metastases, with bone being one of the most prevalent sites3. The occurrence of bone metastasis (BM) is observed in 30–40% of NSCLC patients, heralding a significant decline in prognosis and quality of life3–5. This is primarily mediated through the onset of skeletal-related events (SREs), a constellation of debilitating complications including severe bone pain, pathological fractures, spinal cord compression, and hypercalcemia, which collectively exacerbate morbidity and impose a substantial burden on patients and healthcare resources4,5.
The clinical management paradigm becomes markedly more complex when applied to the growing demographic of elderly patients, defined here as those aged 70 years or older, who present with NSCLC and concomitant bone metastases6. In this vulnerable population, clinicians are confronted with a multifaceted therapeutic dilemma: they must meticulously calibrate the potential survival advantage offered by systemic anti-cancer therapies—such as chemotherapy, immunotherapy, and targeted agents—against the heightened vulnerability to treatment-related toxicities, the presence of comorbid conditions, and the paramount objective of preserving functional status and quality of life7. Central to navigating this dilemma is the ability to generate a reliable estimation of individual life expectancy. In clinical practice, an anticipated survival of less than three months is widely considered a contraindicating factor for aggressive surgical interventions, which are associated with significant morbidity and prolonged recovery periods8. Consequently, the development of a precise and individualized tool for predicting short-term mortality is not merely an academic exercise but a clinical necessity.
In recent years, machine learning (ML) has emerged as a transformative methodology within computational oncology, offering a powerful framework for analyzing high-dimensional and complex biomedical data9–11. Unlike conventional statistical models that often rely on linear assumptions, ML algorithms excel at identifying intricate, non-linear interactions and subtle patterns among variables that may elude traditional analysis. This capability has fueled its successful application across the oncology spectrum, from enhancing early detection and refining diagnostic accuracy to guiding personalized treatment planning and improving prognostic stratification9–11. Notably, a growing body of evidence suggests that well-constructed ML models can achieve predictive performance superior to that of established prognostic criteria and traditional non-ML statistical counterparts in various cancer types12.
Building upon this foundational evidence, we postulated that the analytical power of machine learning could be strategically harnessed to address the specific prognostic challenge in elderly NSCLC patients with bone metastases. We hypothesized that by training a diverse set of ML algorithms on a comprehensive set of clinically relevant variables, we could develop and validate a robust, high-performance prediction model for 3-month all-cause mortality. The ultimate goal of this research is to provide clinicians with a data-driven, objective decision-support tool. By facilitating a more accurate identification of patients at the highest risk of short-term mortality, our model aims to inform personalized clinical strategies, thereby ensuring that the intensity of care is appropriately aligned with individual patient prognosis and goals of care.
Methods
Patients and study design
This retrospective cohort study utilized data from the Surveillance, Epidemiology, and End Results (SEER) database for patients diagnosed between 2010 and 2020. The study included patients aged 70 years or older who were diagnosed with non-small cell lung cancer (NSCLC) and had documented bone metastases. Exclusion criteria encompassed patients who died from unspecified or non-cancer-related causes, had incomplete key data, presented with other primary cancers, or were alive with a follow-up duration of less than three months to mitigate potential immortality bias. A final cohort of 1,773 eligible patients was identified and randomly allocated into a training set (70%, n = 1,240) for model development and a validation set (30%, n = 533) for performance assessment. As the SEER database comprises publicly available, de-identified data, ethical approval and the requirement for informed consent were waived for this study.A total of 1,773 patients who met the inclusion and exclusion criteria were enrolled in the final analysis. The study cohort was then randomly split into a training set (70%, n = 1,240) for model development and a validation set (30%, n = 533) for performance assessment. The data supporting the findings of this study are publicly available from the Surveillance, Epidemiology, and End Results (SEER) database. All analyses were conducted in accordance with the data use agreements stipulated by SEER for the use of its de-identified public data sets.Therefore, ethical approval and informed consent were waived for this study. seven methods, comprising Logistic Regression, Decision Tree, Recurrent Neural Network(RNN), Light Gradient Boosting Machine, Random Forest, Extreme Gradient Boosting, Support Vector Machine were presented. Eighteen criteria were employed to evaluate and contrast the accuracy of models’ forecasts. In order to precisely foretell survival outlook in these individuals, it was theorized in this study that machine learning techniques could generate multiple models, and the most efficient one could be determined through comparisons.
Potential predictors
Our analysis incorporated 18 clinically relevant variables obtained from the database, which were stratified into the following categories: demographics (age, sex, race, marital status); tumor characteristics (primary site, histologic type, laterality, grade, AJCC 6th edition T stage, N stage, tumor size); treatment modalities (surgery, radiation, chemotherapy); metastatic status (brain, liver, lung metastases); and a time-to-treatment metric (months from diagnosis to treatment). The primary outcome was 3-month cancer-specific mortality, defined explicitly as death attributable to lung cancer within 3 months of diagnosis.(Note: cancer-specific survival is the complement of this outcome, i.e., CSS = 1 − mortality.)
Model development
All statistical analyses and model development were performed using R (version 4.1.2) and Python (version 3.9.7).Categorical variables were encoded. Missing data were handled by complete‑case analysis, as the proportion of missing values for each variable was low (< 5%). Prediction models were developed in the training set by employing seven machine learning algorithms, including Logistic Regression, Decision Tree, Recurrent Neural Network (RNN), Light Gradient Boosting Machine, Random Forest, Extreme Gradient Boosting, and Support Vector Machine.
Hyperparameter tuning for each algorithm was conducted via a combination of grid search and random search, with 10-fold cross-validation to prevent overfitting. The model training process incorporated 100 rounds of bootstrapping to ensure robustness.
Model evaluation and selection
The predictive performance of all seven models was rigorously evaluated on the validation cohort using 15 metrics: Area Under the Curve (AUC), Brier Score, Calibration Slope, Calibration Intercept (Intercept-in-the-large), Average Predicted Probability, Discrimination Slope, Accuracy, Sensitivity (Recall), Specificity, Positive Predictive Value (Precision), Negative Predictive Value, F1 Score, Cohen’s Kappa, Matthews Correlation Coefficient (MCC), and Youden’s Index.
Model selection was based on a hierarchical approach. The primary criterion was discriminative ability, measured by the Area Under the ROC Curve (AUC) on the validation set. Among models with similar AUC (difference < 0.01), we further compared calibration (calibration plots and Brier score) and clinical utility (Decision Curve Analysis, DCA). The model with the highest AUC, acceptable calibration (calibration slope close to 1, Brier score low), and favorable net benefit on DCA was selected as the optimal model.Decision Curve Analysis (DCA) was also performed to quantify net clinical benefit.
Model explainability
To enhance the clinical interpretability of the optimal model, we performed a global explainability analysis using SHapley Additive exPlanations (SHAP). This method quantifies the contribution and directional impact of each predictive variable on the model’s output across the entire validation cohort.
Statistical analysis
Data preprocessing and feature selection
Categorical variables were presented as frequencies and proportions, while continuous variables were expressed as mean ± standard deviation. Differences between the training and validation sets were assessed using the Chi-square test (or Fisher’s exact test where appropriate) for categorical variables and the independent samples t-test for continuous variables. Prior to model development, all features were standardized to a mean of zero and a standard deviation of one to ensure comparability across scales.We performed feature selection using multivariable logistic regression, retaining variables with P < .10 for input into the machine learning models. All 18 candidate variables listed in Table 1 were simultaneously entered into the regression model, and variables with adjusted P values below the threshold were selected. This prescreening step aimed to prevent overfitting and promote a more parsimonious model, while using a liberal threshold (P < .10) to avoid excluding potentially important predictors.
Table 1.
Baseline characteristics of the study cohort.
| Characteristics | Overall | 1 (N = 1240) | 2 (N = 533) | p | |
|---|---|---|---|---|---|
| Age | 70–79 | 1276 (72%) | 894 (72.1%) | 382 (71.7%) | 0.269 |
| 80–89 | 472 (26.6%) | 325 (26.2%) | 147 (27.6%) | ||
| >=90 | 25 (1.4%) | 21 (1.7%) | 4 (0.8%) | ||
| Sex | Male | 1025 (57.8%) | 716 (57.7%) | 309 (58%) | 0.970 |
| Female | 748 (42.2%) | 524 (42.3%) | 224 (42%) | ||
| Race | White | 1478 (83.4%) | 1034 (83.4%) | 444 (83.3%) | 0.977 |
| Black | 132 (7.4%) | 93 (7.5%) | 39 (7.3%) | ||
| Other | 163 (9.2%) | 113 (9.1%) | 50 (9.4%) | ||
| Marital status | Married | 1091 (61.5%) | 773 (62.3%) | 318 (59.7%) | 0.313 |
| Unmarried | 682 (38.5%) | 467 (37.7%) | 215 (40.3%) | ||
| Primary Site | Upper lobe | 1048 (59.1%) | 742 (59.8%) | 306 (57.4%) | 0.237 |
| Middle lobe | 83 (4.7%) | 56 (4.5%) | 27 (5.1%) | ||
| Lower lobe | 573 (32.3%) | 388 (31.3%) | 185 (34.7%) | ||
| Other | 69 (3.9%) | 54 (4.4%) | 15 (2.8%) | ||
| Histologic.Type | Adenocarcinoma | 1057 (59.6%) | 734 (59.2%) | 323 (60.6%) | 0.562 |
| ADSQC | 39 (2.2%) | 25 (2%) | 14 (2.6%) | ||
| SQCC | 440 (24.8%) | 318 (25.6%) | 122 (22.9%) | ||
| Other | 237 (13.4%) | 163 (13.1%) | 74 (13.9%) | ||
| Laterality | Right | 1017 (57.4%) | 704 (56.8%) | 313 (58.7%) | 0.478 |
| Left | 756 (42.6%) | 536 (43.2%) | 220 (41.3%) | ||
| Grade | I | 83 (4.7%) | 60 (4.8%) | 23 (4.3%) | 0.747 |
| II | 549 (31%) | 392 (31.6%) | 157 (29.5%) | ||
| III | 1094 (61.7%) | 755 (60.9%) | 339 (63.6%) | ||
| IV | 47 (2.7%) | 33 (2.7%) | 14 (2.6%) | ||
| T | T1 | 203 (11.4%) | 146 (11.8%) | 57 (10.7%) | 0.665 |
| T2 | 528 (29.8%) | 377 (30.4%) | 151 (28.3%) | ||
| T3 | 501 (28.3%) | 346 (27.9%) | 155 (29.1%) | ||
| T4 | 541 (30.5%) | 371 (29.9%) | 170 (31.9%) | ||
| N | N0 | 400 (22.6%) | 281 (22.7%) | 119 (22.3%) | 0.650 |
| N1 | 175 (9.9%) | 129 (10.4%) | 46 (8.6%) | ||
| N2 | 876 (49.4%) | 610 (49.2%) | 266 (49.9%) | ||
| N3 | 322 (18.2%) | 220 (17.7%) | 102 (19.1%) | ||
| Surgery | No | 1714 (96.7%) | 1200 (96.8%) | 514 (96.4%) | 0.826 |
| Yes | 59 (3.3%) | 40 (3.2%) | 19 (3.6%) | ||
| Radiation | No/Unknown | 590 (33.3%) | 408 (32.9%) | 182 (34.1%) | 0.650 |
| Yes | 1183 (66.7%) | 832 (67.1%) | 351 (65.9%) | ||
| Chemotherapy | No/Unknown | 621 (35%) | 428 (34.5%) | 193 (36.2%) | 0.528 |
| Yes | 1152 (65%) | 812 (65.5%) | 340 (63.8%) | ||
| Tumor.size | ≤ 46 | 884 (49.9%) | 619 (49.9%) | 265 (49.7%) | 0.481 |
| 47–70 | 567 (32%) | 404 (32.6%) | 163 (30.6%) | ||
| > 70 | 322 (18.2%) | 217 (17.5%) | 105 (19.7%) | ||
| Months from diagnosis to treatment | 0–1 | 1362 (76.8%) | 954 (76.9%) | 408 (76.5%) | 0.908 |
| >=2 | 411 (23.2%) | 286 (23.1%) | 125 (23.5%) | ||
| Brain metastasis | No | 1419 (80%) | 1006 (81.1%) | 413 (77.5%) | 0.090 |
| Yes | 354 (20%) | 234 (18.9%) | 120 (22.5%) | ||
| Liver metastasis | No | 1430 (80.7%) | 1009 (81.4%) | 421 (79%) | 0.271 |
| Yes | 343 (19.3%) | 231 (18.6%) | 112 (21%) | ||
| Lung metastasis | No | 1290 (72.8%) | 905 (73%) | 385 (72.2%) | 0.789 |
| Yes | 483 (27.2%) | 335 (27%) | 148 (27.8%) |
Model training, validation, and explainability
Seven machine learning algorithms—Logistic Regression, Decision Tree, Recurrent Neural Network(RNN), Light Gradient Boosting Machine, Random Forest, Extreme Gradient Boosting, Support Vector Machine —were implemented. Hyperparameter optimization for each algorithm was conducted via a combined approach of grid search and randomized search, with the optimal parameter set determined through 10-fold cross-validation on the training set to maximize the Area Under the Receiver Operating Characteristic Curve (AUC). The final models were evaluated on the held-out validation cohort. The evaluation of model performance incorporated a suite of 15 metrics, categorized into measures of discrimination (e.g., AUC, Discrimination Slope), calibration (e.g., Brier Score, Calibration Intercept), and classification accuracy (e.g., F1 Score, MCC, Youden’s Index). A composite scoring system (1–7 points per metric) was employed to holistically identify the best-performing model. To interpret the optimal model’s predictions and identify key drivers of mortality, we performed a global explainability analysis using SHapley Additive exPlanations (SHAP), which quantifies the magnitude and direction of each feature’s contribution to the model output.
Software and reproducibility
All analytical procedures, including statistical testing and data visualization, were undertaken within the R programming environment, release 4.1.2.The machine learning pipeline, including model training, hyperparameter tuning, and explainability analysis, was implemented in Python (version 3.9.7). A two-tailed P-value < 0.05 was considered statistically significant.
Results
Patient’s characteristics
A total of 1,773 participants were included in the final analysis. Figure 1 outlines the patient selection flowchart. Following the application of exclusion criteria to an initial cohort of 2,352 patients, 1,773 eligible individuals were enrolled and subsequently randomized into training (n = 1,240) and validation (n = 533) sets for model development and assessment, respectively.The cohort comprised 1,773 patients aged 70 years or older, with the majority distributed in the 70–79 (72.0%) and 80–89 (26.6%) year-old groups. The primary tumor was most frequently located in the upper lobe (59.1%), followed by the lower lobe (32.3%). Histologically, adenocarcinoma was the predominant type (59.6%), with squamous cell carcinoma (SQCC) representing 24.8% of cases. A significant proportion of patients presented with advanced disease, as 58.8% had T3 or T4 stage and 67.6% had N2 or N3 stage. Concurrent metastases to the brain, liver, and lung were observed in 20.0%, 19.3%, and 27.2% of patients, respectively. The vast majority of patients (96.7%) did not undergo surgical resection of the primary tumor, whereas 66.7% received radiation therapy.
Fig. 1.
Flowchart of the study.
Model development
Random assignment of patients yielded training and validation sets with well-balanced baseline profiles upon comparative assessment (Table 1).In multivariable logistic regression analysis, eleven of the 18 candidate variables demonstrated statistically significant independent associations with 3-month cancer-specific mortality (P < .10). These included sex, primary tumor site, histologic type, T stage, radiation therapy, chemotherapy, tumor size, time from diagnosis to treatment initiation, and the presence of brain, liver, or lung metastases (Table 2, multivariable column). These variables were subsequently incorporated into the development of prediction models. Seven machine learning algorithms—Logistic Regression, Decision Tree, Recurrent Neural Network(RNN), Light Gradient Boosting Machine, Random Forest, Extreme Gradient Boosting, Support Vector Machine—were employed for model training and optimization.
Table 2.
Predictors of 3-month mortality in training cohort patients with lung cancer and bone metastasis: Univariate and Multivariate Analyses*.
| Dependent: CSS | 0 (N = 763) | 1 (N = 477) | OR (univariable) | OR (multivariable) | |
|---|---|---|---|---|---|
| Age | 70–79 | 561 (73.5%) | 333 (69.8%) | ||
| 80–89 | 192 (25.2%) | 133 (27.9%) | 1.17 (0.90–1.51, p=.243) | ||
| >=90 | 10 (1.3%) | 11 (2.3%) | 1.85 (0.78–4.41, p=.163) | ||
| Sex | Male | 413 (54.1%) | 303 (63.5%) | ||
| Female | 350 (45.9%) | 174 (36.5%) | 0.68 (0.54–0.86, p=.001) | 0.59 (0.44–0.78, p<.001) | |
| Race | White | 630 (82.6%) | 404 (84.7%) | ||
| Black | 56 (7.3%) | 37 (7.8%) | 1.03 (0.67–1.59, p=.893) | ||
| Other | 77 (10.1%) | 36 (7.5%) | 0.73 (0.48–1.10, p=.136) | ||
| Marital.status | Married | 482 (63.2%) | 291 (61%) | ||
| Unmarried | 281 (36.8%) | 186 (39%) | 1.10 (0.87–1.39, p=.444) | ||
| Primary.Site | Upper lobe | 460 (60.3%) | 282 (59.1%) | ||
| Middle lobe | 36 (4.7%) | 20 (4.2%) | 0.91 (0.51–1.60, p=.733) | 0.95 (0.49–1.86, p=.884) | |
| Lower lobe | 241 (31.6%) | 147 (30.8%) | 0.99 (0.77–1.28, p=.969) | 1.04 (0.77–1.41, p=.779) | |
| Other | 26 (3.4%) | 28 (5.9%) | 1.76 (1.01–3.06, p=.046) | 1.48 (0.77–2.86, p=.244) | |
| Histologic.Type | Adenocarcinoma | 474 (62.1%) | 260 (54.5%) | ||
| ADSQC | 15 (2%) | 10 (2.1%) | 1.22 (0.54–2.74, p=.639) | 1.17 (0.46–2.98, p=.748) | |
| SQCC | 190 (24.9%) | 128 (26.8%) | 1.23 (0.94–1.61, p=.136) | 0.92 (0.66–1.28, p=.633) | |
| Other | 84 (11%) | 79 (16.6%) | 1.71 (1.22–2.41, p=.002) | 1.19 (0.79–1.79, p=.402) | |
| Laterality | Right | 426 (55.8%) | 278 (58.3%) | ||
| Left | 337 (44.2%) | 199 (41.7%) | 0.90 (0.72–1.14, p=.397) | ||
| Grade | I | 38 (5%) | 22 (4.6%) | ||
| II | 275 (36%) | 117 (24.5%) | 0.73 (0.42–1.30, p=.288) | ||
| III | 431 (56.5%) | 324 (67.9%) | 1.30 (0.75–2.24, p=.347) | ||
| IV | 19 (2.5%) | 14 (2.9%) | 1.27 (0.53–3.03, p=.586) | ||
| T | T1 | 102 (13.4%) | 44 (9.2%) | ||
| T2 | 242 (31.7%) | 135 (28.3%) | 1.29 (0.86–1.95, p=.221) | 1.03 (0.62–1.71, p=.910) | |
| T3 | 210 (27.5%) | 136 (28.5%) | 1.50 (0.99–2.27, p=.054) | 1.02 (0.60–1.74, p=.937) | |
| T4 | 209 (27.4%) | 162 (34%) | 1.80 (1.19–2.70, p=.005) | 1.12 (0.65–1.93, p=.689) | |
| N | N0 | 184 (24.1%) | 97 (20.3%) | ||
| N1 | 86 (11.3%) | 43 (9%) | 0.95 (0.61–1.47, p=.814) | ||
| N2 | 364 (47.7%) | 246 (51.6%) | 1.28 (0.96–1.72, p=.098) | ||
| N3 | 129 (16.9%) | 91 (19.1%) | 1.34 (0.93–1.93, p=.117) | ||
| Surgery | No | 733 (96.1%) | 467 (97.9%) | ||
| Yes | 30 (3.9%) | 10 (2.1%) | 0.52 (0.25–1.08, p=.080) | ||
| Radiation | No/Unknown | 284 (37.2%) | 124 (26%) | ||
| Yes | 479 (62.8%) | 353 (74%) | 1.69 (1.31–2.17, p<.001) | 0.63 (0.45–0.87, p=.005) | |
| Chemotherapy | No/Unknown | 142 (18.6%) | 286 (60%) | ||
| Yes | 621 (81.4%) | 191 (40%) | 0.15 (0.12–0.20, p<.001) | 0.12 (0.09–0.16, p<.001) | |
| Tumor.size | ≤ 46 | 423 (55.4%) | 196 (41.1%) | ||
| 47–70 | 231 (30.3%) | 173 (36.3%) | 1.62 (1.25–2.10, p<.001) | 1.64 (1.18–2.26, p=.003) | |
| > 70 | 109 (14.3%) | 108 (22.6%) | 2.14 (1.56–2.93, p<.001) | 1.92 (1.27–2.91, p=.002) | |
| Months from diagnosis to treatment | 0–1 | 521 (68.3%) | 433 (90.8%) | ||
| >=2 | 242 (31.7%) | 44 (9.2%) | 0.22 (0.15–0.31, p<.001) | 0.21 (0.15–0.31, p<.001) | |
| Brain.metastasis | No | 649 (85.1%) | 357 (74.8%) | ||
| Yes | 114 (14.9%) | 120 (25.2%) | 1.91 (1.44–2.55, p<.001) | 1.69 (1.19–2.41, p=.004) | |
| Liver.metastasis | No | 640 (83.9%) | 369 (77.4%) | ||
| Yes | 123 (16.1%) | 108 (22.6%) | 1.52 (1.14–2.03, p=.004) | 1.48 (1.05–2.09, p=.027) | |
| Lung.metastasis | No | 577 (75.6%) | 328 (68.8%) | ||
| Yes | 186 (24.4%) | 149 (31.2%) | 1.41 (1.09–1.82, p=.008) | 1.46 (1.04–2.05, p=.027) |
*Variable selection for model development was based on the multivariable analysis (right‑most column). Univariable results are shown for descriptive purposes only.
Selection of model predictors
We employed both univariable and multivariable logistic regression on the training cohort to determine independent covariates associated with 3-month cancer-specific mortality. Eleven clinically relevant variables demonstrated significant associations in the univariable analysis. The subsequent multivariable model, refined for parsimony and clinical relevance, retained key predictors including sex, primary tumor site, histologic type, T stage, radiation therapy, chemotherapy, tumor size, time from diagnosis to treatment initiation, and the presence of brain, liver, or lung metastases. Variables such as age, race, and marital status, which did not exhibit independent prognostic value after adjustment, were excluded from the final set of model predictors. Notably, a longer interval from diagnosis to treatment (≥ 2 months) was identified as a strong protective factor (OR, 0.21; 95% CI, 0.15–0.31; P < .001), likely serving as a proxy for less aggressive disease and a patient condition stable enough to tolerate delayed intervention.
Model development and estimation
Seven machine learning algorithms were developed and rigorously evaluated on the validation set. A comprehensive assessment across multiple dimensions consistently identified the logistic regression model as the optimal predictive tool.
Discriminatory performance and comprehensive metrics
The logistic regression model achieved the highest Area Under the ROC Curve (AUC) of 0.79 (Fig. 2), indicating superior discriminative ability. Precision-Recall analysis (Fig. 3) further supported its robust performance. For descriptive purposes, a performance heatmap comparing all seven models across 13 metrics is shown in Fig. 4; it illustrates that logistic regression performed consistently well across multiple metrics, although model selection was not based on this composite comparison. Its calibration curve was near‑perfect (Fig. 5), and Decision Curve Analysis (Fig. 6) demonstrated favorable net clinical benefit across a range of threshold probabilities. Based on these standard criteria — prioritizing AUC, followed by calibration and clinical utility — the logistic regression model was selected as the optimal tool.
Fig. 2.

Receiver operating characteristic (ROC) curves of the seven machine learning models on the validation set.
Fig. 3.

Precision-recall (PR) curves of the seven machine learning models on the validation set.
Fig. 4.

Comprehensive performance evaluation of the seven machine learning models.
Fig. 5.
Calibration plots of the seven machine learning models on the validation set.
Fig. 6.

Decision curve analysis (DCA) of the seven machine learning models on the validation set.
Calibration and clinical utility
Critically, the logistic regression model demonstrated near-perfect calibration (Fig. 5), with its predicted probabilities closely aligning with observed outcomes across the entire risk spectrum, ensuring trustworthiness for individual-level risk estimation. Decision Curve Analysis (DCA, Fig. 6) confirmed its superior clinical utility, demonstrating the highest net benefit across the majority of clinically reasonable threshold probabilities (5%–60%). The established optimal risk cut-off of 40.0% resides within this high-utility range.
Model stability
Finally, 10-fold cross-validation (Fig. 7) affirmed the exceptional stability of the logistic regression model, which achieved the highest mean accuracy (0.78) with the smallest standard deviation (± 0.02). This indicates consistent and reliable performance, robust to variations in the training data.
Fig. 7.

Model stability assessment through 10-fold cross-validation on the training set.
Based on its exceptional and consistent performance across all evaluation frameworks—encompassing discrimination, calibration, clinical utility, and stability—the logistic regression model was selected as the final model for clinical application.
Model explainability
The mean absolute SHAP value analysis (Fig. 8A) established a clear hierarchy of predictive importance among the 18 clinical variables. Chemotherapy emerged as the most influential predictor, demonstrating the largest average impact on model output. This was closely followed by time from diagnosis to treatment and patient sex, which constituted the secondary tier of important predictors. Conventional tumor characteristics including tumor size, T stage, and N stage displayed moderate but substantial contributions, while the presence of brain and liver metastases also ranked among the clinically significant predictors. Notably, demographic factors such as age, race, and marital status, along with AJCC stage and surgical intervention, exhibited minimal impact on the model’s predictions, consistent with the advanced disease status of this specific cohort.
Fig. 8.
Interpretability of the optimal logistic regression model using SHAP (SHapley Additive exPlanations).
The SHAP value distribution plot (Fig. 8B) provided crucial insights into the directional relationship between each feature and mortality risk. Chemotherapy displayed exclusively negative SHAP values (blue distribution), unequivocally confirming its role as a potent protective factor. The time from diagnosis to treatment revealed a particularly insightful pattern: longer intervals (represented by red points) were consistently associated with negative SHAP values, indicating that delayed treatment initiation serves as a surrogate marker for less aggressive disease biology rather than a causal factor for poor outcomes.
Conversely, liver metastases and brain metastases exhibited strongly positive SHAP values (red distributions), identifying them as primary risk factors for early mortality. For continuous variables, we observed clear dose-response relationships: increasing tumor size (color gradient from blue to red) demonstrated a progressive transition from negative to positive SHAP values, indicating a graded association with elevated mortality risk.
Discussion
In this retrospective cohort study, we developed and validated a machine learning-based prediction model for 3-month cancer-specific mortality in a vulnerable population of elderly patients (≥ 70 years) with NSCLC and bone metastases. By rigorously evaluating seven machine learning algorithms, we identified the logistic regression model as the optimal predictive tool. It demonstrated superior and stable performance across a comprehensive set of metrics, including discrimination (AUC 0.79), calibration, and clinical utility. Furthermore, SHAP analysis provided crucial model interpretability, confirming chemotherapy as the strongest protective factor while revealing that shorter treatment intervals serve as proxies for more aggressive disease rather than causal factors. Key predictors identified—such as the administration of chemotherapy, and the presence of liver or brain metastases—align with established clinical knowledge and provide a transparent, interpretable framework for risk assessment13. Compared with existing models such as that by Yap et al.14, our work focuses specifically on elderly NSCLC patients, uses 3‑month cancer‑specific mortality as the endpoint, relies solely on routine SEER variables, and provides machine‑learning‑based interpretability via SHAP.
The superior performance of the logistic regression model, compared to more complex algorithms like XGBoost and neural networks, can be attributed to several factors. First, in scenarios where the relationships between predictors and outcome are approximately linear or monotonic, logistic regression is a highly efficient and robust estimator15. Second, its inherent simplicity confers significant advantages in terms of model interpretability and lower risk of overfitting, especially with a moderate sample size and a curated set of clinically relevant predictors. While ensemble methods like LightGBM showed competitive discriminative ability (AUC 0.784), their performance exhibited greater variability in cross-validation, and they demonstrated poorer calibration, systematically overestimating risk. Our findings resonate with a growing body of literature suggesting that for many clinical prediction tasks, well-specified linear models can outperform or match the performance of “black-box” algorithms, while offering crucial advantages in transparency and clinical trust16–19.The predictors identified by our model hold significant clinical and biological plausibility. The strong protective effect of chemotherapy is consistent with its role in controlling systemic disease burden, including micrometastases, thereby prolonging survival even in the palliative setting19.Conversely, the presence of liver and brain metastases were among the most influential risk factors20. Liver metastasis often indicates a high-volume, aggressive disease phenotype with compromised organ function, while brain metastasis involves a privileged anatomical site where the blood-brain barrier can limit therapeutic efficacy21,22.A particularly insightful finding was the association between a shorter time from diagnosis to treatment initiation and a higher risk of mortality. This likely does not imply that treatment delay is beneficial, but rather acts as a proxy for a more aggressive and symptomatic disease presentation, necessitating immediate intervention. Patients with indolent disease or better performance status can often tolerate a longer work-up period, thus being categorized in the lower-risk group23.
Several limitations of our study warrant consideration. First, its retrospective design is inherently susceptible to selection and unmeasured confounding biases. Second, while the SEER database provides a large, population-based sample, our model currently lacks external validation in independent cohorts — a critical step to confirm generalizability across different populations, time periods, and healthcare settings. Therefore, the model should be considered preliminary; its predictions are best used as a research tool or an adjunct to clinical judgment rather than as the sole basis for treatment decisions. We are actively planning external validation using our institutional data and other public databases (e.g., NCDB) to further strengthen the model’s reliability and clinical utility. Third, the database lacks detailed information on several potentially crucial prognostic factors, such as patient comorbidities, performance status, biochemical markers, specific bone metastasis characteristics (e.g., number, location), and the use of modern systemic therapies like immunotherapy. Finally, despite robust internal validation, the model’s performance must be confirmed in a prospective, multi-center setting before widespread clinical implementation.
Notwithstanding these limitations, our study presents a readily implementable tool for personalizing management in a clinically challenging scenario. The model can assist clinicians in identifying patients at high risk of early death, for whom the focus may justifiably shift towards aggressive palliative and supportive care to maximize quality of life. Conversely, for patients predicted to have a better prognosis, more definitive treatments can be considered with greater confidence. Future research should focus on the external validation of this model, the integration of additional clinical and molecular variables (e.g., genomic markers, circulating tumor DNA), and the development of a user-friendly clinical decision support interface to facilitate its adoption in routine practice.
Conclusion
In conclusion, we have successfully developed and validated an interpretable machine learning model that accurately predicts 3-month mortality in elderly NSCLC patients with bone metastases. The logistic regression model, distinguished by its discriminative power, excellent calibration, and clinical applicability, provides a reliable means of risk stratification. This preliminary tool holds the potential to inform shared decision-making, optimize resource allocation, and ultimately, improve the quality of care for this vulnerable patient population, but should only be used for research purposes or as a hypothesis‑generating aid until external validation confirms its reliability.
Author contributions
1. T.G., X.Y., W.W., and L.Q. collectively designed the study and established its conceptual framework. C.R., X.Z., and Q.L. carried out the analytical procedures, interpreted the findings, and prepared the initial manuscript draft. Supervision and project administration were handled by X.X., A.G., and Y.S. Every author has reviewed the final manuscript and attests to their personal accountability for all aspects of the work.
Funding
This work was funded by the Nanjing Medical University(NMUB20230039) and the Fourth Affiliated Hospital of Nanjing Medical University(23YJRC24). This work was supported by the Scientific Research Project of Nanjing Municipal Health Commission (YKK24239).
Data availability
The datasets produced and analyzed in the course of this research are not publicly deposited due to specific restrictions. Nonetheless, reasonable access requests will be evaluated by the corresponding author, and data may be shared accordingly.
Declarations
Competing interests
The authors declare no competing interests.
Ethical approval
The institutional ethics committee waived the mandate for formal ethical approval and written informed consent for this investigation. The study was conducted under the provisions of local legislation and institutional guidelines, which grant this exemption.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Tian Gao, Li Qian, Cheng Rao and Xinjian Zhou.
Contributor Information
Wei Wang, Email: wangwei60215_njmu@163.com.
Jie Yin, Email: njyinjie@soho.com.
References
- 1.Kim, S. Y., Park, H. S. & Chiang, A. C. Small Cell Lung Cancer: A Review. JAMA333 (21), 1906–1917. 10.1001/jama.2025.0560 (2025). [DOI] [PubMed] [Google Scholar]
- 2.Wolf, A. M. D. et al. Screening for lung cancer: 2023 guideline update from the American Cancer Society. CA Cancer J. Clin.74 (1), 50–81. 10.3322/caac.21811Arabi (2024). [DOI] [PubMed] [Google Scholar]
- 3.Xue, M. et al. New insights into non-small cell lung cancer bone metastasis: mechanisms and therapies. Int. J. Biol. Sci.20 (14), 5747–5763. 10.7150/ijbs.100960 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Del Conte, A. et al. Bone metastasis and immune checkpoint inhibitors in non-small cell lung cancer (NSCLC): microenvironment and possible clinical implications. Int. J. Mol. Sci.10.3390/ijms23126832 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Knapp, B. J., Devarakonda, S. & Govindan, R. Bone metastases in non-small cell lung cancer: a narrative review. J. Thorac. Dis.14 (5), 1696–1712. 10.21037/jtd-21-1502 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Montrone, M. et al. Immunotherapy in elderly patients affected by non-small cell lung cancer: a narrative review. J. Clin. Med.10.3390/jcm12051833 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Tsukita, Y. et al. Immunotherapy or Chemoimmunotherapy in Older Adults With Advanced Non-Small Cell Lung Cancer. JAMA Oncol.10 (4), 439–447. 10.1001/jamaoncol.2023.6277 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Morimoto, T. et al. A new era in the management of spinal metastasis. Front. Oncol.14, 1374915. 10.3389/fonc.2024.1374915 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Rozera, T., Pasolli, E., Segata, N. & Ianiro, G. Machine learning and artificial intelligence in the multi-omics approach to gut microbiota. Gastroenterology169 (3), 487–501. 10.1053/j.gastro.2025.02.035 (2025). [DOI] [PubMed] [Google Scholar]
- 10.Zhu, S. et al. Artificial intelligence, machine learning and big data in radiation oncology. Hematol. Oncol. Clin. North. Am.39 (2), 453–469. 10.1016/j.hoc.2024.12.002 (2025). [DOI] [PubMed] [Google Scholar]
- 11.Sartori, F. et al. A Comprehensive review of deep learning applications with multi-omics data in cancer research. Genes (Basel). 16 (6), 648. 10.3390/genes16060648 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Yan, X. et al. Development and validation of a novel prognostic prediction system based on GLIM-defined malnutrition for colorectal cancer patients post-radical surgery. Front. Nutr.11, 1425317. 10.3389/fnut.2024.1425317 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Valencia, G. et al. First-line (1L) treatment decision patterns and survival of hormone receptor (HR)-Positive/HER2-negative advanced breast cancer (abc) patients in a Latin American (LATAM) Public institution. Curr. Oncol.31 (12), 7890–7902. 10.3390/curroncol31120581 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Yap, W. K. et al. Development and validation of a nomogram for assessing survival in patients with metastatic lung cancer referred for radiotherapy for bone metastases. JAMA Netw. Open.1 (6), e183242. 10.1001/jamanetworkopen.2018.3242 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Chan, J. J. L. et al. Inequalities in the prevalence of cardiovascular disease risk factors in Brazilian slum populations: A cross-sectional study. PLOS Glob Public. Health. 2 (9), e0000990. 10.1371/journal.pgph.0000990 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Han, F. et al. Artificial intelligence in orthopedic surgery: current applications, challenges, and future directions. MedComm6 (7), e70260. 10.1002/mco2.70260 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Marquand, A. F. et al. Learning latent profiles via cognitive growth charting in psychosis: design and rationale for the PRECOGNITION project. Schizophr Bull. Open.6 (1), sgaf007. 10.1093/schizbullopen/sgaf007 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Pahar, M., Miranda, I., Diacon, A. & Niesler, T. Automatic non-invasive cough detection based on accelerometer and audio signals. J. Signal. Process. Syst.94 (8), 821–835. 10.1007/s11265-022-01748-5 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Tian, D. et al. Surgical resection of primary tumors improved the prognosis of patients with bone metastasis of non-small cell lung cancer: a population-based and propensity score-matched study. Ann. Transl Med.9 (9), 775. 10.21037/atm-21-540 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Tomioka-Inagawa, R. et al. The impact of neutrophil-to-lymphocyte ratio after two courses of pembrolizumab for oncological outcomes in patients with metastatic urothelial carcinoma. Biomedicines10 (7), 1609. 10.3390/biomedicines10071609 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Ni, X. et al. From biology to the clinic - exploring liver metastasis in prostate cancer. Nat. Rev. Urol.21 (10), 593–614. 10.1038/s41585-024-00875-x (2024). [DOI] [PubMed] [Google Scholar]
- 22.Xu, B. et al. Efficacy and safety of East Asian herbal medicine for brain metastases in non-small cell lung cancer: a systematic review and meta-analysis protocol to identify specific herbs. Integr. Cancer Ther.22, 15347354221150001 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Karthaus, M. et al. Subgroup analyses from patients with pre-treated metastatic colorectal cancer receiving trifluridine/tipiracil: results of the TALLISUR trial. BMC Cancer. 24 (1), 887. 10.1186/s12885-024-12599-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The datasets produced and analyzed in the course of this research are not publicly deposited due to specific restrictions. Nonetheless, reasonable access requests will be evaluated by the corresponding author, and data may be shared accordingly.



