Abstract
Background
Head and neck squamous-cell carcinoma (HNSCC) accounts for ∼5.3% of cancer-related mortality worldwide, with an estimated 890 000 new diagnoses and 450 000 deaths annually. Despite curative-intent therapy, 10% to 50% of patients experience recurrence. Prognosis for recurrent or metastatic disease is poor, with limited treatment options, underscoring the need for accurate prognostic models to guide treatment escalation or de-escalation and avoid over-treatment.
Methods
We conducted a multicenter prognostic study of patients undergoing curative-intent surgery at Samsung Medical Center and Massachusetts Eye and Ear Infirmary/Massachusetts General Hospital from 2008 to 2024. Baseline clinicopathologic variables were integrated with longitudinal laboratory measurements from surveillance. A random 80/20 split defined development and internal-validation cohorts. Using XGBoost, we trained two models to predict recurrence-free survival (RFS) and overall survival (OS) at 1, 2, 3, 4, and 5 years from each visit.
Results
A total of 975 patients with HNSCC (oral cavity, oropharyngeal, hypopharyngeal, and laryngeal subsites) were included. The areas under the curve (AUCs) for predicting 1-, 2-, 3-, 4-, and 5-year RFS from the surveillance time point were 0.785 (sensitivity, 72.8%; specificity, 71.5%), 0.831 (79.7%; 73.7%), 0.788 (74.0%; 73.3%), 0.769 (72.6%; 70.5%), and 0.795 (72.1%; 74.7%), respectively. For OS prediction, AUCs were 0.788 (72.1%; 73.6%), 0.797 (75.7%; 71.8%), 0.796 (81.0%; 68.4%), 0.820 (77.5%; 76.5%), and 0.815 (75.8%; 75.8%), respectively. In subgroup analysis, the model showed strong OS prediction in human papilloma virus (HPV)-positive oropharyngeal cancer, with AUCs of 0.943, 0.736, 0.699, 0.835, and 0.765 at 1-, 2-, 3-, 4-, and 5-years, respectively. In non-HPV-positive HNSCC, OS AUCs ranged from 0.780 to 0.813 and RFS AUCs from 0.774 to 0.830 across the same time points.
Conclusions and relevance
In this multicenter study, an artificial intelligence (AI)-powered model using multimodal and longitudinal data accurately predicted RFS and OS at multiple time points following curative-intent surgery for HNSCC.
Key words: head and neck cancer, multicenter, real-time multimodal model, artificial intelligence
Highlights
-
•
AI model integrates baseline clinicopathology and longitudinal laboratory data after curative HNSCC surgery.
-
•
Dynamic, time-updated 1-5 year predictions for RFS and OS.
-
•
Multicenter cohort (SMC, MEEI/MGH; N = 975) with 80/20 split for development and validation.
-
•
Strong performance: RFS AUC 0.769-0.831; OS AUC 0.788-0.820 with good calibration.
-
•
Enables risk-adapted surveillance and potential electronic medical records integration for real-time support.
Introduction
Head and neck squamous-cell carcinoma (HNSCC) encompasses a group of epithelial malignancies originating in the mucosal surfaces of the oral cavity, oropharynx, hypopharynx, and larynx. With >890 000 new cases and 450 000 deaths reported annually, HNSCC remains a significant global health concern.1 Curative-intent surgery, often combined with adjuvant radiotherapy or chemoradiotherapy, is the standard treatment of local or locally advanced diseases. However, recurrence remains a critical challenge, occurring in up to 50% of patients depending on the risk level, and survival outcomes remain highly variable. These patterns underscore the need for improved tools to identify high-risk individuals who may benefit from closer monitoring or early therapeutic intervention. Beyond TNM (tumor–node–metastasis) staging, no validated predictive model currently exists to accurately stratify recurrence risk. Such models are essential to refine surveillance strategies and to enable evidence-based escalation or de-escalation of therapy, thereby minimizing the risks of under- and over-treatment.
The current standard for post-operative surveillance in HNSCC involves routine imaging and clinical assessments, guided largely by consensus guidelines rather than individualized risk profiles. Although regular clinical examinations are clearly recommended, the guidance for imaging beyond the initial scans after treatment remains relatively vague. This approach, although useful for detecting radiographic recurrence, offers limited prognostic value and may delay intervention in patients experiencing early molecular or biochemical relapse. To address these limitations, recent research has focused on novel biomarkers, such as circulating tumor DNA (ctDNA) or minimal residual disease assays for earlier recurrence detection.2,3 Although promising, these approaches remain limited by high cost, technical complexity, and restricted availability in real-world settings.
Traditional risk stratification tools in HNSCC have focused on static clinical or pathological features such as the American Joint Committee on Cancer (AJCC) TNM stage, resection margin, extra-nodal extension, or other pathologic features. However, these variables fail to account for longitudinal changes in a patient’s biological state—such as trends in inflammatory markers, nutritional status, or organ function—during the post-operative period. Incorporating such dynamic information could provide a more accurate and clinically useful assessment of recurrence risk.
Artificial intelligence (AI) has the potential to overcome limitations of conventional prognostic tools by integrating diverse high-dimensional data and enabling individualized, time-updated risk estimation.
In oncology, there have been attempts to apply AI for diagnostic purposes using radiomic, histologic, and molecular data.4, 5, 6 However, no prior studies have applied AI in the longitudinal setting after treatment for HNSCC, particularly using both baseline and follow-up data to provide real-time predictions of recurrence and survival during clinical surveillance. This study developed a multimodal AI-powered prediction model that integrates baseline clinicopathologic variables with serial laboratory data to estimate recurrence-free survival (RFS) and overall survival (OS) following curative-intent surgery for HNSCC. To enhance generalizability of the prediction tool, we developed the AI model with an international dataset comprising patients from Samsung Medical Center (SMC, Republic of Korea) and Massachusetts Eye and Ear Infirmary/Massachusetts General Hospital (MEEI/MGH, United States of America), representing different health care systems and diverse ethnic groups. Unlike static risk scores, our approach enables dynamic, time-updated risk estimation, providing clinicians with actionable insights at each follow-up visit. By leveraging routinely available clinical and laboratory data, this model has the potential for seamless integration into electronic medical records (EMRs), facilitating personalized surveillance strategies without the need for additional testing or cost.
Methods
Study design, patient selection and data collection
This retrospective, multicenter cohort study was designed to develop and validate an AI-based prognostic model for patients with AJCC stage I-IVB HNSCC who underwent curative-intent surgical resection. This study included patients treated at SMC or MEEI/MGH between January 2008 and March 2024. Eligibility criteria included histologically confirmed HNSCC of the oral cavity, oropharynx, hypopharynx, or larynx, treated with curative-intent surgery. Patients who received definitive non-surgical treatment or had evidence of distant metastasis at initial diagnosis were excluded.
For each patient, detailed baseline clinical and pathological characteristics were collected, including demographic variables, Eastern Cooperative Oncology Group (ECOG) performance status, smoking and alcohol history, AJCC TNM stage, primary tumor site, and adjuvant treatment status and modality (chemotherapy, radiotherapy). In addition, longitudinal laboratory data—including complete blood counts, liver and renal function tests, and inflammatory markers—were extracted from EMRs at multiple time points during post-operative surveillance (Supplementary Table S1, available at https://doi.org/10.1016/j.esmoop.2025.106046).
Patients in both cohorts were followed until death or last known clinical contact, and those with missing essential variables—including survival status, date of surgery, or core clinical predictors—were excluded from the analysis.
Statistical analysis
The primary outcomes of this study were OS and RFS from the surveillance monitoring point after surgery. OS was defined as the time from the date of curative-intent surgery to death from any cause or last follow-up. RFS was defined as the time from surgery to either documented disease recurrence or death from any cause, whichever occurred first. Patients without events were censored at the time of their last follow-up.
Predictive model performance was assessed using the time-dependent area under the receiver operating characteristic curve (ROC) and calculating the area under the curve (AUC) at multiple timepoints. These metrics were calculated at multiple clinically relevant time points (1-, 2-, 3-, 4- and 5-year OS and RFS). The mean AUC and standard deviation (SD) values were generated along with the F1 score. Calibration plots were generated to evaluate the agreement between predicted probabilities and observed outcomes, particularly at the 2-year landmark.
To enhance generalizability and reduce overfitting, five-fold cross-validation was carried out during model development using the SMC and MEEI/MGH cohorts. Interpretability of the machine learning model (XGBoost) was assessed using SHapley Additive exPlanations (SHAP) values to determine the contribution of each feature. In the SHAP plot, the values represent both the direction and magnitude of each feature’s contribution to the model’s prediction. Positive SHAP values indicate that the feature increases the predicted risk, negative values indicate that it decreases the predicted risk, and values close to zero suggest little to no influence on the prediction.
Model development and validation
This developed prediction models for RFS and OS from the surveillance monitoring point after surgery (Figure 1). An Extreme Gradient Boosting (XGBoost) model was configured with a learning rate of 0.1, 1000 estimators, and a maximum tree depth of 4. A total of 68 features were used for model training, including 29 baseline clinical variables, 38 laboratory variables measured over time, and 1 time-dependent variable indicating the interval from surgery to each follow-up. Before model development, all variables were processed in their original clinical format without imputation. Missing values were retained because XGBoost inherently accounts for missingness by learning optimal default split directions during training, thereby allowing the model to exploit informative missing patterns frequently encountered in real-world post-operative surveillance data. These prediction settings were considered: (i) post-surgery prediction using only baseline variables and (ii) surveillance prediction incorporating longitudinal laboratory data. To facilitate longitudinal prediction and enable clinical interpretability, we generated a dynamic risk summary metric termed the ‘Recurrence And Death AI-based Risk’ (RADAR) score, derived from the XGBoost model output at each timepoint. The RADAR score represents the model-predicted probability of recurrence or death within a specified prediction window (i.e. at 1-, 2-, 3-, 4- and 5-years) starting at time of surgery and was recalculated at each surveillance visit using updated laboratory and time-dependent inputs. These scores were subsequently used to construct calibration plots and guide interpretation of individual patient trajectories.
Figure 1.
Study scheme.
OS, overall survival; RADAR, Recurrence And Death AI-based Risk; RFS, recurrence-free survival.
Ethics statement
This study was conducted in accordance with the principles of the Declaration of Helsinki and was approved by the Institutional Review Board (IRB) of Samsung Medical Center (SMC IRB no. 2025-01-011) and Massachusetts General Hospital (IRB no. 2025P001221). The requirement for informed consent was waived due to the retrospective nature of the study and minimal risk to participants.
Results
Baseline characteristics
This study included 1062 patients with HNSCC who underwent curative-intent surgery at SMC and 95 patients in MEEI/MGH between 2008 and 2024. Among them, 182 patients were excluded due to follow-up loss within 1 year. A total of 975 patients with HNSCC were included in this study, of whom 283 (29.0%) died during the follow-up period and 692 (71.0%) survived (Table 1). The median age at surgery was 59 years (range, 18-88), and 71.3% were male. Most patients had primary tumors located in the mucosal lip and oral cavity (59.2%), followed by the larynx (28.5%) and oropharynx (12.3%). Pathologic staging revealed that 37.9% were AJCC stage I, 14.6% stage II, 10.8% stage III, and 20.1% stage IV. Regarding smoking history, 43.4% of patients were never smokers (defined as having smoked fewer than 100 cigarettes in their lifetime), 38.1% were former smokers (ever smokers who had quit smoking before the index date), and 14.1% were current smokers (ever smokers who were actively smoking at or within 30 days prior to the index date). Most patients had ECOG performance status scores of 0 or 1 (5.1% and 39.6%, respectively), and tumors were moderately (34.8%) or well differentiated (27.7%) in most cases. A total of 33.9% received adjuvant radiotherapy, and 14.5% received adjuvant chemotherapy.
Table 1.
Baseline characteristics
| Variables | Patients (N = 975) |
|---|---|
| Age, median years (range) | 59 (18-88) |
| Sex, n (%) | |
| Male | 695 (71.3%) |
| Female | 280 (28.7%) |
| Pathologic TNM stage, n (%) | |
| Stage I | 370 (37.9%) |
| Stage II | 142 (14.6%) |
| Stage III | 105 (10.8%) |
| Stage IVA-IVB | 196 (20.1%) |
| Smoking, n (%) | |
| Never-smoker | 423 (43.4%) |
| Ex-smoker | 371 (38.1%) |
| Current smoker | 137 (14.1%) |
| Primary organ, n (%) | |
| Larynx | 278 (28.5%) |
| Lip & oral cavity | 577 (59.2%) |
| Oropharynx | 120 (12.3%) |
| ECOG PS, n (%) | |
| Score 0 | 50 (5.1%) |
| Score 1 | 386 (39.6%) |
| Score 2 | 29 (3.0%) |
| Differentiation, n (%) | |
| Well | 270 (27.7%) |
| Moderately | 339 (34.8%) |
| Poorly | 77 (7.9%) |
| Adjuvant radiotherapy, n (%) | 331 (33.9%) |
| Adjuvant chemotherapy, n (%) | 141 (14.5%) |
ECOG, Eastern Cooperative Oncology Group; PS, performance status; TNM, tumor–node–metastasis.
Longitudinal surveillance-based prediction model
To evaluate the predictive utility of a dynamic, longitudinal surveillance-based model, we developed an XGBoost classifier incorporating both baseline and serial laboratory variables. Model performance was evaluated across multiple prediction windows (1 to 5 years) using five-fold cross-validation on combined datasets from SMC and MEEI/MGH.
OS prediction
For 1-year OS prediction, the model achieved a mean AUC of 0.788 (SD 0.052), with a sensitivity of 72.1%, specificity of 73.6%, and F1 score of 0.549. At 2 years, the AUC was 0.797 (SD 0.044), with sensitivity of 75.7%, specificity of 71.8%, and F1 score of 0.667 (Table 2). The model maintained strong performance over time, with AUCs of 0.796, 0.820, and 0.815 at 3, 4, and 5 years, respectively. Sensitivity ranged from 72.1% to 81.0%, specificity from 68.4% to 76.5%, and the F1 score reached 0.790 at 5 years. Figure 2A presents the ROC curves for OS prediction at each time point. To further evaluate model performance, we calculated the concordance index (C-index) using five-fold cross-validation. For OS, the mean C-index was 0.828 ± 0.008, with 95% confidence intervals (CIs) across folds ranging from 0.807 to 0.847, indicating strong and stable concordance between predicted and observed survival outcomes (Supplementary Table S2, available at https://doi.org/10.1016/j.esmoop.2025.106046). These results indicate consistently high concordance between predicted and observed survival times across folds, demonstrating robust discriminative ability of the model. We carried out Decision Curve Analysis (DCA) to further assess the clinical utility of the model. For OS, our model demonstrated a higher net benefit than both the Treat-All and Treat-None strategies across a wide threshold probability range (0.15-0.85) (Supplementary Figure S1, available at https://doi.org/10.1016/j.esmoop.2025.106046). This indicates that the model provides meaningful benefit when used to guide clinical decisions related to survival prediction.
Table 2.
Performance metrics of XGBoost models for overall survival and recurrence-free survival prediction in head and neck squamous-cell carcinoma patients across Samsung Medical Center and Massachusetts General Hospital datasets
| OS | AUC | Sensitivity | Specificity | F1 score | RFS | AUC | Sensitivity | Specificity | F1 score | ||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 12-months OS | Mean (SD) | 0.788 (0.052) | 72.1% (4.4%) | 73.6% (6.1%) | 0.549 (0.083) | 12-months RFS | Mean (SD) | 0.785 (0.044) | 72.8% (7.7%) | 71.5% (3.7%) | 0.587 (0.079) |
| 24-months OS | Mean (SD) | 0.797 (0.044) | 75.7% (3.1%) | 71.8% (7.7%) | 0.667 (0.046) | 24-months RFS | Mean (SD) | 0.831 (0.013) | 79.7% (2.3%) | 73.7% (1.8%) | 0.726 (0.027) |
| 36-months OS | Mean (SD) | 0.796 (0.041) | 81.0% (3.8%) | 68.4% (8.7%) | 0.726 (0.023) | 36-months RFS | Mean (SD) | 0.788 (0.011) | 74.0% (2.8%) | 73.3% (2.5%) | 0.718 (0.020) |
| 48-months OS | Mean (SD) | 0.820 (0.048) | 77.5% (7.7%) | 76.5% (5.1%) | 0.771 (0.050) | 48-months RFS | Mean (SD) | 0.769 (0.050) | 72.6% (6.3%) | 70.5% (7.5%) | 0.728 (0.057) |
| 60-months OS | Mean (SD) | 0.815 (0.049) | 75.8% (7.5%) | 75.8% (1.9%) | 0.790 (0.047) | 60-months RFS | Mean (SD) | 0.795 (0.059) | 72.1% (4.7%) | 74.7% (7.0%) | 0.770 (0.053) |
AUC, area under the curve; OS, overall survival; RFS, recurrence-free survival; SD; standard deviation.
Figure 2.
Receiver operating characteristic curve. (A) Overall survival. (B) Recurrence-free survival.
AUC, area under the curve.
RFS prediction
For 1-year RFS prediction, the model achieved a mean AUC of 0.785 (SD 0.044), with a sensitivity of 72.8%, specificity of 71.5%, and F1 score of 0.587. At 2 years, performance improved with an AUC of 0.831 (SD 0.013), sensitivity of 79.7%, specificity of 73.7%, and F1 score of 0.726. At 3, 4, and 5 years, the AUCs were 0.788, 0.769, and 0.795, respectively, with F1 scores increasing to 0.770 at 5 years (Table 2). Figure 2B shows the ROC curves for RFS prediction. For RFS, the model achieved a mean C-index of 0.796 ± 0.009 (95% CI 0.766-0.821), demonstrating similarly robust discriminative ability for recurrence prediction (Supplementary Table S2, available at https://doi.org/10.1016/j.esmoop.2025.106046). For RFS of the DCA, the model consistently outperformed both Treat-All and Treat-None strategies across the entire threshold range, further reinforcing its potential usefulness as a decision-support tool in post-operative surveillance (Supplementary Figure S2, available at https://doi.org/10.1016/j.esmoop.2025.106046). The Treat-None strategy (i.e. no patient treated) consistently yielded lower net benefit relative to our model. Because net benefit reflects the balance between identifying true recurrence (true positives) and minimizing unnecessary interventions (false positives), these findings suggest that our model offers practical clinical value and may support more informed decision-making in real-world practice.
Subgroup analysis by cancer site and human papilloma virus status
Subgroup analyses were conducted to compare model performance in human papilloma virus (HPV)-positive oropharyngeal cancer versus non-HPV-positive HNSCC (HPV-negative oropharyngeal cancer and non-oropharyngeal cancer) (Table 3).
Table 3.
OS and RFS by cancer sites (HPV status): HPV-positive oropharynx versus non-HPV-positive HNSCC
| OS | HPV-positive oropharynx |
Non-HPV-positive HNSCC |
||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AUC | Sensitivity | Specificity | F1 score | AUC | Sensitivity | Specificity | F1 score | |||
| 12-months OS | Mean (SD) | 0.943 (—) | 100.0% (—) | 84.3% (—) | 0.545 (—) | Mean (SD) | 0.780 (0.059) | 72.8% (4.4%) | 71.7% (6.9%) | 0.556 (0.081) |
| 24-months OS | Mean (SD) | 0.736 (0.264) | 100.0% (0.0%) | 68.4% (31.6%) | 0.697 (0.303) | Mean (SD) | 0.798 (0.044) | 78.1% (2.7%) | 69.9% (8.8%) | 0.683 (0.035) |
| 36-months OS | Mean (SD) | 0.699 (0.166) | 91.7% (8.3%) | 60.0% (11.0%) | 0.195 (0.013) | Mean (SD) | 0.778 (0.056) | 77.7% (6.0%) | 67.5% (11.5%) | 0.717 (0.028) |
| 48-months OS | Mean (SD) | 0.835 (0.139) | 75.2% (15.2%) | 98.0% (2.0%) | 0.817 (0.067) | Mean (SD) | 0.813 (0.050) | 78.2% (7.7%) | 75.1% (4.9%) | 0.780 (0.051) |
| 60-months OS | Mean (SD) | 0.765 (0.135) | 78.8% (21.1%) | 86.5% (7.5%) | 0.451 (0.189) | Mean (SD) | 0.801 (0.053) | 72.4% (6.5%) | 78.4% (7.9%) | 0.784 (0.040) |
| RFS | HPV-positive oropharynx |
Non-HPV-positive HNSCC |
||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| AUC | Sensitivity | Specificity | F1 score | AUC | Sensitivity | Specificity | F1 score | |||
| 12-months RFS | Mean (SD) | 0.773 (0.155) | 78.6% (13.8%) | 74.9% (12.0%) | 0.432 (0.094) | Mean (SD) | 0.781 (0.046) | 71.9% (7.3%) | 71.2% (4.3%) | 0.595 (0.076) |
| 24-months RFS | Mean (SD) | 0.647 (0.207) | 76.8% (14.4%) | 65.6% (20.9%) | 0.333 (0.136) | Mean (SD) | 0.830 (0.021) | 79.9% (2.7%) | 73.4% (3.2%) | 0.737 (0.014) |
| 36-months RFS | Mean (SD) | 0.618 (0.231) | 93.4% (10.6%) | 50.4% (19.8%) | 0.334 (0.200) | Mean (SD) | 0.786 (0.010) | 75.6% (3.2%) | 70.4% (2.6%) | 0.734 (0.011) |
| 48-months RFS | Mean (SD) | 0.677 (0.209) | 94.5% (5.8%) | 54.3% (21.0%) | 0.416 (0.233) | Mean (SD) | 0.774 (0.047) | 70.5% (6.0%) | 74.0% (8.8%) | 0.733 (4.8%) |
| 60-months RFS | Mean (SD) | 0.592 (0.238) | 68.0% (23.1%) | 68.4% (22.8%) | 0.469 (0.284) | Mean (SD) | 0.794 (0.062) | 72.2% (3.7%) | 75.5% (9.3%) | 0.782 (0.046) |
AUC, area under the curve; HPV, human papillomavirus; OS, overall survival; RFS, recurrence-free survival; SD; standard deviation.
OS prediction
In HPV-positive oropharyngeal cancer, the model demonstrated strong short-term predictive performance for OS. At 1 year, the AUC reached 0.943, with a sensitivity of 100.0% and specificity of 84.3%, resulting in an F1 score of 0.545. At 2 years, performance was lower, with an AUC of 0.736, sensitivity of 100.0%, specificity of 68.4%, and an F1 score of 0.697. At 3 years, although sensitivity remained high (91.7%), specificity decreased to 60.0%, resulting in a reduced F1 score of 0.195 with an AUC of 0.699. At 4 and 5 years, AUCs were 0.835 and 0.765, respectively. Sensitivity and specificity were 75.2% and 98.0% at 4 years, and 78.8% and 86.5% at 5 years, yielding F1 scores of 0.817 and 0.451.
Non-HPV-positive HNSCC demonstrated more consistent performance over time. The AUCs for OS prediction were 0.780, 0.798, 0.778, 0.813, and 0.801 at 1-, 2-, 3-, 4-, and 5-year time points, respectively. Sensitivity ranged from 72.4% to 78.2%, and specificity from 67.5% to 78.4%. The F1 score gradually improved over time, with values of 0.556, 0.683, 0.717, 0.780, and 0.784 at 1 through 5 years, respectively, reflecting balanced classification performance in the non-HPV subgroup.
RFS prediction
For HPV-positive oropharyngeal cancer, the model’s performance for RFS prediction was more variable. At 1 year, the AUC was 0.773, with a sensitivity of 78.6%, specificity of 74.9%, and an F1 score of 0.432. At 2 years, the AUC dropped to 0.647, with a sensitivity of 76.8% and specificity of 65.6%, resulting in an F1 score of 0.333. At 3 years, despite a high sensitivity of 93.4%, specificity fell to 50.4%, and the F1 score remained low at 0.334 (AUC 0.618). At 4 years, the AUC was 0.677, with sensitivity and specificity of 94.5% and 54.3%, respectively, and an F1 score of 0.416. At 5 years, the model achieved an AUC of 0.592, sensitivity of 68.0%, specificity of 68.4%, and an F1 score of 0.469.
In non-HPV-positive HNSCC, the model showed higher and more stable RFS prediction performance across all time points. AUCs were 0.781, 0.830, 0.786, 0.774, and 0.794 at 1 to 5 years, respectively. Sensitivity ranged from 70.5% to 79.9%, and specificity from 70.4% to 75.5%. F1 scores remained consistently high across time points, with values of 0.595, 0.737, 0.734, 0.733, and 0.782, indicating robust and balanced performance in the non-HPV subgroup.
Model interpretability
Figure 3 presents calibration curves illustrating the agreement between predicted and observed probabilities for 2-year OS (Figure 3A) and RFS (Figure 3B) based on the RADAR scores. The curves demonstrate good calibration for both outcomes, indicating that the model’s predicted risks closely align with actual clinical outcomes at this time point, thereby supporting the reliability of the probability estimates.
Figure 3.
Calibration curve. (A) 2-year overall survival rate. (B) 2-year recurrence-free survival rate.
RADAR, Recurrence And Death AI-based Risk.
Supplementary Figure S3, available at https://doi.org/10.1016/j.esmoop.2025.106046 shows SHAP summary plots from the XGBoost models, highlighting the most influential predictors across 1- to 5-year prediction windows. Consistently important variables included ECOG performance status, age, primary tumor size (long and short axis), T and N classification, serum albumin, hemoglobin, neutrophil count, lymphocyte count, and inflammatory markers such as C-reactive protein. N classification appeared to play a more prominent role in predicting long-term outcomes, particularly for 4- and 5-year OS and RFS predictions.
Discussion
This study demonstrates the feasibility of using AI-driven models to generate dynamic real-time risk estimates of recurrence and survival following curative-intent surgery in patients with HNSCC.
The novelty of our model lies in its multimodal and temporally adaptive architecture. Most existing prognostic models depend on static variables such as TNM stage or treatment modality; our approach dynamically incorporates serial trends in routine laboratory biomarkers (e.g. hemoglobin, albumin, inflammatory markers), thereby capturing ongoing physiological changes that may precede clinical relapse. The model’s ability to continuously update risk estimates over time represents a meaningful advancement beyond conventional baseline-only approaches. One limitation is that longitudinal biomarker changes may reflect factors beyond recurrence, such as infection or treatment effects, introducing variability. Nonetheless, our model uses serial data and machine learning to identify patterns predictive of recurrence, reducing noise from unrelated changes. Future studies should explore methods to further improve specificity by addressing these confounders.
The model demonstrated strong predictive performance across multiple time points. For OS, AUCs ranged from 0.788 to 0.820, and for RFS, from 0.785 to 0.795 across 1- to 5-year prediction windows. Subgroup analysis by HPV status further supports the robustness of the model. In HPV-positive oropharyngeal cancer, the model demonstrated excellent short-term performance for OS prediction, with AUCs of 0.943 and 0.736 at 1 and 2 years, respectively. Non-HPV-positive HNSCC showed more stable performance across time points, with 5-year AUCs reaching 0.801 for OS and 0.794 for RFS. These findings suggest that the model can capture meaningful survival patterns in biologically distinct subgroups and may assist in tailoring follow-up strategies accordingly. These results were consistent across the combined datasets from SMC and MEEI/MGH, underscoring the robustness and generalizability of the model across different institutions and patient populations of diverse ethnic backgrounds.
Compared with other recently published AI-based prognostic models for head and neck cancer, our model demonstrated superior and more consistent performance. For instance, Mansouri et al. reported a C-index of ∼0.73 ± 0.15 using a model that combined computed tomography (CT) radiomics, dose distribution, and clinical data. Although their approach highlighted the value of imaging and treatment planning data, it relied on handcrafted features and exhibited relatively high variability of the C-index.7 In contrast, our model achieved higher predictive accuracy using routinely collected clinical and laboratory data, without the need for complex imaging processing—enhancing its practicality and scalability in real-world settings. Kazmierski et al. evaluated 12 AI-based prognostic models using data from 2552 patients with head and neck cancer and reported the best performance (C-index up to 0.78) with a multitask learning approach. However, their study showed notable variability across institutions, raising concerns about generalizability.8 Our model notably incorporated longitudinal laboratory data through a dynamic surveillance framework, which further improved prediction accuracy. As shown in prior internal testing, 1-year AUCs increased from 0.784 to 0.902 for OS and from 0.788 to 0.883 for RFS, with sustained gains across extended follow-up intervals. These findings suggest that the inclusion of follow-up biomarker trends adds independent prognostic value beyond baseline features. Because these laboratory variables are already part of routine clinical workflows and EMRs, our model can be seamlessly embedded into existing digital infrastructure with minimal additional burden. Our findings are consistent with recent studies leveraging dynamic input modalities in oncology. For example, modeling tumor kinetics using response parameters has shown superior survival prediction in advanced HNSCC, supporting the premise that time-updated features capture disease trajectory more effectively than static baseline metrics.9 Similar approaches using serial imaging in nasopharyngeal carcinoma have also demonstrated the utility of longitudinal data for survival prediction.10 Moreover, a scoping review highlighted the growing application of AI methods on EMR longitudinal data for cancer prediction.11 These precedents reinforce the value of real-time surveillance data to enhance the temporal resolution of prognostic modeling.
From a methodological perspective, this study leveraged XGBoost to capture complex, nonlinear interactions among variables and employed SHAP to enhance model interpretability. As shown in Supplementary Figure S1, available at https://doi.org/10.1016/j.esmoop.2025.106046, SHAP summary plots consistently identified key predictors across 1- to 5-year OS and RFS models. The most influential features included ECOG performance status, tumor length, T and N classification, albumin, hemoglobin, neutrophil count, lymphocyte count, and C-reactive protein. These variables reflect known clinical and biological determinants of prognosis in HNSCC, supporting both the transparency and clinical validity of the model. Similarly, recent studies applying explainable AI in HNSCC have shown that models incorporating microRNA profiles and clinicopathological data can also successfully identify survival-associated features, further supporting the value of interpretable machine learning in oncologic prognostication.12
Practically, our model enables risk stratification at any point during surveillance, allowing for dynamic tailoring of imaging frequency, follow-up intervals, and consideration of adjuvant therapies. This flexible, data-driven approach has the potential to improve early relapse detection while reducing unnecessary testing in low-risk patients, thereby enhancing both clinical outcomes and cost-efficiency.
Nonetheless, several limitations must be acknowledged. Firstly, our model did not incorporate potentially valuable radiologic or genomic features such as radiomics. Secondly, although the model demonstrated robust retrospective accuracy, its true clinical utility remains to be established. Prospective validation studies are necessary to determine whether integration of the model into clinical workflows can meaningfully improve recurrence detection, guide surveillance strategies, or ultimately impact survival outcomes. Lastly, our algorithm was trained on a retrospective dataset collected prior before 2025, before perioperative immunotherapy became standard practice for high-risk patients. As such, future iterations will need to incorporate this treatment paradigm as well as relevant biomarkers, including programmed death-ligand 1 expression levels. Conceptually, our work aligns with the Continuous Individualized Risk Index framework, which dynamically aggregates biomarkers over time to generate personalized risk estimates in other malignancies.13 In parallel, radiomics-based studies in HNSCC using baseline CT features have shown promising prognostic performance.14 Integration of radiomic and longitudinal clinical features may therefore represent a logical next step toward more comprehensive risk stratification.
Future directions include integrating multimodal datasets—such as imaging, pathology, and molecular diagnostics—and conducting prospective interventional trials.15 Beyond risk stratification, the model may also inform future interventional trial designs. For example, high-risk patients identified through dynamic surveillance could be considered for treatment escalation strategies, such as intensified adjuvant therapy or early introduction of immunotherapy. Conversely, patients with favorable predicted outcomes may be eligible for treatment de-escalation trials, potentially reducing radiation or chemotherapy intensity without compromising outcomes, which has been successfully done for surgically-treated HPV-positive oropharyngeal cancer based on the ECOG3311 data, although any future de-escalation attempts would need to be done cautiously in light of multiple randomized trials that showed inferior outcomes in de-escalated treatment arms (De-ESCALaTE HPV, RTOG 1016, REACT). Embedding the model into EMRs could support continuous recalibration and real-time alerts, transforming current fixed-interval surveillance into a responsive, risk-adaptive paradigm. For example, the model could help identify patients who may benefit from intensified monitoring strategies, such as ctDNA testing or advanced imaging, thereby facilitating more personalized and timely follow-up interventions tailored to individual risk profiles.
In conclusion, this study presents an AI-powered, clinically interpretable, and dynamically adaptive prediction model for post-operative survival and recurrence in HNSCC. By bridging static baseline data with serial follow-up variables, our model offers a scalable and practical tool for individualized patient management. Continued research should explore integration with radiogenomic data, prospective clinical validation, and broader implementation across health care systems.
Acknowledgments
Funding
This work was supported by a grant from the National Research Foundation of Korea (NRF), funded by the Ministry of Science and ICT (MSIT) [grant number IRIS RS-2025-00521527].
Disclosure
The authors have declared no conflicts of interest.
Supplementary data
References
- 1.Johnson D.E., Burtness B., Leemans C.R., Lui V.W.Y., Bauman J.E., Grandis J.R. Head and neck squamous cell carcinoma. Nat Rev Dis Primers. 2020;6(1):92. doi: 10.1038/s41572-020-00224-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Flach S., Howarth K., Hackinger S., et al. Liquid BIOpsy for MiNimal RESidual DiSease Detection in Head and Neck Squamous Cell Carcinoma (LIONESS) – a personalised circulating tumour DNA analysis in head and neck squamous cell carcinoma. Br J Cancer. 2022;126(8):1186–1195. doi: 10.1038/s41416-022-01716-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Yang R., Li T., Zhang S., Shui C., Ma H., Li C. The effect of circulating tumor DNA on the prognosis of patients with head and neck squamous cell carcinoma: a systematic review and meta-analysis. BMC Cancer. 2024;24(1):1434. doi: 10.1186/s12885-024-13116-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Broggi G., Maniaci A., Lentini M., et al. Artificial intelligence in head and neck cancer diagnosis: a comprehensive review with emphasis on radiomics, histopathological, and molecular applications. Cancers (Basel) 2024;16(21) doi: 10.3390/cancers16213623. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Song B., Yadav I., Tsai J.C., Madabhushi A., Kann B.H. Artificial intelligence for head and neck squamous cell carcinoma: from diagnosis to treatment. Am Soc Clin Oncol Educ Book. 2025;45(3) doi: 10.1200/EDBK-25-472464. [DOI] [PubMed] [Google Scholar]
- 6.Kolla L., Parikh R.B. Uses and limitations of artificial intelligence for oncology. Cancer. 2024;130(12):2101–2107. doi: 10.1002/cncr.35307. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Mansouri Z., Salimi Y., Amini M., et al. Development and validation of survival prognostic models for head and neck cancer patients using machine learning and dosiomics and CT radiomics features: a multicentric study. Radiat Oncol. 2024;19(1):12. doi: 10.1186/s13014-024-02409-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Kazmierski M., Welch M., Kim S., et al. Multi-institutional prognostic modeling in head and neck cancer: evaluating impact and generalizability of deep learning and radiomics. Cancer Res Commun. 2023;3(6):1140–1151. doi: 10.1158/2767-9764.CRC-22-0152. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Atsou K., Auperin A., Guigay J., Salas S., Benzekry S. Mechanistic learning for predicting survival outcomes in head and neck squamous cell carcinoma. CPT Pharmacometrics Syst Pharmacol. 2025;14(3):540–550. doi: 10.1002/psp4.13294. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Kou J., Peng J.Y., Lv W.B., et al. A serial MRI-based deep learning model to predict survival in patients with locoregionally advanced nasopharyngeal carcinoma. Radiol Artif Intell. 2025;7(2) doi: 10.1148/ryai.230544. [DOI] [PubMed] [Google Scholar]
- 11.Moglia V., Johnson O., Cook G., de Kamps M., Smith L. Artificial intelligence methods applied to longitudinal data from electronic health records for prediction of cancer: a scoping review. BMC Med Res Methodol. 2025;25(1):24. doi: 10.1186/s12874-025-02473-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Kausar N. Machine learning and explainable artificial intelligence reveals the MicroRNAs associated with survival of head and neck squamous cell carcinoma patients. Comput Biol Chem. 2025;118 doi: 10.1016/j.compbiolchem.2025.108503. [DOI] [PubMed] [Google Scholar]
- 13.Kurtz D.M., Esfahani M.S., Scherer F., et al. Dynamic risk profiling using serial tumor biomarkers for personalized outcome prediction. Cell. 2019;178(3):699–713.e619. doi: 10.1016/j.cell.2019.06.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Vallières M., Kay-Rivest E., Perrin L.J., et al. Radiomics strategies for risk assessment of tumour failure in head-and-neck cancer. Sci Rep. 2017;7(1) doi: 10.1038/s41598-017-10371-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Duwe G., Mercier D., Wiesmann C., et al. Challenges and perspectives in use of artificial intelligence to support treatment recommendations in clinical oncology. Cancer Med. 2024;13(12) doi: 10.1002/cam4.7398. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



