Skip to main content
BMC Medical Informatics and Decision Making logoLink to BMC Medical Informatics and Decision Making
. 2026 Jun 13;26:321. doi: 10.1186/s12911-026-03619-6

Beyond classical models: LLM-driven survival analysis for breast cancer prognosis using European cancer registry data

Sergio Consoli 1,✉, Dimitris Katsimpokis 2, Antonello Meloni 3, Diego Reforgiato Recupero 3, Matthijs Sloep 2
PMCID: PMC13495157  PMID: 42288844

Abstract

Background

Survival analysis is a fundamental tool in clinical prognosis, yet traditional statistical models often struggle to capture complex, high-dimensional relationships in modern healthcare data. Recent advances in Large Language Models (LLMs) offer new opportunities for flexible and context-aware modeling. At the same time, access to real-world clinical data remains restricted due to privacy constraints, motivating the use of synthetic datasets as a privacy-preserving alternative for model development and evaluation.

Methods

We propose a survival analysis framework based on fine-tuned LLMs, evaluated on a large-scale synthetic breast cancer dataset derived from European cancer registry data. A synthetic dataset emulating a national population-based cancer registry and comprising 60,000 breast cancer patients was used following feature engineering and data imputation. We compared traditional survival analysis methods, including Cox regression and gradient boosting, with a range of fine-tuned LLMs representing encoder-only, decoder-only, and encoder–decoder architectures. Model performance was evaluated using standard survival analysis metrics accounting for censoring. To assess generalizability, the best-performing models were deployed on a real-world cohort of 183,304 patients from Dutch cancer registries.

Results

The proposed LLM-based models demonstrate competitive performance across multiple survival metrics, with consistent differences observed across model architectures. When applied to real-world registry data, models trained on synthetic data and those trained on real data retained strong performance in retrospective evaluation settings, while performance decreased under standard inference conditions where survival status was unavailable. This study provides large-scale empirical evidence that models trained on synthetic cancer registry data can generalize to real-world populations, achieving strong performance in retrospective evaluation settings and moderate performance under standard inference conditions.

Conclusion

This study demonstrates that LLMs can serve as flexible models for survival analysis. The use of high-fidelity synthetic data enables privacy-preserving model development while maintaining strong predictive performance. These findings support the integration of Generative AI methods into survival modeling pipelines, particularly in data-constrained clinical settings.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12911-026-03619-6.

Keywords: Synthetic health data, Large language models, Privacy-preserving analytics, Survival analysis, Breast cancer, Health data governance, Synthetic health data validation

Background

The integration of Artificial Intelligence (AI) into healthcare has shown transformative potential for medical research and clinical practice [1–5], with particularly notable advances in oncology [6–10]. However, its shift from controlled research environments to real-world clinical settings presents critical challenges regarding performance evaluation and clinical utility. Among these advances, AI applications in survival analysis have shown promise [11, 12], especially for breast cancer data [13–15]. Nevertheless, the use of Generative AI and Large Language Models (LLMs)—the most rapidly advancing AI technologies transforming numerous fields [16–18]—remains underexplored in the context of survival analysis for breast cancer. Generative AI and LLM-based technologies can generate human-like text and interpret complex data patterns [19, 20], offering the potential to significantly impact survival analysis research by enhancing diagnostic accuracy and informing treatment decisions in real-world practice.

LLMs are predominantly based on the Transformer architecture [16], an influential design in machine learning and AI that processes sequential data efficiently through self-attention mechanisms capable of modeling long-range dependencies in input sequences [21]. These models have greatly enhanced the ability of AI systems to process and understand complex data patterns, and may improve survival prediction tasks by providing valuable insights into time-to-event outcomes in medical research.

In this paper, we aim to address the complex challenges of survival analysis by leveraging Generative AI to enhance our understanding and predictive capabilities in breast cancer research. However, the successful deployment of these advanced AI technologies in healthcare must also operate within established regulatory and data governance frameworks.

The policy context of this work is shaped by evolving health data governance frameworks worldwide. In Europe, the European Network of Cancer Registries (ENCR)1 and the European Health Data Space (EHDS)2 provide a structured environment for data standardization and sharing, with EHDS regulation published on March 20253. Similar initiatives exist globally: in the United States, the Health Insurance Portability and Accountability Act (HIPAA)4 establishes privacy standards for health information, with ongoing developments in AI-specific guidance from the FDA’s Digital Health Center of Excellence5; internationally, the World Health Organization has published ethics and governance guidance for AI in health6, emphasizing transparency, accountability, and human oversight. Despite these frameworks, deploying AI models on real-world data presents universal challenges regarding privacy and data protection. The General Data Protection Regulation (GDPR)7 in Europe, HIPAA in the U.S., and similar regulations in other jurisdictions impose stringent requirements on data usage, including restrictions on data sharing, cross-border transfer limitations, explicit patient consent mandates, necessitating innovative solutions such as federated learning and synthetic data generation to address these constraints. Adopting these privacy-preserving approaches is essential for achieving the necessary scalability and generalizability of AI systems across diverse clinical institutions. By enabling privacy-preserving, transparent, and reproducible health data analysis across institutional and national boundaries, we provide empirical evidence that synthetic data, when rigorously validated against population-based registries, can serve as a foundational resource for trustworthy AI development while supporting SDG 168 principles of transparent and accountable data governance.” Synthetic data [22, 23] provides a fast and effective solution for researchers who need timely access to data. Synthetic data generation techniques [23] create artificial datasets that mimic the statistical properties of real data, thereby eliminating the risks associated with personal data handling. Beyond improving data accessibility, synthetic health data raise fundamental questions regarding fidelity, validation, and downstream model reliability. While synthetic datasets are increasingly used to accelerate methodological development, their suitability for training and validating complex predictive models—particularly in time-to-event analysis—remains an open research challenge.

In this study, we focus on synthetic data generation as it provides complete data accessibility for comprehensive analysis while greatly reducing privacy concerns, making it particularly suitable for exploring the full potential of Generative AI in survival analysis and promoting transparency and reproducibility of our methods.

For this purpose we used a synthetic breast cancer dataset from the Netherlands Comprehensive Cancer Organisation (IKNL)9, generated to emulate the Netherlands Cancer Registry (NCR) and comprising 60,000 patient records, available for research use10. A critical requirement for the responsible use of synthetic health data is evidence-based validation: models trained on synthetic data must demonstrate consistency with real-world clinical patterns through rigorous empirical testing and robustness when applied to external populations. Without such validation, synthetic data risk becoming methodologically convenient but scientifically unreliable. This study addresses this challenge by systematically evaluating whether survival models trained exclusively on synthetic cancer registry data can generalize to large-scale, real-world registry cohorts, thereby providing evidence for the reliability and applicability of synthetic data in medical informatics research.

In particular, we leverage Generative AI for survival analysis on the IKNL synthetic dataset, providing a rich foundation for exploring the capabilities of LLMs in predicting long-term survival outcomes for real-world clinical applications. By employing feature engineering and data imputation techniques, we aim to refine the accuracy of survival predictions and assess the potential of synthetic datasets in this domain.

In summary, this work makes three primary contributions:

  • Large-scale validation of synthetic-to-real transfer. We provide one of the largest empirical validations of synthetic-to-real transfer for survival analysis, demonstrating that models trained exclusively on synthetic cancer registry data generalize to a real population-based cohort of 183,304 breast cancer patients. We further compare models trained on synthetic data with those trained directly on real registry data, showing that the performance gap remains limited and supporting the use of synthetic data as a reliable and privacy-preserving surrogate for model development.

  • LLM-based framework for survival modeling. We introduce a novel LLM-based pipeline for survival prediction on breast cancer registry data, based on a natural-language representation of structured clinical variables. We benchmark multiple LLM architectures (encoder-only, decoder-only, and encoder–decoder) against established baselines, including Cox regression and Gradient Boosting, demonstrating competitive and, in controlled settings, superior predictive performance.

  • Censoring-aware analysis of LLM behavior. We conduct a controlled ablation study to quantify the role of censoring information, showing that survival status acts as a strong auxiliary signal in LLM-based survival prediction. By evaluating both retrospective (with censoring information) and standard inference settings (without it), we highlight the gap between controlled benchmarking and realistic deployment, and identify limitations of regression-based formulations for censored data.

Additionally, we release open-source code and a fully reproducible pipeline for LLM-based survival analysis, and provide empirical evidence supporting the use of synthetic cancer registry data as a robust resource for training, benchmarking, and methodological development in privacy-constrained healthcare settings. The rest of the paper is structured as follows. Section "Related work" provides an overview of related work in the field, highlighting previous studies and methodologies. Section "Synthetic breast cancer data" describes the synthetic breast cancer data used in this research, including data generation processes and statistical properties. Section "Methods for survival analysis" outlines the methods employed for survival analysis, including the use of LLMs and other machine learning techniques. Section "Results" presents the results of our computational experiments, offering insights into the performance of various models, including their deployment on real-world clinical data. Finally, Section "Conclusions" discusses the implications of our findings and suggests directions for future research.

Related work

Traditional approaches in survival analysis have relied on the Cox (proportional hazards) regression model [24–27]. The Cox model’s popularity is due to its proportional hazards assumption, which allows estimation of model coefficients without specifying the baseline hazard function [25, 28], making it computationally efficient and giving it a “semi-parametric” characterization [24]. However, the proportionality assumption does not always hold, either empirically or theoretically [29]. To address this limitation, more flexible models have been proposed for survival prediction. Regularized versions of the Cox model, which set some coefficients to zero during estimation, have demonstrated improved predictive performance and generalizability [26, 30]. Additionally, machine learning algorithms such as Random Survival Forests [31, 32] and Gradient Boosting Trees [33, 34] can capture non-linear patterns that the Cox model may miss, further enhancing predictive accuracy and generalizability [35–37]. Recent comparative studies have further demonstrated the advantages of data-driven survival modeling approaches, particularly for breast cancer prognostics, where machine learning methods have shown improved discrimination and calibration compared to traditional Cox regression in various clinical settings [38]. In addition to classical and tree-based methods, deep learning approaches for survival analysis have been proposed, such as DeepSurv [39] and DeepHit [40], which extend neural networks to model hazard functions and time-to-event distributions. These approaches demonstrate the potential of representation learning for survival prediction, although their adoption in large-scale registry-based studies remains limited.

The broader potential of LLMs in oncology has been increasingly recognized, with recent reviews highlighting their applications across cancer screening, diagnosis, treatment recommendations, and clinical decision support [41, 42]. Specifically for breast cancer, comprehensive surveys have documented the transformative impact of LLMs in improving diagnostic accuracy and treatment planning [43].

Beyond clinical applications, recent research has shown that LLMs are capable of solving (non-linear) mathematical problems [44, 45], including performing regression analysis [46] with the use of in-context learning [47, 48]. These models achieved comparable or even better performance, in some cases, compared to popular ML algorithms [46], and were able to generalize to new regression problems that could not be part of their training set. Given the deep learning structure of LLMs, learning very complex functions even without weight updating (fine-tuning) comes naturally from their architecture.

However, while LLMs show promise in processing clinical narratives and providing cancer information [49, 50], their application to quantitative survival prediction tasks remains relatively underexplored. Recent work in ovarian cancer research has advocated for developing task-specific LLM models tailored to particular cancer types and analytical objectives [51], an approach that aligns with our LLMs methodology specifically for breast cancer survival analysis. The present work aims to test whether such LLM regression analysis capabilities can also be generalized in the domain of survival analysis.

Importantly, along with novel methods for prediction modeling, a parallel line of work has focused on data anonymization technologies [52, 53], such as synthetic data [22, 23, 54]. Synthetic data play a vital role in cases where the interested parties cannot have access to the real data but can rely on a (synthetic) data set that retains the general statistical properties of the real data. Special care is given so that privacy is not breached (e.g., in cases where patients have a rare feature that makes them easily identifiable in the synthetic data) [55, 56]. Various methods for generating and evaluating synthetic patient data have been developed and validated, demonstrating their utility in advancing machine learning applications while preserving patient privacy [57].

Previous studies [58] have shown that synthetic data can approximate real data effectively for survival prediction modeling. However, the fidelity of synthetic data can vary depending on the generation method, and predictions for specific subgroups may not match the fidelity observed for entire datasets [59]. Synthetic data generation and fidelity remain active areas of research [60, 61], with ongoing debates about privacy preservation, computational requirements, and training stability [62]. In this work, we perform survival analyses on synthetic data using LLMs and benchmark these results against LLM performance on real data to evaluate the impact of data synthesis on model performance.

The integration of AI and machine learning into cancer clinical trials and research workflows [63] further underscores the importance of developing robust methodologies that can leverage synthetic data effectively while maintaining predictive accuracy and clinical relevance. While synthetic data generation methods have matured and machine learning approaches for survival analysis have advanced significantly, a critical gap remains: establishing the reliability and applicability across medical domains of synthetic data through systematic validation frameworks. Individual studies have demonstrated promise in specific contexts, but rigorous assessment of how models trained on synthetic data generalize to real-world populations is essential for broader adoption in clinical research. This study addresses this gap by providing comprehensive empirical validation across both synthetic benchmarks and large-scale registry data.

Synthetic breast cancer data

A cancer registry serves as a vital, disease-specific resource essential for oncological research. Although it is possible to obtain data from a cancer registry, the process typically demands that the requester possess specific qualifications and a thorough understanding of the available data to prepare a proper request. This requirement greatly restricts the number of individuals who can access cancer registry data for research purposes.

To facilitate research, IKNL has developed a synthetic dataset that closely mirrors the structure and statistical properties of the Netherlands Cancer Registry11. Containing absolutely no data on real patients, the IKNL synthetic dataset enables researchers to use record-level cancer data safely, while knowing that there is no risk of breaching patient confidentiality12. This effort addresses the growing need to make high-quality health data more accessible while preserving patient privacy.

Traditional privacy-preserving technologies have repeatedly been shown to be vulnerable to re-identification, which limits data sharing and scientific progress [64]. In the adopted synthetic data generation process, IKNL has applied advanced techniques including Generative Adversarial Networks (GANs) [65] and Bayesian Networks [66] to address these vulnerabilities and generate synthetic datasets that carefully reflect real-world data distributions without containing any information about actual individuals [67].

The probabilistic process used to generate the IKNL synthetic data is based on the PrivBayes method [68]. It synthesizes data via a Bayesian network with differentially private (DP) conditional distributions [69]. A detailed technical description of the PrivBayes generation process, including node-conditional value generation, variable-type handling, survival time modelling, and censoring generation, is provided in Supplementary Material 1. The same pipeline has been employed in the context of cancer registry data by Elvatun et al. [70], who applied PrivBayes to generate synthetic external control arms from Norwegian health registries, providing additional methodological validation of this approach.

The used combination of de-identification and differential privacy techniques guarantees that the synthetic data do not leak more information than intended [64, 67]. This ensures that we can safely release synthetic data, while protecting the privacy of the individuals in the real data. The IKNL synthetic datasets generated with this process undergo also rigorous privacy tests obtained by comparing the statistical properties and model outputs with the original data, ensuring that they retain utility for research while minimizing privacy risks.

The IKNL synthetic dataset is available for research purposes on request at: https://iknl.nl/en/ncr/synthetic-dataset. This first release of the IKNL synthetic data contains a subset of the items registered only for breast cancer patients, and it is planned to include other types of tumor in the future as well. It is possible to obtain a regular version of the synthetic dataset with the variables in the standard NCR format, or alternatively a dataset version compliant to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) [71]. Although not intended to perfectly replicate the NCR, the synthetic data are designed to produce comparable outcomes for tasks like survival prediction, making them suitable for methodological development and exploratory research without the constraints of accessing sensitive data. The ultimate goal is to enable the safe and widespread use of the NCR for innovation and collaboration in cancer research. However, it is worth noting to the reader that because this data is synthetic, it should not be used for clinical decision-making in real-world situations or for deriving actual cancer statistics.

Data preparation

The IKNL synthetic dataset comprises 60,000 breast cancer patient records with a rich set of clinical and demographic variables. However, like most real-world medical datasets, it contains missing values and a potentially high-dimensional feature space that requires careful preprocessing before model training. To prepare the data for survival analysis, we implemented a systematic pipeline involving data imputation, feature engineering, and cross-validation splitting.

Missing data is a common challenge in healthcare datasets, arising from incomplete medical records, patient non-compliance, or administrative inconsistencies. In our dataset, 10,381 rows (17.3%) contained missing values across multiple features, totaling 21,778 missing entries (see Supplementary Material 1 for more details). To address this issue, we employed the MissEnsemble imputation method13, a generalization of the MissForest algorithm [72] that leverages ensemble learning to predict missing values iteratively.

MissEnsemble operates through an iterative imputation strategy that leverages ensemble learning methods, such as Random Forests, to predict missing values. The algorithm begins by performing an initial imputation (e.g., mean for numerical variables, mode for categorical variables) and then iteratively refines these estimates using trained ensemble models. Original variables were appropriately categorized based on their data types, specifically: numerical features (leeft, incjr, tum_afm, ond_lymf, pos_lymf) were treated as continuous variables; categorical features (gesl, tumsoort, topo_sublok, later, morf, gedrag) were handled with classification-based imputation; and ordinal features (diffgrad, er_stat, pr_stat, her2_stat) were processed respecting their inherent ordering. Afterwards, in each iteration, for every variable with missing values, a predictive model was trained on the complete cases using the other variables as predictors. The trained model was then used to predict and replace the missing values. This process continues until a convergence criterion is met or a maximum number of iterations is reached. For our implementation, we configured MissEnsemble with the following specifications:

  • ens_method=‘forest’: Random Forest ensemble method;

  • n_estimators = 100: Number of trees in each Random Forest;

  • n_iter = 10: Maximum number of imputation iterations;

  • tol = 1e-4: Convergence tolerance threshold;

  • random_state = 42: Fixed random seed for reproducibility.

Given the high dimensionality of the original feature space, we also implemented a feature selection strategy to identify the most predictive variables for our survival analysis task. This step serves multiple purposes: reducing computational complexity, mitigating the risk of overfitting, and improving model interpretability. We also adopted an ensemble-based feature selection approach that leverages the collective knowledge of multiple machine learning algorithms. Specifically, we trained three complementary models on the complete feature set: Random Forest Regressor, Gradient Boosting Regressor, and Decision Tree Regressor. Each model was configured to predict survival time and was evaluated using standard regression metrics.

From each trained model, we extracted feature importance scores. Random Forest and Gradient Boosting provide built-in feature importance measures based on the cumulative decrease in node impurity (Gini importance) across all trees in the ensemble. The Decision Tree model similarly quantifies feature importance based on the total reduction in the splitting criterion brought by each feature. To obtain a robust and consensus-based feature ranking, we employed a majority voting mechanism. For each candidate feature, we counted the number of models (out of three) that ranked it among their top-k most important features, where k was determined through cross-validation experiments. Features that were consistently selected by at least two out of the three models were retained for downstream analysis. This conservative strategy ensures that only features with broad predictive power across different model architectures are included, thereby reducing the risk of model-specific artifacts.

The final selected feature set comprised 14 features, a substantial reduction from the original 46 variables, while preserving the essential predictive information for accurate survival prediction (see Table 1). These features include key demographic characteristics (age, gender), temporal information (incident year), tumor characteristics (type, location, morphology, behavior, differentiation grade), biomarker status (ER, PR, HER2), and survival outcomes (time to event and event status), capturing the most clinically relevant prognostic factors for breast cancer survival analysis.

Table 1.

Selected features for survival analysis and their descriptions

Feature Description
gesl Gender (1: Male, 2: Female)
tumsoort Tumor type (501300: Invasive breast carcinoma, 502,200: Ductal carcinoma in situ, 503,200: Lobular carcinoma in situ)
topo_sublok Topography including sub-localization (anatomical location within the breast, e.g., C500: nipple/areola, C501: central part, C502–C505: quadrants)
later Lateralization (1: Left, 2: Right, X: Unknown)
morf Morphology code indicating the histological type of tumor (e.g., 8500: Ductal carcinoma NOS, 8520: Lobular carcinoma NOS)
gedrag Behavior (2: In situ, 3: Malignant)
diffgrad Differentiation grade (1: Well differentiated/low grade, 2: Moderately differentiated/intermediate, 3: Poorly differentiated/high grade, 4: Undifferentiated/anaplastic, 9: Unknown)
er_stat Estrogen receptor status (0: Negative, 1: Positive, 9: Unknown/not assessable)
pr_stat Progesterone receptor status (0: Negative, 1: Positive, 9: Unknown/not assessable)
her2_stat HER2 status (0: Negative, 1+: Negative, 2+: Equivocal, 3+: Positive, 4/7/9: Not determined/unknown)
leeft Age at incidence date (in years)
incjr Incident year (year of diagnosis)
time_to_os Time to overall survival event (in days) - survival time
os Overall survival status (0: Alive/censored, 1: Deceased/event observed)

To ensure robust model evaluation and prevent overfitting, we also partitioned the dataset using stratified 5-fold cross-validation. The stratification was performed based on the event status (censored vs. uncensored) to maintain consistent proportions of observed events across all folds. This approach is particularly important in survival analysis, where the distribution of censored and uncensored observations can significantly impact model training and evaluation [73].

Each fold consisted of 12,000 test samples and 48,000 training samples. All preprocessing steps, including data imputation and feature selection, were performed independently within each fold using only the training data. This strict separation ensures that no information from the test set influences the training process, thereby providing unbiased estimates of model performance.

The same 5-fold split configuration was consistently applied across all modeling approaches evaluated in this study, including traditional statistical methods (Cox Proportional Hazards), machine learning algorithms (Gradient Boosting), and the various LLM architectures, as will be described in the next section. This unified evaluation framework enables fair and direct comparison of model performance across different methodological paradigms.

Methods for survival analysis

In this section, we illustrate how two machine learning methods were employed for survival analysis and provide details on the LLMs, the data representation, and the prompting techniques used in our approach. The synthetic dataset is used as a privacy-preserving environment for training and evaluating all survival analysis models considered in this study.

Cox regression

We employed the Cox proportional hazards regression model [74] to analyze the relationship between covariates and the time-to-event outcome. The Cox model is a semi-parametric method which estimates the hazard function as:

graphic file with name d33e845.gif 1

where h(t ∣ X) is the hazard at time t given covariates X , Inline graphic is the baseline hazard function, and Inline graphic are the coefficients corresponding to each covariate Inline graphic.

One of the key strengths of the Cox model is its ability to naturally handle right-censored data, which occurs when the event of interest has not been observed for some subjects by the end of the study period. Rather than discarding these incomplete observations, the Cox model incorporates them through the partial likelihood formulation, which relies on the relative ordering of event times and includes censored observations in the risk set up to their censoring time, without requiring assumptions about their outcome afterward. This ensures that the estimates of hazard ratios remain unbiased despite incomplete follow-up for some individuals.

Model fitting was performed using partial likelihood maximization, as implemented in the CoxPHSurvivalAnalysis estimator from scikit-survival version 0.24.114. This implementation is based purely on classical statistical modeling and does not involve any LLM or deep learning component.

We used the default settings of the estimator that fit the Cox model without applying penalization or regularization terms. Covariate significance was assessed using Wald statistics, and the proportional hazards assumption was evaluated using residual-based diagnostics.

Gradient boosting

In addition to the Cox model, we employed a gradient boosting approach for survival analysis [75], implemented via the GradientBoostingSurvivalAnalysis estimator from scikit-survival version 0.24.115. This method builds an ensemble of decision trees sequentially, optimizing a Cox partial likelihood loss function to model the hazard function in a flexible, non-linear manner.

Like the Cox model, the gradient boosting approach naturally handles right-censored data by incorporating censored observations into the partial likelihood calculation. Specifically, censored individuals contribute to the risk set up to their censoring time, ensuring that the model leverages all available survival information without making assumptions about outcomes beyond censoring. This allows the boosting algorithm to estimate hazard functions and risk scores accurately despite incomplete follow-up for some subjects.

The gradient boosting model estimates the hazard function by combining multiple weak learners to improve predictive performance, while handling complex relationships and interactions among covariates. Unlike deep learning or LLMs, this approach relies on classical machine learning techniques without requiring extensive feature engineering.

Hyperparameters such as the number of boosting iterations, learning rate, and tree depth were left at their default values.

Large language models

To explore the potential of pretrained LLMs for survival analysis, we selected representative models from each of the three principal architectural paradigms: encoder-only, decoder-only, and encoder-decoder. Specifically, we utilized BioBert [76], an encoder-only model specialized in biomedical text; LLaMA [77, 78] and Qwen [79], two decoder-only models; and T5 [80], a versatile encoder-decoder architecture known for its strong generalization capabilities across NLP tasks.

Given that LLMs are inherently designed to process natural language, we reformulated structured clinical data into textual prompts that encapsulate patient-specific features, such as age, tumor characteristics, and receptor status, as coherent sentences. This prompt-based representation ensures compatibility with the models’ pretraining objectives and enables the application of generative and discriminative language modeling techniques to the regression task of survival time prediction. By expressing structured inputs in natural language, we bridge the gap between tabular medical data and text-based Generative AI models, thereby facilitating their deployment in clinical prognostic tasks.

Detailed descriptions of the fine-tuning procedures of the considered LLMs, including hyperparameter configurations, training schedules, and implementation details for each model architecture, are provided in Supplementary Material 2 to ensure reproducibility. An overview of the methodological architectures and experimental setup is provided in the following.

Representing censored data in the LLMs

An essential aspect of survival analysis is the correct representation of censored observations, i.e., cases where the event (e.g., death) has not occurred by the end of follow-up. This information is represented in the dataset by the os (overall survival status) field, where a value of 1 denotes that the patient is deceased, and 0 indicates a censored observation.

Classical regression models (such as linear, logistic, polynomial, ridge, or lasso regression) are not inherently designed to handle censored observations, as they assume that the target variable represents the true event time [81]. This is because these models minimize the error between predicted and true labels, whereas censored samples do not have a known event time, only a lower bound on it. Including censored data as if the observed time were the true event time introduces bias and violates the assumptions of standard regression techniques.

In contrast, survival-specific models such as the Cox Proportional Hazards model or Gradient Boosting Survival Analysis incorporate censoring directly into their learning objectives. These models accept both the survival time and the event indicator (censoring status) as input:

  • Cox models optimize a partial likelihood that accounts for censored observations by including them in the risk set without treating them as events.

  • Survival gradient boosting methods use extensions of the Cox partial likelihood or other specialized loss functions (e.g., log-rank loss) that natively handle censored data.

Contrary to classical supervised regression, we do not discard censored samples; instead, we incorporate survival status as an explicit input feature to encode censoring information. During training, the Patient status variable (event indicator) is included in the prompt, allowing the model to distinguish between observed events and right-censored cases. We consider two evaluation settings. In a controlled retrospective benchmarking setting, survival status is also provided at inference time. This setting allows us to study how the model conditions on censoring information and to assess its performance when such information is available, effectively acting as a form of privileged input [82]. In addition, to reflect realistic deployment conditions, we evaluate model performance without providing survival status at inference time. In this standard setting, predictions are generated using only covariates available prior to the event, ensuring comparability with classical survival models.

Survival status is included in the prompt using a natural language statement:

  • Patient status: Deceased (for os = 1);

  • Patient status: Alive (for os = 0);

  • Patient status: Unknown (when the information is missing).

This design choice enables the model to distinguish between censored and uncensored examples when predicting survival durations [83]. This is especially important in our setting, where:

  1. The dataset is synthetic and explicitly encodes both event times and censoring information. In this context, censoring information is operationalized through the survival status variable included in the prompt.

  2. Including survival status in the input does not correspond to standard inference conditions; rather, it is used only in controlled retrospective benchmarking analyses to study how models condition on censoring information.

  3. It improves interpretability and performance by allowing the model to condition its prediction based on the known event status.

Example Prompt (Censored Case):

  • Predict the survival time in days based on the following patient information:

  • Patient age: 60 years old.

  • Gender: Female.

  • Incident year: 2010.

  • Tumor type: Lobularcarcinoma in situ,

  • morphology: Lobular carcinoma, NOS,

  • behavior: In situ.

  • Location: Left breast, sublocation: Lower-outer quadrant.

  • Differentiation grade: Moderately differentiated (intermediate).

  • ER status: Positive, PR status: Positive,

  • HER2 status: 0 (negative).

  • Patient status: Alive.

  • Survival time in days: 3592

During training, the prompt includes the patient covariates, the censoring indicator encoded as Patient status, and the target survival time in days. At prediction time, the target survival time is always omitted.

Qwen

Qwen 2.5–7B Instruct is a decoder-only language model comprising approximately 7 billion parameters, designed for instruction-following and open-ended text generation. In our experiments, Qwen was fine-tuned on the survival prediction task using prompt-based input representations derived from structured clinical records, following the same protocol adopted for other generative models in the study (see Supplementary Material 2 for more details). The model was trained with a standard cross-entropy loss to generate survival durations in natural language, which were subsequently converted into numerical values for evaluation.

Unlike encoder-based architectures, Qwen required no transformation of the target variable, as it learns to predict interpretable time durations directly in free text. Its ability to process the entire input prompt as a coherent textual unit allowed it to model complex interactions among clinical variables without architectural adjustments. The inclusion of survival status as a textual feature further enhanced its capacity to handle censored data, aligning with the supervised setting established for this benchmark.

LLaMA

Our study employed three models from the LLaMA family: LLaMA-3.1–8B, LLaMA-3.1–70B, and MedLLaMA-3–8B. All are decoder-only transformer architectures designed for efficient language generation.

LLaMA-3.1–8B serves as a compact yet capable model for general NLP tasks, while LLaMA-3.1–70B represents a significantly larger variant with increased capacity for modeling complex linguistic dependencies. MedLLaMA-3–8B, in contrast, is a domain-specific model pretrained on biomedical corpora, selected to investigate whether exposure to medical text improves performance on clinical prediction tasks.

All LLaMA-based models were fine-tuned (Supplementary Material 2) using our custom prompt-based formulation of survival analysis, in which structured patient data is represented as natural language input. The models were trained to generate survival times directly as free-text outputs, optimized with the standard cross-entropy loss, and evaluated after converting the generated durations into numerical form. As with other decoder-only models, no transformation such as quantile normalization was applied to the target variable.

The inclusion of MedLLaMA in our experiments was primarily exploratory, aimed at assessing the potential benefit of domain adaptation in a generative survival prediction setting. However, its performance did not surpass that of the general-purpose LLaMA-3.1–8B, suggesting that domain specialization in pretraining does not necessarily translate to improved downstream performance in this task configuration. All models leveraged the full prompt format, including survival status, to appropriately represent censored and uncensored cases during training (and in controlled retrospective analyses).

BioBert

BioBert is an encoder-only transformer model based on the BERT architecture and consists of approximately 110 million parameters. It is pre-trained on large-scale biomedical corpora, including PubMed abstracts and PMC full-text articles, and has demonstrated state-of-the-art performance on a wide range of biomedical natural language processing (NLP) tasks such as named entity recognition, question answering, and relation extraction. Its domain-specific pretraining makes it particularly well-suited for processing clinical and biomedical texts, where specialized terminology and context are crucial for understanding.

In our work, we adapted BioBert for survival analysis by replacing its original classification head with a regression layer designed to output a continuous estimate of survival time, expressed in days. The model was fine-tuned end-to-end (see Supplementary Material 2) using the mean absolute error (MAE) as the loss function, which is particularly appropriate for regression tasks involving skewed or noisy targets.

To further improve the model’s robustness and convergence during training, we applied a Quantile Transformer to the survival times. This preprocessing step transformed the raw durations into a standard normal distribution, mitigating issues related to skewness and heteroscedasticity often encountered in survival data.

Structured patient data were encoded into natural language prompts to create inputs compatible with the model’s text-based architecture. This prompt-based approach allowed us to leverage BioBert’s pretraining on biomedical text while preserving semantic richness in the representation of clinical features. The minimal architectural modification, limited to the output layer, enabled an effective adaptation of BioBert to the regression setting, demonstrating the feasibility of repurposing pretrained language models for survival prediction in the biomedical domain.

T5

Our study explored several variants of the T5 (Text-to-Text Transfer Transformer) model: t5-small (60 M parameters), t5-base (220 M), t5large (770 M), and flan-t5-large (770 M). T5 is an encoder-decoder architecture designed by Google to unify all NLP tasks under a text-to-text framework, where both input and output are textual. This formulation makes T5 particularly versatile, enabling it to handle tasks such as translation, summarization, question answering, and classification with minimal task-specific modifications.

T5 is pre-trained on a large corpus using a masked span corruption objective, which helps it learn rich language representations suitable for downstream fine-tuning. The flan-t5-large model further benefits from instruction tuning, where it is exposed to a diverse set of NLP tasks framed as instructions, enhancing its few-shot and generalization capabilities.

In our survival prediction setting, we first fine-tuned the T5 models in their original sequence-to-sequence configuration (see Supplementary Material 2 for more details). Structured patient information was formatted as a natural language prompt, and the model was trained to generate the survival duration in free-text form using standard cross-entropy loss. Predictions were then post-processed to obtain numerical survival times for evaluation.

To bridge the gap between discrete language generation and continuous outcome prediction, we also adapted T5 into an encoder-only regression model. In this configuration, only the encoder was used to process the natural language input. The resulting contextual embeddings were pooled, via mean or attention-based pooling, and passed to a regression head trained with mean absolute error loss. This adaptation preserved the benefits of T5’s rich language modeling while enabling direct optimization on continuous survival times.

This hybrid use of T5 demonstrates its flexibility: while originally conceived for generative tasks, its encoder can also serve as a powerful feature extractor for regression tasks, especially when combined with natural language prompts that contextualize structured input data.

Results

This section presents the outcomes of our computational experiments, focusing on the performance of various models in survival analysis using synthetic breast cancer data. Through evidence-based validation against both synthetic benchmarks and real-world registries, we assess whether LLM-based approaches can reliably capture survival patterns and maintain performance when deployed in clinical settings. We first describe the main evaluation metrics for survival analysis, then present our experiments on the IKNL synthetic breast cancer data, and finally discuss a practical deployment on an existing population-based cancer registry of 183,304 patients to demonstrate the utility and challenges of our approach in real-world practice.

Evaluation metrics for survival analysis

We outline here the evaluation metrics employed to assess the performance of survival analysis models. These metrics are crucial for determining the accuracy, discriminative ability, and robustness of the models when predicting survival outcomes.

The time-dependent Brier Score [84] evaluates the precision of survival probability estimates at specific time points. It is calculated as the weighted mean squared difference between predicted survival probabilities and actual survival status at that time, that is:

graphic file with name d33e1218.gif

where n is the number of subjects, Inline graphic is the indicator of whether the event has occurred by time t for subject i, Inline graphic is the predicted survival probability for subject i at time t, and Inline graphic is the inverse probability weight to account for censoring [84]. The Brier score was evaluated over a predefined grid of time points between 1 and 10 years of follow-up, restricted to the maximum observed follow-up time in the training data. It assesses predictive accuracy throughout the follow-up period, providing comprehensive temporal coverage and aligning with standard clinical practice of annual survival evaluation. Lower Brier Scores indicate better predictive performance.

The Concordance Index (C-Index) [85] measures the model’s ability to correctly predict the order of events. It is calculated as:

graphic file with name d33e1267.gif

where Inline graphic and Inline graphic are the observed survival times, Inline graphic and Inline graphic are the predicted survival times, and Inline graphic is the indicator function [85]. A C-Index of 0.5 indicates random prediction, while a value of 1.0 denotes perfect prediction.

To handle censored data, the C-Index with Inverse Probability of Censoring Weights (IPCW) [85] adapts the traditional C-Index by incorporating inverse probability weighting:

graphic file with name d33e1303.gif

where Inline graphic is the Kaplan-Meier estimate of the censoring distribution, which provides a non-parametric estimate of the survival function that helps determine the probability of a subject being uncensored up to time t [73]. This adjustment improves accuracy in survival analysis by addressing biases from censored observations.

The C-Index Not Censored calculates the C-Index without considering censored data, assuming all events are observed. This metric provides insights into the model’s performance in an ideal scenario where censoring does not occur, though it may not reflect real-world conditions.

Finally, the integrated time-dependent Area Under the Receiver Operating Characteristic Curve (AUC) [85, 86] quantifies a model’s ability to discriminate between classes. In particular, in survival analysis, it evaluates the model’s ability to distinguish between subjects with different event times:

graphic file with name d33e1335.gif

where TPR is the true positive rate, representing the proportion of actual positive events (e.g., observed survival beyond a threshold) correctly identified by the model, and FPR is the false positive rate, indicating the proportion of negative events (e.g., non-survival within the threshold) incorrectly identified as positive by the model [85]. In the context of survival analysis, these rates help measure how well the model can predict and discriminate between individuals who will experience the event and those who will not.

The time-dependent AUC was computed as a single time-dependent measure, using the cumulative/dynamic definition over a sequence of evaluation time points defined at yearly intervals across the observed follow-up period. We report the integrated (mean) AUC across all time points, as implemented in the adopted cumulative_dynamic_auc function from the scikit-survival (sksurv) python library16, integrating patient-specific risk scores over annual intervals between the earliest and latest observed survival times. An AUC of 0.5 suggests no discrimination, while a value of 1.0 indicates perfect discrimination, highlighting the model’s ability to identify at-risk individuals.

The time-dependent AUC and Brier score were evaluated over comparable follow-up horizons within the observed data, ensuring consistency in time-dependent performance assessment. By contrast, the C-index–based metrics (standard, IPCW, and not-censored) are global concordance measures and are not defined over explicit time grids.

Computational experiments on synthetic data

To assess the predictive capabilities of LLMs on survival analysis tasks, we designed a rigorous experimental protocol using synthetic breast cancer data derived from a European cancer registry. As described in Section "Data preparation", each model was evaluated through 5-fold cross-validation to ensure robustness and mitigate overfitting. Specifically, in each iteration, we trained on four folds (48,000 samples) and tested on the remaining one (12,000 samples), rotating the test fold across five rounds.

Fine-tuning the generative models required 1 to 2 epochs, depending on model size and convergence behavior. For the LLaMA-3.1-8B, MedLLaMA-8B, and Qwen-7B models, training and inference could be effectively carried out on a virtualized environment equipped with a single NVIDIA Quadro RTX 8000 GPU (48 GB VRAM) and 128 GB of RAM. In contrast, the LLaMA-3.1-70B model required access to a more powerful cluster infrastructure with NVIDIA H100 GPUs due to its significantly larger memory footprint and compute requirements. Each training session for the 70B model exceeded 24 hours per fold on average. Inference was also time-consuming, though substantially faster than training, with durations ranging from several minutes to a few hours, particularly for the largest model.

In addition to LLMs, we also experimented with classical machine learning approaches, including Decision Trees and Random Forests. However, in this study, we reported only the results of the Cox Proportional Hazards model and Gradient Boosting, as they consistently outperformed the other traditional models.

Tables 2, 3, 4 and 5 summarize the predictive performance of the different models for survival prediction across all folds evaluated using the metrics mentioned in Section "Evaluation metrics for survival analysis".

Table 2.

Performance comparison of two models for survival prediction: cox (statistical) and Gradient Boosting (ML). Gradient Boosting achieves better results across all metrics

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
COX 0.1322 0.7538 0.7435 0.5915 0.7881
Gradient Boosting 0.1275 0.7817 0.7642 0.6706 0.8104

Table 3.

Performance metrics of decoder-only generative models fine-tuned for survival prediction using a standard cross-entropy loss. Both LLaMA3.1-8B and LLaMA3.1-70B exhibit strong overall performance, with LLaMA3.1-8B achieving the highest C-Index and AUC scores

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
Qwen 2.5 0.1184 0.7520 0.7278 0.8374 0.8389
LLaMA3.1-8B 0.1200 0.7934 0.8272 0.8560 0.8858
LLaMA-3.1-70B 0.1347 0.7860 0.8336 0.7514 0.8751
MedLLaMA3-8B 0.1386 0.7714 0.8492 0.6901 0.8288

Table 4.

Performance of encoder–decoder models (T5 variants) fine-tuned for survival prediction. T5-base achieves the best overall results across all metrics

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
T5-small 0.1223 0.7218 0.8325 0.8464 0.8391
T5-base 0.1131 0.8440 0.9139 0.8626 0.9295
T5-large 0.1180 0.8354 0.9064 0.8593 0.9163
flan-T5-large 0.1184 0.8026 0.8760 0.8533 0.8799

Table 5.

Performance of encoder-only models (BioBert and T5 variants) fine-tuned and adapted for survival prediction. All models were trained using the mean absolute error (MAE) loss function. For T5 models, only the encoder part was used. BioBert, originally designed for classification tasks, was modified with a regression head for survival analysis. Architectural variants include: MP, which uses mean pooling followed by a linear output layer; and AP, which employs an attention-based pooling mechanism for prediction

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
BioBert 0.1081 0.8556 0.9155 0.8689 0.9359
T5-small MP 0.1117 0.8256 0.9080 0.8676 0.9164
T5-small AP 0.1154 0.8449 0.9188 0.8678 0.9323
T5-base MP 0.1140 0.8463 0.9203 0.8678 0.9340
T5-base AP 0.1152 0.8530 0.9181 0.8670 0.9360
T5-large MP 0.1152 0.8389 0.9109 0.8686 0.9301

In particular, in Table 2 we can see that the Gradient Boosting model consistently outperforms the classical Cox proportional hazards model across all evaluated metrics. Notably, the Gradient Boosting approach achieves a lower Brier Score, indicating better calibration and prediction accuracy, as well as higher C-Index and AUC values, reflecting improved discriminative ability.

Table 4 reports the performance of the decoder-only generative models fine-tuned with cross-entropy loss under the controlled retrospective setting. All the models demonstrate competitive performance, with the LLaMA3.1-8B model achieving the highest C-Index and AUC, which suggests superior risk ranking and classification at various time points. Interestingly, while Qwen 2.5 shows the lowest Brier Score, its C-Index IPCW is lower than that of the LLaMA models, which might reflect differences in how censoring is handled or prediction calibration.

Encoder–decoder models based on the T5 architecture (Table 4) further improve prediction quality under the controlled retrospective setting. In this configuration, T5-base achieves the best overall performance, including the highest C-Index IPCW and AUC scores, indicating strong discriminative ability when conditioning on censoring information. This suggests that leveraging both encoder and decoder representations enhances the model’s capacity to capture complex temporal dependencies in survival data.

Finally, Table 5 reports the results obtained by the encoder-only models, including BioBert and T5 variants, fine-tuned using a Mean Absolute Error loss under the controlled retrospective setting. These models achieve very strong results. In particular, BioBert attains the best performance overall, with the lowest Brier Score and highest C-Index and AUC values. This indicates that specialized pretraining in biomedical contexts combined with architectural adaptations for survival regression can yield highly effective models. Among T5 encoder variants, the differences between mean pooling (MP) and attention pooling (AP) are minor; though attention pooling slightly improves AUC. The results obtained by BioBert significantly outperform those obtained by the COX model, considered a classical benchmark in survival prediction. As shown in Table 6, the paired t-tests using scipy.stats.ttest_rel on these results confirm the statistical significance of the improvements for both Brier Score (t = -7.4209, p = 0.0018) and C-Index (t = 20.3405, p < 0.0001). Similarly, when compared to the Gradient Boosting model, BioBert achieves statistically significant gains, with t = -6.4152 (p = 0.0030) for Brier Score and t = 14.6069 (p = 0.0001) for C-Index.

Table 6.

Comparison of survival prediction performance between the COX model and BioBert across five folds. Metrics include Brier Score (lower is better) and concordance Index (C-Index, higher is better). Paired t-tests assess statistical significance of performance differences

Metric Model Fold 1 Fold 2 Fold 3 Fold 4 Fold 5
Brier Score COX 0.1363 0.1378 0.1285 0.1297 0.1290
BioBert 0.1081 0.1081 0.1105 0.0995 0.1143
Paired t-test: t = -7.4209 p = 0.0018
C-Index COX 0.7535 0.7479 0.7554 0.7590 0.7531
BioBert 0.8410 0.8404 0.8610 0.8695 0.8659
Paired t-test: t = 20.3405 p = 0.0001

To disentangle modeling capacity from censoring signal exploitation, we conducted a controlled ablation study in which the survival status field (os) was excluded from training. The best-performing representative model from each primary architecture—BioBERT (encoder-only), T5-Base (encoder–decoder), and LLaMA-3.1-8B (decoder-only)—was re-trained under otherwise identical conditions. The results, reported in Table 7, show a performance degradation across all three architectures when survival status is removed, confirming that censoring information acts as a meaningful conditioning signal. The effect is most pronounced for LLaMA-3.1-8B, which exhibits a substantial drop across all metrics, while BioBERT and T5-Base show more moderate but consistent reductions in predictive performance. Despite this degradation, BioBERT and T5-Base retain reasonable predictive ability and remain competitive with respect to baseline models on selected metrics, particularly the C-Index (Not Censored), where they outperform Cox regression. However, Gradient Boosting continues to achieve stronger overall performance across most standard survival metrics, including C-Index and AUC. These findings are consistent with the design rationale described in Section "Representing censored data in the LLMs": without explicit censoring information, the model cannot distinguish patients known to have experienced the event from those with incomplete follow-up. The observed performance degradation can be explained by the formulation of the task as a standard regression problem. In the absence of explicit censoring information, the models are trained to predict survival time directly, implicitly assuming that the observed time corresponds to the true event time. However, for censored observations, the recorded survival time represents only a lower bound, as the event has not yet occurred. As a result, the predicted values may deviate substantially from the observed times, particularly for patients who are still alive, reducing predictive accuracy.

Table 7.

Full ablation: mean 5-fold performance metrics when censoring information is excluded from both training and evaluation prompts. Results should be compared with those in Tables 3, 4, 5

Metric BioBERT T5-Base LLaMA-3.1-8B
Brier Score (Average) 0.1310 0.1320 0.3166
C-Index 0.6649 0.5920 0.5533
C-Index (IPCW) 0.6923 0.6032 0.5601
C-Index (Not Censored) 0.8412 0.8301 0.6467
AUC 0.7956 0.7385 0.6158

Overall, our findings suggest that LLM-based approaches possess interesting potential for addressing survival analysis tasks, particularly in modeling complex relationships inherent in time-to-event data. Unlike traditional statistical methods that rely on predefined hazard functions or parametric assumptions, LLMs leverage their powerful representation learning capabilities to automatically extract meaningful patterns from high-dimensional and potentially unstructured data. This ability enables them to capture nonlinear interactions and temporal dependencies that may be difficult to specify explicitly.

Moreover, the generative and contextual understanding nature of LLMs allows these models to integrate heterogeneous information sources, such as clinical notes, laboratory results, and patient histories, which are often present in survival datasets but underutilized by conventional methods. By fine-tuning LLMs for survival prediction, our results indicate competitive discriminative performance compared to classical and machine learning baselines. However, while promising, it is worth noting that LLM-based survival models require careful consideration of challenges such as censoring, computational complexity, and interpretability. The training process must incorporate appropriate loss functions and evaluation metrics that account for censored observations to ensure valid and clinically meaningful predictions.

To further investigate whether ensemble methods could improve predictive performance, we conducted additional experiments combining the predictions of our top-performing models (BioBert, T5-base with cross-entropy loss, and T5-base with attention pooling). Details of these ensemble strategies and their comparative results are presented in Supplementary Material 3. While ensemble methods are often employed to reduce variance and improve robustness, our results (see Table S1 in Supplementary Material 3) indicate that the individual fine-tuned models already achieve such strong performance that ensembling provides minimal additional benefit in this setting. More information on this experiment is provided in Supplementary Material 3.

Computational experiments on real-world data

To assess the real-world applicability of the models fine-tuned on synthetic data, we deployed two of the best-performing LLMs—specifically BioBert and T5-base (attention pooling variant, namely T5-base AP)—on a large de-identified clinical dataset provided by the Netherlands Comprehensive Cancer Organisation who maintains the Netherlands Cancer Registry. The real-world registry data were provided in the same format as the synthetic training data, ensuring full compatibility with the prompt-based input structure used during fine-tuning. For the inference phase on real data, only the input features were supplied to the models; the Survival time in days variable was omitted as input and used as the prediction target for evaluation.

We focused our deployment on these two top-performing models from the synthetic data experiments to maintain computational feasibility and optimize resource utilization in the restricted on-site environment. Additionally, to provide a comprehensive comparative analysis, we also evaluated the classical Cox Proportional Hazards model and the Gradient Boosting approach, which served as established baseline methods throughout our study. Due to strict privacy and data-governance regulations, access to this dataset required performing all computations on-site at the facilities of the IKNL. For this experiment, the models were fine-tuned on the entire synthetic dataset, comprising all five folds and accounting to the overall 60,000 breast cancer patient records, to leverage the maximum amount of training data available and enhance their generalization capabilities, before being securely transferred to the controlled environment for local execution.

The use of anonymized real-world cancer registry data for this deployment was conducted in accordance with IKNL’s statutory mandate and privacy framework17. As a population-based cancer registry, IKNL is legally authorized to collect and process cancer patient data for scientific research and statistics in the interest of public health. This authorization is based on the legal basis of public interest research under GDPR, which does not require individual informed consent for population-based cancer surveillance activities that serve public health objectives, as outlined in IKNL’s privacy statement. Patients retain the right to opt out of data inclusion in the registry. All data requests from external researchers, including this study, are reviewed by the NCR’s independent Supervisory Committee, which includes patient representatives, to ensure appropriate use and protection of patient privacy. The deployment of our models on these data was approved through this oversight mechanism and conducted entirely within IKNL’s secure on-site facilities, with data never leaving those premises.

The real-world cohort consisted of 183,304 patients extracted from regional cancer registry records in the Netherlands. Our goal was to evaluate whether models trained exclusively on synthetic data could generalize to real clinical survival data and produce reliable risk predictions when applied unchanged to a real registry.

When deployed in this environment, all four models—the Cox regression, Gradient Boosting, BioBert, and T5-base AP—achieved the performance metrics shown in Table 8.

Table 8.

Performance of LLMs and classical machine-learning models (Cox and Gradient Boosting) on the real-world cancer registry dataset (183,304 patients). LLMs were fine-tuned, and machine-learning models were trained, on synthetic data. All models were executed on-site at IKNL on the real-world data without additional retraining

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
COX 0.1686 0.7304 0.7249 0.6203 0.7749
Gradient Boosting 0.1799 0.7130 0.7096 0.6707 0.7589
BioBert (retrospective) 0.1675 0.8268 0.9276 0.8936 0.9279
T5-base AP (retrospective) 0.1679 0.8148 0.9240 0.8938 0.9222
BioBert (standard) 0.1543 0.5206 0.5446 0.8411 0.6979
T5-base AP (standard) 0.1277 0.5035 0.5290 0.8371 0.6846

The obtained results reveal several important insights about model performance on real-world clinical data. First, the classical Cox model maintains reasonable predictive performance with a C-Index of 0.7304 and AUC of 0.7749, confirming its role as a robust baseline in survival analysis. The Gradient Boosting model achieves a slightly lower C-Index (0.7130) compared to Cox, although it shows better discrimination for uncensored cases (C-Index Not Censored = 0.6707). While both traditional methods provide acceptable performance on real data, they are substantially outperformed by the LLM-based approaches when survival status is included at inference time in a controlled retrospective setting, consistently with the trends observed in the synthetic data experiments. In this setting, both BioBert and T5-base AP exhibit markedly superior performance across nearly all metrics. BioBert achieves a C-Index of 0.8268 and an AUC of 0.9279, while T5-base AP attains comparable results with a C-Index of 0.8148 and AUC of 0.9222. These results demonstrate that LLMs fine-tuned exclusively on synthetic data can effectively transfer to real clinical populations and outperform traditional statistical and machine learning approaches under controlled evaluation conditions. Additionally, we evaluated the distributional similarity between the synthetic and real datasets using Jensen–Shannon divergence [87], observing consistently low divergence across all variables. This supports the validity of the synthetic data as a proxy for the original distribution, although minor discrepancies in higher-order dependencies may still influence model behavior. Further details are provided in Supplementary Material 1. Furthermore, the performance achieved by BioBert and T5-base AP in this setting shows only a moderate decrease compared to the synthetic benchmarks and remains within a high-performance range. In particular, the strong C-Index and AUC values indicate that the models retain robust discriminative ability on real data. Interestingly, the C-Index IPCW on real patient data slightly exceeds the value obtained in the synthetic data deployment. This may be attributed to the different training context, as the real-world experiment utilizes models fine-tuned on the entire synthetic dataset rather than on a subset of folds, potentially providing a broader basis for handling censored observations. However, this observation warrants further investigation, as differences in censoring patterns between synthetic and real data may also play a role. When survival status is excluded at inference time to reflect standard deployment conditions, a substantial degradation in ranking-based metrics is observed for both BioBert and T5-base AP.

This behavior is expected given the formulation of the task as a regression problem. This limitation is not specific to LLM-based approaches, but reflects a broader challenge in applying standard regression formulations to censored survival data. In the absence of explicit censoring information, the models are trained to predict observed survival times as exact targets, whereas for censored patients, these values represent only lower bounds of the true event time. This introduces systematic noise that disrupts pairwise ordering across individuals, leading to near-random performance in ranking-based metrics. Despite this degradation, the models retain strong performance on metrics that are less sensitive to censoring ambiguity. In particular, the C-Index computed on non-censored cases remains high (above 0.83 for both models), indicating that the models successfully capture meaningful survival patterns when true event times are available. Similarly, the Brier Score remains competitive, and in the case of T5-base AP even improves (0.1277), suggesting that predicted survival times remain reasonably calibrated on average. These results highlight that the observed performance drop is primarily driven by the mismatch between regression-based learning and censored data, rather than a failure to learn clinically relevant patterns. Overall, this analysis confirms that while survival status provides a strong conditioning signal in retrospective settings, the models maintain partial predictive capability under realistic inference conditions, particularly when evaluated on uncensored outcomes.

To further validate our results, we also trained the same models directly on the real cancer registry dataset, so as to compare the obtained results against those obtained by models trained on synthetic data. Training on real data followed exactly the same procedure and settings as in the synthetic data experiments (Section "Computational experiments on synthetic data"). The real cancer registry dataset of 183,304 patients was divided into five folds, each containing approximately 36,660 patients, and each model was evaluated using 5-fold cross-validation (training on four folds and testing on the remaining one, with the test fold rotated across five rounds). The obtained results are shown in Table 9. Overall, models trained and evaluated directly on real data achieve strong performance, with results broadly consistent with those obtained in the synthetic setting. As expected, performance is slightly higher when both training and testing are conducted on real data, due to the absence of distributional shift.

Table 9.

Performance of LLMs and classical machine-learning models (Cox and Gradient Boosting) on the real-world cancer registry dataset (183,304 patients). LLMs were fine-tuned, and machine-learning models were trained, on the real-world cancer registry data. All models were trained using the mean absolute error (MAE) loss function. For T5-base AP, which employs an attention-based pooling mechanism for prediction, only the encoder part was used. BioBert was modified with a regression head for survival analysis

Brier Score C-Index C-Index IPCW C-Index Not Censored AUC
COX 0.1341 0.7782 0.7500 0.5677 0.8025
Gradient Boosting 0.1219 0.7965 0.7612 0.6155 0.8238
BioBert (Retrospective) 0.0868 0.8720 0.9169 0.9009 0.9413
T5-base AP (Retrospective) 0.0919 0.8590 0.9044 0.9000 0.9330
BioBert (standard) 0.1345 0.5216 0.5309 0.8416 0.6598
T5-base AP (standard) 0.1660 0.5152 0.5159 0.8313 0.6478

In the retrospective setting, where survival status is included at inference time, the LLM-based models (BioBert and T5-base AP) substantially outperform the classical Cox and Gradient Boosting baselines across all metrics. In particular, BioBert achieves a C-Index of 0.8720 and an AUC of 0.9413, while T5-base AP attains similarly strong performance (C-Index of 0.8590 and AUC of 0.9330). These results confirm that, when provided with censoring information, LLMs are highly effective at capturing survival patterns and risk stratification signals in real clinical data.

When survival status is excluded at inference time to reflect standard deployment conditions, a marked degradation is observed in ranking-based metrics for both models. This behavior is consistent with the results observed in the synthetic-to-real setting and can be explained by the formulation of the task as a regression problem. Despite this degradation, the models retain strong discriminative ability when evaluation is restricted to non-censored cases, with C-Index (Not Censored) remaining high (above 0.83 for both models). This indicates that the models successfully learn meaningful survival patterns when true event times are observed. Additionally, the Brier Score remains competitive, particularly for BioBert, suggesting that predictions remain reasonably calibrated on average even in the absence of censoring information. These findings confirm that the observed performance drop is primarily driven by the mismatch between regression-based learning and censored data, rather than a failure to capture clinically relevant relationships.

As anticipated, the results obtained by models trained and evaluated on real data also outperform the setting in which models are trained on synthetic data and evaluated on real-world data, where a modest performance drop is observed due to residual discrepancies between the synthetic and real data distributions. Nevertheless, the gap remains limited, confirming that models trained on synthetic data retain a substantial degree of generalization capability when applied to real patient data.

Taken together, these findings provide additional evidence that the synthetic dataset constitutes a reliable proxy for model development. While direct training on real data yields the best performance, the comparable results obtained in the synthetic-only setting support the use of synthetic data as a practical and effective alternative in scenarios where access to real data is restricted.

Overall, these experiments demonstrate that LLMs fine-tuned on both synthetic and real datasets can generalize effectively to real patient populations and outperform traditional Cox and Gradient Boosting survival models in controlled settings. While direct training and evaluation on real data yield the strongest performance, the results obtained under standard inference conditions highlight the importance of explicitly accounting for censoring in regression-based formulations. The performance gap observed when removing survival status is consistent and theoretically grounded, and does not undermine the validity of the learned survival representations.

Our three-tier validation approach—comprising (i) statistical fidelity of the synthetic data during generation, (ii) strong performance in synthetic-to-synthetic evaluation, and (iii) consistent generalization to real-world data, complemented by direct real-to-real benchmarking—provides a comprehensive assessment of the validity of synthetic data as a practical and privacy-preserving surrogate for training survival models in healthcare settings, where access to real data is restricted.

This approach is particularly pertinent to emerging European health data initiatives such as the EHDS. While EHDS is designed to facilitate access to real-world data, it operates under data minimization principles. Our methodology demonstrates that models can be initially trained on synthetic datasets and subsequently achieve strong performance when deployed on real data, offering a complementary strategy that could reduce the volume of real patient data required during model development phases while maintaining predictive utility in cross-border health data utilization.

Conclusions

This study demonstrates the potential of Generative AI as a flexible and effective approach for survival analysis in oncology. Using a large synthetic breast cancer dataset designed to emulate a national population-based cancer registry, we evaluated the performance of advanced survival prediction methods leveraging Large Language Models (LLMs), and compared them against established statistical and machine-learning baselines.

Our results show that LLM-based models can achieve strong performance in retrospective benchmarking settings when censoring is provided as privileged information. LLM-based survival models trained exclusively on synthetic data are able to capture clinically meaningful time-to-event patterns and generalize effectively to real-world cancer registry data, achieving strong discriminative performance when deployed unchanged on a large population-based cohort of 183,304 patients from Dutch cancer registries. This empirical validation addresses a central challenge in the use of synthetic health data, namely whether models developed on synthetic datasets retain reliability and applicability when applied to real clinical populations.

However, these results should not be interpreted as directly representative of deployable performance, as survival status is not available at prediction time in real-world clinical settings. These findings suggest that current prompt-based LLM formulations can exploit censoring information effectively, but also highlight the need for censoring-aware training objectives that do not require survival status at prediction time.

Beyond predictive performance, the use of high-fidelity synthetic data in this context, enables privacy-preserving model development and evaluation, facilitating experimentation in settings where access to real-world clinical data is restricted. Our comprehensive validation framework demonstrates that synthetic data can support reliable innovation when subjected to rigorous empirical testing. The observed consistency between synthetic-trained and real-world model performance supports the use of synthetic data for early-stage model development, methodological comparison, and reproducible research, while reserving access to real data for final validation and deployment.

Some limitations should be acknowledged. Although the synthetic dataset closely mirrors the statistical properties of the underlying cancer registry, it cannot capture all sources of bias, heterogeneity, or rare clinical trajectories present in real populations. Residual distributional differences between synthetic and real data may still influence model behavior, particularly for underrepresented subgroups, even when overall fidelity is high. Furthermore, the retrospective evaluation setting used in this study does not fully reflect prospective clinical deployment, where additional constraints related to interpretability, fairness, and clinical integration must be considered.

Further future work should focus on three priorities. First, establishing community-consensus validation protocols including standardized metrics for model performance, subgroup-level evaluation, and longitudinal consistency checks. Second, extending this evaluation framework to additional cancer types and chronic diseases to assess the generalizability of LLM-based survival models across medical domains. Third, developing best-practice guidelines for post-deployment monitoring that ensure synthetic-trained models maintain performance as real-world clinical practices evolve. These efforts will strengthen the evidence base for using synthetic data as a reliable alternative to restricted datasets in healthcare innovation.

Overall, our findings support the integration of LLM-based approaches into survival analysis pipelines and highlight their potential in data-constrained clinical environments. By demonstrating that rigorous validation enables privacy-preserving innovation without compromising predictive utility, this work provides a pathway for responsible AI development in healthcare that balances research advancement with patient privacy protection.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 2 (151.7KB, pdf)
Supplementary Material 3 (149.4KB, pdf)

Acknowledgements

We would like to thank the colleagues of the Digital Health Unit (JRC.F7) and the Disease Prevention Unit (JRC.F1) at the Joint Research Centre of the European Commission for guidance and support. The views and opinions expressed herein are the authors’ own and do not necessarily state or reflect an official position of the European Commission. We would like to thank also the colleagues from Netherlands Comprehensive Cancer Organisation (IKNL) for the helpful suggestions and support during the development of this work.

Author contributions

S.C. conceptualized the work, helped in the software development, and created the LLMs releases. A.M. developed the software and main pipelines, and analysed the results. D.R.R. conceptualized the experiments, analysed the results, and supervised the work. D.K. and M.S. conceived the data collection and feature engineering, analysed the results, and curated the datasets. All authors have read, edited, and approved the original manuscript.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or nonprofit sectors.

Data availability

The data that support the findings of this study are available from the Netherlands Comprehensive Cancer Organisation (IKNL). These data are used under license for the current study and are therefore not publicly available. In particular, the used synthetic dataset is available for research purposes on request to IKNL at: https://iknl.nl/en/ncr/synthetic-dataset. It is possible to obtain a regular version of the synthetic dataset with the variables in the standard NCR format (https://iknl.nl/en/ncr), or alternatively a dataset version compliant to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) [71]. The de-identified clinical data used in the deployment on real-world cancer registry cases consist of 183,304 patient records from the Netherlands Cancer Registry (NCR), collected and maintained by the Netherlands Comprehensive Cancer Organisation (IKNL). As a population-based cancer registry operating under Dutch and European Union data protection law, IKNL is statutorily mandated to collect cancer patient data from all hospitals in the Netherlands for scientific research and statistics in the interest of public health. The legal basis for this data collection, as detailed in IKNL’s privacy statement (https://iknl.nl/en/privacystatement), is grounded in the principle that requesting individual informed consent is not possible or appropriate, given the burden such requests would place on patients during the difficult period of cancer diagnosis and the critical societal need for comprehensive, unbiased cancer epidemiology data. This exemption operates within the framework of the EU General Data Protection Regulation (GDPR), specifically under the provision for scientific research and statistics in the public health interest. Patients are informed about the registry and retain the right to opt out of data inclusion at any time. Access to NCR data for research purposes is strictly controlled through an independent Supervisory Committee that includes patient representatives, medical professionals, and data protection experts. This Committee evaluates all data requests to ensure scientific merit, appropriate data use, and adequate privacy protection measures. For this study, all models were trained exclusively on synthetic data and subsequently deployed on real patient data only after approval by the Supervisory Committee. Due to the sensitive nature of individual-level patient data and strict privacy and data-governance regulations, direct access to the real-world dataset is not possible. All computations involving real patient data were performed on-site at IKNL’s secure facilities under controlled conditions, with data never leaving those premises. This regulatory approach to cancer registry data access reflects common practices in population-based cancer surveillance worldwide, including frameworks employed by registries such as the U.S. SEER Program, the Canadian Cancer Registry, and the Australian Cancer Database, which similarly balance research utility with stringent privacy protections under their respective jurisdictions. Researchers interested in conducting similar analyses using NCR data may submit requests through IKNL’s official data request process, as outlined in the IKNL privacy statement and data access policy. Access may be granted upon reasonable request and subject to approval by IKNL and its governance bodies.

Code availability

All software was developed in Python and is publicly available to ensure reproducibility. Complete code, scripts, and deployment instructions on the various Generative AI techniques are available at the repository: https://github.com/paper-support-materials/llm-driven-survival-prediction. The code used to generate the IKNL synthetic dataset can be accessed instead at https://github.com/daanknoors/synthetic_data_generation.

Declarations

Ethical approval

This retrospective study was conducted using (i) a fully synthetic breast cancer dataset generated by the Netherlands Comprehensive Cancer Organisation (IKNL), which contains no data relating to real or identifiable individuals, and (ii) analyses on pseudonymized and aggregated real-world population-based cancer registry data from the Netherlands Cancer Registry (NCR), performed within the secure research environment of IKNL. The use of fully synthetic data does not fall within the scope of human subjects research as defined by applicable national and European regulations. For analyses involving real-world registry data, all data were pseudonymized at source and processed in accordance with the governance framework of the Netherlands Cancer Registry. Data release by IKNL is subject to strict regulatory conditions and oversight by a Privacy Review Board and an independent scientific committee (NCR data application number: 25-00799). Ethical oversight for registry-based research using data from the Netherlands Cancer Registry is provided by the Netherlands Comprehensive Cancer Organisation, which acts as the responsible institutional authority for the Netherlands Cancer Registry. According to Dutch legislation, including the Medical Research Involving Human Subjects Act (WMO), this type of retrospective registry-based research does not require approval from a medical ethics review committee. The requirement for formal ethics approval was therefore waived.

Ethical standards

The research met all ethical guidelines, including adherence to the legal requirements of the study country.

Informed consent

Informed consent from individual patients was not required for this study. The synthetic dataset contains no real patient data. For the real-world registry component, data were collected as part of routine cancer registration under statutory mandate. The secondary use of these data for scientific research purposes is permitted under Dutch law and the General Data Protection Regulation (GDPR), subject to the governance and oversight procedures of the Netherlands Comprehensive Cancer Organisation (IKNL). The requirement for informed consent was therefore waived in accordance with national regulations and institutional governance procedures.

Declaration of Helsinki statement

All research involving human data was conducted in accordance with the ethical principles of the Declaration of Helsinki and its later amendments. The study relied exclusively on secondary use of registry data and synthetic data, involved no direct patient contact, and included no interventions. Appropriate safeguards for data protection, privacy, and confidentiality were applied throughout the study in compliance with applicable ethical and legal standards.

Competing interests

The authors declare no competing interests.

Footnotes

References

  • 1.Bhuyan SS, Sateesh V, Mukul N, Galvankar A, Mahmood A, Nauman M, et al. Generative artificial intelligence use in healthcare: opportunities for clinical excellence and administrative efficiency. J Med Syst. 2025;49(1). 10.1007/s10916-024-02136-1. [DOI] [PMC free article] [PubMed]
  • 2.Ullah W, Ali Q. Role of artificial intelligence in healthcare settings: a systematic review. J Med Artif Intell. 2025;8:24. 10.21037/jmai-24-294. [Google Scholar]
  • 3.Moulaei K, Yadegari A, Baharestani M, Farzanbakhsh S, Sabet B, Reza Afrash M. Generative artificial intelligence in healthcare: a scoping review on benefits, challenges and applications. Int J Multiling Med Inf. 2024;188:105474. 10.1016/j.ijmedinf.2024.105474. [DOI] [PubMed] [Google Scholar]
  • 4.Alowais SA, Alghamdi SS, Alsuhebany N, Alqahtani T, Alshaya AI, Almohareb SN, et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med Educ. 2023;23(1). 10.1186/s12909-023-04698-z. [DOI] [PMC free article] [PubMed]
  • 5.Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan, Ting DSW, et al. Large language models in medicine. Nat Med. 2023;29(8):1930–40. 10.1038/s41591-023-02448-8. [DOI] [PubMed] [Google Scholar]
  • 6.Khalighi S, Reddy K, Midya A, Pandav KB, Madabhushi A, Abedalthagafi M. Artificial intelligence in neuro-oncology: advances and challenges in brain tumor diagnosis, prognosis, and precision treatment. npj Precis Onc. 2024;8(1). 10.1038/s41698-024-00575-0. [DOI] [PMC free article] [PubMed]
  • 7.Wang Z, Liu Y, Niu X. Application of artificial intelligence for improving early detection and prediction of therapeutic outcomes for gastric cancer in the era of precision oncology. Semin Cancer Biol. 2023;93:83–96. 10.1016/j.semcancer.2023.04.009. [DOI] [PubMed] [Google Scholar]
  • 8.Lotter W, Hassett MJ, Schultz N, Kehl KL, Van Allen EM, Cerami E. Artificial intelligence in oncology: current landscape, challenges, and future directions. Cancer Discov. 2024;14(5):711–26. 10.1158/2159-8290.CD-23-1199. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Kolla L, Parikh RB. Uses and limitations of artificial intelligence for oncology. Cancer. 2024;130(12):2101–07. 10.1002/cncr.35307. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Fountzilas E, Pearce T, Baysal MA, Chakraborty A, Tsimberidou AM. Convergence of evolving artificial intelligence and machine learning techniques in precision oncology. NPJ Digit Med. 2025;8(1). 10.1038/s41746-025-01471-y. [DOI] [PMC free article] [PubMed]
  • 11.Mohammadzadeh I, Hajikarimloo B, Niroomand B, Eini P, Ghanbarnia R, Habibi MA, et al. Application of artificial intelligence in forecasting survival in high-grade glioma: systematic review and meta-analysis involving 79,638 participants. Neurosurg Rev. 2025;48(1):240. 10.1007/s10143-025-03419-y. [DOI] [PubMed] [Google Scholar]
  • 12.Feng Y, Wang Z, Cui R, Xiao M, Gao H, Bai H, et al. Clinical analysis and artificial intelligence survival prediction of serous ovarian cancer based on preoperative circulating leukocytes. J Ovarian Res. 2022;15(1). 10.1186/s13048-022-00994-2. [DOI] [PMC free article] [PubMed]
  • 13.Uwimana A, Gnecco G, Riccaboni M. Artificial intelligence for breast cancer detection and its health technology assessment: a scoping review. Comput Biol Med. 2025;184:109391. 10.1016/j.compbiomed.2024.109391. [DOI] [PubMed] [Google Scholar]
  • 14.Javanmard Z, Zarean Shahraki S, Safari K, Omidi A, Raoufi S, Rajabi M, et al. Artificial intelligence in breast cancer survival prediction: a comprehensive systematic review and meta-analysis. Front Oncol. 2025;14. 10.3389/fonc.2024.1420328. [DOI] [PMC free article] [PubMed]
  • 15.Ahn JS, Shin S, Yang SA, Park EK, Kim KH, Cho SI, et al. Artificial Intelligence in breast cancer diagnosis and personalized medicine. J Breast Cancer. 2023;26(5):405–35. 10.4048/jbc.2023.26.e45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;5999–6009. 10.5555/3295222.3295349.
  • 17.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. Adv Neural Inf Process Syst. 2020;6(159):1877–901. https://dl.acm.org/doi/abs/10.5555/3495724.3495883. [Google Scholar]
  • 18.Kumar A, Shankar A, Hollebeek LD, Behl A, Lim WM. Generative artificial intelligence (GenAI) revolution: a deep dive into GenAI adoption. J Educ Chang Bus Res. 2025;189:115160. 10.1016/j.jbusres.2024.115160. [Google Scholar]
  • 19.Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA. 2025;333(4):319. 10.1001/jama.2024.21700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhang K, Meng X, Yan X, Ji J, Liu J, Xu H, et al. Revolutionizing health care: the transformative impact of large language models in medicine. J Med Internet Res. 2025;27:e59069. 10.2196/59069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Sutskever I, Vinyals O, Le QV. Sequence to sequence learning with neural networks. Adv Neural Inf Process Syst. 2014;3104–12. https://dl.acm.org/doi/10.5555/2969033.2969173.
  • 22.Pezoulas VC, Zaridis DI, Mylona E, Androutsos C, Apostolidis K, Tachos NS, et al. Synthetic data generation methods in healthcare: a review on open-source tools and methods. Comput Struct Biotechnol J. 2024;23:2892–910. 10.1016/j.csbj.2024.07.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Miletic M, Sariyar M. Synthetic data generation methods for longitudinal and time series health data: a systematic review. BMC Med Inf Decis Mak. 2025;26(1):early access. 10.1186/s12911-025-03326-8. [DOI] [PMC free article] [PubMed]
  • 24.Lee SW. Kaplan-meier and Cox proportional hazards regression in survival analysis: statistical standard and guideline of Life Cycle Committee. Life Cycle. 2023;3:e8. 10.54724/lc.2023.e8.
  • 25.Zhang Y, Muller S. Robust variable selection methods with Cox model—a selective practical benchmark study. Briefings Bioinf. 2024, 10;25(6):bbae508. 10.1093/bib/bbae508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Ojeda FM, Müller C, Börnigen D, Trégouët DA, Schillert A, Heinig M, et al. Comparison of Cox model methods in a low-dimensional setting with few events. Genomics Proteomics Bioinf. 2016, Aug;14(4):235–43. 10.1016/j.gpb.2016.03.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Asghar N, Khalil U, Ahmad B, Alshanbari HM, Hamraz M, Ahmad B, et al. Improved nonparametric survival prediction using CoxPH, Random Survival Forest & DeepHit Neural Network. BMC Med Inf Decis Mak. 2024;24(1):120. 10.1186/s12911-024-02525-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Collett D. Modelling survival data in medical research. 3rd. New York, US: Chapman and Hall/CRC; 2014. Available from: 10.1201/b18041. [Google Scholar]
  • 29.Stensrud MJ, Hernán MA. Why use methods that require proportional hazards? Am J Epidemiol. 2025, 01;194(6):1504–06. 10.1093/aje/kwae361. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhang HH, Lu W. Adaptive Lasso for Cox’s proportional hazards model. Biometrika. 2007;94(3):691–703. 10.1093/biomet/asm037. [Google Scholar]
  • 31.Pickett KL, Suresh K, Campbell KR, Davis S, Juarez-Colunga E. Random survival forests for dynamic predictions of a time-to-event outcome using a longitudinal biomarker. BMC Med Res Methodol. 2021;21(1). 10.1186/s12874-021-01375-x. [DOI] [PMC free article] [PubMed]
  • 32.Qiu X, Gao J, Yang J, Hu J, Hu W, Kong L, et al. A comparison study of machine learning (random survival forest) and classic statistic (Cox proportional hazards) for predicting progression in high-grade glioma after Proton and carbon ion radiotherapy. Front Oncol. 2020;10. 10.3389/fonc.2020.551420. [DOI] [PMC free article] [PubMed]
  • 33.Taghavi Razavizadeh N, Salari M, Jafari M, Sabaghian E, Ghavami V. Comparison of two methods, gradient boosting and extreme gradient boosting to predict survival in covid-19 data. J Retailing Biostat Epidemiol. 2023;9(3):378–87. 10.18502/jbe.v9i3.15450. [Google Scholar]
  • 34.Shabani N, Yaseri M, Alimi R, Nazemian F, Zeraati H. Dynamic survival analysis via a landmarking-gradient boosting approach and its application to kidney transplant data. BMC Med Inf Decis Mak. 2025;25(1):368. 10.1186/s12911-025-03205-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Moncada-Torres A, van Maaren MC, Hendriks MP, Siesling S, Geleijnse G. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci Rep. 2021, 3;11(1):6968. 10.1038/s41598-021-86327-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Kantidakis G, Putter H, Lancia C, Boer J, Braat AE, Fiocco M. Survival prediction models since liver transplantation - comparisons between Cox models and machine learning techniques. BMC Med Res Methodol. 2020, 11;20(1):277. 10.1186/s12874-020-01153-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Katsimpokis D, van Odenhoven AEC, van Erp MAJM, Wenzel HHB, van der Aa MA, van Swieten MMH, et al. Ovarian cancer recurrence prediction: comparing confirmatory to real world predictors with machine learning. medRxiv. 2025. 10.1101/2025.03.04.25321571. [DOI] [PMC free article] [PubMed]
  • 38.Baidoo TG, Rodrigo H. Data-driven survival modeling for breast cancer prognostics: a comparative study with machine learning and traditional survival modeling methods. PLoS One. 2025;20(4):1–18. 10.1371/journal.pone.0318167. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. 2018;18(1):24. 10.1186/s12874-018-0482-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Lee C, Zame W, Yoon J, van der Schaar M. DeepHit: a deep learning approach to survival analysis with competing risks. Proc AAAI Conf Artif Intel. 2018;32(1). 10.1609/aaai.v32i1.11842.
  • 41.Liang S, Zhang J, Liu X, Huang Y, Shao J, Liu X, et al. The potential of large language models to advance precision oncology. eBiomedicine. 2025;115:105695. 10.1016/j.ebiom.2025.105695. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Chen D, Parsa R, Swanson K, Nunez JJ, Critch A, Bitterman DS, et al. Large language models in oncology: a review. BMJ Oncol. 2025;4(1):e000759. 10.1136/bmjonc-2025-000759. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Ghorbian M, Ghobaei-Arani M, Ghorbian S. Transforming breast cancer diagnosis and treatment with large language models: a comprehensive survey. Methods. 2025;239:85–110. 10.1016/j.ymeth.2025.04.001. [DOI] [PubMed] [Google Scholar]
  • 44.Romera-Paredes B, Barekatain M, Novikov A, Balog M, Kumar MP, Dupont E, et al. Mathematical discoveries from program search with large language models. Nature. 2024;625(7995):468–75. 10.1038/s41586-023-06924-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Yuksekgonul M, Bianchi F, Boen J, Liu S, Lu P, Huang Z, et al. Optimizing generative AI by backpropagating language model feedback. Nature. 2025;639(8055):609–16. 10.1038/s41586-025-08661-4. [DOI] [PubMed] [Google Scholar]
  • 46.Vacareanu R, Negru VA, Suciu V, Surdeanu M. From words to numbers: your large language model is secretly a capable regressor when given In-context examples. arXiv. 2024. p. 2404.07544. Available from: https://arxiv.org/abs/2404.07544.
  • 47.Dong Q, Li L, Dai D, Zheng C, Ma J, Li R, et al. In: Al-Onaizan Y, Bansal M, Chen YN, editors. A survey on In-context learning. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, Florida, USA: Association for Computational Linguistics; 2024. p. 1107–28. Available from: https://aclanthology.org/2024.emnlp-main.64/.
  • 48.Wang L, Chen X, Deng X, Wen H, You M, Liu W, et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. NPJ Digit Med. 2024;7(1). 10.1038/s41746-024-01029-4. [DOI] [PMC free article] [PubMed]
  • 49.Iannantuono GM, Bracken-Clarke D, Floudas CS, Roselli M, Gulley JL, Karzai F. Applications of large language models in cancer care: current evidence and future perspectives. Front Oncol. 2023;13. 10.3389/fonc.2023.1268915. [DOI] [PMC free article] [PubMed]
  • 50.Menz BD, Modi ND, Abuhelwa AY, Ruanglertboon W, Vitry A, Gao Y, et al. Generative AI chatbots for reliable cancer information: evaluating web-search, multilingual, and reference capabilities of emerging large language models. Eur J Criminol Cancer. 2025;218:115274. 10.1016/j.ejca.2025.115274. [DOI] [PubMed] [Google Scholar]
  • 51.Laios A, Theophilou G, Jong DD, Kalampokis E. The future of AI in ovarian cancer research: The large language models perspective. Cancer Control: J Moffitt Cancer Cent. 2023;30:10732748231197915. 10.1177/10732748231197915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Pau D, Bachot C, Monteil C, Vinet L, Boucher M, Sella N, et al. Comparison of anonymization techniques regarding statistical reproducibility. PLoS Digit Health. 2025;4(2):1–19. 10.1371/journal.pdig.0000735. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Wirth FN, Meurers T, Johns M, Prasser F. Privacy-preserving data sharing infrastructures for medical research: systematization and comparison. BMC Med Inf Decis Mak. 2021;21(1):242. 10.1186/s12911-021-01602-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Rollo C, Pancotti C, Birolo G, Rossi I, Sanavia T, Fariselli P. SYNDSURV: a simple framework for survival analysis with data distributed across multiple institutions. Comput Biol Med. 2024;172:108288. 10.1016/j.compbiomed.2024.108288. [DOI] [PubMed] [Google Scholar]
  • 55.El Emam K, Mosquera L, Bass J. Evaluating identity disclosure risk in fully synthetic health data: model development and validation. J Med Internet Res. 2020, Nov;22(11):e23139. 10.2196/23139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Qian Z, Callender T, Cebere B, Janes SM, Navani N, van der Schaar M. Synthetic data for privacy-preserving clinical risk prediction. Sci Rep. 2024;14(1):25676. 10.1038/s41598-024-72894-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Goncalves A, Ray P, Soper B, Stevens J, Coyle L, Sales AP. Generation and evaluation of synthetic patient data. BMC Med Res Methodol. 2020;20(1):108. 10.1186/s12874-020-00977-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Chatzichristos C, Katsimpokis D, Lai C, Van Santvliet L, Camarrone F, Knoors D, et al. Quality assessment of synthetic data in healthcare: a critical appraisal. JMIR Preprint. 2024; 10.2196/preprints.69390.
  • 59.Rujas M, Martín Gómez Del Moral Herranz R, Fico G, Merino-Barbancho B. Synthetic data generation in healthcare: a scoping review of reviews on domains, motivations, and future applications. Int J Multiling Med Inf. 2025;195:105763. 10.1016/j.ijmedinf.2024.105763. [DOI] [PubMed] [Google Scholar]
  • 60.Smith A, Lambert PC, Rutherford MJ. Generating high-fidelity synthetic time-to-event datasets to improve data transparency and accessibility. BMC Med Res Methodol. 2022;22(1). 10.1186/s12874-022-01654-1. [DOI] [PMC free article] [PubMed]
  • 61.Tucker A, Wang Z, Rotalinti Y, Myles P. Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. NPJ Digit Med. 2020;3(1). 10.1038/S41746-020-00353-9. [DOI] [PMC free article] [PubMed]
  • 62.Goyal M, Mahmoud QH. A systematic review of synthetic data generation techniques using Generative AI. Electronics. 2024;13(17):10.3390/electronics13173509. 10.3390/electronics13173509.
  • 63.Kang J, Chowdhry AK, Pugh SL, Park JH. Integrating artificial intelligence and machine learning into cancer clinical trials. Semin Radiat Oncol. 2023;33(4):386–94. Advances in Clinical Trial Design, Execution, and Implementation. 10.1016/j.semradonc.2023.06.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Auñón JM, Hurtado-Ramírez D, Porras-Díaz L, Irigoyen-Peña B, Rahmian S, Al-Khazraji Y, et al. Evaluation and utilisation of privacy enhancing technologies—a data spaces perspective. Data Brief. 2024;55:110560. 10.1016/j.dib.2024.110560. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative adversarial networks. Commun ACM. 2020;63(11):139–44. 10.1145/3422622. [Google Scholar]
  • 66.Jospin LV, Laga H, Boussaid F, Buntine W, Bennamoun M. Hands-on Bayesian neural networks—A tutorial for deep learning users. IEEE Comput Intell Mag. 2022;17(2):29–48. 10.1109/MCI.2022.3155327. [Google Scholar]
  • 67.Vallevik VB, Babic A, Marshall SE, Elvatun S, Brøgger HMB, Alagaratnam S, et al. Can I trust my fake data – a comprehensive quality assessment framework for synthetic tabular data in healthcare. Int J Multiling Med Inf. 2024;185:105413. 10.1016/j.ijmedinf.2024.105413. [DOI] [PubMed] [Google Scholar]
  • 68.Zhang J, Cormode G, Procopiuc CM, Srivastava D, Xiao X. PrivBayes: private data release via Bayesian networks. ACM Trans Database Syst. 2017;42(4):1–41. 10.1145/3134428. [Google Scholar]
  • 69.Ouadrhiri AE, Abdelhadi A. Differential privacy for deep and federated learning: a survey. IEEE Access. 2022;10:22359–80. 10.1109/ACCESS.2022.3151670. [Google Scholar]
  • 70.Elvatun S, Knoors D, Brant S, Jonasson C, Nygård JF. Synthetic data as external control arms in scarce single-arm clinical trials. PLoS Digit Health. 2025;4(1):e0000581. 10.1371/journal.pdig.0000581. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Wang L, Wen A, Fu S, Ruan X, Huang M, Li R, et al. A scoping review of OMOP CDM adoption for cancer research using real world data. NPJ Digit Med. 2025;8(1). 10.1038/s41746-025-01581-7. [DOI] [PMC free article] [PubMed]
  • 72.Stekhoven DJ, Bühlmann P. MissForest—non-parametric missing value imputation for mixed-type data. Bioinformatics. 2012;28(1):112–18. 10.1093/bioinformatics/btr597. [DOI] [PubMed] [Google Scholar]
  • 73.Uno H, Cai T, Pencina MJ, D’Agostino RB, Wei LJ. On the C-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Stat Med. 2011;30(10):1105–17. 10.1002/sim.4154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Cox DR. Regression models and life-tables. J R Stat Soc Ser B Stat Methodol. 1972;34(2):187–202. 10.1111/j.2517-6161.1972.tb00899.x. [Google Scholar]
  • 75.Chen Y, Jia Z, Mercola D, Xie X. A gradient boosting algorithm for survival analysis via direct optimization of concordance Index. Comput Math Method M. 2013;2013(1):1–8. 10.1155/2013/873595. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234–40. 10.1093/bioinformatics/btz682. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA: open and efficient foundation language models. arXiv. 2023. p. 2302.13971. Available from: https://arxiv.org/abs/2302.13971.
  • 78.Grattafiori A, Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, et al. The LLaMA 3 herd of models. arXiv. 2024. p. 2407.21783. Available from: https://arxiv.org/abs/2407.21783.
  • 79.Yang A, Yang B, Zhang B, Hui B, Zheng B, Yu B, et al. Qwen2.5 technical report. arXiv; 2025. p. 2412.15115. Available from: https://arxiv.org/abs/2412.15115.
  • 80.Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res. 2020;21(1):5485–551. [Google Scholar]
  • 81.Klein JP, Moeschberger ML. Survival analysis: techniques for censored and truncated data. In: Statistics for biology and health. New York, NY: Springer; 1997.
  • 82.Karlsson RKA, Willbo M, Hussain Z, Krishnan RG, Sontag D, Johansson FD. Using time-series privileged information for provably efficient learning of prediction models. Proceedings of Machine Learning Research. 2022. p. 5459–84. vol. 151.
  • 83.George B, Seals S, Aban I. Survival analysis and regression models. J Nucl Cardiol. 2014;21(4):686–94. 10.1007/s12350-014-9908-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Jeanselme V, Yoon CH, Tom B, Barrett J. Neural fine-gray: monotonic neural networks for competing risks. Proc Mach Learn Res. 2023;209(379):–392.
  • 85.Huang Y, Li W, Macheret F, Gabriel RA, Ohno-Machado L. A tutorial on calibration measurements and calibration models for clinical prediction models. J Am Med Inf Assoc. 2021;27(4):621–33. 10.1093/JAMIA/OCZ228. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Hartman N, Kim S, He K, Kalbfleisch JD. Pitfalls of the concordance index for survival outcomes. Stat Med. 2023;42(13):2179–90. 10.1002/sim.9717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Nielsen F. On a generalization of the Jensen–Shannon divergence and the Jensen–Shannon centroid. Entropy. 2020;22(2):10.3390/e22020221. 10.3390/e22020221. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 2 (151.7KB, pdf)
Supplementary Material 3 (149.4KB, pdf)

Data Availability Statement

The data that support the findings of this study are available from the Netherlands Comprehensive Cancer Organisation (IKNL). These data are used under license for the current study and are therefore not publicly available. In particular, the used synthetic dataset is available for research purposes on request to IKNL at: https://iknl.nl/en/ncr/synthetic-dataset. It is possible to obtain a regular version of the synthetic dataset with the variables in the standard NCR format (https://iknl.nl/en/ncr), or alternatively a dataset version compliant to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) [71]. The de-identified clinical data used in the deployment on real-world cancer registry cases consist of 183,304 patient records from the Netherlands Cancer Registry (NCR), collected and maintained by the Netherlands Comprehensive Cancer Organisation (IKNL). As a population-based cancer registry operating under Dutch and European Union data protection law, IKNL is statutorily mandated to collect cancer patient data from all hospitals in the Netherlands for scientific research and statistics in the interest of public health. The legal basis for this data collection, as detailed in IKNL’s privacy statement (https://iknl.nl/en/privacystatement), is grounded in the principle that requesting individual informed consent is not possible or appropriate, given the burden such requests would place on patients during the difficult period of cancer diagnosis and the critical societal need for comprehensive, unbiased cancer epidemiology data. This exemption operates within the framework of the EU General Data Protection Regulation (GDPR), specifically under the provision for scientific research and statistics in the public health interest. Patients are informed about the registry and retain the right to opt out of data inclusion at any time. Access to NCR data for research purposes is strictly controlled through an independent Supervisory Committee that includes patient representatives, medical professionals, and data protection experts. This Committee evaluates all data requests to ensure scientific merit, appropriate data use, and adequate privacy protection measures. For this study, all models were trained exclusively on synthetic data and subsequently deployed on real patient data only after approval by the Supervisory Committee. Due to the sensitive nature of individual-level patient data and strict privacy and data-governance regulations, direct access to the real-world dataset is not possible. All computations involving real patient data were performed on-site at IKNL’s secure facilities under controlled conditions, with data never leaving those premises. This regulatory approach to cancer registry data access reflects common practices in population-based cancer surveillance worldwide, including frameworks employed by registries such as the U.S. SEER Program, the Canadian Cancer Registry, and the Australian Cancer Database, which similarly balance research utility with stringent privacy protections under their respective jurisdictions. Researchers interested in conducting similar analyses using NCR data may submit requests through IKNL’s official data request process, as outlined in the IKNL privacy statement and data access policy. Access may be granted upon reasonable request and subject to approval by IKNL and its governance bodies.

All software was developed in Python and is publicly available to ensure reproducibility. Complete code, scripts, and deployment instructions on the various Generative AI techniques are available at the repository: https://github.com/paper-support-materials/llm-driven-survival-prediction. The code used to generate the IKNL synthetic dataset can be accessed instead at https://github.com/daanknoors/synthetic_data_generation.


Articles from BMC Medical Informatics and Decision Making are provided here courtesy of BMC

RESOURCES