Skip to main content
HemaSphere logoLink to HemaSphere
. 2026 Apr 14;10(4):e70336. doi: 10.1002/hem3.70336

Machine learning enhances risk stratification and treatment failure prediction in diffuse large B‐cell lymphoma

Mikkel Werling 1,2, Alexander D Fuglkjær 3, Peter Brown 1,4, Carsten U Niemann 1,2,4,✉, Rudi Agius 1,2,✉
PMCID: PMC13077786  PMID: 41987904

Abstract

Diffuse large B‐cell lymphoma (DLBCL) is the most common lymphoma subtype worldwide. Existing prognostic models, including the National Comprehensive Cancer Network International Prognostic Index (NCCN IPI), rely on small predefined variable sets, discretized inputs, and do not incorporate longitudinal clinical histories or nonlinear relationships. To address these limitations, we present MLAll, a machine learning model developed using clinical and laboratory data from 14,832 patients from the Danish Lymphoid Cancer Research (DALY‐CARE) resource (2005–2021). The primary objective was fixed‐time risk prediction of treatment failure within 2 years of first‐line therapy. Compared with the NCCN IPI at matched specificity, MLAll improved recall relatively by 22% (0.32–0.39), precision by 7% (0.56–0.60), and precision–recall AUC by 28% (0.46–0.59) in a blinded control population. Survival analyses were performed as a complementary evaluation and demonstrated reduced age‐related bias in risk stratification. MLAll identified more than four times as many low‐risk patients compared to the NCCN IPI, largely by correctly classifying older patients with favorable outcomes. Prognostic performance was nearly identical when evaluated on DLBCL alone and on a cohort including multiple aggressive lymphoma subtypes, suggesting that shared prognostic signals can be leveraged through joint modeling. These findings show that addressing key limitations of existing prognostic models enables more accurate and individualized risk stratification in DLBCL. MLAll provides a general framework for registry‐based prognostic modeling that can be adapted to other lymphoma subtypes and healthcare systems.


graphic file with name HEM3-10-e70336-g005.jpg

INTRODUCTION

Diffuse large B‐cell lymphoma (DLBCL) is the most common lymphoma worldwide. 1 Although outcomes have improved over recent decades with the introduction of rituximab and CHOP‐like (e.g., CHOP, MAXICHOP, and CHOEP) chemotherapy regimens, 2 , 3 30%–50% of patients still experience treatment failure. 4 , 5 Early risk assessment of patients is important for treatment selection and clinical trial enrollment, since frontline therapy is often guided by disease stage and prognostic indices. 1 , 6 , 7

Risk assessment in DLBCL remains dominated by the International Prognostic Index (IPI). 8 While several refinements have been proposed in the rituximab era, 8 , 9 , 10 , 11 , 12 , 13 , 14 , 15 the National Comprehensive Cancer Network International Prognostic Index (NCCN IPI) 9 is generally regarded as the most robust risk stratification tool. 11 , 16 Nevertheless, both the IPI and the NCCN IPI are limited by relying on a small set of baseline diagnostic variables and on discretization of continuous variables, 8 , 9 despite evidence that inclusion of long‐term histories and continuous measures can improve discrimination. 14 In addition, survival‐based indices may overweight the contribution of age 17 and competing mortality in contemporary cohorts, potentially masking other clinically relevant disease‐related factors. 18

In contrast, modern machine learning (ML) approaches can alleviate many of these limitations by incorporating larger numbers of predictors, capturing nonlinear relationships, and learning from longitudinal patient histories now routinely stored in electronic health records (EHRs). 19 Prior studies have shown that such methods can outperform rule‐based prognostic scores in heterogeneous diseases such as DLBCL 20 , 21 , 22 and support real‐time, individualized risk estimation. 23

In this study, we use ML as a targeted strategy to overcome key limitations of existing prognostic scoring systems, such as IPI and NCCN IPI. Specifically, we evaluate whether ML‐based models trained on longitudinal, routinely collected health data can improve fixed‐time prediction of treatment failure among patients receiving CHOP‐like first‐line therapy. We further examine whether joint training across multiple aggressive lymphoma subtypes can maintain or enhance prognostic performance, reflecting the hypothesis that related diseases share clinically relevant prognostic signals. Similar cross‐disease modeling strategies have shown promise in oncology, including pan‐cancer mortality prediction. 24 Together, this work aims to provide a more empirically grounded and flexible risk stratification framework that others can build upon as both data sources and computational methods continue to evolve. 7 , 14 , 16

METHODS

A consort diagram detailing our inclusion and exclusion criteria is shown in Figure 1. We initially identified 22,115 patients registered in the Danish National Lymphoma Registries (LYFO) 25 who received a lymphoma diagnosis from 2005 to 2023. After applying our exclusion criteria (Supporting Information S1: A.1), our cohort consisted of 14,832 eligible patients who initiated first‐line treatment before 2022, ensuring at least two years of follow‐up. Data were split into a training set (85%) and a blinded test set (15%), stratified by death, relapse, and lymphoma subtype. A small number of patients (<1%) were registered with a preliminary identification number, which precluded linkage to national demographic registries. These patients were excluded from all analytic data sets (Supporting Information S1: A.1).

Figure 1.

Figure 1

Consort diagram of patients. We excluded patients with no first‐line treatment information, as we could not define a prediction point for these patients. We thereafter removed patients who received treatment after 2021 to ensure that we had a full 2‐year follow‐up for all patients included in the study. Patients registered with preliminary identification numbers (n = 134) were excluded from the analytic train and test sets due to missing demographic linkage. The subtypes of lymphomas included were diffuse large B‐cell lymphoma (DLBCL), follicular lymphoma (FL), classical Hodgkin lymphoma (cHL), marginal zone lymphoma (MZL), mantle cell lymphoma (MCL), Waldenstrom macroglobulinemia (WM), non‐Hodgkin lymphoma not otherwise specified (nHL), small lymphocytic lymphoma (SLL), and lymphocyte‐predominant Hodgkin's disease (HD‐LP). See Supporting Information S1: A.1 for further details.

For the primary evaluation, we assessed model performance in DLBCL patients receiving CHOP‐like first‐line therapy in the test set. Unless otherwise specified, performance metrics are reported for this population (Figure 1). Sensitivity analyses were performed for patients receiving non‐CHOP‐like treatments, patients receiving R‐CHOP‐like treatments, and all patients regardless of treatment (Supporting Information S1: B.6).

We defined a composite endpoint of treatment failure as either death, relapse, or initiation of second‐line treatment within 2 years of first‐line treatment. The time of treatment failure was the earliest occurrence of any of these events. Performance metrics for death and relapse as individual outcomes, as well as models trained specifically on each endpoint, are reported in Supporting Information S1: B.5. As a sensitivity analysis, we evaluated models using overall survival as the endpoint without retraining, with results presented in Supporting Information S1: C.10.

Data sources

For all patients, we gathered data from 8 different data sources using the Danish Lymphoid Cancer Research (DALY‐CARE) data resource 26 : prescription medication, laboratory tests, surgical and other procedures, pathology codes, diagnosis codes, blood cultures, microbiology charts, and quality registers 25 , 27 (Figure 2).

Figure 2.

Figure 2

Schematic overview of our data‐driven feature generation process. (A) Using the quality registers, we identify relevant patients and define outcomes. We then extract data from the other seven different data sources. (B) These data sources are then processed to create different features. An example of laboratory tests is used to illustrate our approach. If a variable is recorded for more than 10% of patients in the training data within 3 years before first‐line treatment, we aggregate the data within all lookback (LB) windows using the relevant aggregation functions for each data source. This process results in 3881 candidate features. (C) Using lasso regression, we eliminate uninformative features. We reinclude the variables used by the NCCN IPI if they were removed by the LASSO regression procedure. This results in a final feature matrix with 87 features.

Feature creation and selection

We developed a data‐driven pipeline to systematically extract features from the longitudinal patient histories across all EHR modalities (Figure 2). We accounted for the irregular time series of EHR data by using the timeseriesflattener framework 28 to aggregate predictors within different lookback windows before initiation of first‐line treatment. For all 7 EHR data modalities (excluding the quality registry), we included three different lookback windows: 3 months, 1 year, and 3 years (Supporting Information S1: A.2 and Supporting Information S1: Table 3). Subsequently, we used lasso regression for feature selection, resulting in a final set of 87 features used for model training and evaluation (Supporting Information S1: Table 4). All NCCN IPI variables were retained regardless of LASSO selection.

Model design overview

To investigate how feature complexity and training population affect clinical risk prediction, we developed three ML models with increasing scope:

MLIPI used only the continuous versions of the five NCCN IPI variables to isolate the effect of modeling technique. MLDLBCL used 87 features from patient records but was trained solely on DLBCL cases to assess whether disease‐specific training improves performance. MLAll, the primary model, used the same feature set of 87 features but was trained across all aggressive lymphoma subtypes to leverage shared clinical patterns and potential gains from pooled learning (Supporting Information S1: Table 1).

To further explore model specialization, we trained subtype‐specific models for other lymphomas (e.g., FL, MCL, and cHL) using the same 87 features. These models permitted direct comparison of subtype‐specific training with broader pooled training. Subtype‐specific models are described in detail in Supporting Information S1: A.5.

Decision thresholds for all models are denoted using superscripts (e.g., MLAll0.5 for a 0.5 threshold). Adjusting the decision thresholds allows for more targeted comparisons with, for instance, matching the recall or precision of another model (e.g., NCCN IPI). All models were trained using the XGBoost algorithm, 29 a gradient boosting framework well‐suited for structured clinical data. 30 Model hyperparameters are provided in Supporting Information S1: A.4.

Baseline comparisons

The NCCN IPI was used as the main clinical comparator to the ML models. Our primary objective was fixed‐time risk prediction, where the goal was to estimate the probability of treatment failure within two years of first‐line therapy. As the NCCN IPI is applied in practice as a categorical risk score rather than a probability model, we evaluated it both as a binary classifier and as a risk stratification tool.

For binary classification, patients with an NCCN IPI score of 6 or higher were considered high risk (treatment failure), while all others were classified as low risk (treatment success). For threshold‐independent metrics such as area under the precision–recall curve (PR‐AUC), each integer score was treated as a decision threshold. Uncertainty estimates for performance metrics were obtained using nonparametric bootstrapping with 1000 resamples of the test set to derive 95% confidence intervals. Statistical significance of pairwise differences in performance was assessed using paired bootstrap resampling 31 , 32 (Supporting Information S1: A.9).

To enable comparison of risk stratification, we grouped patients into four risk categories based on the predicted probability of treatment failure estimated by MLAll. These categories were designed to align with the four NCCN IPI risk groups. Cutoffs were selected according to two criteria: (i) the “Low”‐risk group should have a comparable event rate to the NCCN IPI's low‐risk group and (ii) each group should be large enough to be clinically meaningful. While these cutoffs are somewhat arbitrary and can be adjusted for specific use cases, we used the following thresholds: <0.1 for “Low,” 0.1–0.3 for “Low‐Intermediate,” 0.3–0.65 for “Intermediate‐High,” and ≥0.65 for “High.”

Survival analysis

The NCCN IPI was originally developed for risk stratification of time‐to‐event outcomes, such as overall and progression‐free survival at five years, rather than for fixed‐time binary classification of treatment failure. 9 To allow comparison within that framework and to assess how MLAllrisk groups behave over time, we examined Kaplan–Meier Estimates (KMEs) 33 of treatment‐free survival across the corresponding risk categories. KMEs are reported at two years after first‐line treatment in the main text, and at five and ten years in Supporting Information S1: C.3 and C.6. To evaluate calibration and discrimination in a time‐to‐event setting, we fit Cox proportional hazards models using risk group assignment as the sole predictor and computed concordance indices and Integrated Brier scores based on these models.

Explainability

To explain what variables contributed to the predictions made by the MLAll model, we calculated Shapley Additive Explanations (SHAPs) 34 across all features for patients in the test set. We provide SHAP analyses for the MLDLBCL and MLIPI in Supporting Information S1: C.5.2.

Data sharing and reproducibility

The data used for this project are a subset of the Danish Lymphoid Cancer Research (DALY‐CARE) data resource. We refer to the paper 26 introducing the data resource for details concerning the data and how to access them. All scripts and feature engineering pipelines for this project are publicly available on GitHub to enable replication and testing in external cohorts. To support easier clinical integration and external validation, we derived a more minimal model using recursive feature elimination (RFE), retaining only 30 features (Supporting Information S1: Table 4) while preserving near‐equivalent performance to the full model (Supporting Information S1: C.8). This reduced model is well‐suited for settings with more limited data availability.

RESULTS

A detailed overview of clinical characteristics for the included patients is shown in Table 1, demonstrating similar characteristics between the training and the unseen test set. At the same time, age and risk factors differ between CHOP‐like‐treated DLBCL patients and DLBCL patients receiving other treatments.

Table 1.

Summary of baseline characteristics for the 14,698 patients included in the analytical training and test set, stratified by diagnosis and first‐line treatment regimen.

DLBCL + CHOP‐like DLBCL ‐ CHOP‐like Other lymphomas
N = 5848 N = 1209 N = 7641
Train Test Train Test Train Test
N = 4951 N = 897 N = 1046 N = 163 N = 6490 N = 1151
Age ≤40 265 (5.35) 52 (5.80) 31 (2.96) 5 (3.07) 918 (14.14) 155 (13.47)
41–60 1128 (22.78) 213 (23.75) 178 (17.02) 21 (12.88) 1578 (24.31) 268 (23.28)
61–75 2259 (45.63) 401 (44.70) 393 (37.57) 66 (40.49) 2562 (39.48) 461 (40.05)
>75 1299 (26.24) 231 (25.75) 444 (42.45) 71 (43.56) 1432 (22.06) 267 (23.20)
Sex Males 2804 (56.64) 519 (57.86) 588 (56.21) 83 (50.92) 3638 (55.06) 630 (54.74)
AA stage I 941 (19.01) 194 (21.63) 406 (38.81) 63 (38.65) 1221 (18.18) 220 (19.11)
II 859 (17.35) 148 (16.50) 96 (9.18) 10 (6.13) 1046 (16.12) 185 (16.07)
III 964 (19.47) 177 (19.73) 100 (9.56) 18 (11.04) 994 (15.32) 161 (13.99)
IV 2143 (43.28) 369 (41.14) 408 (39.01) 67 (41.10) 3155 (48.61) 581 (50.48)
NA 44 (0.89) 9 (1.00) 36 (3.44) 5 (3.07) 74 (1.14) 4 (0.35)
ECOG PS ≥2 808 (16.32) 146 (16.28) 391 (37.39) 65 (39.88) 565 (8.71) 104 (9.03)
NA 21 (0.42) 1 (0.11) 3 (0.29) 0 (0.00) 22 (0.34) 3 (0.26)
B Symptoms Present 2015 (40.70) 378 (42.14) 336 (32.12) 69 (42.33) 2293 (35.33) 398 (34.59)
NA 119 (2.40) 18 (2.01) 74 (7.07) 9 (5.52) 109 (1.68) 25 (2.17)
LDH ≤ULN 2203 (44.50) 410 (45.71) 504 (48.18) 83 (50.92) 4346 (67.86) 781 (67.85)
1–3 × ULN 2174 (43.91) 395 (44.04) 433 (41.40) 53 (32.52) 1874 (28.88) 327 (28.41)
>3 × ULN 452 (9.13) 73 (8.14) 75 (7.17) 24 (14.72) 103 (1.59) 14 (1.22)
NA 122 (2.46) 19 (2.12) 34 (3.25) 3 (1.84) 167 (2.57) 29 (2.52)
ALC ≤0.84 × 109/L 1124 (22.70) 193 (21.52) 319 (30.50) 48 (29.45) 868 (13.37) 164 (14.25)
NA 202 (4.08) 30 (3.34) 46 (4.40) 10 (6.14) 248 (3.82) 35 (3.04)
Albumin ≤35 g/L 1482 (29.93) 251 (27.98) 407 (38.91) 68 (41.72) 1480 (22.80) 261 (22.68)
NA 825 (16.66) 166 (18.51) 188 (17.97) 28 (17.18) 1053 (16.22) 186 (16.16)
HGB <10 g/dL 551 (11.13) 92 (10.26) 98 (9.37) 26 (15.95) 764 (11.77) 146 (12.68)
NA 35 (0.71) 5 (0.56) 14 (1.34) 2 (1.23) 49 (0.76) 10 (0.87)
Bulky disease ≥10 cm 914 (18.46) 171 (19.06) 122 (11.66) 17 (10.43) 669 (10.31) 111 (9.64)
NA 1191 (24.04) 219 (24.41) 323 (30.88) 52 (31.9) 1857 (28.61) 350 (30.41)
IPI 0 366 (7.39) 75 (8.36) 47 (4.49) 5 (3.07) 871 (13.42) 145 (12.60)
1 1029 (20.78) 192 (21.40) 216 (20.65) 34 (20.86) 1609 (24.79) 286 (24.85)
2 1189 (24.02) 215 (23.97) 246 (23.52) 41 (25.15) 2008 (30.94) 373 (32.41)
3 1194 (24.12) 218 (24.30) 253 (24.19) 44 (26.99) 1302 (20.06) 238 (20.68)
4 757 (15.29) 130 (14.49) 150 (14.34) 20 (12.27) 384 (5.92) 61 (5.30)
5 228 (4.61) 39 (4.35) 65 (6.21) 12 (7.36) 55 (0.85) 10 (0.87)
NA 188 (3.80) 28 (3.12) 69 (6.60) 7 (4.29) 261 (4.02) 38 (3.30)
NCCN IPI 0 52 (1.05) 12 (1.34) 0 (0.00) 1 (0.61) 362 (5.58) 53 (4.60)
1 258 (5.21) 52 (5.80) 15 (1.43) 1 (0.61) 596 (9.18) 106 (9.21)
2 583 (11.78) 101 (11.26) 59 (5.64) 8 (4.91) 999 (15.39) 183 (15.90)
3 953 (19.25) 188 (20.96) 180 (17.21) 20 (12.27) 1419 (21.86) 253 (21.98)
4 1126 (22.74) 206 (22.97) 260 (24.86) 50 (30.67) 1490 (22.96) 282 (24.50)
5 984 (19.87) 171 (19.06) 240 (22.94) 38 (23.31) 995 (15.33) 179 (15.55)
6 585 (11.82) 105 (11.71) 144 (13.77) 23 (14.11) 307 (4.73) 48 (4.17)
7 201 (4.06) 28 (3.12) 66 (6.31) 9 (5.52) 58 (0.89) 11 (0.96)
8 28 (0.57) 7 (0.78) 13 (1.24) 6 (3.68) 7 (0.11) 0 (0.00)
NA 181 (3.66) 27 (3.01) 69 (6.60) 7 (4.29) 257 (3.96) 36 (3.13)

Note: Patients are divided into training and held‐out test sets across three cohorts: (1) DLBCL patients treated with CHOP‐like regimens, (2) DLBCL patients treated with other regimens, and (3) patients with other lymphoma subtypes. Percentages are column‐wise within each subgroup. Missing values are reported as NA.

Abbreviations: AA stage, Ann Arbor Stage; ALC, Absolute Lymphocyte Count; ECOG PS, ECOG performance status; HGB, hemoglobin; IPI, International Prognostic Index; LDH, lactate dehydrogenase; NCCN IPI, National Comprehensive Cancer Network IPI; ULN, Upper Limit of Normal.

MLAll outperforms NCCN IPI across all considered metrics

We compared the performance of MLAll with the NCCN IPI in CHOP‐like‐treated DLBCL patients in the test set. MLAll0.5 outperformed the NCCN IPI across all considered threshold‐independent and classification metrics (PR‐AUC: 0.59 vs. 0.46; ROC‐AUC: 0.80 vs. 0.72; MCC: 0.34 vs. 0.27) (Table 2). Additional comparator models, including a logistic regression baseline and the transformer‐based TabPFN 35 classifier, outperformed the NCCN IPI but showed poorer performance than MLAll across all performance metrics (Supporting Information S1: Material B.7). As a secondary evaluation, we assessed prediction of relapse and death in isolation. Performance for relapse‐only prediction was consistently lower than for the composite endpoint across all models, with only marginal gains from outcome‐specific training (Supporting Information S1: B.5, C.1).

Table 2.

Performance metrics on the test set.

Metric NCCN IPI
MLIPI0.5
MLDLBCL0.5
MLAll0.5
PR‐AUC 0.46 (0.42–0.51) 0.51 (0.46–0.57) ★ 0.58 (0.53–0.64) ★ 0.59 (0.52–0.63) ★
ROC‐AUC 0.72 (0.69–0.76) 0.74 (0.71–0.78) 0.79 (0.76–0.82) ★ 0.80 (0.77–0.82) ★
MCC 0.27 (0.20–0.34) 0.28 (0.21–0.34) 0.37 (0.30–0.44) ★ 0.33 (0.26–0.40)
Precision 0.56 (0.49–0.64) 0.57 (0.50–0.64) 0.63 (0.57–0.69) ★ 0.60 (0.54–0.67)
Recall 0.32 (0.26–0.37) 0.33 (0.27–0.39) 0.43 (0.37–0.49) ★ 0.39 (0.33–0.44) ★
Specificity 0.90 (0.88–0.92) 0.90 (0.88–0.92) 0.90 (0.88–0.92) 0.90 (0.87–0.92)

Note: Superscripts for ML models refer to the decision threshold used for classification. The subscript refers to different models. Bold numbers indicate the highest scores across all models for the specific performance metric. A star (★) denotes a statistically significant improvement compared with the NCCN IPI, based on paired bootstrap resampling (2000 replicates). Differences were considered significant when the 95% bootstrap confidence interval for ∆ excluded 0. Note that metrics for NCCN IPI are only calculated for patients who did not have missing values in the variables used to calculate the indices.

Abbreviation: MAC, the Matthews correlation coefficient.

MLAll improves precision–recall trade‐offs compared with NCCN IPI

At a fixed specificity, corresponding to the NCCN IPI high‐risk threshold, MLAll0.5 improved recall from 0.32 to 0.39 (22%) and precision from 0.56 to 0.60 (7%). The precision–recall curve for MLAll was consistently above that of the NCCN IPI across all thresholds (Figure 3). Varying the classification threshold allowed for different trade‐offs between recall and precision: MLAll0.4 increased recall from 0.32 to 0.51 while maintaining NCCN IPI's precision; MLAll0.59 increased precision from 0.56 to 0.67 at the recall level of NCCN IPI.

Figure 3.

Figure 3

Precision–recall curve of the NCCN IPI and MLAll . The legend shows the average precision (AP) as measured as the area under the curve. The bottom left arrow shows the “High”‐risk group from NCCN IPI and captions report the precision and recall of the model. The other two arrows point to MLAll thresholds that match either the recall (upper left) or the precision (lower right) of the NCCN IPI “High”‐risk group.

To quantify the clinical implications of improved precision, we calculated the number needed to treat (NNT). At matched recall, the high‐risk group identified by MLAll0.59 had a 67.5% event rate versus 56.4% for the NCCN IPI. This corresponds to an effective NNT of 9.0; for every 9 patients flagged as high risk by MLAll, one additional treatment failure is correctly identified compared to the NCCN IPI.

MLAll improves performance by addressing limitations of traditional models

To disentangle the sources of this improvement, we conducted an ablation‐style analysis. A minimal model, MLIPI0.5, using the same variables as NCCN IPI but retaining continuous values, improved PR‐AUC from the NCCN IPI's 0.46 to 0.51. Adding broader features and longitudinal histories in MLDLBCL0.5 increased PR‐AUC to 0.58. Joint training in MLAll0.5 further increased PR‐AUC to 0.59.

MLAllprovides more granular risk stratification than NCCN IPI

Using the derived risk groups as predictors in time‐to‐event analyses, MLAll showed higher concordance and better calibration than the NCCN IPI. The C‐index was 0.71 (95% CI: 0.68–0.73) for MLAll compared with 0.65 (95% CI: 0.62–0.68) for the NCCN IPI. Similarly, the Integrated Brier score was lower for MLAll (0.165 [95% CI: 0.153–0.177]) than for the NCCN IPI (0.182 [95% CI: 0.170–0.195]).

Based on these risk groups, MLAllidentified more than four times as many low‐risk patients compared to NCCN IPI while maintaining comparable treatment failure rates (Figures 4 and 5). Kaplan–Meier estimates showed clearer separation between MLAll risk groups, particularly between the “Low” and “Low‐Intermediate” categories (Figure 5).

Figure 4.

Figure 4

Scatterplot of predictions made by the NCCN IPI and MLAll . The x‐axis shows the estimated probability from the MLAll model and the y‐axis shows the risk groups identified by the NCCN IPI. Points show individual patients and are colored by whether they had treatment failure within 2 years (red) or not (blue). The red background shading shows the different risk groups created based on the estimated probability of the MLAll model.

Figure 5.

Figure 5

Kaplan–Meier curves and estimations for different risk groups. Only patients in the test set with a DLBCL diagnosis who received CHOP‐like treatment were included in this analysis. The first panel shows Kaplan–Meier curves over time. The dashed vertical line highlights the time point 2 years after treatment. Shaded areas show the 95% CI around probability estimates. Crosses on lines show censored patients. Median follow‐up times for the NCCN IPI risk groups were 9.78 years for the “Low”‐risk group, 9.46 years for the “Low‐Intermediate”‐risk group, 8.58 years for the “Intermediate‐High”‐risk group, and 7.70 for the “High”‐risk group. Median follow‐up times for the MLAll risk groups were 8.79 years for the “Low”‐risk group, 10.28 years for the “Low‐Intermediate”‐risk group, 9.60 years for the “Intermediate‐High”‐risk group, and 6.67 years for the “High”‐risk group. The second panel shows a forest plot of the Kaplan–Meier estimates (with estimates converted into percentages) two years after first‐line treatment accompanied by 95% confidence intervals in parentheses.

To examine agreement at the patient level, we performed a concordance–discordance analysis (Supporting Information S1: C.9.1). Patients classified as higher risk by MLAll but lower risk by the NCCN IPI had poorer 2‐year treatment‐free survival (54% [95% CI: 37%–78%]), whereas those classified as lower risk by MLAll but higher risk by the NCCN IPI had more favorable outcomes (77% [95% CI: 72%–82%]).

MLAllreduces age‐related bias in risk stratification

The NCCN IPI risk score showed a stronger correlation with patient age (ρ = 0.54) than the continuous predicted probabilities of MLAll (ρ = 0.37). Among patients classified as high risk, the 2‐year treatment‐free survival rate was similar across older age strata for MLAll (60–75 years: 30% [95% CI: 19%–45%] vs. >75 years: 29% [95% CI: 19%–45%]), whereas the NCCN IPI showed greater variation (57% [95% CI: 45%–71%] vs. 32% [95% CI: 23%–45%)] (Figure 6). MLAll identified 160 low‐risk patients over 60 years of age, whereas the NCCN IPI identified only one. Additional age‐stratified analyses are presented in Supporting Information S1: C.3. Kaplan–Meier estimates of treatment‐free survival at 2, 5, and 10 years across risk groups and cohorts are provided in Supporting Information S1: C.6.

Figure 6.

Figure 6

Heatmaps of Kaplan–Meier estimates and counts grouped by age and risk groups. Only DLBCL patients who received CHOP‐like treatment in the test set were included. The x‐axis shows different risk groups and the y‐axis shows different age groups. Left panels show results from MLAll and the right panel shows the results from NCCN IPI. The numbers and colors in the upper panels show Kaplan–Meier estimates for treatment‐free survival (TFS) 2 years after first‐line treatment with 95% CIs in parentheses. The numbers and colors in the lower panels show the number of patients in each group. Missing squares indicate that no patients were in both the age and risk group.

MLAll shows stable performance under joint modeling

In the cohort including all lymphoma patients, MLAll achieved a PR‐AUC of 0.61 (95% CI: 0.57–0.64) and a ROC‐AUC of 0.79 (95% CI: 0.77–0.81), similar to its performance in CHOP‐like‐treated DLBCL patients (PR‐AUC: 0.59; 95% CI: 0.52–0.63; ROC‐AUC: 0.80; 95% CI: 0.77–0.82) (Supporting Information S1: Table 10 and Table 2). Kaplan–Meier estimates also showed consistent stratification when performed on all lymphoma patients (Supporting Information: C.6). Training jointly across lymphoma subtypes generally improved prognostic performance compared to subtype‐specific models (Figure 7 and Supporting Information S1: Table 14).

Figure 7.

Figure 7

Comparison between jointly trained and subtype‐specific models. Each point compares the AU‐PRC of a subtype‐specific model (x‐axis) with that of the pooled model, MLAll (y‐axis), for a given lymphoma subtype. Point size reflects the number of patients per subtype. Points above the diagonal indicate better performance of the MLAll model. Arrows indicate performance directionality. MLAll achieves comparable or superior performance in most subtypes despite not being tailored to any one group, supporting its utility as a generalizable risk stratification tool. Figure adapted from Bjerregaard‐Michelsen et al. (2025).

In WM, the subtype‐specific model showed higher performance than MLAll. Across all other subtypes, differences in PR‐AUC and ROC‐AUC were never more than 0.02 lower for MLAll than the best subtype‐specific models and were often substantially higher (Figure 7 and Supporting Information S1: B.6). Full subtype‐specific performance results, as well as analyses of all DLBCL patients and DLBCL patients not receiving CHOP‐like regimens, are reported in Supporting Information S1: B.6.

MLAll captures prognostic signals across clinical domains and time windows

Feature importance was assessed using absolute SHAP values. Core NCCN IPI variables such as age, LDH, and ECOG performance status were among the most influential features. Additional disease‐specific factors such as tumor diameter, hemoglobin levels, and platelet counts also ranked highly. Markers of healthcare interactions, including referrals for radiological procedures and outpatient visits, were also associated with model predictions (Figure 8). Feature importance was stable across training and test sets (Supporting Information S1: C.5.1).

Figure 8.

Figure 8

SHAP plot for the 20 most important features for the MLAll model as measured by absolute mean SHAP values. Each point represents a single patient; the x‐axis shows the SHAP value (the feature's contribution to increasing or decreasing the predicted probability of treatment failure) and the y‐axis lists the feature names. Point color reflects the individual feature value (red = higher value, blue = lower value; and gray = missing). Feature names include the aggregation function (e.g., count) if not a static baseline feature, and the lookback window used for feature generation is indicated in parentheses. Sex was encoded as 0 and 1, with males being 1.

DISCUSSION

Our results show that addressing key limitations of existing prognostic models leads to meaningful improvements in predicting treatment‐free survival in DLBCL. Together, these findings support a move toward data‐driven prognostic models that better reflect real‐world patient trajectories.

In the sections that follow, we consider the broader implications of our framework alongside the methodological, practical, and clinical factors relevant to its interpretation and future application.

Addressing key limitations of traditional prognostic models

Across all considered evaluation metrics, MLAll outperformed the NCCN IPI (Table 2 and Supporting Information S1: C.9.1). The ablation‐style analyses indicate that these gains arise from progressively addressing key limitations of existing prognostic models. Models that retained raw clinical information within a flexible modeling framework showed improved performance, with further gains observed when longitudinal data from multiple modalities were incorporated (Table 2 and Supporting Information S1: Table 1). This is notable, as most clinical prognostic models rely exclusively on baseline measurements. 8 , 36 , 37 While the absolute gains in population‐level metrics such as ROC‐AUC and PR‐AUC may appear modest, the ability to generate more nuanced and individualized risk estimates is clinically important.

We view this work as one step in an iterative development process. By showing that prognostic performance improves when specific modeling limitations are addressed, our results outline a clear path for continued refinement. Further gains are likely achievable through targeted changes in data representation and model design to overcome additional limitations.

Limitations of aggregated longitudinal representations

Although our aggregated approach improved prognostic accuracy, it discarded much of the temporal information. As a result, we likely underestimate the full value of long‐term histories. In our benchmarking (Supporting Information S1: B.7), gradient‐boosted trees (XGBoost) outperformed logistic regression and the transformer‐based TabPFN classifier. 35 However, TabPFN, like our current model, operates on aggregated features rather than raw temporal sequences, and its performance should therefore not be taken as evidence against deep learning approaches in this setting. The greatest potential of deep learning methods such as recurrent neural networks 38 or attention‐based models 39 lies in capturing temporal dependencies across a full disease trajectory of a patient. Our findings therefore support explicit temporal learning, incorporating the whole patient trajectory as a logical next step. 40

Improved risk stratification and age‐adjusted prognosis

By incorporating additional disease‐related variables, MLAll generated prognostic estimates that are less driven by age than the NCCN IPI. Despite identifying substantially more older patients as low risk, the model maintained strong discrimination between risk groups, including at five years after first‐line treatment (Supporting Information S1: C.3, C.6). Reduced reliance on age enabled identification of older patients with favorable biological and clinical profiles who may otherwise be classified as high risk. This may have implications for clinical trial eligibility and for the use of intensive treatment strategies in selected elderly patients who might otherwise be undertreated based on age alone. Future work should further distinguish treatment failure driven by toxicity from disease progression.

Clinical utility and use cases of MLAll

Beyond informing future methodological work, this framework also addresses an emerging clinical need. As phase 3 trials of frontline bispecific antibodies and CAR‐T therapies begin to report results, there is growing demand for prognostic tools that can support patient selection and risk stratification. More accurate individualized risk estimates may help identify patients most likely to benefit from intensified strategies, while reducing unnecessary toxicity and cost for lower risk individuals. Pending external validation, a model like MLAll could ultimately support trial enrichment and contribute to more individualized treatment strategies as new therapies are integrated into practice. We do note, however, that effective clinical implementation depends less on universal generalizability than on context‐aware adaptation and validation within each healthcare setting. 41

Expanding data modalities: Toward biological integration

A key limitation of existing prognostic models is their restricted feature scope. Incorporation of additional modalities, such as molecular or genomic data, represents a natural extension that may further refine risk stratification. For instance, molecular subgroupings such as COO, 42 LymphGen, 43 and DLBCLass 44 are known to associate with prognosis, but were not available for the full cohort in this study. Integration of such data could help link ML‐derived risk estimates more directly with tumor biology. That said, prior work in CLL has shown that registry‐based models can capture most of the prognostic signal without genomics, 45 suggesting that molecular data may offer incremental, but valuable improvements.

Generalization across subtypes and learning from related diseases

Training across multiple lymphoma subtypes generally maintained or improved prognostic performance compared with subtype‐specific models (Figure 7 and Supporting Information S1: B.6). This suggests that subtypes share clinically relevant features and supports the idea that learning across related diseases via techniques like transfer learning can enhance performance, particularly for rarer subtypes with limited data. 24 , 26 , 46 In this study, cross‐subtype learning was implemented in a simple form by pooling data across subtypes during model training. More advanced approaches such as meta‐learning 47 , 48 , 49 , 50 or disease‐similarity‐aware sampling 51 could further improve performance, particularly for underrepresented cohorts 52 , 53 and in the presence of domain shifts. 26 , 47 , 54 Accordingly, these results likely represent a conservative baseline: they demonstrate that cross‐subtype learning is viable and promising, 24 and improvements can likely be gained by implementing targeted transfer‐learning strategies.

Model interpretability and handling of missing data

Using SHAP values to assess feature importance, we found features beyond the NCCN IPI variables that were highly informative (Figure 8). Many of these have been proposed as potential prognostic biomarkers for DLBCL disease severity, including platelet count, 55 hemoglobin levels, 56 , 57 , 58 and tumor diameter. 59 , 60 This confirms that improved performance partly reflects incorporation of a broader range of disease‐related features and clinical context, enabling more individualized risk estimation.

Highly ranked predictors also included indicators of clinical assessment (e.g., radiological procedures), comorbidity burden, and healthcare utilization. These likely function as proxies for disease complexity or diagnostic intensity rather than direct causal determinants of outcome. For example, frequent imaging may reflect closer surveillance in patients with more severe or uncertain disease rather than conferring benefit itself.

Importantly, MLAll does not require complete data for any single variable. Tree‐based algorithms can natively accommodate missing values, allowing predictions even when many features are partially unavailable. 29 This contrasts with conventional scores such as the NCCN IPI, which depend on complete data for each component. Although feature importance may generate hypotheses for further study, the model is intended for fixed‐time risk prediction and probabilistic risk stratification, rather than for causal inference or treatment selection.

Generalizability and implementation across health care systems

While MLAll demonstrates strong prognostic performance within the Danish registry setting, it should primarily be viewed as a proof of concept for how routinely collected health data can support individualized prognostication. Some predictors may reflect organization‐ or documentation‐specific factors. Applying this framework in other health systems would therefore require adaptation, including harmonizing variable definitions, adjusting preprocessing pipelines, and retraining or fine‐tuning the model on local data to account for regional and institutional differences. 23 , 41

As others have previously noted, attempts to build universally generalizable models often result in systems that perform only moderately well across many settings, rather than optimally within a single one. 41 Accordingly, MLAll is best understood as a transferable framework for registry‐based prognostic modeling that can be extended and adapted to other health systems where comparable structured data exist. 41 To support this, all modeling code and preprocessing pipelines are openly shared to enable replication, retraining, and integration into local environments.

Although the present work is not intended as a deployable clinical tool, it delineates a clear path toward that goal. While the full set of structured variables used here may not yet be uniformly available across healthcare systems, this study illustrates the potential performance gains achievable when such collected data are leveraged comprehensively.

Similar registry‐based machine‐learning approaches have already been integrated into EHR systems in other hematologic settings. 19 , 23 Together with the fact that the present framework operates entirely on routinely collected structured data and tolerates incomplete records through native handling of missingness, this suggests that implementation is feasible. 19 , 23 Future integration into EHR systems would require prospective validation, but the present work demonstrates a technically feasible and transferable framework for registry‐based prognostic modeling.

CONCLUSION

This study demonstrates that addressing key limitations of traditional prognostic models enables more accurate and individualized prediction of treatment failure in DLBCL. By integrating longitudinal, structured EHR data within a flexible modeling framework, MLAll outperformed the NCCN IPI in predicting treatment‐free survival among patients receiving CHOP‐like therapy. Continuous risk estimates enabled clinically adaptable trade‐offs between precision and recall, supporting context‐specific risk stratification. Moreover, comparable performance in DLBCL‐only and broader lymphoma cohorts indicates that shared prognostic signals can be leveraged through joint modeling. Pending prospective validation, this framework may support clinical trial enrichment, treatment stratification, and earlier identification of patients at increased risk of treatment failure.

AUTHOR CONTRIBUTIONS

Mikkel Werling: Conceptualization; funding acquisition; methodology; visualization; writing—original draft; writing—review and editing; formal analysis; software; data curation; project administration; investigation; validation. Alexander D. Fuglkjær: Conceptualization; writing—review and editing; methodology; validation. Peter Brown: Conceptualization; writing—review and editing; data curation. Carsten U. Niemann: Funding acquisition; conceptualization; writing—review and editing; supervision; data curation; resources; project administration; validation. Rudi Agius: Funding acquisition; conceptualization; supervision; project administration; writing—review and editing; methodology; validation.

CONFLICT OF INTEREST STATEMENT

Carsten Utoft Niemann received research funding and/or consultancy fees from AstraZeneca, Janssen, AbbVie, BeiGene, Genmab, CSL Behring, Octapharma, Takeda, Eli Lily, MSD, and Novo Nordisk Foundation. Peter Brown received research funding and/or consultancy fees from Roche, AbbVie, and BMS. All other authors declare no competing interests to disclose.

ETHICS STATEMENT

The data are available as a collaborator on the DALY‐CARE project. For guidance on how to become a collaborator, please contact one of the corresponding authors. We refer to the paper introducing the data resource for details concerning the data and how to access it. All scripts and feature engineering pipelines are publicly available on GitHub to enable replication and testing in external cohorts.

FUNDING

The project was funded partially by the Alfred Benzon foundation and by Sygeforsikringen “Danmark” (grant number 2022‐0086). This work was based on data analyzed at the national infrastructure for personal medicine hosted at the Danish National Genome Center, which is supported by the Novo Nordisk Foundation (grant agreement NNF18SA0035348 and grant agreement NNF19SA0035486). This work was supported by the Danish Data Science Academy, which is funded by the Novo Nordisk Foundation (NNF21SA0069429) and VILLUM FONDEN (40516). The PERSIMUNE project contributed data and received funding from the Danish National Research Foundation (#126).

Supporting information

Supplemental Information.

Contributor Information

Carsten U. Niemann, Email: carsten.utoft.niemann@regionh.dk.

Rudi Agius, Email: rudi.agius.01@regionh.dk.

DATA AVAILABILITY STATEMENT

The data used for this project are a subset of the Danish Lymphoid Cancer Research (DALY‐CARE) data resource. The data are available as a collaborator on the DALY‐CARE project. For guidance on how to become a collaborator, please contact one of the corresponding authors. The DALY‐CARE protocol has been approved by the Danish Health Data Authority and Danish National Ethics Committee (approvals P‐2020‐561 and 1804410, respectively).

REFERENCES

  • 1. Sukswai N, Lyapichev K, Khoury JD, Medeiros LJ. Diffuse large B‐cell lymphoma variants: an update. Pathology. 2020;52(1):53‐67. [DOI] [PubMed] [Google Scholar]
  • 2. Pfreundschuh M, Trümper L, Österborg A, et al. CHOP‐like chemotherapy plus rituximab versus CHOP‐like chemotherapy alone in young patients with good‐prognosis diffuse large‐B‐cell lymphoma: a randomised controlled trial by the MabThera International Trial (MInT) Group. Lancet Oncol. 2006;7(5):379‐391. [DOI] [PubMed] [Google Scholar]
  • 3. Coiffier B, Sarkozy C. Diffuse large B‐cell lymphoma: R‐CHOP failure—what to do? Hematology. 2016;2016(1):366‐378. 10.1182/asheducation-2016.1.366 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Xie S, Yu Z, Feng A, et al. Analysis and prediction of relative survival trends in patients with non‐Hodgkin lymphoma in the United States using a model‐based period analysis method. Front Oncol. 2022;12:942122. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Sant M, Minicozzi P, Mounier M, et al. Survival for haematological malignancies in Europe between 1997 and 2008 by region and age: results of EUROCARE‐5, a population‐based study. Lancet Oncol. 2014;15(9):931‐942. [DOI] [PubMed] [Google Scholar]
  • 6. Ritter Z, Papp L, Zámbó K, et al. Two‐year event‐free survival prediction in DLBCL patients based on in vivo radiomics and clinical parameters. Front Oncol. 2022;12:820136. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Harkins RA, Chang A, Patel SP, et al. Remaining challenges in predicting patient outcomes for diffuse large B‐cell lymphoma. Expert Rev Hematol. 2019;12(11):959‐973. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. International Non‐Hodgkin's Lymphoma Prognostic Factors Project . A predictive model for aggressive non‐Hodgkin's lymphoma. N Engl J Med. 1993;329(14):987‐994. 10.1056/NEJM199309303291402 [DOI] [PubMed] [Google Scholar]
  • 9. Zhou Z, Sehn LH, Rademaker AW, et al. An enhanced International Prognostic Index (NCCN‐IPI) for patients with diffuse large B‐cell lymphoma treated in the rituximab era. Blood. 2014;123(6):837‐842. 10.1182/blood-2013-09-524108 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Sehn LH, Berry B, Chhanabhai M, et al. The revised International Prognostic Index (R‐IPI) is a better predictor of outcome than the standard IPI for patients with diffuse large B‐cell lymphoma treated with R‐CHOP. Blood. 2007;109(5):1857‐1861. [DOI] [PubMed] [Google Scholar]
  • 11. Jelicic J, Larsen TS, Maksimovic M, Trajkovic G. Available prognostic models for risk stratification of diffuse large B cell lymphoma patients: a systematic review. Crit Rev Oncol Hematol. 2019;133:1‐16. 10.1016/j.critrevonc.2018.10.006 [DOI] [PubMed] [Google Scholar]
  • 12. Sehn LH, Salles G. Diffuse large B‐cell lymphoma. N Engl J Med. 2021;384(9):842‐858. 10.1056/NEJMra2027612 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Westin J, Sehn LH. CAR T cells as a second‐line therapy for large B‐cell lymphoma: a paradigm shift? Blood. 2022;139(18):2737‐2746. 10.1182/blood.2022015789 [DOI] [PubMed] [Google Scholar]
  • 14. Biccler J, Eloranta S, de Nully Brown P, et al. Simplicity at the cost of predictive accuracy in diffuse large B‐cell lymphoma: a critical assessment of the R‐IPI, IPI, and NCCN‐IPI. Cancer Med. 2018;7(1):114‐122. 10.1002/cam4.1271 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Liu X, Shen Y, Wang H, Ge Q, Fei A, Pan S. Prognostic significance of neutrophil‐to‐lymphocyte ratio in patients with sepsis: a prospective observational study. Mediators Inflamm. 2016;2016:(1) 8191254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Jelicic J, Juul‐Jensen K, Bukumiric Z, et al. Prognostic indices in diffuse large B‐cell lymphoma: a population‐based comparison and validation study of multiple models. Blood Cancer J. 2023;13(1):157. 10.1038/s41408-023-00930-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Brieghel C, Rotbain EC, Jelicic J, Thorsteinsdottir S, Brown PN, Niemann CU. One for the ages: impact of age on international prognostic indices in lymphoid cancer. Blood Adv. 2025;9(16):4260‐4264. 10.1182/bloodadvances.2025016534 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Maurer MJ, Jais JP, Ghesquières H, et al. Personalized risk prediction for event‐free survival at 24 months in patients with diffuse large B‐cell lymphoma. Am J Hematol. 2016;91(2):179‐184. 10.1002/ajh.24223 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Agius R, Brieghel C, Andersen MA, et al. Machine learning can identify newly diagnosed patients with CLL at high risk of infection. Nat Commun. 2020;11(1):363. 10.1038/s41467-019-14225-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Rask Kragh Jørgensen R, Bergström F, Eloranta S, et al. Machine learning–based survival prediction models for progression‐free and overall survival in advanced‐stage Hodgkin lymphoma. JCO Clin Cancer Inform. 2024;8:e2300255. 10.1200/CCI.23.00255 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Kim JY, Kahttana I, Yoon H, Chang S, Yoon SO. Exploring the potential of enhanced prognostic performance of NCCN‐IPI in diffuse large B‐cell lymphoma by integrating tumor microenvironment markers: stromal FOXC1 and tumor pERK1/2 expression. Cancer Med. 2024;13(19):e70305. 10.1002/cam4.70305 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Wang J, Sun Y, Lin M, et al. Risk stratification for diffuse large B‐cell lymphoma by integrating interim 18F‐FDG PET‐CT analysis and the NCCN‐IPI: a multicentre retrospective study. Hematol Oncol. 2025; 43(4):e70118. 10.1002/hon.70118 [DOI] [PubMed] [Google Scholar]
  • 23. Agius R, Riis‐Jensen AC, Wimmer B, et al. Deployment and validation of the CLL treatment infection model adjoined to an EHR system. npj Digit Med. 2024;7(1):147. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Bjerregaard‐Michelsen S, Poulsen LØ, Bjerrum A, Bøgsted M, Vesteghem C. Machine learning for prediction of 30‐day mortality in patients with advanced cancer: comparing pan‐cancer and single‐cancer models. ESMO Real World Data Digit Oncol. 2025;8:100146. 10.1016/j.esmorw.2025.100146 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Arboe B, Josefsson P, Jørgensen J, et al. Danish national lymphoma registry. Clin Epidemiol. 2016;8:577‐581. 10.2147/CLEP.S99470 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Brieghel C, Werling M, Frederiksen CM, et al. The Danish Lymphoid Cancer Research (DALY‐CARE) data resource: the basis for developing data‐driven hematology. medRxiv. Preprint posted online April 12, 2024:2024.04.11.24305663. 10.1101/2024.04.11.24305663 [DOI] [PMC free article] [PubMed]
  • 27. Arboe B, El‐Galaly TC, Clausen MR, et al. The Danish National Lymphoma Registry: coverage and data quality. PLoS One. 2016;11(6):e0157999. 10.1371/journal.pone.0157999 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Bernstorff M, Enevoldsen K, Damgaard J, Danielsen A, Hansen L. timeseriesflattener: a Python package for summarizing features from (medical) time series. J Open Source Softw. 2023;8(83):5197. 10.21105/joss.05197 [DOI] [Google Scholar]
  • 29. Chen T, He T. XGBoost | Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2015. Accessed December 13, 2024. https://dl.acm.org/doi/abs/10.1145/2939672.2939785
  • 30. Shwartz‐Ziv R, Armon A. Tabular data: deep learning is not all you need. Information Fusion. 2022;81:84‐90. 10.1016/j.inffus.2021.11.011 [DOI] [Google Scholar]
  • 31. Davison AC, Hinkley DV. Bootstrap Methods and Their Application. Cambridge University Press; 1997. [Google Scholar]
  • 32. Efron B, Tibshirani RJ. An Introduction to the Bootstrap. Chapman and Hall/CRC; 1994. 10.1201/9780429246593 [DOI] [Google Scholar]
  • 33. Kaplan EL, Meier P. Nonparametric estimation from incomplete observations. J Am Stat Assoc. 1958;53(282):457‐481. 10.1080/01621459.1958.10501452 [DOI] [Google Scholar]
  • 34. Lundberg S. A unified approach to interpreting model predictions. ArXiv Prepr ArXiv170507874. Published online 2017.
  • 35. Hollmann N, Müller S, Purucker L, et al. Accurate predictions on small data with a tabular foundation model. Nature. 2025;637(8045):319‐326. 10.1038/s41586-024-08328-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Ruppert AS, Dixon JG, Salles G, et al. International prognostic indices in diffuse large B‐cell lymphoma: a comparison of IPI, R‐IPI, and NCCN‐IPI. Blood. 2020;135(23):2041‐2048. 10.1182/blood.2019002729 [DOI] [PubMed] [Google Scholar]
  • 37. Schmitz N, Zeynalova S, Nickelsen M, et al. CNS international prognostic index: a risk model for CNS relapse in patients with diffuse large B‐cell lymphoma treated with R‐CHOP. J Clin Oncol. 2016;34(26):3150‐3156. [DOI] [PubMed] [Google Scholar]
  • 38. Sherstinsky A. Fundamentals of Recurrent Neural Network (RNN) and Long Short‐Term Memory (LSTM) network. Physica D: Nonlinear Phenomena. 2020;404:132306. 10.1016/j.physd.2019.132306 [DOI] [Google Scholar]
  • 39. Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems. 30. Curran Associates, Inc.; 2017. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html. [Google Scholar]
  • 40. Odgaard M, Klein KV, Thysen SM, Jimenez‐Solem E, Sillesen M, Nielsen M. CORE‐BEHRT: a carefully optimized and rigorously evaluated BEHRT. arXiv. Preprint posted online October 11, 2024. 10.48550/arXiv.2404.15201 [DOI]
  • 41. Futoma J, Simons M, Panch T, Doshi‐Velez F, Celi LA. The myth of generalisability in clinical research and machine learning in health care. Lancet Digit Health. 2020;2(9):e489‐e492. 10.1016/S2589-7500(20)30186-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Klanova M, Sehn LH, Bence‐Bruckler I, et al. Integration of cell of origin into the clinical CNS International Prognostic Index improves CNS relapse prediction in DLBCL. Blood. 2019;133(9):919‐926. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Runge HF, Lacy S, Barrans S, et al. Application of the LymphGen classification tool to 928 clinically and genetically‐characterised cases of diffuse large B cell lymphoma (DLBCL); 2021. [DOI] [PubMed]
  • 44. Chapuy B, Wood T, Stewart C, et al. DLBclass: a probabilistic molecular classifier to guide clinical investigation and practice in diffuse large B‐cell lymphoma. Blood. 2025;145(18):2041‐2055. [DOI] [PubMed] [Google Scholar]
  • 45. Parviz M, Brieghel C, Agius R, Niemann CU. Prediction of clinical outcome in CLL based on recurrent gene mutations, CLL‐IPI variables, and (para) clinical data. Blood Adv. 2022;6(12):3716‐3728. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Agius R, Parviz M, Niemann CU. Artificial intelligence models in chronic lymphocytic leukemia–recommendations toward state‐of‐the‐art. Leuk Lymphoma. 2022;63(2):265‐278. [DOI] [PubMed] [Google Scholar]
  • 47. Zhang XS, Tang F, Dodge HH, Zhou J, Wang F MetaPred: Meta‐Learning for Clinical Risk Prediction with Limited Patient Electronic Health Records. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. KDD’19. Association for Computing Machinery; 2019:2487‐2495. 10.1145/3292500.3330779 [DOI] [PMC free article] [PubMed]
  • 48. Rafiei A, Moore R, Jahromi S, Hajati F, Kamaleswaran R. Meta‐learning in healthcare: a survey. SN Comput Sci. 2024;5(6):791. 10.1007/s42979-024-03166-9 [DOI] [Google Scholar]
  • 49. Trivedi M, Adhikari PK, Jain S, Deshwal V, et al. An empirical study of meta learning for medical image segmentation with transfer learning. In: 2024 2nd International Conference on Disruptive Technologies (ICDT). IEEE; 2024:857‐862.
  • 50. Hossain MS, Shorfuzzaman M. MetaParkinson: a cyber‐physical deep meta‐learning framework for n‐shot diagnosis and monitoring of Parkinson's patients. IEEE Syst J. 2023;17(4):5251‐5260. 10.1109/JSYST.2023.3308333 [DOI] [Google Scholar]
  • 51. Liu L, Liu Z, Wu H, et al. Multi‐task learning via adaptation to similar tasks for mortality prediction of diverse rare diseases. In: Amia Annual Symposium Proceedings. Vol 2020. 2021:763. [PMC free article] [PubMed]
  • 52. Li X, Yu L, Jin Y, Fu CW, Xing L, Heng PA. Difficulty‐aware meta‐learning for rare disease diagnosis. In: Martel AL, Abolmaesumi P, Stoyanov D, et al., eds. Medical Image Computing and Computer Assisted Intervention – MICCAI 2020. Springer International Publishing; 2020:357‐366. 10.1007/978-3-030-59710-8_35 [DOI] [Google Scholar]
  • 53. Banerjee J, Taroni JN, Allaway RJ, Prasad DV, Guinney J, Greene C. Machine learning in rare disease. Nat Methods. 2023;20(6):803‐814. 10.1038/s41592-023-01886-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Lee S, Yin C, Zhang P. Stable clinical risk prediction against distribution shift in electronic health records. Patterns. 2023;4(9):100828. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Li M, Xia H, Zheng H, et al. Red blood cell distribution width and platelet counts are independent prognostic factors and improve the predictive ability of IPI score in diffuse large B‐cell lymphoma patients. BMC Cancer. 2019;19(1):1084. 10.1186/s12885-019-6281-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Yasmeen T, Ali J, Khan K, Siddiqui N. Frequency and causes of anemia in lymphoma patients. Pak J Med Sci. 2019;35(1):61‐65. 10.12669/pjms.35.1.91 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Hong J, Woo HS, Kim H, et al. Anemia as a useful biomarker in patients with diffuse large B‐cell lymphoma treated with R‐CHOP immunochemotherapy. Cancer Sci. 2014;105(12):1569‐1575. 10.1111/cas.12544 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Matsumoto K, Fujisawa S, Ando T, et al. Anemia associated with worse outcome in diffuse large B‐cell lymphoma patients: a single‐center retrospective study. Turk J Haematol. 2018;35(3):181‐184. 10.4274/tjh.2017.0437 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Tout M, Casasnovas O, Meignan M, et al. Rituximab exposure is influenced by baseline metabolic tumor volume and predicts outcome of DLBCL patients: a Lymphoma Study Association report. Blood. 2017;129(19):2616‐2623. 10.1182/blood-2016-10-744292 [DOI] [PubMed] [Google Scholar]
  • 60. Sasanelli M, Meignan M, Haioun C, et al. Pretherapy metabolic tumour volume is an independent predictor of outcome in patients with diffuse large B‐cell lymphoma. Eur J Nucl Med Mol Imaging. 2014;41(11):2017‐2022. 10.1007/s00259-014-2822-7 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Information.

Data Availability Statement

The data used for this project are a subset of the Danish Lymphoid Cancer Research (DALY‐CARE) data resource. The data are available as a collaborator on the DALY‐CARE project. For guidance on how to become a collaborator, please contact one of the corresponding authors. The DALY‐CARE protocol has been approved by the Danish Health Data Authority and Danish National Ethics Committee (approvals P‐2020‐561 and 1804410, respectively).


Articles from HemaSphere are provided here courtesy of Wiley

RESOURCES