Skip to main content
Springer logoLink to Springer
editorial
. 2026 Aug 19;52(10):2198–2200. doi: 10.1007/s00134-026-08568-2

Machine learning for prediction in secondary hemophagocytic lymphohistiocytosis: real progress, real limits, and the potential of synthetic data

Louis Delamarre 1,2,✉, Jan-Inge Henter 3,4,✉
PMCID: PMC13624052  PMID: 42616081

Secondary hemophagocytic lymphohistiocytosis (sHLH) is a rare syndrome associated with high mortality, which is difficult to diagnose and prognosticate, and challenging to investigate due to its rarity and heterogeneity [1–4]. The predictive modeling presented by Ruzicka et al. in a recent issue of Intensive Care Medicine represents a serious and methodologically rigorous attempt to address the prognostic gap in this condition using machine learning (ML) [5]. The strengths and limitations of this work illuminate a broader problem faced by the critical care community: when the disease is rare enough, the data itself becomes the bottleneck.

What the study achieves

Ruzicka et al. assembled 167 adult sHLH patients, spanning 15 years, from 6 centers across three European countries in a valuable multicentric effort. The models predict both Initial Disease Severity (IDS), defined as intensive care unit (ICU) admission or early death, and time-point specific mortality at 30, 60, 90, 180, and 365 days, achieving strong discriminatory performance (area under the curve (AUC) 0.845 for IDS; mean AUC 0.882 for mortality), calibration, and interpretable feature importance. This could broaden the prognostic options available in the field of sHLH, complementing existing diagnostic (HLH-94/HLH-2024 [6–9]; HScore [10]) and single-time-point mortality prediction tools [11].

The authors’ methodological choices could prove useful beyond the walls of ICUs to improve shared decision-making. First, the composite IDS endpoint, encapsulating ICU admission and death < 90 days without ICU admission, addresses a clinically pertinent question at the bedside: is a patient on a trajectory of severe disease requiring immediate escalation of care, whether that escalation takes the form of ICU transfer, initiation of intensive immunosuppression, or early referral to a tertiary center. In this sense, IDS is presented as an early severity signal. Nevertheless, the authors emphasize that IDS is a “marker of an adverse clinical course rather than as a direct ICU triage tool” and that “model performance for the composite endpoint should not be extrapolated to the individual endpoint components”. In a disease where the time-to-treatment is paramount, such a tool is valuable only if it guides an action. Defining such actions is one important topic for future research. Second, the authors chose to use only objective laboratory variables, deliberately excluding clinician-interpreted variables. The omission of the underlying sHLH trigger, for example, could have led to a loss of valuable information, but is well justified by the frequent uncertainty regarding its identification at the time of diagnosis.

Importantly, sIL-2R emerged as a key predictor of both IDS and mortality but was available in only 40% of the sample. The prominence of sIL-2R in feature importance analyses should be interpreted alongside the sensitivity analyses, which show only a modest AUC reduction when sIL-2R is excluded. This apparent tension reflects a property of ensemble models trained on correlated clinical variables: a single biomarker may be the most distinctive contributor to the model’s decision process while remaining partially redundant with its covariates. Thus, the HLH-Risk-Calculator retains meaningful predictive performance in centers where sIL-2R is unavailable, while sIL-2R adds incremental discriminative value where it can be measured. The high missingness of sIL-2R in the training cohort means that both its feature importance and the magnitude of its marginal contribution to the model could be underestimated. There is, however, a risk of bias in the study in that sIL-2R availability might correlate with disease severity. The true predictive gain from systematically available, high-resolution sIL-2R measurements remains to be established in future prospective cohorts.

Where the bottleneck lies and a potential path forward

After excluding patients with excessive missingness, 152 patients remained. The holdout test set for mortality contained 43 patients, and for IDS, 32 patients. Model hyperparameters were explicitly constrained to prevent overfitting. In sHLH, as in most rare critical illness syndromes, the scarcity of high-quality patient data is the real bottleneck.

Generative synthetic data augmentation is a methodological paradigm gaining traction in rare disease research. Its principle is to train a deep generative model on large amounts of representative real patient data, learn the underlying multidimensional correlations between the characteristics of the clinical population, and generate additional synthetic patients that reproduce the distributional properties of the real cohort, including non-linear correlations, without duplicating individual patients [12–14]. Applied to a cohort like that of Ruzicka et al., such an approach could serve two distinct purposes. First, training set augmentation: generating several hundred synthetic patients to supplement the real training set, allowing random forest models to explore deeper feature interactions without the aggressive regularization constraints currently required. Second, class rebalancing: in a six-center dataset where contribution sizes are unequal, synthetic generation conditioned on smaller-center patient characteristics could homogenize center contributions, addressing the site-specific heterogeneity identified by the authors. Both fidelity (statistical similarity to real patients) and utility (performance of models trained on synthetic-augmented data, tested on real patients) of synthetic data are required, justifying rigorous evaluation based on statistical metrics and human expert judgment. A complementary approach would be to leverage new statistical ideas that enable valid statistical inference even when some of the data originates from a potentially inaccurate ML model [15].

sHLH is one instance of a broader category of rare critical illness syndromes with slow accumulation of clinical data despite multicentric efforts. Ruzicka et al. have possibly built the most comprehensive prognostic tool for sHLH to date. Biologically plausible synthetic data augmentation may represent the next methodological frontier for prognostic modeling in rare critical illness syndromes.

Clinical relevance for sHLH

The authors are to be congratulated on developing this HLH-Risk-Calculator. External validation in independent cohorts remains essential to assess the models’ real-world applicability, and the authors’ decision to make the HLH-Risk-Calculator publicly available for research purposes facilitates this important validation. Ultimately, if external validation supports the reported findings, clinicians will have a valuable tool to predict Initial Disease Severity (IDS) and time-point-specific mortality in sHLH, which may enable both better guidance of clinical decisions and prognosis-based stratification of patients for treatment trials.

Data availability statement

Not applicable.

Declarations

Conflicts of interest

The authors declare that they have no conflict of interest.

Footnotes

This editorial refers to the article available online at https://doi.org/10.1007/s00134-026-08515-1

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Louis Delamarre, Email: louis.delamarre.pro@gmail.com.

Jan-Inge Henter, Email: jan-inge.henter@ki.se.

References

  • 1.Henter JI (2025) Hemophagocytic lymphohistiocytosis. N Engl J Med 392:584–598. 10.1056/NEJMra2314005 [DOI] [PubMed] [Google Scholar]
  • 2.Wimmer T, Mattes R, Stemmler HJ, Hauck F, Schulze-Koops H, Stecher SS et al (2023) sCD25 as an independent adverse prognostic factor in adult patients with HLH: results of a multicenter retrospective study. Blood Adv 7:832–844. 10.1182/bloodadvances.2022007953 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Buyse S, Teixeira L, Galicier L, Mariotte E, Lemiale V, Seguin A et al (2010) Critical care management of patients with hemophagocytic lymphohistiocytosis. Intensive Care Med 36:1695–1702. 10.1007/s00134-010-1936-z [DOI] [PubMed] [Google Scholar]
  • 4.Créput C, Galicier L, Buyse S, Azoulay E (2008) Understanding organ dysfunction in hemophagocytic lymphohistiocytosis. Intensive Care Med 34:1177–1187. 10.1007/s00134-008-1111-y [DOI] [PubMed] [Google Scholar]
  • 5.Ruzicka M, Stubbe HC, Fauser J, Trebo M, Wimmer T, Horvath L et al (2026) The HLH-Risk-Calculator is a machine learning-based tool to predict course & mortality of secondary hemophagocytic lymphohistiocytosis. Intensive Care Med 52:1514–1525. 10.1007/s00134-026-08515-1 [DOI] [PMC free article] [PubMed]
  • 6.Henter JI, Horne A, Aricó M, Egeler RM, Filipovich AH, Imashuku S et al (2007) HLH-2004: diagnostic and therapeutic guidelines for hemophagocytic lymphohistiocytosis. Pediatr Blood Cancer 48:124–131. 10.1002/pbc.21039 [DOI] [PubMed] [Google Scholar]
  • 7.Henter JI, Sieni E, Eriksson J, Bergsten E, Hed Myrberg I, Canna SW et al (2024) Diagnostic guidelines for familial hemophagocytic lymphohistiocytosis revisited. Blood 144:2308–2318. 10.1182/blood.2024025077 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lachmann G, Heeren P, Schuster FS, Nyvlt P, Spies C, Feinkohl I et al (2025) Multicenter validation of secondary hemophagocytic lymphohistiocytosis diagnostic criteria. J Intern Med 297:312–327. 10.1111/joim.20065 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Henter JI (2025) Intercontinentally validated diagnostic criteria for secondary hemophagocytic lymphohistiocytosis - so welcome! J Intern Med 297:240–243. 10.1111/joim.20066 [DOI] [PubMed] [Google Scholar]
  • 10.Fardet L, Galicier L, Lambotte O, Marzac C, Aumont C, Chahwan D et al (2014) Development and validation of the HScore, a score for the diagnosis of reactive hemophagocytic syndrome. Arthritis Rheumatol 66:2613–2620. 10.1002/art.38690 [DOI] [PubMed] [Google Scholar]
  • 11.Zhu J, Cao N, Wu F, Ding Y, Jiao X, Wang J et al (2025) Predicting 30-day mortality in hemophagocytic lymphohistiocytosis: clinical features, biochemical parameters, and machine learning insights. Ann Hematol 104:2239–2264. 10.1007/s00277-025-06249-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Chadebec C, Thibeau-Sutre E, Burgos N, Allassonnière S (2023) Data augmentation in high dimensional low sample size setting using a geometry-based variational autoencoder. IEEE Trans Pattern Anal Mach Intell 45:2879–2896. 10.1109/TPAMI.2022.3185773 [DOI] [PubMed] [Google Scholar]
  • 13.Chadebec C, Allassonnière S (2021) Data augmentation with variational autoencoders and manifold sampling. arXiv. 10.48550/arXiv.2103.13751 [DOI]
  • 14.Ferré F, Allassonnière S, Chadebec C, Minville V (2025) Generating artificial patients with reliable clinical characteristics using a geometry-based variational autoencoder: proof-of-concept feasibility study. J Med Internet Res 27:e63130. 10.2196/63130 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Angelopoulos AN, Bates S, Fannjiang C, Jordan MI, Zrnic T (2023) Prediction-powered inference. Science 382:669–674. 10.1126/science.adi6000 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Not applicable.


Articles from Intensive Care Medicine are provided here courtesy of Springer

RESOURCES