Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 18.
Published in final edited form as: JAMA. 2024 Apr 9;331(14):1225–1226. doi: 10.1001/jama.2024.3376

Using Effect Scores to Characterize Heterogeneity of Treatment Effects

Guanbo Wang 1, Patrick J Heagerty 2,3, Issa J Dahabreh 1,3,4
PMCID: PMC13479313  NIHMSID: NIHMS2196727  PMID: 38501213

It is common for treatments to yield different outcomes in different patients. If patient characteristics that predict treatment response can be identified, understanding this heterogeneity of treatment effect (HTE) should inform individual treatment choices.1,2 Effect score analyses are an approach for evaluating HTE, using a model developed to predict treatment effect given a patient’s characteristics, and then examining differences in treatment effects across groups of patients with different effect scores.35 In this issue of JAMA, Buell et al6 report the results of an effect score analysis using data from 2 randomized clinical trials to evaluate HTE for treatment strategies targeting different levels of oxygen saturation in critically ill patients receiving mechanical ventilation. The authors found that patients in different strata defined by the effect score had different treatment effects. They also noted that the stratum showing benefit from a lower target had a higher proportion of patients with acute brain injury, and the stratum showing benefit from a higher target had a higher proportion of patients with sepsis and abnormally elevated vital signs.3

What Is Effect Score Analysis?

The effect score is a prediction of each patient’s treatment effect, given their individual characteristics. The analysis then stratifies or groups patients in a clinical trial based on their effect scores.3,4 The effect score model can be developed internally, based on the same dataset in which HTE will be evaluated, or externally, using data from another study. When the effect score model is developed internally, data from a single trial are usually randomly split into a model development dataset and a treatment effect estimation and HTE evaluation dataset; when the model is developed externally, one dataset is used for model development and another for evaluation.

Effect score analysis proceeds in 4 steps (Figure). First, the development dataset is used to create an effect score model, a statistical model for the treatment effect based on patient characteristics. Second, the resulting model is used to obtain an effect score value for each patient in the evaluation dataset. Third, the patients in the evaluation dataset are sorted by their effect scores and grouped into a few strata, typically 3 to 5. Fourth, in each stratum of the evaluation dataset, the treatment effect and corresponding confidence interval are obtained and heterogeneity across strata is assessed, for example, using a test for interaction.

Figure.

Figure

Using Effect Scores to Assess Heterogeneity of Treatment Effect

The success of effect score analyses depends on how closely the effect score model captures the relationship between patient characteristics and treatment effects. The unknown “true” relationship can be approximated using flexible modeling methods such as machine learning approaches.7,8 The effect score attempts to capture the relationship between multiple patient characteristics and the treatment effect in a single number.

The use of flexible modeling methods requires special care to avoid overfitting—creating a model that reflects random idiosyncrasies in the available data rather than the true relationship between patient characteristics and treatment effects—by making sure the development dataset is kept separate from the evaluation dataset. This separation is important when the effect score model is developed internally, and it must be preserved when it is developed externally (eg, by avoiding using any information from the evaluation dataset during model development).

Why Is Effect Score Analysis Used?

Effect score analysis is used to quantify treatment effects within each patient stratum of the effect score and the extent of heterogeneity across strata. By using separate evaluation data to quantify treatment effects and examine heterogeneity, valid confidence intervals and P values can be obtained. Importantly, this is true even when using flexible modeling methods, and even if the effect score model is not exactly correct, as is usually the case when the relationship between patient characteristics and the treatment effect is highly complex.

Patients and clinicians may use information from effect score analyses to decide if one treatment is likely to be more beneficial (or harmful) vs another for any specific patient by examining the treatment effect in the effect score stratum corresponding to the patient’s characteristics. Furthermore, examining the characteristics of patients in each effect score stratum can generate future research questions about the underlying sources of heterogeneity.

Effect score analyses address some of the limitations of alternative approaches for examining HTE, such as subgroup analyses that stratify patients by one characteristic at a time (eg, sex, age), examinations of heterogeneity that stratify patients by baseline risk of a poor outcome, and machine learning methods that directly estimate heterogeneous effects. One-characteristic-at-a-time subgroup analyses do not integrate information from multiple patient characteristics and thus produce results that are less clinically informative; they also can result in excessive false-negative (due to low precision) and false-positive (from multiple testing) results.2 Analyses examining HTE by baseline risk can address some of the limitations of one-characteristic-at-a-time subgroup analyses because they consider multiple characteristics jointly; however, baseline risk does not fully determine treatment effect.9 Using machine learning methods to directly estimate heterogeneous effects considers multiple variables jointly and appropriately focuses on the treatment effect instead of risk, but different algorithms produce variable results that may not be precise enough for clinical use.

Limitations of the Method

Developing models for effect score analyses can be challenging when the “true” underlying relationship between patient characteristics and treatment effect is complex. Even flexible modeling methods may fail to accurately capture this relationship, limiting the ability of effect score analyses to uncover HTE. Similarly, examining HTE over a relatively small number of effect score strata may fail to fully reveal complex HTE patterns.10 Effect score analysis, like other approaches for detecting HTE, requires large sample sizes to allow for proper development and evaluation of the effect score model. When the model is internally developed, sample splitting requires a large sample size to produce stable results and thus may only be appropriate for relatively large clinical trials or pooled trial analyses. When the model is externally developed, the sources of the development and evaluation datasets need to be sufficiently similar. Finally, clinical interpretation of the effect score model is not straightforward, particularly when using flexible modeling methods; additional steps are often needed to understand how the model uses patient characteristics to predict treatment effects.

How Was the Method Used in This Study?

Two previously conducted trials found that lower or higher oxygen-saturation targets resulted in similar mortality among critically ill patients requiring mechanical ventilation. Buell et al6 developed an effect score model from the first trial’s data by applying 6 different machine learning methods and choosing the best-performing one. They applied this effect score model to patients in the second trial and defined 3 effect score strata in which they compared outcomes from high vs low oxygen-saturation target strategies. They also summarized the characteristics of patients in each stratum, tested for the presence of HTE across strata, and estimated the average outcome that would have been achieved by allocating treatment guided by the effect score rather than randomly.

How Should the Method Be Interpreted?

Buell et al6 reported that the treatment benefit associated with the higher oxygen saturation target, measured as a risk difference for 28-day mortality within the high, medium, and low effect score patient stratum in the evaluation dataset, was 13% (95% CI, 3.5% to 22.6%), 0.2% (95% CI, −9.6% to 10.0%), and −6.1% (95% CI, −16.5% to 4.3%), respectively. These results suggest that the high effect score stratum would benefit most from a higher target oxygen saturation, and the low effect score stratum would benefit most from a lower target oxygen saturation. The authors also tested the null hypothesis of no heterogeneity across all 3 strata and obtained a P value of .02, suggesting the presence of true heterogeneity. By examining the patient characteristics of each stratum, they found that the high effect score stratum had a higher proportion of patients with sepsis and abnormally elevated vital signs and the low effect score stratum had a higher proportion of patients with acute brain injury; however, these clinical differences should not be used instead of the complete effect score for treatment selection, because diagnosis alone may be inadequately predictive of treatment effect.

Funding/Support:

This work was supported in part by Patient-Centered Outcomes Research Institute (PCORI) award ME-2019C3-17875.

Role of the Funder/Sponsor:

PCORI had no role in the preparation, review, or approval of the manuscript or decision to submit it for publication.

Footnotes

Conflict of Interest Disclosures: None reported.

Disclaimer: The content of this article is solely the responsibility of the authors and does not necessarily represent the official views of PCORI, PCORI’s Board of Governors, or the PCORI Methodology Committee.

References

  • 1.Angus DC, Chang CH. Heterogeneity of Treatment Effect: Estimating How the Effects of Interventions Vary Across Individuals. JAMA. 2021. Dec 14;326(22):2312–2313. doi: 10.1001/jama.2021.20552. [DOI] [PubMed] [Google Scholar]
  • 2.Dahabreh IJ, Hayward R, Kent DM. Using group data to treat individuals: understanding heterogeneous treatment effects in the age of precision medicine and patient-centred evidence. Int J Epidemiol. 2016. Dec 1;45(6):2184–2193. doi: 10.1093/ije/dyw125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Chernozhukov V, Demirer M, Duflo E, Fernandez-Val I. Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, With an Application to Immunization in India. National Bureau of Economic Research; 2018. doi: 10.3386/w24678. [DOI] [Google Scholar]
  • 4.Imai K, Li M. Statistical inference for heterogeneous treatment effects discovered by generic machine learning in randomized experiments. arXiv. Preprint posted March 28, 2022. [Google Scholar]
  • 5.Dahabreh IJ, Kazi DS. Toward Personalizing Care: Assessing Heterogeneity of Treatment Effects in Randomized Trials. JAMA. 2023. Apr 4;329(13):1063–1065. doi: 10.1001/jama.2023.3576. [DOI] [PubMed] [Google Scholar]
  • 6.Buell K, Spicer AB, Casey JD, et al. Individualized treatment effects of oxygen targets in mechanically ventilated critically ill adults. JAMA. Published online March 19, 2024. doi: 10.1001/jama.2024.2933. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Künzel SR, Sekhon JS, Bickel PJ, Yu B. Metalearners for estimating heterogeneous treatment effects using machine learning. Proc Natl Acad Sci U S A. 2019. Mar 5;116(10):4156–4165. doi: 10.1073/pnas.1804597116. Epub 2019 Feb 15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wager S, Athey S. Estimation and inference of heterogeneous treatment effects using random forests. J Am Stat Assoc. 2018;113(523):1228–1242. [Google Scholar]
  • 9.VanderWeele TJ, Luedtke AR, van der Laan MJ, Kessler RC. Selecting Optimal Subgroups for Treatment Using Many Covariates. Epidemiology. 2019. May;30(3):334–341. doi: 10.1097/EDE.0000000000000991. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Skrivankova V, Heagerty PJ. Single index methods for evaluation of marker-guided treatment rules based on multivariate marker panels. Biometrics. 2018. Jun;74(2):663–672. doi: 10.1111/biom.12752. Epub 2017 Aug 7. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES