Skip to main content
Clinical Pharmacology and Therapeutics logoLink to Clinical Pharmacology and Therapeutics
. 2025 Mar 6;117(5):1209–1216. doi: 10.1002/cpt.3627

Characterizing Treatment Effect Heterogeneity Using Real‐World Data

Haedi Thelen 1,, Sean Hennessy 1
PMCID: PMC11993293  NIHMSID: NIHMS2063235  PMID: 40047339

Abstract

Characterizing heterogeneity of treatment effects (HTE) is a fundamental goal of pharmacoepidemiology, addressing why medications work differently across patient populations. This paper reviews state‐of‐the‐art methods for studying HTE using real‐world data (RWD), which offer larger study sizes and more diverse patient populations compared to randomized clinical trials. The paper first defines HTE and discusses its measurement. It then examines three leading approaches to studying HTE: subgroup analysis, disease risk score (DRS) methods, and effect modeling methods. Subgroup analyses offer simplicity, transparency, and provide insights into drug mechanisms. However, they face difficulties in resolving which subgroup or combination of characteristics should be the basis for clinical decision making when multiple effect modifiers are present. DRS methods address some of these limitations by incorporating multiple patient characteristics into a summary score of outcome risk but may obscure insights into mechanisms. Effect modeling methods directly predict individual treatment effects, offering potential for precise HTE characterization, but are prone to model misspecification and may not provide mechanistic insights. The methods each have tradeoffs. Subgroup analysis is straightforward but can lead to spurious associations and does not account for multiple characteristics at once. DRS methods are relatively simple to implement and clinically useful, but may not completely describe HTE or provide mechanistic insight. Effect modeling approaches have great potential for characterizing HTE but are still being developed. Understanding HTE is essential for personalizing treatment strategies to improve patient outcomes. Researchers must weigh the strengths and limitations of each approach when using RWD to study HTE.


Why do medications seem to work well for some patients and not for others? Who is at risk for harm with a given medication? Will my patient benefit from this medication? Can mechanistic insights be gained by examining patient factors associated with drug effects? The study of the heterogeneity of treatment effects (HTE), or how the effects of medications vary across different people and treatment contexts, aims to answer these questions. Pharmacoepidemiology is the study of the use and effects of medications in populations, and the application of the resulting knowledge to improve health. It seeks to provide evidence that allows patients and clinicians to maximize the benefit and minimize the harm caused by medications, and inform mechanistic understanding. Characterizing HTE is a fundamental goal of pharmacoepidemiology. This paper reviews state‐of‐the‐art pharmacoepidemiologic methods to study HTE using real‐world data (RWD), defined as data arising outside of conventional pre‐approval trials.

While most of the literature on characterizing HTE focuses on the randomized controlled trial (RCT) setting, there is a growing body of research that expands HTE methods to the non‐randomized, observational setting. In this paper, we first discuss how HTE is defined and measured, then examine why RWD is useful for studying HTE, followed by a discussion of conventional and emerging methods to study HTE.

DEFINING AND MEASURING HTE

In clinical research, groups of individuals are used to study treatment effects. In RCTs and cohort studies, researchers compare the frequency of the outcome of interest in the treated vs. the control group, usually consisting of patients receiving either an inactive placebo or an active comparator. The difference or ratio in frequency between the two study groups is a measure of the overall or average treatment effect (ATE). Specifically, the risk difference (RD) is calculated as the risk (cumulative incidence) in the treated group minus the risk in the comparator group, while the risk ratio (RR) is calculated as the risk in the treated group divided by the risk in the control group. The ATE is often the principal result reported in RCTs and pharmacoepidemiology studies. While informative, the ATE obscures the possibility that some subjects may benefit greatly from treatment, while others may be harmed. For example, a null ATE could arise when harmful effects in one subgroup cancel out beneficial effects in another.

In another scenario, a relatively small observed average benefit could be driven by a large treatment effect in a potentially identifiable subgroup, while the majority of patients experience no effect on the target endpoint. This phenomenon is not uncommon, especially with the risk of adverse events. 1 , 2 When the risk of an adverse event is large in only a subgroup of patients (e.g., those with a particular pharmacogenetic allele), it is critical to identify that subgroup to inform the safe use of the drug and to contribute to our understanding of the mechanisms that lead to harm. When the potential benefit is restricted to a subgroup of patients, the ability to identify and treat only those patients will avoid unnecessary treatment in those who will not benefit, reducing the associated risk of adverse events and cost.

HTE evaluates how a treatment effect changes across levels of a characteristic, known as an effect modifier. 3 , 4 When the estimates of treatment effect are different at different levels of a characteristic, we say there is effect modification by that characteristic. 4 Importantly, this potential effect modifier must be a characteristic that (i) is not a second exposure, such as another intervention, and (ii) is not the result of or affected by treatment, since adjustment for variables that were affected by treatment can introduce bias. 5 If the characteristic is actually a second exposure, such as a second medication, then researchers are instead interested in studying the effects of joint exposures, which requires additional steps and assumptions. 4 , 6 Rather, a potential effect modifier is a characteristic that can be measured at baseline, prior to the initiation of study drug and the start of follow‐up. Common effect modifiers include age, sex, race, genotype, comorbid conditions, or other risk factors for the outcome of interest. If the characteristic is something that can be changed (e.g., randomized to), it is typically a second exposure, and the research question and analysis should be adjusted to study the effects of joint exposures.

Quantifying effect modification depends on the scale on which treatment effects are measured, a phenomenon known as scale dependence. Scale dependence is the possibility that treatment effects can be constant across levels of the effect modifier on one scale (e.g., the difference or additive scale) but vary on another scale (e.g., the ratio or multiplicative scale). For example, consider a hypothetical study of treatment on a yes/no outcome where a single binary characteristic is being evaluated for effect modification of the treatment‐outcome relationship. Table 1 displays two scenarios illustrating the scale dependence of effect modification. In scenario two, while there is no variation in measures of treatment effect on the RR scale (i.e., no HTE on the multiplicative scale), but there is variation in treatment effect on the RD (i.e., additive) scale.

Table 1.

Scale dependance of effect modification

Scenario Treated Comparator Risk difference Risk ratio
1 Drug effect constant on the difference scale, effect modification on ratio scale
Characteristic present 0.40 0.30 0.10 1.33
Characteristic absent 0.50 0.40 0.10 1.25
2 Drug effect constant on the ratio scale, effect modification on difference scale
Characteristic present 0.40 0.32 0.08 1.25
Characteristic absent 0.50 0.40 0.10 1.25

Measuring the effect modification on the additive scale is done by estimating the difference in RDs across strata of the covariate:

RDStrata1RDStrata2.

The measure of effect modification on the multiplicative scale is the ratio of RRs across strata of the effect modifier:

RRStrata1RRStrata2.

For Scenario 2 in Table 1 , the measure of effect modification on the additive scale is

0.080.10=0.02

While the measure of effect modification on the multiplicative scale is

1.251.25=1.

There is wide consensus that the risk or rate difference scale is most informative for clinical decision making and public health interventions. 7 , 8 This is because results on the RD scale can estimate the number of people who would benefit or be harmed from treatment. However, ratio scales are more common because of convenience: statistical modeling software often reports results on the ratio scales, and additional steps are needed to translate results to the difference scale. 8 There can also be HTE on the odds ratio scale, which is described elsewhere, 8 as well as other derived measures of effect modification including the relative excess risk due to interaction and the proportion attributable to the interaction. 8

Best practices for reporting measures of effect modification are to (i) identify the scale on which effect modification was assessed, (ii) report the frequency of outcome in each level of the effect modifier with and without treatment so readers can calculate the measures of effect on either the difference or ratio scales, and (iii) report the measures of effect modification (difference in RD across strata and ratio of RRs across strata) with a common reference group when there are more than two levels of the effect modifier. 7 , 9 , 10 , 11 , 12

WHY USE RWD TO STUDY HTE?

The study of HTE allows us to customize treatment strategies to individual patients and may contribute to understanding of drug mechanisms. For example, observing that patients with genotypes suggestive of poor CYP2C9 metabolism experience higher rates of adverse effects of celecoxib provides epidemiologic confirmation of expectations based on mechanistic evidence. 13 Further, when a medication is first approved by regulatory authorities, clinicians have little idea about how it might affect any specific patient. This is because RCTs, which typically inform regulatory approval, are designed to detect average treatment effects in controlled environments within highly selected patient populations. Consequently, they often have insufficient precision to estimate subgroup‐specific treatment effects or detect rare and late‐onset safety outcomes. 1 , 3 , 14

RWD is a critical resource to study HTE. Compared to data from trials, RWD tends to have larger study sizes, drawing from highly variable clinical settings with real patients in diverse contexts. This allows researchers to generate more statistically precise estimates of subgroup‐specific treatment effects, improving our understanding of variability in treatment response. This is one reason why regulatory agencies sometimes require post‐marketing studies using RWD of newly approved medications in populations not studied in the trial. 15 Further, RWD enables researchers to address questions about comparative effectiveness vs. other active therapies, mimicking the clinical scenario of choosing among available treatments.

Another reason to study HTE using RWD is to evaluate the generalizability of results from trials. Drugs may not have the same beneficial and adverse effects in real‐world settings as they do in RCTs, including in subgroups defined by patient characteristics that modify treatment effect. Understanding the distribution of such characteristics and how treatment varies in their presence is essential for anticipating the generalizability of treatment effects in specific patients. 16

PHARMACOEPIDEMIOLOGIC APPROACHES TO STUDYING HTE

Conducting a high‐quality study of HTE using RWD builds on the foundational concepts of pharmacoepidemiology. These foundational concepts encompass the entire research process ‐ from question formulation, study design, and data collection through analysis and reporting ‐ including critical assessments of the ability to control and measure confounding variables, the completeness and accuracy of confounders, risk factors, and exposure ascertainment, and the reliability of outcome measurements. Detailed descriptions of these methodological considerations have been published elsewhere. 12 , 17 , 18 , 19 , 20 , 21 When designing a study that includes evaluation of HTE, a researcher should first design the study for the average treatment effect over the entire study population, employing best practices. Building on this, the researcher can then incorporate methods described below for studying HTE.

Subgroup analysis methods

Historically, the predominant method for evaluating HTE has been subgroup analysis. 2 , 3 , 15 Subgroup analysis involves estimating treatment effects within groups defined by baseline characteristics. They offer direct insights into treatment effects for specific sub‐populations and can inform our understanding of drug mechanisms. Their simplicity and interpretability contribute to their enduring popularity. Traditional concerns have centered on the potential for spurious associations due to examining multiple subgroups, especially if they are specified posthoc, and limited statistical precision in estimating effects in individual subgroups. 14 , 22 , 23 However, even if these estimates have low precision to quantify effects in subgroups, they may contribute to accumulating evidence that can later be synthesized through meta‐analyses. 15 , 24 , 25

The experience gained from subgroup analyses in randomized controlled trials offers valuable insights into the process of choosing and defining subgroups. Two key principles emerge from this literature that should guide implementation. First, following guidelines developed for the RCT setting, subgroups should be defined by characteristics that are measurable, mechanistically reasonable, specified a priori, and based on clinically actionable patient features. 14 , 22 , 23 , 26 Common subgroup categories include demographic attributes like age, sex, and race and ethnicity, as well as clinical factors such as risk factors for the outcome of interest, history of related outcomes, pharmacogenetic alleles, and drug target expression. Poorly defined variables can be challenging to implement clinically and offer limited insights into mechanistic understanding. For example, differential treatment effects are often observed across racial and ethnic groups. However, these results can be difficult to interpret, since race and ethnicity are social constructs that may reflect a variety of factors, such as socioeconomic status, medication adherence, or genetic differences in drug metabolism. Moreover, the application of racial and ethnic classifications in clinical settings is both difficult and highly variable, highlighting the importance of studying characteristics with clear biological or clinical relevance whenever feasable. 27 Second, when working with continuous variables, definitions should avoid arbitrary cutoffs and instead be based on clinically meaningful thresholds where possible. For instance, guidelines provide established thresholds for staging chronic kidney disease based on measures of kidney function. 28 Similarly, in oncology, PD‐L1 expression thresholds (< 1%, ≥ 1–49%, > 50%) are well‐established in clinical literature and guide choice of treatment with drugs that inhibit PD‐L1. 29 , 30 Additionally, consistent use of thresholds facilitates meta‐analysis of results across studies.

While these principles provide a foundation for defining subgroups in a rigorous and clinically meaningful manner, their implementation in RWD analyses introduces additional challenges, particularly in the presence of confounding within subgroups. While propensity score (PS) methods are often used in RWD to address confounding for the overall effect estimate, when the data are split into subgroups, residual confounding can remain. To address this, researchers using PS matching methods should re‐estimate the PS in subgroups and match within subgroups; researchers using PS‐based weighting methods should re‐estimate PS within subgroups, calculate new weights, and reassess balance within subgroups. 31 , 32 , 33

A major limitation of subgroup analyses is the reference class problem. 1 , 2 , 3 , 34 , 35 A given patient can belong to multiple (actually, an infinite number of) subgroups, and the expected effects can conflict. For example, if a treatment is more effective in females but less effective in older adults, what can be inferred about the effect in older females? When making a clinical decision for a given patient, there is usually no clear rationale for choosing which subgroup to prioritize. Improving the clinical actionability of evidence about HTE requires methods that consider multiple baseline characteristics together. 1 , 34 , 35 This challenge highlights one opportunity where RWD can be particularly helpful compared with RCTs. Large RWD study sizes enable the study of subgroups defined by multiple characteristics simultaneously, such as levels of PD‐L1 expression and prior chemotherapy treatment in oncology patients. However, even when more specific subgroups are defined, the reference class problem for applying the results to a given patient ultimately persists. 3

Disease risk score methods

As an alternative to studying HTE using subgroups, researchers have developed strategies to study HTE that simultaneously incorporate multiple patient characteristics. The first of these methods utilizes a summary score of a patient's risk for the outcome had they not been treated. Such a summary score is often referred to as a disease risk score (DRS) in the pharmacoepidemiology literature and as a prognostic score or risk score in the RCT literature. 10 , 11 , 36 , 37 , 38 , 39 , 40 , 41 , 42 Many types of models can be used to estimate the DRS, including established clinical models externally validated on large datasets, such as the CHA2DS2‐VASc stroke risk score in atrial fibrillation 43 or the 10‐year cardiovascular disease risk estimator. 44 Importantly, the DRS must be estimated using only risk factors measured prior to the start of treatment to avoid introducing bias. 5

When established risk prediction models are unavailable, an internal model can be derived using data from within the study population. In a placebo‐controlled RCT, the control group theoretically represents the risk for outcome if the full trial population were not treated, and therefore can be leveraged to train the DRS model. However, care must be taken not to rely on the DRS in patients in whom the DRS model was fit. Doing so opens the analysis to bias because the DRS model may fit to the specifics or noise of the training data rather than a true pattern found in the general population, a phenomenon known as over‐fitting. 3 , 11 , 45 , 46 DRS models should be fit following best practice guidance for prediction modeling. 10 , 11 , 47 , 48 , 49 The performance of the risk model can be evaluated using metrics such as calibration and discrimination. 50 While these metrics provide important insights into how well the model predicts baseline risk, there is no consensus for what constitutes optimal performance in the context of HTE analysis. Addressing this gap represents an important area for future research.

In the RWD setting, an additional challenge is to address the generalizability of the DRS model. If the DRS model is fit on the control (or active comparator) group, the model may not generalize to the treated group. This is because treated patients and controls may have differences in baseline risks and distributions of characteristics. 51 , 52 , 53 The best approach to this challenge is an area of active research. 50

While DRS scores themselves can be used for confounding control, to study HTE they are used for stratification. Here additional methods must be used for confounding control, such as PS weighting or matching. As with subgroup analyses, there may be imbalance within the subgroups defined by DRS stratification. The PS may then need to be re‐estimated within subgroups to control for within‐stratum confounding. 50

When multiple outcomes are of interest, such as a primary therapeutic outcome and an adverse effect of treatment, DRS stratification methods are well‐suited for evaluating the benefit‐harm tradeoff with treatment. For instance, patients can be stratified by their baseline risk of an adverse outcome (e.g., bleeding events associated with an anticoagulant for stroke prevention in atrial fibrillation). Within each risk stratum, the effects of the anticoagulant on both the primary outcome (e.g., stroke prevention) and the adverse outcome (e.g., bleeding events) can be separately estimated. Analyzing both outcomes within the same strata of risk—defined by either the primary outcome or the adverse event, depending on the research question—enables meaningful comparisons of benefits and harms. 1 , 35

Another advantage of DRS stratification is that if there is an overall average treatment effect and risk varies in the population, there must be HTE by the DRS on either the difference scale, the ratio scale, or both. This HTE is often observed on the more clinically important RD scale. 1 , 35

A potential disadvantage of the DRS method is that examining effects across strata of a multivariable score may not provide the same kind of mechanistic insights provided by examining subgroups defined based on individual variables with mechanistic relevance. An additional limitation is that the predicted DRS is not guaranteed to correspond with treatment effect. 34 , 35 , 36 Thus, while patients with the highest DRS are arguably in the greatest need for treatment, they may not always be the same as the group of patients who will benefit from treatment. 34 For example, the presence of a pharmacogenetic allele may not predict the clinical outcome affected by a drug that exhibits pharmacogenetic variability in patients who are not treated with that drug. Additionally, while established guidelines exist for assessing results of these methods in RCTs for use in clinical practice—including recommendations for external validation and calibration of the DRS model 10 , 11 —comparable guidance has not yet been extended to the RWD setting.

Effect modeling methods

Effect modeling methods estimate HTE by predicting the effect of treatment for each individual in the study. This is accomplished by fitting an effect model to predict a patient's outcome if treated with the study drug rather than the comparator. The effect model can be trained on data external to the study when a representative dataset is available. If such a dataset is not available, the effect model can be fit internally on study data. However, sample‐splitting methods must be used to ensure the effect model is not both fit on and used in the same individuals, which could lead to over‐fitting. 36 When traditional regression methods are used to fit the effect model, the treatment‐characteristic interactions must be properly specified to avoid introducing bias. This is difficult to do, especially when non‐binary characteristics such as age have nonlinear relationships with the effects of treatment on the risk of outcome. 10 , 11 , 54 Flexible nonparametric and machine learning (ML) approaches to fitting the effect model are an appealing alternative and are a rapidly advancing area of research. 55 , 56 While flexible methods help to avoid the problem of properly specifying model interactions, their estimates of HTE lack transparency, so they are difficult to interpret and apply clinically. 57 Identifying and confirming effect modification relationships found using flexible methods is an active area of research. 57

Evaluating the performance of effect modeling methods presents unique challenges, as one of the potential outcomes for each patient (either the outcome under treated or comparator status) is inherently unobserved for each patient in the dataset. While performance metrics such as the concordance statistic for benefit have been proposed and applied in the RCT setting, there is no consensus on the most appropriate metrics for use in observational studies. 10 , 11 , 58 , 59 Further research is needed to establish reliable methods for assessing effect modeling performance in this context.

Once fit, effect models can be used to predict treatment effects for individuals, or they can be used to develop an effect score. 60 , 61 The effect score is typically defined as the difference between the predicted risk of outcome if the patient were treated vs. if they were not treated. 34 , 36 , 61 , 62 Patients are then stratified by the effect score, and estimates of HTE by effect score are made by comparing treatment effects within each level of effect score.

When analyzing multiple outcomes, such as a therapeutic outcome (e.g., stroke prevention) and an adverse effect of treatment (e.g., serious bleeding), separate effect models should be developed for each outcome. Each model should include the main effects and interaction terms relevant to its respective treatment‐outcome relationship, without requiring identical terms across models. In the RWD setting, confounding must be addressed. PS weighting methods have been proposed for such application. 63 For scoring approaches, as with DRS methods, final effect estimates for both outcomes should be analyzed within the same effect score strata definitions, whether based on the primary outcome or the adverse event. This consistency in stratification is essential to evaluate the treatment benefit–harm trade‐off in a clinically meaningful way.

Effect score approaches have the advantage of using multiple characteristics simultaneously to evaluate HTE, resolving the reference class problem associated with single‐variable subgroup analysis. However, because they rely on the effect models, they have many of the same limitations of using the effect model to directly predict treatment effects, such as model specification and performance evaluation. Yet, rather than relying directly on predictions for treatment effect estimates, these methods derive a summary score of the relationships between all covariates and treatment, making them less sensitive to misspecification. 3 Though, if the strata are too wide, this could represent a loss of granularity in risk prediction. 61

Effect score methods were originally developed and have recently been applied in the RCT setting. For example, in a secondary analysis of RCT data, Buell et al. 58 used machine learning to fit an effect score model to evaluate if the effect of high vs. low peripheral oxygen saturation targets on mortality among critically ill adults was modified by effect score. Another recent trial examined HTE using subgroup analysis, risk score approaches, and effect score methods and found similar results across the three methods. 62 Application of effect scores to non‐randomized observational studies is an area of ongoing methodologic research. Standard guidance is needed for assessing the results of these methods for clinical decision making.

COMPARISON OF METHODS TO EVALUATE HTE

The main goal of HTE analysis is to identify patients who will benefit and who will be harmed by treatment. In addition, characterizing HTE with regard to variables of mechanistic importance can provide biological insights into drug action. The methods summarized in this review have distinct advantages and limitations (Table 2 ). The most appropriate method will thus depend on the reason for conducting the HTE analysis. Additionally, factors such as data complexity, number of potential effect modifiers needed to incorporate into decision making, and interpretability will influence the choice of the method.

Table 2.

Summary of Methods to Evaluate Heterogeneity of Treatment Effect

Subgroup analysis Disease risk score (DRS) stratification Effect score approaches
Effect modeling Effect score
Purpose
  • Identify subgroups of individuals who will benefit or be harmed by treatment

  • Gain insights into mechanism of treatment effects

  • Identify how treatment effect varies across strata of baseline risk for outcome if not treated (or treated with comparator) (DRS)

  • Predict whether a patient will experience benefit or harm under each treatment option

  • Identify how treatment effect varies across strata of effect scores, integrating both baseline risks and treatment‐specific outcomes.

Description
  • Estimate treatment effects in subgroups defined by one (or possibly several) effect modifier 1 , 3 , 35

  • Report effect estimates and measures of precision

  • Fit DRS model‐ a model predicting risk for outcome of interest given multiple baseline characteristics if not treated (or treated with comparator)

  • Estimate DRS for individuals

  • Compare treatment effect estimates across strata of DRS 3 , 10 , 11 , 60

  • Fit effect model to estimate treatment effect for each individual given multiple baseline characteristics
  • Predict risk of outcome under either treatment strategy 3 , 10 , 11 , 56
  • Estimate treatment effect score – the difference in risk of outcome if the patient were treated vs. if not treated, given multiple baseline characteristics

  • Compare treatment effect estimates across strata of effect score 3 , 55 , 60 , 61

Strengths
  • Readily interpreted

  • Especially useful for well‐defined, clinically actionable subgroups

  • Results can be meta‐analyzed across studies 15 , 24

  • Simultaneously considers multiple characteristics 3 , 10 , 11 , 60

  • DRS model predicts the outcome under comparator condition only, no interactions between characteristics and treatment need to be specified 3 , 10 , 11 , 60

  • Simultaneously considers multiple characteristics 3

  • Directly estimates treatment effects for individuals 3

  • Same as effect modeling

  • Less sensitive to misspecification of the effect model 3

Limitations
  • Reference Class Problem: estimating effects in narrowly defined subgroups does not account for patients who simultaneously belong to multiple, overlapping subgroups, limiting ability to assess risk of harms or benefits 1 , 3 , 35

  • Examining effects across strata of multivariable summary score obscures mechanistic insights

  • Treatment effect might not correlate to DRS 34 , 36

  • Traditional modeling techniques must properly specify statistical interactions between treatment and characteristics 3

  • Flexible modeling methods to identify model form yield non‐transparent predictions that are difficult to interpret 3

  • Same as effect modeling

  • Loss of predictive granularity when strata are wide 61

Future Directions
  • Applications of effect modeling methods to identify subgroups 57

  • Extensions to real‐ world data settings
  • Guidance needed for evaluation of DRS model performance
  • Methods needed to understand relationships between characteristics and treatment discovered by flexible modeling approaches 57
  • Guidance needed for evaluation of effect model performance
  • Same as effect modeling

Subgroup analyses can give insight on treatment effects in subgroups, helping identify who is at risk and informing mechanistic knowledge. These are especially helpful when the subgroups are well‐defined and clinically relevant. They also have the advantage of being transparent and readily interpreted, especially for subgroups defined based on mechanistically important variables. However, the reference class problem arises when applying the results of subgroup analyses because membership to multiple subgroups makes it challenging to determine which subgroup findings are clinically actionable, potentially leading to conflicting recommendations.

To address these limitations, methods to study HTE by multiple covariates simultaneously have been developed, including DRS and effect modeling approaches. Of these, DRS methods aim to evaluate whether treatment effect varies by baseline risk for the outcome. This approach avoids the need to explicitly model interactions between patient characteristics and treatment, relying instead on fewer modeling assumptions than effect modeling methods. However, DRS methods may fail to fully capture HTE if the DRS does not include treatment effect modifiers.

Effect modeling methods, by contrast, directly estimate treatment effects. This makes them more suited to explicitly characterizing HTE and predicting individual patient responses to treatment. However, this direct modeling approach requires accurately specifying complex interactions, which is both challenging and sensitive to modeling choices. Flexible modeling frameworks, such as the meta‐learner framework 56 and causal forests, 64 offer promise in addressing these challenges, but they remain under active development. Effect score methods are unique because they focus on the difference in predicted outcomes (i.e., treatment vs. comparator), while DRS methods focus only on the baseline risk without treatment. However, because effect score methods are based on a multivariable score rather than individual, mechanistically important variables, the results may lack biological interpretability.

Sparse outcomes, such as rare adverse events, are common in pharmacoepidemiology and pose significant challenges for all methods of HTE analysis. Subgroup analyses and effect modeling methods are particularly vulnerable to instability when events are sparse, as they further divide the data into smaller groups, leading to bias and imprecision in estimates. This issue is exacerbated in high‐dimensional data or when interaction terms are included. DRS methods are also affected by sparse outcomes when the models are derived internally. In such cases, the ability of the DRS model to discriminate outcome risk affects the HTE results, and sparse training data undermines this discrimination ability. 2 , 47 , 48 , 49 However, if the DRS model is highly discriminating—such as a potentially externally validated model—stratification by DRS may be better equipped to detect HTE in scenarios with low outcome rates, providing valuable insights in these challenging settings. 1 , 2 , 10 , 11

In summary, subgroup analyses are simple, transparent, rely on minimal modeling assumptions, and may provide mechanistic insight. However, their application is hampered by the reference class problem when there are multiple effect modifiers. DRS methods incorporate many baseline characteristics and require fewer modeling assumptions, but may fail to capture true HTE if driven by non‐modifier risk factors and may not provide the same kind of mechanistic insight as subgroup analyses. Effect modeling methods are most capable of directly estimating treatment effects and characterizing HTE, but depend heavily on accurate model specification and are computationally complex.

CONCLUSIONS

Different people have different responses to drugs. Studies of HTE using RWD are needed to characterize who is at risk of harm from treatment and who is likely to benefit from treatment and to gain insights into drug mechanisms. Subgroup analyses are valuable for identifying well‐defined groups of individuals with varying treatment effects and variables of mechanistic importance. However, they have significant limitations when multiple characteristics modify treatment effects but are analyzed separately. Approaches that incorporate multiple patient characteristics simultaneously are needed to predict the safety and effectiveness of medications in diverse patient populations, but may lack mechanistic interpretability. These methods include disease risk score methods and effect modeling methods. DRS methods are relatively simple to implement and can be useful tools clinically to predict HTE. Effect modeling approaches have great potential for characterizing HTE; however, they are still being developed.

FUNDING

This research was supported by the National Institutes of Health Grants 5T32GM075766 and 1F32DK141217.

CONFLICTS OF INTEREST

The authors declared no competing interests for this work.

This work has not previously been presented.

References

  • 1. Kent, D.M. , Steyerberg, E. & van Klaveren, D. Personalized evidence based medicine: predictive approaches to heterogeneous treatment effects. BMJ 363, k4245 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Kent, D.M. , Nelson, J. , Dahabreh, I.J. , Rothwell, P.M. , Altman, D.G. & Hayward, R.A. Risk and treatment effect heterogeneity: re‐analysis of individual participant data from 32 large clinical trials. Int. J. Epidemiol. 45, 2075–2088 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Dahabreh, I.J. , Trikalinos, T.A. , Kent, D.M. & Schmid, C.H. Heterogeneity of treatment effects. In Methods in Comparative Effectiveness Research (eds. Gatsonis, C. & Morton, S.C. ) 227–271 (Chapman and Hall/CRC, Boca Raton, FL, 2017). [Google Scholar]
  • 4. van der Weele, T.J. On the distinction between interaction and effect modification. Epidemiology 20, 863–871 (2009). [DOI] [PubMed] [Google Scholar]
  • 5. Schisterman, E.F. , Cole, S.R. & Platt, R.W. Overadjustment bias and unnecessary adjustment in epidemiologic studies. Epidemiology 20, 488 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Hennessy, S. et al. Pharmacoepidemiologic methods for studying the health effects of drug–drug interactions. Clin. Pharmacol. Ther. 99, 92–100 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Vandenbroucke, J.P. et al. Strengthening the reporting of observational studies in epidemiology (STROBE). Epidemiology 18, 805–835 (2007). [DOI] [PubMed] [Google Scholar]
  • 8. van der Weele, T.J. & Knol, M.J. A tutorial on interaction. Epidemiol. Methods 3, 33–72 (2014). [Google Scholar]
  • 9. Knol, M.J. & VanderWeele, T.J. Recommendations for presenting analyses of effect modification and interaction. Int. J. Epidemiol. 41, 514–520 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Kent, D.M. et al. The predictive approaches to treatment effect heterogeneity (PATH) statement. Ann. Intern. Med. 172, 35–45 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Kent, D.M. et al. The predictive approaches to treatment effect heterogeneity (PATH) statement: explanation and elaboration. Ann. Intern. Med. 172, W1–W25 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. European Network of Centres for Pharmacoepidemiology and Pharmacovigilance (ENCePP) . Guide on methodological standards in pharmacoepidemiology (Revision 11). EMA/95098/2010. <http://www.encepp.eu/standards_and_guidance> Accessed September 9, 2024.
  • 13. Theken, K.N. et al. Clinical pharmacogenetics implementation consortium guideline (CPIC) for CYP2C9 and nonsteroidal anti‐inflammatory drugs. Clin. Pharmacol. Ther. 108, 191–200 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Wang, R. , Lagakos, S.W. , Ware, J.H. , Hunter, D.J. & Drazen, J.M. Statistics in medicine—reporting of subgroup analyses in clinical trials. N. Engl. J. Med. 357, 2189–2194 (2007). [DOI] [PubMed] [Google Scholar]
  • 15. Segal, J.B. et al. Assessing heterogeneity of treatment effect in real‐world data. Ann. Intern. Med. 176, 536–544 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Stuart, E.A. Generalizability of clinical trials results. In Methods in Comparative Effectiveness Research (eds. Gatsonis, C. & Morton, S.C. ) 177–201 (Chapman and Hall/CRC, Boca Raton, FL, 2017). [Google Scholar]
  • 17. Wang, S.V. et al. STaRT‐RWE: structured template for planning and reporting on the implementation of real world evidence studies. BMJ 372, m4856 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Public Policy Committee IS of P . Guidelines for good pharmacoepidemiology practice (GPP). Pharmacoepidemiol. Drug Saf. 25, 2–10 (2016). [DOI] [PubMed] [Google Scholar]
  • 19. Langan, S.M. et al. The reporting of studies conducted using observational routinely collected health data statement for pharmacoepidemiology (RECORD‐PE). BMJ 363, k3532 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Wang, S.V. et al. Reporting to improve reproducibility and facilitate validity assessment for healthcare database studies V1.0. Value Health 20, 1009–1022 (2017). [DOI] [PubMed] [Google Scholar]
  • 21. US FDA . Framework for FDA's real‐world evidence program. Published online December 2018. <https://www.fda.gov/media/120060/download>.
  • 22. Assmann, S.F. , Pocock, S.J. , Enos, L.E. & Kasten, L.E. Subgroup analysis and other (mis)uses of baseline data in clinical trials. Lancet 355, 1064–1069 (2000). [DOI] [PubMed] [Google Scholar]
  • 23. Pocock, S.J. , Assmann, S.E. , Enos, L.E. & Kasten, L.E. Subgroup analysis, covariate adjustment and baseline comparisons in clinical trial reporting: current practiceand problems. Stat. Med. 21, 2917–2930 (2002). [DOI] [PubMed] [Google Scholar]
  • 24. Varadhan, R. , Segal, J.B. , Boyd, C.M. , Wu, A.W. & Weiss, C.O. A framework for the analysis of heterogeneity of treatment effect in patient‐centered outcomes research. J. Clin. Epidemiol. 66, 818–825 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Hernán, M.A. Causal analyses of existing databases: no power calculations required. J. Clin. Epidemiol. 144, 203–205 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Schandelmaier, S. et al. Development of the instrument to assess the credibility of effect modification analyses (ICEMAN) in randomized controlled trials and meta‐analyses. CMAJ 192, E901–E906 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Flanagin, A. , Frey, T. , Christiansen, S.L. & AMA Manual of Style Committee . Updated guidance on the reporting of race and ethnicity in medical and science journals. JAMA 326, 621–627 (2021). [DOI] [PubMed] [Google Scholar]
  • 28. Stevens, P.E. et al. KDIGO 2024 clinical practice guideline for the evaluation and Management of Chronic Kidney Disease. Kidney Int. 105, S117–S314 (2024). [DOI] [PubMed] [Google Scholar]
  • 29. National Comprehensive Cancer Network . NCCN Clinical Practice Guidelines in Oncology (NCCN Guidelines): non‐small cell lung cancer. NCCN. 2025;3.2025. <https://www.nccn.org/professionals/physician_gls/pdf/nscl.pdf>.
  • 30. Jaiyesimi, I.A. et al. Therapy for stage IV non–small cell lung cancer without driver alterations: ASCO living guideline, version 2023.3. J. Clin. Oncol. 42, e23–e43 (2024). [DOI] [PubMed] [Google Scholar]
  • 31. Dong, J. , Zhang, J.L. , Zeng, S. & Li, F. Subgroup balancing propensity score. Stat. Methods Med. Res. 29, 659–676 (2020). [DOI] [PubMed] [Google Scholar]
  • 32. Wang, S.V. et al. Relative performance of propensity score matching strategies for subgroup analyses. Am. J. Epidemiol. 187, 1799–1807 (2018). [DOI] [PubMed] [Google Scholar]
  • 33. Yang, S. , Li, F. , Thomas, L.E. & Li, F. Covariate adjustment in subgroup analyses of randomized clinical trials: a propensity score approach. Clin. Trials 18, 570–581 (2021). [DOI] [PubMed] [Google Scholar]
  • 34. VanderWeele, T.J. , Luedtke, A.R. , van der Laan, M.J. & Kessler, R.C. Selecting optimal subgroups for treatment using many covariates. Epidemiology 30, 334–341 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Dahabreh, I.J. , Hayward, R. & Kent, D.M. Using group data to treat individuals: understanding heterogeneous treatment effects in the age of precision medicine and patient‐centred evidence. Int. J. Epidemiol. 45, 2184–2193 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Dahabreh, I.J. & Kazi, D.S. Toward personalizing care: assessing heterogeneity of treatment effects in randomized trials. JAMA 329, 1063–1065 (2023). [DOI] [PubMed] [Google Scholar]
  • 37. Hansen, B.B. The prognostic analogue of the propensity score. Biometrika 95, 481–488 (2008). [Google Scholar]
  • 38. Wyss, R. , Glynn, R.J. & Gagne, J.J. A review of disease risk scores and their application in pharmacoepidemiology. Curr. Epidemiol. Rep. 3, 277–284 (2016). [Google Scholar]
  • 39. Kent, D.M. , Rothwell, P.M. , Ioannidis, J.P. , Altman, D.G. & Hayward, R.A. Assessing and reporting heterogeneity in treatment effects in clinical trials: a proposal. Trials 11, 85 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Arbogast, P.G. & Ray, W.A. Use of disease risk scores in pharmacoepidemiologic studies. Stat. Methods Med. Res. 18, 67–80 (2009). [DOI] [PubMed] [Google Scholar]
  • 41. Arbogast, P.G. & Ray, W.A. Performance of disease risk scores, propensity scores, and traditional multivariable outcome regression in the presence of multiple confounders. Am. J. Epidemiol. 174, 613–620 (2011). [DOI] [PubMed] [Google Scholar]
  • 42. Glynn, R.J. , Gagne, J.J. & Schneeweiss, S. Role of disease risk scores in comparative effectiveness research with emerging therapies. Pharmacoepidemiol. Drug Saf. 21(S2), 138–147 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Lip, G.Y.H. , Nieuwlaat, R. , Pisters, R. , Lane, D.A. & Crijns, H.J.G.M. Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor‐based approach: the euro heart survey on atrial fibrillation. Chest 137, 263–272 (2010). [DOI] [PubMed] [Google Scholar]
  • 44. American College of Cardiology . ASCVD risk estimator plus. August (2023). <http://tools.acc.org/ASCVD‐Risk‐Estimator‐Plus?_ga=2.130677199.1751240316.1722282859‐1234225747.1722282501>.
  • 45. Abadie, A. , Chingos, M.M. & West, M.R. Endogenous stratification in randomized experiments. Rev. Econ. Stat. 100, 567–580 (2018). [Google Scholar]
  • 46. Steyerberg, E.W. Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating 2nd edn. (Cham: Springer, 2019). [Google Scholar]
  • 47. Riley, R.D. et al. Minimum sample size for developing a multivariable prediction model: PART II ‐ binary and time‐to‐event outcomes. Stat. Med. 38, 1276–1296 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Riley, R.D. et al. Calculating the sample size required for developing a clinical prediction model. BMJ 368, m441 (2020). [DOI] [PubMed] [Google Scholar]
  • 49. van Smeden, M. et al. Sample size for binary logistic prediction models: beyond events per variable criteria. Stat. Methods Med. Res. 28, 2455–2474 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Rekkas, A. , van Klaveren, D. , Ryan, P.B. , Steyerberg, E.W. , Kent, D.M. & Rijnbeek, P.R. A standardized framework for risk‐based assessment of treatment effect heterogeneity in observational healthcare databases. NPJ Digit. Med. 6, 1–11 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Altman, D.G. & Royston, P. What do we mean by validating a prognostic model? Stat. Med. 19, 453–473 (2000). [DOI] [PubMed] [Google Scholar]
  • 52. Harrell, F.E. Regression Modeling Strategies: With Applications to Linear Models, Logistic and Ordinal Regression, and Survival Analysis (Cham; Springer International Publishing, 2015). 10.1007/978-3-319-19425-7. [DOI] [Google Scholar]
  • 53. Lash, T.L. & Rothman, K.J. Selection bias and generalizability. In Modern Epidemiology. (eds. Lash T.L., VanderWeele T.J., Haneuse S. & Rothman K.J.), 315–331 (Pennsylvania: Wolters Kluwer; 2021). [Google Scholar]
  • 54. van Klaveren, D. , Balan, T.A. , Steyerberg, E.W. & Kent, D.M. Models with interactions overestimated heterogeneity of treatment effects and were prone to treatment mistargeting. J. Clin. Epidemiol. 114, 72–83 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55. Künzel, S.R. , Sekhon, J.S. , Bickel, P.J. & Yu, B. Metalearners for estimating heterogeneous treatment effects using machine learning. Proc. Natl. Acad. Sci. U.S.A. 116, 4156–4165 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Kennedy, E.H. Towards optimal doubly robust estimation of heterogeneous causal effects. Electron. J. Stat. 17, 3008–3049 (2023). [Google Scholar]
  • 57. Bonvini, M. , Zeng, Z. , Yu, M. , Kennedy, E.H. & Keele, L. Flexibly estimating and interpreting heterogeneous treatment effects of laparoscopic surgery for cholecystitis patients. Published online November 7, (2023) 10.48550/arXiv.2311.04359. [DOI]
  • 58. Buell, K.G. et al. Individualized treatment effects of oxygen targets in mechanically ventilated critically ill adults. JAMA 331, 1195–1204 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. van Klaveren, D. , Steyerberg, E.W. , Serruys, P.W. & Kent, D.M. The proposed ‘concordance‐statistic for benefit’ provided a useful metric when modeling heterogeneous treatment effects. J. Clin. Epidemiol. 94, 59–68 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Rekkas, A. et al. Predictive approaches to heterogeneous treatment effects: a scoping review. BMC Med. Res. Methodol. 20, 264 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Wang, G. , Heagerty, P.J. & Dahabreh, I.J. Using effect scores to characterize heterogeneity of treatment effects. JAMA 331, 1225–1226 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62. Goligher, E.C. , Lawler, P.R. , Jensen, T.P. et al. Heterogeneous treatment effects of therapeutic‐dose heparin in patients hospitalized for COVID‐19. JAMA 329, 1066–1077 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Lipkovich, I. , Svensson, D. , Ratitch, B. & Dmitrienko, A. Modern approaches for evaluating treatment effect heterogeneity from clinical trials and observational data. Stat. Med. 43, 4388–4436 (2024). [DOI] [PubMed] [Google Scholar]
  • 64. Jawadekar, N. et al. Practical guide to honest causal forests for identifying heterogeneous treatment effects. Am. J. Epidemiol. 192, 1155–1165 (2023). [DOI] [PubMed] [Google Scholar]

Articles from Clinical Pharmacology and Therapeutics are provided here courtesy of Wiley and American Society for Clinical Pharmacology and Therapeutics

RESOURCES