Abstract
The estimation of treatment response heterogeneity (TRH) is increasingly important as medicine moves toward personalized approaches. While various statistical methods have been proposed to quantify TRH in parallel-group trials, the standard deviation of individual responses (SDIR) has gained prominence within physiological research. This method is intended to quantify individual response variation by comparing standard deviations of change scores between intervention and control groups. We acknowledge that SDIR represents an improvement over many other flawed approaches that often involve responder counting. However, SDIR has critical limitations: 1) it cannot overcome the fundamental problem of causal inference because the correlation between potential outcomes remains unidentifiable, 2) it is incorrectly predicated on the assumption that TRH is present only when treatment group variance exceeds control group variance, and 3) it is statistically inefficient. We present an alternative framework, which involves assessing heteroskedasticity and estimating the bounds for the standard deviation of treatment effects (SDD). The presence of heteroskedasticity between treatment groups is a sufficient but not necessary condition for the presence of TRH. Further, SDD makes fewer assumptions than SDIR and, therefore, paints a more complete picture of potential TRH. Using data from a published exercise physiology study, we demonstrate how SDD can better characterize uncertainty in TRH estimation. We recommend researchers probe TRH by assessing heteroskedasticity, providing bounds for SDD, and estimating outcome distributions and probabilities while carefully crafting the theoretical rationale for the presence of TRH.
Treatment response heterogeneity (TRH) has become increasingly important with the rise of personalized medicine. While various statistical approaches have been proposed to quantify this heterogeneity in parallel-group trials, one method has gained prominence in exercise physiology research: the standard deviation of individual responses (SDIR). This method, popularized by Atkinson and Batterham (2015) and Hopkins (2015), is intended to quantify “true” individual response variation by comparing the standard deviations of change scores between intervention and control groups.
The SDIR approach has been utilized mainly in exercise physiology research, where it has been used to identify so-called “responders” and “non-responders” to interventions and estimate the proportion of individuals who might benefit from treatment (Bonafiglia, et al., 2021; Swinton et al, 2018). Those utilizing SDIR in situations where there is a negative control (such as a no-treatment or placebo group) implicitly assume that such groups exhibit zero true inter-individual outcome variance, and any observed variance can be attributed solely to measurement error and within-subject variability (e.g., biological variation). However, despite its popularity, the SDIR method rests on several other crucial assumptions that can limit its validity and utility.
In this paper, we critically evaluate the statistical and conceptual foundations of the SDIR method. Specifically, we show that: 1) SDIR has strict assumptions about TRH, 2) this may lead researchers to dismiss the potential of TRH when it is present, and 3) the method is inefficient compared to other heteroskedasticity tests (i.e., statistical methods that assess how residual variance depends on an independent variable, such as group membership).
We then present alternative approaches that address these limitations and allow researchers to estimate TRH without such strict assumptions. Based on bounds for the true standard deviation of treatment effects, these methods provide a more complete picture of potential response variation while making fewer assumptions about its underlying nature.
Our goal is to encourage the investigation of TRH using approaches that rest on sound statistical foundations and have transparent and justifiable assumptions. By acknowledging the limitations of current methods and exploring approaches that explicitly account for uncertainty, researchers can enhance their contributions to the investigation of personalized treatment approaches.
Defining Treatment Response Heterogeneity
In parallel group trials, understanding TRH requires a precise distinction between true treatment effects (i.e., treatment response) and observed changes. Consider a pre-post parallel group design comparing treatment (T) and control (C). By the end of the study, an individual has two potential outcomes: Y(T) if assigned to treatment and Y(C) if assigned to control. The true individual treatment effect is the difference between these potential outcomes: D = Y(T) - Y(C). TRH represents how D varies across individuals. Importantly, the true individual D cannot be directly observed in parallel group trials because each participant receives only one condition. This is true even in a standard 2-condition 2-period crossover trial because each participant receives only one condition in one period of time. Within each study participant, only one outcome can be observed in any period of time; the other (potential) outcome is the counterfactual (i.e., what would have happened had the participant been randomized to the other group).
A common misconception is attributing observed within-group changes to the treatment (Atkinson & Batterham, 2015). Researchers often consider the pre-post change in the treatment group [Y(T)Post – Y(T)Pre] as representing D and the change in the control group [Y(C)Post – Y(C)Pre] as the “placebo effect“. These interpretations erroneously assume that had an individual not received the treatment or placebo, the measured outcome would not have changed over the duration of the study. Without a control group, one cannot separate treatment effects from natural variation, regression to the mean, or other temporal changes of the dependent variable. Even with gold-standard measurement tools, at least some random within-subject variation is inevitable over time. This variation can create the appearance of response heterogeneity when none exists or mask true treatment response differences. As Atkinson and Batterham (2015) described, this often represents “random error masquerading as response differences“. While individual effects remain unobservable, an unbiased estimate of the average causal treatment effect can be obtained by comparing observed outcomes between groups (Laird, 1983).
Quantifying TRH involves estimating the standard deviation of individual treatment effects, denoted SDD for a population and SDD in a sample. However, the standard deviation of individual treatment effects, SDD, cannot be directly computed from observed data in parallel group trials because it requires knowing each participant’s response to both conditions. The observed variation in outcomes or changes within treatment groups includes both true response heterogeneity and other sources of variation, particularly within-subject random variability between measurement timepoints. Using baseline or pre-intervention scores to calculate change scores does not eliminate this issue.
Background on the Standard Deviation of Individual Responses
Given these points, we would like to address a recently popularized method for assessing TRH, referred to as the SDIR (Atkinson & Batterham, 2015; Atkinson et al, 2018, Atkinson et al, 2019; Hopkins, 2015; Swinton, 2023). The proponents of this specific method have claimed that 1) it can be used to detect and estimate the degree of TRH, 2) the estimated TRH represents TRH caused by the treatment, and 3) the estimated TRH can be used to estimate the proportion of “responders”. All three claims have important assumptions that, when violated, rise to the level of errors. Herein, we aim to explain these assumptions and how they can be addressed. However, we would like to note that this SDIR approach is inferentially superior to many other methods for investigating TRH. Simple responder counts by dichotomizing research participants as “responders” or “non-responders” is presumptive, and the superiority of SDIR in this regard is clearly discussed by Atkinson et al. (2019).
The SDIR method has been proposed for studies involving continuous outcomes that are measured both prior to and after receiving some treatment (pre-post measurements in parallel groups). The change scores for an outcome for every individual in the study can then be calculated. The logic proposed by proponents of the SDIR approach is the following: The standard deviation of the change scores only includes the within-subject variation and measurement error . Within-subject variation encompasses multiple sources, including biological fluctuations over time, behavioral changes unrelated to treatment, environmental factors, and the natural stability or instability of the outcome being measured. For example, in a study of blood pressure medications, might include daily fluctuations in blood pressure due to stress, diet, sleep patterns, and seasonal effects, while in a fitness study measuring maximal oxygen consumption, it could include variations due to motivation during testing or circadian rhythms due to timing of the tests. Measurement error reflects the technical precision of the measurement tools. If treatment effects vary across individuals , the treatment group will exhibit greater change score variability than the control group (i.e., ). In simpler terms, the logic of SDIR is that if the treatment impacts people differently, the variability of change scores (improvements or declines) will be greater in the treatment group than in the control group. This assumption is based on the premise that both groups are equally exposed to non-treatment-related factors affecting observed change scores (e.g., random within-person variation, measurement error, true temporal changes unrelated to treatment, etc.), and only the treatment group receives an active intervention. Thus, an observed difference in variances between groups is attributed to individual variation in the treatment response. For this logic to hold, the control group must function as a negative control, meaning it should exhibit no true inter-individual response variance, and all observed variance can be attributed to within-subject variation and measurement error. This logic makes a critical mathematical assumption: individual treatment effects D are independent of the effects of the control (see Supplementary Material).
The SDIR is calculated as,
which relies on the estimated standard deviation of change scores in the treatment (sT) and control groups (sC). Confidence intervals can be formed on the differences in variances using a normal approximation, and the square root of those bounds can be taken to obtain the desired confidence limits (Mills et al, 2021; Hopkins, 2015). The sgn function included in our equation for calculating SDIR serves to always make SDIR positive for situations in which ST is less than SC, when the difference would be negative. Barring the hassle of potential sign errors if the sgn function is not utilized, there is nothing statistically erroneous about calculating SDIR itself because it is just an estimate of the square root of the difference in variances.
Some have advocated using SDIR to estimate the proportion of responders (Swinton et al., 2021; Swinton, 2023). They utilize the standard normal cumulative distribution function to estimate the proportion of individuals whose treatment effect would exceed a specified threshold (typically the minimally clinically important difference or the minimal detectable difference):
Where is the mean difference between the treatment and control groups and t is the threshold indicating the smallest difference of interest.
The Statistical Issues
Issue #1: The Fundamental Problem of Causal Inference
As some have illustrated (Swinton, 2023), SDIR seems akin to a random slope component from a mixed-effects model. As stated by Atkinson et al. (2019), the SDIR “is considered a parameter for the distribution of true responses in the population of interest alongside the mean treatment effect.” This is a major assumption that might often be incorrect (see Supplemental Material for a greater discussion of this assumption). Although it is tempting to interpret this as being a way to estimate the variance of D, it relies on a critical assumption about something we cannot observe: the correlation between an individual’s outcome with the treatment and their outcome with the control (i.e., the correlation between Y(T) and Y(C) for the same person). To illustrate, imagine we could magically give the same person both treatment and control simultaneously. Would someone who has a favorable outcome with the treatment have a similarly favorable outcome with the control? This correlation is mathematically well-defined within the potential outcomes framework and directly affects variance calculations, but it remains fundamentally unobservable in parallel group designs. This is often referred to as the “fundamental problem of causal inference” (Holland, 1986) because we can only observe a single potential outcome for each participant within a parallel groups experimental design. The SDIR calculation implicitly assumes this correlation takes a specific value, but alternative correlation values could yield substantially different estimates of TRH.
When TRH is present, the genuine impact of treatment(s) varies among different individuals within a population. To clarify these points, we shall consider two potential outcomes, Y(C) and Y(T), from a two-arm parallel groups trial mentioned earlier. The variable Y(T)i is the value of the outcome on participant i when exposed to the treatment, and Y(C)i is the value if exposed to the control condition. The two values are imagined to be measured at the same moment in time. In a parallel group trial, only the outcome corresponding to the group to which the participant was assigned is observable. The concept of two potential outcomes aids in conceptualizing a true treatment effect. Assuming other assumptions are met, the true treatment effect has a mean and variance defined by Gadbury & Iyer (2000) as the following:
The marginal distributions of Y(T) and Y(C) are readily available from any parallel group trial, but the correlation between potential outcomes, ρTC, is not (empirically) identifiable. For TRH to be absent (σD = 0), both the variance of Y(C) and Y(T) would have to be equal and the correlation (ρTC) equal to 1. The central challenge in estimating TRH is the inability to estimate this correlation from real, observed data. For this reason, the presence of a difference in the variances between treatment conditions (at the population parameter level, not at the sample statistic level) is a sufficient but not necessary condition to indicate that TRH is present.
Because ρTC must lie within the interval [−1,1], upper and lower bound estimates of σD can be calculated. The lower bound, which occurs when ρTC = 1, for the sample variance is min(SDD) = |sT−sC|. The upper bound, which occurs when ρTC = −1, is max(SDD) = sT + sC. The standard deviation representing TRH could be greater than min(SDD) but cannot exceed max(SDD) (Gadbury & Iyer, 2000). Further, it could be reasonable to speculate that ρTC is closer to 1 than −1, but this is an untestable assumption.
Given the mathematical assumptions laid out by Gadbury et al. (2001), it becomes clear that the SDIR should not uncritically be assumed to be a “parameter [or unbiased estimate thereof] for the distribution of true responses”. The use of SDIR involves making strong assumptions that cannot be empirically tested because we cannot identify correlation between potential outcomes (ρTC). Indeed, these assumptions may be untenable in many exercise science studies (see Supplement). Instead, one can loosen the strict assumptions made by SDIR by presenting bounds of the SDD. The strict assumptions made by SDIR propagate into the calculation of the proportion of responders. If SDIR’s strict assumptions are not met, the true denominator, the SDD, could be considerably larger or smaller. The proportion of those who benefit (Pbenefit = Pr[D>0]), are harmed (Pharm = Pr[D<0]), or even to respond beyond a predetermined threshold (Presponder = Pr[D>t]) can be estimated from the bounds on SDD (Gadbury et al, 2001). Utilizing these bounds and the estimate of , we can then use the standard normal cumulative distribution function to estimate these probabilities.
Issue #2: Sources of Individual Response Variance
Proponents of SDIR have articulated that TRH is only present when the variance is greater in a treatment arm compared to the control arm. It was stated by Atkinson et al. (2018) that, “In a parallel group study, true individual differences in exercise response are present only if the SD of change is substantially larger in the exercise group than the control group.” Yet, it has been discussed elsewhere that TRH can exist even when the variance is not greater in the treatment group (Cortés et al, 2019; Mills et al., 2021). In fact, treatments could “homogenize” outcomes relative to the control group, which arguably still represents TRH (i.e., the degree of treatment effect still varies between individuals).
By extension, this implies that the intervention must cause a difference in population variances. This might be a plausible (but not necessary or verifiable) circumstance for very short placebo-controlled trials wherein an experimental drug might be the only major source of additional variance. However, such a broad assumption would certainly be too bold for comparative effectiveness studies that involve two competing treatments (e.g., new treatment vs standard-of-care), wherein individuals may have varying responses to both treatments. A difference in variance does imply some degree of heterogeneity (though importantly, the converse and inverse are not true), and this conclusion should not be restricted to when the experimental group has a higher variance than the control group (Senn, 2016). Comparisons of variances between the treatment groups should then be two-sided and consider that there could be greater variance in either the control or treatment arm of a parallel group trial.
Beyond this conceptual limitation, the SDIR framework is also restricted in its applicability, as it can only be used in parallel groups trials where pre-post measurements are available to calculate change scores. In contrast, the SDD and other heteroskedasticity tests provide a more versatile framework that can be applied to any parallel groups design, including those where only post-intervention measurements are available. This broader applicability is especially valuable in settings where baseline measurements are impractical, impossible, or simply weren’t collected in an existing dataset, allowing researchers to investigate TRH across a wider range of study designs (see Cortés et al, 2019 and Mills et al., 2021 for examples).
Issue #3: Statistical Efficiency and Testing
Though the SDIR calculation provides useful insights, it lacks statistical efficiency when comparing group variances, typically requiring substantial sample sizes for precise interval estimates. Statistical efficiency here refers to how effectively a method extracts information with minimal variance—essentially the precision of parameter estimation compared to alternative approaches. Furthermore, while variance comparisons across groups can help identify TRH, this methodology, whether using SDIR or other techniques, faces substantial limitations.
Compared to other heteroskedasticity tests (e.g., variance ratio test), the SDIR approach has less statistical power, with this effect being most pronounced at very small sample sizes, where even the nominal alpha level is not maintained (see Supplementary Material). Given that many studies in physiology research have small sample sizes, researchers may overlook possible TRH when further investigation is warranted (i.e., high type II error rates). In particular, differences in variances between groups may be hard to detect based on “statistical significance”. For example, to achieve 80% statistical power when testing for heteroskedasticity between two groups, a study would need 68 participants per group to detect a two-fold variance increase, and nearly 3,500 participants per group to detect a modest 10% difference in variances (PASS 2024, v24.0.1). Most physiological studies (Renwick et al, 2024) and even meta-analyses lack sufficient statistical power to detect such differences in variances between treatment groups. Physiologists should exercise substantial caution when determining TRH presence based on heteroskedasticity tests with limited sample sizes (for example, fewer than 60 participants per group, and certainly, they should not conclude the absence of TRH based on such tests—see Supplemental Material for more details).
Comparing variances to test for TRH dates back to a 1938 letter from Ronald A. Fisher to Henry Daniels (Fisher, 1990), and various tests for detecting heteroskedasticity have been developed since (Mills et al, 2021). While substantial variance differences between groups should warrant additional investigation into why individuals might respond differently to treatments, it is important to note that TRH doesn’t always result in increased variance in the treatment compared to the control arm (Cortés et al, 2019; Mills et al., 2021). Additional caution is also warranted because heteroskedasticity may reflect measurement scale issues. For example, in some cases, a transformation (like the log-transformation) can eliminate heteroskedasticity (Senn, 2004), but the interpretation of TRH on the clinically meaningful scale should not be ignored (see Poulson et al., 2012 for a greater discussion).
Even when group variances are equal, σD could be meaningfully large, indicating true individual differences in treatment response. Consider a scenario where some individuals experience increased benefit while others experience decreased benefit from treatment. Such differential responses could produce similar overall variance between treatment and control groups while still generating substantial σD. Tests of heteroskedasticity could incorrectly be interpreted as suggesting no TRH exists in this case, because they only capture differences in group variances rather than the true variance of individual treatment effects. The variance components from a parallel groups trial cannot be used to definitively rule out the potential of TRH.
Therefore, while finding different variances between groups is sufficient to demonstrate heterogeneous treatment effects, it is not a necessary condition: TRH can exist even when variances are similar between groups. These tests may be useful when the intention is to establish sufficient conditions for claiming TRH, but the failure of these tests to reach statistical significance cannot be used to eliminate the possibility of TRH. Researchers should be cautious about relying solely on variance comparisons when making claims regarding TRH, especially in scenarios where sample sizes are small and tests of heteroskedasticity will be underpowered. This issue is obviously not limited to the SDIR, but it is more pronounced compared to other, more efficient tests. Compared to other issues, the problem of efficiency for heteroskedasticity tests is a relatively minor issue that primarily affects those making dichotomous decisions based on P-values or confidence intervals from heteroskedasticity tests. If researchers avoid making dichotomous decisions regarding the presence or absence of TRH, the seriousness of this issue is diminished.
Summary of the Issues
While the SDIR approach represents an advancement over simple so-called responder counts, it has several important limitations that researchers should consider that are highlighted in Table 1. First, the fundamental problem of causal inference means we cannot identify the correlation between potential outcomes, making it impossible to directly estimate the true standard deviation of individual treatment effects (σD). The SDIR makes a strong assumption concerning the correlation between treatment and control outcomes; when this assumption is not thoughtfully justified, bounds would be more appropriate, even as a sensitivity analysis. Second, the assumption that TRH only exists when treatment group variance exceeds control group variance is overly restrictive. Heterogeneous treatment effects can exist even when variances are equal between groups or when the control group has greater variance. Third, the SDIR lacks statistical efficiency compared to alternative methods like the variance ratio test, potentially leading to more underpowered analyses of TRH, especially in studies with modest sample sizes. Also, researchers should also be cautious in their interpretation of heteroskedasticity tests and how they relate to the presence or absence of TRH.
Table 1.
Summary of Issues with the Standard Deviation of Individual Response
| Claim | Problem | Solution |
|---|---|---|
| “[SDIR] is considered a parameter for the distribution of true responses in the population of interest alongside the mean treatment effect” (Atkinson et al, 2019) | • Makes a strong, rarely justified assumption concerning the correlation between potential outcomes. • We cannot (empirically) identify this correlation, making the assumption untestable. |
• Present bounds for the SDD (Gadbury et al, 2001) • These bounds could also be presented alongside response distributions and bounds for the probability of harm or benefit |
| “In a parallel group study, true individual differences in exercise response are present only if the SD of change is substantially larger in the exercise group than the control group” (Atkinson et al, 2018) | • TRH can exist when the control group shows greater variance or even when the group variances are equal. | • Ensure any comparisons of the variances are two-sided • Do not necessarily disregard the potential of TRH if the control group has greater or equal variance |
| SDIR can be utilized to detect TRH by comparing variances between treatment groups | • SDIR is an inefficient method • Differences in variances are only a sufficient and not necessary condition to declare that there is non-zero TRH. |
• Use other tests or estimates of heteroskedasticity; See Mills et al. (2021) for possible suggestions • Do not conclude or claim that there is no presence of TRH based solely on a non-significant test of heteroskedasticity |
It is important to recognize that traditional parallel group experimental designs, while efficacious for estimating average treatment effects, are not designed to investigate TRH. These designs were primarily developed to answer population-level questions about mean effects, not individual-level questions about response variability. Consequently, researchers should be cautious about making definitive claims regarding the presence or absence of TRH based solely on parallel-group trials a heteroskedasticity test. The absence of statistically significant heteroskedasticity test as evidence for TRH should not be interpreted as definitive evidence that no meaningful TRH exists (only that there is a lack of evidence of TRH). Alternative designs that expose individuals to both treatment and control conditions, such as replicate crossover designs and variations on the Balaam design, may be better suited for investigating true individual differences in treatment response (Senn et al., 2011; Zoh et al., 2023).
An Example of an Appropriate TRH Analysis
Now, let us inspect a previously published study by Plotkin et al. (2022) as an example of how to appropriately discuss and display TRH from a parallel groups trial. This study aimed to assess how muscle size and strength adaptations differ between two exercise prescription patterns, which aimed to progressively increase load or repetitions (LOAD and REPS, respectively). The SDIR approach becomes particularly problematic when applied to comparative effectiveness studies such as these. In such designs, the fundamental assumption that the differential effects of the interventions can be represented by an orthogonal variance component is tenuous. Moreover, in studies comparing two exercise modalities (i.e., LOAD vs. REPS), designating one intervention as the “control” is arbitrary, as the SDIR assumptions seem most valid in cases when there is a negative control (e.g., a placebo condition). Notably, proponents of the SDIR approach have primarily focused their methodological discussions on classic intervention-vs-(negative) control designs, leaving a significant gap in guidance for comparative effectiveness research. Below, we outline an analysis that could be applied to studies parallel group trials with or without a negative “control” group.
An analysis of covariance (ANCOVA) can be utilized to estimate the average treatment effect (). In the case of Plotkin et al. (2022), we estimate the average treatment effect while adjusting for the pre-intervention 1 repetition maximum (1-RM) back squat and the biological sex of the participant. This indicated that the 2.03 kg improvement in LOAD over REPS, 95% C.I. [−3.99, 8.05].
The change score standard deviation estimates are 7.65 and 12.24 for the REPS and LOAD groups, respectively. Heteroskedasticity tests yield P-values ranging from 0.0506 to 0.0775 (see Supplementary Materials). While none of these tests meet the traditional threshold for “statistical significance”, notably, the standard deviation in the LOAD group is 1.6 times greater than that of REPS, which should (but alone is not needed to) prompt greater scrutiny of TRH. Because the study did not have a negative control, the labeling of which group is a control is arbitrary—the direction of heteroskedasticity (i.e., which treatment group has greater variance) should always be ignored in such trials.
From the observed change scores, the bounds of SDD point estimates range from 4.6 to 19.9 kg, with considerably wide confidence intervals for these point estimates (Figure 1a). Based on these point estimates and normality assumptions, we can graphically depict the expected treatment response distribution as a function of the assumed correlation (ρLOAD,REPS). From these individual treatment effect distributions, the probability of harm from the REPS treatment can be estimated. The bounds of Pharmed ranged from 54% to 67% with, again, considerable variability in these estimates due to the sample size (Figure 1c). Using this information, we could then add that given the average treatment effect, and our estimates of SDD, a sizeable number of individuals could have a net negative effect of using the REPS training compared to LOAD. Finally, because there is appreciable uncertainty in SDD, we can depict how this uncertainty propagates into the expected distributions of treatment effects at certain values of ρLOAD,REPS (Figure 1d). Illustrating treatment response distributions provides readers with a clear and transparent picture of the implications of TRH across a range of assumptions that one can make, while SDIR provides an estimate of TRH with a strict assumption concerning the value of ρLOAD,REPS (vertical dot-dash green lines).
Figure 1.

Relationship between the correlation of potential outcomes (ρLOAD,REPS) and estimated treatment response heterogeneity (TRH). (A) TRH can be easily understood as the standard deviation of the treatment effect (SDD), which varies as a function of the assumed correlation of potential outcomes. In the case of Plotkin et al. (2022), the assumed correlation has drastic effects on the expected TRH, and importantly, there is much uncertainty in these estimates (black dotted line = point estimate; grey ribbon = 95% CI; red line = no effect). (B) Using the point estimates of SDD, the expected distributions of treatment responses can be depicted (color = percentile and 1-percentile of the treatment response). This plot makes it apparent that, under certain assumptions, TRH dwarfs the average treatment effect. (C) Based on the treatment response distributions in panel 1b, the probability of harm (Pharmed) can be calculated (Gadbury et al., 2001). Note that there is considerable uncertainty associated with these point estimates, which incorporates the uncertainty of the intervention effect and the SDD. (D) Since there is appreciable uncertainty in the estimation of SDD, this can be incorporated into depictions of the response distributions. Whereas panel B depicts response distributions that use the point estimates of SDD, panel D shows distributions for the point estimate (green) and 95% CI of SDD (dark blue to yellow) for specific values of ρLOAD,REPS. In every panel, the dot-dash style green vertical line represents the value of ρLOAD,REPS assumed by the SDIR.
The authors could then conclude that given the small effect size, especially relative to our different sources of variance, neither treatment is likely to confer much benefit or harm compared to the other. However, given the potential bounds of SDD, the authors may consider the discussion of TRH worthy of further inspection. Potential interactions could be explored, and covariates could be utilized to tighten the bounds on SDD (Gadbury et al, 2001; Zoh et al, 2023). The authors could also provide some speculations concerning unmeasured factors influencing TRH and propose study designs that could address these factors in the future (Zoh et al, 2023). Finally, researchers should avoid using TRH as a default explanation when results are ambiguous or fail to show “statistically significant” effects. Attributing such results to TRH without theoretical justification is not a rigorous interpretation of parallel groups trials.
Conclusions
The investigation of TRH represents an important frontier in understanding how interventions affect different individuals, particularly as the field moves toward personalized medicine. While the SDIR method improved upon other approaches, it makes strong assumptions that restrict its utility for estimating true TRH in many parallel group trials, especially comparative effectiveness trials which are ubiquitous in exercise physiology. Most importantly, the SDIR approach cannot resolve the fundamental problem of causal inference, potentially making erroneous assumptions when attempting to reflect true individual treatment effect variance. This approach’s strict requirements might only be viable under specific circumstances (detailed in Supplementary Materials)—researchers continuing to use this method should therefore justify how their trial satisfies these necessary conditions.
As Rothman, Greenland, and Walker (1980) astutely observed, disagreements over TRH often stem from failure to separate distinct contexts in which heterogeneity is evaluated: statistical, biological, public health, and individual decision-making. Each context has different implications for how treatment heterogeneity should be assessed and discussed. The SDIR approach primarily operates in the statistical context and the other three contexts that are crucial for meaningful translation of research findings.
Rather than relying solely on SDIR, we recommend that researchers think carefully about their goals and employ analyses that address those goals:
If researchers want to establish the presence of TRH, then evidence of the sufficient condition of heteroskedasticity should be enough. Although such tests are sufficient to evidence the presence of TRH, they cannot rule it out, even if they have a very precise estimate that indicates homoskedasticity.
If researchers are interested in estimating the degree of TRH, they are left with SDIR and SDD. SDIR makes strong assumptions about ρTC, which, in many cases, are tenuous. Thus, the safest option is present SDD with the relevant graphical depictions, and perhaps the authors can comment on which ρTC may be more or less reasonable depending on their experiment and theory
Estimate bounds for the probability of harm or benefit or from treatments, which provides more actionable information for clinical decision-making and better aligns with what Rothman et al. identified as essential for both public health and individual contexts
When a salient degree of heterogeneity is detected, carefully investigate potential mechanisms and moderators rather than simply attributing results to undefined “individual differences”
Importantly, researchers should avoid using TRH as a default explanation for “non-significant” differences without a theoretical justification. Asserting that TRH is an explanation, despite negligible average effects, necessarily implies “harm” to certain subgroups—an assumption that warrants scrutiny. The field of exercise physiology would benefit from moving beyond simply detecting heterogeneity and toward understanding its underlying causes and implications for treatment decisions (Zoh et al., 2023). Of particular concern is the widespread practice of classifying individuals as “responders” and “non-responders” based on observed changes, which represents what Senn termed “dichotomania”—the problematic transmutation of continuous outcome measures into binary classifications (Senn 2005). This practice can be misleading due to the fundamental problem of causal inference that we described above and substantial losses in statistical efficiency and interpretability. Rather than speciously pursuing responder identification in small studies, researchers should prioritize collecting large, diverse samples to detect meaningful moderators of treatment effects. Additionally, modern advances in experimental design offer many improvements over the parallel group design for the identification of TRH (Robinson et al, 2024), exemplified by innovative approaches such as replicated within-participant unilateral trials that can more effectively separate true individual differences from measurement error and biological variability (Robinson et al, 2025). This approach shift is imperative to avoid perpetuating analyses that may reflect statistical artifacts rather than true individual variation, and in turn, it will facilitate genuine personalized medicine and exercise prescriptions.
These recommendations aim to advance more rigorous investigation of TRH while acknowledging the inherent limitations in parallel group designs. By adopting more nuanced analytical approaches that explicitly address statistical, biological, public health, and individual decision-making contexts, and by maintaining high standards for causal inference, researchers can better contribute to the development of truly personalized treatment approaches.
Supplementary Material
Acknowledgements
We would like to thank Dan Beavers and Stephen Senn for their feedback in our discussions about treatment response heterogeneity and their feedback on early drafts of this manuscript.
Disclosures
Collectively, the authors and their institutions have received payments for consultations, grants, contracts, in-kind donations, and contributions from multiple for-profit and not-for-profit entities interested in statistical design and analysis of experiments but not directly related to the methodological questions addressed in the present paper.
Data and Supplemental Material Accessibility
The data and code to reproduce the results and figures presented in our manuscript can be accessed at our online repository: https://doi.org/10.5281/zenodo.15492100
References
- Atkinson G, & Batterham AM (2015). True and false interindividual differences in the physiological response to an intervention. Experimental Physiology, 100(6), 577–588. doi: 10.1113/EP085070 [DOI] [PubMed] [Google Scholar]
- Atkinson G, Williamson P, & Batterham AM (2018). Exercise training response heterogeneity: statistical insights. Diabetologia, 61(2), 496–497. doi: 10.1007/s00125-017-4501-2 [DOI] [PubMed] [Google Scholar]
- Atkinson G, Williamson P, & Batterham AM (2019). Issues in the determination of ‘responders’ and ‘non-responders’ in physiological research. Experimental Physiology, 104(8), 1215–1225. doi: 10.1113/EP087712 [DOI] [PubMed] [Google Scholar]
- Bonafiglia JT, Preobrazenski N, & Gurd BJ (2021). A systematic review examining the approaches used to estimate interindividual differences in trainability and classify individual responses to exercise training. Frontiers in Physiology, 12, 665044. doi: 10.3389/fphys.2021.665044 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cortés J, González JA, Medina MN, Vogler M, Vilaró M, Elmore M, … Cobo E (2019). Does evidence support the high expectations placed in precision medicine? A bibliographic review. F1000Research, 7, 30. doi: 10.12688/f1000research.13490.4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fisher RA (1990). Statistical inference and analysis: Selected correspondence of RA Fisher, edited by Bennett JH. pages: 63–64. [Google Scholar]
- Gadbury GL, & Iyer HK (2000). Unit-treatment interaction and its practical consequences. Biometrics, 56(3), 882–885. doi: 10.1111/j.0006-341x.2000.00882.x [DOI] [PubMed] [Google Scholar]
- Gadbury GL, Iyer HK, & Allison DB (2001). Evaluating subject-treatment interaction when comparing two treatments. Journal of Biopharmaceutical Statistics, 11(4), 313–333. doi: 10.1081/bip-120008851 [DOI] [PubMed] [Google Scholar]
- Hecksteden A, Kraushaar J, Scharhag-Rosenberger F, Theisen D, Senn S, & Meyer T (2015). Individual response to exercise training - a statistical perspective. Journal of Applied Physiology (Bethesda, Md.: 1985), 118(12), 1450–1459. doi: 10.1152/japplphysiol.00714.2014 [DOI] [PubMed] [Google Scholar]
- Holland PW (1986). Statistics and Causal Inference. Journal of the American Statistical Association, 81(396), 945. doi: 10.2307/2289064 [DOI] [PubMed] [Google Scholar]
- Hopkins WG (2015). Individual responses made easy. Journal of Applied Physiology (Bethesda, Md.: 1985), 118(12), 1444–1446. doi: 10.1152/japplphysiol.00098.2015 [DOI] [PubMed] [Google Scholar]
- Laird N (1983). Further comparative analyses of pretest-posttest research designs. The American Statistician, 37(4a), 329–330. doi: 10.1080/00031305.1983.10483133 [DOI] [Google Scholar]
- Mills HL, Higgins JPT, Morris RW, Kessler D, Heron J, Wiles N, … Tilling K (2021). Detecting heterogeneity of intervention effects using analysis and meta-analysis of differences in variance between trial arms. Epidemiology (Cambridge, Mass.), 32(6), 846–854. doi: 10.1097/EDE.0000000000001401 [DOI] [PMC free article] [PubMed] [Google Scholar]
- PASS 2024 Power Analysis and Sample Size Software (2024). NCSS, LLC. Kaysville, Utah, USA, ncss.com/software/pass. Version 24.0.1. [Google Scholar]
- Plotkin D, Coleman M, Van Every D, Maldonado J, Oberlin D, Israetel M, … Schoenfeld BJ (2022). Progressive overload without progressing load? The effects of load or repetition progression on muscular adaptations. PeerJ, 10, e14142. doi: 10.7717/peerj.14142 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Poulson RS, Gadbury GL, & Allison DB (2012). Treatment heterogeneity and individual qualitative interaction. The American Statistician, 66(1), 16–24. doi: 10.1080/00031305.2012.671724 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Renwick JRM, Preobrazenski N, Wu Z, Khansari A, LeBouedec MA, Nuttall JMG, … Gurd BJ (2024). Standard deviation of individual response for VO2max following exercise interventions: A systematic review and meta-analysis. Sports Medicine (Auckland, N.Z.), 54(12), 3069–3080. doi: 10.1007/s40279-024-02089-y [DOI] [PubMed] [Google Scholar]
- Robinson ZP, Helms ER, Trexler ET, Steele J, Hall ME, Huang C-J, & Zourdos MC (2024). N of 1: Optimizing Methodology for the Detection of Individual Response Variation in Resistance Training. Sports Medicine, 54(8), 1979–1990. doi: 10.1007/s40279-024-02050-z [DOI] [PubMed] [Google Scholar]
- Robinson ZP, Steele J, Helms ER, Trexler ET, Hall ME, Huang C-J, Pelland JC, Remmert JF, Hinson SR, Mikula SA, Hamaïde AA, & Zourdos MC (2025). The Effect of Resistance Training Volume on Individual-Level Skeletal Muscle Adaptations: A Novel Replicated Within-Participant Unilateral Trial. Cold Spring Harbor Laboratory. BioRxiv. doi: 10.1101/2025.07.24.666533 [DOI] [Google Scholar]
- Senn S (2004). Controversies concerning randomization and additivity in clinical trials. Statistics in Medicine, 23(24), 3729–3753. doi: 10.1002/sim.2074 [DOI] [PubMed] [Google Scholar]
- Senn S (2005). Dichotomania: an obsessive compulsive disorder that is badly affecting the quality of analysis of pharmaceutical trials. Proceedings of the International Statistical Institute, 55th Session, Sydney. [Google Scholar]
- Senn S, Rolfe K, & Julious SA (2011). Investigating variability in patient response to treatment--a case study from a replicate cross-over study. Statistical Methods in Medical Research, 20(6), 657–666. doi: 10.1177/0962280210379174 [DOI] [PubMed] [Google Scholar]
- Senn S (2016). Mastering variation: variance components and personalised medicine. Statistics in Medicine, 35(7), 966–977. doi: 10.1002/sim.6739 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Swinton PA, Hemingway BS, Saunders B, Gualano B, & Dolan E (2018). A statistical framework to interpret individual response to intervention: Paving the way for personalized nutrition and exercise prescription. Frontiers in Nutrition, 5. doi: 10.3389/fnut.2018.00041 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Swinton P (2023). Assessing individual response to training in sport and exercise. doi: 10.51224/srxiv.288 [DOI] [Google Scholar]
- Zoh RS, Esteves BH, Yu X, Fairchild AJ, Vazquez AI, Chapple AG, … Allison DB (2023). Design, analysis, and interpretation of treatment response heterogeneity in personalized nutrition and obesity treatment research. Obesity Reviews: An Official Journal of the International Association for the Study of Obesity, 24(12), e13635. doi: 10.1111/obr.13635 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data and code to reproduce the results and figures presented in our manuscript can be accessed at our online repository: https://doi.org/10.5281/zenodo.15492100
