Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 Sep 27.
Published in final edited form as: Clin Trials. 2023 Apr 24;20(4):341–350. doi: 10.1177/17407745231169692

Finding the (biomarker-defined) subgroup of patients who benefit from a novel therapy: No time for a game of hide and seek

Lisa Meier McShane 1, Mark D Rothmann 2, Thomas R Fleming 3
PMCID: PMC10523858  NIHMSID: NIHMS1887716  PMID: 37095696

Abstract

An important element of precision medicine is the ability to identify, for a specific therapy, those patients for whom benefits of that therapy meaningfully exceed the risks. To achieve this goal, treatment effect usually is examined across subgroups defined by a variety of factors, including demographic, clinical, or pathologic characteristics or by molecular attributes of patients or their disease. Frequently such subgroups are defined by the measurement of biomarkers. Even though such examination is necessary when pursuing this goal, the evaluation of treatment effect across a variety of subgroups is statistically fraught due to both the danger of inflated false-positive error rate from multiple testing and the inherent insensitivity to how treatment effects differ across subgroups.

Pre-specification of subgroup analyses with appropriate control of false-positive (i.e. type I) error is recommended when possible. However, when subgroups are specified by biomarkers, which could be measured by different assays and might lack established interpretation criteria, such as cut-offs, it might not be possible to fully specify those subgroups at the time a new therapy is ready for definitive evaluation in a Phase 3 trial. In these situations, further refinement and evaluation of treatment effect in biomarker-defined subgroups might have to take place within the trial. A common scenario is that evidence suggests that treatment effect is a monotone function of a biomarker value, but optimal cut-offs for therapy decisions are not known. In this setting, hierarchical testing strategies are widely used, where testing is first conducted in a particular biomarker-positive subgroup and then is conducted in the expanded pool of biomarker-positive and biomarker-negative patients, with control for multiple testing. A serious limitation of this approach is the logical inconsistency of excluding the biomarker-negatives when evaluating effects in the biomarker-positives, yet allowing the biomarker-positives to drive the assessment of whether a conclusion of benefit could be extrapolated to the biomarker-negative subgroup.

Examples from oncology and cardiology are described to illustrate the challenges and pitfalls. Recommendations are provided for statistically valid and logically consistent subgroup testing in these scenarios as alternatives to reliance on hierarchical testing alone, and approaches for exploratory assessment of continuous biomarkers as treatment effect modifiers are discussed.

Keywords: Subgroup analyses, hierarchical testing, precision medicine, personalized medicine, predictive biomarker, treatment effect modifier

Introduction

Impressive progress has been made through recognition that benefits and risks for medical interventions often vary meaningfully across patient subgroups. With this important opportunity to enhance clinical care through pursuit of personalized medicine comes inherent scientific challenges resulting from the high level of multiplicity when there is need, not only to determine proper dose and frequency for a regimen, the optimal trial design including outcome measure and analysis method and selection of control groups but also to identify the subgroup in which the intervention is likely to have a favorable benefit-to-risk balance. Traditionally, clinical trials have been designed primarily to test whether the average treatment effect over the population of interest is positive, but advances in personalized medicine have necessitated shifting focus away from average treatment effect.

Many characteristics putatively defining the subgroup of patients benefiting (or potentially harmed) from an investigational intervention can be assessed by the measurement of biomarkers, which would be considered predictive according to the BEST (Biomarkers, EndpointS, and other Tools) Resource glossary.1 For simplicity, discussion here will focus on therapeutic interventions, though the principles apply more broadly, including for prevention. Candidate predictive biomarkers might be suggested based on the mechanism of action of a new therapeutic agent (e.g. a molecular target) or based on evidence from pre-clinical or early phase clinical trials that conducted exploratory biomarker analyses. However, preliminary data may be insufficient to confidently define restricted eligibility for a definitive Phase 3 trial evaluating the experimental therapeutic, and additional biomarker evaluation occurs within Phase 3 trials.

The dual goals of establishing whether there is a treatment benefit overall and determining in which subgroup that treatment effect is present and clinically meaningful require attention to statistical control of type I error. Subgroups could be defined by patient demographic, clinical, or pathologic characteristics (e.g. age, sex, disease subtype, or stage) or by molecular attributes of patients (e.g. germline polymorphisms in drug-metabolizing genes) or their disease (e.g. tumor biomarkers). It is critical that appropriate statistical procedures be in place to reduce the chance of false-positive findings.

Motivating examples from oncology and cardiology will be discussed with particular attention to the pitfalls of certain commonly used hierarchical or nested testing procedures. Such testing may be encountered when analyses aimed at defining the patient population benefiting from a new therapy search for an “optimal” cut-off on a continuous biomarker or when analyses examine subgroups defined by cross-classification of multiple factors. Guidance is provided for how to perform statistically valid and logically consistent subgroup testing in these scenarios, and approaches for exploratory assessment of continuous biomarkers as treatment effect modifiers are discussed.

Biomarker challenges

An optimal treatment-selection biomarker test might not be ready

Many challenges are encountered in the pursuit of predictive biomarkers, single or complex, to identify a patient subgroup likely to receive clinically meaningful benefit from a new therapy (“biomarker-positive”), or a patient subgroup not benefited or even harmed (“biomarker-negative”). General understanding of disease biology or mechanism of action of a therapy may be insufficient to confidently identify a specific predictive biomarker or complex biomarker signature or profile without collecting more empirical evidence associating biomarkers with therapy benefit. Multiple assays all putatively measuring the same biomarker might exist, yet results might differ. Further complications include uncertainty about optimal cut-offs to delineate positive from negative based on a continuous biomarker, or many ways to segregate patients based on a multivariable profile of biomarkers. Refinement of biomarker-based classifications can be data-intensive. Due to these many challenges of developing accurate and reliable predictive biomarkers, biomarker assay development for therapy selection is often not well aligned with development of the associated drug, leaving uncertainty about predictive biomarkers when the drug is ready for Phase 3 evaluation.

Extracting a candidate biomarker from the literature or preliminary data

Gleaning candidate predictive biomarkers from the literature or preliminary data can be a frustrating experience. Not uncommonly, accumulated evidence has been derived using different assays for putatively the same biomarker that may yield different results and on different patient populations. Distribution of a biomarker and its association with treatment benefit might differ according to patient or disease characteristics, for example, sex, age, stage, and disease subtype. These scenarios substantially complicate interpretation of existing evidence to support choice of predictive biomarkers for inclusion in clinical trials in later drug development stages.

Ki67 proliferation biomarker example.

An example of a single biomarker exhibiting lack of reproducibility is the tumor proliferation biomarker Ki67. A study by Polley et al.2 demonstrated considerable variability in measurements on a common set of 100 breast tumors across eight laboratories experienced in Ki67 assessment using their own in-house immunohistochemical assays. Intraclass correlation was 59%, indicating that more than 40% of the total variation in Ki67 scores was due to intra- and inter-laboratory variability rather than true biological differences across tumors. This lack of reproducibility for Ki67 assessments has been noted as a concern3 in the context of a recent approval for the drug abemaciclib with an indication restricted to tumors with Ki6720%.4 Although a companion diagnostic for Ki67 was approved with the drug, pathology laboratories that opt to use their own in-house Ki67 might not identify the same population of patients on which drug approval was based, resulting in unknown impact on drug efficacy.

Tumor mutational burden complex biomarker example.

A notable example of a complex biomarker that has demonstrated variability in measurement across different assays is tumor mutational burden (TMB), which is defined as the number of somatic nonsynonymous mutations normalized to the total number of megabases of genome sequenced and is of interest for the identification of patients with cancer most likely to respond to immunotherapy. Differences in bioinformatic algorithms for calling mutations, and regions of genome sequenced, can introduce substantial variation in TMB measurements.5,6 Furthermore, exploratory studies have suggested that distribution of TMB and associated cut-offs for segregating patients into those whose tumors respond versus not respond to immunotherapy may differ across tumor types.7 Despite an FDA tumor agnostic approval of the immunotherapy drug pembrolizumab with FoundationOne CDx companion diagnostic assay for TMB,8 questions remain about whether the association between TMB and response may also be influenced by the particular TMB assay, patient population, or immunotherapy drug administered. A uniform TMB cut-off of 10 mutations per megabase might not be optimal in all situations;9,10 therefore, use of TMB in a new clinical trial in a more focused tumor type or clinical setting (e.g. treatmentnaïve versus progressed on prior immunotherapy) might warrant further evaluation of its association with treatment benefit within the trial.

Statistical design and analysis consideration for prospective biomarker-based subgroup testing

Examples just discussed illustrate that further refinement of biomarker-based subgroups identifying which patients are most likely to benefit from a new therapy will often be needed, even within the context of a Phase 3 treatment trial. Proper statistical design and analysis strategies are essential in these situations to address the lingering biomarker questions while avoiding misleading or false claims about treatment benefit.

Clinical trial enrichment

The first decision when a biomarker is thought possibly associated with likelihood of benefit from an investigational therapy is whether to restrict trial eligibility to only biomarker-positive patients (e.g. biomarker value above a threshold). Such a strategy would constitute a type of trial enrichment, defined by the FDA11 as “prospective use of any patient characteristic to select a study population in which detection of a drug effect (if one is in fact present) is more likely than it would be in an unselected population.” Specifically, the biomarker would be used for predictive enrichment.

The previous examples illustrate challenges in optimally defining the biomarker-positive subgroup. Even when higher values of a biomarker are thought to be associated with greater treatment benefit, it may be unclear where to “cut” the biomarker for the purposes of trial eligibility determination. Eligibility might cast a wide net (e.g. lower the biomarker cut-off for defining positivity) to improve chances that no patients potentially benefiting from investigational therapy would be excluded from trial participation. Because patients whose biomarkers fall below the cut-off used for eligibility would not be included in the trial, only refinements of cut-offs above that used for trial eligibility will be possible based solely on the trial data. Residual uncertainty in biomarker cut-offs has led to several trials in which treatment effects were assessed within subgroups defined by more than one cut-off applied to the biomarker. Similar issues arise for complex biomarkers, such as those that are declared positive when any of a several individual characteristics are deemed present. An example is the complex biomarker homologous recombination deficiency (HRD) which may be considered positive when function-altering mutations in any of several genes in the homologous recombination repair pathway are detected, yet different HRD assays may include different genes or other markers of downstream consequences of repair defects.12 These complexities require careful management of subgroup analyses in definitive clinical trial settings.

Statistical management of multiple subgroups in a definitive clinical trial

Pervasive incentives to achieve “positive” results in clinical research must be counterbalanced by rigor and transparency in design, conduct, analysis, and reporting of research. An example of an approach frequently used in pursuit of “positive” results has been the conduct of multiple analyses in clinical trials, including the assessment of treatment effects across a wide array of subgroups. While those subgroup analyses would be informative if conducted descriptively, as through graphical forest plots, it would be problematic to view such results as providing reliable conclusions.

When there is an intention for assessments of efficacy to yield reliable results, that is, to be confirmatory rather than exploratory, then pre-specification is key. This includes situations where there is interest in achieving definitive conclusions about treatment effects within subgroups.

Extending efficacy assessments to the biomarker-negative subgroup: Multiplicity adjustments are necessary, but a hierarchical testing approach is not sufficient.

When assessing effects within subgroups, a classic approach to addressing multiplicity would be to use a hierarchical approach, where the false-positive alpha would be spent initially on an analysis within the biomarker-positive subgroup where benefits could be particularly favorable. If statistically significant, benefit would be established in these biomarker-positive patients; then, hierarchically, a full alpha-level test would be conducted in the pooled cohort of biomarker-positive and biomarker-negative patients to determine whether positivity could be extended to the full ITT population.

Although this hierarchical approach achieves the necessary objective of controlling the experimental error rate for the two hypotheses (treatment effect in biomarker-positive and full groups, respectively), it is flawed reasoning to argue that this approach justifies extending the conclusion of benefit to the biomarker-negative subgroup as well. In particular, it is logically inconsistent to pre-specify exclusion of biomarker-negative patients when assessing treatment effect in the biomarker-positive patients, presumably due to an expectation of attenuated benefit in the biomarker-negative patients, yet then allow the inclusion of the biomarker-positive patients in the primary analysis intended to determine whether there is persuasive evidence that treatment effects extend to the biomarker-negative subgroup. This hierarchical strategy presents realistic potential that apparent statistically significant treatment benefit in the full population is driven exclusively by the biomarker-positive patients.13 Several authors have described examples of trials in which hierarchical subgroup testing such as this was performed and highlighted the flaws in this approach.14,15 An additional example is provided here along with the recommendations for an alternative to hierarchical testing alone.

Illustration: CheckMate 649.

In the Checkmate 649 trial16 (NCT02872116), nivolumab plus chemotherapy was compared with chemotherapy alone in advanced gastric, gastro-esophageal junction and esophageal adenocarcinoma. Two-sided significance levels of 0.03 for overall survival (OS) and 0.02 for progression-free survival (PFS) were implemented. Testing was done hierarchically with alpha spent initially on the PD-L1 combined positive score (CPS)5 subgroup, and then both in the PD-L1 CPS1 subgroup and in the ITT cohort of all randomized patients.

In the trial’s primary results publication,16 treatment effects were reported for the CPS5, CPS1, and ITT cohorts. Approximately 60% of patients had CPS5, and they contributed about 60% of OS and PFS events. Patients with CPS1 comprised slightly more than 80% of the patients and also contributed slightly more than 80% of the OS and PFS events. For OS, the estimated hazard ratio (HR) was 0.71, (98.4% CI: 0.59, 0.86; 2p < 0.0001), in the CPS5 biomarker-positive patients and, hierarchically, the estimated HR was 0.77, (99.3% CI: 0.64, 0.92; 2p < 0.0001), in the CPS1 subgroup, and was 0.80, (99.3% CI: 0.68, 0.94; 2p < 0.0002), in the ITT cohort of all randomized patients. (Note, these 98.4% and 99.3% CIs used in the publication reflect adjustments for group-sequential analyses and the dual endpoints and subgrouping by CPS levels). Similar positive results were obtained for PFS: the estimated HR was 0.68, (98% CI: 0.56, 0.81; 2p < 0.0001), in the CPS5 biomarker-positive patients and, in turn, the HR was 0.74, (nominal 2p < 0.0001), in the CPS1 subgroup and was 0.77, (nominal 2p < 0.0001), in the ITT cohort of all randomized patients.

It would be inappropriate to interpret these hierarchical analyses as indicating that nivolumab’s strong positive effect on OS and PFS in the biomarker-positive CPS5 patients (approximately 60% of patients and 60% of OS and PFS events) could be generalized to the biomarker-negative CPS < 5 cohort. In the publication, it was indicated that the addition of nivolumab in CPS < 5 patients, when data were analyzed separately, provided minimal benefit, with an HR of 0.94 for OS corresponding to an approximate 3-week difference in OS, and with an HR of 0.93 for PFS, corresponding to an approximate 2-week difference in PFS.

A preferred alternative to hierarchical testing alone to address efficacy in a biomarker-negative subgroup: the “Rothmann Criteria.”.

In the analyses of the Checkmate 649 trial, there is evidence suggesting the experimental agent is not harmful in the biomarker-negative CPS < 5 patients in those clinical settings. However, to include patient populations in the label for an approved agent, a higher standard of having substantial evidence of efficacy and safety should be applied. Rothmann et al.13 proposed a higher standard to be met to justify the extrapolation of efficacy to the biomarker-negative subgroup, defining four key criteria specified in Table 1.

Table 1.

Additional criteria to justify extrapolation of efficacy.

Suppose a trial is designed first to assess effects in a biomarker-positive subgroup and then, hierarchically, to assess whether the conclusion of benefit could be extrapolated to a broader cohort. Once statistical significance has been established in the biomarker positive patients, the following proposed criteria* should be simultaneously satisfied to justify extrapolation of efficacy to the biomarker-negative patients:
a. The effect on the primary endpoint should be statistically significant when evaluated in the pooled ITT cohort of biomarker-positive and biomarker-negative patients; AND
b. There should be an adequate amount of data within the biomarker-negative subgroup to reliably estimate the level of effect in that subgroup; AND
c. The size of the estimated effect in the biomarker-negative subgroup should be clinically relevant; AND
d. The size of the estimated effect in the biomarker-negative subgroup should be at least as large as what would be needed to achieve “statistical significance” in an analysis conducted in the entire ITT population.

ITT: intent-to-treat.

*

See Rothmann et al. 13

In addition to the requirement, in Rothmann criterion (a), of achieving statistical significance in the pooled ITT cohort of biomarker-positive and biomarker-negative patients, the Rothmann criteria (b) and (c) require that evidence about treatment effects in the biomarker-negative patients be adequately reliable and that estimated effects be clinically meaningful. Assessing whether these criteria are met requires judgment, as in any clinical and regulatory assessment about an experimental intervention’s efficacy and benefit-to-risk. For example, thresholds for clinical meaningful benefit could be enlightened by formal determinations of minimal clinically important differences, by whether efficacy endpoints are direct measures about how a patient “feels, functions or survives,” by whether these endpoints represent irreversible morbidity/mortality outcomes, and by safety, convenience of administration and cost considerations. Regarding whether there is an adequate amount of data within the biomarker-negative subgroup, in many settings, this could require that the biomarker-negative subgroup provides at least approximately one-third of the overall information in the trial, yet this threshold would be influenced by safety considerations and the clinical relevance of estimated treatment effects. Finally, while Rothmann criterion (d) properly does not require achieving statistical significance in the biomarker-negative subgroup, it is well motivated since, if this criterion failed to hold, a treatment effect observed in the full trial cohort equal to that observed in the biomarker-negative patients would not have led to a positive result for the full cohort even in a setting where subgroup analyses had not been performed.

These “Rothmann criteria” in Table 1 can readily be assessed in the analyses of treatment effect in the biomarker-negative patients in CheckPoint 649. Criteria (a) and (b) were satisfied, recognizing that the biomarker-negative patients did provide 40% of the trial information about OS. However, criteria (c) and (d) were not satisfied. Regarding criterion (c), the estimated HR = 0.94 for OS corresponds to increasing the 12-month survival in the control arm in that subgroup by only a few weeks (clearly less than a clinically meaningful threshold). Finally, for the results in the Checkpoint 649 trial to have satisfied criterion (d), the estimated effect of the addition of nivolumab would need to be an HR0.85 for OS and an HR0.85 for PFS. While these are much more modest levels of effects than were documented on these endpoints in the biomarker-positive patients in this trial, this criterion (d) clearly was not achieved.

Importantly, Rothmann criterion (d) is less stringent than requiring statistical significance be achieved in the biomarker-negative patients. For illustration, consider a trial where the biomarker-positive and -negative subgroups each contribute 50% of the information. Modeling this illustration by some aspects of the Checkpoint 649 trial, suppose one expects the true HR for OS in the biomarker-positives to be 0.706 (i.e. 12 versus 17 months) and in the biomarker-negatives to be 0.810 (i.e. 12 versus 14.8 months). If the trial were designed to provide 90% power (testing with a 2.5% false-positive error rate) when the true HR is 0.706 in the biomarker-positives, then 347 OS events would be required in the biomarker-positives, and statistical significance would occur with an estimated HR = 0.810. Since the biomarker-negatives would have only 347 OS events, the power to achieve statistical significance in that subgroup would be only 50%. In contrast, the Rothmann criterion (d) would have required the estimated HR in the biomarker-negatives to be 0.862 which would be achieved with 72% probability.

Assessing treatment effects in relation to continuous biomarkers

Continuous biomarkers are often categorized for eligibility or to form subgroups so that multiple testing procedures can be applied in clinical trials, but often it is more plausible that treatment effects vary in some “smooth” fashion as a function of continuous biomarkers, either directly or indirectly through their association with other treatment effect modifiers. Exploratory analyses to more fully characterize the association between treatment effect and a continuous biomarker may be of interest. Discussed here are several approaches for evaluation of the relationship between a continuous biomarker (or fully specified multivariate score) and treatment effect in the setting of a time-to-event clinical endpoint.

Modeling treatment HR as a function of continuous biomarker

One approach to explore whether a continuous biomarker might be an effect modifier is to fit a model for the HR of the experimental treatment versus control (y-axis) and plot it against the biomarker value (x-axis). Typically, a log scale is used for the y-axis. Considering a simple Cox proportional hazards regression model of the form λ(t;z,x)=λ0(t)exp(β0z+β1x+β2zx), where z is an indicator variable for the experimental treatment and x is the continuous biomarker, it can be seen that the HR for the effect of the experimental treatment relative to control for a patient with biomarker value x relative to baseline hazard for a patient in the control group is given by exp(β0+β2x). Plotted on a log scale, this appears as a straight line. More complex models could be considered, for example, incorporating polynomial terms or splines. Pre-specification of the treatment effect modification curve may be difficult so the statistical plan might require flexibility for modeling.

Illustration: LVEF in PARADIGM-HF and PARAGON-HF Trials.

The biomarker left ventricular ejection fraction (LVEF) is known to associate with the risk of cardiovascular events. Figure 1 depicts HR plots for continuous LVEF in the PARADIGM-HF (NCT01035255) and PARAGON-HF (NCT01920711) trials illustrating also treatment effect modification. The PARADIGM-HF trial randomized 8442 adult patients with symptomatic chronic heart failure with LVEF < 40% to Entresto or enalapril17 and supported the approval of Entestro, a combination of sacubitril and valsartan, to reduce the risk of cardiovascular death and hospitalization for heart failure in patients with chronic heart failure with reduced ejection fraction.18 The PARAGON-HF trial randomized patients between Entresto and valsartan alone to support a claim for heart failure with preserved ejection fraction, defined as LVEF > 45%.18 In each trial, HRs become more favorable for Entresto, the lower the LVEF, and less favorable for larger values of LVEF. Lower HRs reflect greater relative benefit for Entresto compared to control.

Figure 1.

Figure 1.

Treatment effect for composite endpoint of time to first heart failure hospitalization or cardiovascular death by baseline LVEF in PARADIGM-HF and PARAGON-HF.

Treatment effect is represented as HR for Entresto relative to control arm in each trial as a function of percent LVEF. Estimated curve is based on a Cox proportional hazards regression model that included a linear term for LVEF as a continuous variable, a binary treatment indicator, and an interaction between those two. Source: Entresto product label.18

Although the slope of the HR plot is positive in both PARAGON-HF and PARADIGM-HF, it is noteworthy that the slope appears greater in PARAGON-HF. Confidence bands on the HR are entirely below 1 in the PARADIGM-HF study. Interestingly, the estimated HR in PARAGON-HF crosses 1 at a baseline LVEF of about 64–65, and the upper confidence band crosses 1 at about 55–56 but the lower confidence band remains below 1. In its product label, Entresto “is indicated to reduce the risk of cardiovascular death and hospitalization for heart failure in adult patients with chronic heart failure. Benefits are most clearly evident in patients with LVEF below normal.”18

STEPP

Bonetti and Gelber19 proposed a graphical approach, Subpopulation Treatment Effect Pattern Plot (STEPP), to explore whether a continuous covariate might be a treatment effect modifier and developed inference methods to assess statistical significance of apparent treatment effect heterogeneity20 in the context of time-o-event outcomes. A STEPP displays treatment effect in moving windows of a continuous covariate, such as a biomarker. The treatment effect for a time-to-event outcome could be expressed in a variety of ways, including as an HR or difference in survival at a specific landmark timepoint. Yip et al.21 subsequently extended these methods to continuous, binary, and count outcomes.

Illustration: Ki-67 labeling index in the Breast International Group 1–98 trial.

Figure 2 displays STEPP examples constructed using data from the BIG (Breast International Group) 1–98 randomized clinical trial (NCT00004205), which included a randomization of postmenopausal women with hormone receptor-positive, early breast cancer to letrozole or tamoxifen adjuvant therapy.22,23 Treatment effect (letrozole relative to tamoxifen) is calculated in sliding windows of Ki-67 labeling index. Each subpopulation in sliding window contains approximately 150 patients, including 50 overlapping with adjacent windows. Figure 2(a) displays relative treatment effect as difference in 4-year % disease-free survival (DFS) and 2(b) displays treatment HR (DFS). A log scale on the y-axis for display of HRs is recommended for more natural interpretation, but copyright restrictions prevented updating of the originally published figure, which is reproduced in Figure 2(b). STEPP analyses suggest better outcome on letrozole compared to tamoxifen particularly for patients with higher Ki-67 values.

Figure 2.

Figure 2.

STEPP as a Function of Ki-67 Labeling Index in the BIG 1–98 randomized clinical trial.

Treatment effect of letrozole relative to tamoxifen in subpopulations of the BIG 1–98 randomized trial defined by a sliding window of Ki-67 labeling index (LI). Each subpopulation in sliding window contains approximately 150 patients, including 50 overlapping with adjacent windows. Dotted black horizontal line represents the null hypothesis that letrozole and tamoxifen have identical treatment effects overall and by Ki-67, and solid black horizontal line represents the estimated effect of letrozole relative to tamoxifen, under a model assuming no interaction by Ki-67. Relative treatment effect as a function of Ki-67 LI (solid blue curve) is represented as (a) difference in 4-year % DFS or (b) HR for DFS, with corresponding pointwise 95% confidence intervals (dashed blue curves). The p-values are from interaction test.

Source: Figures are reproduced from Lazar et al.23 ((a) and (b) from Figures 1B and 1C of Lazar, respectively), with permission.

Predictiveness curves

Predictiveness curves, an alternative to STEPP, plot survival percentage at a specified landmark time as a function of quantiles of the continuous biomarker, one curve per treatment arm.24,25 Curve modeling may be parametric or nonparametric. Clinical relevance should drive choice of survival endpoint and landmark time. Typically, such curves would be estimated from a randomized trial, involving the treatments of interest, in which a biomarker was measured for all patients. Figure 3 illustrates hypothetical predictiveness curves reflecting 5-year DFS according to biomarker percentile, by treatment arm. Optimal biomarker-guided therapy would aim to select therapy associated with the higher curve (greater survival percentage) according to biomarker value, assuming no substantial risks that would outweigh the benefits. Note that in a real data application, confidence bands around the curves should also be included to represent the uncertainty in the estimates. Under scenario (a), outcome is more favorable under experimental therapy compared to standard across the majority of biomarker values, but essentially equivalent for higher values. When treatments are equivalent, other factors such as toxicity, convenience, or cost might guide therapy selection; otherwise, the treatment associated with better outcome for the patient’s biomarker value would likely be chosen. Under scenario (c), all patients probably would be directed to experimental therapy; under scenario (d), outcomes are essentially equivalent for outcome on the two therapies. Under scenario (b), patients with low biomarker values likely would receive experimental therapy while those with higher values would receive standard therapy. Scenario (b) represents a clinically useful predictive biomarker.

Figure 3.

Figure 3.

Predictiveness curves illustrating example scenarios comparing outcome on different treatments as a function of biomarker percentile.

Outcome expressed as 5-year DFS as a function of biomarker percentile, separately for each treatment arm. The experimental treatment is the black line and the standard treatment is the gray line. (a) The experimental treatment results in 5-year DFS at least as good as for standard treatment across the range of biomarker values, although difference is minimal for the highest biomarker values; (b) experimental treatment yields better 5-year DFS for low biomarker values, but standard treatment is better at higher values; (c) experimental treatment is uniformly superior for 5-year DFS; (d) treatments yield roughly very similar 5-year DFS across the range of biomarker values. The biomarker that likely has the greatest potential utility for treatment selection is the one depicted in scenario B, assuming that its use still confers a favorable benefit-to-risk balance.

Janes et al.26 proposed calculation of a “theta” value to quantify population-level reduction in unfavorable events (according to endpoint of interest at landmark timepoint) by following the biomarker-guided optimal therapy rule, compared to all patients receiving standard therapy. A method to calculate sample size needed to estimate theta with a desired confidence interval width is available.27 The optimal rule and theta might depend on the chosen endpoint or landmark time. The value of a predictive biomarker is also influenced by the distribution of that biomarker in the population of interest. A biomarker that favors selecting one treatment over another only for very large biomarker values may have little utility in a population with biomarkers values heavily skewed low. It is advisable to consider various scenarios to capture different aspects of the diseases process and trajectory.

Discussion

An important element of precision medicine is the ability to identify, for a specific therapy, those patients for whom benefits of that therapy meaningfully exceed the risks. To accomplish this, treatment benefit needs to be examined across ever-smaller patient subgroups defined according to clinical, pathologic, or molecular attributes of patients or their disease. Competing with the need to find the benefiting subgroups is the danger of generating many spurious treatment effects if type 1 error is not adequately controlled. There is wide-spread interest in achieving “positive” results in clinical research, since such results would provide both financial and professional benefits to trial sponsors and more therapeutic options for caregivers to offer to patients. While these are legitimate interests, they create scientific conflicts-of-interest that could adversely impact research integrity if exploratory subgroup analyses are conducted in search for positive results, inducing substantial random high bias leading to overestimation of effects of interventions and misuse of p-values that readily are misinterpreted when the sampling context for their generation is not well-defined.28 Therefore, pre-specification is important to achieving results in subgroups that are both statistically persuasive and biologically plausible. In turn, rigorous control of type 1 error remains essential in the context of subgroup testing, in search of patients most likely to benefit from a treatment, that is claimed to be confirmatory.

Hierarchical testing strategies were called out here as particularly problematic, despite their popularity in situations where subgroups are defined by biomarkers suspected to associate with treatment benefit, such as positive or negative for certain mutations, or above or below a cut-off for a continuous biomarker (e.g. CheckMate 649 trial discussed above). Recommendations were made (Table 1) for logically consistent approaches for the examination of treatment effects in biomarker-positive and -negative subgroups that account for reduced samples size and power in the subgroups while maintaining some protection against spurious findings.

Not elaborated on here are the development and subsequent evaluation of complex scores, predictors, or classifiers intended to serve as treatment effect modifiers or define benefiting subgroups. These efforts, which are increasingly common in the era of big data and omics-driven precision medicine, are prone to even greater pitfalls than conventional subgroup testing when the subgroups or scores themselves are defined using the same data used to evaluate their predictive potential. Resampling approaches such as cross-validation or bootstrapping, or preferably independent datasets, are needed to assess predictive ability and avoid overfitting bias,2932 but they don’t obviate the need for control of type 1 error across many subgroups tested.

Discovery of new therapeutic strategies, potential for more detailed biological characterization of diseases, and greater understanding of therapy mechanisms will continue to advance the field of medicine toward personalized care. Statistical methods for appropriate analysis and interpretation of biomedical data will need to evolve in concert with this changing paradigm.

Funding

The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work (by TRF) was supported, in part, by the National Institutes of Health (NIH)/National Institute of Allergy and Infectious Diseases (NIAID) grant entitled “Statistical Issues in AIDS Research” (R37 AI 29168).

Footnotes

Disclaimer

This article reflects the personal views of the FDA and NIH authors and should not be construed to represent the views, policies, or guidance of the U.S. Food and Drug Administration or the NIH. Views expressed in written materials or publications and by speakers and moderators do not necessarily reflect the official policies of the Department of Health and Human Services nor does any mention of trade names, commercial practices, or organization imply endorsement by the US Government.

Declaration of conflicting interests

The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

References

  • 1.FDA-NIH Biomarker Working Group. BEST (Biomarkers, EndpointS, and other Tools) resource. Silver Spring, MD: Food and Drug Administration, 2016, https://www.ncbi.nlm.nih.gov/books/NBK338448/(2021, accessed 22 September 2022) [PubMed] [Google Scholar]
  • 2.Polley M-YC, Leung SCY, McShane LM, et al. An international Ki67 reproducibility study. J Natl Cancer Inst 2013; 105(24): 1897–1906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Dowsett M, Nielsen TO, Rimm DL, et al. Ki67 as a companion diagnostic: good or bad news? J Clin Oncol 2022; 40(33): 3796–3799. [DOI] [PubMed] [Google Scholar]
  • 4.Royce M, Osgood C, Mulkey F, et al. FDA approvalsummary: abemaciclib with endocrine therapy for high-risk early breast cancer. J Clin Oncol 2022; 40(11): 1155–1162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Merino DM, McShane LM, Fabrizio D, et al. Establishing guidelines to harmonize tumor mutational burden (TMB): in silico assessment of variation in TMB quantification across diagnostic platforms: phase I of the Friends of Cancer Research TMB harmonization project. J Immunother Cancer 2020; 8(1): e000147. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Vega DM, Yee LM, McShane LM, et al. Aligning TumorMutational Burden (TMB) quantification across diagnostic platforms: phase 2 of the Friends of Cancer Research TMB Harmonization Project. Ann Oncol 2021; 32(12): 1626–1636. [DOI] [PubMed] [Google Scholar]
  • 7.Samstein RM, Lee C-H, Shoushtari AN, et al. Tumormutational load predicts survival after immunotherapy across multiple cancer type. Nat Genet 2019; 51(2): 202–206. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.U.S. Food and Drug Administration. FDA approves pembrolizumab for adults and children with TMB-H solid tumors, 2020, https://www.Fda.gov/drugs/drug-approvals-and-databases/fda-approves-pembrolizumab-adults-and-children-tmb-h-solid-tumors (accessed 22 September 2022).
  • 9.Marabelle A, Fakih M, Lopez J, et al. Association oftumour mutational burden with outcomes in patients with advanced solid tumours treated with pembrolizumab: prospective biomarker analysis of the multicohort, open-label, phase 2 KEYNOTE-158 study. Lancet Oncol 2020; 21(10): 1353–1365. [DOI] [PubMed] [Google Scholar]
  • 10.Valero C, Lee M, Hoen D, et al. Response rates to anti–PD-1 immunotherapy in microsatellite-stable solid tumors with 10 or more mutations per megabase. JAMA Oncol 2021; 7(5): 739–743. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.U.S. Food and Drug Administration. Enrichment strategies for clinical trials to support approval of human drugs and biological products: guidance for industry, 2019, http://www.fda.gov/downloads/drugs/guidancecomplianceregu-latoryinformation/guidances/ucm332181.pdf (accessed 22 September 2022).
  • 12.Stewart MD, Merino Vega D, Arend RC, et al. Homologous recombination deficiency: concepts, definitions, and assays. Oncologist 2022; 27(3): 167–174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Rothmann MD, Zhang JJ, Lu L, et al. Testing in a prespecified subgroup and the intent-to-treat population. Drug Inf J 2012; 46(2): 175–179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Freidlin B and Korn EL. A problematic biomarker trialdesign. J Natl Cancer Inst 2022; 114(2): 187–190. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Kim MS and Prasad V. Nested and adjacent subgroups incancer clinical trials: when the best interests of companies and patients diverge. Eur J Cancer 2021; 155: 163–167. [DOI] [PubMed] [Google Scholar]
  • 16.Janjigian Y, Shitara K, Moehler M, et al. First-line nivo-lumab plus chemotherapy versus chemotherapy alone for advanced gastric, gastrooesophageal junction, and oesophageal adenocarcinoma (CheckMate 649): a randomised, open-label, phase 3 trial. Lancet 2021; 398(10294): 27–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Gandotra C, Clark J, Liu Q, et al. Heart failure population with therapeutic response to sacubitril/valsartan, spironolactone and candesartan: FDA perspective. Ther Innov Regul Sci 2022; 56(1): 4–7. [DOI] [PubMed] [Google Scholar]
  • 18.Entresto product label, https://www.accessdata.fda.gov/drugsatfda_docs/label/2021/207620s018lbl.pdf (2015, accessed 22 September 2022).
  • 19.Bonetti M and Gelber R. A graphical method to assesstreatment-covariate interactions using the Cox model on subsets of the data. Stat Med 2000; 19(19): 2595–2609. [DOI] [PubMed] [Google Scholar]
  • 20.Bonetti M and Gelber RD. Patterns of treatment effectsin subsets of patients in clinical trials. Biostatistics 2004; 5(3): 465–481. [DOI] [PubMed] [Google Scholar]
  • 21.Yip W-K, Bonetti M, Cole BF, et al. Subpopulationtreatment effect pattern plot (STEPP) analysis for continuous, binary, and count outcome. Clin Trials 2016; 13(4): 382–390. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Regan MM, Price KN, Giobbie-Hurder A, et al. Interpreting Breast International Group (BIG) 1–98: a randomized, double-blind, phase III trial comparing letrozole and tamoxifen as adjuvant endocrine therapy for postmenopausal women with hormone receptor-positive, early breast cancer. Breast Cancer Res 2011; 13(3): 209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Lazar AA, Cole BF, Bonetti M, et al. Evaluation oftreatment-effect heterogeneity using biomarkers measured on a continuous scale: subpopulation treatment effect pattern plot. J Clin Oncol 2010; 28(29): 4539–4544. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Huang Y, Sullivan Pepe M and Feng Z. Evaluating thepredictiveness of a continuous marker. Biometrics 2007; 63(4): 1181–1188. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Janes H, Pepe MS, Bossuyt PM, et al. Measuring theperformance of markers for guiding treatment decisions. Ann Intern Med 2011; 154(4): 253–259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Janes H, Brown MD, Huang Y, et al. An approach toevaluating and comparing biomarkers for patient treatment selection. Int J Biostat 2014; 10(1): 99–121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Dobbin KK and McShane LM. Sample size methods forevaluation of predictive biomarkers. Stat Med 2022; 41(16): 3199–3210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Fleming TR. Clinical trials: discerning hype from substance. Ann Intern Med 2010; 153(6): 400–406. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Micheel CM, Nass S and Omenn GS (eds) Evolution of translational omics: lessons learned and the path forward (Committee on the Review of Omics-Based Tests for Predicting Patient Outcomes in Clinical Trials; Board on Health Care Services; Board on Health Sciences Policy; Institute of Medicine ). Washington, DC: The National Academies Press, 2012. [PubMed] [Google Scholar]
  • 30.McShane LM, Cavenagh MM, Lively TG, et al. Criteriafor the use of omics-based predictors in clinical trials. Nature 2013; 502(7471): 317–320. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.McShane LM, Cavenagh MM, Lively TG, et al. Criteriafor the use of omics-based predictors in clinical trials: explanation & elaboration. BMC Med 2013; 11: 220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Sachs MC and McShane LM. Issues in developing multivariable molecular signatures for guiding clinical care decisions. J Biopharm Stat 2016; 26(6): 1098–1110. [DOI] [PubMed] [Google Scholar]

RESOURCES