In oncology drug development, biomarker-driven studies are pivotal, enabling the selection of patient populations most likely to benefit from specific therapeutic interventions. These studies have facilitated US Food and Drug Administration approvals for drugs targeting biomarker-positive subgroups, demonstrating their critical role even when initial trials encompass all-comer populations. Such focused research underscores the effectiveness of treatments in biomarker-defined subgroups, often revealing a favorable benefit-risk assessment. However, this selective drug approval process can inadvertently restrict access to potentially beneficial drugs for biomarker-negative patients, potentially depriving them of effective treatments. Thus, it is crucial to balance the benefits observed in biomarker-positive subgroups with the comprehensive evaluation of potential benefits in biomarker-negative subgroups.
Accompanying this editorial, the article “Biomarker-Driven Oncology Trial Design and Subgroup Characterization: Challenges and Potential Solutions” by Wang and colleagues1 offers profound insights into navigating the complexities inherent in the design and analysis of biomarker-driven trials for precision medicine. The study introduces an innovative decision tree that adeptly guides clinical trial design, distinguishing between patient enrichment and stratification strategies. This tool considers various critical factors such as marker type, prevalence, assay readiness, and turnaround times for marker assessments. Moreover, Wang and colleagues1 explore a range of statistical analysis methods for drawing conclusions about the treatment effect in the biomarker-negative subgroup. They advocate for an approach that determines meaningful effects for all subgroups beforehand, recommending the integration of Bayesian borrowing methods to leverage evidence from the biomarker-positive subgroup. It is important to recognize that the Bayesian approach recommended by the authors, while insightful, represents just one analytic approach and is not a universal solution for all biomarker-driven studies. Oncology trials are diverse and complex, requiring a tailored approach to statistical analysis that considers the unique characteristics of each trial. The study by Wang and colleagues1 will facilitate a discussion among various stakeholders, including those involved in drug development strategies, regulatory positions, and clinicians, on the optimal design and analysis of biomarker-driven clinical trials in precision oncology.
We highly respect the work of Wang and colleagues, which has significantly contributed to advancing our understanding of biomarker-driven trials. Their findings and methodologies provide a critical foundation for ongoing research and dialogue. This editorial aims to build on their valuable contributions by introducing additional statistical considerations that can further enhance the design and analysis of such studies.
In biomarker-driven studies, biomarkers may be measured either continuously or in a categorized format. For the purpose of this discussion, we will assume scenarios where two distinct subgroups have been already defined based on biomarker status: biomarker-positive and biomarker-negative. The core of the design and analysis of biomarker-driven trials can be simplified as the problem of statistical inference of treatment effects across three populations: biomarker-positive, biomarker-negative, and the combined all-comer. The goal of the statistical analysis is to provide, for each of the three populations, (1) a qualitative assessment to guide a dichotomous decision on whether the drug is effective or not, and (2) quantitative information about the magnitude of the treatment effect for treatment decision making.
Controlling the Type I Error Rate:
The qualitative assessment typically involves a statistical test with the Type I error rate controlled at a certain level (conventionally two-sided 0.05). While regulatory decisions are based on a totality of evidence rather than solely on p-values, it is imperative to prespecify such a criterion for determining whether the treatment benefits patients in each of the three analysis populations. This specification is essential for appropriately sizing each subgroup to draw meaningful conclusions for each of the three analysis populations. While there are various statistical approaches introduced by Wang and colleagues for the analysis of the biomarker-negative subgroup, it would be critical to employ a statistical method that can maintain the Type I error rate at a constant level for the qualitative assessment. This is because not only inflation but also deflation of the Type I error rate from the nominal level can lead to significant issues regardless of which analysis population the analysis is performed in. Depending on the situation, the threshold would not necessarily be the conventional threshold of 0.05. What criteria generate useful information for the decision-making process at regulatory agencies may also vary on a case-by-case basis. However, maintaining control over it would be generally important to ensure the robustness of the trial results and support regulatory decision-making by providing a reliable basis for evaluating the treatment benefits (or risks) in each subgroup and the combined all-comer population as well.
Choosing Robust and Interpretable Quantitative Summaries of Treatment Effect:
As highlighted in the American Statistical Association’s statement on p-values, which outlines the limitations of relying solely on p-values2, providing quantitative information about the magnitude of the treatment effect can be significantly more informative and useful for treatment decision-making. Regarding the estimation of the treatment effect magnitude, an important consideration is the choice of “estimand” i.e., a summary measure used to quantify the between-group difference. Traditionally, for time-to-event outcomes, Cox’s hazard ratio (HR) has been employed almost universally in oncology clinical trials3, as Wang and colleagues1 also employed it in their paper. However, this traditional measure faces several limitations in providing a robust and interpretable estimate of the treatment effect magnitude for target populations of interest4,5. Particularly in the present setting with two subgroups or strata, the proportional hazards (PH) assumption must hold for both biomarker-positive and biomarker-negative subgroups. Additionally, to integrate the HR from these two subgroups to derive the HR for the combined all-comer population via a stratified Cox’s analysis, there is an underlying assumption that the HR remains constant across both biomarker-positive and biomarker-negative subgroups. Violations of these assumptions can challenge the interpretability of the resulting HR and its generalizability to future patient populations4,5. It is also important to note that, except for some special cases, the PH assumption does not hold in the combined all-comer population when the PH assumption holds in both biomarker-positive and biomarker-negative subgroups6. Another notable limitation of HR involves the absence of absolute hazards from the treatment and control groups that yield the HR. This presents practical difficulties in interpretation because the clinical value of treatment would depend on the baseline hazard in the control group and would not be fully determined by the between-group contrast (i.e., HR) alone. Given these limitations, the traditional Cox’s HR approach may not fit well for achieving the goal of the analysis. It is worth exploring alternative approaches that do not hold these limitations. For example, the approach using restricted mean survival time7,8 is gaining attention as a robust alternative to the traditional HR approach and is beginning to be employed in practice9,10. The approach using average hazard with survival weight11,12 would also be an option to be considered if the investigators are interested in summarizing the treatment effect in terms of hazard. Methods producing robust and interpretable results via stratified analysis are also available for these alternative approaches6,13.14.
Coherency of Statistical Analysis Models Used for Three Analysis Populations:
Employing varied statistical methodologies across different subgroups within the same trial may raise concerns about the validity and consistency of the results. For example, if a hierarchical model is imposed for the analysis to account for heterogeneity in the treatment effect between biomarker-positive and biomarker-negative subgroups, the same analysis should be performed for both subgroups, and the model should also be taken into account for the inference in the all-comer population. Wang and colleagues1, in the primary analysis, analyzed only the biomarker-positive subgroup, which is assumed to have a strong treatment effect. They then analyzed the biomarker-negative subgroup, which is assumed to have a weak treatment effect, using a Bayesian dynamic method that borrows information from the biomarker-positive subgroup. Such asymmetry in the subgroup analysis, as the primary analysis, may be perceived to be too convenient for the pharmaceutical companies when conducting studies for drug indications. Because the treatment effect in the biomarker-positive subgroup is more pronounced, the results of the biomarker-negative subgroup with borrowing from biomarker-positive would appear to be better than the result without borrowing. On the other hand, in the biomarker-positive subgroup, the results without borrowing from the biomarker-negative would appear better than with borrowing. Bayesian dynamic borrowing is indeed a powerful tool to integrate similar studies systematically. A convincing rationale should be provided especially when a different analytic concept or model is applied to the analysis for biomarker-positive, biomarker-negative, and the all-comer population.
Biomarker-driven research has revolutionized the development of cancer therapeutics by enabling targeted treatments for specific patient subgroups. It has also brought to the forefront the challenge of ensuring that effective treatments are appropriately delivered to patients who benefit from them. The research by Wang and colleagues1 highlights the complexities of study design and statistical analysis, exploring a variety of analytical methods. It would be essential to adopt statistical approaches that provide robust, reliable, and interpretable quantitative information regarding the magnitude of treatment effects within each population to support decision-making. This discussion serves as a catalyst for an ongoing dialogue aimed at refining the methodologies of biomarker-driven clinical trials for precision oncology.
Acknowledgments
Supported by the National Institute of General Medical Sciences of the National Institutes of Health under award number R01GM152499 (H.U.)
References
- 1.Wang J, Yu B, Dou Y, et al. : Biomarker-Driven Oncology Trial Design and Subgroup Characterization: Challenges and Potential Solutions. JCO Precision Oncology: TBD [DOI] [PubMed] [Google Scholar]
- 2.Wasserstein RL, Lazar NA: The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129–133, 2016. 10.1080/00031305.2016.1154108 [DOI] [Google Scholar]
- 3.Uno H, Horiguchi M, Hassett MJ: Statistical Test/Estimation Methods Used in Contemporary Phase III Cancer Randomized Controlled Trials with Time-to-Event Outcomes. The Oncologist, 25(2), 91–93, 2020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Uno H, Claggett B, Tian L, et al. : Moving beyond the hazard ratio in quantifying the between-group difference in survival analysis. J Clin Oncol 32, 2380–2385, 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Horiguchi M, Hassett MJ, Uno H: How do the accrual pattern and follow-up duration affect the hazard ratio estimate when the proportional hazards assumption is violated? The Oncologist, 24(7), 867–871, 2019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Tian L, Jiang F, Hasegawa T, et al. : Moving beyond the conventional stratified analysis to estimate an overall treatment efficacy with the data from a comparative randomized clinical study. Statistics in Medicine. 38, 917–932, 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Royston P, Parmar M: The use of restricted mean survival time to estimate the treatment effect in randomized clinical trials when the proportional hazards assumption is in doubt. Statistics in Medicine, 30, 2409–2421, 2011. [DOI] [PubMed] [Google Scholar]
- 8.Royston P, Parmar M: Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC Medical Research Methodology, 13, 152, 2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Gandara DR, Paul SM, Kowanetz M, et al. : Blood-based tumor mutational burden as a predictor of clinical benefit in non-small-cell lung cancer patients treated with atezolizumab. Nature Medicine, 24, 1441–1448, 2018. [DOI] [PubMed] [Google Scholar]
- 10.Guimarães HP, Lopes RD, de Barros E Silva PGM, et al. : Rivaroxaban in patients with atrial fibrillation and a bioprosthetic mitral valve. New England Journal of Medicine, 383, 2117–2126, 2020. [DOI] [PubMed] [Google Scholar]
- 11.Uno H, Horiguchi M: Ratio and difference of average hazard with survival weight: New measures to quantify survival benefit of new therapy. Statistics in Medicine. 42, 936–952, 2023. [DOI] [PubMed] [Google Scholar]
- 12.Uno H, Tian L, Horiguchi M, et al. : Regression models for average hazard. Biometrics, 80(2), 2024. 10.1093/biomtc/ujae037 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Sun R, McCaw Z, Tian L, et al. : Moving beyond conventional stratified analysis to assess the treatment effect in a comparative oncology study. J Immunother Cancer. 2021. Nov;9(11):e003323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Qian Z, Tian L, Horiguchi M, et al. : A Novel Stratified Analysis Method for Testing and Estimating Overall Treatment Effects on Time-to-Event Outcomes Using Average Hazard with Survival Weight. arXiv preprint 2024, arXiv:2404.00788. [Google Scholar]
