Abstract
Before implementing a biomarker in routine clinical care, it must demonstrate clinical utility by leading to clinical actions that positively affect patient-relevant outcomes. Randomly controlled early detection utility trials, especially those targeting mortality endpoint, are challenging due to their high costs and prolonged duration. Special design considerations are required to determine the clinical utility of early detection assays. This commentary reports on discussions among the National Cancer Institute’s Early Detection Research Network investigators, outlining the recommended process for carrying out single-organ biomarker-driven clinical utility studies. We present the early detection utility studies in the context of phased biomarker development. We describe aspects of the studies related to the features of biomarker tests, the clinical context of endpoints, the performance criteria for later phase evaluation, and study size. We discuss novel adaptive design approaches for improving the efficiency and practicality of clinical utility trials. We recommend using multiple strategies, including adopting real-world evidence, emulated trials, and mathematical modeling to circumvent the challenges in conducting early detection utility trials.
Introduction
Improving clinical outcomes for cancer patients is critically dependent on implementing effective early detection and follow-up procedures for those with abnormal screening results. The mission of the National Cancer Institute’s Early Detection Research Network (EDRN) is to translate promising cancer biomarkers into clinical practice, which would allow for better decision-making and improved outcomes. The consortium recognizes that establishing biomarker-based strategies’ clinical validation and clinical utility is critical for effective translation.
Before being implemented in routine clinical care, a biomarker must demonstrate clinical utility (1), i.e., the results of the test should lead to clinical actions that positively affect patient-relevant outcomes. Compared to clinical validity, i.e., test performance measures such as sensitivity and specificity, clinical utility often requires a higher level of evidence, such as reducing late-stage cancer detection and mortality or improving quality of life. With EDRN’s long-standing effort, many biomarkers have evolved from the laboratory to clinical practice (https://edrn.nci.nih.gov/about-edrn/fda-approved-tests/). While these biomarkers have met the performance criteria for clinical validity, limited studies have demonstrated clinical utility and cost-effectiveness. Investigators often fail to appreciate the high bar for a biomarker test to demonstrate its positive impact on clinical outcomes. There is also limited opportunities due to the cost of conducting clinical utility studies. Guidelines and a clear workflow are needed to relieve this bottleneck between biomarker validation and clinical utility.
Significant challenges exist in the conduct of early detection clinical utility studies. Unlike their use in prognostication and treatment selection, biomarkers for early detection often indirectly impact patients by triggering downstream clinical actions such as diagnostic follow-up and treatment, and their effectiveness can depend on available therapeutic strategies. Consequently, there can be a long time between a positive test for early-stage cancer and a clinical outcome such as disease-specific mortality. Various diagnostic-treatment pathways may be undertaken during this time, leading to greater uncertainty on the ultimate survival benefit. In addition, trials with long-term follow-up are impractical and in danger of becoming irrelevant due to evolving treatments. As such, early detection trials for cancers with low incidence even in high-risk populations require a large sample size, resulting in substantial costs and long durations (2).
Randomized controlled trials (RCTs) provide the gold standard to establish clinical utility, particularly assessing whether the new test can produce a meaningful survival benefit for patients beyond the current standard of care. Guidelines for conducting biomarker utility trials are needed to confront the unique challenges of demonstrating the effectiveness of biomarker-based decision rules compared to standard of care to achieve rigor and feasibility. Recognizing the critical value of clinical utility assessment and the enormous challenges in conducting RCTs for early detection biomarkers, the EDRN, during its October 2020 Steering Committee Meeting, devoted a session to the design of clinical utility trials. Molecular biologists, clinicians, population scientists, and biostatisticians participated in the discussion. Significant progress has been made recently in the planning and initiation of cancer early detection clinical utility trials within the EDRN. We report the recommended process of conducting biomarker-driven clinical utility studies and illustrate key points with completed or proposed clinical utility trials.
In this manuscript, we (i) describe the process for conducting clinical utility trials in the context of EDRN’s 5-phases of biomarker development (3); (ii) discuss key components of phase 4 and 5 early detection clinical utility trials, including biomarker performance criteria, endpoint determination, study design and the size of the study; (iii) review various approaches for establishing clinical utility or enhancing an RCT; and (iv) recommend novel design considerations to improve the efficiency and practicality of clinical utility trials. We focus on the design of clinical utility studies for single-organ early detection and briefly discuss the broader implication of our work for the more complex landscape of multi-cancer early detection (MCED) tests.
Clinical Utility Trials in the Context of EDRN 5-Phases of Biomarker Evaluation
Phases of Biomarker Development
Biomarker development is a complex process, progressing through several phases to produce a meaningful clinical tool with analytical validity, clinical validity, and clinical utility, as illustrated in Figure 1. Biomarker research within the EDRN unfolds in a phased approach and follows established standards for reporting results at each stage (3).
Figure 1: EDRN 5-Phase biomarker development.
The process of biomarker development comprises five phases, ranging from discovery to the confirmation of analytical validity, clinical validation and clinical utility. Figure adapted from Pepe et al., J Natl Cancer Inst 2001;93:1054-61.
The process starts with the discovery stage (phase 1) to identify promising markers that are tested in a case-control or cross-sectional study (phase 2). A phase 3 study is to demonstrate the clinical validity of the biomarker test, assessing its accuracy in reflecting the disease status. This phase involves evaluating operating characteristics of the test as applied to the targeted population, measured with sensitivity and specificity.
Phase 3 validation is typically a retrospective assessment of an established biomarker-assisted decision rule without influencing clinical decisions. Transitioning to the phase 4 study, a prospective trial unfolds, wherein a positive test triggers a change in the diagnostic workup. This stage is often regarded as the initiation of true clinical utility assessment. The primary endpoint is the characteristics of the disease (e.g., stages) detected with the test and the false referral rate. The positive and negative predictive values (PPV and NPV) can be derived from such prospective studies.
A phase 5 study aims to demonstrate the reduction in cancer mortality afforded by the real-world adaptation of a new biomarker test. While phase 4 studies stop at cancer diagnosis, phase 5 studies extend to the time from study entry to cancer-specific death, a critical endpoint in the assessment of a biomarker’s effectiveness in a broader clinical context.
Phased Approach for Planning Clinical Utility Studies
The phases of biomarker development serve as successive stepstones, each contributing increasingly robust evidence of clinical utility. Earlier phases’ results are often required to advance to subsequent phases. A phase 3 study finalizes the positivity criterion of the biomarker test for a phase 4 trial. A phase 4 trial is crucial to confirm biomarker-detected signals through diagnostic workup and gather late-stage information. Neglecting this evidence can result in negative outcomes in an expensive phase 5 trial. Thus, the success of each phase ensures a robust foundation for subsequent stages and safeguards against potential setbacks.
The individual marker’s pathway to establish clinical utility may differ and not necessarily go through the entire sequence of phases. For instance, if there is clear evidence from completed RCTs to demonstrate that the standard of care test confers desired clinical outcomes (such as a reduction in mortality), then a phase 3 finding of improved specificity or safety, with comparable sensitivity and detected cases could be sufficient for establishing the clinical utility. However, there is still value in a phase 4 trial, to evaluate other changes in clinical outcomes such as harms (beyond false positive and false negative rates) that are not evaluated in the phase 3 study. It also allows the evaluation of the adherence to the test in real-world practices. This is particularly important for tests in which providers and patients may interpret and act differently than the current standard (e.g., shift from imaging or endoscopic screening to biomarker screening). In the absence of phase 4 study, post-marketing surveillance would be recommended. In the case a stage shift endpoint is demonstrated and its use as a surrogate for mortality is justified, a definite clinical utility benchmark can be reached in a phase 4 validation trial of candidate biomarkers, without need for phase 5 evaluation. However the surrogacy justification of stage shift is complicated with early detection biomarker tests. When evidence connecting stage shift with mortality reduction is lacking, a phase 5 trial should be considered. if resource limitations make a lengthy Phase 5 trial impractical, we advocate against premature conclusions and recommend alternative methodologies to predict Phase 5 endpoints based on Phase 4 trial findings, ensuring informed decision-making despite feasibility constraints.
The phased approach provides a useful framework for planning clinical utility studies and determining whether lengthy studies with long-term outcomes are necessary. This approach ensures limited funding is allocated efficiently. The Prospective Specimen Collection Retrospective Blinded Evaluation (PRoBE) design principle has been established by EDRN to ensure the rigor of a validation study (4), but there is little guidance available for phase 4 and 5 studies. In the following section, we will describe approaches to facilitate the design and interpretation of an early detection clinical utility trial.
Key Components of Early Detection Clinical Utility Studies
Biomarker Tests of Phase 4 or 5 Trials
As phase 4 and 5 trials involve further workup based on biomarker test results, a rigorous PRoBE-compliant phase 3 study is a prerequisite to obtain reliable estimates of test performance. The test needs to meet the minimally acceptable performance criteria to ensure that the consequences of a false negative, leading to a missed opportunity for further evaluation and early cancer detection, and a false positive, triggering unnecessary diagnostic confirmation, are controlled within clinically acceptable limits. In instances where a positive test prompts additional clinical intervention, the sensitivity / (1-specificity) achieved should exceed a predefined target value, aligning with the cost-benefit ratio. A test with phase 3 performance meeting such criteria can then advance for further evaluation (5). Ultimately, test performance criteria for advancing to later phases should be based on the projected downstream impact on endpoints. For example, the utility of PSA screening can be gauged by the number needed to screen per life saved and the number of excess diagnoses per life saved.
To determine a biomarker’s candidacy for a phase 4 trial, it is necessary to project if the testing sensitivity can yield the required effect size for phase 4 endpoints. For example for a biomarker used for a population screening test, it is crucial that sensitivities derived from a phase 3 biomarker validation studies specifically address early-stage preclinical cancer, ideally estimated in a prospective study where the screening biomarker test is performed on an asymptomatic population and followed immediately with the definitive diagnosis for outcome (6). Validation with samples collected from patients with symptoms who are clinically diagnosed with cancer at specific stages likely overestimates the sensitivity required to achieve late-stage incidence reduction endpoint. Conversely, biomarkers ascertained from retrospective longitudinal repository studies may underestimate the sensitivities due to the potential miscounting of false negatives at the time of biomarker collection before cancer becomes detectable and subsequently symptomatic, clinically diagnosed, and staged.
As phase 4 and 5 studies entail real-world implementation, viewing biomarker tests as a potential implementable population program is crucial. Trial planning must extend beyond assessing accuracy and take into consideration other test features, such as the age range for initiation and termination and the interval between repeated tests. These parameters will impact sample sizes and follow-up duration. Furthermore, data derived from utility trials may prompt further refinement of the test regimen, optimizing its efficacy.
Endpoint Consideration
Mortality or quality-adjusted life-year is arguably the ultimate outcome of a medical test in the oncology (2,7). In an early detection setting, a trial following low-risk patients for disease-specific mortality is often lengthy and expensive. Thus, it is tempting to use an outcome such as a reduction of late-stage incidence or recurrence reduction as a surrogate for mortality reduction to shorten the duration of the trial and reduce costs. However, an observed stage shift may not necessarily lead to mortality reduction, particularly in the presence of overdiagnosis (7). The ability of each biomarker test to detect clinically undetected disease might depend on the tumor’s molecular characteristics. Even within the same anatomic stage group, different molecular characteristics might confer different prognoses for a mortality endpoint. Thus, investigations demonstrating that a surrogate is reliable and effective as an endpoint are required before its use in a clinical utility trial (8). The justification of a surrogate endpoint isn’t always straightforward. One approach involves examining evidence from existing therapeutic trials relating stage-specific prognostic subtypes to mortality. For example, Patients with early-stage Hepatocellular carcinoma (HCC) can undergo curative treatments with 5-year survival greater than 60%, whereas late-stage HCC eligible for palliative treatments has a median survival of 1–3 years. Thus, the TRACER study justifies late-stage HCC as a valid surrogate for HCC-related mortality.
Clinical contexts of preventive tests are often broader than those of a therapeutic procedure. The clinical utility relevant to the specific setting often goes beyond the mortality (9). An evaluation of the full range of a test’s effects on patients, including improved safety, timeliness, efficiency, equity, and patient-centeredness, may be needed. For a lung cancer biomarker that better stratifies lung cancer risk among smokers performed prior to low-dose computerized tomography (LDCT) screening, the reduction in the number of unnecessary LDCT or the frequency of LDCT can be considered part of the multi-dimensional outcome measures. In hepatocellular carcinoma (HCC) screening, tests to identify low-risk patients with cirrhosis for less intensive screening would reduce disease burden while avoiding screening-related harms in patients unlikely to otherwise derive any benefit (10). In these examples, a test may have clinical utility for reducing the disease burden to the patient without a significant improvement in mortality. Other possible endpoints include (i) an improvement in screening uptake, (ii) a decrease in the frequency of invasive procedures, including expensive or morbid ones, (iii) a reduction of time to diagnosis, and (iv) an improvement in quality of life, including patient-reported physical and psychological outcomes. These measures can be considered primary endpoints when the impact on mortality of the new test is considered comparable to the current management. Compared with these other utility endpoints, RCTs for early detection utility targeting mortality endpoint are notably challenging due to their high costs and prolonged duration. Thus, we will focus on describing novel designs and approaches to improve such trials’ feasibility.
Design Considerations for Early Detection Biomarker RCT
Standard design:
A standard RCT for biomarker clinical utility involves participants randomized into a group managed according to the novel biomarker (intervention group) finding and a group with the current standard strategies and managed with best practices (control group). All patients are followed for the clinical outcomes of primary interest. The clinical utility can then be derived from comparing the outcomes between the biomarker-tested group and the control group (Figure 2). Such a biomarker strategy design provides the most robust evidence. However, it may require a trial with a large number of participants and a long duration, especially in the context of early detection. For example, the National Lung Screening Trial (NLST) was conducted to investigate the impact of low-dose compute tomography (LDCT) screening on lung cancer mortality. The primary endpoint was the reduction in lung cancer-specific mortality among high-risk individuals who underwent annual LDCT (intervention arm) compared to these screened with standard chest x-rays (control arm). The study enrolled over 53,000 participants with 10-year follow-up (2). To improve efficiency, modification from such a gold standard protocol can be considered but should be done with deliberation to maintain the study rigor (11).
Figure 2: Standard design of a randomized, controlled trial for biomarker utility.
In a standard randomized controlled trial (RCT) for assessing biomarker clinical utility, participants are randomized into two groups: an intervention group managed according to the findings of the novel biomarker, and a control group managed according to current standard practices The outcomes of both groups are then followed and compared to determine the clinical utility of the biomarker.
Adaptive phase 4/5 design:
In settings that require an RCT to demonstrate mortality reduction, using a flexible adaptive design that combines multiple phases of studies is an appealing solution to improve the overall efficiency of the trials, reduce cost, and shorten the duration compared to a standard trial. Subsequent steps can be adjusted based on the outcomes of the earlier phase results. The National Liver Cancer Screening Trial (TRACER) is a phase 4 clinical utility trial recently launched in EDRN to assess the effectiveness of novel markers for HCC. Because of the long natural history of progression from cirrhosis to HCC and involved protocol for active screening, sequentially performing a phase 4 study followed by a phase 5 study would lead to a long delay, so an adaptive seamless design was proposed. It planned to enroll 5,500 patients with liver cirrhosis who are under surveillance for HCC and randomized into two arms. Patients in the intervention arm will use a biomarker test GALAD for semi-annual surveillance, while patients in the control arm will use standard surveillance modality, i.e., semi-annual surveillance using liver ultrasound combined with Alpha Fetoprotein (AFP) test. For the phase 4 component of the trial, the study is sufficiently powered for the early endpoint of the proportion of late-stage HCC (using an upper bound of 5% overdiagnosis to adjust the estimation of the treatment effect) and has sufficient power at a later year for the endpoint of incidence of later-stage HCC. If the early endpoint of Phase 4 is met, the study will continue, and a sample size re-estimation will be conducted to determine the sample size required for the Phase 5 trial with mortality reduction as the endpoint. Such a design allows biomarkers to be approved without a long delay if they meet the phase 4 objective, particularly for HCC, where overdiagnosis is not a common concern, and approval could be withdrawn if the result on the phase 5 endpoint is not favorable. Here, a sequential design with sample sizes powered for the final endpoint can be more efficient than conducting separate trials for each phase. It is also potentially more cost-effective than a single-phase 5 trial, as the study will halt if there is no evidence of stage migration. In other settings, one may consider designs of a phase 5 trial by recruiting only subgroups where an association between the phase 4 endpoint and mortality reduction is lacking. Adaptive trials need to be prospectively planned to ensure that type I error is controlled, sufficient power is preserved for the final endpoint, and inferences account for the adaption procedures.
Master protocol:
A master protocol using a single overarching design to evaluate multiple markers simultaneously can leverage establishing a uniform infrastructure and procedures, thereby improving overall efficiency for a coordinated effort to assess an array of emerging novel biomarkers quickly. For example, the EDRN’s lung cancer collaborative group proposed a master protocol with a platform trial design. The trial allows for the evaluation of multiple biomarkers, separately or in combination, for risk stratification of patients with intermediate-risk pulmonary nodules. In each of the biomarker-based intervention arms, patients will be stratified into high, intermediate, and low risk based on the biomarker measurements, with either a biopsy, standard of care, or CT follow-up at three months, respectively. The intervention groups will be compared with a common control arm that is followed with the standard of care without the knowledge of biomarker results. Master protocols have increased complexity compared with the traditional single-trial design, requiring great collaborative efforts and thoughtful upfront planning. For example, the common standard-of-care control arm may be no longer contemporaneous or nonoverlapping in recruitment time with later biomarker arms. Varying sample size requirements with multiple biomarker arms may require a trial to be adaptively expanded.
Nested enhanced design:
For cancers without standard-of-care screening, the effect of screening on mortality derives directly from cancer detected by the test and curative treatment. To improve trial efficiency, an interesting sampling design has been proposed to focus the comparison on outcomes of those who tested positive in the two arms. By considering mortality proportions of positive cases instead of the absolute mortality rate, the approach may substantially reduce sample size (12). For example, in a stomach cancer screening trial, blood samples were banked and assayed for all in the screening arm, those who developed cancer or died, and a randomly selected fraction in the control arm. Compared with a traditional RCT, the trial required less than half of the sample sizes (13). However, the design may miss other aspects of the regimen, including adherence and the impact of false negatives. An ultimate evaluation of the original purpose of all randomized subjects may still be preferred.
When investigating organs with established standard-of-care screening modalities, studies must be designed to allow for direct comparison to the gold standard method. As illustrated in the Early Diagnosis of Lung Cancer Screening Scotland (ECLS) trial (14), a 7-autoantibody blood test (EarlyCDT) was given to the participants in the intervention arm. Those testing positive were followed by X-ray and LDCT, while test-negative and control arm participants received standard clinical care without LDCT. While the study observed a stage shift in terms of reduced late-stage proportion in the intervention arm, the design missed an opportunity to assess if the new test provided any improvement in LDCT. While these innovative designs offer the promise of accelerated early detection biomarker development, it is imperative to approach them with caution to ensure comprehensive and definitive evaluation. More in-depth discussions that identify the challenges and strategies to address them are needed.
Sample Size Considerations
Sample size calculation of an early detection clinical utility trial starts with the specification of hypotheses. A null hypothesis for a typical trial is that the gain in clinical utility by the new biomarker test is less than the minimally acceptable outcome set for the clinical setting. The alternative hypothesis is set at the targeted desirable improvement in clinical utility by the biomarker arm. A positive conclusion is drawn if the lower bound of a 95% confidence interval of the outcome parameter estimated from the study exceeds the minimally acceptable criteria for clinical utility. Sufficient sample size is needed to ensure that there is a high chance (or power) that such a positive conclusion will be drawn from the study if the observed clinical utility is as good as is anticipated. The study also needs to maintain a low probability (type I error, e.g., 5%) of drawing a positive conclusion if, in fact, there is no gain in clinical utility.
Utility trials may require larger sample sizes than therapeutic trials because participants who tested negative in the biomarker arm will follow the same standard of care as those in the control arms, which might dilute the intervention effect from participants who tested positive. As in the example of the stomach cancer screening trial (13), targeting the group with a substantial fraction of participants for whom biomarkers most likely make alternative management recommendations may improve the trial efficiency. It is essential to approach such analyses with caution, considering the assumptions made and evaluating whether this method diverges from the initial assessment of effectiveness on all randomized subjects.
Stage shift can be defined as a decrease in the conditional proportion of later-stage cancers among all cancer cases detected. However, in the presence of overdiagnosis, this endpoint could result in biased inference regarding the efficacy of a screening strategy. In contrast, inference based on the absolute incidence endpoint (the number of late-stage diseases detected among all tested patients) remains valid. On the other hand, for screening of rare cancers, as demonstrated in simulated studies, using absolute incidence as the endpoint can inflate sample size requirements several-fold more compared with the conditional proportion endpoint (11). For cancers in which overdiagnosis is assumed unlikely, such as pancreatic and ovarian, using the conditional proportion endpoint could save resources (11,15).
For adaptive trials or master protocols, sample size considerations are more complicated. In the example of the TRACER comparing HCC-related mortality rates between a biomarker-assisted rule and the standard-of-care arm, since the study is designed to allow for termination at phase 4 if the stage migration endpoint is not met, adjusted significance levels at phase 4 and 5 for claiming positive results are specified in advance for both tests, and efficacy thresholds selected at both phases will impact the study power for phase 5.
A significant hurdle in designing early detection trials is specifying the effect size. This contrasts therapeutic trials, where the effect size can often be inferred from the Phase 2 trial outcomes, conducted on a smaller scale and with a relatively shorter time frame. For early detection biomarkers, the investigator only has information regarding the biomarker performance, such as sensitivities obtained retrospectively in an early phase biomarker validation study. The expected effect size, such as the stage shift or mortality reduction due to implementing of a novel biomarker screening test, is unavailable without a prospective screening of asymptomatic patients with an extended follow-up period. A probabilistic model that characterizes the cancer screening process can be used to project expected late-stage cancer reduction and mortality from screening test performance. Such a model must capture the time-varying screening effect through the interplay between the disease’s natural history, test sensitivities, frequency and rounds of testing, and the duration of follow-up (16). For trials using stage shift and mortality endpoints, existing natural history models need expansion to encompass pathways leading to both early and late-stage pre-clinical and clinical cancers, extending from diagnosis to eventual cancer-specific death to account for leading time and within-stage benefit.
Other Approaches to Ascertain Clinical Utility
Several alternative approaches can often augment or circumvent the constraints of traditional RCTs for early detection markers, justify the launch of clinical utility trials, and provide interpretation of trial results. As shown in Table 1, these approaches can be complementary in ascertaining clinical utility.
Table 1.
Approaches to Assess Clinical Utility
| Description | Benefit | Constraint | |
|---|---|---|---|
| Randomized Controlled Trial (RCT) | Randomize participants to biomarker-assisted strategy and standard of care arms, follow for patient outcomes | Provide most robust conclusion on biomarker effectiveness | Expensive and with long duration; assesses only limited number of test-treatment strategies |
| Emulated Trial (ET) from observational prospective cohort Study | Analyze by mimicking the features of a target trial using observational data of cohort patients with biomarker tests; emulate RCT on subsets with follow-up matched to the biomarker testing strategy and standard of care | Efficient; reflects real-world management pathways; less expensive than RCT and quick conclusion | Requires the test is marketed for population or stored samples available in the cohort; relies on propensity score modeling; old cohort may not reflect current management |
| Linked Evidence Approach (LEA) | Synthesis of empirical evidence from multiple sources including: test performance from phase 3 studies, treatment choice dictated by the test result, and treatment effect on definite outcome from existing clinical trial | Leverage existing test performance and trial results; useful in justifying surrogate | Assumptions difficult to justify; empirical data in the linkages may be lacking |
| Decision-Analytic Modeling (DAM) | Use natural history model to project outcomes of interest, with inputs of biomarker accuracy under various diagnostic-treatment pathway | Project to settings when RCT are infeasible or when empirical evidence is limited for LEA; support the design of RCT and phase 3 study | Natural history model may not be available for the disease or for testing population; validity relies on assumptions and input parameters |
Emulated trials (ET)
ET is an analytical method that evaluates the effectiveness of an intervention using observational data by intentionally replicating the structure and characteristics of a randomized clinical trial. When a phase 4 or 5 RCT is impractical or untimely, ET offers a valuable alternative. By resorting to analyses using real-world observational data, ET intentionally mimics the features of a target trial (17). for example, in assessing the impact of screening colonoscopy on colorectal cancer incidence, while waiting for trial results, early evidence of efficacy relies on analyses of observational real-world data. Emulating a hypothetical target trial of screening colonoscopy using a comprehensive insurance claims database from the United States Medicare program suggested a moderate benefit of the test (18). This strategy deserves particular attention if contemporary liquid biopsy tests enter the market and data become widely available in healthcare systems while evidence of clinical utility from phase 4 or 5 RCT is anticipated later. In some cases, banked biospecimens from existing prospective screening preventive trials can also be used to derive ET for future novel biomarker utility studies (19). Even in completed screening trials where a clinical decision is not based on the new biomarker test, one can emulate a utility trial by identifying the participants whose subsequent workup aligned with the biomarker follow-up strategy and comparing their outcomes with those following current practice.
For comparative effectiveness studies with observational data, strategies accounting for the absence of randomization include instrumental variable analysis, propensity score matching, and inverse probability weighting. The aptness of these adjustment approaches relies on deriving a robust propensity score based on the probability of having subsequent workup consistent with the medical test suggestion. When appropriately analyzed, results from an ET can offer real-world evidence of biomarker test effectiveness or justify launching a definite utility trial. Such “prospective-retrospective” designs are more efficient than prospective RCTs, and it is recommended the results be validated from multiple independent studies with archived specimens to limit inherent biases.
The linked evidence approach (LEA)
LEA is a method that systematically synthesizes empirical evidence pertinent to test (i) performance from a phase 3 study, (ii) biomarker’s effect on downstream management such as diagnostic workup and treatment, and (iii) treatment effects on surrogate and definite clinical outcomes (20–22). It is used by the United States Preventive Services Task Force (USPSTF) as one of the analytic frameworks in the development of clinical practice recommendations (23). LEA can help decide if an RCT is necessary when the entire chain of data sources is available or provide support for the use of test performance measures or clinically relevant surrogate outcomes for justifying clinical utility. Furthermore, this exercise helps identify the missing pieces of evidence. Subsequently, a future adaptive trial focusing on obtaining that information can be designed efficiently.
Decision-analytic modeling (DAM)
DAM provides a powerful strategy to estimate patient-relevant outcomes and cost-effectiveness of the early detection program with a simulated mathematic model (24). Large screening trials, such as the NLST, utilized DAM to inform and guide their design decisions (2). An early detection model is a mathematical engine representing interactions between the natural history of cancer development and early detection tests, with clinical outcomes of such interactions as the output. Like LEA, the validity of DAM is dependent on parameters derived from reliable information on the performance of all tests and treatment effects on patient outcomes. The results can be misleading if there is a mismatch between the true and simulated scenarios. DAM can be a useful tool to fill the gap when empirical data is insufficient, or the assumption of LEA is unjustified. For a disease with long natural history and where only short-term or surrogate outcomes are feasible, DAM can project long-term benefit and harm and assess the societal cost-effectiveness of early detection programs. DAM can account for uncertainty in the implementation of various test-treatment strategies and thus generalize trial results in broader settings. It can also infer outcomes for heterogeneous populations, particularly under-represented subgroups in studies. Calculating backward from the desired clinical outcomes, DAM can also identify benchmark performance for a medical test. Thus, DAM supports the design of a clinical validity study, giving more confidence for launching an early detection RCT.
Integrating RCTs with alternative approaches like LEA and DAM may offer innovative strategies to expedite the clinical utility assessment. One can implement a hybrid trial design by first employing RCTs to assess the surrogate endpoint, which can be measured faster and with fewer resources than definite outcomes. Leveraging existing data from various sources, the probabilistic natural history model integrates observed surrogate outcomes from the RCT with data from LEA to project the definitive outcome, potentially shortening the time required for comprehensive follow-up. However, the implementation requires careful consideration to address potential limitations and challenges, such as robust data input, appropriate methodologies, and validation of surrogate endpoints in reliability for predicting definite outcomes. It is also imperative to exercise caution and account for real-world nuances in interpreting the projected outcomes. For instance, patients detected from a biomarker test may differ in histology and treatment pathway from those in a standard-of-care arm, which extends beyond the projection based solely on cancer stages (7).
Conclusion
With the rapid evolution of liquid biopsy technologies offering screening for various cancers, there is an amplified need for innovative trial design methodologies, especially those to reduce trial duration and costs (25). Lessons from the methods employed in single-cancer early detection can be applied to advance the development of MCED tests. First, implementing an early detection test in the population similarly requires evidence of benefits and harms through a phased approach characterized by sequences of performance criteria. Secondly, the complexities associated with study designs, including biomarker test performance criteria, endpoint considerations, and trial planning and analyses, are further magnified in the context of the MCED test due to differences in the availability of standard-of-care screening tools and workup spaces across cancer types. Similar intricacies apply to alternative approaches. For example, developing a probabilistic model for projecting survival benefits based on stage shift demands a comprehensive understanding of the natural history of each cancer and the interplay of competing risks among them. Given the limited real-world data from ongoing MCED trials at the moment, numerous assumed model parameters become necessary. Therefore, the field must diligently address the urgent and unmet need for novel design methodologies. By applying foundational design principles gleaned from single-marker studies and integrating emerging information from ongoing MCED test trials, there is hope that we can uphold rigorous standards for MCED tests.
While conducting rigorous clinical utility trials of early detection biomarkers is unarguably critical in translational research, EDRN investigators increasingly recognize that the constraints of RCTs may significantly hamper the advancement of biomarker tests into clinical practice. We have highlighted a broad range of clinical utilities that are particularly meaningful in the context of early detection biomarkers. These clinical utilities can be evaluated in a phased fashion to save resources. We review approaches for ascertaining clinical utility, including considering appropriately designed nonrandomized comparative effectiveness studies and using mathematical modeling strategies to circumvent the challenges with RCTs. When a biomarker RCT is a must, there is a substantial need for novel trial designs that will (i) demonstrate patient-relevant benefit; (ii) minimize the required time to conclusive results or the number of subjects needed for the trial; and (iii) avoid biases such as lead time and overdiagnosis. This review lays the groundwork for further discussion that will improve the efficiency and rigor of biomarker clinical utility validation research.
Acknowledgements:
The authors thank Dr. Steve Skates for insightful discussion during the manuscript preparation. The paper is dedicated to Dr. Pierre P. Massion, who provided valuable discussion of his proposed lung cancer early detection trial and critical review of the manuscript. Dr. Massion passed away before the final completion of this work.
This work is supported by National Institutes of Health (U24 CA086368 to Y.Z., Y.H., Y-Q.Z., T.M., and Z.F; R01 CA236558 to Y.Z.; U01CA194733 to S.H. and Z.F.; U01CA213285 and U01 CA200468 to S.H.; U24 CA086368, U01 CA271887, U01 CA230694 and R01 CA222900 to A.S., Consortium for Study of Chronic Pancreatitis, Diabetes and Pancreatic Cancer (CPDPC) U01 DK126365 and for S.T.C). S.T.C. and Z.F. were additionally supported by the Pancreatic Cancer Action Network. S.H. was additionally supported Cancer Prevention & Research Institute of Texas (CPRIT; RP180505 to S.H.) and the generous philanthropic contributions to The University of Texas MD Anderson Cancer Center Moon Shots Program and the Lyda Hill Foundation.
Role of the funder:
The funders had no role in the preparation, review, or approval of the manuscript and decision to submit the manuscript for publication.
Footnotes
Conflict of Interest Disclosure Statement
A.S. has served as a consultant for Exact Sciences, Fuji Film Medical Sciences, Glycotest, and GRAIL. Z.F has served on Scientific Advisory Board for Guardant Health. S.H. has filed sl IPs for biomarkers for several clinical indications.
Bibliography
- 1.(US) IoM. Genome-Based Diagnostics: Clarifying Pathways to Clinical Use: Workshop Summary. Washington DC: National Academies Press (US); 2012. [PubMed] [Google Scholar]
- 2.Kramer BS, Berg CD, Aberle DR, Prorok PC. Lung cancer screening with low-dose helical CT: results from the National Lung Screening Trial (NLST). Volume 18: SAGE Publications Sage; UK: London, England; 2011. p 109–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Pepe MS, Etzioni R, Feng Z, Potter JD, Thompson ML, Thornquist M, et al. Phases of biomarker development for early detection of cancer. Journal of the National Cancer Institute 2001;93(14):1054–61 doi 10.1093/jnci/93.14.1054. [DOI] [PubMed] [Google Scholar]
- 4.Pepe MS, Feng Z, Janes H, Bossuyt PM, Potter JD. Pivotal evaluation of the accuracy of a biomarker used for classification or prediction: standards for study design. Journal of the National Cancer Institute 2008;100(20):1432–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Pepe MS, Janes H, Li CI, Bossuyt PM, Feng Z, Hilden J. Early-Phase Studies of Biomarkers: What Target Sensitivity and Specificity Values Might Confer Clinical Utility? Clin Chem 2016;62(5):737–42 doi 10.1373/clinchem.2015.252163. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Pinsky P, Lange J, Etzioni R. Estimating stage-specific sensitivity for cancer screening tests. Journal of Medical Screening 2023;30(2):69–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Menon U, Gentry-Maharaj A, Burnell M, Singh N, Ryan A, Karpinskyj C, et al. Ovarian cancer population screening and mortality after long-term follow-up in the UK Collaborative Trial of Ovarian Cancer Screening (UKCTOCS): a randomised controlled trial. The Lancet 2021;397(10290):2182–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Cuzick J, Cafferty FH, Edwards R, Møller H, Duffy SW. Surrogate endpoints for cancer screening trials: general principles and an illustration using the UK Flexible Sigmoidoscopy Screening Trial. Journal of Medical Screening 2007;14(4):178–85. [DOI] [PubMed] [Google Scholar]
- 9.Hayes DF. Defining clinical utility of tumor biomarker tests: a clinician’s viewpoint. Journal of Clinical Oncology 2021;39(3):238–48. [DOI] [PubMed] [Google Scholar]
- 10.Kanwal F, Singal AG. Surveillance for Hepatocellular Carcinoma: Current Best Practice and Future Direction. Gastroenterology 2019;157(1):54–64 doi 10.1053/j.gastro.2019.02.049. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Chari ST, Maitra A, Matrisian LM, Shrader EE, Wu BU, Kambadakone A, et al. Early Detection Initiative: A randomized controlled trial of algorithm-based screening in patients with new onset hyperglycemia and diabetes for early detection of pancreatic ductal adenocarcinoma. Contemp Clin Trials 2021;113:106659 doi 10.1016/j.cct.2021.106659. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Hackshaw A, Berg CD. An efficient randomised trial design for multi-cancer screening blood tests: nested enhanced mortality outcomes of screening trial. The Lancet Oncology 2021;22(10):1360–2. [DOI] [PubMed] [Google Scholar]
- 13.Wald NJ. The treatment of Helicobacter pylori infection of the stomach in relation to the possible prevention of gastric cancer. IARC Helicobacter pylori Working Group Helicobacter pylori Eradication as a Strategy for Preventing Gastric Cancer Lyon, France: International Agency for Research on Cancer (IARC Working Group Reports, No 8) 2014:174–80. [Google Scholar]
- 14.Sullivan FM, Mair FS, Anderson W, Armory P, Briggs A, Chew C, et al. Earlier diagnosis of lung cancer in a randomised trial of an autoantibody blood test followed by imaging. European Respiratory Journal 2021;57(1). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Rosenthal AN, Fraser LSM, Philpott S, Manchanda R, Burnell M, Badman P, et al. Evidence of Stage Shift in Women Diagnosed With Ovarian Cancer During Phase II of the United Kingdom Familial Ovarian Cancer Screening Study. J Clin Oncol 2017;35(13):1411–20 doi 10.1200/JCO.2016.69.9330. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Hu P, Zelen M. Planning clinical trials to evaluate early detection programmes. Biometrika 1997;84(4):817–29. [Google Scholar]
- 17.Hernán MA, Wang W, Leaf DE. Target trial emulation: a framework for causal inference from observational data. Jama 2022;328(24):2446–7. [DOI] [PubMed] [Google Scholar]
- 18.García-Albéniz X, Hsu J, Hernán MA. The value of explicitly emulating a target trial when using real world evidence: an application to colorectal cancer screening. European journal of epidemiology 2017;32:495–500. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Lakhani DA, Chen S-C, Antic S, Muterspaugh A, Cook C, Liu N, et al. Establishing a cohort and a biorepository to identify biomarkers for early detection of lung cancer: the Nashville Lung Cancer Screening Trial Cohort. Annals of the American Thoracic Society 2021;18(7):1227–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Lord SJ, Irwig L, Simes RJ. When is measuring sensitivity and specificity sufficient to evaluate a diagnostic test, and when do we need randomized trials? Annals of internal medicine 2006;144(11):850–5. [DOI] [PubMed] [Google Scholar]
- 21.Committee MSA. Guidelines for the assessment of diagnostic technologies. Canberra: MSAC; 2005. [Google Scholar]
- 22.Merlin T, Lehman S, Hiller JE, Ryan P. The “linked evidence approach” to assess medical tests: a critical analysis. International Journal of Technology Assessment in Health Care 2013;29(3):343–50. [DOI] [PubMed] [Google Scholar]
- 23.Harris RP, Helfand M, Woolf SH, Lohr KN, Mulrow CD, Teutsch SM, et al. Current methods of the US Preventive Services Task Force: a review of the process. American journal of preventive medicine 2001;20(3):21–35. [DOI] [PubMed] [Google Scholar]
- 24.Trikalinos TA, Siebert U, Lau J. Decision-analytic modeling to evaluate benefits and harms of medical tests: uses and limitations. Medical Decision Making 2009;29(5):E22–E9. [DOI] [PubMed] [Google Scholar]
- 25.Etzioni R, Gulati R, Patriotis C, Rutter C, Zheng Y, Srivastava S, et al. Revisiting the standard blueprint for biomarker development to address emerging cancer early detection technologies. JNCI: Journal of the National Cancer Institute 2024;116(2):189–93. [DOI] [PMC free article] [PubMed] [Google Scholar]


