Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 May 28.
Published in final edited form as: Stat Med. 2025 Sep;44(20-22):e70262. doi: 10.1002/sim.70262

Bayesian Adaptive Enrichment Design for Continuous Biomarkers

Yue Tu 1, Yusha Liu 2, Wendy J Mack 1, Lindsay A Renfro 1,3
PMCID: PMC13215621  NIHMSID: NIHMS2173189  PMID: 40947424

Abstract

With the advent of precision medicine and targeted therapies in cancer, new challenges in the statistical design of clinical trials have naturally emerged. Most randomized clinical trial designs incorporating predictive biomarkers (those associated with treatment efficacy) assume biomarkers are dichotomous, or dichotomize naturally continuous biomarkers upfront, or find cut points mid-way through the trial to classify patients as biomarker-positive or biomarker-negative. However, these practices ignore or discard information about continuous and possible nonlinear or non-monotone prognostic or predictive effects. In this article, we propose a novel adaptive enrichment trial design to handle continuous biomarkers with any effect shape, including Bayesian marker-adaptive randomization. We demonstrate that this design can correctly make marker-specific trial decisions with high efficiency, resulting in improved performance and patient-centered decisions compared to adaptive cut-point selection approaches without adaptive randomization that further ignore or oversimplify true underlying marker relationships.

Keywords: adaptive enrichment design, adaptive randomization, Bayesian adaptive design, biomarker-driven design, clinical trial, precision medicine

1 ∣. Introduction

Cancer was once viewed as relatively homogeneous in terms of treatment strategy according to the type or location of malignancy and the stage of disease. But with the rapid development of gene sequencing techniques, some cancers are now better understood to be heterogeneous across specific genomically or biologically defined characteristics, defining patient subpopulations. In turn, this has led to a new treatment paradigm that has shown success and promise in certain clinical settings. However, treatments based on molecular targeting, an approach known as “precision medicine,” have brought about new challenges in statistical design of clinical trials. Specifically, novel clinical trial designs have been developed to reveal or validate subpopulations in which an experimental treatment has enhanced benefit.

In this paper, we specifically focus on adaptive enrichment trial designs, which are randomized clinical trials of a targeted therapy allowing mid-trial restriction to subpopulation(s) that show signs of benefit. Adaptive enrichment designs have the potential to reduce the trial period time by identifying subpopulation(s) that would benefit from the treatment, as well as the potential to validate putative predictive biomarkers, both of which are crucial objectives in the medical treatment research and development process. Any improvements in these areas will translate to significant resource and time savings and avoidance of assigning too many patients to a less efficacious treatment based on marker status.

Currently, however, the literature and practice commonly focus on modeling biomarkers as binary variables that are either defined at trial initiation or dichotomized mid-trial as part of algorithmic searches for cutpoints or thresholds yielding a single subpopulation with enhanced treatment effect. Work by Brannath et al., Jenkins et al., and Krisam and Kieser requires pre-specification of a marker-positive subpopulation before patient enrollment, and Renfro et al.’s work applied a grid search for an optimal marker cutpoint at an interim analysis assuming monotonicity and linearity of marker-treatment interaction effects [1-4]. Since many biomarkers are naturally continuous, there is sufficient evidence to believe that designs incorporating continuous marker modeling—and specifically, those that handle potentially non-monotonic prognostic and predictive relationships between biomarkers, treatment effect, and outcome—may be more sensitive, accurate, and efficient than crude dichotomization-based enrichment designs. Liu et al. (2022) developed a randomized trial design that uses Bayesian modeling to assess a continuous biomarker’s potentially nonlinear or non-monotone prognostic or predictive effects at interim analyses [5]. While novel for its handling of continuous markers, their design focuses on time-to-event outcomes and, crucially, utilizes biomarker information to identify regions of the biomarker range where the treatment effect can still effectively be dichotomized into positive and negative subgroups. This approach, though more flexible than much of the literature, still ultimately guides trial decisions based on a crude categorization of the biomarker range, rather than treating marker values as a continuum for adaptation. To this end, there remains a critical need for clinical trial designs where continuous biomarkers are treated as such throughout the entire decision-making process, and where decisions (such as efficacy, futility, and inferiority) are made directly for specific points along the continuous marker scale, using all the information the biomarker presents, whether the marker is predictive, prognostic, both, or neither.

We develop a novel clinical trial design that includes these features. Additionally, within an ongoing Bayesian assessment of the biomarker’s predictive value and corresponding statistical uncertainty, we apply a marker-specific adaptive randomization procedure. In contrast to traditional equal randomization, response adaptive randomization utilizes the currently estimated posterior information for the (continuous and potentially nonlinear) marker and treatment effects, assesses these against incoming patients’ biomarker profiles, and assigns them to the optimal treatment arms in accordance with their marker values. Extensive simulations are conducted to evaluate the trial design’s performance compared to a common type of frequentist design without adaptive randomization that identifies a marker subgroup at interim or final analysis timepoints using a grid search algorithm [4].

2 ∣. Notation and Models

In order to study the effect of an investigated treatment on a binary endpoint Y, a two-arm randomized clinical trial is conducted. We use Z to represent the treatment assignment, with Z=1 indicating a patient is assigned the experimental treatment and Z=0 for assignment to the control treatment. We assume that a continuous baseline biomarker X exists with values ranging from a to b for some a, bR1, which is potentially prognostic for the disease outcome and predictive for treatment effect based on previous knowledge. Assume the biomarker information is available from each patient at baseline.

2.1 ∣. Model Formulation

We extended Liu et al.’s Bayesian framework for modeling continuous biomarker effects [5] to handle a binary endpoint. To model the effect of the continuous biomarker X and the investigated treatment Z on the binary response outcome Y, a logistic regression model is used, which can be expressed as

logit(P(Y=1))=log(P(Y=1)1P(Y=1))=f(x)+g(x)z (1)

where f(x) represents the marker’s prognostic effect, and g(x) represents the marker’s predictive treatment effect. Cubic B-splines [6] are used to model f(x) and g(x), which allows non-linear and non-monotonic marker effects. We first define a knot sequence τ:τ1,,τM+8 such that

a=τ1=τ2=τ3=τ4<τ5<<τM+4<τM+5=τM+6=τM+7=τM+8=b (2)

Then, we can express f(x) and g(x) for a patient with marker value x as:

f(x)=m=1M+4Bm(x)ηf,m,g(x)=m=1M+4Bm(x)ηg,m (3)

where B1(x),,BM+4(x) are the cubic B-spline basis functions defined by the knot sequence τ and ηf,m,ηg,m are the corresponding coefficients. We can choose the M interior knots τ5,,τM+4 of τ based on the observed biomarker values; for example, using the biomarker’s sample quantiles. We use the same cubic B-spline basis functions for f(x) and g(x).

Model (3) can be expressed for n patients with marker values x1,,xn as below:

f(x)=Bηf,g(x)=Bηg (4)

where x=(x1,,xn), f(x)=(f(x1),,f(xn)), g(x)=(g(x1),,g(xn)), B is the n×(M+4) cubic B-spline design matrix with the (i,m)th entry Bim=Bm(xi), ηf=(ηf,1,,ηf,M+4), and ηg=(ηg,1,,ηg,M+4). To prevent overfitting, we add an L2 penalty on the integral of the squared second-order derivative of f(x) and g(x) [6], which is equivalent to placing the following priors on ηf and ηg in a Bayesian framework:

p(ηf)exp(12σf2ηfΛηf),p(ηg)exp(12σg2ηgΛηg) (5)

where Λ is the a (M+4)×(M+4) penalty matrix composed of the integrated product of second-order derivative of Bm(x) and Bm(x), with σf2 and σg2 as regularization parameters.

Since rank(Λ)=M+2 [6], let d1,,dM+2 denote the M+2 positive eigenvalues of Λ after spectral decomposition, and WΛ denote the (M+4)×(M+2) matrix formed by the corresponding eigenvectors of Λ. We can then rewrite (4) as the following:

f(x)=1nβf,1+xβf,2+WBvfg(x)=1nβg,1+xβg,2+WBvg (6)

where 1n is an n×1 vector consisting of 1’s, βf,1, βf,2, βg,1 and βg,2 are scalars having non-informative Gaussian priors βf,1, βf,2, βg,1, βg,2N(0,H2), where H is a very large number, vf and vg are (M+2)×1 random vectors with mutually independent Gaussian priors vfMVN(0,σf2IM+2) and vgMVN(0,σg2IM+2), and WB=BWΛdiag(d112,,dM+212). The Gaussian priors placed on βf,1, βf,2, βg,1, βg,2, vf and vg reflect the L2 penalty on ηf and ηg in (5) after matrix transformation. A noninformative inverse Gamma prior is used for σf2 and σg2:σf2IG(a0,f,b0,f), σg2IG(a0,g,b0,g).

2.2 ∣. Posterior Computation

Let D={(yi,xi,zi); i=1,,n} be the patient data collected during the study at an interim analysis or at the end of the study, where yi is equal to 1 if patient i is observed to have a response and equal to 0 if no response, xi is the patient is biomarker value, and zi is patient is treatment assignment. Let θf=(βf,1,βf,2,vf), and θg=(βg,1,βg,2,vg), which denote the logistic regression model parameters. With appropriate priors, the model parameters are updated at each MCMC iteration. Given (1) and (6), the joint posterior for (θf, θg, σf2, σg2) given data D is proportional to the full likelihood times the joint prior:

π(θf,θg,σf2,σg2D)L(θf,θgD)×p(vfσf2)p(vgσg2)p(σf2)p(σg2)×p(βf,1)p(βf,2)p(βg,1)p(βg,2)i=1n{[exp(f(xi)+g(xi)zi)1+exp(f(xi)+g(xi)zi)]yi}×{[11+exp(f(xi)+g(xi)zi)]1yi}×m=1M+2[1σfexp(vf,m22σf2)]×m=1M+2[1σgexp(vg,m22σg2)]×exp(βf,122H2)×exp(βf,222H2)×exp(βg,122H2)×exp(βg,222H2)×(σf2)a0,f1exp(b0,fσf2)×(σg2)a0,g1exp(b0,gσg2) (7)

The parameters σf2 and σg2 have a conjugate full conditional distribution and are updated by Gibbs sampling, while θf and θg do not have a closed-form full conditional distribution so are updated by adaptive Metropolis sampling [7]. The posterior samples of θf and θg can also be used for Bayesian inference, along with the credible intervals for f(x) and g(x) with a prespecified significance level.

The MCMC chain is initialized with parameter estimates from maximum-likelihood estimation (MLE) of the logistic regression model. At early stages of the trial with a small sample size observed, it is possible to have complete or quasi–complete separation, which may be due to perfect prediction by continuous covariates. There are some solutions to handle this problem, including exact logistic regression and Firth’s bias-reduced logistic regression [8]. We decide to impose an L2 penalty on model parameter estimation when we initialize the Metropolis random walk process. The shrinkage on parameter estimation at the initialization step is applied throughout the trial. Let Xf=(1,x,vf) and Xg=(1,x,vg), which are (M+4)×1 vectors; then the variance of θf and θg can be expressed as:

Var(θf)=(XfWfXf+λIM+4)1XfWf[Var(Zf)]WfXf(XfWfXf+λIM+4)1Var(θg)=(XgWgXg+λIM+4)1XgWg[Var(Zg)]WgXg(XgWgXg+λIM+4)1 (8)

where

Var(Zf)=Wf1Var(Y)Wf1=Wf1Var(Zg)=Wg1Var(Y)Wg1=Wg1 (9)

λ is the regularization parameter, Var(Zf) and Var(Zg) are variances of the parameter estimates without penalty, and Wf and Wg are (M+4)×(M+4) diagonal matrices with Wi representing the weight for the ith patient, which is the estimated probability of response [9]. Var(θf) and Var(θg) are the proposal variances for the Metropolis random walk’s estimation of θf and θg.

2.3 ∣. Adaptive Randomization

In this section, we describe how the estimated marker effects to date will be used to adaptively update randomization rules at a series of interim analyses, where the randomization probabilities to experimental versus control arms for a newly enrolled patient with a given marker value of x will be updated periodically to preferentially assign that patient to the arm showing benefit. The probability of assignment to an arm is derived from the posterior evidence accumulated to date in the trial.

The randomization ratio for a patient with marker value x will be determined based on the estimated probability of response in each arm. To obtain the probability of response based on the logistic model, we need to invert the logistic function, defined as logit−1:

logit1(x)=ex1+ex=P(Y=1) (10)

In our context, the above equation represents the predictive response rate for a patient with marker value x.

Assume a trial has randomized n patients and observed the response for all enrolled patients, with data available to estimate the posterior distribution of f(x) and g(x), which are the marker’s prognostic and predictive effects, respectively. Let f^(x) and g^(x) denote the estimated posterior mean for the marker’s observed prognostic and predictive effects. Then, for a newly enrolled (n+1)th patient with marker value xn+1, we can write the predicted response rate in the control arm as

R^ctl,n+1=logit1(f^(xn+1)) (11)

and the predicted response rate in the experimental arm as

R^exp,n+1=logit1(f^(xn+1)+g^(xn+1)) (12)

Intuitively, we can assign the (n+1)th patient to the experimental arm with a randomization probability based on the ratio of the point estimate of the response rate in the treatment arm over the sum of the response rates in both arms:

Pexp,n+1=R^exp,n+1R^exp,n+1+R^ctl,n+1 (13)

while the randomization probability to the control arm is 1Pexp,n+1. However, this approach does not consider the uncertainty of marker effect estimation. Instead, we use the posterior credible interval for the currently estimated predictive marker effect, representing the superiority of one treatment over the other for patients with similar marker values, to determine the assignment probabilities for a patient with marker value xn+1. Given the observed data, the posterior probability that the experimental arm has a higher response rate than the control arm for this patient can be expressed as Pr(P^exp,n+1>P^crl,n+1D). Then, we can assign the (n+1)th patient to the experimental arm with probability

Pexp,n+1=Pr(R^exp,n+1>R^ctl,n+1D)cPr(R^exp,n+1>R^ctl,n+1D)c+Pr(R^ctl,n+1>P^exp,n+1D)c (14)

where c is the tuning parameter that controls how much the randomization probability is influenced by the data. Following the recommendation from Thall and Wathen, c can be chosen to be 0.5 or n2N, where n is the number of currently enrolled patients and N is the total planned sample size for the trial [10].

Since we need to observe sufficient data to inform the model posteriors for subsequent randomization adaptations, we need a run-in phase, during which patients are randomized based on a standard 1:1 ratio. Also, to allow data accumulation between updates, we can set the frequency at which the marker effect posteriors and marker-specific randomization probabilities are updated (while maintaining the same marker-specific randomization probabilities between updates). According to the simulation study conducted by Wathen and Thall, in order to have desirable selection properties, the adaptive randomization probabilities should be restricted to a certain interval [rmin, rmax], for example, [0.1,0.9] [10]. Implementation of the adaptive randomization mechanism via simulation can be found in the Appendix A.

3 ∣. Trial Design

We propose a trial design, which combines the features of covariate-adaptive and response-adaptive randomization, given currently enrolled patient responses, and adaptive enrichment if accrual for patients with certain marker values is terminated early for efficacy, inferiority, or futility. Initially, all marker values are classified in the “continuing marker group,” and there is no restriction on enrollment. As sufficient information is obtained to conclude some ranges of marker values are associated with efficacy, inferiority, or futility of the experimental treatment, subsequent accrual will be restricted to patients defined in the “continuing marker group,” which is updated throughout the trial. We assume a one-to-one randomization ratio to each treatment arm before the start of adaptive randomization. After the run-in phase, patients will be adaptively randomized based on the marker-specific randomization ratio calculated by the estimated penalized splines model with the accumulated patient information. Suppose there are K adaptive randomization blocks and we perform K1 interim analyses during the trial. Let Dk={(yi,xi,zi);i=1,,nk} be the patient data observed by analysis k1,,K, where K is the final analysis. At an interim check k, which occurs at the end of a randomization block, we first check whether there is sufficient evidence of early efficacy, inferiority, or futility for the experimental arm among marker values in the continuing marker group. If this occurs, the definition of “continuing marker group” will be updated, and patients with those marker values will no longer be enrolled. Instead, the trial will enrich and continue with enrollment restricted to the continuing marker group to collect more evidence on the remaining marker values’ relationships with treatment and outcome. The trial terminates once there is sufficiently precise evidence for all marker values’ associations with efficacy, inferiority, or futility, or when the planned total sample size has been reached and outcomes observed for all patients. Patients previously enrolled with marker values in the stopped marker group will continue to contribute to subsequent modeling for estimation purposes. We assume that a patient’s response can be observed shortly after receiving a treatment assignment.

3.1 ∣. Interim/Final Analysis Algorithm

At the end of each adaptive randomization block k1,,K corresponding to an interim or final analysis time point, our proposed trial design’s decision algorithms proceed as follows.

Step 1: Update the full posterior marker model:

Using all data observed up to the current check k, we fit or re-fit the full penalized splines model as described in Subsection 2.3 to evaluate the treatment effect based on the continuous marker X.

Step 2a: Interim efficacy check:

If k1,,K1, we check for early efficacy of the experimental treatment effect among the continuing marker group Xc by comparing the posterior probability of g(x)>0 given x with a prespecified efficacy threshold. We define Xeff,k as Xeff,k{x[a,b]:P(g(x)>0Dk)>Peff}, where Peff is a prespecified efficacy threshold. We stop recruitment for any patients with marker values in Xeff,k due to early efficacy and remove Xeff,k from Xc.We add Xeff,k to the set of Xeff, where Xeff={x:xj=1k1Xeff,j}. Xeff then includes treatment sensitive marker values identified up to the current check k. Then, go to Step 3.

Step 2b: Final efficacy check:

If k=K, we check for final efficacy of the experimental treatment effect among the continuing marker group Xc by comparing the posterior probability of g(x)>0 given x with a prespecified efficacy threshold. We define Xeff,K as Xeff,K{X[a,b]:P(g(x)>0DK)>Peff}, where Peff is a prespecified efficacy threshold. We add Xeff,K to the set of Xeff, where Xeff={x:xj=1K1Xeff,j}. Xeff then includes treatment-sensitive marker values showing efficacy at any point in the trial. For marker values in Xc but not in Xeff, final futility is claimed.

Step 3: Interim inferiority check:

If k1,,K1, we check for early inferiority of the experimental treatment effect among the continuing marker group Xc by comparing the posterior probability of g(x)>0 given x with a prespecified inferiority threshold. We define Xinf,k as Xinf,k{x[a,b]:P(g(x)>0Dk)<Pinf}, where Pinf is a prespecified inferiority threshold. We stop recruitment for patients with marker values in Xinf,k for early inferiority and remove Xinf,k from Xc. We add Xinf,k to the set of Xinf, where Xinf={x:xj=1k1Xinf,j}. Xinf then includes inferior experimental treatment effect marker values identified up to the current check k. Then, we go to Step 4.

Step 4: Interim futility check:

If k1,,K1, we check for early futility of the experimental treatment effect among the continuing marker group Xc based on the posterior predictive probability of trial success given x. We calculate the marker-specific posterior predictive probability of trial success by simulating subsequent trial completion in a loop r, for r1,,R. For each simulated iteration r, we randomly sample the remaining patients with marker values from Xc uniformly, which, combined with the current observed data, add up to the total planned sample size N. Based on the currently estimated posterior distribution and marker-specific randomization probabilities, we assign patients to treatment arms and simulate outcomes for remaining patients yet to have observed outcomes in the trial. The posterior penalized splines model is re-fitted with observed data along with simulated data. For marker values in Xc, we calculate an efficacy indicator variable Xc,r. We define Xc,r=1 if xXc:P(g(x)>0Dr)>Peff, where a chosen efficacy threshold Peff is prespecified; Otherwise, Xc,r=0. Then, for each marker value in Xc, we take the mean of the number of marker-specific efficacy indicators observed over R, denoted by Phit,x=r=1RXc,rR. If Phit,x<Pfut, where Pfut is set to be some low probability, we claim early futility for patients with those marker values and add those to the set of Xfut,k. We stop recruitment for patients with marker values in Xfut,k due to early futility and remove Xfut,k from Xc. We add Xfut,k to the set of Xfut, where Xfut={x:xj=1k1Xfut,j}. Xfut then includes futile experimental treatment effect marker values identified up to the current check k. Then, we go to Step 5.

Step 5: Update the adaptive randomization model and continue recruitment

If there are no more marker values in the set of continuing marker group Xc, we terminate the study early.

Otherwise, if k1,,K1, we continue recruitment only for patients with marker values in the set of continuing marker group Xc and enroll patients up to the next randomization block. For each subsequently enrolled patient, we adaptively randomize the patient based on their marker value at enrollment, and the marker-specific randomization ratio is calculated using the current estimated full penalized splines model at this current randomization block.

The proposed trial design schema can be found in Figure 1.

FIGURE 1 ∣.

FIGURE 1 ∣

Proposed Bayesian enrichment trial with adaptive randomization design schema.

3.2 ∣. Comparator Frequentist Design

We compare our proposed trial design with a traditional frequentist adaptive enrichment design in order to assess the advantages of modeling the truly continuous marker effect as continuous and using adaptive randomization. The comparator design will treat the marker effect as linear and monotonic, as is commonly done [11-14]. At a single interim analysis, a grid search algorithm will be used to decide if the marker by treatment interaction is sufficiently significant, and if so, the optimal marker threshold will be used to dichotomize the patient population into marker positive and negative subgroups. Then, early efficacy, inferiority, and futility analyses will be conducted in each marker subgroup, or in the overall group if no optimal marker cutoff point is identified. If one subgroup is stopped early, the remaining sample size will be allocated to the continuing subgroup. When we dichotomize the patient population based on a marker cutpoint, we define an indicator variable Xd, where Xd=1 denotes patients in the marker-positive subgroup and Xd=0 denotes patients in the marker-negative subgroup based on the cutpoint. We assume a one-to-one randomization ratio during the entire trial. Suppose there are K analyses during the trial and let Dk={(yi,xi,zi); i=1,,nk} be the patient data observed by timepoint k1,,K, where K is the final analysis. Also, let Dk,Xd be the patient data observed by timepoint k in the marker subgroup Xd. At an interim or final analysis time point k, the comparator design proceeds as follows.

Step 1: Interim check for identifying marker subgroup:

If the marker cut-off point was identified and defined prior to k and one marker subgroup was stopped previously due to early efficacy, inferiority or futility, continue to Step 2a and perform the following checks for the remaining marker subgroup.

If there were no early stopped marker subgroups, the optimal marker cut-off point would be identified using a grid search algorithm from equally spaced cut-off points within the range of possible marker values. In general, patients showing a favorable treatment effect will comprise the marker-positive subgroup, and the remaining patients the marker-negative subgroup. At each potential cutpoint, a logistic regression model is fitted with treatment assignment, an indicator of marker subgroup based on the potential cut-off point, and a marker by treatment interaction as covariate terms based on Dk. Among all potential cutpoints, if the smallest p-value for the marker by treatment interaction term is less than Pint,freq, where Pint,freq is a prespecified threshold, we will dichotomize the patient population using the cut-off point that results in the smallest p-value and classify patients into marker positive and marker negative subgroups denoted by the binary indicator Xd. Then, we go to Step 2a and evaluate each marker subgroup independently. If no cut-off point is identified, we go to Step 2c and evaluate patients as an overall group.

Step 2a: Interim efficacy check for marker subgroup:

If k1,,K1, we check for early subgroup efficacy of the experimental treatment in subgroup Xd. A logistic regression model is fitted with treatment assignment based on Dk,Xd. If the p-value of the treatment assignment term is less than Peff,grp,freq, where Peff,grp,freq is a prespecified efficacy threshold, and the estimated coefficient is greater than 0, meaning a positive treatment effect, stop recruiting patients in subgroup Xd and claim early efficacy for marker subgroup Xd; otherwise, go to Step 3a.

Step 2b: Final efficacy check for marker subgroup:

If k=K, we check for final subgroup efficacy of the experimental treatment in subgroup Xd. A logistic regression model is fitted with treatment assignment based on Dk,Xd. If the p-value of the treatment assignment term is less than Peff,grp,freq, where Peff,grp,freq is a prespecified efficacy threshold, and the estimated coefficient is greater than 0, meaning a positive treatment effect, we claim final efficacy for Xd; otherwise, we claim final futility for subgroup Xd.

Step 2c: Interim efficacy check for overall population:

If k1,,K1, we check for efficacy of the experimental treatment in the overall patient population. A logistic regression model is fitted with treatment assignment based on Dk. If the p-value of the treatment assignment term is less than Peff,all,freq, where Peff,all,freq is a prespecified efficacy threshold, and the estimated coefficient is greater than 0, meaning a positive treatment effect, stop recruiting all patients and claim early efficacy; otherwise, go to Step 3b.

Step 2d: Final efficacy check for overall population:

If k=K, we check for the final efficacy of the experimental treatment in the overall patient population. A logistic regression model is fitted with treatment assignment based on Dk. If the p-value of the treatment assignment term is less than Peff,all,freq, where Peff,all,freq is a prespecified efficacy threshold, and the estimated coefficient is greater than 0, meaning a positive treatment effect, we claim final efficacy overall; otherwise, we claim final futility overall.

Step 3a: Interim inferiority check for marker subgroup:

Check for early subgroup inferiority of the experimental treatment in subgroup Xd. A logistic regression model is fitted with treatment assignment based on Dk,Xd. If the p-value of the treatment assignment term is less than Peff,grp,freq, where Peff,grp,freq is a prespecified inferiority threshold, and the estimated coefficient is less than 0, meaning a negative treatment effect, stop recruiting patients in subgroup Xd and claim early inferiority for Xd; otherwise, go to Step 4a.

Step 3b: Interim inferiority check for overall population:

Check for early inferiority of the experimental treatment effect in the overall patient population. A logistic regression model is fitted with treatment assignment based on Dk. If the p-value of the treatment assignment term is less than Pinf,all,freq, where Pinf,all,freq is a prespecified inferiority threshold, and the estimated coefficient is less than 0, meaning a negative treatment effect, stop recruiting all patients and claim early inferiority overall; otherwise, go to Step 4b.

Step 4a: Interim futility check for marker subgroup:

If k1,,K1, check for early subgroup futility of the experimental treatment in subgroup Xd. Assuming equal randomization within the subgroup Xd, the conditional power of the test for the difference between the response rates in the experimental arm and the control arm is calculated at interim check k within Xd. The type I error rate α used in the test is determined based on Peff,grp,freq and θalter, which denotes the targeted treatment effect. If the conditional power is lower than Pfut,grp,freq, where Pfut,grp,freq is a prespecified futility threshold, stop recruiting patients in subgroup Xd and claim early futility for Xd; otherwise, we continue recruitment for patients in subgroup Xd.

Step 4b: Interim futility check for overall population:

If k1,,K1, check for early futility of the experimental treatment in the overall patient population. Assuming equal randomization, the conditional power of the test for the difference between the response rates in the experimental arm and control arm is calculated at interim k in the overall population. The type I error rate of α used in the test is based on Peff,all,freq and θalter, which denotes the expected difference under the alternative hypothesis. If the conditional power is lower than Pfut,all,freq, where Pfut,all,freq is a prespecified futility threshold, stop recruiting patients with any marker values and claim early futility; otherwise, we continue recruitment for patients with any marker values.

The comparator trial design schema can be found in Figure 2.

FIGURE 2 ∣.

FIGURE 2 ∣

Comparator frequentist enrichment trial design schema.

4 ∣. Simulation Study

4.1 ∣. Scenarios and Settings

To evaluate our trial design’s performance, simulations were conducted for 12 scenarios described in Table 1, covering a range of combinations of prognostic and predictive markers and treatment effects, including scenarios when the effects are nonlinear and non-monotone. Scenario 1 is the null case. Scenario 2 has a constant treatment effect but no marker effect. Scenario 3 has a linear prognostic marker effect but no treatment effect. Scenarios 4 to 10 all have predictive marker effects with different underlying functions. Scenario 11 is an inferior treatment effect case. Scenario 12 has a linear prognostic marker effect and a nearly dichotomous predictive marker effect. The response rate is set to be 5% for the control arm, except for Scenarios 3, 11, and 12, and the maximum response rate for the experiential arm is set to be 30%, except for Scenarios 1, 11, and 12. We chose the sample size to have sufficient power to detect clinically meaningful marker effects on the level of 5% versus 30% in relatively small ranges of marker values, even though the design is “overpowered” to find an overall effect between two arms. Marker-specific response functions are selected to achieve the target shape of curves under a range of complex scenarios, and parameters are set based on the target response rate for control and treatment arms. The visual representation of simulation scenarios can be found in the Appendix B.

TABLE 1 ∣.

Simulation scenarios with descriptions, marker effect functions on the log odds ratio scale with true parameter values.

Case Scenario Control arm
f(x)
Treatment arm
f(x)+g(x)
Parameter values Control arm
Maximum
response rate
Treatment arm
Maximum
response rate
1 No treat effect
(No marker effect)
−2.95 −2.95 5% 5%
2 Constant treat effect
(No marker effect)
−2.95 Δ2.95 Δ=2.1 5% 30%
3 Prognostic marker effect
(No treat effect)
xΔ2.95 xΔ2.95 Δ=2.1 30% 30%
4 Predictive marker effect
(perfectly dichotomous)
−2.95 1{x>x0}Δ2.95 x0=0.5, Δ=2.1 5% 30%
5 Predictive marker effect
(nearly dichotomous)
−2.95 exp{30(xx0)}1+exp{30(xx0)}Δ2.95 x0=0.5, Δ=2.1 5% 30%
6 Predictive marker effect
(linear)
−2.95 xΔ2.95 Δ=2.1 5% 30%
7 Predictive marker effect
(non-linear and monotone)
−2.95 exp(7(xx0))+Δ2.95) x0=0.106, Δ=2.1 5% 30%
8 Predictive marker effect
(non-linear and monotone)
−2.95 exp(1.13x)+Δ Δ=3.95 5% 30%
9 Predictive marker effect
(non-linear and non-monotone)
−2.95 Ifxx0,0.85Δexp{30(xx1)}1+exp{30(xx1)} x0=0.5, Δ=2.1 5% 30%
−2.95 Ifx>x0,2.95Δexp{30(xx2)}1+exp{30(xx2)} x1=0.2, x2=0.8
−2.95
10 Predictive marker effect
(non-linear and non-monotone)
−2.95 Ifxx0,2.95+Δexp{30(xx1)}1+exp{30(xx1)} x0=0.5, Δ=2.1 5% 30%
−2.95 Ifx>x0,0.85Δexp{30(xx2)}1+exp{30(xx12)} x1=0.2, x2=0.8
11 Constant inferior treat effect
(No marker effect)
Δ2.95 −2.95 Δ=2.1 30% 5%
12 Prognostic and predictive marker effect
(linear; nearly dichotomous)
xΔ2.95 exp{30(xx0)}1+exp{30(xx0)}x1Δ2.95+xΔ x0=0.5, Δ=1.22
x1=1.33
15% 40%

In our simulations, N=500 is the maximum sample size, and there are two treatment arms. Adaptive randomization is started after observing results from the first 100 patients, and during the run-in phase, the randomization ratio is 1:1. The block size is set to be 50, and the tuning parameter is c=n2N. Given the total number of patients and the block size, there are 8 randomization blocks, 8 interim analyses, and a final analysis. The maximum number of allowable interior knots for the splines is 6. For simulations, we fix the pool of possible continuous marker values by a sequence from 0.01 to 0.99 with an increment of 0.01, which gives 99 unique marker values (there is no constraint on the number of possible marker values within a plausible range for general application). During the run-in phase, patients are recruited from an unrestricted population where marker values are sampled with replacement, and the median marker value is set as the reference for the prognostic effect. Once the set of continuing marker values Xc no longer represents the full population due to enriched enrollment, marker values for subsequently enrolled patients are sampled with replacement from Xc. During MCMC sampling, we discard the first 25 000 samples from the burn-in period, and then keep every 50th sample (thin = 50) and eventually collect 10 000 posterior samples of model parameters. The scaling factor to adjust the size of the random walk proposal is adaptively chosen from 0.5, 0.6, 0.9, and 1.1, depending on the regularly monitored Metropolis acceptance probabilities. To reduce computational time, the MCMC chain used during the posterior predictive-based futility check has burn-in length = 20 000, thin = 20, and 2000 posterior samples. The thresholds for the decision rules are as follows: Peff=0.975, Pinf=0.025, Pfut=0.1, R=100, rmin=0.1, rmax=0.9.

We used simulations to evaluate the performance of the comparator frequentist design as well, intentionally choosing comparable settings for a more direct comparison with our proposed design. In the simulation, up to N=500 total patients are enrolled, and there are two treatment arms. Only one interim analysis, conducted when 50% of patients are enrolled and evaluated, and one final analysis are performed. The grid search algorithm for an optimal marker cutpoint is performed on marker values from 0.01 to 0.99 by an increment of 0.01. The thresholds for the decision rules are as follows: Peff,all,freq=0.025, Peff,grp,freq=0.025, Pinf,all,freq=0.025, Pinf,grp,freq=0.025, Pfut,all,freq=0.1, Pfut,grp,freq=0.1, θalter=0.01, chosen to match those used in our proposed design as closely as possible.

4.2 ∣. Simulation Results

Simulation results with 1000 iterations are presented in the following figures and tables. In Figure 3a,b, marker-specific final trial decisions are summarized in stacked bin plots. The x-axis is the marker value, and the y-axis is the cumulative rate of each final trial decision, including efficacy (red), futility (green), and inferiority (blue), which add up to 100% for each marker value. Given that all scenarios from 1 to 10 (except Scenario 3) have no prognostic marker effects, we notice that the marker-specific separation boundaries between claiming efficacy and claiming futility have shapes similar to the predictive marker effect curves in Scenarios 1, 2, and 4–10. In Scenario 12, the marker-specific separation boundary has the same shape as the treatment effect, which is the difference between the experimental arm and the control arm due to the presence of both prognostic and predictive marker effects. The similarity between the separation boundaries and predictive marker effect curves shows that our proposed design could effectively model the predictive marker and treatment interaction.

FIGURE 3 ∣.

FIGURE 3 ∣

Rate of status by marker values for Scenarios. (a) 1–6, and (b) 7–12.

In Scenario 1, since there is no treatment effect, we expect to claim futility across marker values. Our proposed design correctly classifies overall futility about 90% on average across marker values under the null case. Of the other 10%, 7% are misclassified as overall efficacy, and 3% are misclassified as overall inferiority, which are averaged over marker values due to no predictive marker effect. In Scenario 2, there is a constant treatment effect, and almost 100% of the time, the proposed design correctly concludes overall efficacy while only 0.2% of the time concludes overall futility across marker values. Scenario 3 has a constant prognostic marker effect, but since there is no treatment effect, we expect to see a result that is similar to the null case. The proposed design correctly classifies the conclusion as overall futility about 80% on average across marker values. However, we notice that there are higher chances of incorrect conclusions around marker values closer to 0 or 1, which might be due to challenges associated with marker effect estimation near the boundaries. Scenario 4 has a predictive marker effect represented as a step function, which is perfectly dichotomous, and Scenario 5 uses a smooth step function, which is nearly dichotomous, so that the results from these two scenarios are expected to be similar. In both cases, after the marker value exceeds 0.5, where patients start to benefit from the treatment, almost all marker values are classified as having efficacy as shown in Figure 3a,b. Due to the smooth nature of the step function, the rate of efficacy is higher in Scenario 5 than in Scenario 4 as the marker value approaches 0.5 from 0. The small misclassification rate for marker values close to 0.5 in Scenario 4 is caused by the discontinuity of the step function. Scenario 6 has a linear predictive marker effect, and a small curvature in the separation boundary is observed due to our use of cubic splines to model marker effects. Scenarios 7 and 8 both have nonlinear predictive marker effect functions, and the separation boundaries of the final decision, respectively, reflect the shapes of the predictive marker effect. Scenario 9 has no treatment effect around the marker value of 0.5 and maximum treatment effects around the lowest and highest marker values, and consistently, we see low futility rates around marker values close to 0 or 1; Scenario 10 has the opposite, a bell-shaped predictive marker effect, and we see low efficacy rates around marker values close to 0 or 1. Scenario 11 has a constant inferior treatment effect with no prognostic marker effect, and the separation boundary of overall futility and overall inferiority is similar to a horizontal line. The proposed design correctly classifies overall inferiority about 87% on average across the marker values where the experimental arm is worse than the control arm. Scenario 12 has a linear prognostic marker effect and a smooth S-shaped increasing predictive marker effect, with a nearly dichotomous treatment difference. The result is similar to what we observed in Scenario 5, which is expected.

The marker-specific decision times by specific interim analyses (blocks) are shown in Figure 4a,b. In this plot, the x-axis is the marker value, and the y-axis is the cumulative rate of when the final trial decision is made, adding up to 100% for each marker value and timepoint. Since an interim analysis is conducted at the end of each randomization block, of which there are 8 in our setting, there are 8 interim analyses and a total of 9 possible stopping times, including the final analysis. Our proposed design can successfully reach marker-specific conclusions before the final analysis time point at almost all marker values in all different cases. We observe that marker values that have stopped at a later interim check (as they required more information to conclude a final marker-specific trial decision) or those that have reached the final analysis are more likely to be at the boundaries (close to 0 or 1) or to be associated with a relatively small treatment effect (smaller difference in the log of odds for having a response between the experimental arm and control arm), which is expected. For example, there are higher rates of later stopping around the boundary marker values in Scenario 2; in Scenario 9, as the treatment effect decreases when marker values get closer to 0.5, there is a higher chance that final decisions are concluded at a later interim check.

FIGURE 4 ∣.

FIGURE 4 ∣

Rate of interim stopping time by marker values for Scenarios. (a) 1–6, and (b) 7–12.

The corresponding mean sample sizes and the percentage of early trial termination are shown in Table 2. When the treatment effect is constantly inferior or superior and the effect size is large enough, the required sample size is relatively low and 100% of simulated trials can be terminated early. For example, the average final sample size is 157 for Scenario 2 and 115 for Scenario 11, which is much smaller than the pre-planned maximum sample size of 500. When a higher proportion of marker values have a large treatment effect size, the required sample size is also low with a high rate of early termination, for example, Scenario 7 with an average sample size of 294 and 88% early termination. However, when the predictive marker effect is nonlinear or there is higher variation in treatment effect across marker values, the average sample size tends to be larger. With more patients’ information required for each marker value to reach a decision, trials are more likely to run to the end, especially for marker values with a relatively small effect size.

TABLE 2 ∣.

Simulation results for the proposed trial design.

Case Scenario Average sample size Percentage of early termination
1 No treat effect
No marker effect
316.2 0.85
2 Constant treat effect
No marker effect
156.1 1.00
3 Prognostic marker effect
No treat effect
310.4 0.83
4 Predictive marker effect
(perfectly dichotomous)
410.4 0.71
5 Predictive marker effect
(nearly dichotomous)
426.5 0.60
6 Predictive marker effect
(linear)
424.4 0.56
7 Predictive marker effect
(non-linear and monotone)
293.5 0.88
8 Predictive marker effect
(non-linear and monotone)
436.0 0.51
9 Predictive marker effect
(non-linear and non-monotone)
446.6 0.49
10 Predictive marker effect
(non-linear and non-monotone)
447.2 0.50
11 Constant inferior treat effect
(No marker effect)
114.5 1.00
12 Prognostic and predictive marker effect
(linear; nearly dichotomous)
416.3 0.62

Simulation results for the comparator design with 1000 iterations are presented in Table 3. For each scenario, the rates of marker detection are reported, indicating if sufficiently strong evidence existed for dichotomization of the full population into positive/negative marker subgroups. The percentages of each possible trial decision are also presented, which include overall efficacy, overall inferiority, or overall futility when no marker is detected; if marker subgroups are identified, efficacy, inferiority, or futility can be concluded in the positive or negative marker subgroups. Average sample sizes and average marker prevalences, defined as the proportion of marker-positive patients in the full population when a marker is detected, are also reported. For Scenario 1, the comparator design correctly concludes overall futility 85% of the time, while our proposed design does so about 90% of the time on average across marker values. In Scenario 2, the comparator design correctly concludes overall efficacy 89% of the time, while our proposed design does so about 100% of the time on average across marker values. The comparator design concludes overall futility 85% of the time under Scenario 3, which performs better than our proposed design. Scenarios 4 and 5 are “ideal” cases for the comparator design due to the perfectly dichotomous or nearly dichotomous predictive marker effects, and the comparator design identifies a marker around 99% and 98% of the time for Scenarios 4 and 5, respectively. In Scenario 6, the most common final decision is efficacy in the marker positive subgroup and futility in the marker negative subgroup, which occurs 76% and 81% of the time that a marker is identified, respectively. Since the average marker prevalence is around 55%, the selected marker cut-off point is around 0.55 on average. However, patients with marker values close to 0.5 could also benefit from the treatment, which can be detected in our proposed design. Under Scenario 7, patients with marker values greater than 0.12 have the maximum treatment effect, and our proposed design is able to detect that. However, the comparator design claims marker-negative subgroup futility 44% of the time, with the selected marker cut-off point around 0.34 on average. This means that patients with marker values less than 0.34 will be classified as non-responding, but in fact patients with marker values between 0.12 and 0.34 also benefit from the experimental arm; in other words, under the comparator design, around 20% of the full population will miss out on an efficacious treatment. Similarly, in Scenario 8, patients with marker values greater than 0.25 start to benefit from the experimental treatment, but the comparator design claims marker-negative subgroup futility 78% of the time, with the selected marker cut-off around 0.5. In Scenarios 9 and 10, nonlinear and non-monotone predictive effects are present, which can be detected by our design but not the comparator design. The comparator design is most likely to claim efficacy in the marker-positive subgroup and futility in the marker-negative subgroup, with rates of 75% and 81% in Scenarios 9 and 10, respectively. A marker is detected 87% and 82% of the time for that design, respectively, with corresponding cut-off points of 0.26 and 0.77, leading to misclassification of responders into non-responders with marker values close to 1 in Scenario 9 and nonresponders into responders with marker values close to 0 in Scenario 10. Due to the nonlinear and non-monotonic nature of predictive marker effects under Scenarios 9 and 10, the comparator design could not perform well by searching for a marker cut-off value and dichotomizing the patient population. In Scenario 11, the comparator design correctly concludes overall inferiority 92% of the time. For Scenario 12, the comparator design correctly detects a marker 84% of the time, with a cut-off point around 0.5, which is a little worse than its performance in Scenario 5 due to the presence of the prognostic marker effect.

TABLE 3 ∣.

Simulation results for the comparator frequentist trial design.

Scenario Possible final
decision
Overall trial
decision rate
Average
sample size
Average marker
prevalence
1: No treat effect, No marker effect No Marker Detected 0.713 N = 254.8
Marker Detected 0.287 0.505
Overall Efficacy 0.010
Overall Inferiority 0.017
Overall Futility 0.846
M+ Eff/M− Fut 0.063
M+ Eff/M− Inf 0.005
M+ Fut/M− Eff
M+ Fut/M− Inf 0.059
M+ Inf/M− Eff
M+ Inf/M− Fut
2: Constant treat effect, No marker effect No Marker Detected 0.726 N = 250
Marker Detected 0.274 0.505
Overall Efficacy 0.891
Overall Inferiority
Overall Futility
M+ Eff/M− Fut 0.109
M+ Eff/M− Inf
M+ Fut/M− Eff
M+ Fut/M− Inf
M+ Inf/M− Eff
M+ Inf/M− Fut
3: Prognostic marker effect, No treat effect No Marker Detected 0.302 N = 253.5
Marker Detected 0.698 0.472
Overall Efficacy 0.006
Overall Inferiority 0.008
Overall Futility 0.859
M+ Eff/M− Fut 0.066
M+ Eff/M− Inf 0.001
M+ Fut/M− Eff
M+ Fut/M− Inf 0.060
M+ Inf/M− Eff
M+ Inf/M− Fut
4: Predictive marker effect, (perfectly dichotomous) No Marker Detected 0.009 N = 251
Marker Detected 0.991 0.515
Overall Efficacy 0.015
Overall Inferiority
Overall Futility
M+ Eff/M− Fut 0.947
M+ Eff/M− Inf 0.037
M+ Fut/M− Eff
M+ Fut/M− Inf 0.001
M+ Inf/M− Eff
M+ Inf/M− Fut
5: Predictive marker effect, (nearly dichotomous) No Marker Detected 0.025 N = 251
Marker Detected 0.975 0.527
Overall Efficacy 0.030
Overall Inferiority
Overall Futility
M+ Eff/M− Fut 0.946
M+ Eff/M− Inf 0.024
M+ Fut/M− Eff
M+ Fut/M− Inf
M+ Inf/M− Eff
M+ Inf/M− Fut
6: Predictive marker effect, (linear) No Marker Detected 0.189 N = 254.25
Marker Detected 0.811 0.545
Overall Efficacy 0.229
Overall Inferiority
Overall Futility 0.008
M+ Eff/M− Fut 0.759
M+ Eff/M− Inf 0.004
M+ Fut/M− Eff
M+ Fut/M− Inf
M+ Inf/M− Eff
M+ Inf/M− Fut
7: Predictive marker effect, (non-linear and monotone) No Marker Detected 0.428 N = 250.5
Marker Detected 0.572 0.665
Overall Efficacy 0.558
Overall Inferiority
Overall Futility
M+ Eff/M− Fut 0.442
M+ Eff/M− Inf
M+ Fut/M− Eff
M+ Fut/M− Inf
M+ Inf/M− Eff
M+ Inf/M− Fut
8: Predictive marker effect, (nonlinear and monotone) No Marker Detected 0.179 N = 261.5
Marker Detected 0.821 0.498
Overall Efficacy 0.188
Overall Inferiority
Overall Futility 0.024
M+ Eff/M− Fut 0.779
M+ Eff/M− Inf 0.008
M+ Fut/M− Eff
M+ Fut/M− Inf 0.001
M+ Inf/M− Eff
M+ Inf/M− Fut
9: Predictive marker effect, (nonlinear and non-monotone) No Marker Detected 0.128 N = 263.25
Marker Detected 0.872 0.261
Overall Efficacy 0.189
Overall Inferiority
Overall Futility 0.058
M+ Eff/M− Fut 0.753
M+ Eff/M− Inf
M+ Fut/M− Eff
M+ Fut/M− Inf
M+ Inf/M− Eff
M+ Inf/M− Fut
10: Predictive marker effect, (nonlinear and non-monotone) No Marker Detected 0.177 N = 251.25
Marker Detected 0.823 0.767
Overall Efficacy 0.173
Overall Inferiority
Overall Futility 0.007
M+ Eff/M− Fut 0.812
M+ Eff/M− Inf 0.007
M+ Fut/M− Eff
M+ Fut/M− Inf 0.001
M+ Inf/M− Eff
M+ Inf/M− Fut
11: Constant inferior treat effect, No marker effect No Marker Detected 0.763 N = 250
Marker Detected 0.237 0.420
Overall Efficacy
Overall Inferiority 0.916
Overall Futility
M+ Eff/M− Fut
M+ Eff/M− Inf
M+ Fut/M− Eff
M+ Fut/M− Inf 0.084
M+ Inf/M− Eff
M+ Inf/M− Fut
12: Prognostic and predictive marker effect, (linear; nearly dichotomous) No Marker Detected 0.157 N = 255.75
Marker Detected 0.843 0.509
Overall Efficacy 0.088
Overall Inferiority
Overall Futility 0.104
M+ Eff/M− Fut 0.789
M+ Eff/M− Inf 0.016
M+ Fut/M− Eff
M+ Fut/M− Inf 0.003
M+ Inf/M− Eff
M+ Inf/M− Fut

In the comparator design, the distribution of selected marker cutpoints when a binary marker is identified for each scenario is presented in Figure 5a,b. When there is no treatment difference between the control and experimental arms or the underlying treatment effect is independent of the marker value, the selected marker cutpoint distributions tend to be uniformly distributed; for example, Scenarios 1–3 and 11. When the true underlying predictive marker effect function is a step function or a smoothed step function, the comparator design could correctly identify the optimal marker cut-off point around 0.5. However, when the predictive marker effect is nonlinear or non-monotonic, the identified cut-off points tend to center around marker values where the large change in treatment effect occurs. The distributions have a mode around 0.25 in Scenario 9 and 0.75 in Scenario 10, ignoring the potential responders with marker value at the tail in Scenario 9 and including the nonresponders with marker value close to 0 in Scenario 10.

FIGURE 5 ∣.

FIGURE 5 ∣

Distribution of selected marker cutpoints in comparator frequentist trial design for Scenarios. (a) 1–6, and (b) 7–12.

The average sample sizes are low for the comparator design across all cases, and the average sample sizes are close to the minimum possible number of patients, since the only interim analysis occurs around when 50% of patients are enrolled and evaluated. Though the average sample sizes are smaller in the comparator design than the proposed design, the comparator design has lower power to detect a constant treatment effect (Scenario 2) and has a lower rate of correct classification for the null case (Scenario 1). Also, the comparator is much more likely to categorize patients into the nonresponder marker subgroup when they can actually benefit from the treatment based on the underlying true predictive marker effect function, while efficacy in these cases for some patients can be revealed by the proposed design.

In conclusion, our proposed design is able to model the underlying prognostic and predictive marker effects, which can be potentially nonlinear and non-monotonic, and correctly make marker-specific (patient-focused) decisions with efficiency. Though more marker-specific information may be required to reach some decisions, the proposed design has higher power in identifying patients whose marker values have a positive effect and is more likely to reflect the truly broader underlying classification of “responder” than the traditional frequentist adaptive enrichment design.

5 ∣. Conclusions

Within the scope of this paper, we resolved two major weaknesses of currently used biomarker-driven designs in cancer: (1) simplification or misrepresentation of the prognostic or predictive effects of naturally continuous biomarkers (e.g., circulating tumor DNA) that are driving study design and could therefore lead to incorrect conclusions, and (2) fixed rather than adaptive, marker-driven treatment allocation rules to experimental versus control therapy, despite accumulating information on-trial that could inform the treatment of enrolled patients on a personalized level.

We combined the machinery of Bayesian modeling and computing with immediate access to individual patient data, which together informed the development of a clinically intuitive trial design framework that allows for more accurate real-time identification of patient subgroups who are benefiting from an experimental targeted therapy. In contrast with a commonly used comparator design that uses algorithmic dichotomization with grid search for marker threshold, thus cannot model nonlinear or non-monotone marker effects, our flexible Bayesian modeling “engine” uses accumulating patient, molecular, and outcome data to continuously (1) update estimates of design-driving marker relationships, (2) quantify posterior uncertainty around these effects, and (3) translate the statistical information from (1) and (2) to personalized treatment allocation, such that a newly enrolled patient’s probability of randomization to experimental therapy is a function of his/her marker value, which is calculated based on the marker’s predictive effect estimated so far while taking into account the associated estimation uncertainty. This in turn translates into real time (rather than delayed) learning from patient data, shortened trial duration with faster answers to critical questions, and more ethical and personalized treatment of study patients.

This proposed design is best suited for specific trial settings where its unique features can be maximally beneficial and potential risks, such as population drift, are minimized. In practice, outcome-adaptive trials should only be considered in contexts where population drift is an unlikely concern; that is, where the time to observed outcome is relatively short for each enrolled patient, and where the total accrual period is not so long as to encompass likely changes in practice in the treatment landscape during the same time frame as the trial’s conduct. In addition, this design is particularly valuable when there is limited a priori knowledge about the precise nature of biomarker relationships with treatment effect or prognosis. It excels in scenarios where non-monotone, nonlinear, or other complex relationships between the continuous biomarker and the outcome cannot be ruled out.

The current proposed design only applies to a single-biomarker setting but could be extended to a multiple-marker setting in principle, for example, by developing a “risk score” or classification scale that is a function of multiple markers; however, further refinements should be done with caution, as such composite scores may not be clinically meaningful or interpretable.

Currently, our adaptive randomization algorithm only includes patients’ biomarker information to predict the probability of response. The adaptive randomization step could include and utilize some additional key baseline characteristics based on the investigational drug and disease. In addition, for longer-duration trials with multiple phases, the date of enrollment should also be included as a covariate at the adaptive randomization step, which could help mitigate potential bias from time-related trends in outcomes or population drift. These are future directions of this work.

Acknowledgments

This research was supported by the National Cancer Institute of the National Institutes of Health under award number U10CA180899.

Appendix A

Simulation With Adaptive Randomization

Figure A1a,b are randomization probability spaghetti plots of 10 simulated trial iterations, to visualize the randomization process. Simulation parameters are the same as described in Section 4. The x–axis represents indices of enrolled patients and y–axis is the probability of randomization to the experimental arm for a patient with the marker value associated with the biggest treatment effect for that scenario. If there is no marker effect, a patient with a marker value of 0.5 or the median marker value is used. Since in all Scenarios (except for Scenarios 1, 3, and 11), there are predictive marker effects or treatment effects, we expect to see an upward trend for the randomization probabilities associated with the most efficacious marker value. For Scenario 11, since it is a constant inferior treatment case, there is no marker value associated with treatment benefit. In Figure A1a,b, we see that randomization probabilities to the experimental arm go close to 1 as patients are accrued and more information is collected for Scenarios with treatment benefit. For Scenario 1 and Scenario 3, which are cases without treatment effect, the randomization probabilities are around 50%, with some small variance. For Scenario 11, due to the inferior treatment effect, the randomization probabilities fall to the bounded lowest randomization ratio, which is 0.1.

Appendix B

Visual Representation of Simulation Scenarios

The visual representations for each simulation scenario described in Table 1 are shown in Figure B1a,b.

FIGURE A1 ∣.

FIGURE A1 ∣

Randomization probability under different Scenarios for Scenarios. (a) 1–6, and (b) 7–12.

FIGURE B1 ∣.

FIGURE B1 ∣

Response rate as function of biomarker X for Scenarios. (a) 1–6, and (b) 7–12.

Footnotes

Conflicts of Interest

The authors declare no conflicts of interest.

Data Availability Statement

Data sharing is not applicable to this article as no new data were created or analyzed in this study.

References

  • 1.Jenkins M, Stone A, and Jennison C, “An Adaptive Seamless Phase II/III Design for Oncology Trials With Subpopulation Selection Using Correlated Survival Endpoints,” Pharmaceutical Statistics 10, no. 4 (2011): 347–356. [DOI] [PubMed] [Google Scholar]
  • 2.Brannath W, Zuber E, Branson M, et al. , “Confirmatory Adaptive Designs With Bayesian Decision Tools for a Targeted Therapy in Oncology,” Statistics in Medicine 28, no. 10 (2009): 1445–1463. [DOI] [PubMed] [Google Scholar]
  • 3.Krisam J and Kieser M, “Optimal Decision Rules for Biomarker-Based Subgroup Selection for a Targeted Therapy in Oncology,” International Journal of Molecular Sciences 16, no. 5 (2015): 10354–10375. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Renfro LA, Coughlin CM, Grothey AM, and Sargent DJ, “Adaptive Randomized Phase II Design for Biomarker Threshold Selection and Independent Evaluation,” Chinese Clinical Oncology 3, no. 1 (2014): 3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Liu Y, Kairalla JA, and Renfro LA, “Bayesian Adaptive Trial Design for a Continuous Biomarker With Possibly Nonlinear or Nonmonotone Prognostic or Predictive Effects,” Biometrics 78, no. 4 (2022): 1441–1453. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Wand MP and Ormerod JT, “On Semiparametric Regression With O’sullivan Penalized Splines,” Australian & New Zealand Journal of Statistics 50, no. 2 (2008): 179–198. [Google Scholar]
  • 7.Haario H, Saksman E, and Tamminen J, “An Adaptive Metropolis Algorithm,” Bernoulli 7 (2001): 223–242. [Google Scholar]
  • 8.Firth D, “Bias Reduction of Maximum Likelihood Estimates,” Biometrika 80, no. 1 (1993): 27–38. [Google Scholar]
  • 9.Wieringen v W N., “Lecture Notes on Ridge Regression,” arXiv preprint arXiv:1509.09169 (2015). [Google Scholar]
  • 10.Wathen JK and Thall PF, “A Simulation Study of Outcome Adaptive Randomization in Multi-Arm Clinical Trials,” Clinical Trials 14, no. 5 (2017): 432–440. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.“Morphotek Investigation in Colorectal Cancer: Research of MORAb-004 (MICRO),” Details on how published (2012). [Google Scholar]
  • 12.Grothey A, Strosberg JR, Renfro LA, et al. , “A Randomized, Double-Blind, Placebo-Controlled Phase II Study of the Efficacy and Safety of Monotherapy Ontuxizumab (MORAb-004) Plus Best Supportive Care in Patients With Chemorefractory Metastatic Colorectal Cancer,” Clinical Cancer Research 24, no. 2 (2018): 316–325. [DOI] [PubMed] [Google Scholar]
  • 13.Simon N and Simon R, “Adaptive Enrichment Designs for Clinical Trials,” Biostatistics 14, no. 4 (2013): 613–625. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Jiang W, Freidlin B, and Simon R, “Biomarker-Adaptive Threshold Design: A Procedure for Evaluating Treatment With Possible Biomarker-Defined Subset Effect,” Journal of the National Cancer Institute 99, no. 13 (2007): 1036–1043. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing is not applicable to this article as no new data were created or analyzed in this study.

RESOURCES