Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 4;45(18-19):e70693. doi: 10.1002/sim.70693

Conditional Estimations for Seamless Phase II/III Clinical Trials Involving Multi‐Stage Early Stopping

Siyu Zhu 1, Yuxuan Yang 1, Minggang Yin 1, Yixin Luo 1, Xueqing Liang 1, Shijie Yu 1, Chongyang Duan 1,✉
PMCID: PMC13435242  PMID: 42549628

ABSTRACT

To avoid unnecessary resource expenditure resulting from the sample size overestimation in seamless phase II/III designs, multi‐stage early stopping may be incorporated in the phase III component. However, existing estimation methods for seamless phase II/III designs are primarily developed for two‐stage settings and do not directly accommodate designs with multi‐stage early stopping in phase III. In this paper, we develop a Score‐statistics‐based framework for point and interval estimation in this specific class of designs. The framework provides a design‐specific formulation of conditional bias adjustment, conditional median‐unbiased estimation, and Rao–Blackwellization for seamless phase II/III designs with multi‐stage early stopping. Conditional on the trial continuing to phase III, the proposed framework targets valid and efficient estimation of the treatment effect for the selected arm and is applicable to a range of common endpoint distributions. Within this framework, we formulate and evaluate several representative conditional estimation procedures: the multiple iterations‐based conditional bias‐adjusted estimator (CBAE‐MI), the single iteration‐based conditional bias‐adjusted estimator (CBAE‐SI), the conditional median unbiased estimators with the unselected treatment group effects be estimated by their maximum likelihood estimators (CMUE‐MLE) or set as zero (CMUE‐ZERO), and the Rao–Blackwellized (RB) estimator. In addition, to properly quantify uncertainty around the point estimates, we derive confidence intervals based on CMUE‐MLE, CMUE‐ZERO, and RB. Through simulation under various endpoint types and parameter configurations, we recommend employing RB for point estimation and confidence interval, as it demonstrates superior robustness and conservative coverage probability across treatment selection and trial design settings.

Keywords: confidence interval, phase II/III design, point estimation, seamless sequential design, selected treatment

1. Introduction

Designed to expedite drug development timelines and enhance operational efficiency, seamless phase II/III clinical trials have garnered significant attention from researchers in the pharmaceutical industry [1]. These designs employ interim analyses to identify the most promising treatment, dosage, or subpopulation, and perform statistical inference based on the final cumulative data, thereby reducing intermediate decision time and increasing the informational value by consolidating evidence across different development phases. However, as noted by Teng et al. [2], a major challenge in such designs is the early commitment to a potential phase III design before the start of phase III, which may result in resource expenditure due to sample size overestimation or insufficient statistical power due to underestimation. To mitigate this issue, some researchers have proposed initiating the trial with a larger planned sample size and incorporating multiple interim analyses during phase III to allow early stopping for futility or efficacy. However, methodological research on this class of designs has largely emphasized hypothesis testing [3, 4], whereas the development of methods for parameter estimation remains comparatively limited.

A substantial literature on estimation and inference under adaptive designs has been developed in related settings. For group‐sequential designs, proposed point estimators include the bias‐adjusted estimator [5], the Rao–Blackwellized (RB) estimator [6], and the median‐unbiased estimator [7], accompanied by final and repeated confidence intervals. Methodological extensions have also been developed for nonnormal endpoints, including binary [8] and time‐to‐event [9] outcomes. For two‐stage seamless phase II/III designs, conditional on treatment selection and continuation to phase III, existing approaches include the bias‐adjusted estimator [10], the RB estimator [11, 12, 13], and the shrinkage estimator [14], while the performance of confidence interval procedures has been systematically evaluated [15]. These point estimation methods have recently been further extended to survival endpoints [16]. More comprehensive overviews of point and interval estimation methods under adaptive designs are available in recent review articles [17, 18, 19, 20].

For seamless phase II/III designs permitting multi‐stage early termination, only a limited number of estimation procedures have been proposed. Stallard and Todd [10] derived a bias‐adjusted estimator and a confidence region based on sample‐space ordering. Subsequently, Stallard and Kimani [21] proposed the uniformly minimum‐variance conditional unbiased estimator (UMVCUE) for multi‐arm multi‐stage trials; nevertheless, their work was not developed specifically for the seamless phase II/III setting with multi‐stage early termination. More recently, Gao and Li [4] proposed point and interval estimation methods; in their approach, the point estimator relies solely on phase III data, and the interval construction adopts a conservative lower bound that is not directly aligned with the corresponding point estimator. Overall, existing approaches have not demonstrated broad generalizability across endpoint distributions, treatment‐selection rules, or multi‐stage early stopping criteria.

Motivated by these gaps, it is necessary to investigate point and interval estimation methods specifically tailored to seamless phase II/III clinical trials incorporating multi‐stage early stopping to better inform their application in clinical practice. Moreover, because investigators may be reluctant for safety reasons to allow early termination for efficacy at the end of phase II, we focus on designs that permit termination for futility at the end of phase II while allowing termination for both futility and efficacy in phase III. Conditioning on continuation to phase III, we develop a general Score‐statistics‐based framework for inference on the selected arm, propose five conditional point estimators, and construct three corresponding confidence interval procedures in accordance with FDA guidance [22]. The framework is not intended to introduce entirely new estimation principles; rather, it provides a design‐specific formulation that makes these conditional estimation principles operational for seamless phase II/III designs with multi‐stage early stopping. Within this framework, existing conditional estimation principles are reformulated for the selected treatment under continuation to phase III, and their point‐estimation and interval‐estimation performance is systematically evaluated across endpoint types and design configurations.

The remainder of the paper is organized as follows. Section 2 introduces the design and notation and presents the proposed point estimators and confidence intervals. Section 3 reports a comprehensive simulation study comparing the methods. Section 4 illustrates implementation through a worked example, and Section 5 concludes with a discussion.

2. Methods

2.1. Setting and Notation

As Figure 1 described, we will consider a J‐stage seamless phase II/III clinical trial comparing K experimental treatments Tk(k=1,…,K) with a control treatment T0. At the end of stage 1 (i.e., the end of phase II), the most promising experimental treatment TS is selected based on preestablished selection criteria. If the selected treatment is deemed sufficiently inferior to the control, the trial will stop for futility; otherwise, the selected treatment will continue to phase III alongside the control, in which multi‐stage early stopping for both futility and efficacy are conducted using the data accumulated. For notational simplicity, in the following derivations, we assume that S=1, that is, T1 is the selected treatment.

FIGURE 1.

FIGURE 1

Flowchart of phase II/III clinical trials involving multi‐stage early stopping.

Following Kimani et al. [12] and Robertson et al. [13], we condition inference on continuation to phase III, since when early stopping for efficacy is not permitted in phase II, it is practically meaningful to focus on making a claim only in the subset of trials that continue to phase III. Let Q denote the event that T1 is selected and the trial proceeds to phase III, and let θ=θ1,θ2,…,θKT denote the treatment‐effect parameters that quantify the advantage of each experimental treatment over control. At the j‐th interim analysis (j=1,…,J), the futility and efficacy boundaries for the Wald statistic are denoted by lj and uj, respectively. Although termination for efficacy is not allowed at the end of phase II, we introduce the notation u1 (rather than +∞) to simplify subsequent derivations. In addition, we assume that the trial ultimately stops at stage JS with 1<JS≤J. Under this notation, the objective of this study is to obtain a valid and efficient estimator of θ1, based on the cumulative data available up to stage JS, conditional on the event Q.

To facilitate broad applicability across different endpoint distributions, we characterize the interim analyses statistics through Score statistics. At the j‐th (j=1,…,J) interim analysis, the data available for Tk(k=1,…,K) and T0 are summarized in terms of Score statistic Skj and the Fisher information Ikj. Referred to Jennison and Turnbull [7], the sequence S11,S12,…,S1J has a certain mathematical appeal as its increments S11, S12−S11, S13−S12, etc., are independent, and Skj is asymptotically normally distributed with mean θkIkj and variance Ikj. Therefore, under the asymptotic framework described above and conditional on the realized sequence of Ikj values, SK+JS−1=S11,S21,…,SK1,S12−S11,…,S1JS−S1JS−1T is approximately distributed as a K+JS−1‐dimensional multivariate normal distribution MVNμK+JS−1,∑K+JS−1, where μK+JS−1 and ∑K+JS−1 denote the mean vector and covariance matrix, respectively.

μK+JS−1=θ1I11,θ2I21,…,θKIK1,θ1I12−I11,…,θ1I1JS−I1JS−1T, (1)
∑K+JS−1=I11q12⋯q1K0⋯0q21I21⋯q2K0⋯0⋮⋮⋱⋮⋮⋯⋮qK1qK2⋯IK10⋯000⋯0I12−I11⋯0⋮⋮⋮⋮⋮⋱⋮00⋯00⋯I1JS−I1JS−1, (2)

where qkk′ denote the correlation between Sk1 and Sk′1, as all experimental treatments share a common control treatment in Stage 1. This framework accommodates normally distributed outcomes with known variance, binary outcomes, and time‐to‐event outcomes, with θk corresponding to the mean difference, log‐odds ratio, and log‐hazard ratio, respectively. Details on the computation of ∑K+JS−1 for these three endpoint types are provided in the Supporting Information, with closed‐form expressions available for the first two cases and the survival‐endpoint case requiring estimation via Cox regression.

Because trial stopping decisions are made by comparing Wald statistics to prespecified boundaries, we further characterize the distribution at the Wald‐statistic level. Let Zkj denote the Wald statistic for testing θk=0 at j‐th interim analysis for comparing experimental treatment Tk(k=1,…,K) over T0, it can be calculated using

Zkj=SkjIkj.

Thus, transforming the distribution of SK+JS−1 by ZK+JS−1=ASK+JS−1, ZK+JS−1=Z11,Z21,…,ZK1,Z12,…,Z1JST follow a K+JS−1‐dimensional multivariate normal distribution MVNAμK+JS−1,A∑K+JS−1AT, where A denote the K+JS−1×K+JS−1 matrix given in Supporting Information and AT denotes the transpose of A.

Using the notation introduced above, event Q can be expressed as:

Q:maxZ21,…,ZK1,l1<Z11<u1,

and the probability of event Q can be calculated as:

prob(Q)=∫−∞Z11…∫−∞Z11∫l1u1fθZK,QdZ11dZ21…dZK1, (3)

where fθZK,Q denote the probability density function of ZK=Z11,Z21,…,ZK1T under θ. Since the upper and lower bounds for integration prob(Q) are functions of random variables Z11, a linear transformation can be constructed such that the probability can be computed using the cumulative distribution function of the multivariate normal distribution [16], as demonstrated in the Supporting Information.

2.2. Point Estimators

The naive maximum likelihood estimator calculated from the data available at trial termination is generally biased because the selected treatment is chosen based on interim efficacy data and the trial may subsequently stop early for futility or efficacy. The estimation procedures considered below address this problem from three complementary perspectives. First, conditional bias‐adjusted estimators estimate the bias of the naive estimator conditional on treatment selection and continuation to phase III, and then subtract this estimated bias from the naive estimator. Second, conditional median‐unbiased estimators use a prespecified ordering of the sample space and define the estimate as the parameter value for which the observed trial outcome is conditionally median under the selection and continuation event. Third, the RB estimator starts from a stage‐wise estimator that is unbiased under the conditioning event and then reduces its variance by conditioning on a sufficient statistic. These three approaches are formulated below within the same Score‐statistics framework. We first describe the naive estimator, which serves as the common benchmark and the basis for defining the conditional bias.

2.2.1. Naive Estimator

The naive estimator is obtained by applying the standard formula for calculating standard maximum likelihood estimation (MLE) based on the data available at the time of trial termination. Let θ^MLE=θ^1,MLE,θ^2,MLE,…,θ^K,MLET denote the naive estimator of θ, which can be expressed as a function of ZK+JS−1. Specifically,

θ^k,MLE=z1JSI1JS,k=1,zk1Ik1,k≠1., 

where zjk denote a realization of Zjk. Because treatment selection and early stopping are incorporated during the trial, θ^MLE is generally biased for θ, and the corresponding bias expression depends on whether the treatment is selected.

The conditional bias is defined as follows:

bθθ^k,MLE|Q=Eθθ^k,MLE|Q−θk.

For the selected treatment, the estimator may be observed at any phase III stopping stage; therefore, the conditional expectation is decomposed over the possible stopping stages j=2,…,J, with the lower and upper tail integrals corresponding to futility and efficacy stopping at stage j, respectively. Accordingly, the conditional bias bθθ^1,MLE|Q for the selected treatment T1, is given by [10].

bθθ^1,MLE|Q=1prob(Q)∑j=2J1I1j∫−∞ljZ1jfθZ1j,QdZ1j+∫uj+∞Z1jfθZ1j,QdZ1j−θ1, (4)

where fθZ1j,Q is the marginal distribution of Z1j.

For the unselected treatments Tk(k=2,…,K), no phase III data are collected, and the estimator is based only on phase II data. Its conditional bias is therefore obtained by integrating over the region in which T1 is selected and the trial continues to phase III, giving:

bθθ^k,MLE|Q=1prob(Q)∫−∞Z11…∫−∞Z11∫l1u1Zk1Ik1fθZK,QdZ11dZ21…dZK1−θk. (5)

In this expression, the limits l1<Z11<u1 and Zk1<Z11 for k=2,…,K encode the event that treatment T1 is selected and the trial does not stop for futility at the end of phase II.

Equations (4) and (5) follow the conditional expectation argument used by Stallard and Todd [10], with the conditioning event adapted to the present seamless phase II/III design with multi‐stage early stopping.

When computing the conditional bias for the selected treatment, the marginal density fθZ1j,Q can be obtained from the joint distribution of ZK+j−1, that is

fθZ1j,Q=∫lj−1uj−1…∫−∞Z11…∫−∞Z11∫l1u1fθZK+j−1,QdZ11dZ21…dZK1…dZ1,j−1.

Analogous to ZK+JS−1, the vector ZK+j−1 follow a (K+j−1)‐dimensional multivariate normal distribution. Therefore, as in the computation of prob(Q), a linear transformation can be constructed to calculate the above equation using the cumulative distribution function of the multivariate normal distribution.

2.2.2. Conditional Bias‐Adjusted Estimator

Given the expressions for conditional biases provided in Expressions (4) and (5), an unbiased estimator can be obtained by subtracting these biases from the naive estimator [5, 10]. A formulation for this idea would be that the equation

θ=θ^MLE−bθθ^MLE|Q

holds if the conditional bias vector bθθ^MLE|Q is known.

However, in practice, the effect vector θ is unknown, making it challenging to solve this expression directly. To address this issue, previous literature provided two approaches. One option is the multiple iterations (MI) scheme. Let θ^CBAE−MI(r) denote the multiple iterations‐based estimator (CBAE‐MI) of θ at iteration r(r=1,2,…), then the θ^CBAE−MI(r+1) can be given by

θ^CBAE−MI(r+1)=θ^MLE−bθ^CBAE−MI(r)θ^MLE|Q, (6)

where bθ^CBAE−MI(r)θ^MLE|Q can be calculated using the Expression (4) and (5). In our simulations, the initial value θ^CBAE−MI(0) is set to be θ^MLE, and the iteration is terminated when the convergence criterion or the maximum number of iterations is reached. Due to the difficulties associated with convergence in the multiple iteration process and the computational expense, an alternative for CBAE‐MI is the single iteration (SI) scheme, where θ^MLE is used as the true effect to compute the bias vector directly. We use θ^CBAE−SI to denote the estimation obtained from the single iteration process:

θ^CBAE−SI=θ^MLE−bθ^MLEθ^MLE|Q. (7)

2.2.3. Conditional Median Unbiased Estimator

In classical group sequential designs, the idea of stage‐wise orderings is frequently used to construct confidence interval and median unbiased estimator. Based on this idea, Stallard and Todd [10] constructed a confidence interval conditional on the selection treatment, but they did not propose a corresponding point estimator. Building on this, we develop the conditional median unbiased estimator (CMUE) conditional on event Q. Specifically, CMUE can be defined as the solution to:

Pθ^CMUEj′,z′≻JS,z1JS|Q=0.5.

where Pθj′,z′≻JS,z1JS|Q denotes the conditional probability that the trial outcome j′,z′ is more “extreme” than the observed outcome JS,z1JS under the prespecified sample‐space ordering rule.

Although many sample‐space ordering rules have been proposed, we adopt the stage‐wise ordering of Fairbanks and Madsen [23], as recommended by Jennison and Turnbull [7]. Specifically, j′,z′≻(j,z) if any of the following occurs: (i) j′=j and z′≥z; or (ii) j′<j and z′ crosses the efficacy boundary. Then, using the above notation and concept, the conditional p value under θ, can be calculated as follows:

Pθj′,z′≻JS,z1JS|Q=1prob(Q)∫z1Js+∞fθZ1JS,QdZ1Js,Js=21prob(Q)∑j=2Js−1∫uj+∞fθZ1j,QdZ1j+∫z1Js+∞fθZ1JS,QdZ1Js,Js>2,  (8)

Analogous to the computation of bθθ^MLE|Q, Pθj′,z′≻JS,z1JS∣Q can also be evaluated using the cumulative distribution function of a multivariate normal distribution.

Since researchers are often primarily interested in the effect of the selected treatment θ1, Stallard and Todd [10] proposed substituting θk(k=2,…,K) with either θk=θ^k,MLE or θk=0. Following this approach, we also adopt this substitution strategy and denote the CMUE under the two alternatives as θ^1,CMUE−MLE and θ^1,CMUE−ZERO, respectively, which are defined as follows:

Pθ^1,CMUE−MLE=Pθ^1,CMUE−MLEj′,z′≻JS,z1JS|Q=0.5, (9)
Pθ^1,CMUE−ZERO=Pθ^1,CMUE−ZEROj′,z′≻JS,z1JS|Q=0.5. (10)

2.2.4. Rao–Blackwellized Estimator

Based on Stallard and Kimani [21], conditioning an unbiased estimator on a sufficient statistic yields a RB estimator with variance no greater than that of the original unbiased estimator. In our setting, the unbiased estimator used for Rao–Blackwellization estimator is constructed from the stage‐wise increment available at the second interim analysis. Conditional on the selection event Q, this stage‐wise increment provides an unbiased estimator of θ1. As for the sufficient statistics, the joint density of SK+JS−1 under event Q is given by

fθSK+JS−1|Q=1(2π)K+JS−1∑K+JS−1exp−12SK+JS−1−μK+JS−1TPK+JS−1SK+JS−1−μK+JS−1.

where PK+JS−1 denote the inverse of the variance covariance matrix ∑K+JS−1. It can then be shown that the vector W,JST is the sufficient statistic for θ, where W=W1,W2,…,WKT, and

W1=I11∑k=1Kpk1Sk1+S1,JS−S11,
Wk=Ik1∑k′=1Kpk′kSk′1(k=2,…,K).

with piji,j=1,…,K+JS−1 denote the ijth entry in PK+JS−1. Detailed derivation is provided in the Supporting Information. Therefore, the RB estimator of θ1 is obtained by taking the conditional expectation of the unbiased stage‐wise estimator at the second interim analysis given the sufficient statistic:

θ^1,RB=ES12−S11I12−I11|W=w,JS,Q. (11)

where w denote a realization of W.

Given the difficulty in accurately obtaining the conditional probability density function, the following steps are used to compute the integral in expression (11):

  • Step 1: Monte Carlo Sampling. By applying a linear transformation to SK+JS−1, we obtain the joint distribution of W1,W2,…,WK,S12−S11,…,S1JS−S1JS−1T, which is a K+JS−1‐dimensional multivariate normal distribution. Then, given W, we use the conditional distribution of the multivariate normal distribution to sample S12−S11,…,S1JS−S1JS−1T;

  • Step 2: Screen the samples. To ensure the consistency of JS,Q, we use the termination boundary and treatment selection rules to screen the samples from Step 1, removing those that did not select T1 and did not terminate at the stage JS, and marking the remaining samples as acceptable samplings;

  • Step 3: Calculate the sample mean. Calculate S12−S11/I12−I11 for each acceptable sampling, then compute the mean as the estimate of θ^1,RB.

It is worth to note that W,JST is not a complete statistic, which means that the Lehmann‐Scheffé theorem [24], in its usual form does not apply. This implies that the above steps only yield an estimate with variance no greater than the unbiased estimator based on the second interim analysis, and there may exist other unbiased estimates with smaller variance. We discuss this point in more detail in the discussion section.

2.3. Conditional Confidence Intervals

2.3.1. Confidence Interval Based on Naive Estimator

Using the naive estimator and its standard error, we can construct a (1−α)% confidence interval by applying the standard formula of Wald method, which is given by

θ^1,MLE±z1−α/2seθ^1,MLE, (12)

where seθ^1,MLE=I1JS−2 and z1−α/2 denote the (1−α/2)‐quantile of the standard normal distribution. Because this confidence interval ignores the selection and early stopping in the trial, it does not provide the correct coverage of (1−α)%, as demonstrated by the simulations in the next section.

2.3.2. Confidence Intervals Based on Conditional Median Unbiased Estimator

After choosing an ordering of the sample space and observing a sample JS,z1JS, a (1−α)% confidence interval l^CMUE,u^CMUE can be obtained as

Pl^1,CMUEj′,z′≻JS,z1JS|Q=α2, (13)
Pu^1,CMUEj′,z′≻JS,z1JS|Q=1−α2. (14)

Similarly to the point estimation, the focus here is on estimating the selected treatment effect. Therefore, substitutions of θk=θ^k,MLE or θk=0 for the unselected treatment effect θk(k=2,…,K) is also conducted here.

2.3.3. Confidence Interval Based on Rao–Blackwellized Estimator

As stated by Whitehead et al. [25], if θ^1,RB−Eθ^1,RB/Varθ^1,RB is assumed to follow a standard normal distribution, then an approximate (1−α)% confidence interval for θ1 based on the θ^1,RB is given by

θ^1,RB±z1−α/2seθ^1,RB, (15)

where the standard error of θ^1,RB is computed as

seθ^1,RB=1I12−I11−EvarS12−S11I12−I11|W=w,JS,Q.

The calculation of E[var(S12−S11/I12−I11|W=w,JS,Q)] can be reliably estimated by var(S12−S11/I12−I11|W=w,JS,Q). This conditional variance can be calculated using a similar Monte Carlo step as in the previous section, with Step 3 replaced by calculating the sample variance of the stage‐wise unbiased estimator across the acceptable samplings.

3. Comparison of Methods Using a Simulation Study

3.1. Simulation Study Scenarios

In this section, we present simulation studies designed to evaluate the statistical properties of the conditional point estimators and confidence intervals introduced in Section 2. To ensure a comprehensive assessment, we examined a range of conditions, including different endpoint types, patterns of the true parameter values θ, and the numbers of trial stages. Specifically, we considered three types of endpoint distributions: normally distributed endpoints with known variance, binary endpoints, and survival endpoints following a Weibull distribution. For the true parameter values, we investigated four different scenarios including H0, H1, peak and linear scenarios. For the total number of trial stages, designs with two, three, and four stages early stopping were considered, corresponding to J=3,4,5. Each simulation scenario was based on 100 000 independently generated trial realizations, and all simulated datasets were generated at the patient level. At each interim analysis, the corresponding Score statistics, Fisher information, treatment‐selection rule, and stopping boundaries were applied to the cumulatively accrued patient‐level data to mimic the conduct of the trial.

In the simulations with normally distributed endpoints, patient‐level data were generated from normal distribution Nμk,σ2(k=0,…,4), where μ0=0 represents the effect of the control treatment, and σ2=1 denotes the known and common variance across all treatment groups. Following the framework of Khan et al. [16], we evaluated the statistical properties of the proposed estimators across four scenarios:

  • H0 scenarios: μ1=μ2=μ3=μ4=0.

  • H1 scenarios: μ1=μ2=μ3=μ4=0.50.

  • peak scenario: μ1=μ2=μ3=0.25,μ4=0.50.

  • linear scenario: μ1=0.20,μ2=0.30,μ3=0.40,μ4=0.50.

Regrading to the type I error control and sample size determination, we adopted the strategy proposed by Stallard and Todd [26], and the O'Brien–Fleming alpha spending function [27] was applied to calculate stopping boundaries and to maintain the statistical power at the prespecified level of 0.8. Additional implementation details are provided in the Supporting Information.

In the simulations with binary endpoints, patient‐level responses were generated from Bernoulli distributions with treatment‐specific response probabilities determined according to the H0, H1, peak, and linear scenarios. For the time‐to‐event endpoints, patient‐level survival times were generated from Weibull distributions. Treatment effects were introduced through the log‐hazard ratio parameter θk, with parameter configurations specified according to the H0, H1, peak, and linear scenarios. Detailed endpoint‐specific parameter settings, stage‐specific stopping boundaries, and sample size or event‐number calculations are provided in the Supporting Information to keep the main text focused while ensuring reproducibility.

For the calculation procedure of CBAE‐MI, the maximum number of iterations was set to 20, with a convergence criterion of 0.001; that is, convergence was declared when the difference between successive iterations was less than 0.001. The Monte Carlo sampling size for the RB was fixed at 10000. The confidence level of these two‐sided intervals is set at 95%. Moreover, to ensure a more comprehensive comparison, we further included point and interval estimators derived from stage‐wise data at the second interim analysis (stage2) for comparative purposes, which are unbiased but generally less efficient.

In terms of evaluation metrics, conditional bias and conditional root mean square error (RMSE) were calculated for point estimators, while conditional coverage probability (CP) and conditional interval width were used for confidence intervals. Following the description in Khan et al. [16], let 1S=k,JS≥2(i) indicate that the experimental treatment Tk is selected and the ith simulated trial continues to phase III, and let θ^k(i) denote the point estimator for θk at this trial. With n_sim simulations, conditional on the event Q (i.e., S=k and JS≥2), the conditional bias and RMSE of this estimator can be assessed as

Bias=∑i=1n_simθ^k(i)−θk1S=k,JS≥2(i)∑i=1n_sim1S=k,JS≥2(i),
RMSE=∑i=1n_simθ^k(i)−θk21S=k,JS≥2(i)∑i=1n_sim1S=k,JS≥2(i).

Similarly, letting lk(i),uk(i) denote the confidence interval for θk in the ith trial, the conditional CP and width can be given by

CP=∑i=1n_simIlk(i)<θk<uk(i)1S=k,JS≥2(i)∑i=1n_sim1S=k,JS≥2(i),
Width=∑i=1n_simuk(i)−lk(i)1S=k,JS≥2(i)∑i=1n_sim1S=k,JS≥2(i),

where I(.) denotes the indicator function.

3.2. Simulation Results

Since the simulation results for binary and survival endpoints closely resemble those under normal distribution, Sections 3.2.1 and 3.2.2 only focus on a detailed discussion of the point and interval estimates under the normal distribution, and Section 3.2.3 provides a brief summary of the results for the remaining two endpoint types, with the detailed results presented in the Supporting Information. The Monte Carlo standard errors for all performance measures across each endpoint are summarized in the Supporting Information [28].

3.2.1. Simulation Results for Point Estimators

Figure 2 presents the simulation results of point estimates for the case J=4 under the normal distributions. Details on the termination stages of the simulation trials for each scenario, as well as the convergence of CBAE‐MI, are provided in the Supporting Information.

FIGURE 2.

FIGURE 2

The conditional bias and RMSE of the six considered point estimators. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the probability that the corresponding treatment is selected and the trial continues to phase III).

As illustrated Figure 2A,B, under the global null hypothesis (i.e., θ1=θ2=θ3=θ4=0), the probabilities of selecting and continuing with each treatment were approximately 0.20, reflecting the fact that about 80% of trials continued to phase III and selection was nearly balanced across the four experimental treatments. Consistent with this symmetry, the estimated conditional biases and RMSE are also approximately equal across the four treatments. As expected, the naive estimator is conditionally biased and tends to overestimate the treatment effect. The conditional biases of stage2, CBAE‐SI and RB fluctuate around 0, indicating approximately unbiased. However, the biases of the remaining three estimators (CMUE‐MLE, CBAE‐MI, and CMUE‐ZERO) are similar and negative, suggesting that these estimators overcorrect conditional bias and thus underestimate the treatment effects. In terms of conditional RMSE, CBAE‐SI and the naive estimator perform similarly and have the smallest RMSE among all estimators, followed by CBAE‐MI, and then RB, CMUE‐MLE, CMUE‐ZERO and stage2. Therefore, in this scenario, RB as well as CBAE‐SI are recommended to obtain unbiased estimation, with CBAE‐SI providing the smallest conditional RMSE.

In the scenario where the second parameter configuration is applied, with θ1=θ2=θ3=θ4=0.50 and all experimental treatments outperforming the control with equal effects (Figure 2C,D), these probabilities were approximately 0.25 for each treatment, indicating that almost all simulated trials continued to phase III and that treatment selection remained balanced. Unlike the pattern reported in Khan et al. [16], our simulations show that the naive estimator under this configuration exhibits a larger bias than in the H0 scenario. This occurs because, although the bias induced by early termination for futility in phase II diminishes when all experimental treatments are effective, the multi‐stage early stopping in phase III introduces additional bias. Overall, the naive estimator exhibits the largest positive bias, followed by CMUE‐ZERO, which also tends to overestimate the effect. Additionally, both naive estimator and CMUE‐ZERO show comparable conditional RMSE, slightly higher than those of the other methods except stage2. However, in contrast to H0 scenario, the conditional bias and RMSE of CMUE‐MLE decrease in this configuration, behaving similarly to CBAE‐SI, and both estimators slightly overestimate the effect but have the smallest conditional RMSE compared to other estimators. Among the remaining estimators, RB and stage2 maintain unbiasedness, while CBAE‐MI continues to exhibit overcorrection. Consequently, RB is recommended for obtaining unbiased estimates of the mean, whereas CBAE‐SI and CMUE‐MLE may be considered if the researcher prioritizes lower RMSE despite slight overestimation.

The simulation results for peak and linear scenarios are presented in the third and fourth rows of Figure 2. In the peak scenario, selection and continuation were concentrated on treatment 4, with a probability of 0.624, whereas the corresponding probabilities for treatments 1–3 were approximately 0.12. In the linear scenario, this probability increased with the true treatment effect, from 0.060 for treatment 1 to 0.516 for treatment 4. These patterns indicate that, when one treatment has a clear efficacy advantage, selection is more likely to concentrate on the superior treatment. Compared with the equal‐effect H1 scenario, this reduces the selection‐induced component of bias when treatment 4 is selected, although bias induced by multi‐stage monitoring in phase III remains. Therefore, the conditional bias of the naive estimator is smaller in the peak and linear scenarios than in the H1 scenario, but it does not disappear completely. Overall, the seven estimators display nearly identical patterns in these two scenarios. In terms of conditional bias, CBAE‐SI, CMUE‐MLE, and RB remain essentially unbiased, with RB exhibiting the smallest conditional bias among them. By contrast, the naive estimator and CMUE‐ZERO tend to overestimate, whereas CBAE‐MI underestimates the effect. Regarding conditional RMSE, stage2 yields the largest values, CBAE‐SI achieves the smallest, and the other five methods perform similarly. Overall, RB, CBAE‐SI, and CMUE‐MLE are recommended based on their favorable balance of bias and RMSE.

The simulation results for J=3 and J=5 lead to conclusions broadly consistent with those observed for J=4. Overall, in seamless phase II/III designs involving multi‐stage early stopping, RB, CBAE‐SI, and CMUE‐MLE are recommended for point estimation. Nevertheless, these simulation findings underscore the importance of conducting detailed simulation studies to further investigate the potential overestimation associated with CBAE‐SI and CMUE‐MLE.

3.2.2. Simulation Results for the Confidence Intervals

Figure 3 presents the results of the four confidence intervals under a normal distribution with J=4. As expected, the Wald interval from the naive estimator yields the narrowest width but fails to achieve the nominal 95% conditional CP in most simulations; moreover, consistent with the behavior observed for point estimators, the magnitude of coverage deviation is also influenced by treatment selection and early stopping. The interval based on stage2 attains more accurate coverage but at the cost of wider intervals. Intervals based on CMUE‐MLE and RB tend to be conservative. Concretely, under scenario 2, CMUE‐MLE achieves higher coverage than RB while maintaining narrower conditional intervals, whereas in the other three scenarios, the two methods exhibit similar CP and interval width. The interval constructed from CMUE‐ZERO also improves conditional CP relative to the naive‐based interval, though its coverage is slightly deficient in certain cases—for example, under H1 scenario and when noneffective arms are selected in peak and linear scenarios. Nevertheless, CMUE‐ZERO generally produces narrower intervals than RB and CMUE‐MLE.

FIGURE 3.

FIGURE 3

The conditional coverage probability and interval width of the four considered confidence intervals. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the probability that the corresponding treatment is selected and the trial continues to phase III).

When the maximum number of trial stages is set to J=3 or J=5, the above conclusions remain largely unchanged. The only exception occurs at J=3, where the CMUE‐MLE interval is more conservative than RB while still yielding narrower widths. Detailed results are displayed in the Supporting Information. Overall, for confidence intervals in seamless phase II/III designs involving multi‐stage early stopping, intervals based on RB and CMUE‐MLE are recommended, as they achieve coverage probability closer to the nominal level than the naive‐based interval while yielding narrower widths than the stage2‐based interval.

3.2.3. Other Simulation Results

Simulation results for data follow the binomial distribution or Weibull distribution are summarized in the Supporting Information. Consistent with the findings under the normal distribution, the combination of θ under different scenarios affects the statistical properties of each method, while the total number of trial stages has little effect on the statistical properties of each method. Therefore, we will primarily discuss the performance of each method under the four previously mentioned scenarios (H0, H1, peak and linear scenarios).

For point estimates under binary data, the statistical properties of each method are generally comparable to those observed under the normal distribution. Accurately, CBAE‐MI continues to underestimate the treatment effect, with its conditional RMSE being slightly lower than RB. CBAE‐SI still exhibits slight positive bias when the noneffective group is selected, but its conditional RMSE remains smaller than other methods in most cases. Additionally, when the superior experimental treatment is selected in peak and linear scenarios, the conditional RMSE of CMUE‐ZERO and CMUE‐MLE are smaller than that of the CBAE‐SI, and at these settings, the conditional bias of CMUE‐ZERO is close to zero, whereas the degree of underestimation for CMUE‐MLE is comparable to that of CBAE‐MI. Furthermore, compared to simulation results under the normal distribution, RB exhibits unbiasedness for this endpoint type only when θS is not excessively large; otherwise, it tends to underestimate the effect. This behavior likely arises because the normal approximation is reliable only when the log odds ratio is small. For confidence intervals under binary data, results remain consistent with those reported in Section 3.2.2. The confidence intervals derived from RB and CMUE outperform those from the naive estimate in terms of conditional CP and are narrower than those obtained from the stage‐wise data of the second interim analysis. Concretely, the conditional intervals based on RB and CMUE‐MLE tend to be conservative, with CMUE‐MLE being more conservative yet yielding narrower intervals. CMUE‐ZERO achieves coverage close to the nominal 95%, and its coverage increases with increasing θS, potentially resulting in lower coverage at small θS, while its interval width remains smaller than those of RB and CMUE‐MLE.

For survival data generated from the Weibull distribution, the performance of point estimators in terms of conditional bias is broadly consistent with that observed under the normal distribution, with the exception that CMUE‐MLE and CMUE‐ZERO generally exhibit positive bias across most scenarios. Regarding conditional RMSE, the performance of the naive estimator and CMUE‐ZERO is slightly worse than that of the other methods (except stage2) in peak and linear scenarios, whereas in the remaining scenarios, the methods mentioned in Section 2.2 perform similarly and consistently achieve lower RMSE than stage2. For interval estimation, RB‐based intervals achieve coverage probabilities closer to the nominal level, while maintaining interval widths comparable to those of CMUE‐MLE.

In summary, for binary data, RB is preferable when the log odds ratio of the selected treatment is small to ensure unbiasedness. When the log odds ratio is large, CBAE‐SI may be favored for its reduced bias, whereas CMUE‐ZERO can be advantageous when minimizing conditional RMSE is of primary concern and the trial has a high probability of selecting the superior group. For interval estimation, Wald confidence intervals based on RB remain the primary recommendation due to their reliable coverage performance, although CMUE‐ZERO also emerges as a competitive alternative. Nonetheless, comprehensive simulation studies are advisable to thoroughly evaluate the potential risks associated with applying these methods to binary data. For survival data, RB and its corresponding confidence interval are recommended, as it provides robust unbiased estimates with coverage probabilities closer to the nominal level.

4. Example

In this section, utilizing the previously discussed point estimators and confidence intervals, we compute estimates for an example derived from the case study described in Bretz et al. [29]. The trial is designed to compare the effects of three different dose groups against a placebo for treating generalized anxiety disorder, with the primary endpoint being the change from baseline at week 8 in the total score on the Hamilton Rating Scale for Anxiety (HAM‐A). Prior research suggests that the HAM‐A scores for each dose group can be reasonably assumed to follow a normal distribution with a common standard deviation of 6 points, and the clinically significant improvement is defined as a reduction of 2 points more than the placebo group. Based on these setups, Bretz et al. [29] developed a two‐stage seamless phase II/III design and reviewed relevant hypothesis testing methods. However, to align with the trial design of this study, we modified it to a four‐stage seamless phase II/III design, in which the selection rule is to choose the group with the highest positive Wald statistic.

Assuming that each dose group has the same sample size as the control treatment at each stage, under the global null hypothesis θ1=θ2=θ3=0 and least favorable configuration θ1=θ2=0,θ3=2, we calculated that, given the futility and efficacy bounds of L=(0,0.106,1.395,2.192) and U=(∞,3.954,2.766,2.192), respectively, and the cumulative sample sizes for the placebo at the four stages as n=(52,104,156,208), the trial achieved power of 1−β=0.8 while maintaining the familywise error rate at α=0.025 (assuming one‐sided null hypotheses). The trial results are summarized in Table 1. Column 2 reports the true parameter values specified for data generation, while columns 3–5 present the observed mean values at each stage of the trial and Columns 6–8 provide the corresponding Wald statistics calculated based on cumulative data. As shown in Column 6, the Dose 3 treatment exhibited the highest positive Wald statistic and was therefore selected along with the control group to advance to the phase III; that is, S=3. Moreover, the trial was terminated for efficacy; that is, JS=3.

TABLE 1.

Example trial of a seamless phase II/III design incorporating multi‐stage early stopping.

Stage‐wise mean Wald‐statistic
Treatment True value Stage 1 Stage 2 Stage 3
zk1
zk2
zk3
Placebo 0 0.418 0.520 −0.334 — — —
Dose 1 0.8 1.333 — — 0.777 — —
Dose 2 1.5 1.011 — — 0.504 — —
Dose 3 2.6 3.775 2.821 3.362 2.853 3.400 4.590

Note: Stage‐wise means are the observed mean responses within each stage. zk1, zk2, and zk3 denote the Wald statistics for dose k versus placebo calculated at Stages 1, 2, and 3, respectively, using the data available up to the corresponding analysis. At Stage 1, the dose with the largest positive Wald statistic was selected to continue to phase III. Dashes indicate that the corresponding dose was not continued to that stage or that the Wald statistic was not applicable for placebo.

To make the implementation more transparent, we summarize the calculation workflow for this example. First, the selected dose group was relabeled as the selected arm, and the selection/continuation event Q was defined according to the observed stage‐1 Wald statistics and the phase II futility boundary. Second, the naive estimator and the stage‐wise estimator based on the second interim analysis were computed directly from the observed cumulative and stage‐wise mean differences. Third, CBAE‐SI and CBAE‐MI were obtained by evaluating the conditional bias using Equations (4) and (5), with CBAE‐MI iterated until the prespecified convergence criterion was met. Fourth, CMUE‐MLE and CMUE‐ZERO were obtained by numerically inverting the conditional stage‐wise (p) value in Equation (8), with the nuisance parameters for unselected treatments replaced by their MLEs or set to zero, respectively. Finally, the RB estimator was obtained by Monte Carlo sampling from the conditional distribution given the sufficient statistic, retaining only samples compatible with the observed treatment selection and stopping stage. The corresponding confidence intervals were then computed using the procedures described in Section 2.3. The corresponding implementation workflows for CBAE, CMUE, and RB are illustrated in Figure 4. Further numerical details, and annotated code are provided in the Supporting Information.

FIGURE 4.

FIGURE 4

Implementation workflows for CBAE, CMUE, and RB estimators in the seamless phase II/III example. (Panel (A) shows the multi‐iteration conditional bias‐adjusted estimator (CBAE‐MI), and Panel (B) shows the single‐iteration conditional bias‐adjusted estimator (CBAE‐SI). Panels (C) and (D) summarize the implementation workflows for the conditional median‐unbiased estimators (CMUE‐MLE and CMUE‐ZERO) and the Rao–Blackwellized (RB) estimator, respectively.)

Following this workflow and conditioning on the event Q:Z31>max0,Z11,Z21, we computed the point and interval estimators discussed in Section 2, with the results summarized in Table 2. Using a convergence criterion defined as a difference of less than or equal to 0.001 between successive iterations, the CBAE‐MI estimator converged after four iterations. As expected, the resulting estimates differed from the naive estimator, with the naive estimator yielding the largest value and the stage2 yielding the smallest. Regarding confidence intervals, those derived from RB, CMUE‐MLE, and CMUE‐ZERO were all shifted to the left of the naive‐based interval, and the interval of the naive estimator was the narrowest among all methods. The corresponding code is provided in https://github.com/syzhu0408/conditional‐estimation‐multi‐stage‐seamless‐phase2‐3.

TABLE 2.

Point estimation and interval estimation results from the example trial of a seamless phase II/III design incorporating multi‐stage early stopping.

Methods Main computational step Point‐estimation 95% confidence interval
Naive Compute cumulative mean difference at stopping stage 3.118 (1.787, 4.450)
stage2 Compute mean difference using stage‐2 increment data 2.301 (−0.005, 4.607)
CBAE‐MI Iteratively update θ^CBAE−MI(r+1)=θ^MLE−bθ^CBAE−MI(r)θ^MLE|Q until convergence 2.965 —
CBAE‐SI Subtract conditional bias evaluated from Equations (4) and (5) once 2.973 —
CMUE‐MLE Solve Pθ^S,CMUE−MLE=0.5 with nuisance parameters fixed at their MLEs 2.978 (1.428, 4.377)
CMUE‐ZERO Solve Pθ^S,CMUE−ZERO=0.5 with nuisance parameters fixed at 0 2.998 (1.505, 4.380)
RB Average unbiased stage‐wise estimators over accepted conditional Monte Carlo samples 2.841 (1.322, 4.359)

Note: “stage2” denotes the stage‐wise estimator based on the second interim analysis increment only. CBAE‐SI applies the conditional bias correction once, whereas CBAE‐MI iterates the bias‐correction step until convergence. For CMUE‐MLE and CMUE‐ZERO, Pθ^S,CMUE−MLE and Pθ^S,CMUE−ZERO denotes the conditional stage‐wise p value defined in Equation (8); nuisance parameters for unselected treatments are fixed at their MLEs or at zero, respectively. For RB, accepted conditional Monte Carlo samples are draws from the conditional distribution given the sufficient statistic that are compatible with the observed treatment selection and stopping stage. Confidence intervals are not reported for CBAE‐SI and CBAE‐MI because interval procedures were developed only for CMUE‐MLE, CMUE‐ZERO, and RB in this study.

5. Discussion

With increasing pressures in drug development, there is growing interest in seamless phase II/III study designs. Researchers have proposed implementing multi‐stage early stopping in phase III trials to minimize the potential waste associated with excessive sample sizes. However, existing parameter estimation methods lack generalizability across different endpoint distributions and trial designs, and systematic comparisons between methods are limited. To address these challenges, the present study develops a design‐specific Score‐statistics‐based inferential framework for conditional estimation in seamless phase II/III designs with multi‐stage early stopping. Within this framework, we formulate and evaluate five point estimators (CBAE‐MI, CBAE‐SI, CMUE‐MLE, CMUE‐ZERO, and RB) and three confidence interval procedures derived from CMUE and RB, conditional on treatment selection and continuation to phase III.

For point estimation, our simulation results reveal that the naive estimator suffers from pronounced overestimation bias, and the magnitude of bias reflects two interacting components: a selection component, driven by factors such as the true treatment effects and the number of experimental arms, and a monitoring component, determined by the stopping boundaries and the number of interim analyses. The selection component is closely related to the probability that each treatment is selected and that the trial continues to phase III, whereas the monitoring component reflects the probability of stopping at different phase III analyses. In our simulations, balanced selection probabilities under the global null and equal‐effect scenarios led to similar bias patterns across treatments, whereas concentration of selection on the superior treatment in the peak and linear scenarios attenuated the selection‐induced component of bias. Accordingly, it is important to recognize these sources of bias at the design stage and to account for them through appropriate bias adjustment in subsequent statistical inference. The stage‐wise estimator achieves robust unbiasedness but at the cost of higher variance. The other conditional estimators evaluated in this study effectively mitigate these limitations. Across the three endpoint settings, CBAE‐MI tends to overcorrect for bias, yielding conditional RMSE values slightly lower than those of the RB. By contrast, CBAE‐SI achieves smaller conditional bias and superior conditional RMSE, although it exhibits a slight positive bias when the noneffective treatment is selected. The performance of CMUE‐ZERO and CMUE‐MLE is more heterogeneous across scenarios, and CMUE‐ZERO generally shows positive bias in most settings except under the H0 scenarios, suggesting that it does not fully remove overestimation effects, while CMUE‐MLE behaves similarly to CBAE‐SI in most cases. Additionally, it is worth noting that both CMUE‐ZERO and CMUE‐MLE demonstrate promising potential for binary outcomes, particularly when θS is large. Regarding RB estimation, unbiasedness is consistently maintained under both normal and survival data. For binary outcomes, RB remains unbiased only when effect sizes are small; for larger effects, it tends to underestimate. Nevertheless, across the three outcome types, the conditional RMSE of RB is inferior to that of the naive estimator only under the H0 scenarios. For confidence intervals, the Wald interval based on the naive estimator provides poor coverage due to treatment selection and multi‐stage early stopping. In contrast, the stage‐wise intervals achieve nominal coverage at the expense of wider widths. Conditional intervals based on RB and CMUE‐MLE are generally conservative, with CMUE‐MLE being the more conservative of the two yet yielding narrower intervals. CMUE‐ZERO‐based intervals improve upon the naive estimator but still fail to achieve nominal coverage in most scenarios.

Taken together, under typical usage scenarios, we recommend employing the RB estimator and its corresponding confidence interval, as they provide the most reliable and unbiased inference across a broad range of settings. Moreover, CBAE‐based and CMUE‐based approaches also constitute valuable alternatives, especially when aligned with specific study objectives and underlying data assumptions. However, because the magnitude and direction of bias depend on the selection, continuation, and stopping probabilities, estimator choice should be supported by design‐specific simulation studies that report these probabilities together with bias, RMSE, coverage probability, and interval width. For practical applications, however, it remains essential to conduct comprehensive simulation studies prior to trial initiation to ensure the selection of the most appropriate estimation strategy.

This study also has several limitations. First, to ensure broad applicability across different endpoint types, the proposed methods are developed under a framework that relies on the asymptotic multivariate normality of Score statistics. This approximation is justified under large‐sample assumptions, and its accuracy may be reduced in finite‐sample settings. Accordingly, we recommend that investigators conduct design‐specific simulation studies to assess the adequacy of the asymptotic approximation in practice, particularly for binary and time‐to‐event endpoints. Second, as noted by Chapter 17 of Jennison and Turnbull [7], the validity of the independent‐increment property of Score statistics requires that the Fisher information levels be statistically independent of earlier interim statistics. When this condition is violated—for example, if information levels are adaptively modified based on interim results—the assumed independence structure may no longer hold, potentially affecting the validity of the proposed inference procedures.

In deriving the RB estimator, we have identified the sufficient statistic for θ without addressing their completeness. This suggests that alternative sufficient statistics may exist that could yield unbiased estimators with lower variance. Following the approach of Stallard and Kimani [21], in which the sum of data across treatments serves as the sufficient statistic, one can obtain estimators with reduced variance and narrower confidence intervals under both normal and binary endpoints (simulation results are provided in the Supporting Information). Nevertheless, we retain the current statistic in this study to preserve the distributional generalizability of the proposed method, and simulation studies also confirm the exhibits favorable properties of the current RB estimator. However, developing a RB estimator with enhanced distributional generalizability and improved estimation efficiency remains an important direction for future research.

The considered trial design of this study uses the selection rule of choosing the experimental treatment with the largest Wald test statistic as an example. In practical applications, depending on the specific research interests of the investigator, the selection rules may vary considerably. Nevertheless, we expect that the methods described above can be extended to settings where the selection rule is derived from Score statistics of phase II. Examples include selecting the group according to the largest sample mean or the smallest p value, as well as the rule considered by Stallard and Kimani [21], which involves selecting the lowest dose with an observed clinical response rate within two percentage points of the best‐performing dose. It should be emphasized that, as the selection rule becomes more complex, the associated computational burden can increase substantially. In such cases, it may no longer be possible to simplify the required high‐dimensional integrals via linear transformations, and Monte Carlo methods may be needed to approximate the relevant conditional quantities. Moreover, when the selection rule is more flexible, such as those based on safety indicators or short‐term endpoints, the methods proposed in this study may no longer be applicable. In such cases, researchers might need alternative estimation methods, such as the parameter estimation methods for 2‐in‐1 designs proposed by Li et al. [30], which also represents a potential direction for future exploration. Finally, if phase II itself is a multi‐stage design, the framework considered here may require substantial modification, since the selection event becomes considerably more complex and is determined by a higher‐dimensional collection of statistics.

We provide an R program used to conduct the simulations in Section 3 and to implement the example in Section 4, which is publicly available at https://github.com/syzhu0408/conditional‐estimation‐multi‐stage‐seamless‐phase2‐3, where detailed simulation results are also provided. Developing a dedicated statistical package to facilitate wider adoption represents a direction for future work.

Funding

This work was supported by the National Natural Science Foundation of China (No. 82273727).

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Figure S1: The conditional bias and RMSE of the point estimators for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S2: The conditional coverage probability and interval width of the confidence intervals for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S3: The conditional bias and RMSE of the point estimators for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S4: The conditional coverage probability and interval width of the confidence intervals for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S5: The conditional bias and RMSE of the point estimators for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S6: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S7: The conditional bias and RMSE of the point estimators for the binomial distribution under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S8: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J = 4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III.)

Figure S9: The conditional bias and RMSE of the point estimators for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S10: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S11: The conditional bias and RMSE of the point estimators for the survival data under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S12: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S13: The conditional bias and RMSE of the point estimators for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S14: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S15: The conditional bias and RMSE of the point estimators for the survival data under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S16: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S17: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S18: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=4. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S19: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S20: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S21: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=4. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S22: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S23: Flowcharts of the multi‐iteration and single‐iteration CBAE algorithms.

Figure S24: Flowchart of the CMUE‐MLE and CMUE‐ZERO algorithms.

Figure S25: Flowchart of the Rao–Blackwellized estimator algorithm.

Table S1: Stopping boundaries and stage‐wise sample size in the simulation study under the normal distribution.

Table S2: Stopping boundaries and stage‐wise sample size in the simulation study under the binomial distribution.

Table S3: Stopping boundaries and cumulative number of deaths at each interim analysis under the Weibull distribution.

Table S4: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=3.

Table S5: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=4.

Table S6: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=5.

Table S7: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=3.

Table S8: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=4.

Table S9: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=5.

Table S10: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=3.

Table S11: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=4.

Table S12: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=5.

Table S13: The calculation of each performance measure and the corresponding Monte Carlo standard error (MCSE) formula.

Table S14: The summary values of Monte Carlo standard errors (MCSE) for each estimator across all endpoints.

Table S15: Iterative bias correction process of the CBAE‐MI algorithm.

SIM-45-0-s001.docx (3.9MB, docx)

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analyzed during the current study.

References

  • 1. Broglio K., Cooner F., Wu Y., et al., “A Systematic Review of Adaptive Seamless Clinical Trials for Late‐Phase Oncology Development,” Therapeutic Innovation and Regulatory Science 58, no. 5 (2024): 917–929, 10.1007/s43441-024-00670-1. [DOI] [PubMed] [Google Scholar]
  • 2. Teng Z., Tian Y., Liu Y., and Liu G., “Seamless Phase 2/3 Oncology Trial Design With Flexible Sample Size Determination,” Statistics in Medicine 39, no. 18 (2020): 2373–2386, 10.1002/sim.8543. [DOI] [PubMed] [Google Scholar]
  • 3. Stallard N. and Friede T., “A Group‐Sequential Design for Clinical Trials With Treatment Selection,” Statistics in Medicine 27, no. 29 (2008): 6209–6227, 10.1002/sim.3436. [DOI] [PubMed] [Google Scholar]
  • 4. Gao P. and Li Y., “Adaptive Two‐Stage Seamless Sequential Design for Clinical Trials,” Journal of Biopharmaceutical Statistics 35, no. 4 (2025): 565–587, 10.1080/10543406.2024.2342518. [DOI] [PubMed] [Google Scholar]
  • 5. Whitehead J., “On the Bias of Maximum Likelihood Estimation Following a Sequential Test,” Biometrika 73, no. 3 (1986): 573–581, 10.1093/biomet/73.3.573. [DOI] [Google Scholar]
  • 6. Emerson S. S. and Fleming T. R., “Parameter Estimation Following Group Sequential Hypothesis Testing,” Biometrika 77, no. 4 (1990): 875–892, 10.1093/biomet/77.4.875. [DOI] [Google Scholar]
  • 7. Jennison C. and Turnbull B. W., Group Sequential Methods With Applications to Clinical Trials (Chapman and Hall/CRC, 1999), https://www.taylorfrancis.com/books/9781584888581. [Google Scholar]
  • 8. Chang M. N., Wieand H. S., and Chang V. T., “The Bias of the Sample Proportion Following a Group Sequential Phase II Clinical Trial,” Statistics in Medicine 8, no. 5 (1989): 563–570, 10.1002/sim.4780080505. [DOI] [PubMed] [Google Scholar]
  • 9. Shimura M., Gosho M., and Hirakawa A., “Comparison of Conditional Bias‐Adjusted Estimators for Interim Analysis in Clinical Trials With Survival Data,” Statistics in Medicine 36, no. 13 (2017): 2067–2080, 10.1002/sim.7258. [DOI] [PubMed] [Google Scholar]
  • 10. Stallard N. and Todd S., “Point Estimates and Confidence Regions for Sequential Trials Involving Selection,” Journal of Statistical Planning and Inference 135, no. 2 (2005): 402–419, 10.1016/j.jspi.2004.05.006. [DOI] [Google Scholar]
  • 11. Cohen A. and Sackrowitz H. B., “Two Stage Conditionally Unbiased Estimators of the Selected Mean,” Statistics and Probability Letters 8, no. 3 (1989): 273–278, 10.1016/0167-7152(89)90133-8. [DOI] [Google Scholar]
  • 12. Kimani P. K., Todd S., and Stallard N., “Conditionally Unbiased Estimation in Phase II/III Clinical Trials With Early Stopping for Futility,” Statistics in Medicine 32, no. 17 (2013): 2893–2910, 10.1002/sim.5757. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Robertson D. S., Prevost A. T., and Bowden J., “Unbiased Estimation in Seamless Phase II/III Trials With Unequal Treatment Effect Variances and Hypothesis‐Driven Selection Rules,” Statistics in Medicine 35, no. 22 (2016): 3907–3922, 10.1002/sim.6974. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Carreras M. and Brannath W., “Shrinkage Estimation in Two‐Stage Adaptive Designs With Midtrial Treatment Selection,” Statistics in Medicine 32, no. 10 (2013): 1677–1690, 10.1002/sim.5463. [DOI] [PubMed] [Google Scholar]
  • 15. Kimani P. K., Todd S., and Stallard N., “A Comparison of Methods for Constructing Confidence Intervals After Phase II/III Clinical Trials,” Biometrical Journal 56, no. 1 (2014): 107–128, 10.1002/bimj.201300036. [DOI] [PubMed] [Google Scholar]
  • 16. Khan J. N., Kimani P. K., Glimm E., and Stallard N., “Adjusting for Treatment Selection in Phase II/III Clinical Trials With Time to Event Data,” Statistics in Medicine 42, no. 2 (2023): 146–163, 10.1002/sim.9606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Robertson D. S., Choodari‐Oskooei B., Dimairo M., Flight L., Pallmann P., and Jaki T., “Point Estimation for Adaptive Trial Designs I: A Methodological Review,” Statistics in Medicine 42, no. 2 (2023): 122–145, 10.1002/sim.9605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Robertson D. S., Choodari‐Oskooei B., Dimairo M., Flight L., Pallmann P., and Jaki T., “Point Estimation for Adaptive Trial Designs II: Practical Considerations and Guidance,” Statistics in Medicine 42, no. 14 (2023): 2496–2520, 10.1002/sim.9734. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Robertson D. S., Burnett T., Choodari‐Oskooei B., et al., “Confidence Intervals for Adaptive Trial Designs I: A Methodological Review,” Statistics in Medicine 44, no. 18–19 (2025): e70174, 10.1002/sim.70174. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Robertson D. S., Burnett T., Choodari‐Oskooei B., et al., “Confidence Intervals for Adaptive Trial Designs II: Case Study and Practical Guidance,” Statistics in Medicine 44, no. 18–19 (2025): e70202, 10.1002/sim.70202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Stallard N. and Kimani P. K., “Uniformly Minimum Variance Conditionally Unbiased Estimation in Multi‐Arm Multi‐Stage Clinical Trials,” Biometrika 105, no. 2 (2018): 495–501, 10.1093/biomet/asy004. [DOI] [Google Scholar]
  • 22. FDA/CDER , “Adaptive Designs for Clinical Trials of Drugs and Biologics [Internet],” https://www.fda.gov/media/78495/download.
  • 23. Fairbanks K. and Madsen R., “ P Values for Tests Using a Repeated Significance Test Design,” Biometrika 69, no. 1 (1982): 69–74, 10.1093/biomet/69.1.69. [DOI] [Google Scholar]
  • 24. Lehmann E. L. and ScheffÉ H., “Completeness, Similar Regions, and Unbiased Estimation‐Part I,” in Selected Works of E. L. Lehmann [Internet], ed. Rojo J. (Springer US, 2012), 233–268, 10.1007/978-1-4614-1412-4_23. [DOI] [Google Scholar]
  • 25. Whitehead J., Desai Y., and Jaki T., “Estimation of Treatment Effects Following a Sequential Trial of Multiple Treatments,” Statistics in Medicine 39, no. 11 (2020): 1593–1609, 10.1002/sim.8497. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Stallard N. and Todd S., “Sequential Designs for Phase III Clinical Trials Incorporating Treatment Selection,” Statistics in Medicine 22, no. 5 (2003): 689–703, 10.1002/sim.1362. [DOI] [PubMed] [Google Scholar]
  • 27. O'Brien P. C. and Fleming T. R., “A Multiple Testing Procedure for Clinical Trials,” Biometrics 35, no. 3 (1979): 549, 10.2307/2530245. [DOI] [PubMed] [Google Scholar]
  • 28. Morris T. P., White I. R., and Crowther M. J., “Using Simulation Studies to Evaluate Statistical Methods,” Statistics in Medicine 38, no. 11 (2019): 2074–2102, 10.1002/sim.8086. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Bretz F., Koenig F., Brannath W., Glimm E., and Posch M., “Adaptive Designs for Confirmatory Clinical Trials,” Statistics in Medicine 28, no. 8 (2009): 1181–1217, 10.1002/sim.3538. [DOI] [PubMed] [Google Scholar]
  • 30. Li W., Bai X., Deng Q., Liu F., and Chen C., “Estimation of Treatment Effect in 2‐In‐1 Adaptive Design and Some of Its Extensions,” Statistics in Medicine 40, no. 11 (2021): 2556–2577, 10.1002/sim.8917. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1: The conditional bias and RMSE of the point estimators for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S2: The conditional coverage probability and interval width of the confidence intervals for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S3: The conditional bias and RMSE of the point estimators for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S4: The conditional coverage probability and interval width of the confidence intervals for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S5: The conditional bias and RMSE of the point estimators for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S6: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S7: The conditional bias and RMSE of the point estimators for the binomial distribution under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S8: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J = 4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III.)

Figure S9: The conditional bias and RMSE of the point estimators for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S10: The conditional coverage probability and interval width of the confidence intervals for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S11: The conditional bias and RMSE of the point estimators for the survival data under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S12: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=3. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S13: The conditional bias and RMSE of the point estimators for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S14: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S15: The conditional bias and RMSE of the point estimators for the survival data under J=5. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S16: The conditional coverage probability and interval width of the confidence intervals for the survival data under J=4. (Rows 1–4 correspond to the four scenarios, and the horizontal axis of these plots indicates the selected treatment at the end of phase II, with the numbers in parentheses denoting the corresponding probability of continuation to phase III).

Figure S17: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=3. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S18: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=4. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S19: Simulation results comparing θ^RB−KS and θ^RB for the normal distribution under J=5. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S20: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=3. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S21: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=4. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S22: Simulation results comparing θ^RB−KS and θ^RB for the binomial distribution under J=5. (Rows 1–4 correspond to the four scenarios, and Columns 1–4 correspond to different evaluation metrics).

Figure S23: Flowcharts of the multi‐iteration and single‐iteration CBAE algorithms.

Figure S24: Flowchart of the CMUE‐MLE and CMUE‐ZERO algorithms.

Figure S25: Flowchart of the Rao–Blackwellized estimator algorithm.

Table S1: Stopping boundaries and stage‐wise sample size in the simulation study under the normal distribution.

Table S2: Stopping boundaries and stage‐wise sample size in the simulation study under the binomial distribution.

Table S3: Stopping boundaries and cumulative number of deaths at each interim analysis under the Weibull distribution.

Table S4: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=3.

Table S5: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=4.

Table S6: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the normal distribution under J=5.

Table S7: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=3.

Table S8: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=4.

Table S9: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the binomial distribution under J=5.

Table S10: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=3.

Table S11: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=4.

Table S12: Characteristics of the simulated trials and the convergence probability of the CBAE‐MI for the survival data under J=5.

Table S13: The calculation of each performance measure and the corresponding Monte Carlo standard error (MCSE) formula.

Table S14: The summary values of Monte Carlo standard errors (MCSE) for each estimator across all endpoints.

Table S15: Iterative bias correction process of the CBAE‐MI algorithm.

SIM-45-0-s001.docx (3.9MB, docx)

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analyzed during the current study.


Articles from Statistics in Medicine are provided here courtesy of Wiley

RESOURCES