Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2022 Sep 23.
Published in final edited form as: Pharm Stat. 2021 Jun 4;20(6):1235–1248. doi: 10.1002/pst.2143

Bayesian Single-Arm Phase II Trial Designs with Time-to-Event Endpoints

Jianrong Wu 1,*, Haitao Pan 2,*, Chia-Wei Hsu 2
PMCID: PMC9502026  NIHMSID: NIHMS1824474  PMID: 34085764

Abstract

For the cancer clinical trials with immunotherapy and molecularly targeted therapy, time-to-event endpoint is often a desired endpoint. In this paper, we present an event-driven approach for Bayesian one-stage and two-stage single-arm phase II trial designs. Two versions of Bayesian one-stage designs were proposed with executable algorithms and meanwhile, we also develop theoretical relationships between the frequentist and Bayesian designs. These findings help investigators who want to design a trial using Bayesian approach have an explicit understanding of how the frequentist properties can be achieved. Moreover, the proposed Bayesian designs using the exact posterior distributions accommodate the single-arm phase II trials with small sample sizes. We also proposed an optimal two-stage approach, which can be regarded as an extension of Simon’s two-stage design with the time-to-event endpoint. Comprehensive simulations were conducted to explore the frequentist properties of the proposed Bayesian designs and an R package BayesDesign can be assessed via R CRAN for convenient use of the proposed methods.

Keywords: Bayesian design, phase II trial, proportional hazards, sample size calculation, time-to-event endpoint

1. Introduction

For the cancer clinical trials with immunotherapy and molecularly targeted therapy, time-to-event endpoint, such as progression-free survival (PFS) is often a desired endpoint. Single-arm phase II trial designs with time-to-event endpoints have been largely studied by the frequentist approaches prospectively19, while Bayesian counterpart of such trial designs are limited. Thall et al10 proposed a Bayesian phase II design based on posterior probability of the mean survival time. Zhao et al11 developed a Bayesian decision theoretic two-stage phase II survival trial. Shi and Yin12 proposed a Bayesian enhancement two-stage design for single-arm phase II trial with time-to-event endpoints. Yuan et al13 proposed a practical Bayesian design for platform trials with time-to-event endpoints. Zhou et al14,15 developed several optimal Bayesian multi-stage phase II trial designs for the binary endpoint, time-to-event endpoint and other mixed endpoints, which can maximize the power while controlling the type I error control, and later Lin et al16 provided an extension to deal with the late-onset outcome.

One of advantages of the Bayesian methods is that the study designs can be derived from the exact posterior distributions which accommodate the single-arm phase II trial, for example, most of pediatrics cancers are rare, with small sample sizes, and the frequentist designs based on the large sample asymptotic distributions may not be appropriate. However, all the above methods rely on that the time-to-event following a parametric model, such as, an exponential or Weibull distribution, and sample size calculation requires the assumption of an uniform accrual and censoring distributions. Nevertheless, in practice, it is extremely challenges to make the assumption for the accrual distribution due to uncontrollable accrual rate, and the distribution of censoring time is unknown as well. All the above issues would result in an inaccurate estimation of sample size, which elevates the risk of drug development. To overcome these issues, especially parametric assumption for the time-to-event, we adapt a transformed time-to-event method used by Cotterill and Whitehead17, which is based on the proportional hazards assumption compared to the existing methods adopting parametric models, to propose two versions of Bayesian one-stage designs and a two-stage design which are all driven by the number of events instead of number of patients. As a result, the proposed designs are independent of the accrual and censoring distributions. The advantage of such distribution-free design is to avoid possible misspecifications of the time-to-event, accrual and censoring distribution assumptions which are often very difficult to be specified accurately prior to the study. For planning the trial, number of patients required for the trial can be provided by administrative consideration as usual. Simulations are conducted to study frequentist properties of the proposed Bayesian one-stage and two-stage designs.

The rest of this article is organized as follows. In section 2, we introduce the transformed time-to-event model. Section 3 presents two versions of Bayesian one-stage designs. Section 4 develops an optimal Bayesian two-stage design. In section 5, we conducted comprehensive simulations to evaluate the frequentist properties of the proposed one-stage and two-stage designs. The developed R package BayesDesign is shown in section 5 with examples and discussions are given in section 6.

2. Transformed time-to-event model

Suppose during the accrual phase of a trial with a new treatment, n patients have enrolled in the study. Assuming that {Ti, i = 1, . . . , n} are independent failure times with a common survival distribution S(t) and {Ci, i = 1, . . . , n} are independent censoring times, and {Ti} and {Ci} are independent, then, the observed failure times Xi and failure indicator ∆i are

Xi=minTi,CiandΔi=ITiCi,i=1,,n,

where I(·) is an indicator function.

Let S0(t) be a known survival distribution for the reference group. We are interested in testing the following hypothesis

H0:StS0tvs.H1:St>S0tforallt>0. (1)

Instead of assuming a distribution for the failure time, we assume a proportional hazards model

St=S0tδ,

where δ is the hazard ratio of the treatment group vs. reference group.

Then, the hypothesis (1) is equivalent to the following hypothesis:

H0:δ1vs.H1:δ<1. (2)

By using the transformed observations for testing the hypothesis (2), let T˜i=logS0Ti=1δlogSTi, where S(Ti) is uniform(0, 1), thus, T˜i follows an exponential distribution with rate parameter δ.

Let Wi=minT˜i,C˜i=logS0Xi, where C˜i=logS0Ci, we obtain the transformed data {Wi, ∆i, i = 1, . . . , n}, where Δi=ITiCi=IT˜iC˜i.

The likelihood function of the transformed data is given by

L=i=1nfWiΔiSWi1Δii=1nδΔieδWi=δdeδU,

where d=i=1nΔi is the number of events and U=i=1nWi is the total observation time (after transformation).

Since (d, U ) is a sufficient statistic, we can write the observed data by D = (d, U ).

3. Bayesian one-stage design

3.1. Design based on Gamma posterior

Cotterill and Whitehead17 used a gamma prior distribution on δ. Specifically, let δ ∼ Gamma(a, b) with density function π(δ) ∝ δa–1e, where a is a shape parameter b is a rate parameter. Increasing values of a correspond to increasing levels of certainty in the prior distribution of δ given b. Due to the conjugacy relationship between the gamma distribution and the likelihood function of δ, the posterior distribution of δ follows π(δ|D) ∝ δa+d–1e−(b+U)δ, which is again a gamma distribution Gamma(a + d, b + U ).

To testing the hypothesis (2) after obtaining the number of events d = m and total observation time U and its critical value k, from Bayesian analysis prospective, there is a convincing evidence that the new treatment is promising if the posterior probability P (δ < 1|m, U ) is no less than η when rejecting the null Uk, where η is a value chosen to be close to 1.

More specifically, we require that the posterior probability satisfies the following “Go” criterion to reject the null:

infUkP(δ<1m,U)η.

It can be shown that the posterior probability P (δ < 1|m, U ) is a monotone increasing function in U given the number of events d = m and the Gamma prior parameters a and rate b (Appendix 1).

Thus, we have

infUkP(δ<1m,U)=P(δ<1m,k)η. (3)

We declare that the new treatment is promising (rejection of the null) if Uk.

The posterior probability P (δ < 1|m, k) under the Gamma posterior distribution can be calculated by the numerical integration

P(δ<1m,k)=01(b+k)a+mΓ(a+m)ta+m1e(b+k)tdt.

In order to determine the number of events and critical value k for the reject region, given a hazard ratio δ1 (< 1) at the alternative hypothesis, we require that the posterior probability P (δ > δ1|m, U ) in the accepting region satisfies the following “No-go” criterion:

infU<kPδ>δ1m,U=Pδ>δ1m,kζ, (4)

where ζ is a value close to 1. This is because a large posterior probability P (δ > δ1|m, k) provides a convincing evidence that the new treatment is unpromising. Thus, we will not reject the null hypothesis and abandon the new treatment if U < k.

However, there is no explicit solution for the number of events m and critical value k under the Gamma posterior distribution. Therefore, the following search algorithm 1 is proposed to determine the number of events d = m, and the critical value k of the total observation time U to satisfy both equations of (3) and (4).

Algorithm 1

  1. We start with minimum number of events m = 1 up to maximum number of events mmax = 100.

  2. Given m = i, find a minimum total of observation time (after transformation) k0 such that P (δ < 1|m, k0) = η, and calculate the corresponding P (δ > δ1|m, k0).

  3. If P (δ > δ1|m, k0) ≥ ζ, we stop the algorithm and set m = i and k = k0; otherwise, we set m = i + 1 and repeat step 2.

The above proposed Bayesian one-stage design can be executed as follows in practice: enroll patients until m events are observed, then calculate the transformed total observation time U . If Uk, we declare the new treatment is promising and warrants to invest further development in a large scale phase III trial; otherwise, the new treatment is unpromising.

To control the frequentist type I error rate (α) and power (1 – β) for the proposed Bayesian design, we derived theoretical relationships between the α, β and η, ζ (Appendix 2) as follows:

Pδ<1m,m+mz1α=η, (5)
Pδ>δ1m,mz1α+z1β/1δ1=ζ. (6)

Thus, to obtain the desired frequentist type I error rate α and power 1 – β, we can use the above two equations to find out η and ζ to achieve desirable operating characteristics for the Bayesian design.

3.2. Design based on the normal posterior

The maximum likelihood estimate of δ is given by δ^=d/U and its variance estimate can be obtained from the inverse of Fisher information j1δ^=δ^1=d/U2.

Let θ = − log(δ) be the negative log hazard ratio, then, the maximum likelihood estimate of θ is θ^=logδ^. Using the delta method, an approximate variance estimate of θ^ is 1/d. Thus, θ^ is asymptotically normal distributed, that is θ^~Nθ,1/d.

We impose a normal prior distribution on θ, that is θ~Nθ0,σ02. Then, the posterior distribution of θ is Nθ˜,σ˜2, where the posterior mean and variance are given as follows:

θ˜=σ˜2dθ^+θ0σ02 and σ˜2=d+σ021. (7)

To obtain a Bayesian design given the number of events d = m and the total observation time U , we require that the posterior probability P (θ > 0|m, U ) in the rejection region satisfies

infUkP(θ>0m,U)=P(θ>0m,k)Φ(θ˜/σ˜)η. (8)

Meanwhile, given a hazard ratio δ1 (< 1) at the alternative hypothesis or θ1 = − log(δ1), we require that the posterior probability P (δ > δ1|m, U ) in the accepting region satisfies

infU<kPθ<θ1m,U=Pθ<θ1m,kΦθ1θ˜/σ˜ζ. (9)

Combining two equations (8) and (9), we have θ˜/s˜=zη and θ1θ˜/σ˜=zζ which reduces to the following equation

θ1σ˜=zη+zζ.

Substituting σ˜=m+σ021/2 into the above equation and solving m, we obtain the following formula

m=zη+zζ2θ121σ02. (10)

This simple formula is very informative since it provides the number of events required for the trial to achieve the posterior probabilities η and ζ.

Therefore, the trial rejects the null hypothesis if only if Uk with

k=m×exp1mzησ˜θ0σ02. (11)

It can be shown (Appendix 3) that the frequentist number of event m based on the Wald test Z=dθ^θ is given by

m=z1α+z1β2θ12, (12)

where α and β are the frequentist type I and II error rates. The frequentist rejection region is Z > z1–α which is equivalent to U>k˜ where

k˜=m×expz1αm.

Thus, to obtain the desired frequentist type I error rate α and power 1 – β for the Bayesian design under the normal prior, we can set ζ = 1 – β and calibrate η to satisfy the following equation

η=Φmm+σ021/2z1α.

It is easy to verify k=k˜ for a non-informative prior with σ02=. Thus, the Bayesian design with a non-informative prior is same as the frequentist design.

By comparing formulae (10) and (12), a Bayesian design with η = 1 – α and ζ = 1 – β requires a smaller number of events than that of the frequentist design, particular for the informative prior with a small value of σ02.

4. Optimal Bayesian two-stage design

To stop the trial early if the new treatment shows no sign of promising, the two-stage design is a most popular strategy adopted in real practices. In this section, we propose a flexible Bayesian two-stage design for the single arm studies with time-to-event endpoints.

Specifically, at the interim stage, based upon the observed number of events d1 = m1 and a critical value k1 for the total observation time U1, we make the decision of early stopping for futility or continuation of a study. The posterior probability P (δ > δ1|m1, U1) in the accepting region needs to satisfy the following futility stopping criterion:

infU1<k1Pδ>δ1m1,U1=Pδ>δ1m1,k1ζ1, (13)

where ζ1=η1ξt˜ with 0 < ξ < 1 and t˜=m1/m is the information fraction at the interim analysis. It is noted that ζ1 decreases with t˜, which reflects the clinical rationale that at the interim stage with fewer number of events, it would be prudent and conservative to terminate a study early for futility.

At the final stage, upon the observed number of events d = m and critical value k of the total observation time U , we require that the posterior probability P (δ > δ1|m, U ) in the accepting region satisfies the following ”No-go” criterion:

infU<kPδ>δ1m,U=Pδ>δ1m,kζ, (14)

where ζ = η(1 – ξ) since t˜=1 now.

Meanwhile, we also require that the posterior probability P (δ < 1|m, U ) in the rejection region satisfies the following “Go” criterion:

infUkP(δ<1m,U)=P(δ<1m,k)η. (15)

For the final decision-making, we reject the null hypothesis if Uk.

To determine the two-stage design parameters (m1, k1, m, k), we need to pre-specify an information fraction t˜=m1/m for the interim analysis, which determines the number of events m1=mt˜ for the interim stage. The tuning parameter ξ and η can be calibrated to determine the interim and final stage posterior probability cutoffs ζ1=η1ξt˜ and ζ = η(1 – ξ) such that the trial has desirable size and properties.

The proposed two-stage design algorithm 2 can be described as follows:

Algorithm 2

  1. Start with a minimum number of events m = 1 up to a maximum number of events mmax = 100.

  2. Given m = i, find a minimum total of observed survival time (after transformation) k0 such that P (δ < 1|m, k0) = η.

  3. Calculate the P (δ > δ1|m, k0). If P (δ > δ1|m, k0) ≥ ζ, where ζ = η(1 − ξ), we set m = i and k = k0; otherwise, we set m = i + 1 and repeat step 2.

  4. Calculate m1=mt˜ and solve the total observation time k1 such that such that P (δ > δ1|m1, k1) = ζ1, where ζ1=η(1ξt˜).

In the above algorithm, the two tuning parameters η and ξ are pre-specified, however, for implementing the proposed method, we use the following searching algorithm to find an ”optimal” set of η and ξ such that the power would be maximized. Therefore, we call the proposed method the optimal Bayesian two-stage design. The searching algorithm for the η and ξ is shown as below:

Searching Algorithm for η and ξ

  1. Start with minimum η = 0.8 up to maximum η= 0.95 by 0.05

  2. Given each η, loop through ξ from minimum 0.01 up to maximum 0.15 by 0.01.

  3. For each pair of (η, ξ) repeat Algorithm 2 and calculate the type I error and statistical power.

  4. Among all pairs of (η, ξ) which controls the type I error and maintain the statistical power, choose the one which maximizes the power.

In practice, the proposed optimal Bayesian two-stage design can be executed as follows: enroll patients until m1 events are observed, then, calculate the transformed total observation time U1 at first stage. If U1 > k1, the trial goes to second stage, otherwise the trial stops for futility. At the second stage, enroll patients until m events are observed, then calculate the transformed total observation time U . If U > k we declare the new treatment is promising and warrants for further study in a large scale phase III trial, otherwise, the new treatment is unpromising and is not worth for the further study.

5. Simulation

In this section, we illustrate the proposed Bayesian one-stage and two-stage designs and conduct simulations to study their frequentist properties. Assuming a 3-year survival probability of the reference group is S0(3) = 0.53. The trial is designed to detect hazard ratios of δ1 = 0.6, 0.7 or an equivalent negative log hazard ratios θ1 = − log(δ1).

For Bayesian design based on the normal posterior distribution, we set a normal prior distribution for θ with mean θ0 = 0.5θ1 and variance σ02=0.5,0.2,0.1. Then, the posterior distribution of θ is Nθ˜,σ˜2, where θ˜ and σ˜2 are given in equation (7). For Bayesian design based on the Gamma posterior distribution, we take the prior distribution of δ to be a Gamma(a, b) distribution with shape parameter a = 2, 5, 10 and rate parameter b = 2a/(1 +δ1), corresponding the posterior mean is a/b = 0.8 and standard deviation is a/b=0.8/a. Thus, with a small value of a, the standard deviation is large and the prior is a weak or non-informative prior, and with a large value of a, the standard deviation is small and the prior is more informative. As we derived in section 3.1, the posterior distribution of δ is a Gamma(a + d, b + U ), where d and U are the total number of events and total observation times (after transformation).

For the one-stage designs, in order to be confident in the final conclusion drawn from the trial, we choose η = 0.95, 0.9 and ζ = 0.9, 0.85, 0.8. As the cost of falsely stopping for futility is smaller than that of falsely continuing the trial with an inefficacious treatment into further large scale studies, we typically set ηζ.12 The required number of events m and critical value k were calculated using formulae (10) and (11) for the design based on the normal posterior distribution and using searching algorithm 1 for the design based on the Gamma posterior distribution, respectively.

The results were presented in Table 1. For example, with η = 0.9 and ζ = 0.8, a weak prior σ02=0.5 or a = 2, the required number of events was m = 16 and critical values k = 21.11 and k = 21.77 for the Gamma posterior and normal posterior, respectively. We can see that two methods give the almost identically designs. Similar results were also true for other design scenarios. We can also see that with a more informative prior, the required number of events decreases. For example, by using a design with the Gamma posterior, given that δ = 0.6, η = 0.9, ζ = 0.85, as a increase from 2 to 10, which indicates that the prior mean hazard rates are all 0.8 while the prior standard deviation of hazard rates decrease from 0.57 to 0.25, the required number of events decrease from 19 to 11.

Table 1:

Bayesian single-stage designs under various scenarios with confidence levels η = 0.9, 0.95 and ζ = 0.8, 0.85, 0.9, hazard ratio δ = 0.6, 0.7. Sample sizes were calculated under exponential survival distributions: Frequentist empirical type I error α^ and empirical power (EP) were estimated from 10,000 simulated trials under exponential distribution.

Parameters Design Gamma Posterior Design Normal Posterior

a δ η ζ m k n α^ EP m k n α^ EP
2 .6 .90 .80 16 21.11 41 .109 .795 16 21.77 41 .084 .764
5 .6 .90 .80 13 17.36 34 .122 .750 13 17.91 34 .100 .718
10 .6 .90 .80 8 11.11 21 .140 .644 8 11.48 21 .119 .614

2 .6 .90 .85 19 24.55 49 .109 .840 19 25.20 49 .088 .817
5 .6 .90 .85 16 20.80 41 .121 .808 16 21.33 41 .098 .786
10 .6 .90 .85 11 14.55 29 .144 .738 11 14.88 29 .129 .712

2 .6 .95 .85 25 33.58 64 .055 .837 26 35.63 67 .036 .820
5 .6 .95 .85 22 29.83 57 .061 .812 23 31.77 59 .042 .791
10 .6 .95 .85 17 23.58 44 .069 .748 18 25.34 46 .052 .735

2 .6 .95 .90 31 40.48 80 .053 .897 31 41.36 80 .036 .873
5 .6 .95 .90 28 36.73 72 .058 .878 28 37.49 72 .041 .854
10 .6 .95 .90 23 30.48 59 .072 .844 23 31.05 59 .056 .821

2 .7 .90 .80 33 40.26 76 .108 .797 34 41.99 78 .087 .779
5 .7 .90 .80 30 36.51 69 .119 .790 31 38.13 71 .102 .773
10 .7 .90 .80 25 30.26 57 .148 .776 26 31.68 60 .134 .767

2 .7 .90 .85 41 49.09 94 .112 .858 41 49.70 94 .086 .835
5 .7 .90 .85 38 45.34 87 .123 .850 38 45.84 87 .104 .829
10 .7 .90 .85 33 39.09 76 .144 .843 33 39.40 76 .132 .828

2 .7 .95 .85 54 66.35 124 .056 .855 55 68.30 126 .046 .833
5 .7 .95 .85 51 62.60 117 .062 .843 52 64.43 119 .050 .827
10 .7 .95 .85 46 56.35 105 .072 .832 47 57.98 108 .060 .819

2 .7 .95 .90 65 78.51 149 .056 .900 66 80.43 151 .044 .886
5 .7 .95 .90 62 74.76 142 .061 .896 63 76.57 144 .051 .882
10 .7 .95 .90 57 68.51 130 .070 .889 58 70.13 133 .062 .884

We conducted simulations to study frequentist properties of the proposed Bayesian designs. For each design scenario, 10,000 simulated trials (under exponential distribution) were used to estimate the frequentist empirical type I error rate and power. The frequentist type I error rate and power depend on the level of posterior probabilities η and ζ. Both simulations and numerical calculations using equations (5) and (6) showed that the frequentist type I error rate α and power 1 – β are very closer to those of the Bayesian posterior probabilities 1 – η and ζ, respectively, for one-stage design with a weak or non-informative prior. The simulation results also showed that the frequentist type I error rate inflated and power decreased a little bit as the prior became more informative if a same set cutoffs of η and ζ were applied under a given alternative δ. For instance, given δ = 0.6, with η = 0.9 and ζ = 0.8, as a increases from 2 to 10, the type I error rate inflated from 0.109 to 0.140 and from 0.084 to 0.119 for design using the Gamma and normal posteriors, respectively. The power decreased from 0.795 to 0.644 and from 0.764 to 0.614 for design using the Gamma and normal posteriors, respectively. This finding is expected since a same set cutoffs were used through various priors. On the other hand, this finding also instructs the statisticians the importance of calibrating the cutoff parameters of η and ζ. For example, if we still want to use a = 10 with δ = 0.6 with wish to control the type I error rate under 10% and have a power of at least 80%, from Table 1, we can set η = 0.95 and ζ = 0.90, then, the type I error rate is 7.2% and the power is 84% by using the Gamma posterior approach. We admit that this set of η and ζ over-controls the type I error rate, but just show an example and the proposed searching algorithm can be used for finding the η and ζ to achieve the target type I error rate and power goal.

For the two-stage design, we choose η = 0.9 and 0.95 and the tuning parameter ξ = 0.01. We solve the number of events m and critical value k for the second stage from equations (14) and (15) by the searching algorithm 2. Then, given interim information fraction t˜=0.4,0.5,0.6, we calculate the number of events m1=mt˜ for the first stage and solve the critical value k1 using equation (13). The obtained study designs for these three scenarios were given in Table 2. We conducted simulations to study frequentist properties of the proposed flexible Bayesian two-stage designs. For each design scenario, 10,000 simulated trials (under exponential distribution) were generated to estimate empirical type I error rate and power, probability of early termination (PET) at null and alternative, and expected sample size (ES) under the null and alternative. With η = 0.9, ξ = 0.01 and a weak prior a = 2, the type I error rate can be controlled below 10% while the power was kept above 80%; with η = 0.95, ξ = 0.01 and a weak prior a = 2, the type I error rate can be controlled below 5% while the power was kept above 90%. This is because of the relationship between the frequentist type I error rate α and power 1 – β and the Bayesian posterior probabilities η and ζ as derived for the one-stage design. For the two-stage design with the futility rule only for an interim analysis, the frequentist type I error rate and power are slightly reduced. For a more informative prior, the power was decreased and type I error rate was increased. The probabilities of early stopping were below 20% and above 60% under the null and alternative, respectively. The expected sample sizes under the alternative did not reduce much from the total sample size, however the expected sample sizes under the null were reduced more due to the high probability of early stopping. Furthermore, we explored sensitiveness of the proposed Bayesian design against the prior distributions. The results showed that the proposed Bayesian designs are quite robust and stable for the non-informative prior (see supplemental materials). Therefore, we recommend the non-informative prior to being used in real practice for keeping the study designs with a controlled type I error and high power.

Table 2:

Flexible Bayesian two-stage designs under various scenarios with prior parameter a = 2, 5, 10, information fraction t˜=0.4,0.5,0.6, and tuning parameters η = 0.9, 0.95 and ξ = 0.01. Empirical type I error α^ and power (EP), early stopping probabilities under null (PET0) and alternative (PET1), expected sample size under null (ES0) and alternative (ES1) were estimated from 10,000 simulation runs under exponential distribution.

Design parameters Stage I Stage II Freq. operating characteristics

η a t˜ ζ 1 ζ m 1 n 1 k 1 m n k α^ EP PET1 ES1 PET0 ES0
.9 2 .4 .896 .891 10 26 10.637 23 59 29.08 .082 .812 .122 54.97 .632 38.14
2 .5 .896 .891 12 31 13.404 23 59 29.08 .081 .818 .121 55.61 .697 39.48
2 .6 .895 .891 14 36 16.217 23 59 29.08 .082 .828 .119 56.26 .755 41.63

5 .4 .896 .891 8 21 8.253 20 52 25.33 .094 .787 .134 47.85 .590 33.71
5 .5 .896 .891 10 26 11.043 20 52 25.33 .093 .791 .141 48.33 .678 34.37
5 .6 .895 .891 12 31 13.875 20 52 25.33 .092 .798 .143 49.00 .740 36.46

10 .4 .896 .891 6 16 6.165 15 39 19.08 .110 .727 .176 34.95 .592 25.38
10 .5 .896 .891 8 21 9.012 15 39 19.08 .109 .732 .187 35.63 .684 26.69
10 .6 .895 .891 9 23 10.462 15 39 19.08 .108 .735 .193 35.91 .724 27.42

.95 2 .4 .946 .940 16 41 17.080 38 97 48.44 .044 .902 .058 93.75 .642 61.05
2 .5 .945 .940 19 49 21.213 38 97 48.44 .042 .904 .061 94.07 .723 62.30
2 .6 .944 .940 23 59 26.817 38 97 48.44 .043 .908 .062 94.64 .799 66.64

5 .4 .946 .940 14 36 14.683 35 90 44.69 .048 .888 .064 86.54 .612 56.95
5 .5 .945 .940 18 46 20.225 35 90 44.69 .047 .889 .070 86.92 .728 57.97
5 .6 .944 .940 21 54 24.470 35 90 44.69 .046 .892 .072 87.41 .789 61.60

10 .4 .946 .940 12 31 12.539 30 77 38.44 .057 .856 .085 73.09 .610 48.94
10 .5 .945 .940 15 39 16.761 30 77 38.44 .057 .860 .089 73.62 .708 50.10
10 .6 .944 .940 18 46 21.040 30 77 38.44 .056 .866 .090 74.21 .780 52.82

Finally, we conducted simulations to examine the sensitivity of the empirical operating characteristics under the exponential and Weibull distributions with various accrual and censoring distributions. The simulation results showed that the frequentist operating characteristics are robust to the underlying survival, accrual and censoring distributions (simulation results were given in supplemental materials).

In practice, number of patients or study duration may be also important information for planning a trial. For example, to calculate the sample size, by assuming a constant recruitment with accrual duration ta (years), follow-up time tf (years) and no loss to follow-up on the trial, then, failure probability p can be calculated as follows:18

p=11tatfta+tfS0(t)δdt,

where the integration can be calculated numerically using R function “integrate”.19 Thus, sample size is given by n = m/p. For illustration purpose, we show the calculated sample sizes under the exponential distribution for one-stage and two-stage designs given by n in Tables 12.

6. Software

We developed an R package BayesDesign for implementing the proposed design. The package consist of three functions, optimal_OneStage(·) and optimal_TwoStage(·) for implementing the proposed methods, and tot_tim(·) for calculating the transformed total observation time.

Here, we show how to design a two-stage design using the proposed method as an example, other examples can be found in the R package. In this two-stage design example, the endpoint is the PFS, and for the reference group, the 4-month PFS rate is assumed to be 17%. To detect a hazard ratio δ is 0.517 (equivalent to a 4-month PFS rate of 40% for the treatment group). By assuming the accrual duration is 6 months and follow-up duration is 6 months, which means that all patient will be followed 6 months after the last patient enrolled in the study. We implement the following codes (exponential distribution for the survival time was employed):

> library (Bayes Design)
> # 4−month PFS rate is 17%, that is S0 = 0.17,
> # detect a hazard ratio delta = 0 . 517
> # x = 4 , evaluating timepoint
> # frac = 0.5, information fraction is 50% for the interim analysis
>
> optimal_TwoStage (alphacutoff = 0.1, powercutoff = 0.8 , S0 = 0.17 ,
+ x = 4 , ta = 6 , tf = 6 , frac = .5, delta = 0.517 )
m1 n1 k1 m n k typeI power PET1 ES1 PET0 ES0
1 7 9 7.975 13 16 17.49 0.093 0.811 0.125 13.115 0.685 8.933
2 7 9 8.077 13 16 17.49 0.092 0.806 0.132 13.071 0.696 8.864

From the above outputs, we see the codes produced two rows of outputs. As we have mentioned previously, multiple values of optimal tuning parameters of η and ξ may be found, for instance, corresponding results of the two pairs of η and ξ (not shown here) were reported here. Actually, the function optimal_TwoStage(·) will output all identified values of η and ξ of their corresponding results by the empirical powers with decreased orders. In this example, we can see that the design by the 1st row have a power of 81.1% while the 2nd row 80.6% power. If we use the 1st row to design this study, we see that in the interim stage, if we observe m1 = 9 events or approximately n1 = 9 patients, we would compare the transformed observation total time (U1) based on 9 events to the cutoff k1 = 7.975, if U1 < k1, the study will terminate due to futility; otherwise, the study will continue until m=13 events occur, or approximately n=16 patients being enrolled and the last patient being followed eat least 6 months, if the transformed observation total time (U ) based on 13 events is greater than the cutoff k=17.49, we claim this study is a success or the null is rejected.

The above outputs also report the probabilities of early termination under H1 (PET1=12.5%) and H0 (PET0=68.5%), the expected sample sizes under H1 (ES1=13.12) and H0 (ES0=8.93).

To be noted, in the above code, actually we used a default argument of complete = "partial" in optimal_TwoStage(alphacutoff = 0.1, powercutoff = 0.8, S0 = 0.17, x = 4, ta = 6, tf = 6, frac = .5, delta = 0.517, complete = "partial") for the execution, which would hide the identified optimal values of η and ξ. If the statistician wants to have these optimal values, the following codes by using complete = "complete" can be executed.

> optimal_TwoStage ( alpha cutoff = 0.1 , powercutoff = 0.8 , S0 = 0.17 ,
+ x=4, ta = 6, tf = 6, delta = 0.517, complete = ” complete” )
eta xi ml nl kl m n k typeI power PET1 ES1 PET0 ES0
1 0.9 0.01 7 9 7.795 13 16 17.49 0.093 0.811 0.125 13.115 0.685 8.933
2 0.9 0.02 7 9 8.077 13 16 17.49 0.092 0.806 0.132 13.071 0.696 8.864

Since the optimal Bayesian two-stage design needs to compare the transformed observation total time in the interim and final stages to the corresponding cutoffs, for users to easily get the computed the transformed observation total time, we also provide a convenient function tot_time(·) for computing it. For example, if we have the observed time for first 9 patients as shown below in the vector obs_time, we can calculate the corresponding transformed observation time as 11.9833. The code can be executed as below.

> obs time <− c ( 3 . 003 , 4. 987 , 4. 306 , 2 . 561 , 1 . 575 , 0 . 329 , 1 . 940 , + 0 . 869 , 7 . 481 )
> S0 <− 0 . 17
> x <− 4
> tot_time ( obs_time = obs_time , S0 = S0 , x = x )
[ 1 ] 11 . 9833

7. Discussion

The existing frequentist methods for designing single-arm phase II trials with time-to-event endpoints are based on large sample asymptotic distributions which may not be appropriate when sample size is small. Bayesian methods can provide exact posterior distributions by taking the advantage of using conjugate priors.10,12,13,15 However, due to the complexity caused by censoring, the existing methods assumed that the time-to-event follows an exponential or Weibull distribution and focused on estimating the required sample size, while sample size calculation needs the assumption of accrual and censoring distributions. However, in practice, it is extremely difficult to make an accurate assumption for the accrual distribution due to the uncontrollable accrual rate. On the other hand, the distribution of censoring time is also unknown. All of these will compromise the performances of existing approaches. To overcome these difficulties, we present the event-driven approaches for one-stage and two-stage designs. The event-driven designs do not require knowledge of the underlying survival, accrual and censoring distributions, therefore, the proposed methods take great advantage to avoiding potential misspecification for the trial design.

Regulatory agencies commonly show concern if the type I error rate can be controlled by using the Bayesian designs. To demonstrate controlling of type I and II error rates by the proposed Bayesian designs, we derived theoretical relationships between the frequentist type I and II error rates α and β and Bayesian posterior probability cutoffs η and ζ. Thus, to obtain the desired frequentist type I error rate α and power of 1 – β, we can use equations (6) and (7) to adjust η and ζ for Bayesian one-stage designs. For the Bayesian two-stage design, by further calibrating the tuning parameter ξ, we provide an optimal algorithm to achieve the target frequentist type I error rate and power.

The proposed one-stage design using normal posterior distribution has an explicit formula for the number of events, which provides almost identical designs as using the exact gamma posterior distribution. The proposed Bayesian two-stage design provides a counter-part of Simon’s two-stage design which allows early futility stopping when the new treatment shows no sign of efficacy. However, the proposed Bayesian designs take the advantage of using time-to-event endpoints which becomes very popular for the cancer immunotherapy and targeted therapy trials. Furthermore, using the event-driven approach, the study designs are robust without assumptions of the underlying survival, accrual and censoring distributions. Finally, Bayesian designs based on the exact posterior distributions accommodate the single-arm phase II trials with small sample sizes. Therefore, the proposed event-driven Bayesian approaches can be considered as a useful tool for designing single-arm phase II cancer immunotherapy trials and rare cancer trials, like pediatrics cancer studies.

Supplementary Material

Supplemental Material

ACKNOWLEDGMENTS

This research was supported by the Biostatistics and Bioinformatics Shared Resource Facility of the University of Kentucky Markey Cancer Center and National Cancer Institute (NCI) Cancer Center Support Grant (P30CA177558) and American Lebanese Syrian Associated Charities (ALSAC).

Appendix 1: Monotonicity of the posterior probability

Let Xifi = f(x; a, bi), the gamma density with shape a and rate bi, i = 1, 2, where b1 < b2. It can be shown that f2f1 > 0 if only if xM , where M = a(log b2 – log b1)/(b2b1) > 0. We can show that for any constant C > 0, P (X2C) > P (X1C). As

PX2CPX1C=0Cf2f1dx.

If CM , then

0Cf2f1dx>0.

If C > M , then

0Cf2f1dx=0f2f1dxCf2f1dx=Cf2f1dx>0.

Hence, P (X2C) > P (X1C) for any C > 0.

Appendix 2: Relationship between frequentist and Bayesian type I error rates

The maximum likelihood estimate of δ is given by

δ^=dU.

and its variance estimate can be obtained from the inverse of Fisher information j1δ^=δ^1=d/U2. The Wald test statistic δ^ is given by

Z=δ^1V^ar(δ^)=dUd

which is asymptotically standard normal distributed and reject null hypothesis if Z<z1α. Thus, given number of events d = m, the frequentist type I error rate can be calculated as follows:

PZ<z1αH0=PU>m+mz1αH0.

Then, we calculate the Bayesian posterior probability

infUkP(δ<1m,U)=P(δ<1m,k)

with k=m+mz1α and solve the following equation to obtain frequentist type I error rate α for a Bayesian design with number of events d = m and posteriror probability η,

Pδ<1m,m+mz1α=η. (16)

Under alternative hypothesis δ = δ1, given number of events d = m, the frequentist power can be calculated as follows:

1β=PZ<z1αH1Φk1δ1/mz1α

Thus, we have

k=mz1α+z1β1δ1

We now calculate the Bayesian posterior probability

infU<kPδ>δ1m,U=Pδ>δ1m,k

with k=mz1α+z1β/1δ1 and solve following equation to obtain frequentist type II error rate β or power 1 – β for a Bayesian design with number of events d = m and posterior probability ζ,

Pδ>δ1|m,mz1α+z1β/1δ1=ζ. (17)

For example, for the study design with a = 2 (shape parameter of Gamma prior), δ1 = 0.6, m = 16, η = 0.9 and ζ = 0.8 (first case in Table 1), using equation (16), the calculated frequentist type I error is α = 0.109; using equation (17) with α = 0.109, the calculated frequentist power is 1 – β = 0.795.

Appendix 3: Frequentist number of events formula

Let θ = – log(δ) be negative log hazard ratio. Then, the maximum likelihood estimate of θ is θ^=logδ^. Using delta method, an approximate variance estimate of θ^ is 1/d. Thus, the Wald test Z=dθ^θ is approximately standard normal distributed and rejects the null H0:θ=0 if Z>z1α. Therefore, given the number of events d = m and alternative θ1, the power 1 – β of the test Z satisfies

1β=PZ>z1αH1Φmθ1z1α

Therefore, we have z1β=mθ1z1α. Solving m, we obtain

m=z1α+z1β2θ12.

Footnotes

DATA AVAILABILITY STATEMENT

Data sharing is not applicable to this article as no new data were created or analyzed in this study.

References

  • 1.Case LD and Morgan TM. Design of phase II cancer trials evaluating survival probabilities. BMC Medical Research Methodology, 2003; 3:1–12 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Huang B, Talukder E, Thomas N. Optimal two-stage phase II designs with long-termendpoints. Statistics in Biopharmaceutical Research, 2010; 2:51–61. [Google Scholar]
  • 3.Kwak M, Jung SH, Phase II clinical trials with time-to-event endpoints: Optimal two-stage designs with one-sample log-rank test, Statistics in Medicine 2014;33:2004–2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wu J Sample size calculation for the one-sample log-rank test. Pharmaceutical Statistics, 2015, 14:26–33. [DOI] [PubMed] [Google Scholar]
  • 5.Wu J Statistical Methods for Survival Trial Design: With Applications to Cancer Clinical Trials Using R CRC Press, Florida, 2018. [Google Scholar]
  • 6.Wu J, Chen L, Wei J, Weiss H, Chauhan A, Two-stage single-arm phase II survival trial design. Pharmaceutical Statistics, 2019a, 19:214–229. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Schmidt R, Kwiecien R, Faldum A, Berthold F, Hero B, Ligges S. Sample size calculation for the one-sample log-rank test. Statistics in Medicine, 2015;15:1031–1040. [DOI] [PubMed] [Google Scholar]
  • 8.Belin L, Rycke YD, Broet P. A two-stage design for phase II trials with time-to-event endpoint using restricted follow-up. Contemporary Clinical Trials Communications, 2017;8:127–134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Shan G Two-stage optimal designs based on exact variance for a single-arm trial with survival endpoints Journal of Biopharmaceutical Statistics 2020, 30:797–805 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Thall PF, Wooten LH, Tannir NM. Monitoring event times in early phase clinical trials: some practical issues. Clinical Trials 2005; 2:467–478. [DOI] [PubMed] [Google Scholar]
  • 11.Zhao L, Taylor JMG, Schuetze SM. Bayesian decision theoretic two-stage design in phase II clinical trials with survival endpoint. Statistics in Medicine 2012; 31:1804–1820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Shi H, Yin G. Bayesian enhancement two-stage design for single-arm phase II clinical trials with binary and time-to-event endpoints. Biometrics 2018; 74:1055–1064. [DOI] [PubMed] [Google Scholar]
  • 13.Yuan Y, Guo B, Munsell M, Lu K, Jazaeri A. MIDAS: A Practical Bayesian Design for Platform Trials with Molecularly Targeted Agents. Statistics in Medicine 2016; 35:3892–3906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Zhou H, Lee JJ, & Yuan Y (2017). BOP2: Bayesian optimal design for phase II clinical trials with simple and complex endpoints. Statistics in medicine, 36(21), 3302–3314. [DOI] [PubMed] [Google Scholar]
  • 15.Zhou H, Chen C, Sun L, Yuan Y. Bayesian optimal phase II clinical trial design with time-to-event endpoint. Pharmaceutical Statistics 2020. [DOI] [PubMed] [Google Scholar]
  • 16.Lin R, Coleman RL, Yuan Y. TOP: Time-to-Event Bayesian Optimal Phase II Trial Design for Cancer Immunotherapy. J Natl Cancer Inst 2020; 112:38–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Cotterill A, Whitehead J. Bayesian methods for setting sample szies and choosing allocation rates in phase II clinical trials with time-to-event endpoints. Statistics in Medicine 2015; 34:1889–1903. [DOI] [PubMed] [Google Scholar]
  • 18.Wu J Statistical Methods for Survival Trial Design: With Applications to Cancer Clinical Trials Using R CRC Press, Florida, 2018. [Google Scholar]
  • 19.R Core Team. R: A language and environment for statistical computing R Foundation for Statistical Computing, Vienna, Austria. 2014; URL http://www.R-project.org/ [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material

RESOURCES