Abstract
This paper is motivated by a randomized controlled trial to compare an endovascular procedure with conventional medical treatment for stroke patients, in which the endovascular procedure may be effective only in a subgroup of patients. Since the subgroup is not known at the design stage but can be learned statistically from the data collected during the course of the trial, we develop a novel group sequential design that incorporates adaptive choice of the patient subgroup among several possibilities which include the entire patient population as a choice. We define the type I and type II errors of a test in this design and show how a prescribed type I error can be maintained by using the closed testing principle in multiple testing. We also show how asymptotically optimal tests can be constructed by using generalized likelihood ratio statistics for parametric problems and analogous standardized or Studentized statistics for nonparametric tests such as Wilcoxon’s rank sum test commonly used for treatment comparison in stroke patients.
Keywords: adaptive selection, generalized likelihood ratio statistics, group sequential design, Kullback-Leibler information, multiple testing, normalized Wilcoxon statistic
1. Introduction
It is widely recognized that the comparative efficacy of a new treatment can depend on certain characteristics of the patients that are difficult to pre-specify at the design stage. In a randomized trial, ignoring characteristics that account for patient heterogeneity in response may yield a false negative result. On the other hand, narrowly defining the patient characteristics for inclusion and exclusion limits the proven usefulness of the treatment to a small patient subpopulation. A trial may also encounter difficulties in patient accrual when relatively few patients satisfy the stringent inclusion/exclusion criteria. Adaptive (data-dependent) choice of the patient subgroup to compare the new and control treatments is a natural compromise between ignoring patient heterogeneity and using stringent inclusion/exclusion criteria in the trial design and analysis.
In this paper we introduce some methods for adaptively choosing the patient subgroup and testing the efficacy of the new treatment on the chosen subgroup, with prescribed type I error probability for the data-dependent choice. Our work was motivated by a clinical trial suggested by our Stanford Neurology colleagues. The trial aims to compare standard medical therapy with the endovascular procedure to remove the clot in ischemic stroke. There are two baseline characteristics that define the subgroups of interest: One of them, an MRI-based measure referred to as DWI (diffusion weighted imaging) measures the size of the central infarct that is not considered salvageable, relative to the surrounding penumbra that has been affected by the loss of perfusion but is considered salvageable tissue if the area can be reperfused by recanalization of the artery that is blocked by the clot. The other is time since stroke in hours, which may moderate treatment effects by the way that the salvageable tissue is gradually degraded by chronic lack of perfusion. The clinical outcome is change from baseline, within three months, of an ordered categorical score (Rankin) measuring the change in impairment, that ranges from full recovery to complete disability or death. The clinical hypothesis is that in some subgroup of the patient population, defined by a region in the DWI × Time space, the null hypothesis of no treatment difference in 3-month modified Rankin score is false. The investigators hope to reject the null for the largest possible subgroup for which it is false, although they also recognize that the difference may not be significant over the entire patient population.
Further details of the adaptive choice of the treatment subgroup are given in Section 3. The test statistic commonly used for two-sample comparison of Rankin scores is Wilcoxon’s rank sum. To avoid obscuring the main ideas underlying our methodology by technical arguments involving nonparametric statistics, we first consider in Section 2 the parametric case of normally distributed treatment outcomes with common known variance and mean μj (or μ0j) for patient subgroup j and the new (or control) treatment. We develop the basic theory for adaptive choice of patient subgroup to compare the two treatments in this setting, first for designs with fixed sample size and then for a three-stage group sequential design that is used to approximate an adaptive design with mid-course determination of the patient subgroup and sample size re-estimation. Section 4 gives further discussion, extensions and concluding remarks.
2. Basic theory in prototypical normal setting
2.1. Asymptotic theory via Kullback-Leibler information and prevalence
Suppose n patients are randomized to the new and control treatments and the responses are normally distributed, with mean μj for the new treatment and μ0j for the control treatment if the patient falls in pre-defined subgroups Πj for j = 1, …, J. We assume that the responses have common known variance σ2. Extensions to unknown and possibly unequal variances will be discussed in Section 4. Let ΠJ denote the entire patient population on which a traditional randomized controlled trial (RCT) comparing the two treatments focuses. The Kullback-Leibler (KL) information number for ΠJ is
, where x+ denotes max(x, 0), noting that the comparison involves testing the one-sided null hypothesis μJ ≤ μ0J. The KL information number, or relative entropy, quantifies the amount of information in the sample to distinguish the treatment mean μJ from the control mean μ0J, and plays an important role in the asymptotic theory of efficient parametric tests; see [1], [2]. In particular, if
is not large enough, then the RCT may not have sufficient power, i.e., probability of showing a significant mean difference between the treatment and the control.
Let pj be the prevalence of patient subgroup Πj. Therefore npj is the expected number of subjects in Πj for a trial with a total sample size n that randomizes patients to the two treatments. The KL information number for Πj is
, which is the product of the prevalence pj, the KL information
from a pair of patients in Πj receiving the new treatment and control, respectively, and the expected number n/2 of such pairs. Not only does this show the trade-off between the prevalence of the patient subgroup and the magnitude of the difference μj − μ0j in choosing the patient subgroup to compare the two treatments, but it also suggests that an asymptotically optimal choice of subgroup is the maximizer of
over 1 ≤ j ≤ J. We can alternatively regard a sample size of m from Πj as an effective sample size of m/pj from the entire population.
2.2. An efficient approach for fixed sample size trials
Since there is typically little information from previous studies about the subgroup effect size μj − μ0j for j ≠ J, we begin with a standard RCT to compare the new treatment with the control over the entire population. Determination of the sample size n is based on power at an assumed effect size μJ − μ0J for testing HJ : μJ ≤ μ0J, recalling that μJ and μ0J are the mean responses of the new and control treatments, respectively, over the entire patient population. The main innovation of the proposed trial design is that it allows adaptive choice of the patient subgroup Î in the event HJ is not rejected, to continue testing Hi : μi ≤ μ0i with i = Î, and can claim the new treatment to be better than control for the patient subgroup Î if HÎ is rejected. We can interpret Î as an estimate of I = argmaxi≠J
. Letting θj = μj − μ0j and θ = (θ1, …, θJ), the probability of a false claim is the type I error
| (1) |
for θ ∈ Θ0, where Θ0 consists of all “null” parameter vectors θ such that θj ≤ 0 for some j ≤ J. Note that there is no false claim if HJ is rejected when θJ > 0, because in that case we would not go on to test a subpopulation. We want to maintain a type I error constraint α(θ) ≤ α, and will describe a procedure that uses the closed testing principle in the theory of multiple testing [3], [4] to satisfy this constraint.
We also want the procedure to be efficient in the sense of maximizing the power at θ ∉ Θ0. Since the null hypothesis is highly composite, a uniformly most powerful level-α test is not expected to exist. Instead we try to attain asymptotic efficiency as n → ∞. The KL information numbers in the preceding subsection fit nicely with the concept of Bahadur efficiency [5]. In Appendix A we give a precise statement of the asymptotic efficiency result of the proposed procedure together with its derivation. Note that the complement of Θ0 is {θ : θj > 0 for all j ≤ J} and therefore the type II error at θ ∉ Θ0 is
| (2) |
which is the probability of falsely ending up with no positive claim for the new treatment. We focus here on the description of the procedure and its implementation to maintain a prescribed type I error constraint on level α.
The trial randomly assigns n patients to the experimental treatment and the control. We reject HJ if
| (3) |
for i = J, where μ̂i(μ̂0i) is the mean response of patients in Πi from the treatment (control) arm and ni(n0i) is the the corresponding sample size. Otherwise we choose the patient subgroup Î ≠ J with the largest value of the generalized likelihood ratio statistic GLRi, which is the left-hand side of (3), among all subgroups i = 1, …, J − 1, and reject HÎ if GLRÎ ≥ cα. The threshold cα is chosen such that α(θ) ≤ α for all θ ∈ Θ0, and we next describe its computation.
It will be shown in Appendix A that
| (4) |
by making use of the closed testing principle [3], [4]. To compute α(0), we first use the approximations ni ≈ n0i (since study subjects are equally likely to receive the new treatment or control) and ni + n0i ≈ npi, thereby approximating the random variables ni and n0i by the constant npi/2 and nin0i/(ni+n0i) by npi/4. The error probability α(0) can then be computed as a sum of integrands, over certain sets, of the multivariate normal density of (Z1, …, ZJ) under θ = 0, where . The covariance matrix of this multivariate normal distribution is particularly simple in the case Π1 ⊂ ··· ⊂ ΠJ, which we assume in the sequel because of the motivating clinical trial described in the second paragraph of Section 1. In this nested case, for i ≤ j. Therefore, in this case the threshold cα can be determined by solving the equation
| (5) |
where ϕi(x) is the density function of the standard normal Zi.
2.3. A group sequential design with adaptive patient subgroup selection
The effect size μJ − μ0J underlying the sample size calculation in a RCT is typically based on some related studies and also on constraints on funding and study duration, which leads to the notion of “implied alternative” in [6, p.81]. The observed effect size may differ substantially from the assumed effect size during the course of the trial. This has led to adaptive designs with mid-course sample size re-estimation; see Chapter 8 of [7]. Since our design also allows treatment comparisons over smaller patient subgroups and there is usually little information at the beginning of the trial from previous studies about the effect sizes in the subgroups, an adaptive design that can both choose the patient subgroup and re-estimate the sample size is particularly attractive.
Bartroff and Lai [8], [9] have developed a theory of efficient adaptive design for sample size re-estimation that involves 3-stage GLR tests. At the first interim analysis, the sample size for the second stage is estimated. If the GLR test rejects the null hypothesis or stops early for futility at the second interim analysis, the trial stops. Otherwise the trial continues to the third stage which corresponds to the maximum sample size of the trial. These adaptive designs can be approximated by standard group sequential designs that do not estimate the sample size for the second stage; see [8], [10], [11]. We next extend these approximations of the adaptive design to a 3-stage group sequential design in which the last stage corresponds to the maximum sample size and the sample size up to the second stage is near the mid-point of the first-stage and final sample sizes.
As in the preceding section for fixed sample size designs, the maximum sample size is the fixed sample size for testing HJ, or some inflation thereof, determined by the power at some effect size δ for the entire population. As pointed out in [6], this maximum sample size has order 8δ2|log α|/σ2. At an interim analysis, we first test HJ and can then discontinue testing HJ early for efficacy or futility. If early stopping for efficacy occurs, we terminate the trial and claim that the new treatment is better than the control on average over the entire population. If stopping occurs for futility of testing HJ, then we accept HJ and continue the trial with the most promising patient subgroup, that is, the subgroup i ≠ J that maximizes GLRi defined by (3), but with ni and n0i replaced by the corresponding sample sizes at the time of interim analysis. If the test for HJ does not stop, then we continue to the next stage of the 3-stage design and repeat this procedure.
To be more specific, similarly to the square root of the left-hand side of (3), let
| (6) |
denote the test statistic for patient subgroup Πi at the ℓth interim analysis. At this analysis, the trial rejects HJ, terminates, and claims efficacy of the new treatment if
| (7) |
It accepts HJ and proceeds to the subgroup selection if
| (8) |
where is the test statistic of the alternative Ki : μi ≥ μ0i + δ, that is,
| (9) |
Once HJ is accepted at stage ℓ and a subgroup Î, with the largest value of for i ≠ J, is chosen, the future enrollment of the trial will include patients of this subgroup only, while the maximum total sample size N remains the same. Similar to testing HJ, the trial rejects HÎ at stage ℓ′ and terminates with an efficacy claim of the new treatment in this subgroup if
| (10) |
It may also terminate for futility if
| (11) |
If neither (7) nor (8) ever occurs, the trial proceeds to the final stage, in which case HJ is rejected if
| (12) |
and is accepted otherwise. Figure 1 gives a schematic summary of the sequence of decisions in this group sequential trial.
Figure 1.
Flow chart of the 3-stage group sequential design. Events that line up vertically occur at the same analysis.
Note that in the equations above, Î is a random variable with values ranging from 1 to J − 1. Moreover, the upper bound α(0) of the type I error, which will be explained in Appendix A, of this group sequential design can be decomposed into two parts similar to (1). The first part, P0(Reject HJ), is bounded by
The second part, P0(Accept HJ and reject HÎ), is bounded by
As shown in Appendix A, for ℓ ≤ ℓ′, where Nℓ is the total number of patients enrolled up to the ℓth interim analysis. Hence, the probabilities P1b and P1c can be computed by the recursive numerical integration [7, p.86] using the joint normal density of { , 1 ≤ ℓ ≤ 3}.
Recursive numerical integration can also be used to compute the probability
(and similarly for P2c). We use here the joint normal distribution of { , ℓ ≤ ℓ′ ≤ 3} conditional on , with ℓ representing the stage when HJ is accepted, and the joint normal distribution of and , j ∉ {i, J}, conditional on with i ≠ J. The thresholds b, b̃ < 0 and c are determined by solving the equations
| (13) |
for the prescribed maximum type I error α, power 1 − β at the alternative δ, and the proportion 0 < ε < 1 of type I (or type II) error spent at interim analyses.
3. Adaptive design of a randomized controlled trial to test endovascular therapy
3.1. Endovascular versus standard therapy following imaging evaluation for ischemic stroke
The clinical trial mentioned in the second paragraph of Section 1 is a randomized trial for patients experiencing acute ischemic stroke (IS) due to a proximal anterior circulation large vessel occlusion who are treated within 12 hours of IS onset. Patients will undergo MRI studies after enrollment but before randomization. The only FDA approved drug for the treatment of acute IS is IV-tPA, yet only 2 – 5% of IS patients receive it since most patients are outside the treatment window of the drug. The standard medical treatment to be compared with endovascular therapy may include IV-tPA administered up to 4.5 hours after symptom onset. The study involves multiple centers. At each center, a multi-disciplinary team of stroke neurologists and endovascular therapists who treat patients with acute IS will be recruited as investigators. Patients will be screened according to inclusion and exclusion criteria in the protocol, and an informed consent form will be submitted to the local IRB for each center before enrollment begins at the center. However, in certain cases the acute stroke patient may not be asked to give consent. If neurologists of the stroke care team find that the patient’s condition precludes them from giving informed consent, it will be obtained from the patient’s legally recognized representative as governed by local laws and stipulated by the local IRB. Patients will be monitored for symptomatic intracranial hemorrhage and other endovascular complications during hospital admission. Clinical follow-up of the patient will be performed at 30 and 90 days and will include the modified Rankin score and standard Case Report Forms. A Data and Safety Monitoring Board (DSMB) will be appointed for the study, comprising of representatives from multiple disciplines including neurology and biostatistics, and independent of the investigational sites.
The maximum sample size calculation is based on a non-adaptive trial to detect projected difference between endovascular treatment and standard medical management for these patients, using Wilcoxon’s rank sum of the Rankin scores as the test statistic to test the null hypothesis P(Y ≤ X) ≤ 1/2 of no improvement in the Rankin score X by the endovascular treatment compared to the score Y from the standard medical treatment arm in the entire population. We follow the approach in [12, p.419] to use a working model for a fixed-sample-size nonparametric two-sample test at the design stage of an adaptive or group sequential trial, in which the actual pattern of the treatment effects can be adapted to during the course of the trial. As pointed out in [12], the basic idea is to have power at the working model specifying the alternative close to that of the fixed-sample-size test that is asymptotically most powerful at the specified alternative. It is shown in [12] that the design also has good power properties at other alternatives and can also reduce the expected sample size under the type I error probability and maximum sample size constraints. Using this approach, the maximum sample size is planned to be N = 500. The pre-specified treatment effect size, significance level, and power used for this calculation will be described in the second paragraph of Section 3.2. Interim analysis is performed with the results of approximately 300 and 400 patients randomized to the endovascular and the standard treatments. As described in Section 1, the data from the interim analyses can be used to determine early stopping for futility or efficacy, or to identify a patient subgroup based on their baseline covariate in the DWI × Time space.
To keep the dimensionality of the subset selection manageable, we partition the covariate space into six categories, which we refer to as ‘low’ and ‘high’ DWI, and ‘short’, ‘medium’ and ‘long’ Time since stroke. We label these categories by cij that corresponds to the cell defined by DWI level i and Time level j (i = 1,2; j = 1,2,3), where the cell with the shortest time and lowest DWI is c11. The primary objective of the study is to test the null hypothesis HJ of no improvement of endovascular therapy over the standard medical treatment in the entire patient population ΠJ = ∪i,j cij, with J = 6. However, in the case that HJ is accepted, the prior experience of the investigators and the biological rationale behind the endovascular procedure suggest that the search at the interim analysis should begin with Π1 = c11 (for which endovascular treatment is believed to be the most efficacious) and greedily agglomerate neighboring cells into J − 1 = 5 subgroups Π2 = Π1 ∪ c12, Π3 = Π2 ∪ c21, Π4 = Π3 ∪ c22, Π5 = Π4 ∪ c31..
Let be the sum of the ranks of the Rankin scores from the endovascular treatment arm in the combined sample of Πi (consisting of both the endovascular and standard treatment arms) at the ℓth interim analysis.
The standardized Wilcoxon statistic is
| (14) |
In Appendix B, we derive an exact formula for , i ≤ j ≤ J, 1 ≤ ℓ ≤ ℓ′ ≤ 3, and use it to derive the limiting joint normal distributions of the , i ≤ j ≤ J and 1 ≤ ℓ ≤ ℓ′ ≤ 3, under P(Y ≤ X) = 1/2 and under contiguous alternatives P(Y ≤ X) = 1/2 + θ. The limiting joint normal distributions show that the procedures in Sections 2.2 and 2.3 for standardized sample means (6) of normally distributed observations can be applied to the standardized Wilcoxon statistics (14). In particular, in place of the statistic (9) for testing μi ≥ μ0i + δ to allow early stopping for futility, we now define
| (15) |
for testing Ki : P(Y ≤ X | Πi) = 1/2 + θ. With and defined by (14) and (15), we can apply the fixed sample size design in Section 2.2 and the group sequential design in Section 2.3 with σ = 1. The thresholds cα in (3) and b, b̃, c in (7)–(12) remain the same because of the aforementioned limiting multivariate normal distribution for the standardized Wilcoxon statistics.
3.2. Operating characteristics of adaptive clinical trial designs
To demonstrate the advantages of the proposed adaptive choice of patient subgroup, we compare the following RCT designs involving N = 500 patients in two simulation studies that are reported in Table 1 and 2.
Table 1.
Operating characteristics in the case qij = 1/6 for all i, j
| Design A | Design B | Design C | |||||
|---|---|---|---|---|---|---|---|
|
|
|
|
|||||
| Scenario | p (p1, …p6) | N | Pe | Pf | p (p1, …p6) | p | |
|
|
|
|
|||||
| S0 | μ = (0,0,0,0,0,0) | 5 | 424 | 2.2 | 54.6 | 4.8 | 4.8 |
| ν = (1,1,1,1,1,1) | (1.1, 0.7, 0.4, 0.4, 0.2, 2) | (0.9, 1, 0.9, 0.4, 0.2, 1.3) | |||||
|
|
|
|
|||||
| S1 | μ = (0.3,0.3,0.3,0.3,0.3,0.3) | 89 | 371 | 73.6 | 0.6 | 89.8 | 94.5 |
| ν = (1,1,1,1,1,1) | (0.9, 0.9, 0.6, 0.6, 1, 85) | (0.6, 0.6, 0.8, 1, 0.9, 86) | |||||
|
|
|
|
|||||
| S2 | μ = (0.3,0.3,0.3,0.3,0,0) | 80.5 | 432 | 42.2 | 1 | 77.7 | 69.9 |
| ν = (1,1,1,1,1,1) | (3.8, 4.4, 6.7, 13.6, 2.3, 49.7) | (2.1, 3.2, 5.1, 14.3, 2.8, 50.2) | |||||
|
|
|
|
|||||
| S3 | μ = (0.3,0.3,0,0,0,0) | 65.7 | 456 | 30.9 | 3.6 | 49 | 28.8 |
| ν = (1,1,1,1,1,1) | (14.7, 28.5, 6, 2.2, 0.7, 13.6) | (6.5, 20, 5.3, 2.2, 1, 14.1) | |||||
|
|
|
|
|||||
| S4 | μ = (0,0,0.3,0.3,0,0) | 26.5 | 451 | 11.6 | 21.8 | 25.9 | 28.9 |
| ν = (1,1,1,1,1,1) | (0.6, 0.2, 1.4, 7.8, 1.6, 15) | (0.3, 0.2, 1, 7.7, 2.5, 14.3) | |||||
|
|
|
|
|||||
| S5 | μ = (0.6,0,0,0,0,0) | 86.8 | 439 | 46.6 | 0.9 | 72.8 | 28.1 |
| ν = (1,1,1,1,1,1) | (60.7, 9, 2.6, 1, 0.4, 13.1) | (49.9, 5.9, 2.1, 0.9, 0.3, 13.7) | |||||
|
|
|
|
|||||
| S6 | μ = (0.5,0,0,0,0,0) | 78 | 443 | 44 | 2.3 | 57.4 | 22.3 |
| ν = (1,1,1,1,1,1) | (55.4, 8.6, 2.2, 1.4, 0.6, 10.2) | (36.9, 5.9, 2.5, 1.3, 0.5, 10.4) | |||||
|
|
|
|
|||||
| S7 | μ = (0.5,0.4,0.3,0,0,0) | 92.5 | 426 | 48 | 0.2 | 88.7 | 68.9 |
| ν = (1,1,1,1,1,1) | (7.4, 13.4, 20, 3, 0.6, 48.1) | (5.4, 11.7, 18.3, 2.5, 0.8, 50) | |||||
|
|
|
|
|||||
| S8 | μ = (0.5,0.4,0.3,0,0,0) | 83 | 438 | 40.4 | 1.1 | 77.1 | 61.3 |
| ν = (2,1.5,1,0.5,0.5,0.5) | (5.8, 7.4, 24.9, 3.5, 0.9, 40.5) | (3.6, 5, 23, 3.2, 0.9, 41.4) | |||||
|
|
|
|
|||||
| S9 | μ = (0.5,0.4,0.3,0,0,0) | 72.2 | 451 | 33.8 | 2.2 | 61.1 | 39.8 |
| ν = (0.5,1,1.5,2,2,2) | (15.2, 16.2, 13.7, 4.1, 1.3, 21.7) | (8.9, 11.9, 13, 3.6, 1.7, 21.9) | |||||
|
|
|
|
|||||
| S10 | μ = (0,0,0,0.3,0.4,0.5) | 47.9 | 423 | 35 | 12.6 | 50.1 | 69.2 |
| ν = (1,1,1,1,1,1) | (0.2, 0, 0, 0.1, 0, 47.6) | (0.1, 0.1, 0, 0, 0.2, 49.7) | |||||
|
|
|
|
|||||
Table 2.
Operating characteristics in the case q = (0.2, 0.1, 0.3, 0.1, 0.1, 0.2)
| Design A | Design B | Design C | |||||
|---|---|---|---|---|---|---|---|
|
|
|
|
|||||
| Scenario | p(p1, …p6) | N | Pe | Pf | p(p1, …p6) | p | |
|
|
|
|
|||||
| S0 | μ= (0,0,0,0,0,0) | 5.2 | 420 | 2.9 | 55.6 | 4.8 | 4.8 |
| ν = (1,1,1,1,1,1) | (1, 0.8, 0.5, 0.3, 0.2, 2.4) | (1, 0.8, 0.7, 0.5, 0.4, 1.3) | |||||
|
|
|
|
|||||
| S1 | μ= (0.3,0.3,0.3,0.3,0.3,0.3) | 89 | 370 | 73.6 | 0.5 | 89.7 | 94.5 |
| ν = (1,1,1,1,1,1) | (0.9, 0.7, 0.9, 0.8, 1, 84.8) | (0.5, 0.6, 0.6, 0.6, 0.9, 86.6) | |||||
|
|
|
|
|||||
| S2 | μ= (0.3,0.3,0.3,0.3,0,0) | 83 | 425 | 45.6 | 0.6 | 79.8 | 72.9 |
| ν = (1,1,1,1,1,1) | (3.5, 3.8, 7, 11, 3, 54.6) | (1.8, 3.1, 6.3, 10, 3, 55.5) | |||||
|
|
|
|
|||||
| S3 | μ= (0.3,0.3,0,0,0,0) | 64.5 | 454 | 31.9 | 4.5 | 46.1 | 25.6 |
| ν = (1,1,1,1,1,1) | (17.4, 30.7, 2.8, 1.2, 1, 11.3) | (7.9, 20.9,2.9, 1.5, 0.8, 12.2) | |||||
|
|
|
|
|||||
| S4 | μ= (0,0,0.3,0.3,0,0) | 33 | 451 | 15 | 17.5 | 35.1 | 36.5 |
| ν = (1,1,1,1,1,1) | (0.2, 0.1, 2.7, 7.3, 3.1, 19.6) | (0.3, 0.1, 2.6, 8.4, 3.4, 20.2) | |||||
|
|
|
|
|||||
| S5 | μ= (0.6,0,0,0,0,0) | 90.9 | 438 | 46.2 | 0.4 | 81.1 | 35.9 |
| ν = (1,1,1,1,1,1) | (56, 12.7, 1.6, 0.6, 0.5, 19.6) | (50, 9.4, 1.4, 0.8, 0.6, 18.9) | |||||
|
|
|
|
|||||
| S6 | μ = (0.5,0,0,0,0,0) | 82.9 | 441 | 44.9 | 1.9 | 67.3 | 28.4 |
| ν = (1,1,1,1,1,1) | (51.5, 14.1, 2.1, 1, 0.5, 13.7) | (40.2, 9.4, 2, 1, 0.6, 14) | |||||
|
|
|
|
|||||
| S7 | μ = (0.5,0.4,0.3,0,0,0) | 94.1 | 413 | 53.9 | 0.2 | 92.3 | 79 |
| ν = (1,1,1,1,1,1) | (5.7, 8.6, 13.9, 3.4, 0.8, 61.6) | (4.6, 8.3, 12.8, 2.5, 1, 63) | |||||
|
|
|
|
|||||
| S8 | μ= (0.5,0.4,0.3,0,0,0) | 87.9 | 418 | 49.9 | 0.8 | 84.9 | 75.5 |
| ν = (2,1.5,1,0.5,0.5,0.5) | (3.2, 3.7, 17.7, 3.3, 1, 58.9) | (2.2, 2.9, 17.4, 3.1, 1.2, 58.1) | |||||
|
|
|
|
|||||
| S9 | μ = (0.5,0.4,0.3,0,0,0) | 74.3 | 447 | 35.9 | 2.2 | 65.3 | 47.2 |
| ν = (0,5,1,1.5,2,2,2) | (16.1, 14.4, 10.7, 4, 1.8, 27.4) | (9.1, 11.2, 10.1, 3.9, 2, 29) | |||||
|
|
|
|
|||||
| S10 | μ = (0,0,0,0.3,0.4,0.5) | 38 | 430 | 26.8 | 18.2 | 37.5 | 57.2 |
| ν = (1,1,1,1,1,1) | (0.3, 0.1, 0, 0, 0, 37.5) | (0.3, 0.2, 0, 0, 0, 37) | |||||
|
|
|
|
|||||
Design A: The group sequential design described in Section 2.3 with interim and final analyses performed at 300, 400 and 500 patients.
Design B: The fixed sample size design described in Section 2.2, with 500 patients.
Design C: A fixed sample size RCT that compares the endovascular and standard treatments in 500 patients and does not consider subpopulation selection.
The simulation studies, each covering ten scenarios for the alternative hypothesis besides the null hypothesis denoted by S0, assume that the (normalized) Rankin score of patients in the standard treatment arm follows a standard normal distribution regardless of DWI × Time status, except in scenario S8 where the variance is 0.5, and in scenario S9 where the variance is 2. The Rankin scores of patients in the endovascular treatment arm are assumed to be normally distributed with mean μij and variance νij when the DWI × Time status falls in the cell cij, which has a prevalence qij. The μij are given in a vector μ consisting of μ11, μ12, μ21, μ22, μ31 and μ32 (in that order) and the νij are given in a similar vector for each scenario in the tables. The qij are given in a similar vector shown in the heading of each table, which reports the the average sample size N of the trial, the probability pi (reported in %) of rejecting the null hypothesis Hi, and the overall probability . The probabilities pe and pf of early stopping at any interim analyses by Design A for efficacy and futility, respectively, are also reported in %. Each result is based on 5000 simulations. Table 1 considers equal prevalence among individual cell cij. In this study, Design A uses b̃ = −1.46, b = 2.39, c = 2.31 corresponding to α = 0.05, β = 0.2 and ε = 0.5, while Design B uses cα = 2.16. To examine the impact of the prevalence, the patient subgroups considered in the second simulation study have different prevalence rates, with most patients falling in cell c21. In this simulation study, which is reported in Table 2, Design A uses b̃ = −1.46, b = 2.37, c = 2.28, and Design B uses cα = 2.14. Design A assumes the effect size θ = 0.064 to be the projected treatment effect used in the maximum sample size calculation at the design stage as mentioned in Section 3.1.
All three designs have type 1 error controlled at 5% and early stopping for futility can reduce the expected sample size by about 15–20% (scenario S0 in the tables). When the treatment effect is uniform over the patient subgroups (scenario S1), the power of Design A is slightly smaller than that of Design C (and nearly the same as that of Design B), but there is a considerable (>20%) reduction in expected sample size. In scenario S1, both Design A and Design B seldom choose a smaller subpopulation. When the actual distribution of the effect sizes agrees with the assumptions underlying the ordering of the subgroups (scenarios S2, S3, S5, S6, S7, S8, S9), Design A has the highest power, and also results in a sample size that is on average at least 10% less than those of the other two designs. Although Design B provides a chance to test treatment effects in a selected subpopulation after H0J is accepted, its power for testing the selected subpopulation is not as high as that of Design A. When the actual subgroup effects are contrary to the underlying assumptions (scenarios S4 and S10), Designs A and B exhibit a clear loss in power in comparison with Design C.
4. Discussion
This paper proposes novel efficient methods for the design and analysis of RCTs that allow adaptive choice of patient subgroups for confirmatory testing of a new treatment against a control. These patient subgroups can be defined by biomarkers, brain imaging, or other risk factors measured at baseline. Rosenblum et al. [13] have pointed out the usefulness of such methods in connection with testing the efficacy of antiretroviral medications for HIV positive patients. In particular, they note suggestive evidence from two recently completed RCTs of maraviroc that the treatment benefit may differ depending on the suppressive effect, as measured by the phenotypic severity score (PSS) at baseline, of the patient’s background therapy. Specifically, the estimated average treatment benefit among patients with PSS less than 3 was found to be larger than among those with PSS 3 or more. The approach used by [13] is to conduct multiple testing of the average treatment effect for the overall population and for each of the two PSS-defined subpopulations, leading to the hypotheses Hj, 1 ≤ j ≤ 3, with Π3 being the overall population and Π1, Π2 being disjoint subpopulations. Strong control of FWER imposes many constraints on the parametric space when the responses are assumed to be normally distributed, and [13] addresses the problem of constructing multiple testing procedures that satisfy these constraints and optimize power at a given set of alternatives by transforming a fine discretization of the multiple testing problem into a sparse linear program that has over a million variables and constraints and can be solved by recent advances in large, sparse linear programming. In contrast, our approach addresses testing treatment difference in an adaptively chosen patient subgroup, by partitioning the parameter space and defining the corresponding type I and type II errors. Instead of exact optimality, we resort to asymptotic efficiency via Kullback-Leibler information numbers, and use recent advances in group sequential and adaptive designs to further enhance adaptation and efficiency. Multiple testing methodology is used in our approach only as means to an end, rather than as the end (in the sense of controlling FWER) itself.
The basic theory underlying our approach is developed in Section 2 for normally distributed responses with common known variance σ. Since σ is typically unknown in practice, an obvious modification is to replace σ in Section 2.2 or 2.3 by the sample estimate σ̂, or σ̂ℓ at the ℓth interim analysis. Actually the methodology in Section 2 can be readily extended from the one-parameter normal family (with θj = μj − μ0j as the unknown parameter for the jth patient subgroup) to multiparameter exponential families (thereby allowing unknown variances for different patient subgroups and treatments even for normally distributed responses). The GLR statistics which provide a basic tool for our parametric approach in Section 2 can be readily extended to multiparameter families; see Section 3.4 of [6] which shows how the GLR statistics can be used to develop efficient group sequential tests in multiparameter exponential families.
Section 3 extends our approach from parametric to nonparametric settings. We have focused on the Wilcoxon statistic because it is commonly used for two-sample comparison of Rankin scores. Clearly the approach works much more generally for other nonparametric test statistics that have Chernoff-Savage-type representations; see [14]. In fact, as shown in Sections 3.3 and 3.4 of [12], the group sequential testing theory in [6] based on GLR statistics can be extended to nonparametric or semiparametric statistics with fully observable or censored survival outcomes. This also applies to the group sequential tests with adaptive patient subgroup choice developed herein.
Although the group sequential designs in Section 2.3 and in Tables 1 and 2 have substantial savings in expected sample size, the savings may be compromised by the longer duration to recruit patients with narrower eligibility criteria. If the study duration turns out to be longer than that of a fixed sample size design that allows adaptive choice of patient subgroup in the final analysis (Design B in the tables), then Design B may be preferred to the group sequential Design A despite the sample size savings. Simon and Simon [15] have proposed a class of adaptive enrichment designs, quite different and considerably more restrictive than that proposed herein, which allow the eligibility criteria to be adaptively updated at interim analyses of a trial. They also point out that “the more one restricts the eligibility criteria, the longer patient accrual will take” and that their adaptive enrichment design “is most powerful relative to the standard nonadaptive approach when only a small subset of patients benefit, however, this is exactly when the accrual rate is most decreased.” Note that our approach does not seek to find the patient subgroup for which the new treatment differs the most from the control. Instead, it seeks to find the patient subgroup that has the largest Kullback-Leibler information number
(which involves both the prevalence and the mean treatment difference) for j ≠ J, after accepting the null hypothesis HJ: μJ ≤ μ0J for the entire population ΠJ.
Yeatts et al. [16] recently reported a randomized clinical trial comparing IV-tPA followed by endovascular treatment with IV-tPA alone in IS patients. The trial intended to enroll 900 subjects, but was terminated after 656 subjects were randomized, based on a pre-specified criterion for futility stopping at the fourth interim analysis. Challenges facing the DSMB of the trial are described in [16]. One such challenge was the heterogeneity of the study population and the possibility of significant improvements of the endovascular approach in some patient subgroups, defined by severity of the stroke. However, these subgroups were not pre-specified and no statistically significant treatment difference between the two post-hoc severity-based subgroups was observed. It is pointed out in [16, p. 1412] that if it became clear during the study that the assumption of equal treatment difference over the two subgroups was invalid, the trial protocol might be amended to account for this unexpected difference. The new trial design proposed herein, therefore, provides an efficient method to address this difficulty.
Acknowledgments
Lai’s research was supported by the NSF Grant DMS-1106535 and the NIH grant 1 P30 CA124435-01. Lavori acknowledges support by the Clinical and Translational Science Award UL1 RR025744 to the Stanford Center for Clinical and Translational Education and Research (Spectrum) from the National Center for Research Resources, National Institutes of Health to Stanford University.
Appendix A: Bahadur efficiency, closed testing principle, and group sequential extensions
A.1. Bahadur efficiency of proposed procedure in Section 2.2
Recall that θj = μj − μ0j and θ = (θ1, …, θJ). For θ ∉ Θ0, θj > 0 for all 1 ≤ j ≤ J, and define
, i.e., j(θ) corresponds to the patient subgroup j with the largest KL information number
, which depends on θ and which will be denoted by
(θ). Note that
(θ) is linear in n and that the type II error β(θ) defined in (2) satisfies
| (16) |
in view of the asymptotic behavior of Gaussian tail probabilities. In fact,
(θ) is also the KL divergence In(θ, θ0), where θ0 ∈ Θ0 is the minimizer of
| (17) |
over λ ∈ Θ0; the expected value under θ of the logarithm of the likelihood ratio of θ to λ based on a sample of size n/2 is . Bahadur [5, 17] has shown that exp{−(1 + o(1)In(θ, θ0)} is the asymptotically minimal value (as n → ∞) for the type II error, at an alternative θ ∉ Θ0, of any level-α test of H: θ ∈ Θ0 based on a sample of size n/2. A level-α test whose type II error attains this asymptotically minimal rate at every alternative θ is said to be Bahadur efficient. This way of comparing tests via large deviation approximations applies to general parametric families. Even in cases where large deviation approximations to type II errors are difficult to derive, Bahadur [5, 17] suggests to fix the type II error at an alternative θ and to try deriving large deviation approximations for the type I error instead; specifically, for supλ∈Θ0 α(λ). He also suggests the closely related idea of deriving large deviation approximations for the p-value of the test at θ ∉ Θ0. These large deviation approximations lead to a limiting quantity infλ∈Θ0 q(λ|θ), which is called the Bahadur exact slope and is maximized where q = Q given by (17).
A.2. Closed testing principle and proof of (4)
To prove (4), recall that Hj: θj ≤ 0 and Θ0 = {θ: Hj is true for some j}. For θ ∈ Θ0, let J(θ) denote the set of the true hypotheses. The family-wise error rate (FWER) is defined as Pθ{Reject Hj for some j ∈ J(θ)}, and one can use the closed testing principle (CTP) to control FWER [3, 4]. This principle assumes that there is a level-α test of every intersection hypothesis ∩i∈IHi, I ⊂ {1, …, J}, and uses it to test the intersection hypothesis if and only if all intersection hypotheses ∩i∈I*Hi with I ⫋ I* are tested and rejected. Since a type I error is committed if and only if ∩i∈J(θ)Hi is tested and rejected under CTP, it follows that FWER ≤ α.
We now use a variant of CTP to prove (4). First note that the procedure in Section 2.2 is tantamount to (a) first testing HJ by using GLRJ and (b) using max1≤i≤J−1 GLRi to test the intersection null ∩i≠JHi if HJ is accepted; the critical value remains the same for (a) and (b). If we use CTP and apply maxi∈I GLRi to test the intersection hypothesis ∩i∈IHi with I ⊂ {1, …, J − 1}, this gives the FWER for testing Hi, 1 ≤ i ≤ J − 1 after accepting HJ. Note that
| (18) |
and that the square root of the GLR statistics are jointly normal random variables, with means ≤ 0 except for for i ∉ I in the second probability, under the respective intersection hypotheses. It then follows that FWER ≤ α(0), which is the type I error of testing the global null hypothesis under the parameter value 0. This shows that α(θ) ≤ α(0) for θ ∈ Θ0.
A.3. Extension of (4) to group sequential design
CTP is also applicable to group sequential design [18], [19]. In particular, for the 3-stage design in Section 2.3, subgroup selection immediately follows acceptance of HJ at an interim analysis and then proceeds with the selected subgroup, using the maximum of GLRi over i ≠ J as the selection criterion at the interim analysis. Although there seems to be additional complication because the 3-stage design allows early stopping for futility, this can be easily incorporated into the argument to show that an analog of (18) still holds for the group sequential design; see a similar argument in Appendix B of [20] for a closely related problem.
Appendix B: Joint distribution of standardized Wilcoxon statistics over subgroups and time
B.1. Formula for null covariances of standardized Wilcoxon statistics
Consider the standardized Wilcoxon statistics given by (14) for patient subgroup i at the ℓth interim analysis, under the null hypothesis of no treatment difference between the new and the standard treatments. We now derive an explicit formula for , with 1 ≤ i ≤ j ≤ J and 1 ≤ ℓ ≤ ℓ′. To simplify the notation, let W denote the Wilcoxon statistic based on a sample X1, …, Xm, Y1, …, Yn and let W′ denote that based on a larger sample X1, …, Xm′, Y1, …, Yn′ with m ≤ m′ and n ≤ n′. The null hypothesis of no treatment difference translates into the same distribution of Xi and Yj. Since the Xi and Yj are i.i.d., a standard combinatorial argument gives E(W) = m(m + n + 1)/2 and Var(W) = mn(m + n + 1)/12; see [21, p. 70 and pp. 332–333]. The same argument can be used to show that
| (19) |
Let Z = {W − m(m + n + 1)/2}/{mn(m + n + 1)/12}1/2 be the standardized Wilcoxon statistic and define Z′ similarly with W′, m′, n′ in place of W, m, n. Then it follows from (19) that
| (20) |
The preceding formula is the same as that in the case Z = (X̄m − Ȳn)/(m−1 + n−1)1/2 for the sample means X̄m and Ȳn of i.i.d. observations X1, …, Xm, Y1, …, Yn that have common variance 1. Here (m−1 + n−1)−1 = mn/(m + n) is the “information clock” of X̄m − Ȳn, as first observed by Robbins and Siegmund [22].
B.2. Limiting null distribution of ( : 1 ≤ i ≤ J, 1 ≤ ℓ ≤ L)
A key lemma of [22] is that for the difference of sample means X̄m − Ȳn of (standard) normal observations,
| (21) |
where B(t), t ≥ 0 is Brownian motion. Hence the joint distribution has independent increments in the information time mn/(m + n), and this independent increments property is used in the recursive integration algorithms in Section 2.3. For the standardized Wilcoxon statistics, (20) shows the same covariance structure as that for the standardized sample means (X̄m − Ȳn)/(m−1 + n−1)1/2. In fact, (21) still holds asymptotically as min(m, n) → ∞:
| (22) |
In (22), we denote Wilcoxon’s rank sum by Wm,n to emphasize the dependence on the sample sizes, and use ≈ to denote having the same limiting joint distribution. As in (21), we consider the joint distribution in (22) over a prescribed set of information times of the form mknk/(mk + nk), k = 1, …, K. To prove (22), we can use the Chernoff-Savage representation [23] of rank statistics in terms of sample means of i.i.d. random variables, plus an asymptotically negligible remainder term, and its refinement in [14] that gives stronger results on the remainder term.
For contiguous alternatives that satisfy converging to some finite constant μ, where N is the maximum sample size, the standardized Wilcoxon statistics in (14) still have the same limiting covariance structure as in the null case but its mean is of the order
recalling that patients are randomized to treatment and control and that the total number of patients in subgroup i enrolled up to the ℓth interim analysis is of the order piNℓ. Hence the limiting joint distribution of ( : 1 ≤ i ≤ J, 1 ≤ ℓ ≤ L) is that of a multivariate normal as in the null case but with means .
Footnotes
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Chernoff H. A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations. The Annals of Mathematical Statistics. 1952;23(4):493–507. [Google Scholar]
- 2.Chernoff H. Sequential Analysis and Optimal Design. Society for Industrial and Applied Mathematics; 1972. [Google Scholar]
- 3.Marcus R, Peritz E, Gabriel KR. On closed testing procedures with special reference to ordered analysis of variance. Biometrika. 1976;63(3):655–660. [Google Scholar]
- 4.Hochberg Y, Tamhane AC. Multiple Comparison Procedures. John Wiley & Sons, Inc; New York: 1987. [Google Scholar]
- 5.Bahadur RR. Stochastic comparison of tests. Annals of Mathematical Statistics. 1960;31(2):276–295. [Google Scholar]
- 6.Lai TL, Shih MC. Power, sample size and adaptation considerations in the design of group sequential clinical trials. Biometrika. 2004;91(3):507–528. [Google Scholar]
- 7.Bartroff J, Lai TL, Shih MC. Sequential Experimentation in Clinical Trials. Springer; New York: 2013. [Google Scholar]
- 8.Bartroff J, Lai TL. Efficient adaptive designs with mid-course sample size adjustment in clinical trials. Statistics in Medicine. 2008;27(10):1593–1611. doi: 10.1002/sim.3201. [DOI] [PubMed] [Google Scholar]
- 9.Bartroff J, Lai TL. Generalized likelihood ratio statistics and uncertainty adjustments in efficient adaptive design of clinical trials. Sequential Analysis. 2008;27(3):254–276. [Google Scholar]
- 10.Jennison C, Turnbull B. Mid-course sample size modification in clinical trials based on the observed treatment effect. Statistics in Medicine. 2003;22(6):971–993. doi: 10.1002/sim.1457. [DOI] [PubMed] [Google Scholar]
- 11.Jennison C, Turnbull B. Adaptive and nonadaptive group sequential trials. Biometrika. 2006;93(1):1–21. [Google Scholar]
- 12.He P, Lai TL, Liao OY. Futility stopping in clinical trials. Statistics and Its Interface. 2012;5(4):415–423. [Google Scholar]
- 13.Rosenblum M, Liu H, Yan EH. Optimal tests of treatment effects for the overall population and two subpopulations in randomized trials, using sparse linear programming. Journal of American Statistical Association. 2014 doi: 10.1080/01621459.2013.879063. in press. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Chernoff H, Savage R. On Chernoff-Savage statistics and sequential rank tests. Annals of the Statistics. 1975;3(4):825–845. [Google Scholar]
- 15.Simon N, Simon R. Adaptive enrichment designs for clinical trials. Biostatistics. 2013;14(4):613–625. doi: 10.1093/biostatistics/kxt010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Yeatts S, Martin R, Coffey C, Lyden P, Foster L, Woolson R, et al. Challenges of decision making regarding futility in a randomized trial: the Interventional Management of Stroke III experience. Stroke. 2014;45(5):1408–1414. doi: 10.1161/STROKEAHA.113.003925. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Bahadur RR. Rates of convergence of estimates and test statistics. Annals of Mathematical Statistics. 1967;38(2):303–324. [Google Scholar]
- 18.Jennison C, Turnbull B. Confirmatory Phase II/III clinical trials with hypotheses selection at interim: Opportunities and limitations. Biometrical Journal. 2006;48(4):650–655. doi: 10.1002/bimj.200610248. [DOI] [PubMed] [Google Scholar]
- 19.Tang D, Geller N. Closed testing procedure for group sequential trials with multiple endpoints. Biometrics. 1999;55(4):1188–1192. doi: 10.1111/j.0006-341x.1999.01188.x. [DOI] [PubMed] [Google Scholar]
- 20.Wang Y, Lan KG, Li G, Ouyang SP. A group sequential procedure for interim treatment selection. Statistics in Biopharmaceutical Research. 2011;3(1):1–13. [Google Scholar]
- 21.Lehmann E. Nonparametrics: Statistical Methods Based on Ranks. Holden-Day, Inc; San Francisco: 1975. [Google Scholar]
- 22.Robbins H, Siegmund D. Sequential tests involving two populations. Journal of the American Statistical Association. 1974;69(1):132–139. [Google Scholar]
- 23.Chernoff H, Savage R. Asymptotic normality and efficiency of certain nonparametric test statistics. Annals of the Mathematical Statistics. 1958;29(4):972–994. [Google Scholar]

