Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2011 Dec 1.
Published in final edited form as: Biometrics. 2010 Dec;66(4):1174–1184. doi: 10.1111/j.1541-0420.2010.01402.x

Testing a Primary and a Secondary Endpoint in a Group Sequential Design

Ajit C Tamhane 1, Cyrus R Mehta 2, Lingyun Liu 3
PMCID: PMC2925065  NIHMSID: NIHMS187293  PMID: 20337631

Summary

We consider a clinical trial with a primary and a secondary endpoint where the secondary endpoint is tested only if the primary endpoint is significant. The trial uses a group sequential procedure with two stages. The familywise error rate (FWER) of falsely concluding significance on either endpoint is to be controlled at a nominal level α. The Type I error rate for the primary endpoint is controlled by choosing any α-level stopping boundary, e.g., the standard O’Brien-Fleming or the Pocock boundary. Given any particular α-level boundary for the primary endpoint, we study the problem of determining the boundary for the secondary endpoint to control the FWER. We study this FWER analytically and numerically and find that it is maximized when the correlation coefficient ρ between the two endpoints equals 1. For the four combinations consisting of O’Brien-Fleming and Pocock boundaries for the primary and secondary endpoints, the critical constants required to control the FWER are computed for different values of ρ. An ad-hoc boundary is proposed for the secondary endpoint to address a practical concern that may be at issue in some applications. Numerical studies indicate that the O’Brien-Fleming boundary for the primary endpoint and the Pocock boundary for the secondary endpoint generally gives the best primary as well as secondary power performance. The Pocock boundary may be replaced by the ad-hoc boundary for the secondary endpoint with a very little loss of secondary power if the practical concern is at issue. A clinical trial example is given to illustrate the methods.

Keywords: Familywise error rate, Gatekeeping procedures, Multiple endpoints, Multiple comparisons, O’Brien-Fleming boundary, Pocock boundary, Primary power, Secondary power

1. Introduction

Since the pioneering works of Pocock (1977) and O’Brien and Fleming (1979), group sequential designs have been widely studied and are now commonly employed in clinical trials. Jennison and Turnbull (2000) is a comprehensive reference on this topic. Much of the work on group sequential designs deals with a single endpoint. Jennison and Turnbull (1993) discussed a group sequential design for bivariate endpoints (efficacy and safety). They determined two-sided decision boundaries for the test statistics computed for both endpoints so that if the new treatment is shown to be superior on both endpoints then it is accepted; if it shown to be inferior on one of the endpoints then it is rejected; otherwise sampling is continued to the next stage. We are interested in the problem of testing null hypotheses on hierarchically ordered endpoints, e.g., a primary and a secondary endpoint. For such endpoints it is generally required that statistical significance be shown on the primary endpoint before the secondary endpoint can be validly tested (O’Neill, 1997). Gatekeeping procedures (Dmitrienko and Tamhane, 2007; 2009) deal with such endpoints; however, these procedures have been studied only for nonsequential designs. The present paper studies the problem of testing a primary and a secondary endpoint in a two-stage group sequential setting with the former acting as a gatekeeper for the latter. This problem was originally studied by Hung, Wang and O’Neill (2007). While this paper was being revised, we received a manuscript by Glimm, Maurer and Bretz (2009) which addresses the same problem and reaches similar conclusions but does not give explicit analytical results concerning the FWER and the primary and secondary critical boundaries as we derive in this paper. These two works are completely independent and complementary to each other.

The outline of the paper is as follows. Section 2 describes the group sequential procedure (GSP) for the above problem. Section 3 discusses control of the FWER for this GSP. This section also gives a table of the critical constants using either the O’Brien-Fleming (OF) or the Pocock (PO) boundaries for the primary and secondary endpoints resulting in four different combinations. Two additional combinations that use the OF or the PO boundary for the primary and an ad-hoc (AH) boundary for the secondary endpoint are also studied. This ad-hoc boundary is introduced to address a practical concern that is explained in Section 3.2. The powers of these six different boundary combinations are compared in Section 4. Section 5 gives a clinical trial example. Section 6 outlines extensions for future research to deal with unknown correlation coefficient, multiple stages and multiple endpoints. The proofs of all theoretical results are included in supplementary materials under the Paper Information link at http://www.biometrics.tibs.org.

2. Group Sequential Procedure

Consider a GSP with two stages and a primary and a secondary endpoint. Let n1 and n2 be the sample sizes for the two stages. Suppose that the observations on the primary endpoint are i.i.d. N(μ1,σ12) and those on the secondary endpoint are i.i.d. N(μ2,σ22). Further suppose that the two endpoints are jointly distributed as bivariate normal with correlation coefficient ρ ≥ 0. Hung, Wang and O’Neill (2007) considered the problem of testing the null hypotheses H1: μ1 = 0 and H2: μ2 = 0 against one-sided alternatives with H2 being tested iff H1 is rejected either at the first stage or at the second stage. Thus the primary endpoint acts as a gatekeeper for the secondary endpoint. The FWER requirement is

FWER=P{RejectatleastonetrueHi(i=1,2)}α (1)

for specified α.

The standardized cumulative sample means at the two stages are used as the test statistics for both endpoints. Denote by (X1, X2) the cumulative test statistics at the two stages for the primary endpoint and (Y1, Y2) the corresponding test statistics for the secondary endpoint. Let (c1, c2) and (d1, d2) denote the corresponding stopping boundaries. We will use the following GSP.

Stage 1

Take n1 observations and compute (X1, Y1). If X1c1 continue to Stage 2. If X1 > c1, reject H1 and test H2. If Y1 > d1, reject H2; otherwise accept H2. In either case terminate the trial.

Stage 2

Take n2 observations and compute (X2, Y2). If X2c2, accept H1 and stop testing; otherwise reject H1 and test H2. If Y2 > d2, reject H2; otherwise accept H2.

In general, Xj ~ N1j, 1) and Yj ~ N2j, 1) for j = 1, 2, where

Δi1=δin1,Δi2=δin1+n2andδi=μiσi; (2)

under Hi, Δij = 0 (i, j = 1, 2). Frequently, we will denote the noncentrality parameters by Δi = Δi1 = γΔi2 (i = 1, 2), where

γ=n1n1+n2.

The correlations among (X1, X2, Y1, Y2) are given by

corr(X1,X2)=corr(Y1,Y2)=γcorr(X1,Y1)=corr(X2,Y2)=ρcorr(X1,Y2)=corr(X2,Y1)=ργ. (3)

3. Familywise Error Rate Control

3.1 Choice of the Primary Boundary

To determine the critical boundaries, (c1, c2) and (d1, d2), in order to satisfy the FWER control requirement (1), we need to consider three configurations, namely, H1H2, H12 and 1H2, where i denotes that Hi is false (i = 1, 2). If H1 is true (the first two cases) then the FWER is controlled at level α if the type I error probability for H1 is controlled at level α. This is obvious under H12 since there is no type I error for rejecting H2 in that case. It is also true under H1H2 because the event R2 = {Reject H2} is a subset of the event R1 = {Reject H1} and the probability of R1 does not depend on the truth or falsity of H2. Hence, to control PH1H2 (R1R2) = PH1(R1) ≤ α, one must choose (c1, c2) to satisfy

PH1(X1>c1)+PH1(X1c1,X2>c2)α. (4)

Any pair of constants, (c1, c2), that satisfy this inequality will be referred to as an α-level boundary. For the OF boundary, c2 = γc1, while for the PO boundary, c2 = c1.

More generally, we could use the error spending function approach of Lan and DeMets (1983) to calculate the primary boundary (c1, c2). Let α(t) be a nondecreasing function defined over t ∈ [0, 1] such that α(0) = 0 and α(1) = α; here t is called the information fraction. Let t1 = n1/(n1 + n2) and t2 = 1. Then (c1, c2) can be calculated from the following two equations:

PH1(X1>c1)=α(t1)andPH1(X1>c1)+P(X1c1,X2>c2)=α(t2)=α.

For a K-stage GSP (K ≥ 2) with equal information in each stage, i.e., ti = i/K (1 ≤ iK), an error spending function that approximates the OF boundary is given by (Lan and DeMets 1983)

α1(t)=2Φ(zα/2t), (5)

where Φ(·) is the standard normal c.d.f. and zα/2 is the (1−α/2)-quantile of the standard normal distribution. Similarly, an error spending function that approximates the PO boundary is given by (Lan and DeMets 1983)

α2(t)=αln{1+(e1)t}. (6)

These spending functions are used in the example of Section 5.

In conclusion, the FWER control problem is: Given a choice of (c1, c2) satisfying (4), determine (d1, d2) so that (1) is satisfied under H2. We will assume this configuration, which implies that Δ1 ≥ 0 and Δ2 = 0, throughout the next section.

3.2 Choice of the Secondary Boundary

To test multiple hypotheses in a group sequential setup, Tang and Geller (1999) showed that if there exists a GSP to test every intersection hypothesis at level α then application of the closure principle of Marcus, Peritz and Gabriel (1976) leads to a GSP that controls the FWER at level α. With H1 and H2 there are only three intersections: H1H2, H1 and H2. As seen above, an α-level GSP of H1H2 is the same as the GSP of H1 because of the hierarchical testing of H1 and H2. Therefore, if we use any α-level secondary boundary for testing H2 conditional on rejecting H1, then the overall FWER will be controlled. However, it is possible to use a more liberal secondary boundary by exploiting the dependence of the FWER on the correlation coefficient ρ as calculations in Table 1 will show. The calculations were made using the following expression for FWER under H2:

FWER=PH2(X1>c1,Y1>d1)+PH2(X1c1,X2>c2,Y2>d2)=c1Δ11Φ(d1+ρu1ρ2)φ(u)du+c2Δ12Φ(c1Δ11γu1γ2)Φ(d2+ρu1ρ2)φ(u)du,

where φ(·) denotes the p.d.f. of the standard normal distribution; also Δ11 = γΔ12 = Δ1. This expression is obtained by setting Δ21 = Δ22 = 0 in the secondary power expression (8).

Table 1.

Values of Critical Constant d for Secondary Boundary Using Combinations of the O’Brien-Fleming (OF)1, Pocock (PO)2 and Ad-Hoc (AH)3 Boundaries for n1 = n2 and α = 0.05

Primary Boundary Secondary Boundary ρ
0.0 0.2 0.4 0.6 0.8 1.0
OF1 OF2 1.407 (2.041)4 1.428 (1.773) 1.476 (1.478) 1.495 (1.150) 1.551 (0.767) 1.678 (0)
OF1 PO2 1.645 (∞) 1.663 (2.638) 1.686 (2.134) 1.717 (1.663) 1.760 (1.184) 1.876 (0.497)
OF1 AH2 1.645 (∞) 1.671 (2.505) 1.714 (.949) 1.786 (1.430) 1.926 (0.889) 2.515 (0.023)
PO1 PO2 1.645 (∞) 1.655 (2.679) 1.672 (2.080) 1.698 (1.528) 1.742 (0.952) 1.876 (0)
PO1 OF2 1.280 (2.169) 1.304 (1.874) 1.333 (1.553) 1.372 (1.202) 1.429 (0.809) 1.570 (0.216)
PO1 AH2 1.645 (∞) 1.656 (2.635) 1.679 (2.010) 1.720 (1.444) 1.801 (0.876) 2.092 (0.163)
1

OF Primary Boundary: c1=1.6782, c2 = 1.678, OF Secondary Boundary: d1=d2, d2 = d

2

PO Primary Boundary: c1 = c2 = 1.876, PO Secondary Boundary: d1 = d2 = d

3

AH Secondary Boundary: Tabled value is d1 =, d2 = z.05 = 1.645

4

The value in the parenthesis is the value of Δ1 that maximizes FWER.

Hung et al. (2007) investigated the choice d1 = d2 = zα (where zα is the upper α critical point of the N(0, 1) distribution) as Strategy 1 on the basis that the primary and secondary endpoints are tested in a fixed sequence, and so H2 can be tested at level α following rejection of H1 at level α. We show in Proposition 1 that this choice does not control the FWER for all ρ and Δ1.

Proposition 1

If d1 = d2 = zα then for ρ = 0, FWER is increasing in Δ1 and FWER → α as Δ1 → ∞ (and thus FWER is controlled for all Δ1). In fact, FWER → α as Δ1 → ∞ for all ρ ≥ 0. However, for ρ = 1, maxΔ1 FWER > α and thus FWER is not controlled for all ρ and Δ1.

To illustrate the results of Proposition 1, we considered the OF boundary for the primary endpoint with α = 0.05 and n1 = n2 = n which uses (c1,c2)=(1.6782,1.678). Figure 1 shows the behavior of FWER as a function of ρ and Δ1 when d1 = d2 = z.05 = 1.645. We see that the maximum FWER with respect to ρ and Δ1 is attained when ρ = 1 and Δ1 is close to 1, and this maximum is approximately 0.08. On the other hand, when ρ = 0, FWER ≤ α for all Δ1 ≥ 0. Furthermore, FWER → α as Δ1 → ∞ for all ρ ≥ 0.

Figure 1.

Figure 1

FWER as a function of ρ and Δ1 when (c1,c2)=(1.6782,1.678), n1 = n2 and d1 = d2 = z.05 = 1.645

As seen from Figure 1, the choice d1 = d2 = zα can be liberal for some Δ1 > 0 when ρ > 0.4. As an ad-hoc solution to this problem, Hung et al. (2007) suggested Strategy 3 which uses d1 = d2 = zα/2; however, this choice can be shown to be too conservative. Instead of an ad-hoc solution, we can numerically search to find the max FWER with respect to Δ1 and ρ for different choices of d1 = d2 = zα* such that max FWER = 0.05. Figure 2 shows that this is achieved when Δ1 = 0, ρ = 1 and d1 = d2 = 1.876, which corresponds to α* = 0.0303. Note that this is the PO boundary for α = 0.05. We will show this numerical result analytically in Proposition 3.

Figure 2.

Figure 2

FWER as a function of ρ and Δ1 when (c1,c2)=(1.6782,1.678), n1 = n2 and d1 = d2 = z.0303 = 1.876

In Propositions 2–4 that follow, we study three different combinations of the α-level boundaries: c1 = d1, c2 = d2 in Proposition 2, c1 > d1, c2 < d2 in Proposition 3 and c1 < d1, c2 > d2 in Proposition 4. From Tang and Geller (1999) it follows that all three boundaries control the FWER. In the first two cases we show that max FWER = α, while in the third case we show that max FWER < α. Hence it is possible to use a more liberal (d1, d2) boundary with level α′ > α in the third case.

Proposition 2

If (c1, c2) = (d1, d2) is an α-level boundary for the primary and secondary endpoints then for ρ = 1, maxΔ1 FWER = α is attained at Δ1 = 0 and for Δ1 = 0, maxρ FWER = α is attained at ρ = 1.

Note that this proposition does not show that the global maximum of FWER is attained at Δ1 = 0 and ρ = 1 which we conjecture to be true but have been unable to prove. However, plots in Figure 3 of numerically evaluated FWER as functions of Δ1 and ρ when (c1, c2) = (d1, d2) is the OF boundary for α = 0.05 show that the global maximum of the FWER with respect to ρ and Δ1 is indeed attained at ρ = 1 and Δ1 = 0, and this maximum equals 0.05. In fact, as seen from Figures 1 and 2, even when (c1, c2) ≠ (d1, d2), the global maximum of FWER with respect to ρ is attained at ρ = 1 and Δ1 > 0.

Figure 3.

Figure 3

FWER as a function of ρ and Δ1 when (c1,c2)=(d1,d2)=(1.6782,1.678) and n1 = n2

Proposition 3

If (c1, c2) and (d1, d2) are α-level boundaries for the primary and secondary endpoints such that c1 > d1 and c2 < d2 (e.g., if (c1, c2) is the OF boundary and (d1, d2) is the PO boundary) then for ρ = 1, maxΔ1 FWER = α is attained when Δ1 = Δ11 = c1d1.

Proposition 4

If (c1, c2) and (d1, d2) are α-level boundaries for the primary and secondary endpoints such that c1 < d1 and c2 > d2 (e.g., if (c1, c2) is the PO boundary and (d1, d2) is the OF boundary) then for ρ = 1, maxΔ1 FWER < α is attained when Δ1 = γΔ12 = γ(c2d2). Therefore max FWER can be increased to α by decreasing (d1, d2) to ( d1,d2) so that ( d1,d2) forms an α′-level boundary with α′ > α.

Propositions 2–4 are illustrated by calculations in Table 1 of the secondary boundary (d1, d2) that controls the FWER at level α for the two choices of the α-level primary boundary (c1, c2) for α = 0.05, n1 = n2 = n and ρ = 0.0(0.2)1.0. For the OF boundary, (c1,c2)=(1.6782,1.678) and for the PO boundary, (c1, c2) = (1.876, 1.876). The same two choices are used for the secondary stopping boundary in the following way:

OFBoundary:d1=d2,d2=d,POBoundary:d1=d2=d.

Table 1 lists the d-values for the four combinations of the primary and secondary boundaries.

In addition, we considered another boundary to address a referee’s practical concern over having the secondary boundary with a level α′ > α, as shown in Proposition 4. Note that, although the marginal error rate for the secondary endpoint exceeds α, the joint FWER for the primary and secondary endpoints is controlled at α. However, the referee thought that this subtlety may be difficult to explain to clinicians, and in fact suggested that both d1, d2 should be ≥ zα.

For the PO secondary boundary, (d1 = d2 = d), from Proposition 1 we have d = zα for ρ = 0. From Figures 13 we see that maxΔ1 FWER is an increasing function of ρ and so d is an increasing function of ρ; therefore dzα for all ρ ≥ 0, as can be checked from Table 1. Thus the OF1-PO2 and PO1-PO2 combinations satisfy the referee’s practical condition; as will be seen in the next section, the OF1-PO2 combination generally gives higher power for both the primary and the secondary endpoint, and is therefore preferred.

For the OF secondary boundary, ( d1=d2, d2 = d), it is possible to have d2 < zα and still control FWER. If this occurs then we set d2 = zα and then find the smallest d1zα such that max FWER = α is maintained. We refer to this modified boundary as an ad-hoc (AH) boundary. We computed this secondary boundary for two additional combinations with the OF and PO as the primary boundaries. Note that the we did not employ the AH boundary for the primary endpoint in our study. This gave a total of six combinations of the primary and secondary boundaries.

The following observations can be made from the numerical results in Table 1.

  1. The value of Δ1 that maximizes FWER for fixed ρ is a decreasing function of ρ.

  2. If both primary and secondary boundaries are of the same type (e.g., both OF or both PO) then we get (c1, c2) = (d1, d2) and Δ1 = 0, ρ = 1 gives the max FWER as shown in Proposition 2.

  3. If the primary boundary is OF and the secondary boundary is PO so that c1 > d1 and c2 < d2 then for ρ = 1, Δ11=c1d1=1.67821.876=0.497 gives max FWER as shown in Proposition 3. Note that in Figure 3 we found d1 = d2 = 1.876 by numerical search, but it could have been found directly by applying Proposition 3.

  4. If the primary boundary is PO and the secondary boundary is OF so that c1 < d1 and c2 > d2 then for ρ = 1, the maximizing value of Δ1 is given by γΔ12=γ(c2d2)=(1/2)(1.8761.678)=0.140. However, using (c1, c2) = (1.876, 1.876) and (d1,d2)=(1.6782,1.678) gives FWER = 0.0386 < 0.05. So, as shown in Proposition 4, (d1, d2) can be decreased until max FWER increases to 0.05. The final boundary, (d1,d2)=(1.5702,1.570), and the corresponding maximizing Δ1 = 0.216 were found numerically.

4. Power

There are two powers of interest: (i) primary power (Power1), the probability of rejecting a false H1, is just the power of GSP for a single endpoint which has been studied previously in the literature, and (ii) secondary power (Power2), the probability of rejecting a false H2. Note that Power1 ≥ Power2 for the same Δij (i, j = 1, 2). The overall power for rejecting H1 and H2 when both are false or only H2 is false is just the secondary power.

Consider an arbitrary configuration with Δi ≥ 0 (i = 1, 2). The primary power is given by

Power1=PH¯1(X1>c1)+PH¯1(X1c1,X2>c2),

where Xj ~ N1j, 1) (j = 1, 2) with corr(X1, X2) = γ. Noting that the conditional distribution of X2 given U = X1 − Δ11 = u is N12 + γu, 1 − γ2), we get

Power1=Φ(c1+Δ11)+c1Δ11Φ(c2+Δ12+γu1γ2)φ(u)du. (7)

The secondary power is given by

Power2=PH¯2(X1>c1,Y1>d1)+PH¯2(X1c1,X2>c2,Y2>d2),

where Xj ~ N1j, 1), Yj ~ N2j, 1) (j = 1, 2) have the correlation structure shown in (4). In the first term, the conditional distribution of Y1 given U = X1 − Δ11 = u is N21 + ρu, 1 − ρ2). Similarly, in the second term, the conditional distribution of (X1, Y2) given U = X2 − Δ12 = u is bivariate normal with mean vector (Δ11 + γu, Δ22 + ρu) and covariance matrix

[1ργργ1][γρ][γ,ρ]=[1γ2001ρ2],

i.e., conditional on U = X2 − Δ12 = u, X1 and Y2 are independent N11 + γu, 1 − γ2) and N22 + ρu, 1 − ρ2), respectively. Using these results we get

Power2=c1Δ11Φ(d1+Δ21+ρu1ρ2)φ(u)du+c2Δ12Φ(c1Δ11γu1γ2)Φ(d2+Δ22+ρu1ρ2)φ(u)du. (8)

We studied the primary and secondary powers for the six combinations of the boundaries for which Table 1 is given. The powers were numerically evaluated (not simulated) using (7) and (8). It can be shown that the OF boundary is uniformly (in Δ1) more powerful in terms of primary power than the PO boundary. Jennison (2009) has proved this result more generally for any two α-level boundaries, (c1, c2) and ( c1,c2): if the test statistics have the MLR distribution property then (c1, c2) is uniformly more powerful than ( c1,c2) if c1>c1,c2<c2.

The secondary power is more complicated to study since it depends on the particular combination of the boundaries used for the primary and secondary endpoints as well as the values of ρ and Δij. The plots of secondary powers as functions of Δ1 for the six combinations of the boundaries studied in Table 1 for α = 0.05, n1 = n2, ρ = 0.4 and Δ2 = 1, 2, 3 are given in Figures 46; the patterns of the power functions are similar for other values of ρ and hence are not shown here. These figures show that generally the combination of the OF boundary for the primary endpoint and the PO boundary for the secondary endpoint (the OF1-PO2 combination) gives the highest secondary power. With the OF1-PO2 combination, we have d1, d2 ≥ 1.645, so the practical concern raised above is not an issue. The OF1-AH2 has secondary power very similar to that of the OF1-PO2 combination, because their critical boundaries are similar. The reason that the secondary powers of the OF1-PO2 and OF1-AH2 combinations are close is because their critical boundaries are close. For ρ = 0.4, the OF1-PO2 boundaries are ( c1=1.6782, c2 = 1.678) and (d1 = 1.686, d2 = 1.686) while the OF1-AH2 boundaries are ( c1=1.6782, c2 = 1.678) and (d1 = 1.714, d2 = 1.645).

Figure 4.

Figure 4

Secondary powers of the six boundary combinations (OF = O’Brien-Fleming, PO = Pocock, AH = Ad-Hoc, 1 = Primary, 2 = Secondary) for the primary and secondary endpoints as functions of Δ1 for ρ = 0.4 and Δ2 = 1 when α = 0.05 and n1 = n2

Figure 6.

Figure 6

Secondary powers of the six boundary combinations (OF = O’Brien-Fleming, PO = Pocock, AH = Ad-Hoc, 1 = Primary, 2 = Secondary) for the primary and secondary endpoints as functions of Δ1 for ρ = 0.4 and Δ2 = 3 when α = 0.05 and n1 = n2

The power curves show that, for large Δ1, any combination using OF2 is markedly worse than those using either PO2 or AH2. This is explained by the fact that when Δ1 is large, there is a high chance of rejecting H1 at the first look; so OF2 presents a higher threshold for the secondary endpoint to exceed at that look. Also note that for large Δ1, the secondary power of OF1-OF2 is lower than that of PO1-OF2, which in turn is uniformly dominated by OF1-PO2 for all Δ1 as well as the three values of Δ2 that were studied. Thus PO1-OF2 is not a contender. OF1-OF2 combination has a slight edge for small Δ1. However, if OF1-OF2 is disqualified based on the referee’s practical concern then OF1-AH2 may be preferred for small Δ1 as well.

In looking at these plots, one may be intrigued by the fact that the secondary power of each combination of boundaries decreases with Δ1 for large Δ11 ≥ 2) for fixed Δ2. The reason is that the probability of rejecting H1 at the first stage increases as Δ1 increases. Therefore, for large Δ1, H2 gets only one chance to be tested at Stage 1 and so the secondary power decreases as Δ1 increases. approaching P (Y1 > d1) as Δ1 → ∞. On the other hand, for small Δ1, H2 gets two chances to be tested; therefore the secondary power increases as Δ1 increases.

5. Example

To illustrate the methods discussed in this paper we consider the CAPTURE study (The Capture investigators, 1997) in which a randomized trial was conducted to compare abciximab with placebo for coronary intervention in refractory unstable angina. It was planned to enroll 1400 patients to placebo and abciximab in equal proportions. The primary endpoint was a composite of death, myocardial infarction or urgent intervention for treatment of recurrent ischemia, within 30 days of randomization. From previous studies in this patient population the placebo event rate was estimated to be 15%. The investigators wished to provide good power to detect a 5% reduction in the event rate for the abciximab arm. With a sample size of 1400, a single-look design has 81% power to detect this treatment effect, using a one-sided test at the α = 0.025 level of significance. The trial was, however, designed for a group sequential test with three equally spaced looks and OF-type stopping boundaries derived from the error spending function (5). Due to the possibility of early stopping, there is a power loss of about 1% with the group sequential design relative to the corresponding fixed sample design. For illustrative purposes we will ignore the possibility of early stopping at the first look and consider this to be a two-look design. This is a reasonable approximation since the above spending function yields an extremely conservative stopping boundary at the first look with a negligible chance of early stopping even under the alternative hypothesis of a 5% drop in the event rate.

The main secondary endpoint for this trial was death or myocardial infarction within 30 days of randomization. Although in the actual trial there was no formal testing strategy to protect the FWER of the two endpoints, we will consider two such strategies here. Under Strategy 1, we prespecify in the protocol that the secondary endpoint will be tested only if the null hypothesis for the primary endpoint is rejected; moreover, it will be tested at the same look and with the same spending function as the primary endpoint. This may be considered a conservative strategy since the secondary endpoint will face the same conservative OF-type stopping boundary as the primary endpoint. Suppose, however, that abciximab is expected to reduce the risk for the secondary endpoint to a lesser degree than that for the primary endpoint. It might then be advantageous to consider a more aggressive Strategy 2 in which the Type I error for the secondary endpoint is spent at a faster rate, resulting in a less demanding early stopping boundary at the first look. For example, one might prespecify the use of PO-type stopping boundaries derived from the error spending function (6) for the secondary endpoint.

Let π1e and π1c denote the true event rates for the experimental (abciximab) and control (placebo) arms, respectively, with respect to the primary endpoint. Let π2e and π2c be the corresponding true event rates for the secondary endpoint. Define μi = πicπie, i = 1, 2. Let n1 and n2 be the sample sizes per arm at the two stages and N = 2(n1 + n2) be the planned maximum total sample size. The Wald statistics for the primary and secondary endpoints at the interim look are given by

X1=(π^1c(1)π^1e(1))σ^1(1)n12andY1=(π^2c(1)π^2e(1))σ^2(1)n12

where ( π^ic(1),π^ie(1)) denotes the maximum likelihood estimate of (πic, πie) at the interim look and

σ^i(1)=π^ic(1)(1π^ic(1))+π^ie(1)(1π^ie(1))(i=1,2).

The Wald statistics for the primary and secondary endpoint at the final look are given by

X2=(π^1c(2)π^1e(2))σ^1(2)n1+n22andY2=(π^2c(2)π^2e(2))σ^2(2)n1+n22

where ( π^ic(2),π^ie(2)) denotes the maximum likelihood estimate of (πic, πie) at the final look and

σ^i(2)=π^ic(2)(1π^ic(2))+π^ie(2)(1π^ie(2))(i=1,2).

Applying the large sample results of Jennison and Turnbull (2000, Chapter 3), the above sequentially computed Wald statistics have independent increments, so that (X1, X2, Y1, Y2) have the same distributional properties as the corresponding random variables defined in Section 2. We may thus avail of the results presented in Section 3 for controlling the FWER. No data are available to estimate the correlation coefficient ρ, so we will assume the least favorable configuration ρ = 1.

In the CAPTURE study an interim analysis was performed after data were available on 1050 subjects out of a planned 1400 subjects. Thus the relevant information fractions are t1 = 1050/1400 = 0.75, and t2 = 1400/1400 = 1. We compute the stopping boundary (c1; c2) from the spending function (5) by solving

PH1(X1>c1)=α1(t1)andPH1(X1>c1)+PH1(X1c1,X2>c2)=α.

For Strategy 1 we set (c1, c2) = (d1, d2). For α = 0.025, we obtain (c1, c2) = (2.340, 2.012). (The classical OF boundary would be (c1, c2) = (2.327, 2.015)).

For Strategy 2 we compute the stopping boundary (d1, d2) from the spending function (6) by solving

PH2(Y1d1)=α2(t1)andPH2(Y1>d1)+PH2(Y1d1,Y2>d2)=α.

For α = 0.025, we obtain (d1, d2) = (2.040, 2.258) with the same (c1, c2) as above. (The classical PO boundary would be (d1, d2) = (2.126, 2.126)).

The following event rates (presented by K. Anderson at the 2002 Conference on Clinical Trial Data Monitoring Committees, Barnett International) were obtained for the primary endpoint at the time of the interim analysis: 84/532 (15.8%) for the placebo arm and 55/518 (10.6%) for the abciximab arm. Although no data were available for the secondary endpoint, we extrapolated from the published results of Simoons et al. (1997) to give 44/532 (8.3%) for the placebo arm and 26/518 (5%) for the abciximab arm. The test statistics are displayed in Table 2 and the stopping boundaries are displayed in Table 3.

Table 2.

Interim Results and Test Statistics

Endpoint Event Rates Interim Statistic
Placebo Abciximab
Primary 84/532 (15.8%) 55/518 (10.6%) X1 = 2.485
Secondary 44/532 (8.3%) 26/518 (5.0%) Y1 = 2.123

Table 3.

Critical Boundaries for Two Strategies

Look Strategy 1 Strategy 2
Primary: αOF(t) Secondary: αOF(t) Primary: αOF(t) Secondary: αPO(t)
Interim c1 = 2.340 d1 = 2.340 c1 = 2.340 d1 = 2.040
Final c2 = 2.012 d2 = 2.012 c2 = 2.012 d2 = 2.258

The trial was terminated at the interim look since the test statistic X1 for the primary endpoint crossed its early stopping boundary c1. In order to formally claim statistical significance in the product label for the secondary endpoint, the test statistic Y1 would have to cross its early stopping boundary d1. It is seen that the secondary statistic does not cross its early stopping boundary under Strategy 1, but it does so under Strategy 2. Thus, depending on what was prespecified in the protocol, the investigators might or might not be allowed to claim statistical significance for the secondary endpoint.

6. Extensions

The scope of the present work is somewhat limited. First, we have considered only one-sided tests on two endpoints with two stages. Second, the data on the primary and secondary endpoints are assumed to be bivariate normal with known standard deviations and a nonnegative correlation coefficient that is either assumed to be known or treated as a nuisance parameter with the FWER being maximized with respect to it to determine the critical boundaries for the endpoints. These limitations are in part dictated by the analytical and computational difficulties involved. Given the practical importance of this problem, we expect that the present work will serve as a springboard for further extensions some of which are outlined below.

6.1 Unknown Correlation

The plots of FWER in Figures 46 and the critical constants tabulated in Table 1 show that maxΔ1 FWER is an increasing function of ρ with ρ = 1 being the least favorable case. If an upper bound on ρ can be specified then it can be used to find a more powerful boundary for the secondary endpoint than the conservative boundary corresponding to ρ = 1. A practical approach to this problem would be to use the upper limit of a one-sided confidence interval on ρ estimated from the data after Stage 1 to find the Stage 2 boundary with a suitable adjustment to the required FWER to take into account the error probability of the true ρ exceeding the upper confidence limit. The ideas from Berger and Boos (1994) can be used in this context as follows.

Let ρ* be a (1 − ε)-level upper confidence limit on ρ (e.g., calculated using Fisher’s large sample arctan hyperbolic transformation), i.e., P(ρρ*) = 1 − ε. Then, using the property that maxΔ1 FWER(Δ1, ρ) is an increasing function of ρ, the overall maximum FWER can be written as follows. Let Δ1(ρ) be the value of Δ1 that maximizes FWER(Δ1, ρ) for fixed ρ. Then

max{Δ1,ρ}FWER(Δ1,ρ)=max{ρρ}FWER(Δ1(ρ),ρ)×P(ρρ)+max{ρ>ρ}FWER(Δ1(ρ),ρ)×P(ρ>ρ)=FWER(Δ1(ρ),ρ)×(1ε)+FWER(Δ1(1),1)×ε.

We want to determine the smallest possible (d1, d2) (so as to maximize the power) subject to the above expression being ≤ α. This problem can be solved iteratively on a computer. For illustration, suppose we choose the PO boundary for the secondary endpoint, i.e., d1 = d2 = d. For fixed d we can evaluate FWER(Δ1(ρ),ρ)=α(d) (say) and FWER(Δ1(1),1)=α(d) (say) where note that α′(d) < α″(d). Then by numerical search we can find the smallest d such that α′(d)(1 − ε) + α″(d)ε = α. In this case also one may want to restrict d1, d2zα for practical reasons.

6.2 Multiple Stages

In this paper we have restricted to a GSP with two stages. However, most practical applications involve more than two stages. Restricting still to two endpoints but K ≥ 2 stages, denote the cumulative standardized test statistics for the primary and secondary endpoints at the kth stage by Xk and Yk, respectively. The FWER control problem is to determine the critical boundaries (c1, c2, …, cK) and (d1, d2, …, dK) such that the FWER is strongly controlled at α for the decision rule that continues sampling as long as Xkck for kK. If Xk > ck then H1 is rejected and H2 is tested. H2 is rejected if Yk > dk and accepted if Ykdk; sampling terminates in either case or if k = K. If (c1, c2, …, cK) and (d1, d2, …, dK) are α-level boundaries then from Tang and Geller (1999) it follows that the FWER will be controlled at level α. However, a more detailed analysis can be carried out to sharpen the secondary boundary as a function of ρ as was done for K = 2 stages. It would be useful to develop a software that would compute these boundaries for any choice of the primary boundary, e.g., from the class of boundaries based on error spending functions of Lan and DeMets (1983) and Wang and Tsiatis (1987).

6.3 Multiple Endpoints

Multiple endpoints are of less practical interest than multiple stages, but it may be of interest to study three ordered endpoints (primary, secondary and tertiary). Denote the test statistics for the three endpoints at the kth stage by (Xk, Yk, Zk) and let (c1, c2, …, cK), (d1, d2, …, dK) and (e1, e2, …, eK) be the critical boundaries for the three endpoints. Let H1, H2 and H3 be the null hypotheses for the three endpoints. One may consider the following decision rule: Continue sampling as long as Xkck for kK. If Xk > ck then reject H1 and test H2. If Yk > dk then reject H2; otherwise terminate sampling without testing H3. If H2 is rejected then test H3; reject H3 if Zk > ek; otherwise accept H3. Terminate sampling in either case or if k = K. It is easy to show that if all three boundaries are α-level then the FWER will be controlled. A challenging problem is how the knowledge (or estimates) of the correlation coefficients between the three endpoints can be used to sharpen the boundaries for the secondary and tertiary endpoints. This problem is complicated by the fact that there are three correlation coefficients.

Finally we mention another extension for K = 2 in which one may continue sampling to the second stage and test H2 if H1 is rejected but H2 is not rejected at the first stage, effectively offering a second chance for the secondary endpoint to be tested.

Supplementary Material

Appendix

Figure 5.

Figure 5

Secondary powers of the six boundary combinations (OF = O’Brien-Fleming, PO = Pocock, AH = Ad-Hoc, 1 = Primary, 2 = Secondary) for the primary and secondary endpoints as functions of Δ1 for ρ = 0.4 and Δ2 = 2 when α = 0.05 and n1 = n2

Acknowledgments

The authors are grateful to a co-editor and two referees whose comments helped to improve the paper. Prof. Ajit Tamhane’s research was supported by National Heart, Lung and Blood Institute Award 1 RO1 HLO 82725-01A1.

Footnotes

Supplementary Materials

Web Appendix referenced in Section 1 which gives the proofs of all propositions is available under the Paper Information link at the Biometrics website http://www.biometrics.tibs.org.

Contributor Information

Ajit C. Tamhane, Email: atamhane@northwestern.edu, Department of Industrial Engineering and Management Sciences, Northwestern University, Evanston, IL 60208, U.S.A

Cyrus R. Mehta, Email: mehta@cytel.com, Cytel Corporation and Department of Biostatistics, Harvard School of Public Health, Cambridge, MA 02139, U.S.A

Lingyun Liu, Email: lingyun.liu@cytel.com, Cytel Corporation, Cambridge, MA 02139, U.S.A.

References

  1. Berger RL, Boos DD. P-value maximized over a confidence set for the nuisance parameter. Journal of the American Statistical Association. 1994;89:1012–1016. [Google Scholar]
  2. Dmitrienko A, Offen WW, Westfall PH. Gatekeeping strategies for clinical trials that do not require all primary effects to be significant. Statistics in Medicine. 2003;22:2387–2400. doi: 10.1002/sim.1526. [DOI] [PubMed] [Google Scholar]
  3. Dmitrienko A, Tamhane AC. Gatekeeping procedures with clinical trial applications. Journal of Pharmaceutical Statistics. 2007;6:171–180. doi: 10.1002/pst.291. [DOI] [PubMed] [Google Scholar]
  4. Dmitrienko A, Tamhane AC. Gatekeeping procedures in clinical trials. In: Dmitrienko A, Tamhane AC, Bretz F, editors. Multiple Testing Problems in Pharmaceutical Statistics. Chapter 5. Boca Raton, FL: Taylor & Francis; 2009. [Google Scholar]
  5. Glimm E, Maurer W, Bretz F. Hierarchical testing of multiple endpoints in group-sequential trials. Statistics in Medicine, to appear. 2009 doi: 10.1002/sim.3748. [DOI] [PubMed] [Google Scholar]
  6. Hung HMJ, Wang S-J, O’Neill R. Statistical considerations for testing multiple endpoints in group sequential or adaptive clinical trials. Journal of Biopharmaceutical Statistics. 2007;17:1201–1210. doi: 10.1080/10543400701645405. [DOI] [PubMed] [Google Scholar]
  7. Jennison C, Turnbull BW. Group sequential tests for bivariate response: Interim analyses of clinical trials with efficacy and safety endpoints. Biometrics. 1993;49:741–752. [PubMed] [Google Scholar]
  8. Jennison C, Turnbull BW. Group Sequential Methods with Applications to Clinical Trials. Boca Raton, FL: Chapman & Hall/CRC Press; 2000. [Google Scholar]
  9. Jennison C. Personal communication. 2009.
  10. Lan KKG, DeMets DL. Discrete sequential boundaries for clinical trials. Biometrika. 1983;70:659–663. [Google Scholar]
  11. Marcus R, Peritz E, Gabriel KR. On closed testing procedures with special reference to ordered analysis of variance. Biometrika. 1976;63:655–660. [Google Scholar]
  12. O’Brien PC, Fleming TR. A multiple testing procedure for clinical trials. Biometrics. 1979;35:549–556. [PubMed] [Google Scholar]
  13. O’Neill RT. Secondary endpoints cannot be validly analyzed if the primary endpoint does not demonstrate clear statistical significance. Controlled Clinical Trials. 1997;18:550–556. doi: 10.1016/s0197-2456(97)00075-5. [DOI] [PubMed] [Google Scholar]
  14. Pocock SJ. Group sequential methods in the design and analysis of clinical trials. Biometrika. 1977;64:191–199. [Google Scholar]
  15. Tang DI, Geller NL. Closed testing procedures for group sequential clinical trials with multiple endpoints. Biometrics. 1999;55:1188–1192. doi: 10.1111/j.0006-341x.1999.01188.x. [DOI] [PubMed] [Google Scholar]
  16. The Capture investigators. Randomized placebo-controlled trial for abciximab before and during coronary intervention in refractory unstable angina; the CAPTURE study. Lancet. 1997;349:1429–1435. [PubMed] [Google Scholar]
  17. Wang SK, Tsiatis AA. Approximately optimal one-parameter boundaries for group sequential trials. Biometrics. 1987;43:193–200. [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix

RESOURCES