Abstract
A crossover trial is an efficient trial design when there is no carry-over effect. To reduce the impact of the biological carry-over effect, a washout period is often designed. However, the carry-over effect remains an outstanding concern when a washout period is unethical or cannot sufficiently diminish the impact of the carry-over effect. The latter can occur in comparative effectiveness research, where the carry-over effect is often non-biological but behavioral. In this paper, we investigate the crossover design under a potential outcomes framework with and without the carry-over effect. We find that when the carry-over effect exists and satisfies a sign condition, the basic estimator underestimates the treatment effect, which does not inflate the type I error of one-sided tests but negatively impacts the power. This leads to a power trade-off between the crossover design and the parallel-group design, and we derive the condition under which the crossover design does not lead to type I error inflation and is still more powerful than the parallel-group design. We also develop covariate adjustment methods for crossover trials. We evaluate the performance of cross-over design and covariate adjustment using data from the MTN-034/REACH study.
Keywords: causal inference, comparative effectiveness research, covariate adjustment, efficiency
1. INTRODUCTION
A crossover trial is a longitudinal study in which patients are randomized to different sequences of treatments with the intention of comparing the effects of different treatments (Hills and Armitage, 1979; Senn, 1994; Kim et al., 2021, Chapter 8). In the simplest crossover design, patients are randomly allocated to 2 groups: one group receives treatment 1 at period 1 and treatment 0 at period 2, and the other group receives the treatment sequence in the reverse order. The main advantage of a crossover design is that every patient can serve as their own control, which largely eliminates the between-patient variability and makes the crossover design more efficient than the parallel-group design.
A key assumption for the crossover design is that there is no carry-over effect, that is, the treatment assigned during the first period does not interfere with the outcome in the second period. Therefore, a crossover design is most common for treatments whose effect vanishes when discontinued and for non-absorbing endpoints. Examples include early phase trials (eg, pharmacokinetic [PK] studies and dose-finding studies) and phase III studies with chronic conditions (eg, hypertension, pain, and asthma); see Jones and Lewis (1995) for a review. Meanwhile, a washout period is often designed between periods 1 and 2 to effectively reduce the impact of the biological carry-over effect engendered by treatments taken at period 1. The washout can be passive (ie, no treatment is given during the washout period) or active (ie, a treatment is given during the washout period but measurement is delayed until steady state is reached) (Senn, 2002; Araujo et al., 2016). However, the carry-over effect remains an outstanding concern when a washout period is inappropriate or cannot sufficiently diminish the impact of the first period. For example, the MTN-034/REACH study (Nair et al., 2023) is a phase 2a HIV-prevention trial evaluating the safety of and the adherence to the monthly dapivirine vaginal ring (DVR) and daily oral tenofovir disoproxil fumarate plus emtricitabine (TDF/FTC) among adolescent girls and young women aged 16–21 years. This trial uses a 2-period crossover design without a passive washout period since withholding an effective HIV preventive agent from the population at risk is unethical. If participants initially using the monthly DVR (which requires no active effort to remain adherent for a month after inserting the ring) find the transition to daily pill-taking in period 2 particularly burdensome, then the ease and habit formed during the first period with the monthly DVR might carry over into period 2 and affect their adherence behavior in period 2. This type of carry-over effect is non-biological and is hard to be eliminated by washout periods; we term it as the behavioral carry-over effect, which is the impact that a treatment has on the subsequent outcomes due to the altering of a participant’s behavior. It can be an important consideration in other HIV prevention trials, for example, the TRIO study (Minnis et al., 2018), and more broadly in comparative effectiveness research that uses the crossover design (Hemming et al., 2020).
The presence of the carry-over effect can bias the estimation of the treatment effect. This motivates Grizzle’s (1965) 2-stage procedure, which first tests whether the carry-over effect exists or not, and the testing result determines whether to use the period 2 data in the analysis. This 2-stage procedure was then criticized by Freeman (1989) because it has low power and can inflate the type I error of subsequent analysis. Another type of method is to model the carry-over effect (Brown, 1980; Laird et al., 1992; Jones and Donev, 1996; Kunert and Stufken, 2002; Bailey and Kunert, 2006), which is sensitive to model misspecification.
In this paper, we examine the carry-over effect in crossover trials using a potential outcomes framework (Neyman, 1923; Rubin, 1974) and only require minimal statistical assumptions to gain a better understanding of its impact. Under the ideal setting with no carry-over effect, we derive the asymptotic properties of the basic estimator of the treatment effect, based on which we show that the crossover design is typically more efficient than the parallel-group design. Our findings, while aligning with existing efficiency comparisons in the literature (refer to Senn 2002, Section 9.2.1, eg), are more general, because we only assume the no carry-over effect assumption and no other assumptions (eg, modeling assumptions or normality of data). Then we investigate the situation with the carry-over effect, uncovering new insights on its influence on estimation bias, as well as the type I error and power of tests. When the carry-over effect λ1 + λ0 is negative, the basic estimator overestimates the treatment effect, which can inflate the type I error of one-sided tests. Conversely, when the carry-over effect is positive, the basic estimator underestimates the treatment effect, which does not inflate the type I error of one-sided tests but negatively impacts the power. This leads to a power trade-off between the crossover design and the parallel-group design, and we derive the condition under which the crossover design does not lead to type I error inflation and is still more powerful than the parallel-group design.
The rest of the paper proceeds as follows. Section 2 focuses on crossover trials with no carry-over effect, where we present a potential outcomes framework for crossover trials, investigate the basic estimator, and compare the efficiencies of a crossover and a parallel-group design. Section 3 extends the results to the case with the carry-over effect. Section 4 proposes a covariate adjustment method and studies its asymptotic properties. Section 5 examines the power trade-off of the crossover and the parallel-group design. Section 6 presents the application of our methods to the REACH study, as well as power considerations in the design of a future HIV prevention trial using data from the REACH study. Section 7 concludes with a discussion.
2. CROSSOVER TRIALS WITH NO CARRY-OVER EFFECT
2.1. Setup and assumptions
Consider a crossover trial for 2 treatments in 2 periods. A sample of n subjects are randomly allocated to 2 treatment sequences, where Ai = 1 denotes that subject i first receives treatment 1 and then treatment 0, and Ai = 0 denotes the reverse order. Let
be the potential outcome at time 1 had the subject been exposed to treatment j at time 1, for j = 0, 1. Let
be the potential outcome at time 2 had the subject been exposed to treatment j at time 1 and treatment k at time 2, for j, k = 0, 1. The observed outcome for subject i at time t is Yit. Throughout the article, we make the consistency assumption that links the observed outcomes to the potential outcomes: for
and
for
and
. Let
be a vector of observed baseline covariates for subject i. We assume that
are independent and identically distributed with finite second-order moments, and the covariance matrix
is positive definite.
Simple randomization assigns subjects to the 2 treatment sequences completely at random. This is summarized in Assumption 1.
Assumption 1 (Simple randomization)
for j, k = 0, 1, and P(Ai = 1) = π1, where 0 < π1 < 1 is known and π0 = 1 − π1.
Assumption 2 is the key assumption typically imposed in crossover trials: It states that the treatment at time 1 does not have a direct effect on the outcome at time 2; see Figure 1(A) for an illustration. Many crossover trials would plan a sufficiently long washout period between the 2 time periods to make this assumption more plausible.
FIGURE 1.
(A) Six potential outcomes under Assumption 2. (B) Six potential outcomes without Assumption 2.
Assumption 2 (No carry-over effect)
For k = 0, 1,
almost surely.
Under Assumption 2, we can simply let
denote the potential outcome at time 2. We are interested in the average treatment effects 2−1(θ1 + θ2), where
and
. Note that it is common to assume that the treatment effect is time-invariant so that θ1 = θ2 (Hills and Armitage, 1979). However, as shown in Theorem 1, a time-invariant treatment effect assumption is not necessary for studying treatment effect in crossover trials.
2.2. The basic estimator
Using the full crossover period, we can calculate the difference in outcomes for every subject under treatment 1 and treatment 0. The basic estimator is commonly used, which first calculates the arm-specific average outcome difference, and then averages the 2 means from the 2 arms (Hills and Armitage, 1979; Senn, 1994; Kim et al., 2021, Chapter 8):
![]() |
(1) |
where
is the average of Δi’s for subjects with Ai = a, Δi = Yi1 − Yi2, and na is the number of subjects with Ai = a, for a = 0, 1. Theorem 1 summarizes the statistical properties of the basic estimator
.
Theorem 1
(a)
, where
for t = 1, 2.
(b)
, where
.
Proof of Theorem 1 and all other proofs are in the supplementary materials. In the proof of Theorem 1(a), we show that the expected change in outcome for Ai = 1 is the average treatment effect at time 1 minus the expected change in outcome in the absence of the treatment, that is,
, where
denotes the expected temporal change in outcome in the absence of the treatment. Similarly, the expected change in outcome for Ai = 0 is the average treatment effect plus the expected change in outcome in the absence of the treatment, that is,
. At first glance, it appears that both the group-specific average changes in outcome are biased by τ. However, randomization balances out the effect of temporal trend that is not due to the treatment, and thus the overall average change in outcome remains unbiased for 2−1(θ1 + θ2). Moreover, with a given sample size n, Theorem 1(b) guides the optimal choice of π1 to minimize the asymptotic variance. For example, if
, the optimal choice would be equal allocation with π1 = π0 = 1/2. For statistical inference,
can be consistently estimated by
![]() |
(2) |
where
and
are, respectively, the sample variance of Δi = Yi1 − Yi2 for subjects under Ai = 1 and Ai = 0.
There is another estimator that looks similar to the basic crossover estimator
:
. These 2 estimators
and
are the same under randomization schemes that enforce n1 = n0, but they are not the same in other cases (eg, under the simple randomization considered in this article). Specifically, under simple randomization with equal allocation (π1 = π0 = 1/2),
is also unbiased for 2−1(θ1 + θ2); however, the asymptotic variance of
equals
, which is larger than the asymptotic variance of
. The additional variance component 4−1(θ1 − θ2 − 2τ)2 is due to the random effect of the temporal trend and the treatment effect heterogeneity at the 2 time points. Under simple randomization with unequal allocation,
is biased due to the time effect τ. This important point is also discussed in Hills and Armitage (1979). Therefore, we do not consider this alternative estimator
in the rest of the article.
2.3. Efficiency comparison between a crossover and parallel-group design
When there is no carry-over effect, that is, when Assumption 2 holds, it is well known that a crossover design is typically more efficient than a parallel-group one. This is because each subject can serve as their own control, which cuts down half the sample size, and the within-subject comparison can further remove the inter-subject variability (Jones and Kenward, 2015). Prior work such as Senn (2002, Section 9.2.1) has provided efficiency comparisons between a crossover and parallel-group design when there is no carry-over effect and under various additional assumptions (eg, modeling assumptions, time-invariant treatment effect, and normality of data). In this section, we provide a more general efficiency comparison that assumes only Assumption 2 and no other assumptions.
Suppose a randomized controlled trial is being planned to demonstrate the superiority or non-inferiority of an investigational treatment. First, consider a crossover design with the null hypothesis H0: 2−1(θ1 + θ2) = θ* versus the alternative hypothesis HA: 2−1(θ1 + θ2) > θ* for some pre-specified θ*, where θ* = 0 for test of superiority and θ* > 0 for test of non-inferiority. The test statistic based on the basic estimator is
. From Theorem 1,
under H0, and thus, we reject H0 if and only if Tcr > z1 − α, where α is the significance level and z1 − α is the (1 − α)th quantile of the standard normal distribution. Under the local alternative
with a constant γcr > 0, the power of Tcr is Powercr ≈ Φ( − z1 − α + γcr/σcr), where Φ( · ) is the cumulative distribution function of the standard normal distribution and ≈ denotes asymptotic approximation.
Now suppose we only use the data at time 1, which is effectively a parallel-group design. The parallel-group counterpart to
is a simple mean difference
. Under Assumption 1, we have
and
, where
. For testing the null hypothesis H0: θ1 = θ* versus the alternative hypothesis HA: θ1 > θ*, the test statistic is
, where
is a consistent estimator for
. Under the local alternative
with a constant γpr > 0, the power of Tpr is
![]() |
(3) |
Let Powercr = Powerpr = 1 − β, we obtain the required sample sizes to achieve power of 1 − β under the 2 designs:
![]() |
These sample size formulae have been previously derived, for example, in Jones and Kenward (2015, Section 2.4) under additional assumptions such as the time-invariant treatment effect. When θ1 = θ2, the 2 aforementioned null hypotheses are the same, and the ratio of the sample sizes to achieve the same power using the 2 tests, which is also called the Pitman asymptotic relative efficiency, is
![]() |
For illustration, consider a simple case where θ1 = θ2,
for t = 1, 2 and j = 0, 1, and
, where ρ ∈ [0, 1) is the intraclass correlation coefficient (ICC). Then
and
, and ncr/npr = (1 − ρ)/2. This implies that the crossover design only requires (1 − ρ)/2, that is, at most half of the sample size required by the parallel-group design, which is the key advantage of the crossover design over the parallel-group design.
3. CROSSOVER TRIALS WITH CARRY-OVER EFFECT
3.1. Setup and assumptions
When there exists the carry-over effect in a crossover trial, the treatment at time 1 may interfere with the outcome at time 2; see Figure 1(B) for an illustration of the 6 potential outcomes with carry-over effects. In this case, Assumption 1 still holds by the act of randomization, while Assumption 2 is violated.
We present a way of parameterizing the expected values of the 6 potential outcomes in Table 1 that directly extends the parameterization used in Section 2. In Table 1,
denotes the average treatment effect at time 2, which compares the potential outcome had one stayed on treatment 1 to the potential outcome had one stayed on treatment 0;
denotes the expected temporal change in outcome had one stayed on treatment 0;
and
are the expected carry-over effects under the 2 treatment regimes.
TABLE 1.
Parameterization of the 6 potential outcome means with and without Assumption 2.
| Potential outcome means | With Assumption 2 | Without Assumption 2 |
|---|---|---|
|
μ | μ |
|
μ + θ1 | μ + θ1 |
|
μ + τ |
|
|
μ + τ |
|
|
μ + τ + θ2 |
|
|
μ + τ + θ2 |
|
The 2 shadowed rows correspond to the potential outcomes that are never observed.
3.2. The basic estimator
Theorem 2 derives the statistical properties of the basic estimator
defined in Equation 1 without Assumption 2.
Theorem 2
Under Assumption 1,
(a)
.
(b)
, where
.
Similar to the proof of Theorem 1(a), we show that the expected change in outcome for Ai = 1 is
, that for Ai = 0 is
. Hence, the overall average change in outcome is
. Hence, Theorem 2(a) implies that
is still an unbiased estimator of the treatment effect
under a population-level no carry-over effect assumption (
for k = 0, 1), which is weaker than the individual-level no carry-over effect assumption stated in Assumption 2. It is also straightforward to verify that Theorem 1 is a special case of Theorem 2 under Assumption 2. Lastly, note that
defined in Equation 2 is still a consistent estimator of
without Assumption 2.
3.3. Type I error and power analysis
With possible carry-over effects, the first question one would ask is whether using Tcr to analyze crossover trials would lead to type I error inflation. To answer this question, consider the null hypothesis
versus the alternative hypothesis
for some pre-specified θ*. The type I error rate of
is
![]() |
When λ0 + λ1 = 0,
is an unbiased estimator for the treatment effect
and the type I error rate of Tcr is α. When λ0 + λ1 > 0,
underestimates the treatment effect
and the type I error rate is less than α, meaning that the test is conservative. When λ0 + λ1 < 0,
overestimates the treatment effect
and the type I error rate is larger than α, meaning that the test is invalid. Therefore, even with the carry-over effect, the crossover design can still control the type I error rate of a one-sided test when λ0 + λ1 ≥ 0.
However, a carry-over effect that does not inflate the type I error rate can have a negative impact on the power. Specifically, with possible carry-over effects and under the local alternative
with a constant γcr > 0, the power of Tcr is
![]() |
(4) |
Senn (1997, Table 1) briefly notes the impact of carry-over on power but does not provide a power formula. Our Equation 4 provides a general power formula in the presence of carry-over effects, which produces results that are the same as those in Senn (1997, Table 1) (with n = 44,
). Equation 4 also implies that the required sample size to achieve 1 − β power is
![]() |
For illustration, consider a test of superiority and suppose that
and the Pitman asymptotic relative efficiency between the crossover and the parallel-group design is
![]() |
When
, that is, the carry-over effects are non-negative for the purpose of controlling type I error rate and are smaller than the treatment effects, the difference between npr and
equals
![]() |
where c is some positive constant. Therefore, in order for the crossover design to have higher efficiency than the parallel-group one, we need the carry-over effect to be small. Specifically, similar to our discussion at the end of Section 2: When
, we need
. When ρ = 0.3, 0.5, 0.7, we need to have 2−1(λ0 + λ1) > 0 and less than 0.41θAlt, 0.5θAlt, and 0.61θAlt, respectively, so that the type I error is not inflated and the crossover design is more powerful than the parallel-group one. This small carry-over effect condition may be plausible in many scenarios because carry-over effects are usually relatively small compared to the treatment effects. Therefore, a crossover design can still be more powerful than a parallel-group design in many cases even when the carry-over effect exists. In Section S2 of the supplementary materials, we discuss how to control the type I error when λ0 + λ1 < 0 using a sensitivity analysis approach (Rosenbaum, 2020).
4. COVARIATE ADJUSTMENT FOR CROSSOVER TRIALS
Adjusting for prognostic baseline covariates in the analysis of randomized controlled trials is encouraged by regulatory agencies because it has high potential to improve efficiency under approximately the same minimal statistical assumptions that would be needed for unadjusted estimation (FDA, 2023). It often uses a working model between the outcomes and covariates, but its estimand is the same as when using the unadjusted method and its inference does not rely on the working model being correctly specified.
Covariate adjustment for parallel-group trials has been extensively studied recently. From the theory of semiparametrics (Robins et al., 1994; Tsiatis, 2006), Tsiatis et al. (2008) established a general class of consistent and asymptotically normal estimators for the average treatment effect. When linear working models are used, covariate adjustment using an analysis of heterogeneous covariance (ANHECOVA) working model that includes all treatment-by-covariate interaction terms can lead to guaranteed efficiency gain regardless of the model is misspecified or not (Yang and Tsiatis, 2001; Lin, 2013; Ye et al., 2022, 2023). These recent results have not been extended to crossover trials, although covariate adjustment is broadly recommended for crossover trials (Metcalfe, 2010; Mehrotra, 2014; Jemielita et al., 2016).
We consider adjusting for a baseline covariate vector
measured before randomization (ie, pre-randomization covariates) using the following covariate-adjusted ANHECOVA estimator
![]() |
where
is the sample mean of all
’s,
is the sample mean of
’s from subjects with Ai = a, Δi and
are defined in Equation 1, and
is the least squares estimator of
from fitting the linear working model
using subjects with Ai = a.
The following heuristics reveal why ANHECOVA does not change the estimand, often gains but never hurts efficiency even when the linear working model is wrong. As randomization balances the covariate distribution, both
and
estimate the same quantity and thus,
is an “estimator” of zero. Hence,
and
correspond to the same estimand. In addition, as n → ∞,
converges to
in probability, regardless of the linear working model is correct or not. Hence,
is asymptotically equivalent to
, whose variance is
![]() |
Consequently, the asymptotic variance of
is no larger than that of
. These results are formally stated in Theorem 3, which are all in the asymptotic sense.
Theorem 3
Under Assumption 1,
(a)
, where
, and
.
(b) Moreover,
.
Theorem 3 is proved by applying Theorem 1 and Corollary 1 in Ye et al. (2023) with Δi as the outcome. From Theorem 3(b), we see that the asymptotic variance of
is no larger than that of
, where the equality holds if and only if
. This occurs, for example, when
, that is, when the covariates are uncorrelated with the change in outcome Δi. In fact, Theorem 1 of Ye et al. (2023) implies a stronger result that
has the smallest asymptotic variance among all linearly adjusted estimators of the form
, where
are any fixed or random vectors that have the same dimension as
. A consistent estimator of
is
![]() |
where
is the sample variance of
based on subjects under Ai = a, for a = 0, 1, and
is the sample covariance matrix of
based on the entire sample. One can easily construct a Z-test based on the covariate-adjusted estimator
, which from Theorem 3 is guaranteed to be more powerful than the unadjusted counterpart Tcr in the asymptotic sense.
In this article, we have focused on adjusting for pre-randomization covariates. As discussed above, if the covariates are related to the change in outcome Δi, which could happen, for example, when there exist treatment-by-covariates interactions, then adjusting for pre-randomization covariates can reduce the variability of Δi and improve asymptotic efficiency. This is also shown in our simulation study in Section 5. However, in crossover trials, this adjustment typically yields small to modest benefits unless the covariates strongly influence the change in outcome. Previous studies have explored adjusting for period-dependent baseline covariates, that is, covariates that are measured before treatment is given within each period (Kenward and Roger, 2010; Jones and Kenward, 2015). Yet, one important caveat is that the washout period needs to be sufficiently long to ensure that the period-dependent baselines are not influenced by previous treatment (ie, by carry-over). In instances where the carry-over effect is a concern, more sophisticated methods to adjust for period-dependent covariates should be used (eg, the longitudinal G methods [Hernán and Robins, 2020, Chapter 21]). Exploring these methods will be a direction for future research.
5. POWER CALCULATIONS
To compare the power of parallel-group and crossover design, we consider 2 situations. In Case I, we consider the following simple data-generating process for which we can calculate the power using both the formula and simulations:
![]() |
In Case II, we consider a more complex data-generating process of the potential outcomes with treatment-by-covariate interactions:
![]() |
In both cases, Xij, ϵik ∼ N(0, 1) for j = 1, 2, 3 and k = 1, 2, 3, 4. The treatment arm indicator Ai is Bernoulli with π0 = π1 = 1/2. The observed outcomes for subject i are
if Ai = 1 and
if Ai = 0. The observed data are
. For the covariate-adjusted estimator, we adjust for
. We set
, θ* = 0, and
.
In Case I, because of the simple data-generating process, we can directly calculate that
,
,
, and thus
and
. In this case,
is a 3 × 3 identity matrix and
. Thus,
. From Corollary 1 in Ye et al. (2023), it is also easy to calculate the asymptotic variance of the covariate-adjusted ANHECOVA estimator using only the parallel-group design as
. Hence, we calculate the power based on both formulae and simulations. Figure 2 shows the type I error rate and power for 4 one-sided tests Tpr, Tpr, adj, Tcr, and Tcr, adj under Case I when
and λ0 = λ1 = λ ∈ { − 0.1, 0, 0.1, 0.3} based on the formula in (3) and (4) and its covariate-adjusted counterpart. Note that Tpr, adj is based on the ANHECOVA estimator that takes the same form as
but with Δi replaced by Yi1. In this setting, the value of b only affects the power of Tcr but not the other 2 tests, so we present
for Tcr only. In the supplementary materials, we present the type I error rate and power obtained by simulation, which are shown to agree with the power in Figure 2.
FIGURE 2.
Power curves for 4 tests under Case I calculated by formula when λ = −0.1, 0, 0.1, 0.3, and
under Case I. Note that the power of Tpr, Tpr, adj, and Tcr, adj is unaffected by b.
In Figure 2, the power of Tpr and Tpr, adj is not affected by λ and can be used to benchmark the performance of the other tests. When θ = 0, that is, under the null hypothesis of no treatment effect, the type I error rates of Tpr and Tpr, adj are equal to α = 0.025. The type I error rates of Tcr and Tcr, adj are greater than α when λ < 0, equal to α when λ = 0, and smaller than α when λ > 0. In other words, Tcr and Tcr, adj can control the type I error rate when λ ≥ 0; meanwhile, Tcr and Tcr, adj become more and more conservative as λ grows.
When θ > 0, that is, under the alternative hypothesis, the power comparison of Tcr and Tpr depends on the sign of
following the calculations at the end of Section 3.3. Specifically, with b = 0, Tcr is more powerful than Tpr when λ < 0.5θ. Hence, with no carry-over effect (λ = 0), we see that Tcr has substantial power gain compared to Tpr; with a small carry-over effect (λ = 0.1), Tcr is more powerful than Tpr when θ > 0.2; with a large carry-over effect (λ = 0.3), Tpr is more powerful than Tcr for 0 < θ ≤ 0.5. Furthermore, since a larger b leads to a larger ICC, it also results in a slightly higher power of Tcr. Lastly, adjusting for covariates that are related to the change in outcome can always increase power regardless of whether the model is correct or not. Hence, we see that the covariate-adjusted test Tcr, adj is always more powerful than the unadjusted test Tcr; in some cases, the power gain can be up to 17%.
Figure 3 shows the empirical type I error rate and power for Case II obtained by simulation from 10 000 repetitions. The results are similar to the results from Figure 3, and covariate adjustment can improve the power when there exist treatment-by-covariates interactions.
FIGURE 3.
Power curves for 4 tests under Case II calculated by simulations when λ = −0.1, 0, 0.1, 0.3.
6. REAL DATA EXAMPLE
6.1. Application to the REACH study
In this section, we revisit the REACH study (Nair et al., 2023) described in Section 1 and apply our methods to study the adherence difference between monthly DVR and daily oral TDF/FTC.
In the REACH study, a total of 247 participants were enrolled and randomized (1:1) to 2 treatment arms of product use: using the DVR for 6 months and switching to daily oral TDF/FTC for a second 6 months, or using the daily oral TDF/FTC for 6 months and then switching to DVR for 6 months. Participants’ baseline characteristics are summarized in Table 2.
TABLE 2.
Baseline characteristics by randomized crossover sequence in the REACH study.
| DVR first | TDF/FTC first | Overall | |
|---|---|---|---|
| (n = 124) | (n = 123) | (n = 247) | |
| Age in years | |||
| Median (IQR) | 18 (17-19) | 18 (17-19) | 18 (17-19) |
| Site location | |||
| South Africa—Cape Town | 30 (24.2%) | 30 (24.4%) | 60 (24.3%) |
| South Africa—Johannesburg | 34 (27.4%) | 33 (26.8%) | 67 (27.1%) |
| Uganda—Kampala | 30 (24.2%) | 30 (24.4%) | 60 (24.3%) |
| Zimbabwe—Harare | 30 (24.2%) | 30 (24.4%) | 60 (24.3%) |
| Marital status | |||
| Single | 106 (85.5%) | 108 (87.8%) | 214 (86.6%) |
| Married/cohabiting | 16 (12.9%) | 14 (11.4%) | 30 (12.1%) |
| Separated/divorced | 2 (1.6%) | 1 (0.8%) | 3 (1.2%) |
| Highest level of education | |||
| Primary | 15 (12.3%) | 18 (14.6%) | 33 (13.5%) |
| Secondary | 92 (75.4%) | 97 (78.9%) | 189 (77.1%) |
| Higher | 15 (12.3%) | 8 (6.5%) | 23 (9.4%) |
| Earns own income | 22 (17.7%) | 31 (25.2%) | 53 (21.5%) |
| Ever pregnant | 52 (41.9%) | 47 (38.2%) | 99 (40.1%) |
| Currently has a sex partner | 111 (89.5%) | 108 (87.8%) | 219 (88.7%) |
| Worry about HIV infection | 86 (39.4%) | 82 (66.7%) | 168 (68.0%) |
| Diagnosed with STI | 40 (32.3%) | 47 (38.2%) | 87 (35.2%) |
| Syphilis | 3 (2.4%) | 3 (2.4%) | 6 (2.4%) |
| Trichomoniasis | 8 (6.5%) | 5 (4.1%) | 13 (5.3%) |
| Gonorrhea | 13 (10.5%) | 8 (6.5%) | 21 (8.5%) |
| Chlamydia | 30 (24.2%) | 41 (33.3%) | 71 (28.7%) |
| CES depression scale | |||
| ≥12 (indicative of depression) | 28 (22.6%) | 38 (30.9%) | 66 (26.7%) |
| <12 | 89 (71.8%) | 78 (63.4%) | 167 (67.6%) |
| Missing | 7 (5.6%) | 7 (5.7%) | 14 (5.7%) |
| Alcohol use disorder (AUDIT-C score) | |||
| ≥3 (indicative of alcohol use disorder) | 50 (40.3%) | 47 (38.2%) | 97 (39.3%) |
| <3 | 74 (59.7%) | 75 (61.0%) | 149 (60.3%) |
| Missing | 0 (0.0%) | 1 (0.8%) | 1 (0.4%) |
| Preference for products | |||
| DVR | 51 (41.1%) | 43 (35.0%) | 94 (38.1%) |
| oral TDF/FTC | 43 (34.7%) | 57 (46.3%) | 100 (40.5%) |
| Equal preference | 29 (23.4%) | 22 (17.9%) | 51 (20.6%) |
| Did not answer | 1 (0.8%) | 1 (0.8%) | 2 (0.8%) |
Data are presented as median (IQR) or n (%).
AUDIT-C, alcohol use disorders identification test-concise; CES, Center for epidemologic studies; IQR, Interquartile range; STI, sexually transmitted infection.
We apply our method to compare adherence to monthly DVR and daily oral TDF/FTC. The binary adherence endpoint is defined as high use at the end of period 1 and period 2, with high use of daily oral TDF/FTC defined as tenofovir-diphosphate concentrations greater than or equal to 700 fmol/punch (associated with taking an average of 4 or more tablets per week in the previous month), and high use of monthly DVR defined as greater than or equal to 4 mg dapivirine released from the returned ring (continuous use for 28 days in the previous month). These adherence definitions are widely accepted in the literature (Nair et al., 2023). It is important to note, however, that the definitions require continuous daily usage of the ring, but not for oral TDF/FTC. For covariate adjustment, we consider baseline covariates that are likely to be associated with adherence, including age, site location, preference for either product, depression, alcohol use disorder, sexually transmitted infection, and participants’ level of worry about HIV infection. We use the single imputation method to handle the missing baseline covariates (Zhao and Ding, 2022).
To discuss the impact of potential carry-over effects in the analysis of this trial, we first note that for the DVR first arm, using unadjusted sample means, the estimated adherence rate to DVR at period 1 was 55.6% and to TDF/FTC was 44.4% at period 2; for the TDF/FTC first arm, the estimated adherence rate to TDF/FTC at period 1 was 48.0% and to DVR was 52.8% at period 2. Hence, the average adherence difference at period 1 is 55.6%–48.0% = 7.6%, and average adherence difference at period 2 is 52.8%–44.4% = 8.4%. This small difference between the treatment effect at periods 1 and 2 may be attributed to a larger treatment effect in period 2, and/or a negative carry-over effect (λ1) for the reason discussed in Section 1. However, because the difference is small, even if the carry-over effect exists, it is unlikely to substantially alter the result.
Table 3 presents both the parallel and crossover estimates and their covariate-adjusted counterparts, alongside the standard errors and 95% CIs. The parallel estimates use data from period 1 only. Estimates from all methods are similar, indicating that adherence to DVR is about 7.7%–8.5% higher compared to daily oral TDF/FTC. The SEs of the crossover estimators are smaller compared to the SEs of the parallel estimators, demonstrating that the crossover design is more efficient. Adjusting for covariates leads to slightly larger estimates and smaller standard errors. All the 95% CIs (except for the covariate-adjusted crossover one) cover 0, suggesting that the difference is not statistically significant.
TABLE 3.
Covariates-adjusted and unadjusted parallel and crossover estimators for the average treatment effect of DVR on adherence compared to daily oral TDF/FTC, with standard errors and 95% CIs.
| Type | Mean | SE | 95% CI |
|---|---|---|---|
| Parallel, unadjusted | 0.077 | 0.064 | (−0.048, 0.202) |
| Parallel, adjusted | 0.082 | 0.062 | (−0.040, 0.204) |
| Crossover, unadjusted | 0.081 | 0.043 | (−0.003, 0.165) |
| Crossover, adjusted | 0.085 | 0.042 | (0.003, 0.168) |
6.2. Using the REACH study to design a hypothetical trial
Despite being highly effective for HIV prevention, adherence to daily oral TDF/FTC is low among women (Celum et al., 2019). The dual prevention pill (DPP), a daily oral pill combining oral contraceptives and TDF/FTC, has the potential to increase women’s adherence to daily oral TDF/FTC (Friedland et al., 2021). In this section, using data from the REACH study, we use simulations to evaluate the power trade-off between the crossover design and parallel-group design in a hypothetical trial comparing adherence to a single daily DPP versus adherence to 2 daily pills (1 pill of TDF/FTC and 1 pill of oral contraceptive). The central hypothesis is that a DPP regime can increase women’s adherence to TDF/FTC.
Suppose adolescent girls and young women with baseline characteristics similar to those in the REACH study are recruited for this trial, and we use the same design as described in Section 6.1 to compare DPP and 2 daily pills. In this example, we anticipate a non-negative carry-over effect primarily because the DPP is expected to increase adherence in the first period, and those who adhere consistently in the first period tend to develop a routine or habit around taking daily oral pills, which may positively influence their adherence in the second period. From the results in Section 3, using a crossover design with the non-negative carry-over effect does not lead to type I error inflation.
Our simulations are based on the data of 123 participants who took TDF/FTC first in the REACH study. In the simulation study, we set π1 = π0 = 0.5,
,
where θ ∈ {0.05, 0.08, 0.10, 0.15} for different treatment effects. The ICC (ie, ρ) is estimated to be ρ = 0.39 from the 123 participants in the REACH study. Under the simple scenario described at the end of Section 3, when
, the crossover design is less powerful than the parallel one. So here we set λ1 = λ0 = λ ∈ {0, 0.25θ, 0.45θ} for different carry-over effects. We consider 4 test statistics: Tpr, Tpr, adj, Tcr, Tcr, adj. The simulation process involves the following steps:
Denote the baseline covariates and the adherence variable at week 24 from the REACH study as
, with N = 123. The average of
is 0.48. Fit a logistic regression model for the probability of adherence to TDF/FTC using age, site location, STI status, depression, alcohol use disorder, worry about HIV infection and baseline preferences for the products. Single imputation is applied for all missing values as described in Section 6.1. Denote the fitted model as
, where
.Generate
, where
satisfies
.Generate
from
and
, where s0 is chosen (as a function of
) to ensure that
and
; see details in Section S3.3 of the supplementary materials.Generate
from
and
, where s1 is chosen (as a function of
) to ensure that
and
. By now, we have obtained the covariates and 4 potential outcomes
for 123 individuals.A random sample of size n is drawn from the 123 individuals’ covariates and potential outcomes with replacement. Each of the obtained n subjects is assigned randomly to Ai = 0 or 1 with equal probability. For Ai = 1,
and
; for Ai = 0,
and
. Hence, we obtain the observed data
, and calculate the test statistics Tpr, Tpr, adj, Tcr, Tcr, adj.Steps (2)-(5) are repeated 2000 times for each θ and λ, from which we obtain the empirical powers of the test statistics.
Table 4 shows the empirical powers for Tpr, Tpr, adj, Tcr, Tcr, adj for different θ and λ based on 2000 simulations. The sample sizes n = 325, 680, and 1900 are determined such that Powercr is close to 80% when θ = 0.10. From Table 4, first, we can see that the empirical power increases when θ increases. Second, under either the parallel or crossover design, covariate adjustment always leads to a larger power. Third, when λ = 0, the crossover tests have considerably larger power than the parallel-group tests. When λ = 0.25θ, the crossover tests still have larger power, but the power difference is not as pronounced. When λ = 0.45θ, the crossover tests become slightly less powerful than the parallel-group tests.
TABLE 4.
Empirical power of Tpr, Tpr, adj, Tcr, and Tcr, adj for a hypothetical trial using data from the REACH study. For each λ, the sample size is determined such that Powercr is close to 80% when θ = 0.10.
| n | λ | θ | Powerpr | Powerpr, adj | Powercr | Powercr, adj |
|---|---|---|---|---|---|---|
| 325 | 0 | 0.05 | 0.214 | 0.236 | 0.423 | 0.446 |
| 0.08 | 0.342 | 0.379 | 0.670 | 0.681 | ||
| 0.10 | 0.439 | 0.484 | 0.797 | 0.818 | ||
| 0.15 | 0.738 | 0.770 | 0.973 | 0.978 | ||
| 680 | 0.25θ | 0.05 | 0.324 | 0.354 | 0.452 | 0.463 |
| 0.08 | 0.534 | 0.570 | 0.680 | 0.6701 | ||
| 0.10 | 0.679 | 0.720 | 0.799 | 0.808 | ||
| 0.15 | 0.905 | 0.923 | 0.959 | 0.961 | ||
| 1900 | 0.45θ | 0.05 | 0.558 | 0.579 | 0.523 | 0.528 |
| 0.08 | 0.778 | 0.789 | 0.684 | 0.688 | ||
| 0.10 | 0.876 | 0.889 | 0.791 | 0.797 | ||
| 0.15 | 0.988 | 0.989 | 0.929 | 0.935 |
7. DISCUSSION
The crossover is an efficient trial design that uses participants as their own controls. The carry-over effect, especially the behavioral carry-over effect, is an outstanding concern because it can bias the estimation of the treatment effect. Using a potential outcomes framework and minimal statistical assumptions, we investigate the impact of the carry-over effect in a 2-treatment 2-period crossover trial. Our results provide a clear characterization of how the carry-over effect influences estimation bias, as well as the type I error and power of tests. When the carry-over effect λ1 + λ0 is negative, the basic estimator overestimates the treatment effect, which can inflate the type I error of one-sided tests. Conversely, when the carry-over effect λ1 + λ0 is positive, the basic estimator underestimates the treatment effect, which does not inflate the type I error of one-sided tests but negatively impacts the power. Furthermore, when λ1 + λ0 is positive but relatively small compared to the treatment effect, the crossover design can still be more powerful than the parallel-group design. We can further apply covariate adjustment in crossover trials for guaranteed efficiency gain. All the methods in this article can be implemented using the RobinCar package in R (Ye et al., 2023).
We discuss our results in the context of 2 studies and demonstrate that the anticipated carry-over effect can be either non-positive or non-negative: one is the REACH study evaluating monthly DVR and daily oral (TDF/FTC), the other is a hypothetical future trial evaluating a single daily DPP and 2 daily pills. Through these 2 example studies, our key message is that while the crossover design is an efficient design, the carry-over effect needs to be carefully considered in trial planning and execution.
In this article, our primary focus has been on the power trade-off between the parallel-group design and the crossover design, particularly in the context of carry-over effects. However, this is just one among various practical considerations. First, a crossover trial typically requires a longer follow-up time, about twice as long as a parallel-group design, and this duration may be even longer if a washout period is incorporated to reduce carry-over effects (Senn, 2002). The longer follow-up time can increase costs and the risk of dropout. Second, the longer commitment and the need to switch treatments in a crossover design might discourage participants, potentially slowing recruitment. On the other hand, the opportunity to try multiple treatments in one study can make the crossover design more appealing to participants in certain scenarios (Diener et al., 2019). Meanwhile, in some trials, experiencing multiple treatments is essential for the trial objective. For example, in the REACH study, the participants need to experience both treatments before entering the choice period, which allows for the study of their preference between the 2 treatments.
Supplementary Material
Web Appendices referenced in Sections 2-6 are available with this paper at the Biometrics website on Oxford Academic. Code necessary to reproduce simulation and application results can be found in the supplementary materials.
Acknowledgement
We are grateful to the study participants, study staff, and investigators on the MTN-034/REACH study who provided the data for this analysis. We would also like to thank the anonymous referee, an Associate Editor, the Editor, and Professor Elizabeth Brown for their constructive comments.
Contributor Information
Danni Shi, Department of Biostatistics, University of Washington, Seattle, WA 98195, United States.
Ting Ye, Department of Biostatistics, University of Washington, Seattle, WA 98195, United States.
FUNDING
This work was supported by the National Institute of Allergy and Infectious Diseases [NIAID 5 UM1 AI068617].
CONFLICT OF INTEREST
None declared.
DATA AVAILABILITY
The data that support the findings in this paper are available from the Microbicide Trials Network. Restrictions apply to the availability of these data, which were used under license in this paper. Data are available from the authors with the permission of the Microbicide Trials Network.
References
- Araujo A., Julious S., Senn S. (2016). Understanding variation in sets of n-of-1 trials. PLoS One, 11, e0167167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bailey R., Kunert J. (2006). On optimal crossover designs when carryover effects are proportional to direct effects. Biometrika, 93, 613–625. [Google Scholar]
- Brown B. W. Jr. (1980). The crossover experiment for clinical trials. Biometrics, 36, 69–79. [PubMed] [Google Scholar]
- Celum C. L., Delany-Moretlwe S., Baeten J. M., Straten A., Hosek S., Bukusi E. A. et al. (2019). HIV pre-exposure prophylaxis for adolescent girls and young women in Africa: from efficacy trials to delivery. Journal of the International AIDS Society, 22, e25298. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Diener H.-C., Tassorelli C., Dodick D. W., Silberstein S. D., Lipton R. B., Ashina M. et al. (2019). Guidelines of the international headache society for controlled trials of acute treatment of migraine attacks in adults: fourth edition. Cephalalgia, 39, 687–710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- FDA (2023). Adjusting for covariates in randomized clinical trials for drugs and biological products. Guidance for Industry. Center for Drug Evaluation and Research and Center for Biologics Evaluation and Research, Food and Drug Administration (FDA), U.S. Department of Health and Human Services. May 2023.
- Freeman P. (1989). The performance of the two-stage analysis of two-treatment, two-period crossover trials. Statistics in Medicine, 8, 1421–1432. [DOI] [PubMed] [Google Scholar]
- Friedland B. A., Mathur S., Haddad L. B. (2021). The promise of the dual prevention pill: a framework for development and introduction. Frontiers in Reproductive Health, 3, 682689. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Grizzle J. E. (1965). The two-period change-over design and its use in clinical trials. Biometrics, 21, 467–480. [PubMed] [Google Scholar]
- Hemming K., Taljaard M., Weijer C., Forbes A. B. (2020). Use of multiple period, cluster randomised, crossover trial designs for comparative effectiveness research. BMJ, 371, m3800. [DOI] [PubMed] [Google Scholar]
- Hernán M., Robins J. (2020). Causal Inference: What If. Boca Raton: Chapman and Hall/CRC. [Google Scholar]
- Hills M., Armitage P. (1979). The two-period cross-over clinical trial. British Journal of Clinical Pharmacology, 8, 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jemielita T., Putt M., Mehrotra D. (2016). Improved power in crossover designs through linear combinations of baselines. Statistics in Medicine, 35, 5625–5641. [DOI] [PubMed] [Google Scholar]
- Jones B., Donev A. (1996). Modelling and design of cross-over trials. Statistics in Medicine, 15, 1435–1446. [DOI] [PubMed] [Google Scholar]
- Jones B., Kenward M. G. (2015). Design and Analysis of Cross-Over Trials. Boca Raton, FL: Chapman & Hall/CRC Press. [Google Scholar]
- Jones B., Lewis J. (1995). The case for cross-over trials in phase III. Statistics in Medicine, 14, 1025–1038. [DOI] [PubMed] [Google Scholar]
- Kenward M. G., Roger J. H. (2010). The use of baseline covariates in crossover studies. Biostatistics, 11, 1–17. [DOI] [PubMed] [Google Scholar]
- Kim K., Bretz F., Cheung Y. K. K., Hampson L. V. (2021). Handbook of Statistical Methods for Randomized Controlled Trials. New York: Chapman and Hall/CRC. [Google Scholar]
- Kunert J., Stufken J. (2002). Optimal crossover designs in a model with self and mixed carryover effects. Journal of the American Statistical Association, 97, 898–906. [Google Scholar]
- Laird N. M., Skinner J., Kenward M. (1992). An analysis of two-period crossover designs with carry-over effects. Statistics in Medicine, 11, 1967–1979. [DOI] [PubMed] [Google Scholar]
- Lin W. (2013). Agnostic notes on regression adjustments to experimental data: reexamining freedman’s critique. Annals of Applied Statistics, 7, 295–318. [Google Scholar]
- Mehrotra D. V. (2014). A recommended analysis for 2×2 crossover trials with baseline measurements. Pharmaceutical Statistics, 13, 376–387. [DOI] [PubMed] [Google Scholar]
- Metcalfe C. (2010). The analysis of cross-over trials with baseline measurements. Statistics in Medicine, 29, 3211–3218. [DOI] [PubMed] [Google Scholar]
- Minnis A. M., Roberts S. T., Agot K., Weinrib R., Ahmed K., Manenzhe K. et al. (2018). Young women’s ratings of three placebo multipurpose prevention technologies for HIV and pregnancy prevention in a randomized, cross-over study in Kenya and South Africa. AIDS and Behavior, 22, 2662–2673. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nair G., Celum C., Szydlo D., Brown E. R., Akello C. A., Nakalega R. et al. (2023). Adherence, safety, and choice of the monthly dapivirine vaginal ring or oral emtricitabine plus tenofovir disoproxil fumarate for HIV pre-exposure prophylaxis among african adolescent girls and young women: a randomised, open-label, crossover trial. The Lancet HIV, 10, E779–E789. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Neyman J. (1923). On the application of probability theory to agricultural experiments essay on principles. Section 9. Statistical Science, 5, 465–472. Trans. Dorota M. Dabrowska and Terence P. Speed (1990). [Google Scholar]
- Robins J. M., Rotnitzky A., Zhao L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89, 846–866. [Google Scholar]
- Rosenbaum P. R. (2020). Design of Observational Studies 2nd ed. New York: Springer. [Google Scholar]
- Rubin D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 6, 688–701. [Google Scholar]
- Senn S. (1994). The AB/BA crossover: past, present and future?. Statistical Methods in Medical Research, 3, 303–324. [DOI] [PubMed] [Google Scholar]
- Senn S. (2002). Cross-over Trials in Clinical Research. West Sussex: John Wiley and Sons, Ltd. [Google Scholar]
- Senn S. J. (1997). Letter to the editor: the case for cross-over trials in phase iii by B.J. Jones and J. Lewis, statistics in medicine, 14, 1025-1038 (1995). Statistics in Medicine, 16, 2021–22. [DOI] [PubMed] [Google Scholar]
- Tsiatis A. A. (2006). Semiparametric Theory and Missing Data. New York: Springer. [Google Scholar]
- Tsiatis A. A., Davidian M., Zhang M., Lu X. (2008). Covariate adjustment for two-sample treatment comparisons in randomized clinical trials: a principled yet flexible approach. Statistics in Medicine, 27, 4658–4677. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang L., Tsiatis A. A. (2001). Efficiency study of estimators for a treatment effect in a pretest–posttest trial. The American Statistician, 55, 314–321. [Google Scholar]
- Ye T., Bannick M., Yi Y., Bian F. (2023). RobinCar: robust estimation and inference for covariate-adaptive randomization. https://github.com/tye27/RobinCar [Accessed 20 March 2023].
- Ye T., Shao J., Yi Y., Zhao Q. (2023). Toward better practice of covariate adjustment in analyzing randomized clinical trials. Journal of the American Statistical Association, 118, 544, 2370–2382. [Google Scholar]
- Ye T., Yi Y., Shao J. (2022). Inference on the average treatment effect under minimization and other covariate-adaptive randomization methods. Biometrika, 109, 33–47. [Google Scholar]
- Zhao A., Ding P. (2024). To adjust or not to adjust? estimating the average treatment effect in randomized experiments with missing covariates. Journal of the American Statistical Association, 119, 545, 450–460. [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Web Appendices referenced in Sections 2-6 are available with this paper at the Biometrics website on Oxford Academic. Code necessary to reproduce simulation and application results can be found in the supplementary materials.
Data Availability Statement
The data that support the findings in this paper are available from the Microbicide Trials Network. Restrictions apply to the availability of these data, which were used under license in this paper. Data are available from the authors with the permission of the Microbicide Trials Network.






























