Abstract
Pre-post designs are widely used in clinical trials and experimental studies to assess the effectiveness of treatments. Common statistical methods for analyzing pre-post data include analysis of variance (ANOVA) using post-treatment or the change from baseline, analysis of covariance (ANCOVA) with homogeneous or heterogeneous slopes, and linear mixed models (LMM). While numerous studies have compared these methods, limited studies have investigated the impact of adjusting for influential baseline covariates under different randomization approaches. In this study, we conducted a series of comprehensive simulation studies to investigate the impact of adjusting baseline covariates under several randomization approaches: simple randomization, stratified block randomization, and covariate adaptive randomization using the minimization method by Pocock and Simon. Results demonstrated that when no covariates were considered in the randomization approach, the two ANCOVA methods always have good performance. Adjusting for relevant baseline covariates led to substantial power gains, with the extent of these gains depending on the size of the covariate effects and the randomization approach employed. Stratified block randomization and covariate adaptive randomization consistently outperformed simple randomization in terms of power gains after adjusting for covariates, with covariate adaptive randomization becoming more superior as the number of covariates increased.
Keywords: ANCOVA, Covariate adaptive randomization, Linear mixed model, Pre-post design, Stratified block randomization
Introduction
In clinical trials and experimental studies, the pre-post design is widely used to evaluate the effectiveness of interventions or treatments. The design measures the outcome of interest before the intervention (pre-test) and after the intervention (post-test) in both the treatment and control groups. The primary goal of this design is to assess the treatment effect, defined as the difference between the treatment and control groups on specific study endpoints. Some common study endpoints include the post-test score, the change score of the post-test measure from baseline, the percentage change from baseline, and the rate of change from baseline [1, 2]. By comparing these measurements between the two groups, researchers can quantify the treatment effect and determine the impact of the intervention on the outcome variable.
Several statistical methods have been proposed to analyze data from pre-post studies, including analysis of variance (ANOVA), analysis of covariance (ANCOVA), and linear mixed models (LMM). One of the simplest approaches is ANOVA-post, which compares the post-test scores between groups, ignoring the baseline scores [3]. This method is straightforward, but it does not account for any baseline difference between the groups [4]. ANOVA-change, on the other hand, analyzes the change scores from pre-test to post-test, providing a direct measure of the treatment effect as the difference in mean changes between groups [5].
ANCOVA methods adjust for baseline differences by incorporating the pre-test score as a covariate in the model. Including pre-test responses can help minimize error variance, thereby leading to statistical power gain to detect the treatment effect compared to methods that do not utilize pre-test outcomes [6–9]. ANCOVA-hom, which represents the homogeneous ANCOVA model, assumes that the relationship between the pre-test and post-test responses is the same for both groups. ANCOVA-hom models the post-test response with the pre-test score as a covariate and the treatment group as the independent variable. This method may mitigate the bias from the imbalance of baseline responses and can provide more efficient estimates of the treatment effect [10, 11]. ANCOVA-het, representing heterogeneous ANCOVA, allows for the possibility that the relationship between the pre-test and post-test scores differs between the treatment and control groups. ANCOVA-het extends ANCOVA-hom by further including an interaction term of the pre-test score and the group indicator. This method captures differential treatment effects based on the pre-test responses, and the inclusion of the interaction term allows ANCOVA-het to account for varying treatment effects across different baseline status [11].
In addition to the ANOVA and ANCOVA models, LMMs have gained popularity as a flexible alternative to the above methods by accounting for repeated measures within subjects and including random effects to capture individual variability [12]. These approaches can also handle complex data structures, such as varying time points and missing data [13, 14]. LMMs provide a comprehensive framework for analyzing longitudinal data, incorporating both fixed and random effects. In the context of pre-post designs, LMMs can include the treatment group, time (pre-test vs. post-test), and their interaction as fixed effects, while allowing for participant-specific random intercept and slope [1, 5]. This allows for accurate modeling of within-subject correlations, leading to more precise estimates of the treatment effect.
While numerous studies have compared the performance of various statistical methods for analyzing pre-post data [1, 5, 14], very limited studies have considered the impact of adjusting for additional baseline covariates on the treatment effect estimation under different randomization approaches. Adjusting for influential factors can potentially improve the precision of treatment effect estimates and increase the power to detect significant differences between groups [15]. However, the extent to which covariate adjustment can enhance the efficiency of different statistical methods under various randomization approaches remains unclear.
In pre-post studies, simple randomization (SR) is frequently utilized to allocate participants to different treatment groups, with the expectation that this approach will create comparable groups by evenly distributing potential confounders across the groups [16]. This approach is straightforward to implement and minimize the risk of selection bias, as the assignment of participants is entirely based on chance [17]. Nevertheless, when sample size is relatively small or when influential baseline covariates are strongly associated with the outcome variable, simple randomization may not adequately balance these covariates between different treatment groups [18]. This imbalance can lead to biased estimates of the treatment effect and low statistical power. To address this issue, alternative randomization methods, such as stratified block randomization (SBR) and covariate adaptive randomization (CAR), can be employed to improve balance and potentially enhance the accuracy of treatment effect estimation [19]. The details of these two methods and their implementation will be thoroughly discussed in “Randomization techniques to balance baseline covariates” section.
In this study, we aim to evaluate the impact of incorporating influential baseline covariates into the analysis of pre-post studies and investigate the performance of the common statistical methods under different randomization approaches. We first compare the performance of the models without the presence of additional baseline covariates. Following this, we explore the impact of adjusting for one or more covariates on the performance of the five models under different randomization approaches, specifically SR, SBR, and CAR. By considering different magnitudes of the covariate effect, we aim to assess how the strength of the association between the covariate and the outcome influences the benefits of covariate adjustment. Additionally, by comparing the performance of the models under various randomization approaches, we seek to identify the most effective strategies for balancing baseline covariates and optimizing the power of the statistical analyses. Through this comprehensive study, we aim to provide practical guidance for researchers in selecting appropriate statistical methods and randomization approaches when analyzing pre-post data, particularly in the presence of influential baseline covariates. While the statistical methods we examine are well-established, our study provides the first systematic investigation of how their performance depends on the randomization scheme employed.
The remainder of the paper is structured as follows: “Five methods for estimating treatment effects” section details the five statistical methods commonly used for analyzing pre-post data. “Randomization techniques to balance baseline covariates” section elaborates on the randomization approaches and explains how the models can be adjusted to incorporate additional baseline covariates. The results from the simulation studies are presented in “Numerical study” section. Lastly, “Discussion” section concludes the paper with a discussion of our findings.
Methods
Five methods for estimating treatment effects
We consider a pre-post two-arm randomized trial with a total of n subjects. Let
be the response of the i-th participant at time j, where
,
for pre-test and
for post-test. Each participant is allocated to one of the two groups, and the group assignment is represented by
, with
denoting the control group and
denoting the treatment group. We assume
follows the multivariate normal distribution (MVN):
where
and t. Here,
and
denote the common mean and variance of the baseline outcome in the two groups.
and
represent the means of the post-test outcome in the control and treatment groups. The variances of the post-test responses are given by
for the control group and
for the treatment group. The correlation coefficients for the control and treatment arms are
and
, respectively.
To estimate the treatment effect in a pre-post two-arm randomized trial, we consider five common models: ANOVA-post, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM [11, 20].
ANOVA models
ANOVA-post is a statistical method used to compare the means of post-test responses between different groups while ignoring the pre-test responses, with the outcome as
where
. ANOVA-change model compares the change from pre-test to post-test between groups. The change score is
.
The two ANOVA models can be represented as:
where
is the treatment effect and
is the error term that is independently and identically normally distributed with mean 0 and variance
.
ANCOVA models
ANCOVA-het, which represents the heterogeneous ANCOVA model, relaxes the assumption of homogeneity of regression slopes. This model allows for different relationships between the pre-test and post-test scores for the treatment and control groups.
To estimate the treatment effect using ANCOVA-het, we first center the pre-test responses by subtracting the overall mean pre-test response from each individual’s pre-test response, and then model the post-test outcome using the group indicator, the centered baseline outcome, and the interaction between the group and the centered baseline outcome. The model can be expressed as follows:
where
is the parameter of interest to test the group difference.
ANCOVA-hom model accounts for the baseline difference by including the pre-test score as a covariate in the model with
. ANCOVA-het extends ANCOVA-hom by including an interaction term between the pre-test score and the group indicator. This interaction term captures the potential difference in slopes between the groups, allowing for a more flexible model that can account for varying treatment effects across different baseline levels. By incorporating the interaction term, this model can be particularly useful when there is evidence that the treatment effect may differ depending on the initial status of the participants.
LMM
Linear mixed models are commonly used in pre-post study designs because they effectively handle repeated measurements collected from each participant. By incorporating both fixed and random effects, LMMs account for the correlation between pre-test and post-test outcomes within the same individual. In our setting, we specify the LMM as:
where
represents the outcome for participant i at time j,
denotes the treatment group assignment, and
indicates the time point. The parameters
, and
represent the fixed effects for the intercept, group, time, and group-by-time interaction. The model also incorporates subject-specific random effects
and
, representing random intercepts and slopes, respectively. These random effects are assumed to follow a bivariate normal distribution with mean zero and variance-covariance matrix containing variance components
for the random intercepts,
for the random slopes, and their covariance
. The residual errors
are assumed to be independently and identically distributed with mean zero and variance
.
Since our simulation specifically focuses on a pre-post design with exactly two measurements per participant, the LMM naturally fits this longitudinal data structure. By modeling participant-specific intercepts and slopes, we effectively account for within-subject correlations and individual variability, ensuring that our simulated data aligns clearly with the assumptions and structure of the LMM. The model can be fitted using restricted maximum likelihood (REML) or maximum likelihood (ML) methods, using the lme() function in R.
Randomization techniques to balance baseline covariates
In pre-post designs, simple randomization is often used due to its straightforward implementation and its ability to give each participant an equal chance of being assigned to any group. However, when there are influential baseline covariates (e.g., gender, age group, disease severity) that may influence the outcome, SR may lead to imbalances between the groups, especially in small to moderate sample sizes. Imbalanced covariates can reduce the power of the study and potentially confound the treatment effect, therefore affecting the validity of the trial’s results. Therefore, alternative randomization methods can be used to mitigate this problem, such as covariate adaptive randomization and stratified block randomization. SBR improves balance by dividing participants into strata based on key baseline covariates and then performing block randomization within each stratum. This method ensures that the distribution of covariates is balanced across treatment groups within each stratum.
CAR approaches further enhance balance by dynamically adjusting the allocation of participants based on existing covariate distributions. A common CAR approach is the minimization method, which was proposed by Taves [21] and expanded by Pocock and Simon [22]. The minimization method assigns each new participant to the group that minimizes the imbalance in covariates. The process begins by randomly assigning the first participant to one of the groups to initiate the trial. For each subsequent participant, the method calculates the imbalance that would result from assigning the participant to each of the possible groups. Imbalance is typically measured using one of several methods: standard deviation (SD), range, or variance. These three methods quantify imbalance differently for the covariate distributions across treatment groups. The range method calculates imbalance as the difference between the maximum and minimum number of participants assigned to each treatment within a covariate level. The variance method computes the variance of covariate counts across all treatments, and the SD method calculates the standard deviation of the covariate distribution across treatment groups [22, 23]. In this study, we specifically employed the SD method based on recent findings by Shan et al. [24], who demonstrated that the SD approach consistently achieves superior performance in balancing covariate distributions while maintaining lower allocation predictability compared to the variance and range methods.
The participant is then assigned to the group that would minimize the overall imbalance in covariates. Three formulas were proposed by Pocock and Simon for determining the treatment assignment probability, and we refer to them as the
and
. These formulas differ in their approach to calculating and assigning treatment probabilities based on the imbalance scores [24].
The
approach assigns a higher probability, p, to the treatment arm with the lowest total imbalance score, while the remaining probability,
, is equally distributed among the other treatment arms. The probabilities are determined as follows:
where K is the number of treatment arms and
. The value of p is often chosen to be greater than 1/K to reduce the overall imbalance score.
The
approach assigns probabilities in a decreasing order based on the rank of the total imbalance score. Specifically, the probability assigned to the group ranked b is calculated by the following formula:
where q is a predefined constant between 1/K and
. For a study with 2 arms,
is equivalent to
when
.
The
calculates the treatment assignment probabilities by weighting them inversely proportional to the total imbalance scores for each treatment arm. The probabilities are updated dynamically as each new participant is enrolled. The formula for this approach is given by:
![]() |
where
represents the total imbalance score in the treatment arm b, with
. The value of r is a constant that ranges between 0 and 1. We chose
in this article [24]. The
method often has slightly better performance than the other two methods [24]. For that reason, the
method was used in this article.
Numerical study
Analysis without additional covariates
First, we conducted simulation studies to evaluate and compare the statistical power of the five models (ANOVA-post, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM) using simple randomization. These simulations were designed to focus specifically on the efficiency of each model without the presence of additional baseline covariates that might require balancing.
Data simulations
We conducted a series of simulations with varying sample sizes. The baseline score for each participant was generated from a normal distribution with a mean
and a variance
. Participants were assigned to the treatment group with a probability of 0.5, mimicking a simple equal randomization process. The post-test responses were generated from the following linear function:
where
follows a normal distribution with mean 0 and variance
. The coefficients
,
,
, and
in the linear model define the relationships between the baseline and post-test outcomes. In this set of simulations, we chose the values of these coefficients to be (
,
,
)= (0, 1, 0). The values of
and
were selected to achieve different correlation structures between the control and treatment groups. We considered the following combinations of correlations:
(0.2,0.2), (0.5,0.5), (0.8,0.8), (0.3,0.5), (0.5,0.3), and (0.4,0.3). Given
,
, and
, the variance of the error term in the control group is
. For example,
in the control group is 3 when
,
, and
. The variance of the error term in the treatment group can be calculated by replacing
with
. This configuration was designed to simplify the model with the focus on the direct effects of the treatment and baseline scores without additional complexity from interaction effects. The simulations were repeated 5,000 times for each sample size, and the statistical power of each model was calculated as the proportion of simulations in which the model correctly rejected the null hypothesis of no treatment effect at a significance level of 0.05.
Results with
Figure 1 presents the statistical power of the five models under simple randomization for different sample sizes and correlation structures, without adjusting for additional covariates. The results demonstrate that the performance of the models varies depending on the values of
and
. When
(low correlation), ANCOVA-hom and ANCOVA-het exhibit the highest power, followed closely by ANOVA-post, while LMM and ANOVA-change display lower power. As the sample size increases from 100 to 200, the power of all models improves, with ANCOVA-hom and ANCOVA-het maintaining their advantage over the other models. For
(moderate correlation), ANCOVA-hom and ANCOVA-het continue to demonstrate the highest power, with an average power of 0.678. ANOVA-post, ANOVA-change and LMM show similar performance, with a lower average power of 0.563. When
(high correlation), ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM exhibit similar power, with an average power of 0.534. ANOVA-post, however, consistently shows lower power, with an average power at 0.242. In scenarios where the correlations differ between the control and treatment groups (
and
,
and
,
and
), ANCOVA-hom and ANCOVA-het again show the highest power, followed by ANOVA-post. LMM and ANOVA-change have relatively lower power compared to the other models. Across all scenarios, the two ANCOVA models consistently demonstrate high power.
Fig. 1.
Statistical power of five models using simple randomization across different sample sizes for a study with (
,
,
)= (0, 1, 0)
The lower power observed for ANOVA-change compared to ANCOVA methods, particularly under low to moderate baseline-post correlations, reflects their different approaches to using baseline information. ANOVA-change analyzes simple difference scores (post minus pre), which assumes a fixed, equal influence of baseline scores on post-test outcomes. In contrast, ANCOVA explicitly estimates how baseline scores relate to post-test outcomes, allowing the data to determine the optimal weight for baseline adjustment. This flexibility reduces unexplained variance and thus gives ANCOVA better power when correlations are low to moderate.
Impact of
To investigate the effect of
on the statistical power, we considered the following study parameters with
,
, and (
,
,
)= (0, 0.05, 0.1). We present the statistical power for the five methods in Fig. 2 for a study using simple randomization. It can be seen that the two ANCOVA methods have larger statistical power than other methods. The ANOVA-post method could have very low statistical power when the correlation is high. As the correlation goes up, the ANOVA-change method and the LMM method improve these statistical power but the ANCOVA methods are still associated with the highest statistical power. Among the two ANCOVA methods, it is as expected that the ANCOVA-het method has higher power than the ANCOVA-hom as the simulated data does not satisfy the homogeneous slope between the two groups.
Fig. 2.
Statistical power of five models using simple randomization across different sample sizes when (
,
,
)= (0, 0.05, 0.1)
Impact of single covariate adjustment
We conducted simulations to assess the performance of the models when adjusting for one baseline covariate under different randomization approaches: SR, SBR, and CAR. For illustrative purposes, we considered sex (0 = female, 1 = male) as an example covariate in these simulations. The setup of these simulations in this subsection and the following subsections was similar to the one in “Data simulations” section. However, in this set of simulations, we used
, which was achieved by setting
and adjusting the variance of the error term to be
. This choice of correlation structure allows us to focus on the impact of covariate adjustment when the baseline and post-treatment outcomes are strongly correlated in both the control and treatment groups. The sex covariate was generated from a Bernoulli distribution with a probability of 0.5, ensuring an equal distribution of males and females in the simulated data. We set the correlation between the baseline score and sex to be 0, representing a scenario where the covariate is independent of the baseline score. For SBR, we stratified the participants based on sex and then performed block randomization within each stratum using a block size of 8. For CAR, we used Pocock and Simon’s minimization method to balance the sex covariate between the treatment and control groups. We chose the standard deviation (SD) method to measure the imbalance, and the treatment assignment probability was determined using the
method. The post-test responses were then generated using the following modified linear relationship:
where
represents the sex covariate for the i-th participant, and
is the corresponding coefficient. To investigate the impact of the sex covariate on the post-test response, we set
to be 0.5, 1, and 2 to represent small, moderate, and large sex effects, respectively. The models were then fitted with the sex covariate as an additional predictor, and the statistical power was calculated based on 2,000 simulations for each sample size, randomization method, and the
value.
The simulation results presented in Table 1 provide a comprehensive comparison of the average power for each statistical method under different randomization approaches, sex effects, and covariate adjustment strategies (none or adjusted for sex). These power values represent averages across sample sizes ranging from 100 to 300 per group, which ensures adequate power for meaningful comparisons. We observed that ANOVA-post consistently demonstrates the lowest average power, approximately 50%, regardless of the randomization method or sex effect size. In contrast, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM exhibit substantially higher average power ranging from 85% to 88%, with similar performance across configurations. These high power values confirm the efficiency of these methods in the pre-post design setting. As shown in Table 1, SBR and
tend to yield slightly higher power compared to simple randomization, particularly when the sex effect size is moderate or large.
Table 1.
Comparison of average power for statistical models across randomization approaches and covariate effects, unadjusted and adjusted for single covariate
| Models | Sex effect | Average power across sample sizes | |||||
|---|---|---|---|---|---|---|---|
| Unadjusted | Adjusted for covariate | ||||||
| SR1 | SBR | CAR | SR | SBR | CAR | ||
| ANOVA-post | Small | 0.503 | 0.501 | 0.503 | 0.503 | 0.503 | 0.504 |
| Moderate | 0.497 | 0.500 | 0.506 | 0.501 | 0.503 | 0.510 | |
| Large | 0.495 | 0.492 | 0.494 | 0.508 | 0.504 | 0.509 | |
| ANOVA-change | Small | 0.875 | 0.876 | 0.877 | 0.875 | 0.877 | 0.878 |
| Moderate | 0.872 | 0.871 | 0.873 | 0.878 | 0.877 | 0.878 | |
| Large | 0.849 | 0.854 | 0.859 | 0.878 | 0.880 | 0.880 | |
| ANCOVA-hom | Small | 0.874 | 0.874 | 0.876 | 0.875 | 0.876 | 0.877 |
| Moderate | 0.870 | 0.870 | 0.872 | 0.878 | 0.876 | 0.878 | |
| Large | 0.848 | 0.853 | 0.858 | 0.876 | 0.878 | 0.879 | |
| ANCOVA-het | Small | 0.874 | 0.875 | 0.876 | 0.875 | 0.876 | 0.877 |
| Moderate | 0.870 | 0.870 | 0.872 | 0.878 | 0.876 | 0.877 | |
| Large | 0.848 | 0.853 | 0.858 | 0.876 | 0.878 | 0.879 | |
| LMM | Small | 0.875 | 0.876 | 0.877 | 0.875 | 0.877 | 0.878 |
| Moderate | 0.872 | 0.871 | 0.873 | 0.878 | 0.877 | 0.878 | |
| Large | 0.849 | 0.854 | 0.859 | 0.878 | 0.880 | 0.880 | |
Figure 3 provides a detailed view of power trajectories across sample sizes ranging from 40 to 120 per group. While Table 1 presents average power across larger sample sizes (100-300 per group), Fig. 3 and Table 2 focus on a smaller sample size range (40-120 per group) to better illustrate how covariate adjustment impacts power, particularly in studies with limited sample sizes where such adjustments can be most beneficial. As shown in the figure, for larger sample sizes (100-120 participants per group), most methods consistently achieve power levels exceeding 0.6. It can also be seen that adjusting for the sex covariate generally improves the power of all models across most scenarios, with the extent of the power increase varying depending on the sample size, randomization method, and the size of the sex effect. Specifically, the power gains are more pronounced for larger sex effect sizes and smaller sample sizes.
Fig. 3.
Power analysis of five models with covariate adjustment under different randomization methods across varying sample sizes and sex effects (small, moderate, large)
Table 2.
Average power gains from single covariate adjustment for different statistical models and randomization approaches across various covariate effects
| Models | Sex effect | Average power gain | ||
|---|---|---|---|---|
| SR | SBR | CAR | ||
| ANOVA-post | Small | −0.4% | −0.4% | −0.3% |
| Moderate | 0.7% | 0.5% | 2.0% | |
| Large | 1.1% | 5.1% | 5.3% | |
| ANOVA-change | Small | −0.3% | 0.2% | 0.4% |
| Moderate | 0.9% | 2.2% | 2.0% | |
| Large | 7.3% | 9.2% | 8.2% | |
| ANCOVA-hom | Small | −0.4% | 0.5% | 0.7% |
| Moderate | 0.7% | 1.5% | 2.0% | |
| Large | 7.3% | 8.8% | 8.4% | |
| ANCOVA-het | Small | −0.3% | 0.1% | 0.5% |
| Moderate | 0.7% | 1.5% | 2.0% | |
| Large | 7.2% | 8.9% | 8.2% | |
| LMM | Small | −0.3% | 0.2% | 0.4% |
| Moderate | 0.9% | 2.2% | 2.0% | |
| Large | 7.3% | 9.2% | 8.2% | |
Table 2 further quantifies the power gains observed in Fig. 3 by presenting the average power gain achieved by adjusting for the sex covariate as compared to the model without adjusting for sex, for each statistical method under different randomization approaches and sex effect sizes. The results in Table 2 indicate that adjusting for sex generally leads to an increase in power, with the extent of the power gains being dependent on the sex effects. For a small sex effect, the average power gains are minimal, ranging from −0.4% to 0.7% across all randomization approaches and statistical methods. ANOVA-post shows no improvement in power (average power gains of −0.4% to −0.3%) regardless of the randomization method. The other four models demonstrate small but positive gains under SBR and
, with average power gains ranging from 0.1% to 0.7%. As the sex effect size increases to moderate and large, the power gains become more substantial, with average improvements ranging from 0.5% to 2.2% and 1.1% to 9.2%, respectively. In general, ANOVA-change, LMM, and the ANCOVA methods consistently achieve higher average power gains from sex adjustment compared to ANOVA-post, regardless of the randomization approach or the sex effects.
Comparing the impact of different randomization methods, it is evident that both SBR and CAR methods (
) consistently outperform simple randomization in terms of average power gains after adjusting for sex. Simple randomization shows modest power improvements, generally ranging from −0.4% to 7.3%. Block randomization shows slightly higher average power gains, ranging from 0.4% to 9.2%. CAR using
also provides substantial power improvements, with average power gains ranging from −0.3% to 8.4%.
To better understand how these randomization methods perform under different covariate distributions and how covariate imbalance affects the the benefit of adjustment, we conducted additional analyses focusing on varying levels of covariate balance. Figure 4 presents a detailed comparison for moderate sex effect (
) across sample sizes from 60 to 100 per group, comparing balanced (50% male) and moderately imbalanced (70% male) scenarios. In this focused view, several key patterns emerge. Under balanced conditions, the initial power before adjusting for the covariate are generally similar across all randomization methods. After covariate adjustment, all methods show notable power improvements, with SBR and CAR demonstrating slightly higher power than simple randomization in most cases. In contrast, under moderate imbalance, CAR using
consistently demonstrates good performance before covariate adjustment, particularly for ANOVA-change, ANCOVA methods, and LMM under moderate sample sizes. Notably, when combing with covariate adjustment, CAR also achieves the highest absolute power across models for smaller sample sizes under imbalanced conditions.
Fig. 4.
Power comparison of statistical models with covariate adjustment under balanced and imbalanced covariate distributions across randomization methods for moderate sex effect
To further investigate this phenomenon, we examined power gains from covariate adjustment under different levels of covariate balance when the sex effect was large (
). Table 3 presents the average power gains from adjusting for sex across different statistical methods and randomization approaches under balanced (50% male), slight imbalance (60% male), and moderate imbalance (70% male) scenarios. The results reveal that all methods benefit substantially from sex adjustment, with ANOVA-post showing gains of 2.6% to 4.9%, while the other four methods demonstrate larger gains ranging from 6.9% to 9.2%. While all methods show substantial power gains after covariate adjustment, both stratified block randomization and covariate adaptive randomization demonstrate decreasing gains as covariate imbalance increases. For SBR, power gains for ANOVA-change and LMM decrease from 9.2% (balanced) to 7.1% (moderate imbalance), and power gains for ANCOVA models decrease from more than 8% to around 7%. This likely reflects the reduced variability within strata when one stratum becomes small; with fewer females under moderate imbalance, block randomization within the female stratum provides less opportunity for chance imbalances to occur. CAR shows a similar but less pronounced decrease (from 8.7% to around 7%) in power gains after adjusting for the sex covariate. This decreasing gain actually reflects CAR’s proactive balancing during randomization. When covariates are naturally imbalanced, CAR has already addressed much of the imbalance, leaving less room for additional improvement through statistical adjustment. These findings highlight that the benefit of covariate adjustment depends not only on the covariate effect size but also on how much imbalance remains after randomization.
Table 3.
Average power gains from single covariate adjustment for different statistical models and randomization approaches under large sex effect across different covariate distributions
| Models | Sex distributiona | Average power gain | ||
|---|---|---|---|---|
| SR | SBR | CAR | ||
| ANOVA-post | Balanced | 3.0% | 4.2% | 4.6% |
| Slight imbalance | 3.4% | 3.7% | 4.9% | |
| Moderate imbalance | 2.6% | 2.9% | 4.8% | |
| ANOVA-change | Balanced | 7.4% | 9.2% | 8.7% |
| Slight imbalance | 8.3% | 7.4% | 8.5% | |
| Moderate imbalance | 6.9% | 7.1% | 7.1% | |
| ANCOVA-hom | Balanced | 7.9% | 8.1% | 8.7% |
| Slight imbalance | 8.3% | 8.1% | 8.4% | |
| Moderate imbalance | 7.1% | 7.3% | 7.7% | |
| ANCOVA-het | Balanced | 7.9% | 8.4% | 8.7% |
| Slight imbalance | 8.1% | 8.3% | 8.5% | |
| Moderate imbalance | 7.0% | 7.3% | 7.6% | |
| LMM | Balanced | 7.4% | 9.2% | 8.7% |
| Slight imbalance | 8.3% | 7.4% | 8.5% | |
| Moderate imbalance | 6.9% | 7.1% | 7.1% | |
aSex distribution categories: Balanced (50% male), Slight imbalance (60% male), and Moderate imbalance (70% male). All results shown are for large sex effect (
)
Impact of multiple covariate adjustments
To further explore the impact of adjusting for multiple covariates, we performed two more sets of simulations. The first set incorporated sex (0 = female, 1 = male) and the number of sites (0, 1, or 2) as covariates, while the second set included an additional covariate for the severity of disease (0 = absence of the disease, 1 = mild level of the disease, 2 = moderate level, 3 = severe level). The covariates were generated independently of the baseline responses, with equal probabilities assigned to each category within the respective covariate.
For the first set of simulations, the post-treatment responses were modeled as:
where
and
denote the sex and site covariates for the i-th participant, respectively, and
and
represent the corresponding coefficients. We set
and
to 0.5, 1, and 2 to represent small, moderate, and large effects for both covariates. The models were then adjusted to include both the sex and number of sites covariates as additional predictors.
In the second set of simulations with three covariates, the post-treatment responses were generated using the following equation:
where the third covariate
represents the severity of disease for the i-th participant, and
denotes the corresponding coefficient. The values of
,
, and
were varied (0.5, 1, and 2) to represent small, moderate, and large effect sizes for all three covariates. The models were then expanded to include sex, number of sites, and severity of disease as additional predictors. The statistical power was evaluated based on 2,000 simulations for each sample size, randomization method, and combination of
,
, and
values.
The results of the first set of simulations, presented in Table 4, highlight the average power gain achieved by adjusting for two baseline covariates (sex and number of sites) for each statistical method under different randomization approaches and selected combinations of covariate effect sizes. The results demonstrate that adjusting for multiple covariates can lead to substantial improvements in power, particularly when the covariate effects are moderate to large. Comparing the methods, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM consistently show higher power gains than ANOVA-post across all scenarios. As the covariate effect sizes increase, the power gains from adjusting for multiple covariates become larger for all methods and randomization approaches, which is consistent with the findings from adjusting for a single covariate. The impact of randomization approach on power gains is more pronounced when adjusting for multiple covariates compared to adjusting for a single covariate. In general, SBR and CAR using
method yield higher power gains than simple randomization for all methods across all scenarios, particularly when the covariate effects are moderate to large.
Table 4.
Average power gains from dual covariate adjustment for different statistical models and randomization approaches across various covariate effects
| Method | Sex effect | Site effect | Average power gain | ||
|---|---|---|---|---|---|
| SR | SBR | CAR | |||
| ANOVA-post | Small | Small | −0.9% | 1.2% | −0.2% |
| Moderate | Moderate | 1.0% | 4.4% | 2.8% | |
| Large | Large | 10.1% | 18.7% | 16.2% | |
| ANOVA-change | Small | Small | 0.0% | 1.4% | 1.1% |
| Moderate | Moderate | 5.1% | 7.6% | 6.8% | |
| Large | Large | 27.0% | 35.9% | 32.4% | |
| ANCOVA-hom | Small | Small | 0.0% | 1.5% | 1.1% |
| Moderate | Moderate | 5.4% | 7.9% | 7.0% | |
| Large | Large | 26.8% | 36.0% | 33.0% | |
| ANCOVA-het | Small | Small | 0.0% | 1.3% | 1.4% |
| Moderate | Moderate | 5.4% | 7.9% | 7.0% | |
| Large | Large | 26.3% | 36.3% | 33.2% | |
| LMM | Small | Small | 0.0% | 1.4% | 1.1% |
| Moderate | Moderate | 5.1% | 7.6% | 6.8% | |
| Large | Large | 27.0% | 35.9% | 32.4% | |
Table 5 shows the average power gain obtained by adjusting for three baseline covariates (sex, number of sites, and severity of disease) across different statistical methods, randomization approaches, and covariate effect sizes. Similar to the results from the two-covariate adjustment, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM consistently achieve greater power gains than ANOVA-post when adjusting for three covariates. The power gains become more substantial as the covariate effect sizes increase, with the impact being most prominent when the covariate effects are large. Furthermore, when adjusting for three covariates, CAR consistently outperforms SBR and simple randomization in terms of attaining higher average power gains.
Table 5.
Average power gains from triple covariate adjustment for different statistical models and randomization approaches across various covariate effects
| Method | Sex effect | Site effect | Severity effect | Average power gain | ||
|---|---|---|---|---|---|---|
| SR | SBR | CAR | ||||
| ANOVA-post | Small | Small | Small | −0.6% | −0.2% | 0.0% |
| Moderate | Moderate | Moderate | 4.0% | 6.1% | 7.6% | |
| Large | Large | Large | 23.0% | 33.1% | 33.5% | |
| ANOVA-change | Small | Small | Small | 1.6% | 1.2% | 2.4% |
| Moderate | Moderate | Moderate | 14.8% | 15.4% | 16.5% | |
| Large | Large | Large | 66.9% | 73.6% | 80.9% | |
| ANCOVA-hom | Small | Small | Small | 1.6% | 1.1% | 2.7% |
| Moderate | Moderate | Moderate | 15.1% | 14.6% | 17.1% | |
| Large | Large | Large | 67.1% | 73.6% | 80.3% | |
| ANCOVA-het | Small | Small | Small | 1.7% | 1.1% | 2.9% |
| Moderate | Moderate | Moderate | 14.9% | 14.7% | 17.3% | |
| Large | Large | Large | 67.0% | 73.6% | 80.7% | |
| LMM | Small | Small | Small | 1.6% | 1.2% | 2.4% |
| Moderate | Moderate | Moderate | 14.8% | 15.4% | 16.5% | |
| Large | Large | Large | 66.9% | 73.6% | 80.9% | |
The selection of the randomization approach plays a pivotal role in optimizing the power gains when adjusting for multiple covariates. As evident from Table 4, SBR and CAR consistently surpass simple randomization in terms of power gains, with SBR showing a slight advantage, especially when the covariate effects are moderate to large. However, as the number of covariates increases from two to three, as shown in Table 5, CAR demonstrates superior performance compared to SBR and simple randomization. This shift in the relative performance of CAR and SBR highlights CAR’s growing effectiveness in managing multiple influential covariates. As the analysis becomes more complex with the inclusion of a greater number of covariates, CAR’s advantages become more significant.
Discussion
In this study, we investigated the impact of adjusting baseline covariates on the performance of statistical methods for analyzing pre-post data by using various randomization approaches. We assessed five methods: ANOVA-post, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM. Adjusting for a single baseline covariate can improve the power of all methods, with the extent of power gains depending on the size of the covariate effect and the randomization approach employed. Notably, ANOVA-change, LMM, and the ANCOVA methods consistently achieve higher average power gains from covariate adjustment compared to ANOVA-post. Stratified block randomization and covariate adaptive randomization using
consistently outperform simple randomization in terms of average power gains after adjusting for the covariate.
In scenarios where the sex covariate has no effect (effect size of 0), we observed that simple randomization and stratified block randomization had slightly negative average power gains. Conversely, covariate adaptive randomization using
had average power gains close to 0%. These findings suggest that when the covariate has no impact on the outcome, using simple randomization or SBR may lead to a small decrease in power when adjusting for the covariate. However, CAR appears to be more robust in this scenario by maintaining the power unchanged. Therefore, CAR may be preferable when there is uncertainty about the effect of the covariate on the outcome, as it minimizes the potential negative impact on power.
When multiple baseline covariates are considered, adjusting for all relevant covariates leads to even more substantial power improvements, particularly when the covariate effects are moderate to large. Across all scenarios, ANOVA-change, ANCOVA-hom, ANCOVA-het, and LMM consistently demonstrate higher power gains than ANOVA-post. The impact of the randomization approach on power gains becomes increasingly evident when adjusting for multiple covariates, with SBR and CAR using
yielding higher power gains than simple randomization. However, the relative performance of SBR and CAR varies depending on the complexity of the analysis. As the number of covariates increases, CAR tends to emerge as the superior choice, demonstrating its effectiveness in handling a larger number of influential covariates simultaneously. These results underscore the importance of incorporating relevant covariates and selecting appropriate randomization methods to enhance the efficiency of the statistical methods for analyzing pre-post data.
For our study, we chose to employ the
method for CAR and selected an optimal parameter value of
. This choice was informed by the findings in the literature [24]: a t value near 0.8 provides an optimal balance between the imbalance score and allocation predictability. Here, allocation predictability (AP) refered to the likelihood of correctly predicting treatment assignments based on the imbalance score, given that the investigator consistently predicts the treatment with the higher allocation probability [25, 26]. It is an essential metric for assessing the randomness of the allocation process, and a lower AP value signifies a more unpredictable randomization procedure. The
method was preferred over the
and
methods due to its unique approach of using the quantity rather than the rank order of the total imbalance score, thus can better protect allocation predictability. Shan et al. [24] demonstrated that the
method results in significantly lower allocation predictability compared to the
and
methods, and this advantage is maintained across varying sample sizes. This reduction in allocation predictability is particularly crucial in trials susceptible to selection bias, where preserving a high level of allocation randomness is critical to protect the trial’s integrity. Moreover, the
method, especially with a high r value (0.8 or above), has been shown to achieve greater statistical power compared to the
and
methods. Additionally, we opted for the standard deviation (SD) method over the variance and range methods for imbalance measurement. The SD method provides a better balance between imbalance and allocation predictability, and it generally outperforms the range method in reducing allocation predictability values.
One limitation of this simulation study is the specific configurations and assumptions underlying the setup of the post-test response models and the choices for the associated coefficients. While these choices were made to represent a range of scenarios, they may not fully capture the complexities and variations that can arise in real-world clinical trials. Specifically, the setup of the post-treatment responses was based on a linear model with certain choices for the coefficients, assuming a linear relationship between the baseline and post-treatment responses, as well as additive effects of the covariates. However, in real-world scenarios, the relationships between variables may be non-linear, and there could be more complex interactions or non-additive effects of covariates. Moreover, the choices of specific parameter values for the coefficients in the linear model, such as setting
to 0 and
to 1, or using particular values to represent small, moderate, and large covariate effects, may not accurately reflect the full range of potential relationships encountered in real-world settings. Consequently, the conclusions drawn from these simulations may not be applied to situations where the underlying assumptions or parameter values differ significantly from those used in the simulations, or where the relationships between variables are more complex or non-linear in nature. This could potentially limit the applicability of the simulation findings to real-world clinical trials. Therefore, future research should consider a wider variety of model configurations, including non-linear relationships, higher-order interactions, and a wider range of covariate distributions and effect sizes, to better understand how different parameter choices impact the performance of the statistical methods. Moreover, examining how these models perform under different assumptions about the distribution of errors and the presence of outliers would provide a more comprehensive evaluation of their robustness and applicability.
Our study focused exclusively on designs with single pre- and post-treatment measurements. While this represents a common scenario in clinical trials, many modern studies employ multiple baseline assessments or extended follow-up periods. The relative performance of the statistical methods we examined may differ in designs with multiple pre-treatment and post-treatment measurements, where methods like LMM could naturally extend to incorporate multiple repeated measures, further leveraging their ability to explicitly model within-subject correlation and variability over time. Future research should extend our comparative framework to these more complex longitudinal designs, considering additional factors such as the correlation structure over time, the timing of treatment effects, and missing data patterns. Nevertheless, our findings for the single pre-post design provide important insights into the fundamental relationships between randomization approaches and statistical methods, which likely extend to more advanced trial designs.
Authors' contributions
The idea of this article was originally developed by GS and SW. XL, GS, and YZ conducted simulation studies. All authors drafter the manuscript and approved the final version.
Funding
Shan’s research is partially supported by grants from the National Institutes of Health: R03AG083207, and R01AG070849.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Wan F. Statistical analysis of two arm randomized pre-post designs with one post-treatment measurement. BMC Med Res Methodol. 2021;21(1):1–16. 10.1186/S12874-021-01323-9/FIGURES/2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Shan G, Ma C. A Comment on Sample Size Calculation for Analysis of Covariance in Parallel Arm Studies. J Biom Biostat. 2014;05(01):184. 10.4172/2155-6180.1000184. [Google Scholar]
- 3.Frison L, Pocock SJ. Repeated measures in clinical trials: Analysis using mean summary statistics and its implications for design. Stat Med. 1992;11(13):1685–704. 10.1002/sim.4780111304. [DOI] [PubMed] [Google Scholar]
- 4.Vickers AJ. Analysis of Variance Is Easily Misapplied in the Analysis of Randomized Trials: A Critique and Discussion of Alternative Statistical Approaches. Psychosom Med. 2005;67(4): 652-655. https://journals.lww.com/psychosomaticmedicine/fulltext/2005/07000/analysis_of_variance_is_easily_misapplied_in_the.23.aspx [DOI] [PubMed]
- 5.O Connell NS, Dai L, Jiang Y, Speiser JL, Ward R, Wei W, et al. Methods for analysis of pre-post data in clinical research: a comparison of five common methods. J Biom Biostat. 2017;8(1):1–8. 10.4172/2155-6180.1000334. [DOI] [PMC free article] [PubMed]
- 6.Stevens J, et al. Applied multivariate statistics for the social sciences, vol. 4. NJ: Lawrence Erlbaum Associates Mahwah; 2002. [Google Scholar]
- 7.Dimitrov DM, Rumrill Phillip DJ. Pretest-posttest designs and measurement of change. Work. 2003;20:159–65. [PubMed] [Google Scholar]
- 8.Vickers AJ, Altman DG. Analysing controlled trials with baseline and follow up measurements. BMJ. 2001;323(7321):1123–4. 10.1136/bmj.323.7321.1123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Van Breukelen GJP. ANCOVA versus change from baseline had more power in randomized studies and more bias in nonrandomized studies. J Clin Epidemiol. 2006;59(9):920–5. 10.1016/j.jclinepi.2006.02.007. [DOI] [PubMed] [Google Scholar]
- 10.Chen X. The adjustment of random baseline measurements in treatment effect estimation. J Stat Plan Infer. 2006;136(12):4161–75. 10.1016/j.jspi.2005.08.046. [Google Scholar]
- 11.Yang L, Tsiatis AA. Efficiency Study of Estimators for a Treatment Effect in a Pretest-Posttest Trial. Am Stat. 2001;55(4):314–21. 10.1198/000313001753272466. [Google Scholar]
- 12.Winkens B, van Breukelen GJP, Schouten HJA, Berger MPF. Randomized clinical trials with a pre- and a post-treatment measurement: Repeated measures versus ANCOVA models. Contemp Clin Trials. 2007;28(6):713–9. 10.1016/j.cct.2007.04.002. [DOI] [PubMed] [Google Scholar]
- 13.Verbeke G, Molenberghs G. Linear mixed models for longitudinal data. New York: Springer; 2000.
- 14.Liang KY, Zeger SL. Longitudinal data analysis of continuous and discrete responses for pre-post designs. Sankhyā: Indian J Stat Ser B. 2000;62(1):134–148. https://www.jstor.org/stable/25053123?seq=1.
- 15.Kahan BC, Jairath V, Doré CJ, Morris TP. The risks and rewards of covariate adjustment in randomized trials: An assessment of 12 outcomes from 8 studies. Trials. 2014;15(1):1–7. 10.1186/1745-6215-15-139/FIGURES/2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Schulz KF, Grimes DA. Generation of allocation sequences in randomised trials: chance, not choice. Lancet. 2002;359(9305):515–9. 10.1016/S0140-6736(02)07683-3. [DOI] [PubMed] [Google Scholar]
- 17.Kang M, Ragan BG, Park JH. Issues in outcomes research: An overview of randomization techniques for clinical trials. 2008. 10.4085/1062-6050-43.2.215. [DOI] [PMC free article] [PubMed]
- 18.Nguyen TL, Collins GS, Lamy A, Devereaux PJ, Daurès JP, Landais P, et al. Simple randomization did not protect against bias in smaller trials. J Clin Epidemiol. 2017;84:105–13. [DOI] [PubMed] [Google Scholar]
- 19.Lin KA, Choudhury KR, Rathakrishnan BG, Marks DM, Petrella JR, Doraiswamy PM. Marked gender differences in progression of mild cognitive impairment over 8 years. Alzheimers Dement Transl Res Clin Interv. 2015;1(2):103–10. 10.1016/j.trci.2015.07.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Funatogawa I, Funatogawa T. Longitudinal analysis of pre- and post-treatment measurements with equal baseline assumptions in randomized trials. Biom J Biom Z. 2020;62(2):350–60. 10.1002/bimj.201800389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Taves DR. Minimization: a new method of assigning patients to treatment and control groups. Clin Pharmacol Ther. 1974;15(5):443–53. [DOI] [PubMed] [Google Scholar]
- 22.Pocock SJ, Simon R. Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial. Biometrics. 1975;31(1):103–15. 10.2307/2529712. [PubMed] [Google Scholar]
- 23.Coart E, Bamps P, Quinaux E, Sturbois G, Saad ED, Burzykowski T, et al. Minimization in randomized clinical trials. Stat Med. 2023. 10.1002/SIM.9916. [DOI] [PubMed]
- 24.Shan G, Lu X, Li Z, Caldwell JZK, Bernick C, Cummings J. ADSS: a composite score to detect disease progression in Alzheimer’s disease. J Alzheimers Dis Rep. 2024;8(1):307–16. 10.3233/ADR-230043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Atkinson AC. The comparison of designs for sequential clinical trials with covariate information. J R Stat Soc Ser A Stat Soc. 2002;165(2):349–73. [Google Scholar]
- 26.Heritier S, Gebski V, Pillai A. Dynamic balancing randomization in controlled clinical trials. Stat Med. 2005;24(24):3729–41. 10.1002/sim.2421. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No datasets were generated or analysed during the current study.





