Skip to main content
BMC Medical Research Methodology logoLink to BMC Medical Research Methodology
. 2026 Jan 2;26:20. doi: 10.1186/s12874-025-02758-0

Cluster minimal sufficient balance (CMSB): an efficient covariate balancing randomization method for cluster randomized trials

Jiaxin Cai 1,2,3, Shanshan Suo 3, Valirie Ndip 5, Carole Khairallah 5, Binyan Zhang 6, Lingxia Zeng 3, Hong Yan 3, Fang Shao 4,✉, Tao Chen 5, Chao Li 3,✉
PMCID: PMC12857073  PMID: 41484830

Abstract

Background

Cluster randomized trials (CRTs) require balanced baseline covariates to yield unbiased estimates of treatment effects. Existing approaches such as constrained randomization can improve balance but may compromise allocation randomness. We introduce Cluster Minimal Sufficient Balance (CMSB), a cluster randomization method designed to enhance covariate balance while preserving allocation randomness and computational efficiency.

Methods

CMSB integrates dynamic imbalance monitoring with conditional biased randomization into a single procedure. The method accommodates both continuous and categorical covariates and was evaluated through simulation studies comparing its performance with constrained randomization, simple randomization, block randomization, stratified randomization, and minimization across varying numbers of clusters and covariate dimensions. An empirical application further assessed its practical utility.

Results

In high-dimensional settings with 10 covariates, CMSB achieved a 45% greater improvement in covariate balance compared to constrained randomization, while maintaining near-optimal allocation randomness (correct guess probability: 49.79% versus 61.03% for minimization). CMSB reduced mean allocation time from over 100 s under constrained randomization to 0.116 s per allocation, even when simulating up to 2,000 randomization schemes. In the empirical application, CMSB demonstrated a 76% improvement in covariate balance relative to simple randomization.

Conclusions

CMSB effectively balances the competing demands of covariate balance, allocation randomness, and computational efficiency in cluster randomized trials. Its ability to handle high-dimensional covariates makes it particularly suitable for large-scale trials.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12874-025-02758-0.

Keywords: Cluster randomized trial, Covariates balance, Randomization, Cluster minimal sufficient balance, Computational efficiency

Introduction

Randomized controlled trials (RCTs) represent the gold standard for evaluating therapeutic interventions against standard care, placebo, or no treatment [1–3]. A fundamental principle for ensuring an unbiased estimation of the treatment effect is achieving baseline covariate balance between treatment arms—a criterion particularly challenging to meet in trials conducted within naturally occurring groupings such as schools, primary care practices, or communities [4, 5].

Cluster randomization has emerged as a key methodological framework for addressing the limitations of individual-level randomization in complex intervention research, where intact groups (clusters) rather than individuals are allocated to intervention and control arms [6, 7]. This approach offers several advantages: it preserves intervention integrity by minimizing cross-arm contamination, enhances logistical feasibility through centralized implementation, and enables evaluation of inherently cluster-level interventions, including community-wide health promotion programs and policy-based initiatives [8, 9]. Over the past three decades, the use of cluster-randomized designs has grown substantially in health research, driving methodological innovation and establishing cluster randomized trials (CRTs) as essential tools for population-level intervention studies.

Achieving and maintaining baseline covariate balance is a critical methodological imperative in CRTs, requiring equivalence across cluster-level prognostic factors [10]. Systematic reviews indicate that approximately 8% of published CRTs exhibit baseline imbalances, with prevalence rising to ≥ 25% in trials using ≤ 4 clusters per arm [11]. In the seminal cluster randomized trial by Brinkman et al. (2016) [12], designed to prevent teenage pregnancy, a baseline imbalance in socioeconomic status was observed between intervention and control groups. The control arm included a higher proportion of adolescents from socioeconomically advantaged backgrounds—a well-established protective factor against teenage pregnancy. This imbalance in a key prognostic covariate may explain the paradoxical finding of higher pregnancy rates in the intervention group (17% vs. 11%; RR 1.55, 95% CI 1.21–1.98). These findings collectively demonstrate that failure to maintain baseline balance in CRTs can distort treatment effect estimates, potentially compromising internal validity and conclusions about intervention efficacy [13].

Traditional randomization methods such as stratified block randomization are limited by their applicability primarily to categorical covariates and often compromise allocation randomness [14]. In CRTs, restricted or constrained randomization has long been used as a design strategy to improve balance on prognostic cluster-level covariates. Notable approaches include covariate-constrained randomization [15–17], minimization methods [18], and cluster-specific constrained designs [19]. While these constrained methods improve baseline balance, they involve trade-offs: computational burden increases with number of clusters and covariate dimensionality, and overly restrictive allocation schemes may reduce allocation randomness, potentially affecting the validity of statistical inference [20].

To address these challenges, we introduce the Cluster Minimal Sufficient Balance (CMSB) as an integrated randomization framework for CRTs. CMSB builds upon the concept of Minimal Sufficient Balance (MSB) [14] and extends it to cluster settings, embedding balance considerations directly into the randomization mechanism. Importantly, CMSB simultaneously reduces imbalances in cluster-level covariates—encompassing both inherent cluster characteristics (e.g., Gross Domestic Product) and aggregated individual-level measures (e.g., mean baseline blood pressure)—while preserving allocation randomness. This integration of design and allocation aims to achieve robust covariate balance, maintain randomization integrity, and ensure computational efficiency across diverse trial sizes and covariate dimensions.

The primary contributions of this work are threefold. First, CMSB demonstrates superior performance in achieving covariate balance compared to established methods—including constrained randomization, simple randomization, block randomization, stratified randomization, and minimization—across a range of number of clusters and covariate dimensionalities. Second, through an adaptive thresholding mechanism, CMSB maintains allocation randomness within a single-stage randomization process while remaining adaptable to sequential enrollment and cluster substitutions. Third, CMSB exhibits favorable computational performance, enabling scalable application to large and complex CRTs without sacrificing interpretability or reproducibility. Its streamlined framework supports effective implementation as cluster numbers and covariate dimensions increase.

The remainder of the article is structured to systematically validate these contributions. Section Methods provides a formal algorithmic description of CMSB. Section Simulation study presents results from extensive simulation studies benchmarking CMSB against competing methods across key evaluation metrics. In Sect. Empirical study, we illustrate the practical application of CMSB through empirical analysis of data from a real-world CRT. Section Discussion discusses the broader implications of our findings, and Sect. Conclusion offers concluding remarks.

Methods

Cluster minimal sufficient balance

To address these challenges, we adapt and evaluate the CMSB for CRTs—an MSB-inspired extension that embeds balance considerations directly within the randomization mechanism while preserving allocation randomness. This extension enhances the original MSB in the context of CRTs by providing an efficient framework for handling both continuous and categorical covariates. As a key modification, Fisher’s exact test is incorporated for categorical covariates when the number of clusters is low, ensuring valid imbalance assessment across all trial sizes. This approach has the following distinctive features:

  1. Cluster-level covariate balancing

    CMSB prioritizes balancing important cluster-level covariates that may influence trial outcomes while ensuring the randomness of treatment allocation. CMSB analyzes all covariates in their original measurement scales, whether continuous or categorical, cluster-aggregated or not. This dual focus on balance and randomness distinguishes CMSB from traditional cluster randomization methods that often sacrifice one for the other [14]. Covariate balance refers to the similarity in the distribution of baseline covariates between treatment arms. An ideal randomization ensures comparable joint and marginal distributions of covariates across arms. Imbalance occurs when these distributions differ systematically, potentially introducing confounding bias into the treatment effect estimate.

    Allocation randomness quantifies the unpredictability of the treatment assignment sequence. A perfectly random procedure, such as a fair coin toss, ensures that the probability of assigning a cluster to either arm is 0.5, independent of prior allocations or covariate values.

  2. Original scale analysis

    Stratified randomization is restricted to categorical variables and requires arbitrary discretization of continuous covariates, which may result in information loss and potential bias [21]. In contrast, CMSB analyzes all covariates in their original measurement scales, eliminating the need for arbitrary categorization and preventing information loss due to dichotomization or grouping of continuous measures. This enhances sensitivity in detecting imbalance and improves the efficiency of maintaining balance across treatment groups.

  3. Dynamic imbalance monitoring

    Covariate imbalances are continuously assessed throughout the CMSB process using appropriate statistical measures:
    • For continuous covariates: Two -sample t-tests comparing means across groups.
    • For categorical covariates: Pearson’s χ² tests or Fisher’s exact tests.

    These statistical measures serve solely as quantitative indicators (p values) of imbalance during the randomization process and are not used as formal hypothesis tests for inferential conclusions about baseline characteristics.

    Although termed “Dynamic Imbalance Monitoring”, this approach remains applicable in CRTs where all clusters are available for randomization simultaneously. This flexibility allows researchers to evaluate covariate balance in real time, ensuring that any imbalances can be addressed before final treatment allocation, thereby preserving trial design integrity.

  4. Conditional biased randomization

    The method employs a biased coin mechanism with probability ξ only when two strict conditions are met:
    • i)
      One or more covariates exceed pre-specified imbalance thresholds (e.g., p < 0.3 [14] based on statistical tests).
    • ii)
      Assigning the incoming cluster to a specific arm would substantially improve covariate balance across all covariates.
    This dual criterion ensures maximal unpredictability in allocation when:
    • Covariates are sufficiently balanced
    • The current clusters’ characteristics are intermediate between arms
    • Correcting one imbalance would worsen others

Otherwise, the method defaults to simple randomization (ξ = 0.5) to preserve allocation randomness.

The CMSB algorithm uses a sequential dynamic allocation strategy. When a new cluster becomes available for randomization, it first evaluates the current imbalance across pre-specified cluster-level covariates between treatment arms.

Rather than deterministically assigning clusters to correct imbalances—which could compromise allocation randomness—CMSB employs a probabilistic voting system. For each imbalanced covariate, a vote is cast for the treatment arm that would improve balance upon assignment of the cluster. The voting rule varies by variable type. For continuous covariates, the decision is based on the current cluster’s covariate value relative to the mean values among already-assigned clusters, enabling direct correction of existing imbalances. For categorical covariates, the vote depends on both the distribution of categories across previously assigned clusters and the current cluster’s category level.

  1. For continuous covariates, the standardized mean difference between arms for covariate k is computed as:
    graphic file with name d33e487.gif 1

    where Inline graphic and Inline graphic denote the mean and variance of covariate k in arm a (Inline graphic). These summary statistics are calculated as unweighted means across clusters within each arm, ensuring equal weighting of clusters irrespective of their individual sample sizes.

    The voting is as follows:
    graphic file with name d33e513.gif 2

    Where Inline graphic is the value of covariate k for the current unallocated cluster, δ is the standardized difference threshold, derived from a predefined p -value (e.g., 0.3), and the degrees of freedom by referencing the t-distribution table.

  2. For categorical covariates in cluster randomization trials, let Zₖ denote the k-th categorical covariate with G levels (e.g., hospital type with G categories), observed at the cluster level. The allocation procedure proceeds as follows:

    Let Inline graphic represent the expected number of clusters in the l category of Zk to be allocated to the treatment and control groups, respectively.

    The chi-square statistic is
    graphic file with name d33e567.gif 3

Based on the chi-square distribution with (G − 1) degrees of freedom, the covariate is flagged as imbalanced if its associated p value,Inline graphic, falls below the significance threshold α (commonly set at 0.3).

Fisher’s exact test is applied when the number of clusters is fewer than 40 [22]. The probability of the observed allocation is computed using the hypergeometric distribution according to the following formula:

graphic file with name d33e591.gif 4

where Inline graphic denotes the total number of clusters in category l*, which is the category level of covariate Zₖ for the current unallocated cluster. Inline graphic represents the total number of clusters in arm a (a = T or a = C), which is derived by summing over all categories of Zₖ, and N denotes the total number of clusters. The corresponding p value,Inline graphic, is calculated as the sum of probabilities for all possible allocation outcomes at least as extreme as the observed result.

The voting rule is as follows:

graphic file with name d33e646.gif 5

In the “Neutral” state, the covariate abstains from voting. The final allocation decision for the current cluster is made only after all covariates have voted, based on whether one treatment arm receives a majority of votes across both continuous and categorical types. This design ensures that the algorithm avoids correcting one covariate’s imbalance at the expense of worsening another, thereby maintaining a global balance perspective. The “Neutral” outcome serves as a critical mechanism for preserving allocation randomness when the optimal balancing action is unclear. After evaluating all baseline covariate imbalances, the final assignment probability integrates these votes, with dominance established when one treatment arm accumulates more votes than the other:

graphic file with name d33e652.gif 6

with ξ typically set to 0.8 [23] to balance randomness and imbalance control.

Although CMSB is described as a sequential procedure, it is directly applicable to static allocation designs in which all clusters are randomized simultaneously. In such settings, clusters are first randomly ordered and then sequentially allocated according to the CMSB mechanism. This approach retains dynamic imbalance monitoring and corrective allocation within a single randomization sequence. The overall randomization scheme for CMSB is illustrated in Fig. 1, with a detailed flowchart provided in Supplementary Fig. 1.

Fig. 1.

Fig. 1

The flowchart of CMSB

Hypothesis testing under cluster-minimal sufficient balance

The CMSB randomization method introduces sequential dependencies through its dynamic allocation mechanism. We employ Linear Mixed Models (LMMs) for statistical inference. The validity of LMMs lies in their ability to model the primary source of non-independence in CRTs: intra-cluster correlation. By incorporating cluster-level random effects, LMMs account for correlated outcomes within clusters, thereby enabling valid inference. Importantly, the core objective of CMSB is to achieve robust balance on cluster-level covariates between treatment arms. This balance minimizes systematic baseline differences and mitigates confounding bias. Notably, the combination of LMMs’ capacity to model CRT correlation structures and CMSB’s achievement of baseline balance justifies the use of LMMs for valid estimation of treatment effects.

To test the null hypothesis of no treatment effect, we specify the following LMM:

graphic file with name d33e689.gif 7

where Inline graphic is the outcome for individual i in cluster j, Inline graphic is the treatment assignment indicator for cluster j, Inline graphic is the fixed intercept, representing the mean outcome in the control group, and Inline graphic is the fixed treatment effect, which is the primary parameter of interest. The term Inline graphic denotes the vector of fixed effects corresponding to covariates, and Inline graphic represents the vector of cluster-level baseline covariates. Additionally, Inline graphic is the random intercept for cluster j, and Inline graphic is the individual-level residual error.

The null hypothesis of no treatment effect is Inline graphic, tested against the alternative Inline graphic.

The significance of the treatment effect Inline graphic is assessed using either the Wald F-test or the Wald χ2-test. Both testing procedures are implemented within the linear mixed model framework following CMSB randomization to evaluate the statistical significance of the treatment effect.

Simulation study

Simulation settings

A comprehensive simulation study was conducted to evaluate the performance of the CMSB method under diverse experimental conditions. The simulation framework employed a factorial design incorporating two key parameters: (1) number of clusters (10, 30, or 50 per arm), and (2) covariate dimensionality, with scenarios spanning low -dimensional (4 covariates) and high -dimensional (10 covariates) settings. This design enabled a robust assessment of the method’s performance across a range of trial configurations, reflecting conditions from small-scale pilot studies to large multicenter trials.

Cluster-level covariates, including both binary and continuous variables, were generated to reflect realistic clinical trial settings. In the low dimension setting, binary covariates Inline graphic were simulated from Bernoulli distributions with success probabilities of 0.3 and 0.4, respectively, while continuous covariates Inline graphic followed uniform distributions Inline graphic. The outcome variable Y was generated according to the linear mixed model:

graphic file with name d33e782.gif 8

where Inline graphic represents cluster-level random effects, Inline graphic denotes individual-level errors, and Inline graphic indicates the treatment assignment. The intraclass correlation coefficient (ICC) was set to 0.01, with Inline graphic and Inline graphic adjusted accordingly to maintain the desired variance components. The treatment effect size δ was specified as 0 under the null hypothesis and 0.5 under the alternative.

For high-dimensional settings, binary covariates Inline graphic were generated from independent Bernoulli distributions with success probability 0.2, while continuous covariates Inline graphic followed uniform distributions Inline graphic. The outcome variable Y was simulated based on the following linear mixed model:

graphic file with name d33e829.gif 9

The CMSB procedure evaluates balance for each covariate using a two-sample test. For each covariate, a single vote is cast in favor of the treatment arm whose assignment would improve balance on that covariate. If one arm receives a strict majority of votes, the current cluster is assigned to that arm with biased-coin probability ξ; otherwise, in the case of a tie, the cluster is allocated with equal probability (representing an overall ‘Neutral’ state). For comparative evaluation, we implemented five alternative randomization methods: simple randomization, block randomization, stratified randomization (using covariates x1 and x2 as stratification factors), minimization randomization [18], and constrained randomization [19]. Notably, the constrained randomization method randomly simulated 2,000 randomization schemes by default when the total number of possible schemes exceeded this threshold, rather than exhaustively generating all potential allocations, to reduce computational burden and accelerate the process. Each scenario was replicated 1,000 times to assess covariate balance performance. All simulations were conducted in R 4.4.2 using an Intel Core i9-14900 K processor and 32 GB RAM (Table 1).

Table 1.

Simulation settings of six scenario

Simulation Scenario Number of Clusters (per arm) Covariate Dimension Covariate Type
Scenario 1 10 4 2 Binary + 2 Continuous
Scenario 2 30 4 2 Binary + 2 Continuous
Scenario 3 50 4 2 Binary + 2 Continuous
Scenario 4 10 10 5 Binary + 5 Continuous
Scenario 5 30 10 5 Binary + 5 Continuous
Scenario 6 50 10 5 Binary + 5 Continuous

Simulation results

The performance of the proposed CMSB method was benchmarked against established randomization approaches across three primary domains: (a) covariate balance, (b) computational efficiency, and (c) allocation randomness. Additionally, hypothesis testing performance was evaluated in terms of (d) Type I error control and statistical power.

  1. Covariate balance across randomization methods

    The effectiveness of each randomization method in maintaining covariate balance was assessed using imbalance ratios and distributions of p values. The imbalance ratio is defined as the total number of imbalance events across all covariates divided by the product of the number of simulation replicates and the number of covariates. An imbalance event occurs when a statistically significant imbalance (p < 0.05), as determined by a significance test, is observed for a given covariate between the two arms.

    Tables 2, 3, 4, 5, 6 and 7 summarize the performance of different randomization methods under low-dimensional (4 covariates) and high-dimensional (10 covariates) scenarios, with varying cluster sizes (10, 30, 50):
    • Low-dimensional setting
      In low-dimensional scenarios (Simulations 1–3 with 10, 30, 50 clusters), CMSB demonstrated performance comparable to constrained randomization at the cluster level. Figure 2 displays the mean covariate imbalance ratio, calculated as the number of times imbalance occurred across 1,000 simulations for different cluster randomization methods in the low-dimensional setting. As shown in Table 2, for 10 clusters, CMSB achieved imbalance ratios of 0.128–0.156 compared to 0.115–0.143 for constrained randomization, while maintaining mean p values of 0.591–0.651 versus 0.601–0.651 for constrained randomization. Minimization showed superior performance for categorical variables but exhibited substantial variability for continuous variables.
      Increasing the cluster size to 30 improved CMSB’s performance, reducing imbalance ratios to 0.057–0.101 compared to 0.135–0.167 for constrained randomization. The advantage became more pronounced at 50 clusters, where CMSB achieved imbalance ratios of 0.041–0.050 compared to 0.127–0.155 for constrained randomization. Simple and block randomization consistently performed worst across all cluster sizes, while stratified randomization showed good balance for stratified variables but poor performance for others (up to 0.295).
    • Low-dimensional set
      In high-dimensional scenarios (Simulations 4–6 with 10, 30, 50 clusters), CMSB demonstrated clear superiority over all comparison methods. Figure 3 presents the mean covariate imbalance ratio across different cluster randomization methods in the high-dimensional setting.
      With 10 clusters, CMSB maintained imbalance ratios of 0.194–0.259 compared to 0.175–0.238 for constrained randomization. The performance of minimization deteriorated substantially in high dimensions, with imbalance ratios reaching up to 0.311 for continuous variables. When cluster size increased to 30, CMSB’s advantage became more evident, achieving 22–35% lower imbalance ratios than constrained randomization. As shown in Table 7, at 50 clusters, CMSB achieved a 45% improvement in covariate balance across all variables relative to constrained randomization, with imbalance ratios of 0.073–0.119 versus 0.177–0.227, respectively.
      CMSB maintained more stable p -value distributions, with smaller discrepancies between median and mean values, indicating greater robustness. Although minimization demonstrated excellent balance for categorical covariates, its performance markedly declined for continuous covariates. Similarly, stratified randomization showed limited effectiveness for non-stratified variables. Both simple and block randomization consistently underperformed across all metrics (imbalance ratios ranging from 0.253 to 0.327), confirming their inadequacy for trials requiring stringent covariate balance control.
  2. Computational efficiency and complexity

    Among all methods, constrained randomization demonstrated the most competitive covariate balancing after CMSB. However, it incurred the highest computational cost in scenarios with 50 clusters and 10 covariates (Table 8). Constrained randomization suffers from prohibitively high computational demands, requiring over 101.177 s per allocation when simulating up to 2,000 randomization schemes.

    In contrast, CMSB outperformed constrained randomization in complex scenarios—particularly those involving large cluster sizes and high-dimensional settings—with a mean allocation time of 0.116 s. This represents a greater than 99.9% improvement in computational speed while simultaneously delivering superior balancing performance. Simple, block, and stratified randomization were computationally efficient but provided inferior balance control, limiting their practical utility.

  3. Randomness and predictability

    We also evaluated the inherent randomness and unpredictability of the cluster-level randomization methods. Predictability was quantified through computational simulations that attempted to guess the next treatment assignment based on accumulated allocation patterns and covariate imbalances. Following Zhao et al. (2015) [14], we calculate the correct guess probability. CMSB demonstrated a correct guess probability of 49.8%, whereas block randomization exhibited higher predictability, with correct guess probabilities of 66%, 68%, 71%, and 75% for block sizes of 8, 6, 4, and 2, respectively. This comparison indicates that CMSB maintains superior randomness compared to block randomization across all block size configurations.

  4. Statistical inference under CMSB randomization in cluster randomized controlled trials

    Treatment effects were evaluated using Wald F-tests and Chi-square tests, with the null hypothesis stating no treatment effect. Statistical performance was assessed over 1,000 simulation replicates, examining both Type I error rates and power characteristics of CMSB. Comparative analyses using Wald F-tests [24] and Chi-square tests [25] are summarized in Table 9, assuming an effect size of 0.5. Both the F-test and Chi-square test effectively controlled Type I error rates near the nominal 0.05 level across various experimental conditions, with a margin of error of approximately ± 0.015. Statistical power approached 1.000 in scenarios with larger numbers of clusters, indicating reliable detection of the treatment effect. Moreover, as shown in Supplementary Tables 1 and 2, we evaluated different cluster sizes (50, 100), and the results indicate that larger sizes enhance statistical power while maintaining Type I error rates close to the nominal level.

Table 2.

Covariate balance metrics across different cluster randomization methods in low dimension (10 clusters), simulation data

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.164 0.128 0.313 0.317 0.015 0.023
Mean p value 0.591 0.651 0.515 0.501 0.771 0.775
Median p value 0.606 0.660 0.557 0.557 0.807 0.860
x2 Imbalance ratio 0.156 0.115 0.267 0.269 0.004 0.007
Mean p value 0.596 0.644 0.509 0.525 0.777 0.797
Median p value 0.606 0.660 0.398 0.628 0.732 0.860
x3 Imbalance ratio 0.161 0.136 0.314 0.298 0.308 0.304
Mean p value 0.587 0.601 0.488 0.490 0.493 0.505
Median p value 0.605 0.630 0.463 0.487 0.496 0.511
x4 Imbalance ratio 0.175 0.143 0.324 0.293 0.305 0.162
Mean p value 0.576 0.619 0.488 0.513 0.500 0.599
Median p value 0.590 0.644 0.470 0.531 0.506 0.626

Table 3.

Covariate balance metrics across different randomization methods in low dimension (30 clusters)

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.068 0.156 0.313 0.335 < 0.001 < 0.001
Mean p value 0.634 0.630 0.514 0.496 0.832 0.856
Median p value 0.652 0.599 0.549 0.443 0.858 0.898
x2 Imbalance ratio 0.056 0.167 0.321 0.305 < 0.001 < 0.001
Mean p value 0.640 0.607 0.497 0.514 0.853 0.876
Median p value 0.652 0.605 0.445 0.498 0.831 0.915
x3 Imbalance ratio 0.076 0.135 0.287 0.313 0.286 0.312
Mean p value 0.629 0.625 0.512 0.491 0.509 0.500
Median p value 0.636 0.656 0.521 0.487 0.516 0.498
x4 Imbalance ratio 0.101 0.151 0.298 0.300 0.295 0.077
Mean p value 0.609 0.604 0.507 0.501 0.501 0.665
Median p value 0.615 0.627 0.496 0.496 0.495 0.697

Table 4.

Covariate balance metrics across different randomization methods in low dimension (50 clusters)

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.044 0.155 0.331 0.310 < 0.001 < 0.001
Mean p value 0.654 0.621 0.499 0.506 0.863 0.891
Median p value 0.665 0.660 0.504 0.513 0.892 0.924
x2 Imbalance ratio 0.044 0.127 0.291 0.292 < 0.001 < 0.001
Mean p value 0.652 0.619 0.496 0.501 0.886 0.900
Median p value 0.657 0.681 0.534 0.539 0.877 0.931
x3 Imbalance ratio 0.041 0.132 0.295 0.303 0.315 0.259
Meanp value 0.659 0.611 0.502 0.501 0.495 0.523
Median p value 0.668 0.635 0.494 0.491 0.491 0.535
x4 Imbalance ratio 0.050 0.155 0.291 0.295 0.312 0.051
Mean p value 0.647 0.604 0.508 0.511 0.482 0.694
Median p value 0.665 0.640 0.507 0.516 0.478 0.724

Table 5.

Covariate balance metrics across different randomization methods in high dimension (10 clusters)

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.225 0.198 0.321 0.304 0.014 0.085
Mean p value 0.565 0.595 0.501 0.506 0.763 0.700
Median p value 0.589 0.628 0.557 0.557 0.785 0.660
x2 Imbalance ratio 0.198 0.200 0.237 0.288 0.005 0.050
Mean p value 0.572 0.572 0.542 0.498 0.786 0.731
Median p value 0.604 0.644 0.628 0.398 0.785 0.673
x3 Imbalance ratio 0.259 0.238 0.339 0.324 0.287 0.106
Mean p value 0.518 0.579 0.504 0.516 0.522 0.671
Median p value 0.512 0.557 0.557 0.557 0.512 0.628
x4 Imbalance ratio 0.201 0.175 0.261 0.278 0.290 0.030
Mean p value 0.571 0.577 0.536 0.510 0.513 0.733
Median p value 0.604 0.660 0.660 0.398 0.512 0.673
x5 Imbalance ratio 0.199 0.182 0.308 0.274 0.321 0.035
Mean p value 0.565 0.597 0.489 0.514 0.503 0.749
Median p value 0.589 0.660 0.398 0.557 0.512 0.673
x6 Imbalance ratio 0.232 0.193 0.302 0.294 0.292 0.311
Mean p value 0.549 0.567 0.497 0.493 0.503 0.494
Median p value 0.559 0.595 0.497 0.478 0.507 0.478
x7 Imbalance ratio 0.213 0.192 0.294 0.308 0.320 0.292
Mean p value 0.550 0.572 0.504 0.498 0.496 0.501
Median p value 0.556 0.593 0.505 0.495 0.490 0.507
x8 Imbalance ratio 0.234 0.196 0.308 0.317 0.297 0.192
Mean p value 0.544 0.572 0.498 0.490 0.514 0.563
Median p value 0.546 0.581 0.491 0.486 0.526 0.564
x9 Imbalance ratio 0.200 0.179 0.277 0.305 0.291 0.253
Mean p value 0.557 0.581 0.507 0.508 0.505 0.539
Median p value 0.569 0.604 0.504 0.503 0.507 0.554
x10 Imbalance ratio 0.194 0.190 0.319 0.324 0.295 0.300
Mean p value 0.557 0.577 0.478 0.482 0.498 0.494
Median p value 0.565 0.607 0.454 0.474 0.488 0.482

Table 6.

Covariate balance metrics across different randomization methods in high dimension (30 clusters)

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.123 0.208 0.319 0.322 < 0.001 0.013
Mean p value 0.596 0.578 0.503 0.508 0.830 0.802
Median p value 0.588 0.581 0.527 0.549 0.858 0.791
x2 Imbalance ratio 0.119 0.215 0.327 0.318 < 0.001 0.015
Mean p value 0.606 0.577 0.498 0.502 0.857 0.824
Median p value 0.610 0.605 0.447 0.574 0.843 0.799
x3 Imbalance ratio 0.144 0.199 0.319 0.323 0.298 0.031
Mean p value 0.591 0.581 0.495 0.493 0.510 0.775
Median p value 0.608 0.567 0.497 0.497 0.512 0.759
x4 Imbalance ratio 0.101 0.167 0.245 0.259 0.277 0.007
Mean p value 0.616 0.586 0.518 0.507 0.505 0.834
Median p value 0.612 0.612 0.447 0.447 0.464 0.800
x5 Imbalance ratio 0.127 0.256 0.344 0.314 0.288 0.007
Mean p value 0.607 0.557 0.485 0.498 0.510 0.816
Median p value 0.609 0.591 0.434 0.447 0.518 0.795
x6 Imbalance ratio 0.138 0.191 0.278 0.300 0.315 0.328
Mean p value 0.589 0.572 0.514 0.508 0.492 0.486
Median p value 0.604 0.584 0.527 0.520 0.493 0.463
x7 Imbalance ratio 0.140 0.194 0.318 0.306 0.282 0.279
Mean p value 0.597 0.582 0.480 0.498 0.516 0.499
Median p value 0.610 0.625 0.484 0.496 0.528 0.495
x8 Imbalance ratio 0.136 0.225 0.290 0.311 0.304 0.155
Mean p value 0.589 0.556 0.504 0.492 0.505 0.593
Median p value 0.600 0.590 0.501 0.497 0.503 0.612
x9 Imbalance ratio 0.132 0.196 0.292 0.309 0.299 0.280
Mean p value 0.609 0.575 0.503 0.497 0.503 0.515
Median p value 0.623 0.598 0.511 0.495 0.500 0.532
x10 Imbalance ratio 0.141 0.214 0.291 0.289 0.277 0.322
Mean p value 0.591 0.559 0.492 0.503 0.512 0.495
Median p value 0.593 0.581 0.487 0.515 0.513 0.496

Table 7.

Covariate balance metrics across different randomization methods in high dimension (50 clusters)

Variable Metric CMSB Constrained Simple Block Stratified Minirandom
x1 Imbalance ratio 0.088 0.205 0.310 0.329 < 0.001 0.006
Mean p value 0.627 0.579 0.497 0.493 0.867 0.840
Median p value 0.637 0.633 0.504 0.504 0.892 0.834
x2 Imbalance ratio 0.073 0.194 0.268 0.302 < 0.001 0.002
Mean p value 0.634 0.571 0.511 0.499 0.886 0.850
Median p value 0.632 0.551 0.534 0.531 0.874 0.842
x3 Imbalance ratio 0.119 0.183 0.282 0.278 0.293 0.005
Mean p value 0.612 0.591 0.501 0.510 0.499 0.819
Median p value 0.637 0.621 0.481 0.481 0.504 0.814
x4 Imbalance ratio 0.080 0.177 0.274 0.253 0.269 < 0.001
Mean p value 0.633 0.567 0.499 0.522 0.506 0.861
Median p value 0.653 0.553 0.547 0.551 0.539 0.843
x5 Imbalance ratio 0.075 0.194 0.311 0.320 0.294 0.001
Mean p value 0.637 0.579 0.506 0.499 0.506 0.854
Median p value 0.653 0.633 0.525 0.521 0.521 0.839
x6 Imbalance ratio 0.105 0.223 0.309 0.327 0.296 0.315
Mean p value 0.623 0.566 0.487 0.484 0.498 0.484
Median p value 0.634 0.606 0.478 0.492 0.503 0.478
x7 Imbalance ratio 0.098 0.199 0.301 0.307 0.287 0.307
Mean p value 0.620 0.570 0.493 0.501 0.503 0.501
Median p value 0.635 0.578 0.488 0.502 0.505 0.503
x8 Imbalance ratio 0.110 0.216 0.300 0.292 0.290 0.119
Mean p value 0.611 0.575 0.503 0.501 0.503 0.623
Median p value 0.624 0.605 0.498 0.508 0.499 0.652
x9 Imbalance ratio 0.098 0.207 0.296 0.290 0.298 0.233
Mean p value 0.622 0.564 0.499 0.497 0.494 0.533
Median p value 0.624 0.598 0.501 0.481 0.495 0.539
x10 Imbalance ratio 0.119 0.227 0.316 0.278 0.318 0.272
Mean p value 0.608 0.556 0.495 0.497 0.498 0.513
Median p value 0.624 0.560 0.489 0.487 0.499 0.513

Fig. 2.

Fig. 2

Covariates mean imbalance ratio across different cluster randomization methods in low dimension (simulation data). Panel (a), (b), and (c) are results for 10, 30, and 50 clusters, respectively. Note: ‘CMSB’ represents Cluster Minimal Sufficient Balance, ‘Constrained’ represents Constrained Randomization, ‘Simple’ represents Simple Randomization, ‘Block’ represents Block Randomization, ‘Stratified’ represents Stratified Randomization, ‘Minirandom’ represents Minimization Randomization

Fig. 3.

Fig. 3

Covariate mean imbalance ratio across different cluster randomization methods in high dimension (simulation data). Panel (a), (b), and (c) are results for 10, 30, and 50 clusters, respectively. Note: ‘CMSB’ represents Cluster Minimal Sufficient Balance, ‘Constrained’ represents Constrained Randomization, ‘Simple’ represents Simple Randomization, ‘Block’ represents Block Randomization, ‘Stratified’ represents Stratified Randomization, ‘Minirandom’ represents Minimization Randomization

Table 8.

The total running time across 1000 replications of different randomization methods

Randomization Method Running Time (seconds)
CMSB randomization 115.701
Constrained randomization* 101177.000
Simple randomization 0.053
Block randomization 0.212
Stratified randomization 560.045
Minimization randomization 21.096

*The number of randomization schemes to simulate is up to 2,000

Table 9.

Type I error rates and statistical power across different hypothesis testing methods

Covariate dimension Cluster size Type I Error Rate Power
Wald F-test Wald χ2-test Wald F-test Wald χ2-test
4 10 0.041 0.059 0.598 0.655
30 0.045 0.049 0.992 0.992
50 0.055 0.055 1.000 1.000
10 10 0.040 0.057 0.620 0.685
30 0.042 0.047 0.986 0.987
50 0.042 0.044 1.000 1.000

Empirical study

The empirical evaluation of the CMSB method used data from a double-blind cluster randomized controlled trial conducted between August 2002 and January 2006 in Shaanxi Province, Northwest China [26]. The original trial evaluated the impact of antenatal micronutrient supplementation on birth outcomes among 5,828 pregnant women across 531 village clusters. Key covariates included maternal sociodemographic characteristics (e.g., age, education, occupation), household wealth indices, and anthropometric measures, aggregated at the cluster level. For the current analysis, eight cluster-level covariates were selected to reflect these heterogeneities, enabling rigorous evaluation of CMSB’s balancing performance under real-world conditions involving high-dimensional and socioeconomically stratified data.

To evaluate the performance of the six randomization methods, we implemented a nonparametric bootstrap resampling procedure. In each of the 1,000 simulation runs, 100 clusters were resampled with replacement from the original dataset. This framework was chosen to estimate imbalance ratios while accounting for variability in cluster-level characteristics. The evaluation was based on eight key covariates representing heterogeneity across socioeconomic and anthropometric domains.

Figures 4 and 5 illustrate the imbalance ratio for each covariate and the mean imbalance ratio, respectively. CMSB achieved the lowest average imbalance ratio across all covariates (0.073), representing a 76% reduction compared to simple randomization (0.309) and a 57% improvement over constrained randomization (0.170) (Table 10). This superior balancing performance was achieved with modest computational requirements of 0.092 s per randomization. CMSB was approximately 500 times faster than constrained randomization and delivered better balance. This performance advantage was especially evident in scenarios with larger cluster sizes. Minimization, while computationally efficient (0.020 s), showed variable performance across covariate types, performing well for wealth quintile (0.136) but less consistently for other measures.

Fig. 4.

Fig. 4

Different Covariates Imbalance Ratios across Randomization Methods (empirical data) Note: ‘CMSB’ represents Cluster Minimal Sufficient Balance, ‘Constrained’ represents Constrained Randomization, ‘Simple’ represents Simple Randomization, ‘Block’ represents Block Randomization, ‘Stratified’ represents Stratified Randomization, ‘Minirandom’ represents Minimization Randomization

Fig. 5.

Fig. 5

Mean Imbalance Ratios across Randomization Methods (empirical data) Note: ‘CMSB’ represents Cluster Minimal Sufficient Balance, ‘Constrained’ represents Constrained Randomization, ‘Simple’ represents Simple Randomization, ‘Block’ represents Block Randomization, ‘Stratified’ represents Stratified Randomization, ‘Minirandom’ represents Minimization Randomization

Table 10.

Imbalance ratios across randomization methods

Covariate CMSB Constrained Simple Block Stratified Minimization
Mean income 0.086 0.181 0.311 0.287 0.269 0.235
Mean age 0.085 0.218 0.315 0.306 0.271 0.232
Mean education level 0.077 0.169 0.309 0.303 0.268 0.194
Mean bmi 0.070 0.146 0.294 0.312 0.269 0.256
Mean wealth 0.056 0.158 0.309 0.293 0.297 0.136
Mean height 0.070 0.180 0.304 0.269 0.262 0.251
Mean weight 0.059 0.144 0.320 0.315 0.257 0.247
Agricultural employment rate 0.084 0.166 0.306 0.304 0.283 0.112
Average 0.073 0.170 0.309 0.299 0.272 0.208

Table 11 reveals substantial variation in computational demands across different randomization methods, reporting total costs for 1,000 bootstrap resamples. Constrained randomization incurred substantial computational costs. Traditional methods such as stratified and block randomization showed intermediate performance in both balance and speed, whereas simple randomization, despite its computational simplicity, produced unacceptable imbalance levels for rigorous trials.

Table 11.

Running time across randomization methods with the total cost for 1000 bootstrap resamples

Randomization Method Running Time (seconds)
CMSB randomization 91.960
Constrained randomization 46446.680
Simple randomization 0.047
Block randomization 0.069
Stratified randomization 552.082
Minimization randomization 20.083

The number of randomization schemes to simulate is up to 2,000

The results demonstrate that CMSB provides an optimal compromise between covariate balance and computational efficiency, particularly in studies with larger numbers of clusters. Its stable performance across varying number of clusters makes it especially suitable for complex cluster randomized trials where both balance and practical implementation are critical considerations.

-Discussion

In this study, we developed the CMSB method, a randomization approach specifically designed to optimize covariate balance between treatment and control arms in cluster randomized trials. By extending the core principles of the MSB method, CMSB introduces modifications in how balance is assessed for different variable types, thereby enhancing both computational efficiency and balance quality in practical applications. Moreover, we provided a compatible statistical inference framework tailored for CMSB, ensuring valid hypothesis testing. CMSB extends the MSB framework to cluster randomized trials. It preserves MSB’s core approach to dynamic imbalance monitoring and conditional biased-coin randomization, while introducing key adaptations for the CRT context. Coupled with its computational efficiency, these features make CMSB a reliable randomization tool for CRTs across a broad range of sizes and complexities.

The primary advantage of CMSB lies in its improved ability to preserve covariate balance across diverse settings. Unlike conventional minimization methods, which are largely effective only for categorical covariates [27], CMSB achieves robust balance control for both continuous and categorical variables, demonstrating methodological flexibility across variable configurations. While achieving balance quality comparable to constrained randomization in low-dimensional scenarios (4 covariates), CMSB demonstrates substantially superior performance in high-dimensional settings (10 covariates), with approximately 45% improvement in balancing efficacy. However, this observed advantage depends on preserving a sufficiently informative covariate set and avoiding aggressive dimension reduction. In practice, high-dimensional scenarios often require preprocessing steps such as variable selection or dimensionality reduction. When applied, these steps may diminish CMSB’s net gain, particularly if essential prognostic information is inadvertently discarded. Therefore, the reported advantage should be interpreted conditionally, and sensitivity analyses varying the covariate subset or reduction method are advisable.

The advantages of CMSB are particularly pronounced in complex CRTs. Contemporary studies often require balancing 3–9 covariates simultaneously, with recent evidence from 33 major cluster trials involving 50–200 clusters each [10]. Under these demanding, real-world conditions, constrained randomization imposes a prohibitive computational burden without guaranteeing superior covariate balance.

Beyond its primary objective of ensuring rigorous balance, CMSB exhibits strong computational efficiency. CMSB reduces the computationally intensive runtime of constrained randomization (> 100 s per allocation with up to 2,000 randomization schemes simulated) to 0.116 s, yielding a 99.9% improvement without sacrificing balance quality. This computational breakthrough makes CMSB especially advantageous for large-scale trials where resource constraints might otherwise necessitate methodological compromises. It should be noted that the computational burden of constrained randomization partly depends on the number of simulated randomization schemes, capped at 2,000 by default. Reducing this cap could lower computation time, albeit potentially at the expense of balance quality.

Another central motivation of this study was to examine how a dynamic randomization procedure such as CMSB performs in static allocation settings typical of CRTs, where all clusters are available and a single allocation is drawn. Although CMSB was designed as a dynamic, imbalance-monitoring method, its mechanism can be applied in a one-shot fashion by evaluating imbalance once before final allocation. This allows assessment of whether the benefits of dynamic designs—continuous monitoring, flexible adjustment, and balance improvement—can translate into the static allocation context dominant in CRT practice. By positioning CMSB alongside established randomization approaches—including static designs enforcing balance through enumeration or constraints, and other covariate-adaptive methods—our study contributes not only a methodological extension of MSB but also a systematic comparison across fundamentally different design philosophies. The findings highlight the potential of minimization-style, dynamic procedures to perform competitively in static CRT settings, offering new insights into the practical scope of such methods.

Crucially, these gains are not achieved at the expense of randomization integrity. CMSB maintains desirable randomness, with a near-ideal correct-guess probability of 49.79%—outperforming minimization (61.03%) and block randomization (66–75%). High predictability is a critical flaw, as it undermines trial validity by increasing risks of selection and chronological biases and potentially inflating Type I error rates [27, 28]. The capacity of CMSB to simultaneously optimize covariate balance while satisfying the fundamental requirement of proper randomization is a key feature.

To complete the framework, we validated the linear mixed model Wald F-test and χ2-test as a suite of compatible statistical inference methods. Through extensive simulations, all methods preserved the nominal Type I error rate while achieving adequate power for clinically meaningful effect sizes, confirming their appropriateness for analyzing CMSB-based randomized trials. These validated analytical approaches provide researchers with flexible options tailored to specific study characteristics and distributional assumptions.

While rerandomization is a well-established method for improving covariate balance, its application in cluster randomized trials can encounter practical limitations in certain settings [17, 29–32]. These include considerable computational demands in high-dimensional covariate scenarios, as well as potential challenges in accommodating sequential cluster enrollment. In contrast, CMSB’s sequential algorithm provides a computationally efficient and dynamically adaptive alternative that is well-suited for contemporary CRT designs requiring real-time allocation.

Future work could extend the CMSB framework to employ threshold criteria based on effect size measures, such as the Standardized Mean Difference (SMD) for continuous covariates. While the current p value thresholds maintain consistency with the original MSB approach and facilitate comparisons, subsequent versions could adopt SMD-based rules (e.g., SMD > 0.1) to define imbalance, offering more intuitive and scale-invariant control. Another promising direction is to explore alternative decision rules for the biased-coin mechanism. Rather than relying on the current discontinuous step-function, continuous allocation functions that map imbalance margins directly to assignment probabilities could be investigated, potentially enabling finer control over balance.-

Conclusion

In summary, the CMSB method provides a viable randomization strategy for cluster-randomized trials by effectively optimizing covariate balance without compromising allocation randomness or computational efficiency. Its integration of robust performance with practical feasibility enhances the design of trials involving numerous clusters or high-dimensional covariate structures.

Supplementary Information

Supplementary Material 1. (478.2KB, docx)

Acknowledgements

The authors appreciate the very helpful comments and suggestions from the referees and the editor.

Authors’ contributions

Jiaxin Cai, Chao Li, Fang Shao, and Tao Chen proposed the study concept and design; Jiaxin Cai drafted the manuscript; Jiaxin Cai, Shanshan Suo, and Binyan Zhang performed statistical analysis; Valirie Ndip, Carole Khairallah provided critical revision of the manuscript; Hong Yan and Lingxia Zeng provided data curation and empirical study support; Chao Li, Fang Shao, and Tao Chen conducted study supervision; and all authors critically revised the manuscript for important intellectual content.

Funding

This research was supported by the National Natural Science Foundation of China (Nos. 82322060 and 82473732), Open Project Fund of Key Laboratory of Biosafety Defense, Ministry of Education (No. KLBD-2024-005) and Open Project Fund from Key Laboratory of Coal Environmental Pathogenicity and Prevention (Shanxi Medical University), Ministry of Education, China (MEKLCEPP/SXMU-2025MEKLCEPP/SXMU-202501).

Data availability

All data involved in the current empirical study were obtained from a double-blind cluster randomized controlled trial conducted between August 2002 and January 2006 in two rural counties in Shaanxi Province, northwest China.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Fang Shao, Email: shaofang@njmu.edu.cn.

Chao Li, Email: lcxjtu_research@163.com.

References

  • 1.Bondemark L, Ruf S. Randomized controlled trial: the gold standard or an unobtainable fallacy? EORTHO. 2015;37:457–61. 10.1093/ejo/cjv046. [DOI] [PubMed] [Google Scholar]
  • 2.Strayhorn JM. Virtual controls as an alternative to randomized controlled trials for assessing efficacy of interventions. BMC Med Res Methodol. 2021;21:3. 10.1186/s12874-020-01191-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zabor EC, Kaizer AM, Hobbs BP. Randomized controlled trials. Chest. 2020;158:S79–87. 10.1016/j.chest.2020.03.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wang X, Chen X, Goldfeld KS, Taljaard M, Li F. Sample size and power calculation for testing treatment effect heterogeneity in cluster randomized crossover designs. Stat Methods Med Res. 2024;33:1115–36. 10.1177/09622802241247736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Ivers NM, Halperin IJ, Barnsley J, Grimshaw JM, Shah BR, Tu K, et al. Allocation techniques for balance at baseline in cluster randomized trials: a methodological review. Trials. 2012;13:120. 10.1186/1745-6215-13-120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Hemming K, Eldridge S, Forbes G, Weijer C, Taljaard M. How to design efficient cluster randomised trials. BMJ. 2017;358. 10.1136/bmj.j3064. [DOI] [PMC free article] [PubMed]
  • 7.Weijer C, Grimshaw JM, Taljaard M, Binik A, Boruch R, Brehaut JC, et al. Ethical issues posed by cluster randomized trials in health research. Trials. 2011;12:100. 10.1186/1745-6215-12-100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ivers NM, Taljaard M, Dixon S, Bennett C, McRae A, Taleban J, et al. Impact of CONSORT extension for cluster randomised trials on quality of reporting and study methodology: review of random sample of 300 trials, 2000-8. BMJ. 2011;343. 10.1136/bmj.d5886. [DOI] [PMC free article] [PubMed]
  • 9.Eldridge S, Ashby D, Bennett C, Wakelin M, Feder G. Internal and external validity of cluster randomised trials: systematic review of recent trials. BMJ. 2008;336:876–80. 10.1136/bmj.39517.495764.25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Wright N, Ivers N, Eldridge S, Taljaard M, Bremner S. A review of the use of covariates in cluster randomized trials uncovers marked discrepancies between guidance and practice. J Clin Epidemiol. 2015;68:603–9. 10.1016/j.jclinepi.2014.12.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Nevins P, Davis-Plourde K, Pereira Macedo JA, Ouyang Y, Ryan M, Tong G, et al. A scoping review described diversity in methods of randomization and reporting of baseline balance in stepped-wedge cluster randomized trials. J Clin Epidemiol. 2023;157:134–45. 10.1016/j.jclinepi.2023.03.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Brinkman SA, Johnson SE, Codde JP, Hart MB, Straton JA, Mittinty MN, et al. Efficacy of infant simulator programmes to prevent teenage pregnancy: a school-based cluster randomised controlled trial in Western Australia. Lancet. 2016;388:2264–71. 10.1016/S0140-6736(16)30384-1. [DOI] [PubMed] [Google Scholar]
  • 13.Bolzern J, Mnyama N, Bosanquet K, Torgerson DJ. A review of cluster randomized trials found statistical evidence of selection bias. J Clin Epidemiol. 2018;99:106–12. 10.1016/j.jclinepi.2018.03.010. [DOI] [PubMed] [Google Scholar]
  • 14.Zhao W, Hill MD, Palesch Y. Minimal sufficient balance—a new strategy to balance baseline covariates and preserve randomness of treatment allocation. Stat Methods Med Res. 2015;24:989–1002. 10.1177/0962280212436447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Zhou Y, Turner EL, Simmons RA, Li F. Constrained randomization and statistical inference for multi-arm parallel cluster randomized controlled trials. Stat Med. 2022;41:1862–83. 10.1002/sim.9333. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Li F, Turner EL, Heagerty PJ, Murray DM, Vollmer WM, DeLong ER. An evaluation of constrained randomization for the design and analysis of group-randomized trials with binary outcomes. Stat Med. 2017;36:3791–806. 10.1002/sim.7410. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Li F, Lokhnygina Y, Murray DM, Heagerty PJ, DeLong ER. An evaluation of constrained randomization for the design and analysis of group-randomized trials. Stat Med. 2016;35:1565–79. 10.1002/sim.6813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Raab GM, Butcher I. Balance in cluster randomized trials. Statist Med. 2001;20:351–65. 10.1002/1097-0258(20010215)20:3%3C;351::AID-SIM797%3E;3.0.CO;2-C. [DOI] [PubMed] [Google Scholar]
  • 19.Crisp AM, Halloran ME, Longini IM, Vazquez-Prokopec G, Dean NE. Covariate-constrained randomization with cluster selection and substitution. Clin Trials. 2023;20:284–92. 10.1177/17407745231160556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Ciolino JD, Diebold A, Jensen JK, Rouleau GW, Koloms KK, Tandon D. Choosing an imbalance metric for covariate-constrained randomization in multiple-arm cluster-randomized trials. Trials. 2019;20:293. 10.1186/s13063-019-3324-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Sullivan TR, Morris TP, Kahan BC, Cuthbert AR, Yelland LN. Categorisation of continuous covariates for stratified randomisation: how should we adjust? Stat Med. 2024;43:2083–95. 10.1002/sim.10060. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Campbell I. Chi-squared and Fisher-Irwin tests of two‐by‐two tables with small sample recommendations. Stat Med. 2007;26:3661–75. [DOI] [PubMed] [Google Scholar]
  • 23.Kuznetsova OM, Johnson VP. Approaches to expanding the two-arm biased coin randomization to unequal allocation while preserving the unconditional allocation ratio. Stat Med. 2017;36:2483–98. 10.1002/sim.7290. [DOI] [PubMed] [Google Scholar]
  • 24.Molenberghs G, Verbeke G. Likelihood ratio, score, and Wald tests in a constrained parameter space. Am Stat. 2007;61:22–7. 10.1198/000313007X171322. [Google Scholar]
  • 25.Nihan ST. Karl Pearsons chi-square tests. Educ Res Rev. 2020;15:575–80. 10.5897/err2019.3817. [Google Scholar]
  • 26.Zeng L, Cheng Y, Dang S, Yan H, Dibley MJ, Chang S, et al. Impact of micronutrient supplementation during pregnancy on birth weight, duration of gestation, and perinatal mortality in rural western China: double blind cluster randomised controlled trial. BMJ. 2008;337(nov07 4):a2001–a2001. 10.1136/bmj.a2001. [DOI] [PMC free article] [PubMed]
  • 27.Kahan BC, Morris TP. Improper analysis of trials randomised using stratified blocks or minimisation. Stat Med. 2012;31:328–40. 10.1002/sim.4431. [DOI] [PubMed] [Google Scholar]
  • 28.Berger VW, Bour LJ, Carter K, Chipman JJ, Everett CC, Heussen N, et al. A roadmap to using randomization in clinical trials. BMC Med Res Methodol. 2021;21:168. 10.1186/s12874-021-01303-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Li X, Ding P, Rubin DB. Asymptotic theory of rerandomization in treatment-control experiments. Proc Natl Acad Sci USA. 2018;115:9157–62. 10.1073/pnas.1808191115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Li X, Ding P, Lin Q, Yang D, Liu JS. Randomization inference for peer effects. J Am Stat Assoc. 2019;114:1651–64. 10.1080/01621459.2018.1512863. [Google Scholar]
  • 31.Li X, Ding P. Rerandomization and regression adjustment. J R Stat Soc Ser B Stat Methodol. 2020;82:241–69. 10.1111/rssb.12353. [Google Scholar]
  • 32.Morgan KL, Rubin DB. Rerandomization to improve covariate balance in experiments. Ann Statist. 2012;40:1263–82. 10.1214/12-AOS1008. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (478.2KB, docx)

Data Availability Statement

All data involved in the current empirical study were obtained from a double-blind cluster randomized controlled trial conducted between August 2002 and January 2006 in two rural counties in Shaanxi Province, northwest China.


Articles from BMC Medical Research Methodology are provided here courtesy of BMC

RESOURCES