Skip to main content
Sage Choice logoLink to Sage Choice
. 2026 Aug 5;35(9):1838–1856. doi: 10.1177/09622802261454833

Stratification in small randomised clinical trials and analysis of covariance: Some simple theory and recommendations

Stephen Senn 1,2,3, Franz König 1, Martin Posch 1,✉
PMCID: PMC13558804  PMID: 42555140

Abstract

A simple device for balancing for a continuous covariate in clinical trials is to stratify by whether the covariate is above or below some target value, typically the predicted median. The object is to improve balance of the covariate and hence efficiency of the treatment estimate particularly if the trial is small, as may be the case if the disease in question is rare. This raises an issue as to which model should be used for modelling the effect of treatment on the outcome variable, Y. Should one fit, the stratum indicator, S, the continuous covariate, X, both or neither? When a covariate is added to a linear model there are three consequences for inference: (a) the mean square error effect, (b) the variance inflation factor and (c) second order precision. We consider that it is valuable to consider these three factors separately, even if, ultimately, it is their joint effect that matters. We present some simple theory, concentrating in particular on the variance inflation factor, that may be used to guide trialists in their choice of model. We also consider the case where the precise form of the relationship between the outcome and the covariate is not known. We conclude by recommending that the continuous covariate should always be in the model but that, depending on circumstances, there may be some justification in fitting the stratum indicator also.

Keywords: Efficiency, variance inflation factor, degrees of freedom, mean square error, truncated normal

1. Introduction

In what follows we shall consider small two-arm parallel-group trials in which the target outcome variable is continuous and may be assumed (perhaps after suitable transformation) to be approximately Normally distributed and it is proposed to use a single covariate in analysis. Although the theory we develop applies to trials of any size, its practical importance is greatest for small trials. Thus, it may be helpful for the reader to assume that the context is rare diseases. We shall consider the case of one covariate only and for which the treatment-by-covariate interaction is not fitted. However, in the discussion we shall consider briefly what happens if trials become large, if they have more than two arms, if more than one covariate is used or if the interaction is fitted.

A clarification of terminology may be helpful. The term stratification usually refers to stratified allocation (sometimes referred to as pre-stratification), which is typically accompanied by an analysis taking strata into account. However, even where such stratification has not taken place, it is possible to create strata for analysis. This technique is sometimes referred to as post-stratification (also known as stratification after selection 1 ). In this paper, we are mainly concerned with the former form of stratification, which is to say, trials that have been stratified by design, and unless we explain otherwise, this is the sense that should be understood. We shall compare two designs. Both designs will have equal numbers on each arm and both designs will be randomised. The first design will use stratified randomisation and we shall refer to this as the stratified design. The second design will have no further restriction and we refer to this as the randomised design. Furthermore, we shall assume that the following specific approach to stratification has been applied to a single continuous covariate, for example the value measured at baseline of the target outcome variable. Alternatively, some sort of prognostic score or ‘supercovariate’ might be considered. 2 A predicted median value of the variable is used and when the patients are entered onto the trial they are assigned either to the low or high stratum depending on how their value compares to the median. In allocating patients to treatment, an attempt is made to balance numbers on each treatment arm within each stratum. In practice, this would have to be achieved using some method such as permuted blocks within strata. 3 In the case of randomisation this would be used to balance numbers for the trial as a whole. In developing the theory, we shall assume that the blocks are merely a balancing device and can otherwise be ignored. We shall pick this point up in the discussion.

Qu has suggested the following possible classification of models 4 (p233):

Model(A):Y∼ZModel(B):Y∼Z+XModel(C):Y∼Z+SModel(D):Y∼Z+X+S.

Here the continuous covariate is X, the outcome variable is Y, the treatment indicator is Z and the stratum indicator is S. We shall adopt this classification system and mainly consider the effect of stratification when it has been decided to fit Model (B). We shall then consider the effect of fitting Model (D).

Note that all four models are main-effect models only. They do not include a treatment-by-covariate interaction. One could consider an expanded set of models in which the interaction of Z with S,X or both was considered. In this paper we shall not cover such models but they will be briefly considered in the discussion. However, as regards stratification there is one important technical point. The trials we are considering are stratified by design. In stratifying, there is a common choice in analysis between either fitting a factor with as many levels as strata in a linear model or fitting the treatment effect stratum-by-stratum and then combining the estimates. The latter also implicitly removes the treatment-by-stratum interaction from the mean square error and raises issues as to how the stratum estimates should be combined, a matter to which the long-standing debate regarding type II and type III sums of squares is relevant.5,6 Considering this latter issue would detract from the main points we wish to make, so we shall avoid it by not considering interactions. However, we shall pick this matter up briefly in our conclusions.

Previous investigations of modelling the effect of covariate adjustment for median stratified trials have relied on simulation and have looked at the net total effect of covariate adjustment. What is novel in this paper is that we shall show that it is useful to consider the effect of covariate adjustment in terms of three components (to be discussed in the next section). Furthermore, we develop a theory to address these components that does not rely on simulation. This provides insight regarding the role of stratification and modelling.

The outline of this paper is as follows. In section 2 we introduce the framework we propose to use for discussing the effects of fitting covariates. Sections 3 and 4 cover two components that depend on the model employed but are largely independent of the design: the expected mean square error and the residual degrees of freedom. Section 5 covers the effect of covariate balance through the variance inflation factor. This is the key to understanding what the value, if any, is of stratification. Section 6 presents a brief example. Section 7 is a discussion of the results and Section 8 provides some recommendations,

In what follows, it may be helpful to the reader to be able to refer to definitions of terms we shall use. These are given in Table 1.

Table 1.

Key statistical terms and conceptual frameworks used in the paper.

Term/framework Explanation
Stratification Assigning participants so that treatment groups are balanced within predefined baseline groups. In this paper, median stratification refers to dividing participants into ‘low’ and ‘high’ strata using the predicted median of a continuous covariate. The paper distinguishes between:
  • Pre-stratification: Stratification built into the trial design before treatment allocation, so that participants are assigned within predefined strata. This is also referred to a as ‘ stratified allocation ’.

  • Post-stratification: Stratification introduced after allocation or selection, usually for analysis rather than for controlling treatment assignment. This is also referred to as ‘ stratification after selection ’.

Two-arm parallel-group trial A clinical trial design in which participants are allocated to one of two treatment groups. The paper distinguishes two types of trial design:
  • Stratified design: a RCT design in which treatment allocation is balanced within predefined strata, such as low and high baseline covariate groups.

  • Randomised design: a RCT design with random treatment allocation and equal arm sizes, but no additional balancing by covariate strata.

Continuous covariate A numerical baseline variable used either for stratification at allocation and/or adjustment in the analysis.
Qu's model classification A four-model framework proposed by Qu4 distinguishing whether the analysis adjusts for the continuous covariate, the stratum indicator, both, or neither. All four models are main-effect models and do not include treatment-by-covariate interactions.
  • Model (A): A treatment-only model with no adjustment for either the continuous covariate or the stratum indicator.

  • Model (B): A model adjusting for treatment and the underlying continuous covariate; this is the paper's main reference model.

  • Model (C): A model adjusting for treatment and the stratum indicator, but not the original continuous covariate.

  • Model (D): A model adjusting for treatment, the continuous covariate, and the stratum indicator.

Three-components framework Proposed framework separating precision effects into mean square error, imbalance/VIF, and second-order precision.
  • Mean square error effect: The precision gain from reducing unexplained outcome variation by adjusting for a prognostic covariate.

  • Imbalance effect: The loss of precision that occurs when a covariate included in the analysis is unevenly distributed between treatment arms. This can be expressed through the VIF, which quantifies the resulting non-orthogonality penalty by showing how much the variance of the treatment effect estimate is increased.

  • Second-order precision effect: The precision cost (i.e. loss of degrees of freedom) of estimating extra model terms, especially relevant in small samples.

Mean Square Error (MSE) A measure of the unexplained variation left after fitting a model;
Variance Inflation Factor (VIF) A multiplier showing how much the variance of the treatment effect estimate increases because the treatment groups are not perfectly balanced on a covariate included in the model.
Non-orthogonality The situation in which treatment assignment and a covariate are not perfectly balanced with respect to one another in the design matrix.

2. A three components framework

When considering the standard linear model, there are three effects on precision of a treatment effect estimate due to fitting a covariate. 7 First, there is the mean square error (MSE) effect. To the extent that the covariate is prognostic, the expected value of the MSE to be used in calculating variances of estimates of treatment effects will be reduced. Second, there is the imbalance effect. To the extent that the covariate is imbalanced between the treatment arms, the variance will be increased and this is allowed for in regression approaches by multiplying the estimated variance by the appropriate variance inflation factor (VIF) 8 or non-orthogonality penalty.9,10 Third, there is the second order precision effect. A degree of freedom will be lost in estimating the MSE.

The relationship of these three components to precision is very different and this makes it desirable when considering them to separate their different roles. When the decision is made whether or not to use the baseline in an ANCOVA, the prime motivation, of course, is the expected reduction in the estimated MSE by fitting the baseline. This effect is usually much more important than the other two. However, as we discuss below, this component is not relevant when making a decision to stratify or not. For this, it is the VIF that matters. In consequence, since the question that is addressed in this work is the value or otherwise of stratification and what analysis to use if one has stratified, it is the VIF on which we shall concentrate attention.

It is important to understand the relationship the three components have to the model, the design and the outcome. These are summarised in Table 2.

Table 2.

Influence of three aspects, design (stratified or randomised), model (linear models without or with adjustment for different covariates) and outcome (different correlations of the outcome variables and the covariates) on the three components of precision.

Effect
Mean square error effect Imbalance effect or VIF Second order Precision
Influence Design No Yes No
Model Yes Yes Yes
Outcome Yes No No

Note that although all three components are influenced by the model, as regards second order precision, it is only the rank of the model that matters, since this determines the number of degrees of freedom lost for estimating the error term.

As regards the VIF, it is not affected by the outcome, because we restrict our attention here to the simple linear model with estimation using ordinary least squares. Under such circumstances, as is well known in the literature on design of experiments, design efficiency only depends on the design matrix. 11 For extensions of the linear model, this does not apply and the efficiency of chosen designs depends on unknown parameters that have to be estimated, 12 perhaps by iterative reweighting. For example, for binary outcomes the variance of a response depends on its expected value and for hierarchical models, such as might be applied to incomplete blocks designs, on the variance–covariance matrix. These parameters have to be estimated using the outcomes.

Finally, what the table states about the mean square error effect is true about the estimated mean square error effect given an appropriate model. An appropriate model, in this context, is a model that includes every prognostic covariate whose distribution has been restricted (as opposed to being randomised). Further unrestricted covariates that are prognostic may be included and restricted covariates may be excluded if not prognostic given other covariates in the model. If this is the case, only the model is relevant for considering the effect on MSE. However, if one restricts allocation in some way that is not reflected in the ANCOVA model, the actual mean square error is reduced (if the covariate is prognostic for the outcome) but the estimated mean square error from the ANCOVA model will be biased and larger than it should be, a point which RA Fisher raised some 90 years ago in his discussion 13 of Student's non-randomised designs.14,15 Indeed, in a recent paper using simulations to examine strategies for analysing stratified designs Sullivan et al. state, 16 ‘Since stratified randomisation induces a correlation between treatment groups, failure to adjust for stratification variables when estimating treatment effects can lead to overly wide confidence intervals, type I error rates less than α, and reduced statistical power’ (p. 2). As discussed above, if one stratifies, one should adjust in the analysis. However adjustment is possible whether or not one has stratified (provided that the covariate has been measured) and it also follows that if one has decided to adjust, stratification of itself plays no further role as regards two of the three components summarised in Table 1 above: the MSE effect and the second order precision effect are the same. For example, if, not having stratified, one simply recoded a covariate as 0 or 1 depending on whether its value was below or above the median, as Sullivan et al. did in their simulation (Section 3.1) and adjusted for it, the expected consequences for these two effects would be the same as if one had stratified. This then raises the question, what is the value of stratification? As we shall show, the key to understanding the value of stratification and the consequence of fitting various models is the VIF. We shall discuss this in due course. However, first we divert to discuss briefly the effect of the other two components.

3. The expected mean square error effect

This has been often discussed. (See, for example, p16 of the famous paper by Hills and Armitage. 17 ) We can consider this by supposing that the continuous predictor X has been standardised as has the outcome variable Y. The regression of Y on X is then the same as the correlation between the two, ρ , and, asymptotically, the reduction in the variance of the estimate of the treatment effect by fitting Model (B) rather than Model (A) is given by the ratio 1−ρ2 . A very common covariate to consider is the baseline value corresponding to the outcome and it is then interesting to compare covariate adjustment to simple analysis of the change-score (sometimes called the gain-score), that is to say, the difference of outcome from baseline. This is illustrated in Figure 1, which is based on Figure 7.4 of Statistical Issues in Drug Development 9 and takes as a reference analysing the ‘raw’ outcomes alone and ignoring the baseline altogether.

Figure 1.

Figure 1.

Variance of the treatment effect estimate for Model(B) (ANCOVA using X) as a function of the correlation coefficient compared to Model (A) (analysis of raw outcome only) as well as the common strategy of using a change score.

Note that, whatever the correlation coefficient, the covariate adjusted analysis always yields a lower or equal variance of the treatment effect estimate. However, the result is asymptotic. The fact that an extra parameter, the correlation coefficient, has to be estimated means that there is some loss to which neither alternative analysis (raw outcome or change score) is subjected. This suggests that it could be possible that for very low values of the correlation coefficient, raw outcomes could be preferable. Similarly, for high values, the change score could be preferred. Understanding this point requires looking at the other two components in Table 1, second order precision and the VIF.

4. Second order precision

When an extra term is fitted, the residual degrees of freedom for estimating the residual variance are reduced by one to the extent that the fitted term is not redundant. That is to say, if there are N patients in total and a matrix of predictors, P , the degrees of freedom ν are given by ν=N−rank(P) . We thus have ν=N−2 for Model (A), where a degree of freedom is lost for the treatment indicator in addition to the intercept, ν=N−3 for Models (B) and (C), where a further degree of freedom is lost either for X or for S and ν=N−4 for Model (D), where two degrees of freedom are lost for fitting both.

For any type-one error or level of confidence chosen, the residual degrees of freedom will determine the critical values of the t-distribution that has to be used. However, a way of indexing second order precision that is not linked to any given level chosen, is to use the variance of the t-distribution, which is

Var(tν)=νν−2,ν≥3.

Given this formula all that is necessary is to substitute the relevant values for ν , that is to say N−2,3or4 as the case may be. The results as functions of N are shown in Figure 2.

Figure 2.

Figure 2.

Variance of the t-distribution for various models (A–D).

Since it may be of interest to compare the variances of the different models an alternative representation is in terms of the ratio of the variances compared to the model with no covariates. This is given in Figure 3.

Figure 3.

Figure 3.

Ratio of variances of the t-distribution fitting one (Variance1) and two covariates (Variance2) compared to fitting none (Variance0) for various sample sizes.

Once the sample size reaches 20, the variance of the distribution of the t-statistic when fitting two covariates is only 1.6% higher than that of the distribution when fitting none. The variance when fitting one covariate is only 0.7% higher than that when fitting none. As already noted, this component depends on the number of covariates fitted and the sample size. It neither depends on the design nor on the nature of the covariates.

5. The variance inflation factor and the value of stratification

5.1. Why the variance inflation factor is important

We regard the VIF as central to understanding why exactly it is that balance matters. Various recent reviews of covariate adjustment stress that balance is valuable without really explaining why. For instance, Van Lancker et al. 18 in a thorough investigation of many consequences of using ANCOVA state, ‘In finite samples, adjusting for covariates with a very weak or null association with the outcome may not lead to efficiency gains and can even result in a precision loss’ and with respect to stratified biased-coin designs, ‘Such designs create balance of treatment arms with respect to the stratified variables in order to improve efficiency of the treatment effect estimate’. However, they do not pinpoint or quantify why and how balance contributes to efficiency. Here we explain how the VIF holds the key to understanding this.

If it has been decided to fit Model (B), the value of stratification is that it reduces the non-orthogonality penalty due to imbalance. Qu provides a careful examination of the effect of stratification on covariate balance. 4 However, what we are suggesting here is to take it one step further. Rather than considering the effect of balancing strategies on the covariate itself, we should consider how balance or imbalance affects the variance of the treatment effect estimate. If the covariate is in the model, its effect is adjusted for: the treatment effect estimate is unbiased and the expected mean square error is reduced, whether or not the covariate is unbalanced, to the degree that it is prognostic. Furthermore, the residual degrees of freedom are unaffected by the degree of imbalance. What changes is the VIF This is what we examine in this section,

For a trial with n subjects per arm i,i=1,2 , and thus N=2n in total, in which there is adjustment for a linear predictor, X, the variance inflation factor (VIF) will be:

VIF=∑i=12∑j=1n(Xij−X¯..)2∑i=12∑j=1n(Xij−X¯..)2−n2(X¯2.−X¯1.)2 (1)

(See Appendix.) This is the factor by which the usual multiplier for the MSE of 2/n must be further multiplied to produce the variance of the treatment effect. Unless X¯1.=X¯2. , which is to say unless the mean values of the covariate are identical in the two arms of the trial, this factor will be greater than one. The value of stratification is that it forces the expected value of this difference to be smaller than it otherwise would be.

5.2. Expected values of the variance inflation factor

If the covariate is Normally distributed, then in the absence of stratification, the expected value of the imbalance effect can be shown to be19,20

E[VIF]=1+1N−4=N−3N−4. (2)

(This formula applies if the covariate is Normally distributed and was given by Cochran 19 in 1957. It is derived in an appendix because it will be modified below for the case where the covariate has been stratified.) For moderately large values of N this is quite modest, a fact that is possibly not well understood. For example, for N=200 , which would be a small size for a Phase III clinical trial, one has E[VIF]=1.005 to three decimal places, so that the expected variance inflation would be about half of one percent. This calls into question the value of stratification and also other balancing strategies such as minimisation for such trials.

It is hard to give general rules for what benefit stratification will bring but one will have

1≤E[VIFS]≤N−3N−4=1+1N−4,

where VIFS is the variance inflation factor value given stratification.

Sullivan et al. 16 simulate covariate values, X, from a standard Normal distribution, stratifying at X=0 . (See Section 3.1 of their paper.) The effect of this can be judged approximately by considering that the variance of a Normal distribution truncated at 0 is 1−2/π , which is 0.363 to three decimal places (see Appendix). In each stratum, values can be regarded as being sampled from a distribution with this variance and this means that whereas such means have a variance proportional to 1 if sampled from the parent Normal distribution, they have a variance proportional to 0.363 when sampled from the truncated distribution. The resulting distribution is, however, not itself Normally distributed, so that standard distribution theory does not apply. Nevertheless, a value of

E[VIFs]=1+1−2/πN−4≈1+0.363N−4 (3)

should apply approximately, although it must be understood, that in practice it will not be possible to predict perfectly when recruiting patients what the median is for the population from which theoretically they could be considered to be drawn.

It is also interesting to consider the effect of stratification on using Model (D). The following argument can be used. If, having stratified, we fit Model (C), there will be no loss of orthogonality since we are only fitting the stratum indicator, which is balanced by design. If we now add as a further covariate the continuous predictor X, then it is the component of X that is orthogonal to S that is relevant to determining the VIF. However, given that it is orthogonal the expected penalty is that indicated by (2) but adjusted for the fact that we have lost one further degree of freedom for fitting the stratum. The formula for the penalty is

E[VIFS,X]=1+1N−5. (4)

One further formula may be of interest and that is the VIF that applies to a completely randomised design fitting Model (D), that is to say not only the continuous covariate X but also a dichotomy of it at the median even though stratification has not taken place. The problem is that there is a finite probability that the allocation will be completely confounded. For a trial with 2n patients, and subject to the constraint that there are n patients per arm, there are

(2n)!n!n!

possible allocations. Of these allocations, there are two for which either all values for group 1 (say) will be above the median and all those for group 2 will be below the median or vice versa. Thus, the probability of complete confounding is

2n!n!(2n)!. (5)

For moderately large values this becomes very small. For example, for n=10 , that is to say 20 patients in total, it is less than one in 92,000. Nevertheless, a purist might claim that this means that the expected value of the VIF is undefined, since there is a finite probability of an infinite variance. Note that, in theory, a similar problem could arise with median stratification, since the ‘median’ used for allocation could (unbeknown to the trialist when the trial was designed) be higher (or lower) than the value for any patient. Then, effectively, there would only be one stratum and the baseline values would be as if completely randomised. Obviously, however, this is unlikely. In any case, this is not a problem for validity of marginal modelling inferences, since the stratum indicator can be eliminated from the model and since that choice will have been made based on the covariate distribution only, inferences will not be biased. It will, of course, be a problem for conditional inferences.

If this objection can be ignored, then for moderately large samples we can treat both covariates as if they were Normally distributed. A further paper of ours treats adjustment for any number of covariates, both binary and continuous and shows that the general formula given by Cox and McCullagh 20 in 1982 (p547) can be applied . For k covariates we have that

E[VIF(k)]=1+kN−k−3. (6)

may be used as an approximation, even though one of the predictors is binary. Substituting k=2 we have

E[VIF(2)]=1+2N−5. (7)

Expected VIFs for models B and D for both stratified and randomised designs are given in Figure 4 as are means from 106 simulation runs sampling the covariate from a Normal distribution. The theory seems to hold up well. Of course, one can argue that the Normal distribution might not accurately cover any given covariate distribution. However, there are three answers to this. The first is that the Normal distribution is what others have assumed in simulating. If such simulations are justified, it is surely of interest to look at what theory has to say. The second is that considering the Normal distribution is a start. It is useful to have established what the theory says about this standard case. The third is that the stratum indicator itself, which is included in model D, is not Normally distributed but the theory appears to hold. In fact, our investigations of this point, which will be presented in a further paper on this topic, suggest that for moderate sample sizes, even for binary covariates the theory holds up quite well.

Figure 4.

Figure 4.

Variance inflation factors for randomised (red, dashed) and stratified (blue, solid) designs, theoretical values (lines) and empirical (symbols) results for models B and D (106 simulation runs per scenario).

6. An example

We present a small example that has been adapted from data given in Table 12.9 of Piantadosi's famous book 21 and which concerns a trial in familial adenomatous polyposis (FAP) that compared sulindac to placebo. 22 We have reduced the data to that for 16 patients only in order to provide an example that permits division of patients into four equally sized groups produced by median stratification and treatment. The trial may seem to be rather small but Hee et al., in their survey of trials in rare diseases, 23 found a median size of 15 for phase II trials. One value has been slightly modified to avoid ties compared to the median. The data we shall use are polyp size at baseline. Since we are interested in the VIF, the values at outcome are not relevant. Note that we are re-allocating using these data not resampling.

The data are given in Table 3. We re-allocated from these data in two ways. First by considering all possible allocations that would give eight patients on each treatment. Second, by creating two strata by noting whether values are above or below the median and then allocating four patients to each treatment in each stratum. The former approach yields 16!/(8!8!)=12,870 possible allocations. The latter approach yields (8!/(4!4!))2=702=4900 possible allocations. We have simulated 5000 allocations from the former and performed a complete enumeration of all 4900 possible allocations for the latter.

Table 3.

Modified data from a trial in FAP reported by Piantadosi showing polyp size at baseline for 16 subjects.

Patient ID Treatment Baselinesize
12 Sulindac 1.7
8 Placebo 2.0
7 Sulindac 2.2
14 Placebo 2.3
15 Placebo 2.4
13 Placebo 2.5
5 Sulindac 3.0
16 Sulindac 3.0
22 Sulindac 3.1
4 Placebo 3.4
23 Sulindac 4.0
6 Placebo 4.2
9 Placebo 4.2
10 Placebo 4.8
3 Sulindac 5.0
11 Sulindac 5.5

The data have been sorted by polyp size. The value for patient 22 has been increased from 3.0 to 3.1 in order to yield a unique median.

Table 4 below gives the results of the calculation of the VIF for treatment for re-allocating for this example.

Table 4.

Results for the VIF obtained by taking the baseline values as fixed and re-allocating treatment.

Allocation Method Number of Allocations Expected Observed Mean Observed Median Standard Error of Mean
Completely Randomised 5000 1.083 1.086 1.035 0.0021
Stratified 4900 1.030 1.022 1.011 NA

The results for completely randomised allocations reflect a true simulation. For the stratified case, the VIF has been calculated for every one of the 4900 possible allocations. The result is thus not subject to random simulation error and the standard error of the mean is not quoted.

It can be seen that the theoretical expected values, which strictly speaking apply to sampling from a Normal distribution, agree fairly well with the empirical results based on re-allocating. It can also be seen that there is a modest reduction in the VIF as a result of median stratification. Figure 5 gives density plots for the two types of allocation. It can be seen that the distributions are highly skewed towards large values, which explains that the mean VIF is larger than the median for both types of allocation. The VIF distribution under randomisation is much more skewed than under median stratification.

Figure 5.

Figure 5.

Violin plots for VIFs for randomisation (not stratifying) and stratification. The randomisation values are based on a simulation with 5,000 re-samplings. The stratification distribution is based on complete evaluation of all 4,900 possible allocation.

7. Discussion

7.1. The value of balance for covariate adjustments

As was pointed out in Statistics in Medicine 30 years ago, 24 (p1721) ‘having decided to condition, we minimise the standard errors of our treatment effects if our covariate is orthogonal to treatment. This is the true value of balance, which is an issue of efficiency, not validity’. This ought to be more widely understood and it also ought to be understood that the potential for increasing efficiency in this way is limited. What is important is to condition on that which is prognostic.

We have concentrated on the expected values of the VIF here but a referee has suggested that, in choosing to stratify, one is not just seeking to avoid expected loss due to imbalance but one might also be concerned to avoid cases of extreme loss. Indeed, our small, simulated example in Section 6 shows that extreme values for the VIF are far less common for the stratified design than for the randomised design. See in particular Figure 5. The further paper of ours already mentioned, in addition to giving expected values for VIFs for a randomised design using any number of Normally distributed covariates, also gives their variances. We do not give the theory here but in general, for a completely randomised design, if VIF(k) is the VIF for model (B) fitting k Normally distributed covariates, its variance is

Var[VIF(k)]=2k(N−3)(N−k−3)2(N−k−5),N>k+5. (8)

Expression (8) can be adapted to be applied to all the cases we have considered as follows. For a simply randomised design for Model (B) or Model (C) set k=1 . For Model (D) set k=2 . For a stratified design for Model (C) the VIF is always 1 and its variance is 0. For Model (D) set k=1 and replace N by N−1 . For Model (B) set k=1 but multiply the result by

4π+π2−202π2≈0.1234. (9)

Expression (9) is derived from the general formula for the variance of a variance (see Stuart and Ord 25 section 2.15) and using the 4th moment about the origin of a standard Normal truncated at 0. This term shows that in addition to reducing the VIF, stratification has a considerable effect on its variance.

It should be noted that for randomisation and Model (C) and Model (D) the stratum indicator is not balanced by design, so this theory relies on being able to treat a dichotomised covariate as being Normally distributed. Rather surprisingly, our limited investigations suggest that this is not a problem.

7.2. Implications for allocation and choosing a model

It is interesting to note that an early paper on stratification for a continuous covariate by Finney considered trying to make mean covariate values close. 26 This is the relevant target for a continuous covariate for minimising the VIF. This was possible in the animal studies he was considering, where the animals to be treated were available as a group before starting the study. However, for clinical trials this will rarely be possible and the sort of stratification considered here will have to be used. As Sullivan et al. 16 point out, given that patient accrual has to be sequential in clinical trials, balance of total numbers and, by extension, balance in any stratum, can only be achieved by some device such as permuted blocks.3,27

Thus, usually, subjects will be balanced by blocks that are smaller than the strata. Yet, typically one does not adjust for this extra balancing factor: ‘block’ will not be in the model. 27 A justification for not putting block in the model is that it is just a device for balancing the dichotomised covariate represented by the strata and not expected to have any prognostic value of its own. The exception to this would be if it did have a prognostic value because some time trend was expected.27,28

Thus, if block can be ignored when balancing strata, because it is not expected to bring additional prognostic value, it would seem logical that stratum can be ignored if it is not expected to bring additional modelling value to the covariate from which it was created. In other words, it would be logical to use Model (B) of Qu's, 4 rather than Model (C). Another argument can be given as follows. Although there is no penalty in terms of VIF for fitting Model (C), as expression (3) shows, the penalty for fitting Model (B) is small. It is likely, however, that as regards expected MSE, whatever the true relationship is between the baseline and the outcome, it is likely to be better approximated by a straight line than a step-function. This agrees with Sullivan et al. 16 who state, ‘We therefore advocate adjustment for the underlying continuous values of stratification variables rather than adjustment for the categories used for stratification.’ (p. 10).

7.3. Is there an argument for using Model (D)?

As regards Model (D), things are not quite so clear and it is useful to consider what the consequences will be of fitting this model, rather than Model (B) on the three components we have identified as relevant. For large trials, the advantage might lie with D but for small trials with B.

The second component, the VIF, will be increased in expectation from expression (3) to expression (4). The third, second order precision will be affected by the loss of one degree of freedom. The first, mean square error might be reduced if the regression of Y on X is not perfectly linear. For example, one might suppose an unidentified form that might be approximated by some modelling scheme such as splines, 29 fractional polynomials,30,31 orthogonal polynomials 32 or regular polynomials. As an example of the latter, it is interesting to consider the correlation of the stratum indicator with various powers of X. (We are not suggesting that such a polynomial model is one most likely to be true or useful but just using it as an illustration as to what might happen with a more complex relationship between covariate and outcome than a simple linear one.) In order to make progress in developing some simple theory we shall make the following assumption that can apply only approximately at best, namely that the sample is very well behaved as regards the stratification plan and that all standardised values below the median are negative and all those above the medians are positive. Where this is the case, we have.

S={−1X<X~1X>X~, (10)

where X~=0 is the median of X. The values are given in Figure 6 and have been calculated as follows. We note that a positive and negative stratum indicator will be associated with equal probability with negative and positive values of X. For any even power of X, the expectation is thus zero, since an even power will be positive and associated with a negative or positive stratum indicator with equal probability. However, a negative value of an odd power will always be associated with a negative stratum indicator and a positive one with a positive stratum indicator, so the expected product of indicator and power is given by the corresponding moment of the standard Normal distribution truncated at zero. As is explained in the appendix, this may be adapted to provide expected correlations between the stratum indicator and the odd powers of X.

Figure 6.

Figure 6.

Correlation of the stratum indicator with various powers of X (squares) and the partial correlation (dots) of the stratum indicator with various powers of X given X. For even degrees, both the simple and partial correlations are 0 and the dots overlap.

If Figure 6 is examined, it will be noted that the correlation with all even powers of X is zero and that for each successive odd power of X, the correlation is about one half of the previous odd power. However, as we saw when moving from Model (B) to Model (D), it is really the partial correlation with the treatment indicator that is relevant. Many schemes for models are possible and we have already mentioned some. However, in a randomised design, a scheme in which successive polynomial terms to order m were used would require a value of k=m to be used in expression (6). It seems intuitively reasonable that however much could be achieved in reducing the VIF by median stratification, the maximum effect would equivalent to setting k=m−1 and that this is unlikely to be achieved in practice.

We prefer not to be proscriptive regarding the choice between Model (B) and Model (D). We suspect that in many cases choosing D rather than B will bring little by way of benefit but also will not cause much harm. Whatever choice is made, we advise that it should be agreed and pre-specified.

So far, we have not dealt with situations in which the treatment-by-baseline interaction is considered important. Regulatory guidance has generally stressed using a model without an interaction for a primary investigation and looking at interactions in a secondary one.33,34 Models that do not include interactions are generally robust even if an interaction is present but since they pool the interaction term with the error term in the residual mean square, may lose some power. That is a matter, however, that is not directly related to this investigation. However, in the discussion we provide a formula that may be used for models with interactions.

7.4. What is the value of median stratification?

As regards the value of median stratification as opposed to randomisation if, say, model (B) is envisaged, our opinion is that this might be valuable for a small trial, such as the one illustrated in Section 6. Comparison of the VIFs in Figure 5 for the small trial we considered supports this. (Note, however, that stratifying on the true median, as opposed to the predicted median, which would have to be used in practice, overstates the value.) For large phase III trials it will be of minimal advantage. For example, for 40 patients in total the expected value of the VIF given randomisation is 1.028 and for a stratified design is 1.010. Expression (3) may help trialists decide whether stratification is worthwhile.

The theory we have given applies to stratification for one covariate. Given that the context we have considered is small trials, for example in rare diseases, fitting more than one covariate may be rather ambitious. Formula (6) suggests the VIF that will apply. However, one possible strategy is to construct a prognostic score from historical data.2,7,35 Allocation could then be stratified by this score using its predicted median values, in which case, the theory we have presented applies. A further paper of ours considers expected values and variances of VIFs given randomisation when many covariates are fitted but we do not think this is of much practical relevance to the rare disease situation.

A referee has drawn our attention to the fact that stratification may be in terms of what is deemed a clinically relevant cut-point and this may not be, even by intention, close to the median. Our view is that ideally, as suggested by expression (1), it is the mean values of the covariate that we should like to equal between groups. 26 If a single cut-point is used to define groups to stratify allocation, then the median will come closest to achieving such mean balance. Any other cut-point will be less effective in approximating mean balance. Thus, the formula we have given should be regarded as providing a lower bound of what is achievable by stratification. Of course, it may be that the trial population contains a mixture of two very different populations with (say) a minority having extreme values. This, however, would raise the question of subgroup analysis and interactive effects which go beyond the goals of this paper.

Trials with three or more arms raise other interesting considerations. Variance inflation increases as the number of columns in the design matrix increases relative to the number of rows. Now consider a trial with t≥3 arms, a single covariate and n patients per arm, where n is small. If, when comparing two treatments, the data from other treatments are ignored, the design matrix is reduced from having 1+(t−1)+1=t+1 columns, where the three terms on the left hand side are respectively the degrees of freedom for intercept, treatments and covariate to having three columns only, However the rows of the matrix are reduced from tn to 2n and the net effect will be to increase the VIF. It may thus be advantageous to use the data from all treatments, imposing the assumption of common slopes across all treatments, including ones not involved in a given contrast.

However, as the value of n increases, such an assumption of homogeneity of slopes may do more harm than good, and Ye et al. have suggested that allowing for heterogeneity may be beneficial where sample sizes are large and trials have three or more arms. 36 A similar issue arises with variance estimation. In analysing clinical trials, many use a linear model that pools the variance estimates from all treatments. As was pointed out by Senn, 27 this makes sense in the context of agricultural field trials, where degrees of freedom are scarce but less so for large phase III clinical trials. We consider, however, that the large multi-armed trials are not really relevant to the rare disease context we have considered.

7.5. Allowing for treatment by covariate interaction

As we explained in the introduction, we have limited our investigation to main effect models. However, the three- component framework can be extended to models with treatment by covariate interactions. For example, we might be interested in models

Model(Bi):Y∼Z+X+Z×XModel(Ci):Y∼Z+S+Z×S.

where the previous Models (B) and (C) have been modified to include interaction terms indicated by × . To the extent that our stratification strategy has succeeded in balancing perfectly by the stratum indicator, Model (Ci) incurs no variance inflation and the effect of fitting the interaction is simply to remove the interaction from the residual mean square error at the cost of losing one degree of freedom for second order precision. Equivalent mean square error and second order precision effects will also occur for Model (Bi). However, for this model there will be an increase in the VIF compared to Model (B). We can argue as follows. To the VIF for Model (B) one degree of freedom has to be removed from the denominator because an extra term is being fitted. However, what is relevant to the VIF numerator is that a further element orthogonal to the covariate already fitted must be accounted for. We thus need to modify the term (1−2/π)/(N−4) in (3) to become (2−2/π)/(N−5) and thus obtain,

E[VIFBi,S]=1+2−2/πN−5≈1+1.363N−5 (11)

where VIFBi,S is the VIF for Model(Bi) given stratification. Our simulations suggest that it works well. Our feeling is that given the small sample sizes such models will not usually be of interest. If stratification has not been applied but a randomised design has been used, then the relevant formula for the VIF is

E[VIFBi,R]=1+2N−5. (12)

This is the same formula as expression (7). However, others who disagree may find consideration of expression (11) useful.

The reader who is interested in advice on models with interactions allowing for heterogeneity should consult the papers by Ye et al., Wang et al. and the paper by Van Lancker et al. already cited.18,36,37 We are assuming that the sample average treatment effect is of primary interest even if an interaction is fitted. Looking for subgroup effects has the consequence of making a rare disease even rarer.

8. Recommendations

Finally, we suggest the following way of looking at covariate adjustment and design of clinical trials.

  1. The intended analysis guides the design. The main reason for balancing a covariate is to reduce the non-orthogonality penalty, which is to say, to make the VIF closer to 1 than it otherwise would be, given that the covariate is in the model.

  2. Since the intended analysis guides the design and since the reason for stratification is to make this intended analysis more efficient (by balancing to the extent possible), one is not committed to putting the strata in the model.

  3. However, one is committed to putting the covariate in the model, since that is why the design was stratified. Thus Model (B), which is the approach recommended by Sullivan et al. 16 in their discussion is a consistent approach.

  4. Nevertheless, stratifying can contribute to reducing the conditional bias given an observed distribution of a covariate, X if the effect of the covariate is not approximately linear. This is a point that Qu has made. 4 (p235) Under such circumstances it will also have an effect on the true standard error of the estimate. To be consistent, such a reduction ought to be reflected in the estimated standard error for the treatment effect and this can be achieved by also fitting the stratum indicator. Thus, using Model (D) can also constitute a consistent position.

  5. However, one is neither limited to Model (B) nor to Model (D) and one should use a form (transformation, etc.) of the covariate that it is believed is valuable prognostically.

  6. Although it is often the case that a trial will be conducted with a treatment that has never or only rarely been studied before, it is rarely the case that a prognostic covariate is used that has not often been studied before. Thus, previous experience will be valuable in choosing a functional form and one should not assume that it is necessary to run a trial in order to know how to adjust the outcome.

  7. For small trials, using splines to model continuous covariates can be expensive in terms of degrees of freedom and hence in terms of the effect on the VIF and second order precision. An alternative that may be attractive, because it is less expensive in degrees of freedom, is to use fractional polynomials. 38

  8. Given that the intended analysis guides the design and given that pre-specification increases confidence in the results obtained, the intended form of adjustment should be pre-specified, as is required in any case by regulatory guidance. 39

  9. We have suggested that in the context of rare diseases, fitting the covariate interaction will usually not be of interest. However, if this is done, our advice is that users should take care that the definition of the treatment effect is that for the patient with the average baseline score and that the model is fitted accordingly. See the discussion in the Many Modes of Meta 40 pp542–545.

Acknowledgements

The authors thank our colleague Robin Ristl for drawing our attention to the paper by Qu4 and Tim Morris and two referees for helpful comments on earlier versions of our paper. FK and MP are members of RealiseD, which is supported by the Innovative Health Initiative Joint Undertaking (IHI JU) under grant agreement No 101165912. The JU receives support from the European Union's Horizon Europe research and innovation programme and COCIR, EFPIA, Europa Bío, MedTech Europe, and Vaccines Europe. Views and opinions expressed are those of the authors only. This publication reflects the author's views. Neither IHI nor the European Union, EFPIA, or any Associated Partners are responsible for any use that may be made of the information contained herein.

Appendix

Variance inflation factor (VIF)

We suppose for simplicity that we have n patients per arm with N=2n in total. The covariate values may be indicated by Xij,i=1,2,j=1…n , where i=1,2 indicates treatment group membership. The treatment dummy variable may be coded as −1/2,i=1,1/2,i=2 . Without loss of generality the covariate may be recoded in terms of deviations from the overall mean as xij=Xij−X¯.. .

The treatment indicator and recoded covariate sum to zero and are therefore orthogonal to the intercept indicator, which we may therefore ignore. We make no use of the stratum indicator at this stage of the argument. We can now formulate the predictors as a 2n×2 matrix, where the first column carries the treatment dummy variable and the second the covariate. The matrix of sums of square and cross-products is

x′x=(n2n2(x¯2.−x¯1.)n2(x¯2.−x¯1.)∑i=12∑j=1nxij2).

If this is inverted, then the element of the first row and column will be found to be VIF×(2/n) , where 2/n is the usual multiplier for the variance of the treatment effect when a covariate is not fitted and VIF is the variance inflation factor due to non-orthogonality and is given by:

VIF=∑i=12∑j=1nxij2∑i=12∑j=1nxij2−(n/2)(x¯2.−x¯1.)2.

However, in analysis of variance terms, the numerator is the total sum of squares and the denominator, being the difference between the total sum of squares and the between sum of squares, is the within sum of squares. Thus, we may re-write this expression as

VIF=SSWithin+SSBetweenSSWithin=SSWithinSSWithin+1(N−2)SSBetweenSSWithin/(N−2)=1+1(N−2)F1,N−2,

where the last term on the right is an F statistic. In general, an Fν,ω statistic has expectation ω/(ω−2) , so substituting N−2 for ω in the rightmost term above we get that the expected VIF is given by

E[VIF]=1+1N−4=N−3N−4.

Approximate variance inflation factor for the stratified case

The pdf of a standard Normal truncated at 0 is

ϕtrunc(Z=z)=2πe−z2/2,z≥0.

If a standard untruncated Normal has pdf ϕ() , then the mean of such a truncated Normal is

μtrunc=2π≈0.798

and the variance is

Vtrunc=1−μtrunc2=1−2π≈0.363.

The mean square between will have expectation proportional to 0.363, whereas without stratification it would have been proportional to one. However, under H0, the expectation of the total sum of squares will be unchanged. Thus, there will be an increase in the expected value of the sum of squares within. However, this effect may be expected to be small and given other uncertainties we suggest ignoring it.

A further issue is that the ratio of the mean square between to that within will not, however, have an F distribution. Nevertheless, as a rough guide to the expected maximum gain from this form of stratification the expression

E[VIFs]≈1+0.363N−4

seems not unreasonable.

An alternative argument for the multiplier of 0.363 proceeds like this. Stratification will achieve a median balance between the two treatment groups and this will go some way to achieving mean balance. If one is sampling from a Normal distribution, the mean is the sufficient statistic. Where that is the case, the coefficient of regression of median on mean must be one. It thus follows that the covariance between the median and the mean is equal to the variance of the mean. On the other hand, the variance of the median is asymptotically equal to π/2 times the variance of the mean. Thus, the asymptotic regression of mean on median is 2/π and the proportion of variation unexplained by the median is 1−2/π≈0.363 .

In small samples, the efficiency of the median is greater than the asymptotic efficiency. On the other hand, given the sequential nature of clinical trial recruitment, perfect stratification on the median will not be possible anyway. In practice it will be difficult to do better than this.

Correlation of the stratum indicator with polynomials

The following formula for the moments about the origin of a random variable Xt having the truncated Normal distribution is inspired by one for the standard Normal given by Rabbani 41 and uses the fact that moments of even powers will be the same as for the untruncated Normal.

E(Xtm)={2m2(m−12)!1πmodd2−m/2m!(m/2)!meven (13)

To calculate correlations, we shall also need moments about the mean: covariances and variances. However, calculation of this is simplified by the fact that the stratum indicator S has mean zero and variance 1 and that the mean of any odd power of X is also zero. The variance of an odd power will involve an even power, since these powers must be squared to calculate it. However, all these quantities can be given by moments of the truncated standard Normal. The formula for the correlation is then given by

ρ(m)={E(Xtm)E(Xt2m)modd0meven, (14)

where the moments in this formula are as given by (13).

The variance of Xm and correlation of the centred stratum indicator S and Xm are given by Var(Xm)=E(X2m)=(2m−1)!! and Var(S)=1 we obtain for odd m

Cor(S,Xm)=2∫0∞xmϕ(x)dxVar(Xm)=2m/2((m−1)/2)!π(2m−1)!!=2(m−1)!π(2m−1)!!. (15)

Here, ‘!!’ indicates the double factorial. (See, for example, Zwillinger and Kokoska 42 section 18.7.) For even m the correlation is 0. Thus, the correlations are decreasing in m for m odd. For example, for m=1,3,5 the correlations are

2π≈0.80,22(15π)≈0.41,and832105π=0.21.

In fact, as noted before for each successive odd power the correlation coefficient is approximately halved. To obtain the partial correlations, we use the recursive formula,

ρXY.W=ρXY.W∖(W0)−ρXW0.W∖(W0)ρW0Y.W∖(W0)1−ρXW0.W∖(W0)21−ρW0Y.W∖(W0)2 (16)

where W0∈W and W is a set of covariates. The partial correlations of X3,X5 given are

−13(π−2)≈−0.54,−7610(π−2)≈−0.35.

The partial correlation of X5 given X3 and X is

325(3π−7)≈0.43.

Footnotes

Ethical approval and informed consent statements: This work does not involve original research in any humans or animals and ethical approval was not required.

Funding: The authors disclosed receipt of the following financial support for the research, authorship and/or publication of this article: This work was supported by the Innovative Health Initiative Joint Undertaking (IHI JU) (grant number 101165912).

The authors declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.

Data availability statement: The data used in this paper are included in Section 6.

References

  • 1.Dodge Y. The Oxford dictionary of statistical terms. Oxford: Oxford, 2003. [Google Scholar]
  • 2.Holzhauer B, Adewuyi ET. “Super-covariates”: using predicted control group outcome as a covariate in randomized clinical trials. Pharm Stat 2023; 22: 1062–1075. [DOI] [PubMed] [Google Scholar]
  • 3.Rosenberger WF, Lachin JM. Randomization in clinical trials: theory and practice. 2nd ed. Hoboken: John Wiley & Sons, 2016. [Google Scholar]
  • 4.Qu Y. Issues for stratified randomization based on a factor derived from a continuous baseline variable. Pharm Stat 2011; 10: 232–235. [DOI] [PubMed] [Google Scholar]
  • 5.Gallo P. Center-Weighting issues in multicenter clinical trials. J Biopharm Stat 2001; 10: 145–163. [DOI] [PubMed] [Google Scholar]
  • 6.Nelder JA. A reformulation of linear models. J R Stat Soc A 1977; 140: 48–77. [Google Scholar]
  • 7.Siegfried S, Senn S, Hothorn T. On the relevance of prognostic information for clinical trials: a theoretical quantification. Biom J 2023; 65: 2100349. DOI: 10.1002/bimj.202100349 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Marquardt DW. Generalized inverses, ridge regression, biased linear estimation, and nonlinear estimation. Technometrics 1970; 12: 591–612. [Google Scholar]
  • 9.Senn SJ. Statistical issues in drug development. 3rd ed. Chichester: John Wiley & Sons, 2021, p.616. [Google Scholar]
  • 10.Harrell FE. Regression modeling strategies. New York: Springer, 2001. [Google Scholar]
  • 11.Cox DR, Reid N. The theory of the design of experiments. Boca Raton: Chapman and Hall CRC, 2000, p.323. [Google Scholar]
  • 12.Fedorov VV, Leonov SL. Optimal design for nonlinear response models. Boca Raton: CRC Press, 2013, p.371. [Google Scholar]
  • 13.Fisher RA. Co-operation in large-scale experiments. Suppl J R Stat Soc 1936; 3: 122–124. [Google Scholar]
  • 14.Gosset WS. Co-operation in large-scale experiments. Suppl J R Stat Soc 1936; III: 115–122. [Google Scholar]
  • 15.Senn SJ. Added values: controversies concerning randomization and additivity in clinical trials. Stat Med 2004; 23: 3729–3753. [DOI] [PubMed] [Google Scholar]
  • 16.Sullivan TR, Morris TP, Kahan BC, et al. Categorisation of continuous covariates for stratified randomisation: how should we adjust?. Stat Med 2024; 43: 2083–2095. DOI: 10.1002/sim.10060 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Hills M, Armitage P. The two-period cross-over clinical trial. Br J Clin Pharmacol 1979; 8: 7–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Van Lancker K, Bretz F, Dukes O. Covariate adjustment in randomized controlled trials: general concepts and practical considerations. Clin Trials (London, England) 2024; 21: 399–411. [DOI] [PubMed] [Google Scholar]
  • 19.Cochran WG. Analysis of covariance: its nature and uses. Biometrics 1957; 13: 261–281. [Google Scholar]
  • 20.Cox DR, McCullagh P. Some aspects of analysis of covariance. Biometrics 1982; 38: 541–554. [PubMed] [Google Scholar]
  • 21.Piantadosi S. Clinical trials: a methodologic perspective. 1st ed. New York: Wiley, 1997, p.687. [Google Scholar]
  • 22.Giardiello FM, Hamilton SR, Krush AJ, et al. Treatment of colonic and rectal adenomas with sulindac in familial adenomatous polyposis. N Engl J Med 1993; 328: 1313–1316. [DOI] [PubMed] [Google Scholar]
  • 23.Hee SW, Willis A, Tudur Smith C, et al. Does the low prevalence affect the sample size of interventional clinical trials of rare diseases? An analysis of data from the aggregate analysis of clinicaltrials.gov. Orphanet J Rare Dis 2017; 12: 44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Senn SJ. Testing for baseline balance in clinical trials. Stat Med 1994; 13: 1715–1726. [DOI] [PubMed] [Google Scholar]
  • 25.Stuart A, Ord K. Kendall’s advanced theory of statistics. 6th ed. London: Edward Arnold, 1994, p.676. [Google Scholar]
  • 26.Finney DJ. Stratification, balance and covariance. Biometrics 1957; 13: 373–386. [Google Scholar]
  • 27.Senn SJ. Consensus and controversy in pharmaceutical statistics (with discussion). J R Stat Soc Series D, Stat 2000; 49: 135–176. [Google Scholar]
  • 28.Altman DG, Royston JP. The hidden effect of time. Stat Med 1988; 7: 629–637. [DOI] [PubMed] [Google Scholar]
  • 29.Harrell FE, Jr., Lee KL, Pollock BG. Regression models in clinical studies: determining relationships between predictors and response. J Natl Cancer Inst 1988; 80: 1198–1202. [DOI] [PubMed] [Google Scholar]
  • 30.Royston P, Altman DG. Regression using fractional polynomials of continuous covariates - parsimonious parametric modeling. Appl Stat-J R Stat Soc Series C 1994; 43: 429–467. [Google Scholar]
  • 31.Royston P, Sauerbrei W. Stability of multivariable fractional polynomial models with selection of variables and transformations: a bootstrap investigation. Stat Med 2003; 22: 639–659. [DOI] [PubMed] [Google Scholar]
  • 32.Wetherill GB. Intermediate statistical methods. Berlin: Springer Science & Business Media, 2012. [Google Scholar]
  • 33.International Conference on Harmonisation. Statistical principles for clinical trials (ICH E9). Stat Med 1999; 18: 1905–1942. [PubMed] [Google Scholar]
  • 34.FDA. Adjusting for covariates in randomized clinical trials for drugs and biological products: guidance for industry. Silver Springs, MD: FDA; 2023, p. 9.
  • 35.Agency EM. Qualification opinion for prognostic covariate adjustment (PROCOVA). In: (CHMP) CfMPfHU (ed.). Amsterdam; 2022, pp. 1–33.
  • 36.Ye T, Shao J, Yi Y, et al. Toward better practice of covariate adjustment in analyzing randomized clinical trials. J Am Stat Assoc 2023; 118: 2370–2382. [Google Scholar]
  • 37.Wang B, Susukida R, Mojtabai R, et al. Model-robust inference for clinical trials that improve precision by stratified randomization and covariate adjustment. J Am Stat Assoc 2023; 118: 1152–1163. [Google Scholar]
  • 38.Binder H, Sauerbrei W, Royston P. Comparison between splines and fractional polynomials for multivariable model building with continuous covariates: a simulation study with continuous response. Stat Med 2013; 32: 2262–2277. [DOI] [PubMed] [Google Scholar]
  • 39.Lewis JA. Statistical principles for clinical trials (ICH E9): an introductory note on an international guideline. Stat Med 1999; 18: 1903–1942. [DOI] [PubMed] [Google Scholar]
  • 40.Senn SJ. The many modes of meta. Drug Inf J 2000; 34: 535–549. [Google Scholar]
  • 41.Rabbani S. Moments of the standard normal probability density function. https://www.srabbani.com/moments.pdf (2007, accessed June 12, 2024).
  • 42.Zwillinger D, Kokoska S. CRC standard probability and statistics tables and formulae. Boca Raton: Chapman and Hall, 2000, p.554. [Google Scholar]

Articles from Statistical Methods in Medical Research are provided here courtesy of SAGE Publications

RESOURCES