Skip to main content
American Journal of Epidemiology logoLink to American Journal of Epidemiology
. 2019 May 30;188(12):2213–2221. doi: 10.1093/aje/kwz136

Diagnosing Covariate Balance Across Levels of Right-Censoring Before and After Application of Inverse-Probability-of-Censoring Weights

John W Jackson 1,2,
PMCID: PMC7212402  PMID: 31145432

Abstract

Covariate balance is a central concept in the potential outcomes literature. With selected populations or missing data, balance across treatment groups can be insufficient for estimating marginal treatment effects. Recently, a framework for using covariate balance to describe measured confounding and selection bias for time-varying and other multivariate exposures in the presence of right-censoring has been proposed. Here, we revisit this framework to consider balance across levels of right-censoring over time in more depth. Specifically, we develop measures of covariate balance that can describe what is known as “dependent censoring” in the literature, along with its associated selection bias, under multiple mechanisms for right censoring. Such measures are interesting because they substantively describe the evolution of dependent censoring mechanisms. Furthermore, we provide weighted versions that can depict how well such dependent censoring has been eliminated when inverse-probability-of-censoring weights are applied. These results provide a conceptually grounded way to inspect covariate balance across levels of right-censoring as a validity check. As a motivating example, we applied these measures to a study of hypothetical “static” and “dynamic” treatment protocols in a sequential multiple-assignment randomized trial of antipsychotics with high dropout rates.

Keywords: attrition, covariate balance, dependent censoring, informative censoring, inverse-probability-of-censoring weights, IPCW, per-protocol effect, selection bias


Covariate balance is a central concept in the potential outcomes literature (1). In expectation, randomization balances all covariates in experimental studies, and propensity scores balance measured covariates in observational studies. When treatment groups have similar distributions of prognostic covariates that jointly affect or associate with outcomes and treatments, marginal contrasts between groups identify causal effects of treatment. Many experimental and observational studies of time-fixed treatments that employ propensity scores inspect covariate balance as a validity check (2, 3). See Jackson (4) for extensions to time-varying treatments. Although balance checks are limited under unmeasured confounding and misspecified causal models, they can be substantively meaningful and helpful in evaluating potential residual bias after adjustments for confounding (e.g., matching, subclassification, weighting). For example, imbalance can reflect positivity violations and effects of weight truncation (4, 5).

With missing data, balance across treatment groups can be insufficient for estimating marginal treatment effects (610). Those with complete versus incomplete data might have different distributions of covariates that correlate with the outcome, which can bias estimates of absolute and relative causal contrasts. Here, we revisit a previous framework of covariate balance for time-varying and other multivariate exposures in the presence of right censoring (4). Our contribution is to develop extensible measures of covariate balance that can substantively describe what is known as “dependent censoring” (7) under multiple mechanisms for right-censoring that evolve over time. Furthermore, we provide weighted versions that can depict how well such dependent censoring has been eliminated when inverse-probability-of-censoring weights (IPCW) (11) are applied to adjust for selection bias, as a conceptually grounded validity check.

As a motivating example, we applied these measures to a study of symptom trends under alternate antipsychotic treatment protocols in a randomized trial. This study had 2 censoring mechanisms, one of which is study attrition. The other is induced when the investigators estimate a per-protocol effect by artificially censoring those that deviate from protocol. We show how to describe both censoring mechanisms in the context of intention-to-treat, static, and dynamic treatment protocols. Our balance measures for (in)dependent censoring extend beyond intention-to-treat and per-protocol analyses of a trials (12) and their emulations (13) to other cohort designs (14).

We begin with notation and formal definitions of (in)dependent censoring and their role in causal identification. We apply these definitions to justify extensible measures of covariate balance across levels of censoring in unweighted and weighted data. We then demonstrate their use in our motivating example and discuss generalizations. The justifications in Web Appendixes 1–4 (available at https://academic.oup.com/aje) are given to highlight implications for balance, but are subsumed by the seminal work of Robins and colleagues (1518).

NOTATION

Let V(x) represent the value that the random variable V would take had the random variable X been set to value x. Let t equal discrete time (i.e., study visit number) with baseline indexed at t=0 and the end of follow-up indexed at t=T. Let an overbar represent the entire history vector of X from baseline through the indexed time (i.e., X¯t={X0,,Xt}). Let an underbar represent the vector of realizations from the indexed time through the end of follow-up (i.e., X_t+1={Xt+1,,XT}.) Let the expression VX|W denote conditional independence between V and X given W. Throughout, outcomes measured at time t+1 will be denoted as continuous Yt+1, treatment actually taken at time t will be denoted as binary At, and prognostic covariates, which might include earlier measures of the outcome, will be denoted as continuous or binary Lt. U represents unmeasured prognostic factors not used to guide treatment decisions. Censoring mechanisms for study dropout St by time t and artificial censoring by the investigator St by time t are indicator variables (1 = censored, 0 = not censored). We assume the time-specific order St,Lt,At,St such that if St=1 the rest are missing. Last, G is describes a treatment protocol of interest to the investigator, a collective set of time-specific treatment rules gt (i.e., g¯T=[g0(),,gT()]), that a person’s treatment history A¯t might be observed to be consistent with through time t (19). Static treatment protocols do not depend on Lt whereas dynamic ones do (20). See Web Appendix 1 for further details.

DEFINITION OF (IN)DEPENDENT CENSORING

We define a censoring mechanism to be a particular way in which a person’s data ceases to contribute to the analysis. One mechanism might be exit from a study for administrative reasons or loss to follow-up. Another mechanism might be due to the analytical strategy. For example, investigators might censor someone once their data are no longer consistent with a protocol of interest (17, 18, 21). The reasons that persons leave a study or switch treatments might be associated with outcomes. In what follows, we formally express this dependence between censoring and outcomes. Our arguments align with those used by Robins and other investigators in the context of survival analysis (6, 7), as well as missing data mechanisms (22) advanced by Rubin (23).

We define independent censoring as when, for all follow-up times, no censoring mechanism is associated with the history of prognostic covariates, measured or unmeasured, that are associated with subsequent outcomes. We define dependent censoring as when, for any follow-up time, any censoring mechanism is associated with the history of any prognostic covariate, measured or unmeasured. Under these definitions, censoring that is dependent will be associated with the history of prognostic covariates as well as future outcomes. Conversely, censoring that is independent will be associated with neither the history of prognostic covariates nor future outcomes. There is no way to empirically verify that independent censoring holds, even in randomized trials.

Implicitly, these definitions condition on those who have remained uncensored by any mechanism up to the considered time, because only those previously uncensored can become censored. Furthermore, for causal identification, these definitions condition on the history of treatment. Here, this simply means that definitions of (in)dependent censoring are specific to the treatment protocol arm, capturing the plausible possibility that the nature of dependent censoring varies by protocol arm. Participants in a placebo arm might exit for lack of efficacy, but participants in a protocol arm exit for intolerability. Last, these definitions condition on time because the way in which censoring mechanisms relate to the history of prognostic covariates might evolve over time. For example, attrition might be driven by adverse events and treatment efficacy early in a trial but towards the end be driven by social support or quality of life.

We can equivalently define (in)dependent censoring via the (in)dependence between censoring and either 1) the history of prognostic covariates or 2) potential outcomes, with factual outcomes realized under consistency with observed treatment history. With 2 mechanisms, attrition and departure from treatment protocol, we can define independent censoring as independence from the history of measured prognostic factors L¯t and unmeasured prognostic factors U¯t:

{L¯t,U¯t}St+1|S¯t=0¯,S¯t=0¯,G=g (1a)
{L¯t,U¯t}St|S¯t=0¯,S¯t1=0¯,G=g. (1b)

Equivalently, we can define independent censoring as independence from potential outcomes:

Y_t+1(g,s¯T=0¯,s¯T1=0¯)St+1|S¯t=0¯,S¯t=0¯,G=g (2a)
Y_t+1(g,s¯T=0¯,s¯T1=0¯)St|S¯t=0¯,S¯t1=0¯,G=g. (2b)

Dependent censoring is defined as when either of the above independencies do not hold. When we have measured all the prognostic factors responsible for the violation of independent censoring, we can define conditional forms of independent censoring:

U¯tSt+1|L¯t,S¯t=0¯,S¯t=0¯,G=g (3a)
U¯tSt|L¯t,S¯t=0¯,S¯t1=0¯,G=g. (3b)

Equivalently, we can also define conditional independent censoring as:

Y_t+1(g,s¯T=0¯,s¯T1=0¯)St+1|L¯t,S¯t=0¯,S¯t=0¯,G=g (4a)
Y_t+1(g,s¯T=0¯,s¯T1=0¯)St|L¯t,S¯t=0¯,S¯t1=0¯,G=g. (4b)

In 4a and 4b we include independence with the potential outcomes of the response Y_t+1(g,s¯T=0¯,s¯T1=0¯) alone, which is sufficient for causal identification of per-protocol effects under static protocols (19). We can alternatively define conditional independent censoring, also equivalent to 3a and 3b, that includes independence with potential outcomes of prognostic factors L_t+1(g,s¯T=0¯,s¯T1=0¯) for identification of effects under dynamic protocols: (19)

{Y_t+1(g,s¯T=0¯,s¯T1=0¯),L_t+1(g,s¯T=0¯,s¯T1=0¯)}St+1|L¯t,S¯t=0¯,S¯t=0¯,G=g (4c)
{Y_t+1(g,s¯T=0¯,s¯T1=0¯),L_t+1(g,s¯T=0¯,s¯T1=0¯)}St|L¯t,S¯t=0¯,S¯t1=0¯,G=g (4d)

(IN)DEPENDENT CENSORING AND CAUSAL IDENTIFICATION

Intention-to-treat estimand

These definitions are not just useful for substantively describing the censoring mechanisms at play. They are critical assumptions that underlie any causal analysis of right-censored data. For example, suppose we are interested in estimating the outcome trend under an intention-to-treat protocol in a baseline randomized trial with study attrition, as represented by the causal graph in Figure 1A and Web Figure 1. There is only one censoring mechanism at play here, that of study dropout. We want to estimate the outcome trend had, counter to fact, everyone been observed for the duration of the study and followed the intention-to-treat protocol. That is, we want to estimate Y¯T(g,s¯T=0¯). Those who initiate different protocols are exchangeable by virtue of the randomization. If time-varying prognostic factors Lt and At (as well as Ut) are independent of dropout St+1 given the protocol arm, the uncensored are exchangeable with the censored. This sequential exchangeability would be reflected by a lack of backdoor paths between G and Y¯T, and also between St+1 and Y_t+1 given S¯t=0¯ and G=g. Under dependent censoring where Ut is still independent of St+1, such backdoor paths would be present. But the paths between St+1 and Y_t+1 would be blocked upon further conditioning on Lt and At, rendering the uncensored and censored conditionally exchangeable. There are several ways to leverage such conditional exchangeability to identify our causal estimand. Interestingly, IPCW for study dropout accomplishes this by producing a pseudopopulation where there is independent censoring with respect to the measured covariates (15). More directly, if conditional independent censoring holds in the crude study data along with positivity, independent censoring with respect to measured and unmeasured covariates holds in the weighted population. It does this by reweighting the data in such a way that, for each level of Lt and At among the previously uncensored, the proportion of those who depart versus remain in the study does not depend on Lt or At. In our context, the stabilized weights (11) Wt are:

Wt=j=0tPr[Sj+1=s|S¯j=0¯,G=g]Pr[Sj+1=s|Lj,Aj,S¯j=0¯,G=g] (5)

Figure 1.

Figure 1.

Causal graph relating protocol arm at baseline G, actual treatment At taken, prognostic covariates Lt, unmeasured prognostic factors U, subsequent outcomes Yt+1, and censoring due to study dropout St and/or artificial censoring St for an intention-to-treat analysis (A), static per-protocol analysis (B), dynamic per-protocol analysis (C).

Static per-protocol estimand

Suppose that we were now interested in estimating the outcome trend under adherence to a static treatment protocol (always take a given treatment) in a baseline randomized trial with study attrition, as represented by the causal graph in Figure 1B and Web Figure 2, where there is no unmeasured confounding of time-varying treatment At. We want to estimate the outcome trend had, counter to fact, everyone been observed for the duration of the study and followed the protocol. That is, we want to estimate Y¯T(g,s¯T=0¯,s¯T1=0¯) and do so by artificially censoring persons when they their actual treatment conflicts with the protocol. With treatments randomized at baseline, those who initiate the protocol versus those who do not are exchangeable at baseline. However, there are 2 censoring mechanisms at play here, that of study dropout and also that of protocol nonadherence. If time-varying factors Lt (and also Ut) are independent of dropout St+1 and nonadherence St in each arm, as in the case of independent censoring, then for both censoring mechanisms the censored are exchangeable with the uncensored. Under dependent censoring where Ut is still independent of St+1 and also independent of St through its lack of association with At, we have Lt associated with St+1 or with St through its association with At. This leads to backdoor paths between one or both censoring mechanisms (St+1 or St) and the outcome Y¯t+1. However, conditional on Lt, these pathways would be blocked, such that the censored and uncensored are conditionally exchangeable. Interestingly, joint IPCW for study dropout and static protocol nonadherence identifies our counterfactual trend by producing a pseudopopulation where there is independent censoring with respect to measured covariates. It does this by reweighting the data in such a way that, for each level of Lt among those previously uncensored by any mechanism, the proportion of those who drop out versus remain in the study, and adhere versus do not adhere to protocol, does not depend on Lt. In our context, the stabilized weights (17, 18) are:

Wt=j=0tPr[Sj+1=s|S¯j=0¯,S¯j=0¯,G=g]Pr[Sj+1=s|Lj,S¯j=0¯,S¯j=0¯,G=g]×j=0tPr[Sj=s|S¯j=0¯,S¯j1=0¯,G=g]Pr[Sj=s|Lj,S¯j=0¯,S¯j1=0¯,G=g] (6)

Dynamic per-protocol estimand

These concepts also extend to the case for estimating the outcome trend under adherence to a dynamic treatment protocol that depends on prognostic factors (e.g., initiate a given treatment, switch to another when a prognostic factor drops below a certain threshold), as in Figure 1C and Web Figure 3, where there is no unmeasured confounding of time-varying treatment At. Suppose we are interested in estimating this counterfactual trend by censoring persons when their treatment conflicts with the protocol. Even though treatments are randomized, the protocol of interest depends on prognostic factors Lt, so those who initiate protocol versus those who do not are only conditionally exchangeable at baseline given L0. As with the static case, there are 2 censoring mechanisms at play here, that of study dropout and that of protocol nonadherence, which depends on Lt. Thus, censoring will always be dependent in the crude data, even when the treatments At are randomized. Provided that Ut is independent of St+1 and At, conditioning on Lt blocks backdoor paths between each censoring mechanism (St+1 and St) and the outcome Y_t+1, making the uncensored exchangeable with the censored. The joint IPCW for study dropout and dynamic protocol nonadherence identify our counterfactual trend by producing a pseudopopulation with independent censoring with respect to measured covariates. In our context, the stabilized weights are of the same form as equation 6.

MEASURES OF COVARIATE BALANCE FOR (IN)DEPENDENT CENSORING

The definitions outlined above have direct implications for using covariate balance to describe (in)dependent censoring and the nature of its associated selection bias. The general approach is, for each censoring mechanism, to compare at each time the distribution of covariate history among those who become censored versus those who do not, among those following the same treatment protocol who have not yet been censored by any mechanism. This definition implies an ordering of mechanisms within each time interval that is known a priori, could be ascertained, or could be arbitrarily chosen for mechanisms whose timing within time intervals coincides. They also highlight why it is important to distinguish mechanisms: The prognostic factors involved in dependent censoring plausibly vary by mechanism.

Using the standardized mean difference, we can define the following balance measures. The first measure assesses balance across levels of censoring due to study dropout, among those in arm G=g who have not dropped out of the study or deviated from protocol up to that point:

E[Lt×Wt,st+1=1,st=0×I(St+1=1,St=0)|S¯t=0¯,S¯t1=0¯,G=g]E[Lt×Wt,st+1=0,st=0×I(St+1=0,St=0)|S¯t=0¯,S¯t1=0¯,G=g]σSt+1=0|S¯t=0¯,S¯t=0¯,G=g (7a)

where Wt,st+1=s',st=0=1 for assessment of balance in the unweighted data. If we used the weights specified in equation 6 evaluated at the person’s observed values for censoring due to study dropout and protocol nonadherence, we would have:

Wt,st+1=s',st=0=Pr[St+1=s'|S¯t=0¯,S¯t=0¯,G=g]Pr[St+1=s'|Lt,S¯t=0¯,S¯t=0¯,G=g]×j=0t1Pr[Sj+1=0|S¯j=0¯,S¯j=0¯,G=g]Pr[Sj+1=0|Lj,S¯j=0¯,S¯j=0¯,G=g]×j=0tPr[Sj=0|S¯j=0¯,S¯j1=0¯,G=g]Pr[Sj=0|Lj,S¯j=0¯,S¯j1=0¯,G=g].

The second measure assesses balance across levels of censoring due to protocol nonadherence, among those in arm G=g who have not dropped out the study or deviated from protocol up to that point:

E[Lt×Wt,st+1=s,st=1×I(St+1=s,St=1)|S¯t=0¯,S¯t1=0¯,G=g]E[Lt×Wt,st+1=s,st=0×I(St+1=s,St=0)|S¯t=0¯,S¯t1=0¯,G=g]σSt=0|S¯t=0¯,S¯t1=0¯,G=g. (7b)

where Wt,st+1=s,st=s'=1 for assessment of balance in the unweighted data. If we used the weights specified in equation 6 evaluated at the person’s observed values for censoring due to study dropout and protocol nonadherence, we would have:

Wt,st+1=s,st=s=Pr[St+1=s|S¯t=0¯,S¯t=0¯,G=g]Pr[St+1=s|Lt,S¯t=0¯,S¯t=0¯,G=g]×j=0t1Pr[Sj+1=0|S¯j=0¯,S¯j=0¯,G=g]Pr[Sj+1=0|Lj,S¯j=0¯,S¯j=0¯,G=g]×Pr[St=s'|S¯t=0¯,S¯t1=0¯,G=g]Pr[St=s'|Lt,S¯t=0¯,S¯t1=0¯,G=g]×j=0t1Pr[Sj=0|S¯j=0¯,S¯j1=0¯,G=g]Pr[Sj=0|Lj,S¯j=0¯,S¯j1=0¯,G=g].

Note that for both measures, for efficiency we estimate the standard deviation σ of Lt among the currently uncensored in G=g, but other choices are possible, such as pooling across censoring levels. Note also that the weights in 7b are depicted as they are to make clear that assessing balance across levels of protocol nonadherence is not restricted to those present at the next study visit. They also make clear that the measures capture the effects of the joint weights on balance, not just their individual components. As for interpretation, following Ho et al. (24), independent censoring would call for minimizing imbalance without limit.

We need measures 7a and 7b when both mechanisms for censoring are present. In the case of censoring due only to study dropout, we would need only the relevant balance measure: 7a for study dropout or 7b for nonadherence. If there were more than 2 censoring mechanisms, we could expand our approach by evaluating time-specific balance for each form of censoring, conditional on the protocol arm and not having been previously censored by any mechanism. We argue in Web Appendix 5 that, although possible (4), it is not necessary to evaluate balance across finer gradations of censoring levels (e.g., Web Figure 4), such as documented reasons for study dropout or reasons for treatment switching.

APPLICATION TO MOTIVATING EXAMPLE

Clinical management of schizophrenia requires balancing treatment response and adverse effects, such as weight gain and onset of movement disorders. Because treatment-specific response unpredictably varies by patient, clinicians usually adapt a treatment strategy until a suitable course is found. The Clinical Antipsychotic Trial of Intervention Effectiveness schizophrenia study (CATIE; NCT00014001) was a landmark comparative effectiveness trial (n = 1,462) in the United States that evaluated effects of first- and second-generation antipsychotics over 18 months of follow-up of persons with schizophrenia (25, 26). Treatment choices were randomized at baseline and again upon treatment failure or intolerability at the patient and treating physician’s discretion. Despite this robust design, high attrition rates across treatment protocol arms (44%–57%) might bias assessments of symptom trajectories, such as the Positive and Negative Syndrome Scale (PANSS) (27). High attrition rates are common in trials of schizophrenia patients (28, 29).

A previous study (30) used the same study’s data to study the effects of several hypothetical treatment protocols on 18-month PANSS scores among patients without tardive dyskinesia, a movement disorder that occurs more frequently with first-generation antipsychotics (31). We consider 3 of these protocols that trade between PANSS score reduction and risk of tardive dyskinesia. The first protocol is static and always treats with the first-generation antipsychotic perphenazine. The second protocol, also static, always treats with a second-generation antipsychotic. The third protocol depends on a threshold PANSS score, θ. If baseline PANSS score is above threshold, treatment begins with perphenazine. Once the PANSS score drops below threshold, treatment continues with a second-generation antipsychotic from that point forward. Conversely, if baseline PANSS score is below threshold, the third protocol always treats with a second-generation antipsychotic. Protocols 2 and 3 allow patients to switch between second-generation antipsychotics as needed. For estimation, we used the “clone-and-censor” strategy with IPCW for attrition and protocol deviation (18) to examine trends in PANSS score under these protocols with the threshold θ set to 50 in protocol 3 (see Li et al. (32) for an overview). Following the previous study (30), we excluded those with tardive dyskinesia at baseline, but we differed by choosing to not censor patients who were randomized to clozapine or ziprasidone or who entered the open-label phase of the study. Our use of IPCW rather than imputation to address attrition also differs from the previous study (30). The purpose of our application is demonstrative. Web Appendix 6 and Web Table 1 provide further analytical details.

Table 1 shows that, with the exception of baseline PANSS scores, baseline covariates had similar distributions across protocol arms. Figure 2A and 2B show, respectively, the absolute distributions of study dropout and protocol nonadherence over time in each arm for selected times. Study dropout is right-skewed over time. The nonadherence patterns show that for those who initiated perphenazine (arm 1) or second-generation antipsychotics (arm 2), very few departed from these strategies until month 3. The distribution of the weights (Web Figure 5) show somewhat well-behaved weights for arm 2 but not the others. Figures 3 and 4 show, respectively, the arm-specific covariate balance across study dropout and protocol nonadherence, before and after weighting. Due to space limitations, these display data for months 1, 6, and 12 (see Web Figures 6 and 7 for all months). Each plot reports standardized mean differences (x-axis) comparing the means of those who drop out/depart from protocol with those who do not at a given month. Panels arrange the data by month in increasing order. The standardized means for each covariate (y-axis) are further distinguished by protocol arm (shapes). These distinctions allow us to examine whether the covariate associations with censoring vary over time and protocol arm.

Table 1.

Baseline Characteristics and Censoring Frequency of Those Initiating Different Hypothetical Treatment Protocols, Using Data From the Clinical Antipsychotic Treatment Intervention Effectiveness Studya, United States, 2000–2004

Characteristic Protocol 1b Protocol 2c Protocol 3d
No. % Mean (SD) No. % Mean (SD) No. % Mean (SD)
Sample size 151 100 1,055 100 205 100
Age, years 38.8 (11.0) 39.3 (10.9) 39.9 (10.5)
Male sex 103 68 781 74 145 70
Female sex 48 32 274 26 61 30
Race/ethnicity
 White 90 60 633 60 120 58
 Black 52 34 372 35 75 36
 AIAN/Asian/NHPI/Mixed race 9 6 50 5 11 5
High-school graduate 113 75 787 75 159 78
Unemployed 124 82 882 84 161 78
Recent hospitalization 46 30 295 28 56 27
Phase 1A stratume 0 0 0 0 0 0
Ziprasidone stratumf 151 100 591 56 179 87
PANSSg 75.1 (18.5) 75.6 (17.5) 66.5 (21.6)
CGI severityg 3.9 (0.9) 4 (1.0) 3.5 (1.0)
Calgary depression 1.9 (0.8) 1.9 (0.9) 1.7 (0.8)
Quality of life 2.7 (1.1) 2.8 (1.1) 3 (1.2)
Body weight, lb 192.6 (46.3) 195.8 (47.7) 194.9 (46.5)
Simpson-Agnes EPS 0.2 (0.3) 0.2 (0.3) 0.2 (0.3)
Drug use scale 1.4 (0.7) 1.4 (0.7) 1.3 (0.6)
Censoring reason
 Protocol nonadherence 57 38 136 13 69 33
 Study discontinuation 74 49 464 44 93 45
 Either reason 120 79 586 56 152 74

Abbreviations: AIAN, American Indian and Alaskan Native; CGI, Clinical Global Impressions; EPS, extrapyramidal symptoms; NHPI, Native Hawaiian and Other Pacific Islander; PANSS, Positive and Negative Syndrome Scale; SD, standard deviation.

a A multisite pragmatic trial (ClinicalTrials.gov: NCT00014001).

b Protocol 1: Always take perphenazine, a first-generation antipsychotic.

c Protocol 2: Always take a second-generation antipsychotic.

d Protocol 3: Initiate perphenazine if PANSS ≥50, switch to second-generation antipsychotic when PANSS <50; always second-generation antipsychotic if PANSS <50.

e Patients with tardive dyskinesia.

f Patients randomized after approval of (and thus eligibility to be assigned) ziprasidone.

g Positive and Negative Syndrome Scale (27) and Clinical Global Impressions (45).

Figure 2.

Figure 2.

Histogram of censoring events over time, using data from Clinical Antipsychotic Trials of Intervention Effectiveness, United States, 2000–2004. A) Study dropout; B) protocol departure. White: protocol arm 1, always perphenazine; gray: protocol arm 2, always second-generation antipsychotic; black: protocol arm 3, dynamic.

Figure 3.

Figure 3.

Covariate balance across levels of censoring at months 1, 6, and 12, using data from Clinical Antipsychotic Trials of Intervention Effectiveness, United States, 2000–2004. A) Due to study dropout in the crude data; B) after applying inverse-probability-of-censoring weights. Reference lines for the standardized mean difference are located at −0.25, 0, and 0.25. Some standardized mean differences are beyond the limits of (−1,1); see Web Figure 6 for all times and for all months. Others could not be calculated due to lack of covariate variation among the uncensored (see Web Table 2 for these unstandardized differences). Circles: protocol arm 1, always perphenazine; triangles: protocol arm 2, always second-generation antipsychotic; squares: protocol arm 3, dynamic. CGI, Clinical Global Impressions (45); EPS, extrapyramidal symptoms; PANSS, Positive and Negative Syndrome Scale (27).

Figure 4.

Figure 4.

Covariate balance across levels of censoring at months 1, 6, and 12, using data from Clinical Antipsychotic Trials of Intervention Effectiveness, United States, 2000–2004. A) Due to protocol nonadherence in the crude data; B) after applying inverse-probability-of-censoring weights. Reference lines for the standardized mean difference are located at −0.25, 0, and 0.25. Some standardized mean differences are beyond the limits of (−1,1); see Web Figure 7 for all times and for all months. Others could not be calculated due to lack of covariate variation among the uncensored (see Web Table 2 for these unstandardized differences). Circles: protocol arm 1, always perphenazine; triangles: protocol arm 2, always second-generation antipsychotic; squares: protocol arm 3, dynamic. CGI, Clinical Global Impressions (45); EPS, extrapyramidal symptoms; PANSS, Positive and Negative Syndrome Scale (27).

In Figure 3A, we see that those who dropped out had higher levels of illness severity, higher use of illicit substances, lower declines in PANSS scores, and less weight gain. Speculatively, these clinical parameters are consistent with poorer treatment effectiveness, which can occur with concomitant illicit substance use. These relationships vary across arm and also over time, with larger differences between dropouts versus nondropouts towards the end of the study. In Figure 3B we see that after applying IPCW, these differences were ameliorated in many cases. In other cases, they persisted or were exacerbated but there is no clear pattern. In Figure 4A we see that, initially, those who deviated from the dynamic protocol had lower depression symptoms, illness severity, and weight gain, suggesting that those who adhere to protocol early on are sicker. From month 6 on, those who depart from protocol had, at times, higher levels of illness severity, higher drug use, and lower quality of life. Those who departed from the static protocol to always treat with a second-generation antipsychotic did not often differ greatly from those who adhered, although there were some exceptions across covariates at certain times. Those who departed from the static protocol to always treat with perphenazine also showed fewer differences earlier on in the trial. Later, those who departed had higher levels of depression, illicit drug use, and weight gain. In Figure 4B, we can see that after applying IPCW these differences are sometimes ameliorated. Particularly for those following protocol 1, the smallest group, differences persist, or are exacerbated or qualitatively change direction.

We did not observe significant differences in PANSS score trends across protocol arms at the α=0.05 significance level, either before or after weighting (see Web Figure 8 for estimated trends). However, both sets of results likely suffer from some degree of selection bias and should be interpreted with caution. The inability of IPCW to fully remove covariate differences between the censored and uncensored is not altogether unsurprising, given the small size of the groups initiating protocols 1 and 3. It can be difficult to model heterogeneous selection mechanisms over time well with small samples, and positivity violations might also be present. Our findings support the decision of the previous studies of these regimes (30, 33) to pursue imputation methods for dealing with the high rates of study attrition in the Clinical Antipsychotic Trial of Intervention Effectiveness schizophrenia study.

DISCUSSION

Assumptions about the censoring processes are essential for identification, and the (in)dependent nature of censoring will be reflected in balance of covariates across levels of censoring. We can never be sure that independent censoring holds in the raw data. However, we can reject it when measured prognostic factors associate with censoring, which can substantively enrich our understanding of it. Further, we can use substantive knowledge to consider the causal architecture of censoring and measure sufficient prognostic factors to achieve conditional independent censoring. Finally, if we use IPCW to correct for dependent censoring, we can verify whether our estimated weights produce independent censoring with respect to measured covariates in the pseudopopulation. We have theoretically justified covariate balance measures across levels of multiple censoring mechanisms for each of these purposes and demonstrated their utility.

The present work could be expanded in several ways. Our measures of covariate balance assess only the balance of marginal covariate distributions at the mean, which is appropriate only when the underlying structural model lacks nonlinearities and interactions in the covariates (34). Given that the measures need only condition on the protocol arm and being previously uncensored, it would be straightforward to adapt more detailed assessments of covariate balance that have been considered in the propensity-score literature (35, 36). These would include balance of product terms, cumulative distributions, the Kolmogorov distance, and kernel-based estimators, among others. Second, the plots themselves might become unwieldy when there are several points of follow-up and rich covariate history. Previous work for plotting and summarizing covariate balance with high-dimensional measurements could easily be adapted here (4). The mean values and distribution of time-specific stabilized weights (5) are compact, easy to compute, and informative about aggregate balance (4). However, they do not substantively describe how the censored differ from the uncensored. Another important avenue for future work is to explore potential for bias, given that there might be little bias when the proportion of dropout is very low. There is an emerging literature exploring the presence, direction, and magnitude of bias under various causal structures for selection bias, including the absence of scale-dependent treatment effect heterogeneity (on censoring and/or outcomes) (9, 10, 37, 38). This literature could be used to define bias metrics for dependent censoring that incorporate a priori substantive knowledge about the censoring mechanisms. Concerning our restriction to a time-fixed exposure, our balance measures could be expanded to condition on a function of time-varying exposure history (4), as would be relevant for epidemiologic cohort designs that do not emulate trials (14). Last, our results could also be incorporated into efforts that use covariate balance to directly optimize estimation of causal effects (3943).

Covariate balance can be useful for describing multiple censoring mechanisms over time, not only for describing potential selection bias but also residual selection bias in the application of IPCW. See the example code in the Web Appendix for implementation in R (R Foundation for Statistical Computing, Vienna, Austria) (44).

Supplementary Material

Jackson_Web_Material_Final_kwz136

ACKNOWLEDGMENTS

Author affiliations: Department of Epidemiology, Johns Hopkins Bloomberg School of Public Health, Baltimore, Maryland (John W. Jackson); and Department Mental Health, Johns Hopkins Bloomberg School of Public Health, Baltimore, Maryland (John W. Jackson).

This study used data from the National Institutes of Mental Health National Data Archive and was partially supported by the National Center for Advancing Translational Science (grant UL1 TR001102) and National Heart, Lung, and Blood Institute (grant K01HL145320).

I thank Erin Schnellinger for assistance with assembling the confoundr R package, as well as Dr. Linda Valeri, Dr. Sharon-Lise Normand, Jake Spertus, Dr. Hardeep Ranu, Dr. Gary Gray, and Dr. Marsha Willcox for advice and support as part of the Open Translational Science in Schizophrenia (OPTICS) project.

The funders had no role in study design, conduct, interpretation or reporting of this study.

Conflict of interest: none declared.

Abbreviations

IPCW

inverse-probability-of-censoring weights

PANSS

Positive and Negative Syndrome Scale

REFERENCES

  • 1. Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika. 1983;70(1):41–55. [Google Scholar]
  • 2. Imai K, King G, Stuart EA. Misunderstandings between experimentalists and observationalists about causal inference. J R Stat Soc Ser A Stat Soc. 2008;171(2):481–502. [Google Scholar]
  • 3. Jackson JW, Schmid I, Stuart EA. Propensity scores in pharmacoepidemiology: beyond the horizon. Curr Epidemiol Rep. 2017;4(4):271–280. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Jackson JW. Diagnostics for confounding of time-varying and other joint exposures. Epidemiology. 2016;27(6):859–869. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Cole SR, Hernán MA. Constructing inverse probability weights for marginal structural models. Am J Epidemiol. 2008;168(6):656–664. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Robins JM, Rotnitzky A. Recovery of information and adjustment for dependent censoring using surrogate markers In: AIDS Epidemiol. Boston, MA: Birkhäuser Boston; 1992:297–331. [Google Scholar]
  • 7. Robins JM, Finkelstein DM. Correcting for noncompliance and dependent censoring in an AIDS Clinical Trial with inverse probability of censoring weighted (IPCW) log-rank tests. Biometrics. 2000;56(3):779–788. [DOI] [PubMed] [Google Scholar]
  • 8. Greenland S. Response and follow-up bias in cohort studies. Am J Epidemiol. 1977;106(3):184–187. [DOI] [PubMed] [Google Scholar]
  • 9. Hernán MA. Invited commentary: selection bias without colliders. Am J Epidemiol. 2017;185(11):1048–1050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Daniel RM, Kenward MG, Cousens SN, et al. Using causal diagrams to guide analysis in missing data problems. Stat Methods Med Res. 2012;21(3):243–256. [DOI] [PubMed] [Google Scholar]
  • 11. Robins JM, Hernán MÁ, Brumback B. Marginal structural models and causal inference in epidemiology. Epidemiology. 2000;11(5):550–560. [DOI] [PubMed] [Google Scholar]
  • 12. Hernán MA, Robins JM. Per-protocol analyses of pragmatic trials. N Engl J Med. 2017;377(14):1391–1398. [DOI] [PubMed] [Google Scholar]
  • 13. Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758–764. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Jackson JW, García-Albéniz X. Studying the effects of nonindicated medications on cancer: etiologic versus action-focused analysis of epidemiologic data. Cancer Epidemiol Biomarkers Prev. 2018;27(5):520–524. [DOI] [PubMed] [Google Scholar]
  • 15. Murphy SA, van der Laan MJ, Robins JM, et al. Marginal mean models for dynamic regimes. J Am Stat Assoc. 2001;96(456):1410–1423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Orellana L, Rotnitzky A, Robins JM. Dynamic regime marginal structural mean models for estimation of optimal dynamic treatment regimes, part II: proofs of results. Int J Biostat. 2010;6(2):Article 9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Hernán MA, Lanoy E, Costagliola D, et al. Comparison of dynamic treatment regimes via inverse probability weighting. Basic Clin Pharmacol Toxicol. 2006;98(3):237–242. [DOI] [PubMed] [Google Scholar]
  • 18. Cain LE, Robins JM, Lanoy E, et al. When to start treatment? A systematic approach to the comparison of dynamic regimes using observational data. Int J Biostat. 2010;6(2):Article 18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Hernán MA, Robins JM. Causal Inference. Forthcoming ed Boca Raton, FL: Chapman & Hall/CRC, 2019. [Google Scholar]
  • 20. Young JG, Herńan MA, Robins JM. Identification, estimation and approximation of risk under interventions that depend on the natural value of treatment using observational data. Epidemiol Methods. 2014;3(1):1–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Seeger JD, Williams PL, Walker AM. An application of propensity score matching using claims data. Pharmacoepidemiol Drug Saf. 2005;14(7):465–476. [DOI] [PubMed] [Google Scholar]
  • 22. Seaman S, Galati J, Jackson D, et al. What is meant by “missing at random”? Stat Sci. 2013;28(2):257–268. [Google Scholar]
  • 23. Rubin DB. Inference and missing data. Biometrika. 1976;63(3):581–592. [Google Scholar]
  • 24. Ho DE, Imai K, King G, et al. Matching as nonparametric preprocessing for reducing model dependence in parametric causal inference. Polit Anal. 2007;15(3):199–236. [Google Scholar]
  • 25. Stroup TS, McEvoy JP, Swartz MS, et al. The National Institute of Mental Health Clinical Antipsychotic Trials of Intervention Effectiveness (CATIE) project: schizophrenia trial design and protocol development. Schizophr Bull. 2003;29(1):15–31. [DOI] [PubMed] [Google Scholar]
  • 26. Swartz MS, Stroup TS, McEvoy JP, et al. What CATIE found: results from the schizophrenia trial. Psychiatr Serv. 2008;59(5):500–506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Kay SR, Fizbein A, Opler LA. The Positive and Negative Syndrome Scale (PANSS) for schizophrenia. Schizophr Bull. 1987;13(2):261–276. [DOI] [PubMed] [Google Scholar]
  • 28. Wahlbeck K, Tuunainen A, Ahokas A, et al. Dropout rates in randomised antipsychotic drug trials. Psychopharmacology (Berl). 2001;155(3):230–233. [DOI] [PubMed] [Google Scholar]
  • 29. Kemmler G, Hummer M, Widschwendter C, et al. Dropout rates in placebo-controlled and active-control clinical trials of antipsychotic drugs: a meta-analysis. Arch Gen Psychiatry. 2005;62(12):1305–1312. [DOI] [PubMed] [Google Scholar]
  • 30. Shortreed SM, Moodie EE. Estimating the optimal dynamic antipsychotic treatment regime: evidence from the sequential multiple assignment randomized CATIE schizophrenia study. J R Stat Soc Ser C Appl Stat. 2012;61(4):577–599. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Margolese HC, Chouinard G, Kolivakis TT, et al. Tardive dyskinesia in the era of typical and atypical antipsychotics. Part 2: incidence and management strategies in patients with schizophrenia. Can J Psychiatry. 2005;50(11):703–714. [DOI] [PubMed] [Google Scholar]
  • 32. Li X, Young JG, Toh S. Estimating effects of dynamic treatment strategies in pharmacoepidemiologic studies with time-varying confounding: a primer. Curr Epidemiol Rep. 2017;4(4)288–297. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Shortreed SM, Laber L, Scott Stroup T, et al. A multiple imputation strategy for sequential multiple assignment randomized trials. Stat Med. 2014;33(24):4202–4214. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. de los Angeles Resa M, Zubizarreta JR. Evaluation of subset matching methods and forms of covariate balance. Stat Med. 2016;35(27):4961–4979. [DOI] [PubMed] [Google Scholar]
  • 35. Austin PC, Stuart EA. Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Stat Med. 2015;34(28):3661–3679. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Zhu Y, Savage JS, Ghosh D. A kernel-based metric for balance assessment. J Causal Inference. 2018;6(2):20160029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Nguyen TQ, Dafoe A, Ogburn EL. The magnitude and direction of collider bias for binary variables [published online ahead of print March 12, 2019]. Epidemiol Methods. ( 10.1515/em-2017-0013). [DOI] [Google Scholar]
  • 38. Shahar DJ, Shahar E. A theorem at the core of colliding bias. Int J Biostat. 2017;13(1). [DOI] [PubMed] [Google Scholar]
  • 39. Hainmueller J. Entropy balancing for causal effects: a multivariate reweighting method to produce balanced samples in observational studies. Polit Anal. 2012;20(1):25–46. [Google Scholar]
  • 40. Zubizarreta JR. Stable weights that balance covariates for estimation with incomplete outcome data. J Am Stat Assoc. 2015;110(511):910–922. [Google Scholar]
  • 41. Imai K, Ratkovic M. Robust estimation of inverse probability weights for marginal structural models. J Am Stat Assoc. 2015;110(511):1013–1023. [Google Scholar]
  • 42. Wong RKW, Chan KCG. Kernel-based covariate functional balancing for observational studies. Biometrika. 2018;105(1):199–213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Griffin BA, McCaffrey D, Almirall D, et al. Chasing balance and other recommendations for improving nonparametric propensity score models. J Causal Inference. 2017;5(2):20150026. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Jackson JW, Schnellinger E confoundr R package. https://github.com/jwjackson/confoundr Accessed May 10, 2019.
  • 45. Schneider LA, Olin JT, Doody RS, et al. Validity and reliability of the Alzheimer’s Disease Cooperative Study—Clinical Global Impression of Change. The Alzheimer’s Disease Cooperative Study. Alzheimer Dis Assoc Disord. 1997;11(suppl 2):S22–S32. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Jackson_Web_Material_Final_kwz136

Articles from American Journal of Epidemiology are provided here courtesy of Oxford University Press

RESOURCES