Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Dec 1.
Published in final edited form as: Biometrics. 2023 Apr 3;79(4):3126–3139. doi: 10.1111/biom.13850

Efficient and flexible estimation of natural direct and indirect effects under intermediate confounding and monotonicity constraints

Kara E Rudolph 1,*, Nicholas Williams 1, Iván Díaz 2
PMCID: PMC11037503  NIHMSID: NIHMS1984132  PMID: 36905172

Abstract

Natural direct and indirect effects are mediational estimands that decompose the average treatment effect and describe how outcomes would be affected by contrasting levels of a treatment through changes induced in mediator values (in the case of the indirect effect) or not through induced changes in the mediator values (in the case of the direct effect). Natural direct and indirect effects are not generally point-identified in the presence of a treatment-induced confounder; however, they may be identified if one is willing to assume monotonicity between the treatment and the treatment-induced confounder. We argue that this assumption may be reasonable in the relatively common encouragement-design trial setting where the intervention is randomized treatment assignment and the treatment-induced confounder is whether or not treatment was actually taken/adhered to. We develop efficiency theory for the natural direct and indirect effects under this monotonicity assumption, and use it to propose a nonparametric, multiply robust estimator. We demonstrate the finite sample properties of this estimator using a simulation study, and apply it to data from the Moving to Opportunity Study to estimate the natural direct and indirect effects of being randomly assigned to receive a Section 8 housing voucher—the most common form of federal housing assistance—on risk developing any mood or externalizing disorder among adolescent boys, possibly operating through various school and community characteristics.

Keywords: Efficient influence function, Natural indirect effects, Mediation, Monotonicity

1. Introduction

Researchers are frequently interested in summarizing mediational effects across individuals (or other units). For example, in trying to understand reasons underlying an unexpected treatment effect, it may be helpful to decompose the overall effect into the portion operating through some intermediate variables (i.e., mediators) hypothesized to be responsible for the unexpected overall effect—called the indirect effect—and the portion not operating through those intermediate variables—called the direct effect.

Natural direct and indirect effects (NDE, NIE) are a type of mediation estimand that allows for such a decomposition. The NIE describes how individuals’ outcomes would be affected by changes in their mediator values induced by contrasting levels of a treatment/exposure, while keeping the treatment fixed. The NDE describes how individuals’ outcomes would be affected by contrasting levels of a treatment/exposure not operating through the mediator (by keeping the mediator fixed to its values under no treatment). To formally define the NIE and NDE, we first define a counterfactual outcome as an outcome variable had treatment been set to some value a and the mediator been set to some value, m, possibly contrary to fact, denoted YA=a,M=m (we provide a more detailed description of notation in Section 2). Similarly, we define a counterfactual mediator as a mediator variable had treatment been set to some value a, possibly contrary to fact, denoted MA=a. The NDE and NIE are then defined in terms of nested counterfactual outcomes, such as YA=a,MA=a (Pearl, 2001). Using this counterfactual notation, the NDE and NIE decompose the average treatment effect (ATE) of a binary treatment/exposure as follows:

EY1,M1-Y0,M0=EY1,M1-Y1,M0naturalindirecteffect(throughM)+EY1,M0-Y0,M0naturaldirecteffect(notthroughM).

Generally, the NDE/NIE are not point-identifiable in the presence of treatment-induced confounders of the mediator-outcome relationship (Avin et al., 2005). Such confounders are variables that are affected by treatment/exposure and in turn affect both the mediator(s) and the outcome. In Figure 1, Z represents such a variable. To give intuition for the lack of identification of the NDE/NIE in the presence of treatment-induced confounders, we focus on the counterfactual outcome YA=a,MA=a. This counterfactual outcome involves two counterfactual worlds simultaneously: one in which A=a for the first portion of the counterfactual outcome, and one in which A=a for the nested portion of the counterfactual outcome. Setting A=a induces a counterfactual treatment-induced confounder, denoted ZA=a. Setting A=a induces another counterfactual treatment-induced confounder, denoted ZA=a. The two treatment-induced counterfactual confounders, ZA=a and ZA=a share unmeasured common causes, UZ, which creates a spurious association. Because ZA=a is also causally related to YA=a,M=m, and ZA=a is also casually related to MA=a, the path through UZ means that the backdoor criterion is not met for identification of YA=a,MA=a, i.e., MaYA=a,mW, where W denotes baseline covariates. We depict this in Figure A1 in the Supporting Information.

Figure 1:

Figure 1:

Directed acyclic graph of the structural causal model considered.

This general lack of identification in the presence of a treatment-induced confounder presents a problem for applied mediational research, because such variables are near-ubiquitous. For example, in trials where an individual or community is randomized to a particular treatment but cannot be forced to take the treatment, treatment take-up/ adherence represents a treatment-induced confounder. In observational data, there may be multiple intermediate variables linking treatment/exposure to the outcome, but only a subset are of interest to examine as mediators (for example, in the case discussed above where one is interested in examining which mediators contribute to the unexpected overall effect). The remaining intermediate variables would be considered as treatment-induced confouders.

However, particular cases exist when the NDE and NIE are identifiable even in the presence of one or more treatment-induced confounders (Miles, 2022; Tchetgen Tchetgen and VanderWeele, 2014). One of these cases is when one can assume monotonicity between a treatment and one or more binary treatment-induced confounders (Tchetgen Tchetgen and VanderWeele, 2014). Assuming monotonicity is also sometimes referred to as assuming “no defiers”—in other words, assuming that there are no individuals who would do the opposite of the encouragement (Angrist et al., 1996).

Monotonicity may seem like a restrictive assumption, but, in fact, it may be reasonable in some common scenarios. For example, consider the relatively common trial setting where the intervention is randomized treatment assignment and the treatment-induced confounder is whether or not treatment was actually taken. In this setting, we may feel comfortable assuming that there are no “defiers”—no individuals who would take the treatment if they were assigned to the control group but would not take the treatment if they were assigned to the treatment group. Indeed, monotonicity in this context is frequently assumed when using instrumental variables (IV) to identify causal effects (Angrist et al., 1996). Using IV as an identification strategy is common in the context of randomized trials where randomized assignment is the IV and treatment uptake or adherence is the exposure of interest. However, we consider so-called intent-to-treat effects here, where randomized assignment is our treatment of interest (not an IV), and treatment uptake/adherence is the treatment-induced confounder.

We consider this same context in our motivating example, the Moving to Opportunity study (MTO) (Sanbonmatsu et al., 2011), where families were randomized to receive a Section 8 housing voucher that would allow them to move out of public housing and into a rental on the private market. In MTO, randomized assignment to receive the voucher is frequently used as an IV for moving to a lower poverty neighborhood. Assuming monotonicity in MTO would mean that we assume that families fall into one of following groups: 1) would not move to a lower-poverty neighborhood regardless of whether they received a housing voucher or not (“never takers”), 2) would move to a lower-poverty neighborhood regardless of whether they received a housing voucher or not (“always takers”), or 3) would move to a lower-poverty neighborhood if they received a housing voucher but would not move if they did not receive a voucher (“compliers”). In this motivating example, we use randomized assignment as the treatment of interest and moving to a lower-poverty neighborhood as the treatment-induced confounder.

Although monotonicity may be a reasonable assumption in trial settings, we know of no non-parametric estimators of natural (in)direct effects in this context (i.e., in the presence of treatment-induced confounders, assuming monotonicity). Consequently, we develop efficiency theory for the NDE/NIE under monotonicity, and use it to propose a nonparametric, multiply robust estimator based on the efficient influence function (EIF). Of relevance for real-world analyses, our estimator allows for continuous and/or multiple mediators and is cross-fitted. This estimator should be applicable whenever natural direct or indirect effects of treatment assignment are of interest in the presence of potentially imperfect compliance acting as a single treatment-induced confounder.

This paper is organized as follows. We give notation, the structural causal model, and review the definition and identification of the NDE and NIE in Section 2. We propose a nonparametric, robust estimator using flexible, data-adaptive regression methods in Section 3. Section 4 details a limited simulation study evaluating the finite sample performance of the proposed estimator. In Section 5, we apply the estimator to our motivating example estimating the NDE and NIE of being randomly assigned to receive a Section 8 housing voucher—the most common form of federal housing assistance—on risk developing any mood or externalizing disorder among adolescent boys—-possibly operating through various school and community characteristics, using longitudinal data from the Moving to Opportunity Study (MTO) (Sanbonmatsu et al., 2011). Section 6 concludes.

2. Notation, structural causal model, estimand definition

Let O=(W,A,Z,M,Y) denote the observed data, and let O1,,On denote a sample of n i.i.d. observations of O. Let W denote a vector of observed baseline covariates, W=fUW, where UW is unobserved exogenous error on W and where function f is assumed deterministic but unknown (Pearl, 2009). Let A denote a binary treatment/exposure variable, A=fW,UA. Let Z denote single, binary treatment-induced confounder of the mediator-outcome relationship, e.g., initiation of treatment/ adherence to treatment assignment, Z=fW,A,UZ. Let M denote a set of mediating variables, M=fW,A,Z,UM, that may be multiple, multi-valued, and/or continuous, and that are considered as a “bundle” (as their joint distribution) in what follows. Finally, let Y denote a continuous or binary outcome, Y=fW,A,Z,M,UY. As in Tchetgen Tchetgen and VanderWeele (2014), we assume mutually independent error terms, U=UW,UA,UZ,UM,UY. We depict this structural causal model graphically in Figure 1.

We use P to denote the distribution of O. We let P be an element of the nonparametric statistical model defined as all continuous densities on O with respect to some dominating measure ν.·..Let p denote the corresponding probability density function. We let E denote the corresponding expectation operator, and define Pf=f(o)dP(o) for a given function f(o). We use g(a|w) to denote the probability mass function of A=a conditional on W = w, and e (a|m,z,w) to denote the probability mass function of A=a conditional on (M,Z,W)=(m,z,w). We use μ(m,z,a,w) to denote the outcome regression function E(Y|M=m,Z=z,A=a,W=w). We use q(z|a,w) and r(z|a,m,w) to denote the corresponding conditional densities of Z.

We define counterfactual variables in terms of interventions on the nonparametric structural equation model (NPSEM). For simplicity of notation, we will use the random variable with its corresponding intervention in the index to denote the counterfactual. For example, YA=1 denotes the random variable fYW,1,Z,M,UY. In what follows, we drop the random variable from the index of the counterfactual whenever the variable that is being intervened on is clear from context. For example, we simply use Ya,m to denote YA=a,M=m, and Ma to denote MA=a.

We are interested in the NDE and NIE, which we defined in the Introduction, and which decompose the ATE. However, we note that our proposed estimator can apply to an alternative decomposition of the total effect into the pure indirect effect EY0,M1-Y0,M0 and the total direct effect (EY1,M1-Y0,M1 (Robins and Greenland, 1992; VanderWeele, 2013).

Tchetgen Tchetgen and VanderWeele (2014) identified the NDE and NIE from observed data O in the presence of treatment-induced confounders, Z. For a single, binary Z, the parameter EYa,Ma, is identified as

θa,a=E(Ya,m,1,w)dPma,1,wPZ=1a,wdP(w)+E(Ya,m,1,w)dPma,0,wP(Z=1a,w)-PZ=1a,wdP(w)+EYa,m,0,wdPma,0,wPZ=0a,wdPw,

under consistency and the following assumptions:

A1 (Monotonicity). If a<a, then ZaZa for all i;

A2 (Positivity of treatment/exposure, treatment-induced confounder, and mediator mechanisms).

Assume:

  • p(w)>0 implies p(aw)>0 and paw>0;

  • p(w)>0 and PZ=za,w>0 for z{0,1} implies P(Z=1a,w)>0

  • p(w)>0 and P(Z=za,w)>0 for z{0,1} implies PZ=0a,w>0

  • p(w)>0 and PZ=1a,w>0 and pma,1,w>0 implies p(ma,1,w)>0

  • p(w)>0 and P(Z=0a,w)>0 and pma,0,w>0 implies p(ma,0,w)>0

  • p(w)>0 and P(Z=1a,w)>0 or PZ=0a,w>0 and pma,0,w>0 implies p(ma,1,w)>0. (We note that the positivity assumptions were not given in the original Tchetgen Tchetgen and VanderWeele (2014).) We discuss where each positivity assumption stems from and interpret each in Section A1 of the Supporting Information.

As in Tchetgen Tchetgen and VanderWeele (2014), we assume mutually independent error terms, U=UW,UA,UZ,UM,UY. We note that sequential exchangeability and cross-world exchangeability are implied by the mutually independent errors in the NPSEM (Tchetgen Tchetgen and VanderWeele, 2014).

3. Estimation

Our goal is to construct a multiply robust, nonparametric, n-consistent estimator of the parameter EYa,Ma.

To construct such an estimator, θ(P) needs to be smooth in the sense that the following expression, known as path-wise differentiability, holds:

ϵθPϵϵ=0=EP{s(O)D(O)} (1)

for a parametric submodel Pϵ that covers the non-parametric model such that Pϵ=0=P, where s(O) is the score of the model Pϵ at ϵ=0. The function D(O) is known as the efficient influence function (EIF), which we discuss next. The EIF will typically depend on P, and can be denoted D(O,P). Importantly, path-wise differentiable parameters will be estimable at n-consistency rates (Bickel et al., 1997).

The EIF is a central object for the construction of frequentist estimators in a non-parametric model, because 1) it provides an expression for the first-order bias of a plug-in estimator, and 2) it characterizes the efficiency bound in the non-parametric model. To see this, let P and Q denote any two distributions in the non-parametric model. Then it can be proved that θ(Q)-θ(P)=-EP[D(O,Q)]+R2(P,Q), where R2(P,Q) is a second-order term that will depend on differences between P and Q, the specific form of this term is given in Lemma 1. Evaluating the above expression at Q=P^ shows that the first-order bias of a plug-in estimator θ(P^) is equal to -EP{D(O,P^)}, where the outer expectation can be estimated using an empirical mean. Therefore, a bias-corrected one-step estimator can be constructed as θ(P^)+1ni=1nD(Oi,P^), for a preliminary estimator P^. The sample variance of the EIF is the efficiency lower bound of any regular estimator of θ(P) in the nonparametric model (Van der Vaart, 2000).

It remains to derive the specific form of the EIF for our problem. There are multiple ways to do this; we refer the reader to van der Laan and Rose (2011) and Kennedy (2022) for examples on the algebraic steps to derive the EIF. We followed an approach proposed by van der Laan and Rose (2011), where we first assume that all the data are discrete, and use the Delta method to find the influence function of the non-parametric MLE of θ. This influence function is then proved to be the EIF for general data structures by confirming that it satisfies expression (1).

The EIF for this parameter is given by the following expressions. It uses a reparameterization similar to that used in Lemma 1 of Díaz et al. (2021) p(ma*,w)p(ma,z,w) was reparameterized as g(aw)g(a*w)q(za,w)r(za,m,w)p(a*m,w)p(am,w) to avoid modeling the conditional joint distribution of mediators. Define

HY,1,1=1{Z=1,A=a}PaWPaM,1,WP(aM,1,W)HY,1,0=1{Z=1,A=a}PaWPZ=0a,WPaM,1,WP(aM,1,W)PZ=0M,a,WPZ=1M,a,WP(Z=1a,w)-PZ=1a,wHY,0,0=1{Z=0,A=a}P(aW)PaM,0,WP(aM,0,W)P(Z=0a,W)PZ=0a,WHM,1,1=1A=a,Z=1PaWHM,1,0=1A=a,Z=0PZ=0a,WPaWP(Z=1a,w)-PZ=1a,wHM,0,0=1{A=a,Z=0}PZ=0a,WPaWP(Z=0a,w)HZ,1,1=1A=aPaWE(Ya,m,1,W)dPma,1,WHZ,1,0=1{A=a}P(aW)-1A=aPaWE(Ya,m,1,W)dPma,0,WHZ,0,0=-1{A=a}P(aW)E(Ya,m,0,W)dPma,0,WHW,1,1=E(Ya,m,1,w)dPma,1,wPZ=1a,wHW,1,0=E(Ya,m,1,W)dPma,0,WP(Z=1a,W)-PZ=1a,WHW,0,0=E(Ya,m,0,W)dPma,0,WP(Z=0a,W)

Then, for z,z{(1,1),(1,0),(0,0)}, we can define the EIF of θz,za,a is equal to Dz,z(O;η)=Dz,z(O;η)-θz,za,a, where

Dz,z(O)=HY,z,z{Y-E(YA,M,Z,W)}+HZ,z,z{Z-P(Z=1A,W)}+HM,z,zE(Ya,M,z,W)-E(Y,a,m,z,W)dPma,z,W+HW,z,z.

The EIF of θa,a is then D(O)=D(O)-θa,a, where D(O)=D1,1(O)+D1,0(O)+D0,0(O).

We propose a one-step estimator for the statistical parameter θa,a that is the sample average of the estimated uncentered EIF. Let ρ denote E(Ya,m,z,w)dPma,z,w. Let η=(g,q,e,r,μ,ρ), where ρ is defined above and where the remaining parameters are defined as in Section 2. Let ηˆ denote an estimator of η. Let D(O,P)=D(O,η), as η contains all the relevant features of P. Let D(O,η^) denote an estimator of D(O,η). The one-step estimator we propose is the sample average of D(O,η^).

We use a cross-fitted version of this estimator. Cross-fitting is a data-splitting technique that weakens some of the technical assumptions (i.e., Donsker-type assumptions) required for asymptotic normality (Klaassen, 1987; Zheng and van der Laan, 2011; Chernozhukov et al., 2018). We perform crossfitting for estimation of all the components of η as follows. Let 𝒱1,,𝒱J denote a random partition of data with indices i{1,,n} into J prediction sets of approximately the same size such that j=1J𝒱j={1,,n}. For each j, the training sample is given by 𝒯j={1,,n}\𝒱j.η^j denotes the estimator of η, obtained by training the corresponding prediction algorithm using only data in the sample 𝒯j, and j(i) denotes the index of the validation set which contains observation i. We then use these fits, η^j(i)(Oi) in computing each efficient influence function, i.e., we compute D(Oi,η^j(i)).

Thus, our one-step estimator of θa,a can be calculated in the following steps:

  1. Let the components of η be defined as above. With the exception of ρ, each can be estimated by cross-fitting a regression of the dependent variable on the independent variables and generating predicted probabilities (if the dependent variable is binary) or predicted values otherwise, setting the values of independent variables where indicated. For example, g^aw can be estimated by cross-fitting a logistic regression model of A on W and generating predicted probabilities that A=a for all observed w. One could also use machine learning in model fitting, which is what we do in the simulations and data analysis.

  2. To estimate ρ,we use the predicted values from μ(a,m,z,W) as a pseudo-outcome in a new regression on variables A,Z,W, and then generate predicted values setting A=a and Z=z.

  3. The estimator of θa,a is θ^a,a=1ni=1nD(Oi,η^j(i)).

  4. The variance can be estimated as the sample variance of D(O,η^).

An R package to implement this estimator is included (see Supporting Information).

3.1. Estimator properties

Let η~ denote the probability limit of the estimator η^ in L2(P) norm. The above estimator is multiply robust such that θa,a is expected to be estimated consistently under one of the following:

(μ~,ρ~,q~)=(μ,ρ,q),or 1)
(μ~,ρ~,g~)=(μ,ρ,g), 2)
(μ~,g~,q~)=(μ,g,q),or 3)
(g~,e~,q~,r~)=(g,e,q,r). 4)

The monotonicity assumption upon which this identification rests may be most tenable in encouragement-design randomized trials like the MTO study used in the illustrative application. In that case, estimation of g may not be necessary, so one would expect consistent estimation if either (μ,q) or (μ,ρ) or (e,q,r) were consistently estimated. These robustness conditions are a consequence of the following Lemma, the proof of which is given in the Supporting Information.

Lemma 1.

Let Dz,z(O;η~) denote the EIF evaluated at a given value η~ of the nuisance parameters. Then we have θz,za,a(η~)-θz,za,a(η)=-EDz,z(O;η~)+Rz,z(η,η~), where Rz,z(η,η~) is a second order term given by:

Rz,z(η,η~)=E{H~Y,z,z-HY,z,z}{E(YA,M,Z,W)-E~(YA,M,Z,W)}+EH~M,z,z-HM,z,zE~(Ya,M,z,W){dPma,z,W-dP~ma,z,W}+E{H~Z,z,z-HZ,z,z}{P(Z=1A,W)-P~(Z=1A,W)}+Rz,zη,η~,

where, for each z,z,Rz,z(η,η~) is equal to:

R1,1(η,η~)={E(Ya,m,1,w)-E~(Ya,m,1,w)}PZ=1a,w-P~Z=1a,wdP~ma,1,wdP(w)+E(Ya,m,1,w)PZ=1a,w-P~Z=1a,wdPma,1,w-dP~ma,1,wdP(w)R1,0(η,η~)={E(Ya,m,1,w)-E~(Ya,m,1,w)}qZ(w)-q~Z(w)dP~ma,0,wdP(w)+E(Ya,m,1,w)qZ(w)-q~Z(w)dPma,0,w-dP~ma,0,wdP(w)R0,0(η,η~)={E(Ya,m,0,w)-E~(Ya,m,0,w)}PZ=1a,w-P~Z=0a,wdP~ma,0,wdP(w)+E(Ya,m,0,w)PZ=0a,w-P~Z=0a,wdPma,0,w-dP~ma,0,wdP(w),

for qZ(w)=P(Z=1a,w)-PZ=1a,w.

Inspection of the remainder terms reveals that the one-step estimator should be consistent in the configurations outlined above, as in such cases all the remainder terms become null. In addition, this lemma will allow us to prove an asymptotic normality result stating that the estimator converges in distribution to a normal random variable with variance equal to the non-parametric efficiency bound. Importantly, this result holds even when the nuisance parameters η are estimated using using flexible regression techniques such as those available in the machine learning literature, as long as the regressions are cross-fitted as detailed above. The only requirement is that the second order terms of the previous lemma converge to zero in probability at rate n-1/2. This can occur, for example, if all the nuisance regressions are estimated consistently at rate n-1/4. This rate is much slower than the convergence rate of parametric models and is achievable under a flexible regression framework.

Theorem 1

(Asymptotic normality of estimators). Let R=R1,1+R1,0+R0,0. Assume R(η,η^)=oPn-1/2. Let H^Y,z,z,H^M,z,z,H^Z,z,z denote the estimates of HY,z,z,HM,z,z,HZ,z,z constructed by plugging in estimates η^. Assume that there exists a constant c such that P(H^Y,z,z<c)=P(H^M,z,z<c)=P(H^Z,z,z<c)=1. Then we have

θ^a,a-θa,a=1ni=1nDOi,η+oP(n-12).

The proof of this result for estimators that allow a second order expansion as in Lemma 1 follows standard empirical process theory and is given, for example, in Zheng and van der Laan (2011).

This result together with the central limit theorem shows that n1/2{θ^a,a-θa,a} converges to a normal distribution with variance Var{D(O,η)} equal to the non-parametric efficiency bound. Furthermore, this result is useful to compute standard errors and Wald-type confidence intervals for contrasts of the parameter θa,a for varying values of a,a, by simply applying the Delta method to the corresponding contrast.

3.2. Contrast of proposed estimator with existing estimators

Tchetgen Tchetgen and VanderWeele (2014) proposed two estimators for the NDE and NIE under monotonicity. First, they proposed a nonparametric maximum likelihood estimator, but one that can only be used if the data, including covariates, are categorical and low-dimensional. They state that such a scenario is unlikely in real-world applications (Tchetgen Tchetgen and VanderWeele, 2014), so we focus our discussion on their second, more practical, parametric estimator. Their parametric estimator relies on correctly specifying: 1) a parametric linear model for the outcome as a function of covariates (W), binary treatment (A), binary treatment-induced confounder (Z), and mediator (M); 2) a parametric linear model for the mediator as a function of , A, and Z; and 3) a parametric logistic model for Z as a function of W and A. Then, the NDE/NIE are estimated using the expressions given in Tchetgen Tchetgen and VanderWeele (2014) that are generalizations of the “product method” of Baron and Kenny (1986), and the variance is estimated by the delta method or bootstrap.

In contrast, our proposed estimator is developed in a nonparametric model, and consequently, makes fewer assumptions on the statistical model as compared to the estimator of Tchetgen Tchetgen and VanderWeele (2014), which assumes the analyst can correctly pre-specify a parametric functional form for each of the models. Our estimator is efficient in the nonparametric model under slower-than-parametric consistency rates in estimation of the nuisance parameters. This allows the use of flexible, data-adaptive regression methods to fit these nuisance quantities. Our estimator is also fully cross-fitted (discussed above). Although the estimator we propose fits more models of nuisance parameters as compared to that of Tchetgen Tchetgen and VanderWeele (2014)—six as compared to three—it only relies on consistent estimation of 3–4 of these parameters, depending on the robustness condition. In the randomized encouragement design scenario we consider, the estimator we propose relies on consistent estimation of 2–3 nuisance parameter models. Lastly, we note that the “product method” of Tchetgen Tchetgen and VanderWeele (2014) does not allow for multiple mediating variables, whereas a set of mediating variables, considered together as a “bundle” can be incorporated in our estimator.

4. Simulation

We conducted a limited simulation study to: 1) verify that our programmed estimator had the theoretical properties of: a) consistency under the robustness conditions established in Lemma 1, b) efficiency, and c) confidence interval coverage; 2) illustrate the estimator’s finite sample performance; and 3) contrast its finite sample performance with a parametric g-computation estimator, similar to (but slightly more flexible than) that of Tchetgen Tchetgen and VanderWeele (2014). The parametric g-computation estimator we use in the simulations is more flexible than the one proposed by Tchetgen Tchetgen and VanderWeele (2014), because it does not use the “product method”; instead, it allows for multiple mediating variables incorporated as a ‘bundle’ by using sequential regression. Namely, we use a cross-fitted plug-in estimator constructed as the empirical mean of HW,1,1+HW,1,0+HW,0,0, where parametric regressions are fitted for the regression functions. This is a generalization of the estimator of Tchetgen Tchetgen and VanderWeele (2014); when the models are specified as in Tchetgen Tchetgen and VanderWeele (2014), this corresponds exactly to their estimator. We use 500 bootstrapped replicates to estimate standard errors. We conducted 1,000 simulations for sample sizes n{1,000,10,000}.

We considered three data-generating mechanisms: 1) where all variables are binary; 2) where W,A,Z are binary and generated as above, and M and Y are continuous; and 3) where W,A,Z, are binary and generated as above; Y is binary and M is multivariate. Each is detailed in Table 1.

Table 1:

Data-generating mechanisms considered in the simulation.

All variables are binary W, A, Z are binary and M, Y are continuous W, A, Z, Y are binary and M is multivariate

PW1=1=0.6
PW2=1=0.3
PW3=1=0.4
P(A=1)=0.5
P(Z=1A,W)=expit-log(1.3)×W1+W2+W3/3+2A-1
P(M=1Z,W)=expit-log(1.1)W3+2Z-0.9 E(MZ,W)=1.1-log(1.1)W3+2Z+ϵ1 PM1=1||Z,W=expit-log(1.1)W3+2Z-0.9
PM2=1Z,W=expitlog(2)W3-Z/2-0.2
P(Y=1M,Z,W)=expit-log(1.3)×W1+W2+W3/3+Z+M E(YM,Z,W)=5-log(1.3)×W1+W2+W3/3+Z+M+ϵ2 P(Y=1M,Z,W)=expit-log(1.3)×W1+W2+W3/3+Z+M1-0.3M2

These DGMs were constructed to align with the illustrative application in which: A is randomly assigned and adheres to the exclusion restriction assumption by only having an effect on M and Y through Z (Angrist et al., 1996), and Z is monotonic with respect to A. We depict this structural causal model graphically in Figure 2.

Figure 2:

Figure 2:

Directed acyclic graph of the structural causal model considered in the simulation and illustrative application.

For the one-step estimator we propose, we considered the following model specification scenarios: 1) correct specification of nuisance parameters in η; 2) correct specification of (g, e, q, r) but misspecification of (μ,ρ); 3) correct specification of (μ,ρ,g) but misspecification of (e,q,r); and 4) correct specification of (μ,q,g) but misspecification of (e,r,ρ). We would still expect consistent estimation of θa,a under each of these misspecifications given the robustness results that are a consequence of Lemma 1. Also, note that A is exogenous in this simulation, so although we estimate g in this simulation, its estimation is not necessary, and it is assumed correctly specified with a simple mean model.

The parametric g-computation estimator relies on estimation of (g,q,ρ,μ), though because A is exogenous, it is correctly specified by a simple mean model. For this estimator, we considered the following model specification scenarios: 1) correct specification of all parameters; 2) correct specification of (g,q,μ) but misspecification of ρ; 3) correct specification of (g,q,ρ) but misspecification of μ; 4) correct specification of (g,ρ,μ) but misspecification of q.

We fit parameters with the highly adaptive lasso (Hejazi et al., 2020; Benkeser and van der Laan, 2016; van der Laan, 2017) for the correctly specified scenarios, and fit parameters using an intercept-only model for the incorrectly specified scenarios (and for the correctly specified g model).

We considered estimator performance in terms of absolute bias, absolute bias scaled by n, influence curve-based standard error relative to the Monte Carlo-based standard error, standard deviation of the estimator relative to the efficiency bound scaled by n, mean squared error relative to the efficiency bound scaled by n, and 95% confidence interval (CI) coverage.

Table 2 shows simulation results for the first data-generating mechanism where all variables are binary. This simple case best allows us to verify that our proposed estimator is implemented properly. We see that the one-step estimator has very little bias and its 95% confidence interval coverage is close to 95% for both the NDE and NIE under correct specification of all nuisance parameters. This estimator also has very little bias under the misspecifications we consider here, as expected given the robustness conditions. However, aspects of performance related to efficiency and coverage deteriorate slightly in some of the misspecification scenarios for the NIE. For example, when q,e, and r are misspecified the 95% CI coverage for the NIE includes the true value less than 80% of the time. Comparing our proposed one-step estimator to the parametric g-computation estimator, we see that the parametric g-computation estimator does not exhibit robustness to model misspecification. We also see that it is more efficient, which is expected given the additional assumptions in the parametric model. Performance under the third data-generating mechanism with multi-variate M is similar, and results are given in Table A1 in the Supporting Information.

Table 2:

Simulation results for the first data-generating mechanism where all variables are binary. Estimation of the natural direct effect, θ(1,0)-θ(0,0)=0.099, nonparametric efficiency bound = 0.997, and the natural indirect effect θ(1,1)-θ(1,0)=0.034, nonparametric efficiency bound = 0.153.

Correctly specified N Bias relse relsd relmse 95% cov. N×Bias

One-step Estimator

NIE
All correct 1000 0.00 0.97 0.93 0.92 0.95 0.02
All correct 10000 0.00 0.94 0.97 1.06 0.93 0.03
g, q, e, r 1000 0.00 0.91 0.83 0.89 0.91 0.11
g, q, e, r 10000 0.00 0.87 0.87 1.20 0.86 0.18
μ, ρ, g 1000 0.00 0.59 0.47 0.62 0.74 0.02
μ, ρ, g 10000 0.00 0.61 0.49 0.65 0.78 0.04
g, q, μ 1000 0.01 1.08 1.03 1.55 0.90 0.31
g, q, μ 10000 0.00 1.07 1.05 1.92 0.86 0.38

NDE
All correct 1000 0.00 0.97 0.99 1.04 0.94 0.04
All correct 10000 0.00 1.00 0.99 0.99 0.95 0.05
g, q, e, r 1000 0.00 1.05 1.01 0.92 0.96 0.07
g, q, e, r 10000 0.00 1.01 1.01 1.05 0.95 0.18
μ, ρ, g 1000 0.00 0.92 0.90 0.95 0.94 0.03
μ, ρ, g 10000 0.00 0.99 0.90 0.82 0.95 0.05
g, q, μ 1000 0.01 1.13 1.01 0.90 0.96 0.31
g, q, μ 10000 0.00 1.05 1.01 1.08 0.94 0.39

Parametric g-computation estimator

NIE
All correct 1000 0.00 0.98 0.58 0.34 0.94 0.01
All correct 10000 0.00 0.96 0.57 0.35 0.93 0.02
μ, q 1000 0.03 1.02 0.00 7.54 0.00 1.08
μ, q 10000 0.03 1.03 0.00 75.43 0.00 3.40
ρ, q 1000 0.03 1.00 0.00 7.54 0.00 1.08
ρ, q 10000 0.03 0.98 0.00 75.43 0.00 3.40

NDE
All correct 1000 0.00 1.03 0.52 0.26 0.96 0.01
All correct 10000 0.00 1.03 0.52 0.26 0.97 0.00
μ, q 1000 0.01 0.92 0.50 0.36 0.90 0.25
μ, q 10000 0.01 0.95 0.50 0.92 0.62 0.80
ρ, q 1000 0.10 0.97 0.00 9.83 0.00 3.13
ρ, q 10000 0.10 0.97 0.00 98.29 0.00 9.89

Abbreviations: relse = influence curve-based standard error relative to the Monte-Carlo-based standard error, relsd = standard deviation of the estimator relative to the efficiency bound scaled by N, relrmse = mean squared error relative to the efficiency bound scaled by N, 95% CI cov = 95% confidence interval coverage.

Table 3 shows simulation results for the second data-generating mechanism where M and Y are continuous. We see that the one-step estimator has good coverage, particularly under correct specification of all models. We also see, as expected, that it is consistent under correct specification and each of the misspecifications considered here that fall under the robustness conditions. However, performance is more sensitive to the sample size in this case, particularly for measures related to the efficiency of the estimator. In contrast, performance of the parametric g-computation estimator is similar in this scenario as in the first all-binary data-generating mechanism, and relatively insensitive to sample size.

Table 3:

Simulation results for the second data-generating mechanism where the mediator and outcome are continuous. Estimation of the natural direct effect, θ(1,0)-θ(0,0)=0.46, nonparametric efficiency bound = 32.98, and the natural indirect effect θ(1,1)-θ(1,0)=0.921, nonparametric efficiency bound = 37.05.

Correctly specified N Bias relse relsd relmse 95% cov. N×Bias

One-step estimator

NIE
All correct 1000 0.03 0.36 1.26 12.44 0.94 0.87
All correct 10000 0.01 0.78 1.38 3.15 0.96 0.72
g, q, e, r 1000 0.58 0.06 5.01 7745.91 0.57 18.32
g, q, e, r 10000 0.07 0.76 1.65 6.20 0.67 6.95
μ, ρ, g 1000 0.00 0.75 0.46 0.38 0.86 0.05
μ, ρ, g 10000 0.00 0.73 0.47 0.42 0.85 0.41
g, q, μ 1000 0.03 0.86 0.56 0.46 0.88 1.07
g, q, μ 10000 0.01 0.87 0.57 0.48 0.89 1.34

NDE
All correct 1000 0.02 0.40 1.31 10.91 0.94 0.59
All correct 10000 0.00 0.82 1.40 2.96 0.95 0.24
g, q, e, r 1000 0.58 0.06 5.05 6833.14 0.79 18.19
g, q, e, r 10000 0.07 0.82 1.78 6.08 0.74 7.25
μ, ρ, g 1000 0.00 0.86 0.68 0.62 0.93 0.05
μ, ρ, g 10000 0.00 0.90 0.67 0.56 0.93 0.41
g, q, μ 1000 0.03 1.09 0.82 0.58 0.97 0.98
g, q, μ 10000 0.01 1.09 0.82 0.59 0.96 1.03

Parametric g-computation estimator

NIE
All correct 1000 0.00 1.00 0.47 0.23 0.95 0.03
All correct 10000 0.00 1.02 0.47 0.21 0.95 0.01
μ, q 1000 0.92 1.03 0.00 25.71 0.00 29.12
μ, q 10000 0.92 0.85 0.00 257.13 0.00 92.09
ρ, q 1000 0.92 1.03 0.00 25.71 0.00 29.12
ρ, q 10000 0.92 1.01 0.00 257.13 0.00 92.09

NDE
All correct 1000 0.00 1.02 0.45 0.20 0.95 0.04
All correct 10000 0.00 0.97 0.45 0.22 0.94 0.15
μ, q 1000 0.00 1.01 0.46 0.21 0.95 0.12
μ, q 10000 0.00 0.96 0.45 0.22 0.94 0.02
ρ, q 1000 0.46 1.03 0.00 5.73 0.00 14.56
ρ, q 10000 0.46 0.93 0.00 57.27 0.00 46.05

Abbreviations: relse = influence curve-based standard error relative to the Monte-Carlo-based standard error, relsd = standard deviation of the estimator relative to the efficiency bound scaled by N, relrmse = mean squared error relative to the efficiency bound scaled by N, 95% CI cov = 95% confidence interval coverage.

5. Empirical Illustration

Next, we apply our proposed one-step estimator to estimate natural direct and indirect effects among adolescent boys whose families participated in the Moving to Opportunity Study (MTO) (Sanbonmatsu et al., 2011). We also estimate the total direct effect and pure indirect effect for comparison, and also apply the parametric g-computation estimator of Tchetgen Tchetgen and VanderWeele (2014) for comparison.

Briefly, MTO randomized families living in public housing in five cities in the United States (US), who volunteered to participate, to either receive a Section 8 housing voucher or not. (Families were actually randomized to one of three groups—two of which received a housing voucher, but we collapse these for simplicity, as others have done (Rudolph et al., 2021, 2018; Osypuk et al., 2012).) Section 8 housing vouchers are the primary form of federal housing assistance, and they serve to subsidize rents on the private rental market for low-income households, thus making it more financially viable for a family to move out of public housing. In this example, we are interested in decomposing the average effect of being randomized to receive a Section 8 housing voucher (A) in early childhood on 10–15-year risk of developing any psychiatric disorder by adolescence (Y) (defined using the Diagnostic and Statistical Manual, Version IV (DSM-IV) criteria, as measured by the CIDI-SF (Kessler et al., 1998; Kessler and Üstün, 2004)) into the portion that operates through changes in the school environment and residential and school instability, considered together as a bundle, (M)—the indirect effect—and the portion that does not—the direct effect.

MTO is an example of an encouragement intervention, because receipt of a Section 8 housing voucher (the randomized intervention) encourages families to move out of public housing and into a rental on the private market by subsidizing their rent. In the context of an encouragement intervention study design, the monotonicity assumption is reasonable. We do not expect there to be any individuals who would not move to a lower poverty neighborhood within the first year if they had received a voucher but who would move if they had not received a voucher. This is also referred to as assuming that there are no “defiers”, individuals who would do the opposite of the encouragement. We assume this illustrative application adheres to the structural causal model depicted in Figure 2.

In this example, we only consider boys whose families participated in MTO. There have been marked differences in MTO’s health effects between girls and boys in which being in the intervention group resulted in higher rates of post-traumatic stress disorder, mood disorder, any psychiatric disorder, smoking, and problematic drug use for boys, but not for girls (Clampet-Lundquist et al., 2011; Sanbonmatsu et al., 2011; Schmidt et al., 2017; Kessler et al., 2014; Rudolph et al., 2018; Kling et al., 2007). We also exclude the Baltimore site, because of evidence that the intervention was meaningfully different in this city versus the others (Rudolph et al., 2018), likely due to concurrent housing related-interventions in that city. These restrictions resulted in a rounded sample size of N=2,100. We adjust for numerous baseline covariates, W, which we detail in the Supporting Information. We provide additional detail of mediator variables, M, considered together as a bundle (jointly distributed without assumptions about independence/dependence), in the Supporting Information as well. For simplicity, we used one imputed dataset, which was imputed using multiple imputation by chained equations (Buuren and Groothuis-Oudshoorn, 2010). The mediator of school rank had the most missingness at 12%, other mediators had 8–9% missingness or no missingness. The outcome of any DSM-IV disorder had 8% missingness. The randomized voucher assignment variable and moving variable had no missingness. Two baseline covariates had 2% missing data (race/ethnicity and baseline neighborhood poverty), and the rest had no missing data.

We estimated the natural direct and indirect effects of being randomized to the Section 8 voucher group on risk of having a psychiatric disorder disorder in adolescence, 10–15 years later, mediated through features of the neighborhood and school environments (in the case of the indirect effect), and not (in the case of the direct effect), using the cross-fitted one-step estimator proposed here, with 2 folds. We also estimated the pure indirect effect and the total direct effect, which are the alternative direct and indirect effects that decompose the ATE. In the Supporting Information, we interpret the components of the positivity assumption in terms of this MTO analysis. We used the SuperLearner ensemble method of combining machine learning algorithms in fitting the nuisance parameters, implemented with the sl3 package (Coyle et al., 2021); this approach weights the algorithms to minimize the 10-fold cross-validated prediction error (Van der Laan et al., 2007). We included the following algorithms: intercept-only regression, generalized linear regression, lasso (Tibshirani, 1996), and gradient boosted machines (Chen and Guestrin, 2016). For comparison, we also estimated the NDE, NIE, total direct effect, and pure indirect effect using the parametric g-computation estimator of Tchetgen Tchetgen and VanderWeele (2014). Columbia University determined this analysis of deidentified data to be non-human subjects research.

Figure 3 shows the point estimates and 95% confidence intervals for each effect. While the natural direct effect appears null, the natural indirect effect shows evidence of an unintended harmful path from being randomized to the housing voucher group on later risk of developing a psychiatric disorder operating through voucher-induced changes in the school and neighborhood environments and the instability of those environments, among boys. This indirect effect was estimated to contribute to a 3.69 percentage point increased risk of developing such a disorder (95% CI: 0.91, 6.47 percentage points), which is in-line with previous estimates using population interventional indirect effects as the estimand (Rudolph et al., 2021). The estimates from the parametric g-computation estimator are qualitatively the same but point estimates are attenuated and confidence intervals are narrower.

Figure 3:

Figure 3:

Estimates of the natural indirect effect, natural direct effect, pure indirect effect, and total direct effect of being randomized to the Section 8 voucher group on risk of having a psychiatric disorder disorder in adolescence, 10–15 years later, mediated through features of the neighborhood and school environments, among boys. Estimates and 95% CIs. All results were approved for release by the U.S. Census Bureau, authorization number CBDRB-FY23-CES018-001.

Estimates of the alternative decomposition components, the pure indirect effect and total direct effect, are the same as their natural counterparts for the parametric g-computation estimator. The estimates differ for the one-step estimator, though they remain in the same direction. We see that the pure indirect effect and total direct effects estimated with the one-step estimator have markedly wider confidence intervals than the natural indirect and direct effects estimates. This is a result of estimating the parameter θ(0,1). The Y component of the EIF of θ(0,1) is estimated using only those families who moved to neighborhoods with ≥ 5% lower poverty levels but who did not receive the Section 8 housing voucher. This was a relatively rare occurrence—comprising less than 10% of the sample. Consequently, the alternative decomposition of the ATE into the pure indirect effect and total direct effect is not as well matched to the reality of MTO as the NDE and NIE decomposition. This example highlights not only differences in interpretation of the two decompositions—with the NIE and NDE most closely reflecting the MTO data—but also highlights how the different decompositions can have implications on efficiency.

6. Conclusion

Although natural direct and indirect effects are not generally identifiable in the presence of treatment-induced confounders, assuming a monotonic relationship between the treatment and treatment-induced confounder achieves identifiability, as shown by Tchetgen Tchetgen and VanderWeele (2014). However, we are not aware of epidemiologists or other applied researchers using this identification strategy in practice. One reason could be that such scenarios are believed to be rare, niche circumstances. However, the monotonicity assumption is likely satisfied in any trial or intervention where the intervention assigns individuals to a treatment condition, but it is up to the individual whether or not to comply with the treatment they were assigned. Indeed, such study designs are common, and have been the setting for much mediation-related research (Vo et al., 2020). In this encouragement design setting, intervention take-up or adherence becomes the treatment-induced confounder, and one can then make the assumption that treatment assignment is monotonic with treatment take-up. If the remaining treatment-induced variables are treated as mediators, as we do in the applied example, then one can use our proposed one-step estimator.

Strengths of our proposed estimator include that it is cross-fitted, multiply robust and efficient, incorporates data-adaptive machine learning into model fitting, and is available as an R package. Limitations of our proposed estimator include that relevant pairs of nuisance parameters may not be congenial with each other—g(a|w) may not be congenial with eam,z,w) and qza,w) may not be congenial with rza,m,w). This lack of congeniality would be an issue in the scenario where μ and ρ are misspecified, so achieving consistency through double robustness would mean that g(a|w),e(a|m,z,w),q(z|a,w) and r(z|a,m,w) all converge at rate n-1/4. While congeniality is not technically required for Theorem 1 (instead, consistency is required), it is plausible that lack of congeniality may affect the estimator’s finite sample performance (Dukes et al., 2019; Vansteelandt et al., 2011). How to enforce congeniality in semiparametric estimation approaches remains an open question and important area for future work.

Supplementary Material

appendix

Acknowledgements:

This work was supported through a Patient-Centered Outcomes Research Institute (PCORI) Project Program Award (ME-2021C2-23636-IC) and R00DA042127. This research was conducted as a part of the U.S. Census Bureau’s Evidence Building Project Series. The Census Bureau has reviewed this data product to ensure appropriate access, use, and disclosure avoidance protection of the confidential source data used to produce this product (Data Management System (DMS) number: P-7504667, Disclosure Review Board (DRB) approval number: CBDRB-FY23-CES018-001).

Footnotes

Supporting Information

Web Appendices, Tables, and Figures referenced in Sections 1, 2, 3.1, 4, and 5, and code to implement the proposed estimator and replicate the simulation results, are available with this paper at the Biometrics website on Wiley Online Library. Code is also available at https://github.com/kararudolph/NDENIE-undermonotonicity.

Data Availability Statement:

Researchers may apply for data access through the U.S. Census Bureau https://www.census.gov/newsroom/press-releases/2023/standard-application-process.html. The data that support the findings in this paper are available in the Supporting Information section of this paper.

References

  1. Angrist J, Imbens G, and Rubin D (1996). Identification of causal effects using instrumental variables. Journal of the American Statistical Association 91, 444–555. [Google Scholar]
  2. Avin C, Shpitser I, and Pearl J (2005). Identifiability of path-specific effects. In IJCAI International Joint Conference on Artificial Intelligence, pages 357–363. [Google Scholar]
  3. Baron RM and Kenny DA (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of personality and social psychology 51, 1173. [DOI] [PubMed] [Google Scholar]
  4. Benkeser D and van der Laan M (2016). The highly adaptive lasso estimator. In 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), pages 689–696. IEEE. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Bickel PJ, Klaassen CA, Ritov Y, and Wellner JA (1997). Efficient and Adaptive Estimation for Semiparametric Models. Springer-Verlag. [Google Scholar]
  6. Buuren S. v. and Groothuis-Oudshoorn K (2010). mice: Multivariate imputation by chained equations in r. Journal of statistical software pages 1–68. [Google Scholar]
  7. Chen T and Guestrin C (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ‘16, pages 785–794, New York, NY, USA. ACM. [Google Scholar]
  8. Chernozhukov V, Chetverikov D, Demirer M, Duflo E, Hansen C, Newey W, and Robins J (2018). Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning. The Econometrics Journal 21,. [Google Scholar]
  9. Clampet-Lundquist S, Edin K, Kling JR, and Duncan G (2011). Moving teenagers out of high-risk neighborhoods: How girls fare better than boys. American Journal of Sociology 116, 1154–1189. [DOI] [PubMed] [Google Scholar]
  10. Coyle JR, Hejazi NS, Malenica I, Phillips RV, and Sofrygin O (2021). sl3: Modern pipelines for machine learning and Super Learning. https://github.com/tlverse/sl3. R package version 1.4.2. [Google Scholar]
  11. D´ıaz I., Hejazi NS., Rudolph KE., and van Der Laan MJ. (2021). Nonparametric efficient causal mediation with intermediate confounders. Biometrika 108, 627–641. [Google Scholar]
  12. Dukes O, Martinussen T, Tchetgen Tchetgen EJ, and Vansteelandt S (2019). On doubly robust estimation of the hazard difference. Biometrics 75, 100–109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Hejazi NS, Coyle JR, and van der Laan MJ (2020). hal9001: Scalable highly adaptive lasso regression inr. Journal of Open Source Software 5, 2526. [Google Scholar]
  14. Kennedy EH (2022). Semiparametric doubly robust targeted double machine learning: a review. arXiv preprint arXiv:2203.06469. [Google Scholar]
  15. Kessler RC, Andrews G, Mroczek D, Ustun B, and Wittchen H-U (1998). The world health organization composite international diagnostic interview short-form (cidi-sf). International journal of methods in psychiatric research 7, 171–185. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Kessler RC, Duncan GJ, Gennetian LA, Katz LF, Kling JR, Sampson NA, Sanbonmatsu L, Zaslavsky AM, and Ludwig J (2014). Associations of housing mobility interventions for children in high-poverty neighborhoods with subsequent mental disorders during adolescence. JAMA 311, 937–948. [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  17. Kessler RC and Üstün TB (2004). The world mental health (wmh) survey initiative version of the world health organization (who) composite international diagnostic interview (cidi). International Journal of Methods in Psychiatric Research 13, 93–121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Klaassen CA (1987). Consistent estimation of the influence function of locally asymptotically linear estimators. The Annals of Statistics pages 1548–1562. [Google Scholar]
  19. Kling JR, Liebman JB, and Katz LF (2007). Experimental analysis of neighborhood effects. Econometrica 75, 83–119. [Google Scholar]
  20. Miles CH (2022). On the causal interpretation of randomized interventional indirect effects. arXiv preprint arXiv:2203.00245. [Google Scholar]
  21. Osypuk TL, Schmidt NM, Bates LM, Tchetgen-Tchetgen EJ, Earls FJ, and Glymour MM (2012). Gender and crime victimization modify neighborhood effects on adolescent mental health. Pediatrics 130, 472–481. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Pearl J (2001). Direct and indirect effects. In Proceedings of the seventeenth conference on uncertainty in artificial intelligence, pages 411–420. Morgan Kaufmann. [Google Scholar]
  23. Pearl J (2009). Myth, Confusion, and Science in Causal Analysis. Technical Report R-348, Cognitive Systems Laboratory, Computer Science Department University of California, Los Angeles, Los Angeles, CA. [Google Scholar]
  24. Robins JM and Greenland S (1992). Identifiability and exchangeability for direct and indirect effects. Epidemiology pages 143–155. [DOI] [PubMed] [Google Scholar]
  25. Rudolph KE, Gimbrone C, and Díaz I (2021). Helped into harm: Mediation of a housing voucher intervention on mental health and substance use in boys. Epidemiology 32, 336–346. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Rudolph KE, Sofrygin O, Schmidt NM, Crowder R, Glymour MM, Ahern J, and Osypuk TL (2018). Mediation of neighborhood effects on adolescent substance use by the school and peer environments. Epidemiology 29, 590–598. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Sanbonmatsu L, Ludwig J, Katz LF, Gennetian LA, Duncan GJ, Kessler RC, Adam E, McDade TW, and Lindau ST (2011). Moving to Opportunity for Fair Housing Demonstration Program–Final Impacts Evaluation. US Department of Housing and Urban Development, Office of Policy Development and Research, Washington, DC. [Google Scholar]
  28. Schmidt NM, Glymour MM, and Osypuk TL (2017). Adolescence is a sensitive period for housing mobility to influence risky behaviors: An experimental design. Journal of Adolescent Health 60, 431–437. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Tchetgen Tchetgen EJ and VanderWeele TJ (2014). On identification of natural direct effects when a confounder of the mediator is directly affected by exposure. Epidemiology (Cambridge, Mass.) 25, 282–291. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Tibshirani R (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological) pages 267–288. [Google Scholar]
  31. van der Laan M (2017). A generally efficient targeted minimum loss based estimator based on the highly adaptive lasso. The international journal of biostatistics 13,. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Van der Laan MJ, Polley EC, and Hubbard AE (2007). Super learner. Statistical Applications in Genetics and Molecular Biology 6,. [DOI] [PubMed] [Google Scholar]
  33. van der Laan MJ and Rose S (2011). Targeted Learning: Causal Inference for Observational and Experimental Data. Springer, New York. [Google Scholar]
  34. Van der Vaart AW (2000). Asymptotic statistics, volume 3. Cambridge university press. [Google Scholar]
  35. VanderWeele TJ (2013). A three-way decomposition of a total effect into direct, indirect, and interactive effects. Epidemiology (Cambridge, Mass.) 24, 224–232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Vansteelandt S, Bowden J, Babanezhad M, and Goetghebeur E (2011). On instrumental variables estimation of causal odds ratios. Statistical Science 26, 403–422. [Google Scholar]
  37. Vo T-T, Superchi C, Boutron I, and Vansteelandt S (2020). The conduct and reporting of mediation analysis in recently published randomized controlled trials: results from a methodological systematic review. Journal of clinical epidemiology 117, 78–88. [DOI] [PubMed] [Google Scholar]
  38. Zheng W and van der Laan MJ (2011). Cross-validated targeted minimum-loss-based estimation. In Targeted Learning, pages 459–474. Springer. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

appendix

Data Availability Statement

Researchers may apply for data access through the U.S. Census Bureau https://www.census.gov/newsroom/press-releases/2023/standard-application-process.html. The data that support the findings in this paper are available in the Supporting Information section of this paper.

RESOURCES