Skip to main content
Biometrika logoLink to Biometrika
. 2011 Dec 20;99(1):115–126. doi: 10.1093/biomet/asr059

Directed acyclic graphs with edge-specific bounds

Tyler J Vanderweele 1, Zhiqiang Tan 2
PMCID: PMC3412607  PMID: 23049135

Summary

We give a definition of a bounded edge within the causal directed acyclic graph framework. A bounded edge generalizes the notion of a signed edge and is defined in terms of bounds on a ratio of survivor probabilities. We derive rules concerning the propagation of bounds. Bounds on causal effects in the presence of unmeasured confounding are also derived using bounds related to specific edges on a graph. We illustrate the theory developed by an example concerning estimating the effect of antihistamine treatment on asthma in the presence of unmeasured confounding.

Some key words: Bayesian network, Bound, Causal inference, Confounding, Directed acyclic graph

1. Introduction

Building on Wellman (1990), VanderWeele & Robins (2009, 2010) developed theory for signed causal directed acyclic graphs and derived results that relate signed edges to causal effects, to covariance amongst variables, and to the sign of the bias that results when unmeasured confounding is present. Signed edges amount to statements about ratios of survivor probabilities, that the ratios are bounded either between 0 and 1 for negative edges or between 1 and ∞ for positive edges, but in certain cases these bounds may be too restrictive. If, for example, the bounds for the ratio are of the form (a, b) with a < 1 < b and thus include 1 rather than being bounded above or below by 1, a signed edge cannot be assigned. In other cases, the ratios might be known to lie in ranges of the form (c, 1) or (1, 1/c), where 0 < c < 1. In this paper, we generalize the definitions for a weak monotonic effect and a signed edge given by VanderWeele & Robins (2010) to the case of bounded edges and derive results concerning these bounded edges.

2. Causal directed acyclic graphs and signed edges

A directed graph (Spirtes et al., 1993; Pearl, 1995, 2000; Dawid, 2002) consists of a set of nodes and directed edges amongst nodes. A path is a sequence of distinct nodes connected by edges regardless of arrowhead direction; a directed path is a path which follows the edges in the direction indicated by the graph’s arrows. A directed acyclic graph is a directed graph in which no node has a directed path back to itself. The nodes with directed edges into a node A are said to be the parents of A; the nodes into which there are directed edges from A are said to be its children. We say that node A is an ancestor of node B if there is a directed path from A to B, and B is then called a descendant of A. A node is called a collider for a particular path if both the preceding and subsequent nodes on the path have directed edges going into it. A path between two nodes, A and B, is said to be blocked given some set of nodes C if either there is a variable in C on the path that is not a collider for the path or if there is a collider on the path such that neither the collider itself nor any of its descendants are in C. For disjoint sets of nodes A, B and C, we say that A and B are d-separated given C if every path from any node in A to any node in B is blocked given C. Directed acyclic graphs are sometimes used as statistical models to encode independence relationships amongst variables represented by the nodes on the graph (Lauritzen, 1996). We will use the notation AB | C to denote that A is conditionally independent of B given C. The variables corresponding to the nodes on a graph are said to satisfy the global Markov property for the directed acyclic graph if for any disjoint sets of nodes A, B, C we have that AB | C whenever A and B are d-separated given C.

Directed acyclic graphs can be interpreted as representing causal relationships (Spirtes et al., 1993; Pearl, 1995, 2000; Dawid, 2002). Let Ya denote the counterfactual value of Y under an intervention to set A to a. Pearl (1995) defined a causal directed acyclic graph as a directed acyclic graph with nodes (X1, …, Xn) corresponding to variables such that each variable Xi is given by its nonparametric structural equation Xi = fi (pai, ∊i) where pai are the parents of Xi on the graph and the ∊i are mutually independent. These structural equations generalize the path analysis and linear structural equation models (Pearl, 1995, 2000) developed by Wright (1921) in the genetics literature and Haavelmo (1943) in the econometrics literature. The structural equations encode counterfactual relationships amongst the variables represented on the graph, themselves representing one-step ahead counterfactuals, with other counterfactuals given by recursive substitution. On a causal directed acyclic graph, a node C is said to be a common cause of A and Y if there exists a directed path from C to Y not through A and a directed path from C to A not through Y. The requirement that the ∊i be mutually independent is essentially a requirement that there is no variable absent from the graph which, if included on the graph, would be a parent of two or more variables (Pearl, 1995, 2000). A causal directed acyclic graph defined by nonparametric structural equations satisfies the global Markov property as stated above (cf. Verma & Pearl, 1988; Geiger et al., 1990; Lauritzen et al., 1990; Pearl, 2000). We will say that (V1, …, Vn) constitutes an ordered list if i < j implies that Vi is not a descendant of Vj. On some graphs there will be more than one possible ordering. The results below will apply provided that there is some ordering under which the conditions of the propositions and theorems are satisfied. We will use k to denote (V1, …, Vk), with 0 = ∅. For further discussion of the causal interpretation of directed acyclic graphs see Spirtes et al. (1993), Pearl (1995, 2000), Dawid (2002) and Robins (2003).

The directed acyclic graph causal framework has proved to be particularly useful in determining whether conditioning on a given set of variables, or none at all, is sufficient to control for confounding. The most important result in this regard is the back-door path criterion (Pearl, 1995). A back-door path from some node A to another node Y is a path into Y which begins with a directed edge into A; a front-door path from A to Y is a path into Y which begins with a directed edge emanating from A. Pearl (1995) showed that for intervention variable A and outcome Y, if a set of variables X is such that no variable in X is a descendant of A and such that X blocks all back-door paths from A to Y then YaA | X, so that conditioning on X suffices to control for confounding for the estimation of the causal effect of A on Y. We will use the idea of a back-door path throughout the paper.

Some of the results below only make reference to conditional probability distributions rather than counterfactuals and thus do not require a causal interpretation of directed acyclic graphs. Nevertheless, we believe the results in this paper will be of greatest interest in drawing inferences concerning causal effects.

If υ is a vector, then we will say some function f (υ) is nondecreasing in υ if it is nondecreasing in each component of υ. If A is a parent of Y, we will use paYA to denote the parents of Y other than A. Throughout, we assume that the conditional distributions used are well defined (e.g., Billingsley, 1995, § 33). For example, we assume that pr(Yy | A = a, paYA) as a function of y for fixed a and paYA is a proper cumulative distribution function. By convention, we interpret the fraction y/0 as infinite for any y > 0.

VanderWeele & Robins (2009, 2010) gave the following definitions for a weak monotonic effect and a signed edge on a causal directed acyclic graph (cf. Wellman, 1990).

Definition 1. We say that A has a weak positive monotonic effect on Y if the survivor function S(y | a, paYA) = pr(Y > y | A = a, paYA) is such that whenever a1a0 we have S(y | a1, paYA) ⩾ S(y | a0, paYA) for all y and all paYA, and a weak negative monotonic effect if whenever a1a0 we have S(y | a1, paYA) ⩽ S(y | a0, paYA) for all y and all paYA.

Corresponding to the notion of a weak monotonic effect is that of a signed edge.

Definition 2. An edge on a causal directed acyclic graph from A to Y is said to be of positive or negative sign if, respectively, A has a weak positive or negative monotonic effect on Y; otherwise, it is without sign. The sign of a path is the product of the signs of the edges that constitute that path; the path is without sign if one of the path’s edges is without sign.

The definition of a weak positive monotonic effect requires that for all a1a0 such that at least one of S(y | a1, paYA) or S(y | a0, paYA) is nonzero, 1 ⩽ S(y | a1, paYA) / S(y | a0, paYA) ⩽ ∞; a weak negative monotonic effect requires that for all a1a0 such that at least one of S(y | a1, paYA) or S(y | a0, paYA) is nonzero, 0 ⩽ S(y | a1, paYA) / S(y | a0, paYA) ⩽ 1 whenever a1a0. The restriction that such inequalities only need to be satisfied when either the numerator or the denominator is nonzero will be assumed to apply to all ratios throughout and this condition will not be repeatedly stated.

3. Bounded edges and the propagation of bounds

We now give a definition for a bounded edge that generalizes the notion of a signed edge.

Definition 3. For some node Y with parent A we will say that the AY edge is stochastically bounded by, Λ+) if

Λpr(Y>y|A=a1,paYA)pr(Y>y|A=a0,paYA)Λ+,

for all a1 > a0, y and paYA. If the AY edge is stochastically bounded by, Λ+), then we write Θ(A, Y) = (Λ, Λ+) and place the ordered set, Λ+) on the AY edge of the directed acyclic graph. If there is no edge from A to Y on the graph then we define Θ(A, Y) = (1, 1).

We will refer to a directed acyclic graph with bounds placed on the edges of the graph as a bounded directed acyclic graph. For any edge we must have that 0 ⩽ Λ < 1 and 1 < Λ+ ⩽ ∞; the bounds (0, ∞) can be placed on any edge, equivalent to an edge without bounds. By letting y → −∞ in the ratio in Definition 3, we see that the interval (Λ, Λ+) must contain 1. If Θ(A, Y) = (1, b), then A has a weak positive monotonic effect on Y and the A → Y edge will be of positive sign; if Θ(A, Y) = (a, 1), then the AY edge will be of negative sign. If Y is binary, then Θ(A, Y) = (1, Λ+) simply requires that

1pr(Y=1|A=a1,paYA)pr(Y=1|A=a0,paYA)Λ+,

for all a1 > a0, paYA. Definitions 1–3 all apply to any statistical graphical model that satisfies the global Markov property for a directed acyclic graph.

To develop theory on the propagation of bounds on a directed acyclic graph, we rely on the following proposition. Its proof and those of other results are given in the Appendix.

Proposition 1. For fixed a0, a1 and q, assume that the following conditions hold.

  1. For i = 1, …, n,
    Λipr(Vi>υi|V¯i1=υ¯i1,A=a1,Q=q)pr(Vi>υi|V¯i1=υ¯i1,A=a0,Q=q)Λi+,
    for all υi and ῡi−1;
  2. pr(Vi > υi | V̄i−1 = i−1, A = a1, Q = q) and pr(Vi > υi | i−1 = i−1, A = a0, Q = q) are nondecreasing in ῡi−1 (i = 2, …, n); and

  3. G(ῡn) is nondecreasing in ῡn.

    Then
    i=1nΛipr{G(V¯n)>g|A=a1,Q=q}pr{G(V¯n)>g|A=a0,Q=q}i=1nΛi+, (1)
    for all g.

Proposition 1 does not require reference to a directed acyclic graph; the conclusion holds for any ordered sequence V1, …, Vn that satisfies (a) and (b). The next theorem essentially says that if we can find some set X that blocks all back-door paths from a node A to a node Y, and if certain directed paths into Y are signed, then bounds for the effect of A on Y can be derived from the bounds on the edges emanating from A.

Theorem 1. Suppose that A is an ancestor of Y and that some set X of nondescendants of A blocks all back-door paths from A to Y. Let (V1, …, Vn) be an ordered list of all nodes on directed paths from A to Y and let Vn+1 = Y. We may denote the bounds on the edges from A to Vi by Θ(A,Vi)=(Λi,Λi+) for i = 1, …, n + 1. If for all i such that there is an edge from A to Vi all directed paths from Vi to Y are of positive sign, then for all G(n+1) nondecreasing in ῡn+1,

i=1n+1Λipr{G(V¯n+1)>g|A=a1,X=x}pr{G(V¯n+1)>g|A=a0,X=x}i=1n+1Λi+,

for a1 > a0 and all g.

Theorem 1 allows for settings in which there is some node Vj on a directed path from A to Y such that there is no edge from A to Vj; in this case Λj=Λj+=1. If for some i, all directed paths from Vi to Y are of negative, rather than positive, sign and G(n+1) is nonincreasing, rather than nondecreasing, in υi then, the result could still be applied by replacing Vi with −Vi so that all directed paths from −Vi to Y are of positive sign. The bounds Λi, Λi+ will then have to be specified so that Θ(A,Vi)=(Λi,Λi+). If Θ(A, Vi) is of form (a, 1), then Θ(A, −Vi) is of form (1, b), but b is not in general 1/a.

VanderWeele & Robins (2009) showed that if X is a set of nondescendants of A that blocked all back-door paths from A to Y and if all directed paths from A to Y are of positive sign then, pr(Y > y | a, x) is nondecreasing in a. If in Theorem 1 Θ(A,Vi)=(Λi,Λi+) is of the form (1, ∞) for all i, then, equivalently, under the assumptions of Theorem 1, all directed paths from A to Y are positive. This special case of Theorem 1 is still a generalization of the aforementioned result of VanderWeele & Robins (2009). The reason is that under these assumptions, Theorem 1 would allow one to conclude that pr{G(n+1) > g | a, x} was nondecreasing in a for any choice of the function G(n+1) nondecreasing in n+1, rather than simply for the special choice of G(n+1) as G(n+1) = y.

If G(n+1) is taken as G(n+1) = y, then by Pearl’s back-door path adjustment theorem (Pearl, 1995), it follows immediately from Theorem 1 that

i=1n+1Λipr(Ya1>y|X=x)pr(Ya0>y|X=x)i=1n+1Λi+.

We illustrate the use of Theorem 1 in the following example.

Example 1. Consider the bounded directed acyclic graph given in Fig. 1. All directed paths from V1 to Y are of positive sign and all directed paths from V2 to Y are of positive sign. The path consisting of the edge V1Y is of positive sign since Θ(V1, Y) = (1, 4); the path V1V3Y is of positive sign since the edge V1V3 is of negative sign and the edge V3Y is of negative sign and thus the product of the signs of these edges is positive. Finally, the path consisting of the edge V2Y is also of positive sign. There are edges emanating from A into V1, V2 and Y with bounds Θ(A, V1) = (1, 3), Θ(A, V2) = (1, 2), Θ(A, Y) = (2/3, 2), and by Theorem 1, we have for all c and all a1 > a0 that 2/3 = (1)(1)(2/3) ⩽ pr(Y > y | A = a1, C = c)/pr(Y > y | A = a0, C = c) ⩽ (3)(2)(2) = 12 since C blocks all back-door paths from A to Y; similarly, 2/3 ⩽ pr(Y > y | A = a1, X = x)/pr(Y > y | A = a0, X = x) ⩽ 12 since X also blocks all back-door paths from A to Y.

Fig. 1.

Fig. 1

Example of the propagation of bounds.

Others have tried to generalize the notion of a signed edge in order to account for additional information or numeric bounds. Parsons (1995) provides a set of possible axiomatic rules to govern the propagation of influences on networks that can be strongly or weakly positive or negative, rather than simply positive or negative. Parsons thus extended Wellman’s qualitative influence to the notion of categorical influence. Renooij & van der Gaag (2008) have recently further developed Parsons’ approach, but neither provide the generality of our notion of a bounded edge. A more related approach is that of Liu & Wellman (1998, 2004) who provide bounds for a cumulative distribution function, not by bounding the ratios of survivor probabilities but by postulating that certain cumulative distribution functions stochastically dominate those in fact governing various signed edges on a graph. Their results also differ from ours in another way: while in Theorem 1 bounds of the form pr(Y > y | A = a1, X = x)/pr(Y > y | A = a0, X = x) are obtained by information on bounded edges emanating from A, their results require bounds on the cumulative distribution functions corresponding to edges pointing into Y. Additional research might consider if further inferences concerning bounds would be possible if their results were combined with ours.

4. Bounds for causal effects in the presence of unmeasured confounding

The following result allows us to give bounds for a causal effect in the presence of unmeasured confounding.

Theorem 2. Suppose that for some variable A and some nonnegative outcome Y, the set X = CU of nondescendants of A blocks all back-door paths from A to Y, where C consists of measured covariates and U unmeasured covariates. Let Sa = ∑c E(Y | a, c)pr(c). If for some (u and some ΛL, ΛH > 0,

ΛLE(Y|a,u,c)E(Y|a,u,c)ΛHE(Y|a,u,c), (2)

for all a, c and u, then

ΛLΛHSaE(Ya)ΛHΛLSa.

For (2) to hold for u = u, the interval (ΛL, ΛH) must contain 1.

Corollary 1. Under the assumptions of Theorem 2 we have

ΛLΛHSa1ΛHΛLSa0E(Ya1)E(Ya0)ΛHΛLSa1ΛLΛHSa0 (3)

and

(ΛLΛH)2Sa1Sa0E(Ya1)E(Ya0)(ΛHΛL)2Sa1Sa0.

In Theorem 2, if ΛL and ΛH are known a priori from subject matter knowledge then Sa = ∑c E(Y | a, c)pr(c) can be estimated from data and thus, by using Corollary 1, one can obtain bounds for the causal effect, E(Ya1) − E(Ya0) on the additive scale, or E(Ya1) / E(Ya0) on the multiplicative scale. If ΛL and ΛH are unknown, Corollary 1 could still be used in sensitivity analysis (Cornfield et al., 1959) by varying ΛL and ΛH. The assumptions required for the use of these results in sensitivity analysis are much weaker than those required for other techniques (e.g., Lin et al., 1998). Bounded edges in conjunction with Theorem 1 can also be used to yield bounds for inequality (2) by means of the following proposition.

Proposition 2. For nonnegative Y, if for all y, pr(Y > y | u, l) ⩽ Λpr(Y > y | u, l), then E(Y | u, l) ⩽ ΛE(Y | u, l).

If in Theorem 2, U is univariate and (A, C) contain all parents of Y other than U then the bounds for the UY edge, say Θ(U, Y) = (ΛL, ΛH), will, by Proposition 2 with L = (A, C), be bounds that satisfy (2) provided there is some minimum value u of U. However, although the application of Proposition 2 could be used to draw conclusions about the inequality in (2), (2) is in fact weaker than what is required for a bounded edge and Theorem 2 can thus be employed more generally. For example, suppose (A, C) are nondescendants of U and block all back-door paths from U to Y, although U lies on a back-door path from A to Y. Then the bounds relating U and Y from Theorem 1, with U and (A, C) taking the roles of A and X, can be used to give bounds that satisfy (2) by Proposition 2. In the next section, we apply Theorem 2 to a problem concerning the effect of antihistamine treatment on asthma in the presence of unmeasured confounding.

5. Application and further discussion

We present an example adapted from Greenland et al. (1999) and discussed by VanderWeele et al. (2008). Using bounds for edges rather than signed edges allows us to derive bounds for the causal effect under fewer assumptions.

Example 2. Consider a hypothetical study of the relation of antihistamine treatment, denoted by A, and asthma incidence, denoted by Y, among first-grade children attending public schools. Suppose that air pollution level, denoted by W, is independent of sex, denoted by C. Suppose further that sex influences the administration of antihistamine only through its relation to bronchial reactivity, denoted by U, but directly influences asthma risk; suppose also that air pollution leads to asthma attacks only through its influence on antihistamine use and bronchial reactivity; and that there are no important confounders beyond air pollution, bronchial reactivity and sex. The causal relationships amongst these variables are then those given in Fig. 2.

Fig. 2.

Fig. 2

Example concerning bounds for the effect of antihistamine use on asthma in the presence of unmeasured confounding.

Under the assumptions given above, by Pearl’s back-door path criterion conditioning on C and U suffices to control for confounding; conditioning on C, U and W or on U and W also suffices. If data were only available for antihistamine use A, asthma Y and sex C, then we could not produce valid estimates of the causal effect of A on Y because controlling only for C does not suffice to control for confounding. Suppose now that, for the purposes of this study, asthma and bronchial reactivity can be considered binary, comparing high versus low, and that 1 ⩽ pr(Y = 1 | a, U = 1, c)/pr(Y = 1 | a, U = 0, c) ⩽ 2 for all a, c. In other words, high bronchial reactivity increases the likelihood of asthma by a factor somewhere between 1 and 2 for all levels of antihistamine treatment for both males and females. Suppose that in the analysis of the available data it was found that S1 = ∑c E(Y | A = 1, C = c)pr(C = c) = 0.06 and (S0 = ∑c E(Y | A = 0, C = c)pr(C = c) = 0.25. From Corollary 1 we have that

12S12S0E(Ya=1)E(Ya=0)2S112S0,

and thus, −0.47 ⩽ E(Ya=1) − E(Ya=0) ⩽ −0.005. We could then conclude from this study that antihistamine use truly had a beneficial effect on asthma. To draw conclusions about bounds for the causal effect using signed edges, VanderWeele et al. (2008) had to assume that the WU, WA, UA and UY edges all had positive sign. From these assumptions, they concluded that E(Ya=1) − E(Ya=0) ⩽ S1S0 and from this it follows that E(Ya=1) − E(Ya=0) ⩽ 0.06 − 0.25 = −0.19. The application of Theorem 2 required assumptions about bounds related to only one edge, namely the UY edge.

As is clear from Example 2, Theorem 2 in this paper allows the researcher to draw conclusions about bounds on causal effects in the presence of unmeasured confounding by making assumptions concerning fewer edges than were previously required. VanderWeele et al. (2008) showed that to draw conclusions about the sign of the bias in the presence of unmeasured confounding using signed edges the treatment had to be binary or comparison had to be made between the minimum and maximum levels of treatment. Intuition about the sign of the bias could fail if an intermediate level of treatment was considered. In contrast, Theorem 2 makes no assumptions on whether A is binary, ordinal or continuous.

The approach employed in Example 2 to address bounds under unmeasured confounding by use of Theorem 2 is broadly applicable. In cases in which the outcome Y is binary, as in many epidemiologic studies, the approach is straightforward because specifying (ΛL, ΛH) that satisfy (2) consists only of specifying bounds on the risk ratio for the effect of U on Y and the inequalities in (3) of Theorem 2 then give bounds on the causal effect. By specifying bounds on the risk ratio for the effect of U on Y one immediately obtains bounds for the effect of A on Y.

Acknowledgments

The authors thank the editor and two referees for helpful comments. VanderWeele acknowledges support from the National Institutes of Health, U.S.A. Tan acknowledges support from the National Science Foundation, U.S.A.

Appendix

VanderWeele & Robins (2009) proved Lemmas A1–A3 below, which will be used in this Appendix. Lemmas A2 and A3 are given in a somewhat more general form in VanderWeele & Robins (2009) but these special cases will suffice for our purposes here.

Lemma A1. If h(z2, z1, q) is nondecreasing in z1 and in z2 and pr(Z2 > z2 | Z1 = z1, Q = q) is nondecreasing in z1, for all z2, then E{h(Z2, z1, q) | Z1 = z1, Q = q} is nondecreasing in z1.

Lemma A2. Let X denote some set of nondescendants of Z1 that block all back-door paths from Z1 to Z2. If all directed paths between Z1 and Z2 are of positive sign then pr(Z2 > z2 | z1, x) is nondecreasing in z1 for all z2 and all x.

Lemma A3. Suppose that Z1 is a nondescendant of Z2 and let Q denote the union of (i) the ancestors of Z1 and (ii) the ancestors of Z2 which are not descendants of Z1. Let V0 = Z1 and Vn = Z2 and let (V1, …, Vn−1) be an ordered list of all the nodes on directed paths from Z1 to Z2 exclusive of Z1 and Z2 then pr(Vk > υk | z1, k−1, q) = pr(Vk > υk | paυk) for k = 1, …, n.

Proof of Proposition 1. Note that 0Λi< and 0<Λi+(i=1,,n). We give a proof for the inequality involving Λi. The inequality involving Λi+ can be proved similarly by exchanging the roles of a0 and a1. Recall that k denotes (V1, …, Vk).

We first establish the inequality for n = 1. Let G−1(g) = sup{υ1 : G(υ1) ⩽ g}. A useful result is that (i) if G{G−1(g)} ⩽ g, then G(υ1) >g if and only if υ1 > G−1(g) and that (ii) if G{G−1(g)} > g, then G(υ1) >g if and only if υ1G−1(g). We prove (i) by contradiction. Suppose υ1 > G−1(g). If G(υ1) ⩽ g, then υ1G−1(g) by the definition of G−1(g), which is a contradiction. Suppose G(υ1) > g. If υ1G−1(g), then G(υ1) ⩽ G{G−1(g)} ⩽ g since G is nondecreasing in υ1, again a contradiction. To prove (ii), if υ1 > G−1(g), then G(υ1) > g, because if G(υ1) ⩽ g, then υ1G−1(g) by the definition of G1(g), which is a contradiction, and if υ1 = G−1(g) then G(υ1) >g. To show the converse, suppose G(υ1) >g. If υ1 < G−1(g), then there exists u1 such that υ1 < u1 < G−1(g) and G(u1) ⩽ g by the definition of G−1(g) and this contradicts that G is nondecreasing. Due to (i) and (ii), we have that {υ1: G(υ1) > g} equals either {υ1 : υ1 >G−1(g)} or {υ1 : υ1G−1(g)}.

Consider the case {υ1 : G(υ1) >g} = {υ1 : υ1 >G−1(g)}. Then pr{G(V1) >g | A = a1, Q = q} = pr{ V1 >G−1(g) | A = a1, Q = q} and pr{G (V1) > g | A = a0, Q = q} = pr{ V1 >G−1(g) | A = a0, Q = q}. The desired inequality is trivial by condition (a). Next, consider the case {υ1 : G(υ1) > g} = {υ1 : υ1G−1(g)}. Then pr{G(V1) > (g | A = a1, Q = q} = pr{V1G−1(g) | A = a1, Q = q} and pr{G(V1) > g | A = a0, Q = q} = pr{V1G− 1(g) A = a0, Q = q}. Note that p1/p0Λ1 if and only if p1Λ1p0, for 0Λi< and 0 ⩽ p0, p1 ⩽ 1. Let υ1j be a sequence increasing to G−1(g). Then pr(V1 > υ1j | A = a1, Q = q) ⩾ Λ1 pr(V1 > υ1j | A = a0, Q = q) by condition (a). Let j → ∞, we obtain pr{V1G−1(g) | A = a1, Q = q} ⩾ Λ1 pr{V1G−1(g) | A = a0, Q = q} and thus pr{G(V1) > g | A = a1, Q = q} / pr{G(V1) > g | A = a0, Q = q} ⩾ Λ1.

Suppose that inequality (1) involving Λi holds for n = k. We show that it then holds for n = k + 1. Let Gυ¯k1(g)=sup{υk+1:G(υk+1,υ¯k)g}. Similarly as in the case for n = 1, we have that {υk+1 : G(υk+1, k) >g} equals either {υk+1:υk+1>Gυ¯k1(g)} or {υk+1:υk+1Gυ¯k1(g)} and, by condition (a), pr{G(Vk+1, ῡk) > g | k = k, A = a1, Q = q} ⩾ Λk+1 pr{G(Vk+1, k) > g | k = k, A = a0, Q = q}. Therefore,

pr{G(V¯k+1)>g|A=a1,Q=q}=E[pr{G(Vk+1,V¯k)>g|V¯k,A=a1,Q=q}|A=a1,Q=q]Λk+1E[pr{G(Vk+1,V¯k)>g|V¯k,A=a0,Q=q}|A=a1,Q=q]=Λk+101pr{Ψg(V¯k)>z|A=a1,Q=q}dz,

where Ψg(k) = pr{G(Vk+1, k) > g | k = k, A = a0, Q = q}. Since G(υk+1, k) is nondecreasing in υk+1 and k and since pr(Vk+1 > υk+1 | k = k, A = a0, Q = q) is nondecreasing in k, we have by Lemma A1 that Ψg(k) is nondecreasing in k. Since (1) holds for n = k we have that

pr{G(V¯k+1)>g|A=a1,Q=q}=Λk+101pr{Ψg(V¯k)>z|A=a1,Q=q}dzΛk+101(i=1kΛi)pr{Ψg(V¯k)>z|A=a0,Q=q}dz=(i=1k+1Λi)pr{G(V¯k+1)>g|A=a0,Q=q},

and so (1) holds for n = k + 1.

Proof of Theorem 1. Fix a1 > a0. Since Θ(A, Vi) = ( Λi, Λi+) we have

Λipr(Vi>υi|A=a1,paYA)pr(Vi>υi|A=a0,paYA)Λi+,

for all υi and paYA, where paYA denote the parents of Vi other than A. Let Q denote the union of (i) the ancestors of A and (ii) the ancestors of Y which are not descendants of A then by Lemma A3, pr(Vi > υi | a, paYA) = pr(Vi > υi | i−1, a, q) and thus

Λipr(Vi>υi|V¯i1=υ¯i1,A=a1,Q=q)pr(Vi>υi|V¯i1=υ¯i1,A=a0,Q=q)Λi+(i=1,,n+1),

for all υi, i−1, q. Under the assumption that for all i such that there is an edge from A to Vi all directed paths from Vi to Y are of positive sign, relevant nodes can be replaced by their negations so that for all i such that there is an edge from A to Vi all edges on all directed paths from Vi to Y are of positive sign. By Lemma A2, pr(Vi > υi | i−1 = i−1, A = a, Q = q) is nondecreasing in i−1 for all υi, a, q since (A, Q) will block all back-door paths from i−1 to Vi. By Proposition 1 we thus have that

i=1n+1Λipr(G(υ¯n+1)>g|A=a1,Q=q)pr(G(υ¯n+1)>g|A=a0,Q=q)i=1n+1Λi+, (A1)

for all g and q. Furthermore,

pr{G(υ¯n+1)>g|A=a,X=x)}=E[pr{G(υ¯n+1)>g|A=a,X=x,Q}|A=a,X=x]=E[pr{G(υ¯n+1)>g|A=a,Q}|A=a,X=x]=E[pr{G(υ¯n+1)>g|A=a,W}|A=a,X=x],

where W is the subset of Q which are parents of one of the nodes in n+1. There can be no unblocked front-door paths from A to W given X since the nodes in W are not descendants of A and thus any front-door path from A to W will be blocked given X by a collider. All back-door paths from A to W are blocked given X since X blocks all back-door paths from A to Y. From this it follows that all paths from A to W are blocked given X and so W is conditionally independent of A given X and so we have

E[pr{G(υ¯n+1)>g|A=a,W}|A=a,X=x]=E[pr{G(υ¯n+1)>g|A=a,W}|X=x]=E[pr{G(υ¯n+1)>g|A=a,Q}|X=x].

We have thus shown that

pr{G(υ¯n+1)>g|A=a,X=x}=E[pr{G(υ¯n+1)>g|A=a,Q}|X=x]. (A2)

By (A1) we have that

(i=1n+1Λi)pr{G(υ¯n+1)>g|A=a0,Q=q}pr{G(υ¯n+1)>g|A=a1,Q=q} (A3)

and

pr{G(υ¯n+1)>g|A=a1,Q=q}(i=1nΛi+)pr{G(υ¯n+1)>g|A=a0,Q=q} (A4)

for all g and q. Taking conditional expectations given X = x of (A3) and (A4) and making use of relation (A2) we have that

(i=1n+1Λi)pr{G(υ¯n+1)>g|A=a0,X=x}pr{G(υ¯n+1)>g|A=a1,X=x}

and

pr{G(υ¯n+1)>g|A=a1,X=x}(i=1nΛi+)pr{G(υ¯n+1)>g|A=a0,X=x},

for all g and x. This completes the proof.

Proof of Theorem 2. Let Sa = ∑cE(Y | A = a, C = c)pr(C = c). We then have that

Sa=cE(Y|A=a,C=c)dF(c)=c{cE(Y|A=a,C=c,U=u)dF(u|A=a,C=c}dF(c)c{uΛHE(Y|A=a,C=c,U=u)dF(u|A=a,C=c)}dF(c)=ΛHcE(Y|A=a,C=c,U=u)dF(c)=ΛHc{uE(Y|A=a,C=c,U=u)dF(u|C=c}dF(c)ΛHc{u1ΛLE(Y|A=a,C=c,U=u)dF(u|C=c}dF(c)=ΛHΛLc,uE(Y|A=a,C=c,U=u)dF(u,c)=ΛHΛLE(Ya).

Thus, (ΛLH)SaE(Ya). The proof that E(Ya) ⩽ (ΛHL)Sa is similar.

Proof of Proposition 2. We have that E(Y|u,l)=0pr(Y>y|u,l)dy0Λpr(Y>y|u,l)dy=ΛE(Y|u,l).

References

  1. Billingsley P. Probability and Measure. 3rd ed. Wiley; New York: 1995. [Google Scholar]
  2. Cornfield J, Haenszel W, Hammond EC, Lilienfeld AM, Shimkin MB, Wynder LL. Smoking and lung cancer: recent evidence and a discussion of some questions. J Nat Cancer Inst. 1959;22:173–203. [PubMed] [Google Scholar]
  3. Dawid AP. Influence diagrams for causal modelling and inference. Int Statist Rev. 2002;70:161–89. [Google Scholar]
  4. Geiger D, Verma TS, Pearl J. Identifying independence in Bayesian networks. Networks. 1990;20:507–34. [Google Scholar]
  5. Greenland S, Pearl J, Robins JM. Causal diagrams for epidemiologic research. Epidemiology. 1999;10:37–48. [PubMed] [Google Scholar]
  6. Haavelmo T. The statistical implications of a system of simultaneous equations. Econometrica. 1943;11:1–12. [Google Scholar]
  7. Lauritzen S. Graphical Models. Oxford: Oxford University Press; 1996. [Google Scholar]
  8. Lauritzen SL, Dawid AP, Larsen BN, Leimer HG. Independence properties of directed Markov fields. Networks. 1990;20:491–505. [Google Scholar]
  9. Lin DY, Psaty BM, Kronmal RA. Assessing the sensitivity of regression results to unmeasured confounding in observational studies. Biometrics. 1998;54:948–63. [PubMed] [Google Scholar]
  10. Liu C-L, Wellman MP. Proc 14th Conf Uncertainty Artif Intel. Madison: Morgan Kaufmann; 1998. Using qualitative relationships for bounding probability distributions. In; pp. 346–53. [Google Scholar]
  11. Liu C-L, Wellman MP. Bounding probabilistic relationships in Bayesian networks using qualitative influences: methods and applications. Int J Approx Reason. 2004;36:31–73. [Google Scholar]
  12. Parsons S. Proc 11th Conf Uncertainty Artif Intel. San Francisco: Morgan Kaufman; 1995. Refining reasoning in qualitative probabilistic networks; pp. 427–34. [Google Scholar]
  13. Pearl J. Causal diagrams for empirical research. Biometrika. 1995;82:669–88. [Google Scholar]
  14. Pearl J. Causality: Models, Reasoning, and Inference. Cambridge: Cambridge University Press; 2000. [Google Scholar]
  15. Renooij S, van der Gaag LC. Enhanced qualitative probabilistic networks for resolving trade-offs. Artif Intel. 2008;172:1470–94. [Google Scholar]
  16. Robins JM. Semantics of causal DAG models and the identification of direct and indirect effects. In: Green P, Hjort NL, Richardson S, editors. Highly Structured Stochastic Systems. New York: Oxford University Press; 2003. pp. 70–81. [Google Scholar]
  17. Spirtes P, Glymour C, Scheines R. Causation, Prediction and Search. New York: Springer; 1993. [Google Scholar]
  18. VanderWeele TJ, Hernán MA, Robins JM. Causal directed acyclic graphs and the direction of unmeasured confounding bias. Epidemiology. 2008;19:720–8. doi: 10.1097/EDE.0b013e3181810e29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. VanderWeele TJ, Robins JM. Properties of monotonic effects on directed acyclic graphs. J Mach Learn Res. 2009;10:699–718. [Google Scholar]
  20. VanderWeele TJ, Robins JM. Signed directed acyclic graphs for causal inference. J. R. Statist. Soc. B. 2010;72:111–27. doi: 10.1111/j.1467-9868.2009.00728.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Verma T, Pearl J. Causal networks: semantics and expressiveness. In: Shachter R, Levitt TS, Kanal LN, editors. Proc 4th Workshop on Uncertainty in Artif Intel. Vol. 4. Amsterdam: Elsevier; 1988. pp. 352–9.pp. 69–76. Reprinted in Uncertainty in Artificial Intelligence. [Google Scholar]
  22. Wellman MP. Fundamental concepts of qualitative probabilistic networks. Artif Intel. 1990;44:257–303. [Google Scholar]
  23. Wright S. Correlation and causation. J Agric Res. 1921;20:557–85. [Google Scholar]

Articles from Biometrika are provided here courtesy of Oxford University Press

RESOURCES