Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Sep 17.
Published in final edited form as: Stat Biosci. 2021 Mar 21;13(3):524–542. doi: 10.1007/s12561-021-09305-7

Intergenerational Associations Between Maternal Diet and Childhood Adiposity: A Bayesian Regularized Mediation Analysis

Yu-Bo Wang 1, Cuilin Zhang 2, Zhen Chen 2
PMCID: PMC12439109  NIHMSID: NIHMS2109973  PMID: 40964517

Abstract

Growing evidence supports a positive association between childhood obesity and chronic diseases in later life. It is also suggested that childhood obesity is more prevalent for children born from pregnancies complicated by metabolic disorders such as gestational diabetes, and can be related to maternal dietary factors during gestation. Extending conventional analyses that report only the marginal associations within non-causal mediation frameworks, we present mediation analysis in the case of multiple exposures and multiple mediators using a regularized two-stage approach. By placing shrinkage priors on each parameter relating to direct and indirect effects, a parsimonious model can be obtained, and consequently, the most relevant pathways will be selected to inform the development of efficient prevention programs. We apply this method to data from the Danish site of the Diabetes & Women’s Health Study, Danish National Birth Cohort (DNBC), and find 6 significant maternal risk factors either directly or indirectly affecting childhood body mass index z score at age 7. Simulations with data-generating mechanisms similar to the DNBC data demonstrate good performance of the proposed model.

Keywords: Childhood adiposity, Mediation analysis, Potential outcome, Regularization, Lasso

1. Introduction

Over last decades, the escalating burden of childhood obesity has received increasing attention and becomes a significant public health concern. Childhood obesity is also found related to substantially increased risk of adulthood obesity, type 2 diabetes, stroke, and coronary heart disease afterwards, resulting in large medical and societal costs. Hence, identifying risk factors, in particular modifiable ones, that can inform the development of efficient prevention programs is imperative for tackling this global epidemic.

Accumulating evidence supports that childhood obesity is more prevalent among children born from pregnancies complicated by metabolic disorders such as gestational diabetes mellitus (GDM), one of the most common pregnancy complications that affects more than 6% of pregnancies in the US and up to 12% worldwide [1]. Several studies have also found that in this particular group, maternal dietary factors such as greater intakes of refined grains and artificial sweetened beverages are related to increased risk of obesity from birth through childhood [2,3].

Although many studies have suggested and discussed these intergenerational associations between maternal dietary risk factors and childhood obesity, most considered these factors one at a time, without jointly considering them in a single framework. Moreover, few investigated these associations in the presence of mediators. To fill these gaps, we propose an approach where a collection of risk factors is jointly investigated and pathways of their associations to childhood adiposity are delineated in relation to select mediators. The proposed method will be illustrated using a subset of data from the Danish site of the Diabetes & Women’s Health Study [4], Danish National Birth Cohort (DNBC), where risk factors include maternal demographics, diets and lifestyle, mediators consist of birth outcomes (weight, height, and head circumference), and the outcome relates to ageand sex-specific body mass index z score (BMI z score) at age 7. Further detailed descriptions of DNBC have been provided elsewhere [5].

Mediation analysis assesses the effects of study exposures on an outcome either through or around select mediators. The classic Baron–Kenny (B–K) model [6] is one of the widely used and has been greatly extended in recent literature. These extensions include non-linear regressions and non-continuous outcomes in [710], a compositional mediation analysis in [11], a non-parametric Bayesian approach in [12], developments involving multiple mediators in [1317], and a framework for both multiple exposures and multiple mediators in [18]. In particular, [18] considered difference-of-coefficient approach and proposed shrinkage priors (Laplace priors [19]) directly on the effects of interest with the assumption that only few direct or indirect effects of exposures have significant impact on the outcome. This estimation procedure with regularizations facilitates stable parameter estimations in mediation analysis while allows effects selection. However, this method may fail to obtain unbiased estimates of effects, especially when mediators are highly correlated. In this work, we propose a regularized two-stage approach to circumvent this limitation. Specifically, we consider independent Laplace priors directly on regression coefficients in an extended B–K model, where coefficients associated with exposures and mediators have their own penalizing/shrinkage parameters. Under such a setting, the proposed method allows different regularizations on direct and indirect effects and entertains the property of Lasso. The proposed method also enjoys model parsimony and simultaneous estimation and effects selection.

The remainder of the paper is organized as follows. Section 2 introduces the proposed strategy along with the traditional ways of evaluating direct and indirect effects when multiple exposures and mediators are present. In Sect. 3, we apply this strategy to data from DNBC. We further conduct simulations in Sect. 4 to examine the performance of the methodology. Finally, Sect. 5 summarizes and discusses future research directions.

2. The Method

2.1. Premise

Suppose there are p>1 risk factors, q>1 mediators, and r covariates in the study. Let y=y1,y2,,ynT denote a continuous outcome vector of n subjects, and X=x1,x2,,xnT, M=m1,m2,,mnT, and Z=z1,z2,,znT be the corresponding n×p risk factor, n×q mediator, and n×r covariate matrices, where the superscript T denotes the transposition of a vector or a matrix, xiT=xi1,xi2,,xip, miT=mi1,mi2,,miq, and ziT=zi1,zi2,,zir for i=1,2,,n. To define the total, direct, and indirect effects, we first let mkmkxT for k=1,2,,q denote the kth potential mediator of a subject when risk factors are at the level of xT. Accordingly, yxT,mT is the potential outcome of a subject when the levels of risk factors and mediators are set at xT and mT=m1xT,m2xT,,mqxT, respectively. Assuming x~T=xT+(0,,0,1,0,,0)=x1,xj-1,xj+1,xj+1,,xp, m~T=m1x~T,m2x~T,,mqx~T, and m~kT=m1xT,,mk-1xT,mkx~T,mk+1xT,,mqxT, the total effect of the jth risk factor is then defined as

Eyx~T,m~T-EyxT,mT, (1)

which is the expected difference of the outcome with one unit increase in the jth risk factor while other factors are held constant. If yx~T,mT and yxT,m~T are available, (1) can be rewritten as the sum of Eyx~T,m~T-EyxT,m~T and EyxT,m~T-EyxT,mT, the direct and total indirect effects of the jth risk factor, respectively. This decomposition can go further to individual indirect effects when yxT,m~kT for k=1,2,,q are obtainable. Note that for both direct and indirect effects, yx~T,mT, yxT,m~T, and yxT,m~kT are unobservable and known as counterfactual outcomes since only one level of risk factors can be assigned to a subject. To estimate these effects under the potential outcome framework, we require the usual stable unit treatment value assumption (SUTVA) and several additional ones [14, 20, 21]: (A1) no unmeasured confounding for exposure–outcome relation; (A2) no unmeasured confounding for mediator–outcome relation; (A3) no unmeasured confounding for exposure–mediator relation; and (A4) exposure does not confound mediator–outcome relation for any values of the mediators. These four assumptions are required to hold with respect to the whole set of exposures and mediators.

2.2. The Baron–Kenny Model

To study the underlying causal mechanisms, the B–K model [6], a structural–equation modeling approach that implements a one-exposure-at-a-time procedure as shown in Fig. 1a, is widely used in current literature. This model aims at delineating the direct and indirect effects of the jth exposure as follows:

yi=μj+αjxij+miTβj+εij,mi1=μj1+γj1xij+εij1,miq=μjq+γjqxij+εijq, (2)

where μj and μjk, for k=1,2,,q, are intercepts, αj,βj=βj1,βj2,,βjqT, and γjk are regression coefficients, εij~iidN0,σj2 and εij1,εij2,,εijq~iidN0,Σj are independent error terms. For ease of presentation, we have omitted here (and hereinafter) the additive covariate components in each equation of (2), i.e, ziTξj,ziTφj1,,ziTφjq, respectively. In this framework, the collection of αj and ηjβj1γj1,βj2γj2,,βjqγjqT for j=1,2,,p over p separate models represents p direct effects as well as q×p indirect effects through q mediators. Although easy to implement, this procedure ignores the correlations among the exposures in each analysis and only reveals the marginal causal mechanisms.

Fig. 1.

Fig. 1

Mediation analysis frameworks with outcome Y and (a) a single exposure X and multiple mediators M1,M2,,Mq, and (b) multiple exposures X1,X2,,Xp and multiple mediators M1,M2,,Mq, where dashed lines associate with (γ11,γ12,,γ1q), dotted lines with (γ21,γ22,,γ2q), and dashdotted lines with γp1,γp2,,γpq. In both cases, α’s, γ’s, and β’s are regression coefficients, the solid lines represent the direct effects of X’s on Y, and the dashed/dotted/dashdotted lines are effects of X’s on Y transmitted through mediators M’s

2.3. The Regularized Two-Stage (regTS) Approach

In pursuit of an overall picture of the mechanisms when multiple exposures and mediators exist (as shown in Fig. 1b), one can extend (2) to include all exposures as follows:

yi=μ+xiTα+miTβ+εi,mi1=μ1+xiTγ1+εi1,miq=μq+xiTγq+εiq, (3)

where μ and μk, for k=1,2,,q, are intercepts, α=α1,α2,,αpT, β=β1,β2,,βqT, and γk=γ1k,γ2k,,γpkT are regression coefficients, εi~iidN0,σ2 and εi1,εi2,,εiq~iidN(0,Σ) are independent error terms. Under (3), the effects of all pathways in Fig 1b will be measured. Detailed definitions of these effects are provided in Supplemental Materials. However, potential multi-collinearity among exposures and mediators may result in estimation and inference problems. Moreover, it is desirable to allow variable selection and effects estimation simultaneously when multiple risk factors and multiple mediators are present. Consequently, we combine (2.3) with Bayesian least absolute shrinkage and selection operator (Lasso) [19] to achieve model parsimony, a belief that most but few of pathways are ignorable. Specifically, for α and β, we assign p+q conditional Laplace priors, each with its own shrinkage parameter (λ1j or λ2k):

αjσ2~iidLaplace0,σ2λ1j,
βkσ2~iidLaplace0,σ2λ2k,

for j=1,2,,p and k=1,2,,q. The corresponding density functions are

πασ2=j=1pλ1j2σ2exp-j=1pλ1jαjσ2 (4)

and

πβσ2=k=1qλ2k2σ2exp-k=1qλ2kβkσ2, (5)

respectively. For γj1,γj2,,γjq, we propose q Laplace priors with common shrinkage parameter λ1j to link to the jth exposure:

πγj1,γj2,,γjqσ12,σ22,,σq2=k=1qλ1j2σk2exp-λ1jk=1qγjkσk2,j=1,2,,p, (6)

where σk2 is the kth diagonal element of Σ. For μ and μk, vague conjugate priors are assigned with μσ2~N0,σ02σ2 and μkΣ~N0,σ02σk2, where σ02 is a pre-specified hyperparameter. We also assign non-informative priors to σ2 and σk2:

πσ2=1σ2

and

πσk2=1σk2.

To allow the shrinkage parameters λ1j and λ2k to reflect information from the data, we let each follow Gamma θ10,δ10 and Gamma θ20,δ20, respectively, where θ10, θ20, δ10, and δ20 are pre-specified shape and rate parameters. The setting ensures that the proposed method has flexibility in penalizing the two types of effects (direct and indirect). As a result, the proposed framework entertains the property of adaptive Lasso. Overall, this Bayesian regularization approach enjoys three desirable features: (1) it avoids the problem of multi-collinearity, (2) it achieves model parsimony, and (3) it conducts simultaneous variable selection and estimation of effects. We provide all detailed steps of posterior computations in Supplemental Materials.

3. Childhood Adiposity Study

In the efforts of preventing childhood obesity, it is important to identify early life modifiable factors such as maternal dietary intakes and lifestyle characteristics during pregnancy that may intergenerationally affect the cardiometabolic outcomes of offsprings. This information is particularly useful in dietary guideline and nutritional program formulations, especially for women with pregnancy complications. In recent literature, several studies (for example, [2, 3]) have demonstrated that maternal food consumptions such as a higher intake of refined grains or artificially sweetened beverages are significantly related to the elevation of offspring risk of overweight or obesity among children who were born from pregnancies complicated by GDM. To have a further understanding of these intergenerational associations and to strengthen causality inference, we conduct mediation analysis of data of women complicated by GDM and their children born from the index pregnancy from DNBC under the B–K model framework with Bayesian regularizations.

DNBC is a longitudinal cohort that recruited 91,827 women from 1996 to 2002 in Denmark and collected 101,402 cases of pregnancy. Besides sociodemographic characteristics as well as perinatal and medical information, a food-frequency questionnaire (FFQ) for maternal dietary intakes during pregnancy was collected via phone interview, and a follow-up questionnaire about the child’s health status was administered when the child was 7 years old. In this paper, we focus on a subset of women–offspring pairs (n=481), where women had a singleton pregnancy and were diagnosed with GDM, and their children were born from the index pregnancy. We investigate the direct and indirect effects of maternal diets on offspring BMI z score at age 7, a quantity commonly used to measure childhood adiposity. Specifically, we examine 15 dietary factors (g/day): sugar-sweetened beverage, artificially sweetened beverage, red meat, processed meat, whole grains, refined grains, cereal, vegetables, legumes, potato, fruits, oil, egg, diary, and desserts. We also consider total energy intake (kcal/day), maternal age (year), gestational age at delivery (week), maternal pre-pregnancy BMI (kg/m2), maternal physical activity (min/week), parity (0: first live birth; 1: had previous live birth), and smoking (0: no; 1: yes) as potential risk factors. To explore the mechanisms of intergenerational effects, we consider three birth variables of offspring as mediators: neonatal weight (g), height (cm), and head circumference (cm). Detailed distributions of these maternal and offspring characteristics are provided in Table 1.

Table 1.

Characteristics of women during the index pregnancy and offsprings in the subset data of DNBC. SD stands for standard deviation, and Q1, Q2, Q3 for the first, second, and third quantiles, respectively

Variable Mean SD Q1 Q2 Q3
BMI z score 0.286 1.162 − 0.490 0.250 0.950
Maternal age (year) 31.674 4.429 29.000 32.000 35.000
Gestational age at delivery (week) 39.589 1.661 38.571 39.714 40.714
Pre-pregnancy BMI (kg/m2) 27.060 5.697 22.590 26.219 30.444
Maternal physical activity (min/week) 30.177 81.657 0.000 0.000 30.000
Total energy (kcal/day) 10056.681 2663.157 8207.122 9744.579 11517.246
Maternal dietary factors during pregnancy
 Sugar-sweetened beverage (g/day) 457.188 459.314 157.775 326.161 600.098
 Artificially sweetened beverage (g/day) 118.337 273.521 0.000 19.968 122.161
 Red meat (g/day) 71.277 33.352 47.480 66.532 87.569
 Processed meat (g/day) 17.199 13.065 7.641 13.705 23.876
 Whole grains (g/day) 179.421 92.235 111.133 166.473 241.140
 Refined grains (g/day) 100.616 50.152 63.553 88.766 134.673
 Cereal (g/day) 24.726 27.458 1.786 12.857 47.143
  Vegetables (g/day) 131.730 107.264 64.883 102.092 153.948
 Legumes (g/day) 12.600 18.686 3.226 7.860 17.923
 Potato (g/day) 133.096 84.689 72.491 110.425 179.257
 Fruits (g/day) 143.412 107.706 55.662 100.269 236.830
 Oil (g/day) 27.332 20.783 12.846 21.631 35.511
 Egg (g/day) 14.672 12.555 6.802 11.980 19.024
 Diary (g/day) 661.127 437.332 338.762 607.678 824.873
Desserts (g/day) 36.156 30.157 17.410 28.807 48.416
Birth outcomes
 Weight (g) 3747.000 586.728 3390.000 3780.000 4120.000
 Height (cm) 52.630 2.521 51.000 53.000 54.000
 Head circumference (cm) 35.520 1.709 35.000 36.000 37.000
n %
Parity
 First live birth 187 38.877
 Had previous live birth 294 61.123
Smoking
 No 411 85.447
 Yes 70 14.553

Due to right skewed distributions of dietary variables and the existence of zeros, we take log transformation of them after 1 unit location shift. Also, to improve Markov chain Monte Carlo (MCMC) convergence, we rescale all continuous exposures and mediators by their corresponding standard deviations. With the hyper-parameters σ02=100, θl0=1, and δl0=0.1 for l=1,2, we generate an MCMC sample of 15, 000 iterations with a 5000 burn-ins and present the posterior means and 95% highest posterior density (HPD) intervals of direct and indirect effects in Tables 2, 3, 4, and 5, which also include comparisons with the general B–K model in (3). An effect is significant when its 95% HPD interval does not contain zero. We find that regTS produces similar results of direct effects as B–K, where pre-pregnancy BMI, smoking, and refined grains are positively associated with BMI z score at age 7 while total energy is negatively associated (Table 2). There are some differences, however, in the indirect effects between these two models. While there are seven significant indirect effects through birth weight in the B–K model (Table 3), only three remain significant in the regTS approach. These are gestational age at delivery, pre-pregnancy BMI, and parity. Smoking, total energy, consumption of whole grains, and diary are significant in B–K but not in regTS. This reduction in the number of significant indirect effects is both expected and desirable, since the regTS approach simultaneously conducts variable selection and effects estimation. Tables 4 and 5, where estimates of indirect effects through birth height and head circumference are presented, reveal similar conclusions. While there are three significant indirect effects through birth height in the B–K model, only one remains significant in regTS. Neither B–K nor regTS produces a significant indirect effect through head circumference.

Table 2.

Posterior means and 95% HPD interval estimates of direct effects of maternal dietary intakes and other risk factors on offspring BMI z score at age 7 in the presence of birth outcomes as mediators

Exposure B–K regTS
Mean 95% HPD Int. Mean 95% HPD Int.
Lower Upper Lower Upper
Maternal age (year) 0.001 − 0.096 0.113 − 0.012 − 0.106 0.088
Gestational age at delivery (week) − 0.049 − 0.173 0.113 − 0.044 − 0.177 0.026
Pre-pregnancy BMI (kg/m2) 0.314 0.199 0.410 0.301 0.207 0.374
Maternal physical activity (min/week) 0.013 − 0.081 0.113 0.014 − 0.076 0.103
Parity (0: first baby; 1: above one) − 0.194 − 0.420 0.030 − 0.154 − 0.361 0.048
Smoking (0: non-smoker; 1: smoker) 0.522 0.248 0.800 0.461 0.186 0.754
Total energy (kcal/day) − 0.242 − 0.369 − 0.056 − 0.099 − 0.171 − 0.009
Sugar-sweetened beverage (g/day) 0.086 − 0.014 0.191 0.048 − 0.041 0.154
Artificially sweetened beverage (g/day) 0.045 − 0.059 0.152 0.056 − 0.042 0.159
Red meat (g/day) 0.078 − 0.024 0.176 0.045 − 0.038 0.142
Processed meat (g/day) − 0.005 − 0.113 0.104 − 0.011 − 0.111 0.086
Whole grains (g/day) 0.044 − 0.071 0.179 0.001 − 0.111 0.115
Refined grains (g/day) 0.133 0.039 0.247 0.098 0.006 0.199
Cereal (g/day) − 0.036 − 0.143 0.068 − 0.043 − 0.142 0.052
Vegetables (g/day) 0.108 − 0.014 0.236 0.077 − 0.033 0.210
Legumes (g/day) − 0.034 − 0.167 0.085 − 0.023 − 0.140 0.090
Potato (g/day) − 0.004 − 0.098 0.102 − 0.014 − 0.108 0.072
Fruits (g/day) 0.021 − 0.081 0.125 0.011 − 0.082 0.103
Oil (g/day) 0.083 − 0.045 0.209 0.033 − 0.069 0.137
Egg (g/day) 0.010 − 0.093 0.112 0.008 − 0.082 0.113
Diary (g/day) 0.082 − 0.032 0.192 0.043 − 0.056 0.143
Desserts (g/day) 0.035 − 0.082 0.145 0.017 − 0.081 0.111

Table 3.

Posterior means and 95% HPD interval estimates of indirect effects of maternal dietary intakes and other risk factors on offspring BMI z score at age 7 through birth weight

Exposure B–K regTS
Mean 95% HPD Int. Mean 95% HPD Int.
Lower Upper Lower Upper
Maternal age (year) − 0.002 − 0.035 0.038 − 0.003 − 0.040 0.031
Gestational age at delivery (week) 0.163 0.086 0.230 0.160 0.095 0.245
Pre-pregnancy BMI (kg/m2) 0.048 0.015 0.090 0.046 0.000 0.098
Maternal physical activity (min/week) − 0.009 − 0.045 0.025 − 0.010 − 0.044 0.023
Parity (0: first baby; 1: above one) 0.168 0.081 0.262 0.164 0.059 0.275
Smoking (0: non-smoker; 1: smoker) − 0.108 − 0.212 − 0.009 − 0.097 − 0.207 0.005
Total energy (kcal/day) − 0.056 − 0.098 − 0.014 − 0.027 − 0.123 0.017
Sugar-sweetened beverage (g/day) − 0.001 − 0.038 0.034 − 0.006 − 0.039 0.028
Artificially sweetened beverage (g/day) 0.027 − 0.012 0.067 0.024 − 0.013 0.062
Red meat (g/day) − 0.007 − 0.043 0.029 − 0.010 − 0.048 0.024
Processed meat (g/day) − 0.017 − 0.057 0.019 − 0.014 − − 0.052 0.026
Whole grains (g/day) 0.053 0.018 0.094 0.041 − 0.006 0.097
Refined grains (g/day) 0.010 − 0.025 0.042 − 0.004 − 0.038 0.033
Cereal (g/day) 0.011 − 0.025 0.051 0.010 − 0.023 0.046
Vegetables (g/day) 0.008 − 0.039 0.057 − 0.002 − 0.047 0.043
Legumes (g/day) 0.001 − 0.044 0.043 0.004 − 0.034 0.046
Potato (g/day) 0.019 − 0.019 0.055 0.010 − 0.025 0.046
Fruits (g/day) 0.001 − 0.033 0.033 0.002 − 0.034 0.043
Oil (g/day) 0.010 − 0.032 0.058 0.005 − 0.038 0.051
Egg (g/day) 0.002 − 0.037 0.039 0.002 − 0.035 0.039
Diary (g/day) 0.046 0.007 0.085 0.036 − 0.008 0.088
Desserts (g/day) 0.040 − 0.000 0.086 0.030 − 0.011 0.081

Table 4.

Posterior means and 95% HPD interval estimates of indirect effects of maternal dietary intakes and other risk factors on offspring BMI z score at age 7 through birth height

Exposure B–K regTS
Mean 95% HPD Int. Mean 95% HPD Int.
Lower Upper Lower Upper
Maternal age (year) − 0.002 − 0.022 0.016 − 0.002 − 0.020 0.016
Gestational age at delivery (week) − 0.081 − 0.138 − 0.014 − 0.077 − 0.179 − 0.011
Pre-pregnancy BMI (kg/m2) − 0.005 − 0.023 0.011 − 0.002 − 0.019 0.022
Maternal physical activity (min/week) 0.005 − 0.011 0.024 0.005 − 0.012 0.023
Parity (0: first baby; 1: above one) − 0.050 − 0.101 − 0.002 − 0.045 − 0.103 0.007
Smoking (0: non-smoker; 1: smoker) 0.074 0.001 0.143 0.065 − 0.008 0.153
Total energy (kcal/day) 0.006 − 0.007 0.023 − 0.003 − 0.032 0.030
Sugar-sweetened beverage (g/day) 0.007 − 0.011 0.028 0.009 − 0.006 0.030
Artificially sweetened beverage (g/day) − 0.005 − 0.024 0.014 − 0.005 − 0.025 0.014
Red meat (g/day) 0.005 − 0.013 0.023 0.007 − 0.011 0.031
Processed meat (g/day) − 0.001 − 0.019 0.020 − 0.001 − 0.019 0.019
Whole grains (g/day) − 0.020 − 0.044 0.000 − 0.015 − 0.054 0.005
Refined grains (g/day) 0.011 − 0.003 0.027 0.012 − 0.005 0.034
Cereal (g/day) − 0.002 − 0.021 0.015 − 0.002 − 0.022 0.014
Vegetables (g/day) − 0.012 − 0.035 0.007 − 0.005 − 0.032 0.015
Legumes (g/day) 0.007 − 0.012 0.031 0.003 − 0.018 0.025
Potato (g/day) 0.001 − 0.017 0.020 0.002 − 0.019 0.023
Fruits (g/day) 0.003 − 0.012 0.021 0.002 − 0.015 0.021
Oil (g/day) − 0.006 − 0.030 0.015 − 0.003 − 0.028 0.019
Egg (g/day) − 0.003 − 0.022 0.014 − 0.003 − 0.025 0.012
Diary (g/day) − 0.014 − 0.036 0.002 − 0.009 − 0.039 0.008
Desserts (g/day) − 0.016 − 0.038 0.005 − 0.012 − 0.041 0.008

Table 5.

Posterior means and 95% HPD interval estimates of indirect effects of maternal dietary intakes and other risk factors on offspring BMI z score at age 7 through birth head circumference

Exposure B–K regTS
Mean 95% HPD Int. Mean 95% HPD Int.
Lower Upper Lower Upper
Maternal age (year) 0 − 0.007 0.008 0 − 0.005 0.004
Gestational age at delivery (week) 0 − 0.041 0.049 − 0.007 − 0.040 0.020
Pre-pregnancy BMI (kg/m2) 0 − 0.010 0.011 − 0.001 − 0.008 0.005
Maternal physical activity (min/week) 0 − 0.006 0.007 0 − 0.005 0.005
Parity (0: first baby; 1: above one) 0 − 0.028 0.035 − 0.005 − 0.030 0.019
Smoking (0: non-smoker; 1: smoker) 0 − 0.038 0.032 0.004 − 0.017 0.029
Total energy (kcal/day) − 0.001 − 0.014 0.006 0.002 − 0.004 0.016
Sugar-sweetened beverage (g/day) 0 − 0.006 0.007 0 − 0.007 0.004
Artificially sweetened beverage (g/day) 0 − 0.012 0.012 − 0.001 − 0.009 0.006
Red meat (g/day) 0.001 − 0.008 0.009 0 − 0.006 0.007
Processed meat (g/day) 0 − 0.013 0.010 0.001 − 0.005 0.011
Whole grains (g/day) 0 − 0.008 0.006 − 0.001 − 0.010 0.004
Refined grains (g/day) 0.001 − 0.007 0.011 − 0.002 − 0.014 0.004
Cereal (g/day) 0 − 0.011 0.013 − 0.002 − 0.013 0.006
Vegetables (g/day) 0 − 0.012 0.013 0.002 − 0.007 0.014
Legumes (g/day) 0 − 0.011 0.011 − 0.002 − 0.013 0.006
Potato (g/day) 0 − 0.010 0.012 − 0.001 − 0.009 0.005
Fruits (g/day) 0 − 0.006 0.007 0 − 0.005 0.005
Oil (g/day) 0 − 0.012 0.014 0.001 − 0.006 0.008
Egg (g/day) 0 − 0.006 0.007 − 0.001 − 0.006 0.004
Diary (g/day) 0 − 0.008 0.007 0 − 0.004 0.005
Desserts (g/day) 0 − 0.013 0.011 − 0.002 − 0.014 0.007

4. Simulations

To study the performance of the proposed approach, we conduct a simulation study with data generated from a mechanism that mimics the DNBC study. Specifically, we consider 22 exposures and 3 mediators with 2 direct and 10 indirect effects non-zero. The detailed values of these direct and indirect effects are provided in the second column of Tables 6, 7, 8, and 9. To generate the mediators and outcome, we use the same exposure matrix as well as the estimated variance terms σˆ2 and Σ^ from the DNBC data.

Table 6.

Simulation results of direct effects (DE) based on the B–K and regTS models in the presence of 3 mediators and 22 exposures. #Sig denotes the number of times the corresponding effect is significant (zero falling outside 95% HPD interval) among the 200 replicates

DE True B–K regTS
Bias SD rMSE #Sig Bias SD rMSE #Sig
α1 0 − 0.009 0.056 0.056 13 − 0.007 0.048 0.049 7
α2 0 − 0.005 0.071 0.071 38 0.002 0.060 0.060 31
α3 0 0.007 0.055 0.056 15 0.007 0.049 0.049 9
α4 0 0 0.060 0.060 14 0.003 0.054 0.054 8
α5 0 − 0.012 0.057 0.058 11 − 0.010 0.050 0.051 9
α6 0 0 0.070 0.070 22 0.001 0.058 0.058 14
α7 0 0.004 0.064 0.064 15 0.006 0.053 0.053 9
α8 0.3 − 0.002 0.062 0.062 199 − 0.004 0.060 0.060 200
α9 0 0.001 0.052 0.052 11 0 0.046 0.046 5
α10 0 0.014 0.111 0.112 8 0.014 0.090 0.091 4
α11 0.5 − 0.010 0.151 0.151 186 − 0.053 0.150 0.159 176
α12 0 − 0.004 0.054 0.054 13 − 0.005 0.048 0.049 6
α13 0 0.006 0.114 0.113 59 0.002 0.084 0.084 39
α14 0 0.001 0.070 0.070 19 0 0.058 0.058 7
α15 0 − 0.003 0.067 0.066 11 − 0.002 0.056 0.055 5
α16 0 − 0.001 0.049 0.049 11 0.001 0.044 0.044 5
α17 0 − 0.001 0.055 0.054 12 − 0.002 0.049 0.049 9
α18 0 − 0.006 0.060 0.060 18 − 0.005 0.053 0.053 11
α19 0 − 0.003 0.060 0.060 15 − 0.002 0.050 0.050 5
α20 0 − 0.004 0.066 0.066 17 − 0.003 0.057 0.057 9
α21 0 0 0.057 0.056 15 0 0.050 0.050 9
α22 0 0.006 0.057 0.057 14 0.005 0.045 0.045 6

Table 7.

Simulation results of indirect effects (IE) intermediated through the first mediator based on the B–K and regTS models in the presence of 3 mediators and 22 exposures

IE True B–K regTS
Bias SD rMSE #Sig Bias SD rMSE #Sig
δ1,1 0 − 0.002 0.016 0.017 5 0 0.014 0.014 3
δ2,1 0.16 0.001 0.043 0.043 199 − 0.009 0.039 0.040 198
δ3,1 0 0.002 0.020 0.020 12 0.001 0.015 0.015 4
δ4,1 0 − 0.002 0.020 0.020 10 − 0.001 0.017 0.017 10
δ5,1 0 0.002 0.021 0.021 12 0.002 0.018 0.018 6
δ6,1 0.04 0 0.026 0.026 93 − 0.005 0.021 0.022 85
δ7,1 0 0.001 0.025 0.025 19 0 0.019 0.019 12
δ8,1 0.08 0.004 0.028 0.028 197 − 0.003 0.026 0.026 198
δ9,1 0 0.001 0.017 0.017 8 0 0.015 0.015 5
δ10,1 0.16 0.007 0.056 0.056 196 − 0.011 0.051 0.052 192
δ11,1 − 0.12 − 0.006 0.061 0.061 133 0.013 0.056 0.057 104
δ12,1 0 0.002 0.020 0.020 9 0.002 0.017 0.017 5
δ13,1 0 − 0.005 0.043 0.043 65 0 0.031 0.031 44
δ14,1 0 − 0.001 0.026 0.026 7 − 0.002 0.021 0.021 4
δ15,1 0 0 0.022 0.022 7 0.001 0.017 0.017 2
δ16,1 0 0.001 0.019 0.019 12 0.001 0.016 0.016 9
δ17,1 0 0.002 0.020 0.020 9 0.002 0.017 0.017 9
δ18,1 0 − 0.001 0.022 0.022 17 − 0.001 0.018 0.018 13
δ19,1 0.04 − 0.001 0.023 0.023 95 − 0.007 0.020 0.021 76
δ20,1 0 0.002 0.024 0.024 9 0 0.020 0.020 8
δ21,1 0 0 0.021 0.020 9 0 0.018 0.017 4
δ22,1 0 0 0.021 0.021 12 − 0.001 0.017 0.017 5

The indirect effect δj,1=β1γj1 for j=1,2,,22. #Sig denotes the number of times the corresponding effect is significant (zero falling outside 95% HPD interval) among the 200 replicates

Table 8.

Simulation results of indirect effects (IE) intermediated through the second mediator based on the B–K and regTS models in the presence of 3 mediators and 22 exposures

IE True B–K regTS
Bias SD rMSE #Sig Bias SD rMSE #Sig
δ1,2 0 0 0.009 0.009 3 − 0.001 0.008 0.008 2
δ2,2 0.02 0 0.013 0.013 88 − 0.002 0.012 0.012 76
δ3,2 0 − 0.001 0.010 0.010 1 − 0.001 0.008 0.008 2
δ4,2 0 0.001 0.011 0.011 6 0.002 0.009 0.009 4
δ5,2 0 − 0.001 0.012 0.011 7 − 0.001 0.010 0.010 4
δ6,2 0 0 0.013 0.013 11 0 0.010 0.010 7
δ7,2 0 0.001 0.014 0.014 15 0 0.011 0.011 8
δ8,2 0 − 0.002 0.010 0.010 11 − 0.001 0.009 0.009 7
δ9,2 0 0 0.008 0.008 1 0 0.007 0.007 2
δ10,2 − 0.06 − 0.006 0.034 0.034 121 0.005 0.030 0.030 97
δ11,2 0.08 0.004 0.044 0.044 118 − 0.011 0.038 0.040 93
δ12,2 0 0 0.010 0.010 3 0 0.009 0.009 1
δ13,2 − 0.02 − 0.002 0.024 0.024 64 0.003 0.019 0.019 41
δ14,2 0 0 0.013 0.013 6 0 0.010 0.010 3
δ15,2 0 0.001 0.012 0.012 6 0 0.009 0.009 1
δ16,2 0 − 0.001 0.010 0.010 6 − 0.001 0.008 0.008 3
δ17,2 0 0 0.011 0.011 5 0 0.010 0.010 4
δ18,2 0 0.001 0.011 0.012 10 0 0.010 0.010 6
δ19,2 0 0.001 0.011 0.011 6 0 0.009 0.009 2
δ20,2 0 0 0.013 0.013 5 0 0.010 0.010 4
δ21,2 0 0 0.011 0.011 8 0 0.009 0.009 3
δ22,2 0 0 0.010 0.010 5 − 0.001 0.009 0.009 7

The indirect effect δj,2=β2γj2 for j=1,2,,22. #Sig denotes the number of times the corresponding effect is significant (zero falling outside 95% HPD interval) among the 200 replicates

Table 9.

Simulation results of indirect effects (IE) intermediated through the third mediator based on the B–K and regTS models in the presence of 3 mediators and 22 exposures

IE True B–K regTS
Bias SD rMSE #Sig Bias SD rMSE #Sig
δ1,3 0 0 0.003 0.003 0 0 0.002 0.002 0
δ2,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ3,3 0 0 0.003 0.003 0 0 0.002 0.002 0
δ4,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ5,3 0 0 0.003 0.003 0 0 0.002 0.002 0
δ6,3 0 0 0.004 0.004 0 0 0.003 0.003 0
δ7,3 0 0 0.004 0.004 0 0 0.003 0.003 0
δ8,3 0 0 0.003 0.003 0 0 0.002 0.002 0
δ9,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ10,3 0 0 0.007 0.007 0 0 0.005 0.005 0
δ11,3 0 0.001 0.007 0.007 0 0.001 0.005 0.006 0
δ12,3 0 0 0.003 0.003 1 0 0.003 0.003 0
δ13,3 0 0 0.007 0.007 0 0 0.004 0.004 0
δ14,3 0 0 0.004 0.004 0 0 0.003 0.003 0
δ15,3 0 0 0.003 0.003 0 0 0.003 0.002 0
δ16,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ17,3 0 0 0.003 0.003 0 0 0.002 0.002 0
δ18,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ19,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ20,3 0 0 0.004 0.004 1 0 0.003 0.003 0
δ21,3 0 0 0.003 0.003 0 0 0.003 0.003 0
δ22,3 0 0 0.003 0.003 0 0 0.003 0.003 0

The indirect effect δj,3=β3γj3 for j=1,2,,22. #Sig denotes the number of times the corresponding effect is significant (zero falling outside 95% HPD interval) among the 200 replicates

For each generated data, we apply B–K and regTS with the same prior settings and pre-specified parameters as in the real data analysis, and run a single MCMC chain of 10, 000 iterations after 5000 burn-ins for each model. Based on 200 replicates, we present the average biases, standard deviations (SD), and square root of mean square errors (rMSE) of each effect. We also use 95% HPD interval to decide whether an effect is significant or not in a particular replicate, and report its frequency across simulation replicates as #Sig.

Tables 6, 7, 8, and 9 summarize the simulation results of 22 direct as well as 66 indirect effects through each of the three mediators, respectively. Overall, regTS has smaller rMSEs than B–K in all direct and indirect effects estimates. The proposed regTS also performs well in identifying significant direct and indirect effects. For example, while B–K correctly identifies the two non-zero direct effects (α8 and α11) 199 and 186 times, respectively, regTS does that 200 and 176 times correspondingly (Table 6). We note that, compared to B–K, regTS slightly underperforms in sensitivity (correctly identifying true non-zero effects), but outperforms in specificity (correctly identify true zero effects), as indicated by the #Sig columns in Tables 6, 7, 8, and 9. We view this a desirable feature in that it can efficiently narrow down the targeted risks factors while retain reasonable sensitivity. While the proposed approach might not be suitable for signal detection in early-phase studies, it can avoid unnecessary waste in prevention studies, especially when the resources are limited. Also note that the relatively large biases in regTS can be partly attributed to the use of shrinkage priors, which causes attenuated effect estimates. As shown in Tables 6, 7, 8, and 9, elevated negative biases are observed in true positive non-zero effects while elevated positive biases observed in true negative non-zero effects. The shrinkage priors do not cause more biases in true zero effects.

To further evaluate the performance of regTS under different scenarios, we also repeat the above simulation by varying the number of exposures. In particular, we consider the simulation where we halve or double the number of exposures so that p=11 or 44 while q remains at 3. These results are reported in Tables S1S4 for p=11 and S5S8 for p=44. Similar findings are observed as in p=22 above, with regTS producing smaller rMSE in almost all of the effect estimates and achieving slightly lower sensitivities but higher specificities. These findings suggest that the proposed regTS works well with different numbers of exposures.

A third simulation is also carried out to demonstrate the performance of regTS in comparison to regDOC of [18]. Using a similar set up as in the first simulation, we considered 3 mediators either highly correlated or uncorrelated. In the former, Pearson correlations of 0.75, 0.625, and 0.5 are used between mediator 1 and 2, 1 and 3, and 2 and 3, respectively. The results are reported in Tables S9S12 for highly correlated and S13S16 for uncorrelated mediators. When mediators are correlated, regDOC produces larger rMSE in almost all effect estimates, especially in indirect effects, or in effects that are truly non-zero. For example, estimate of the indirect effect of the second exposure mediated through the first mediator (δ2,1 in Table S10) has an rMSE of 0.088 under regDOC but 0.040 under the proposed regTS. The inferior performance of regDOC is more pronounced when we consider effects selection. While regDOC is comparable with regTS in correctly identifying true non-zero direct effects, it mis-identifies true zero direct effects as non-zero (Table S9) much more frequently. It does much worse in selecting indirect effects. Tables S10S12 show that regDOC rarely identifies any indirect effects as significant, no matter whether they are true zero or not. In comparison, the proposed regTS has quite reasonable sensitivities and specificities. When mediators are uncorrelated, regDOC has improved rMSEs, although these rMSEs are still higher than those from regTS. However its poor performance in identifying significant effects, especially indirect effects, remains; see Tables S13S16.

5. Discussion

In this paper, we combine regularizations and the general B–K model in the case of multiple mediators and multiple exposures. The shrinkage priors enable simultaneous variable selection and effects estimation. As shown in the data analysis and simulation study, the proposed method tends to be parsimonious, can capture the most relevant effects, and yield more precise effect estimates. The proposed method may serve as a useful tool in developing targeted prevention programs in obesity control.

Although the illustrated example consists of continuous mediators and outcome, the proposed method can be easily modified and extended to other types of data including categorical exposures or mediators and time-to-event or discrete outcomes. Unlike the early work that places shrinkage priors on the effects of interest [18], the proposed method does not require working with direct or indirect effects. Instead, it puts L1 penalty on the regression coefficients in each model component. This avoids the pitfall in the difference-of-coefficient approach where it can be difficult to evaluate indirect effects if the mediators are highly correlated.

In this paper, we consider continuous BMI z score to investigate the intergenerational effects of maternal dietary risk factors on childhood adiposity. Based on the fitted model and the MCMC sample, it is possible to obtain the predictive distributions of BMI z score for any levels of risk factors and then assess the probability of childhood obesity at that given level simply by comparing the distribution with the known cutoff point.

Considering the rich literature on regularizations, it is also worthwhile to consider other priors such as the elastic Lasso and the spike and slab. Finally, investigating the trajectory of BMI z score over time may provide useful insights into the development of prevention programs, as compared to considering BMI z score at a single time point.

Supplementary Material

Supp

Supplementary Information The online version contains supplementary material available at https://doi.org/10.1007/s12561-021-09305-7.

Acknowledgements

This research was supported by the Intramural Research Program of Eunice Kennedy Shriver National Institute of Child Health and Human Development.

References

  • 1.McIntyre HD, Catalano P, Zhang C, Desoye G, Mathiesen ER, Damm P (2019) Gestational diabetes mellitus. Nat Rev Dis Primers 5(1):47–47 [DOI] [PubMed] [Google Scholar]
  • 2.Zhu Y, Olsen SF, Mendola P, Halldorsson TI, Yeung EH, Granström C, Bjerregaard AA, Wu J, Rawal S, Chavarro JE, Hu FB, Zhang C (2017) Maternal dietary intakes of refined grains during pregnancy and growth through the first 7 y of life among children born to women with gestational diabetes. Am J Clin Nutr 106(1):96–104 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zhu Y, Olsen SF, Mendola P, Halldorsson TI, Rawal S, Hinkle SN, Yeung EH, Chavarro JE, Grunnet LG, Granström C, Bjerregaard AA, Hu FB, Zhang C (2017) Maternal consumption of artificially sweetened beverages during pregnancy, and offspring growth through 7 years of age: a prospective cohort study. Int J Epidemiol 46(5):1499–1508 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Zhang C, Olsen SF, Hinkle SN, Gore-Langton RE, Vaag A, Grunnet LG, Yeung EH, Bao W, Bowers K, Liu A, Mills JL, Sherman S, Gaskins AJ, Ley SH, Madsen CM, Chavarro JE, Hu FB (2019) Diabetes & Women’s Health (DWH) Study: an observational study of long-term health consequences of gestational diabetes, their determinants and underlying mechanisms in the usa and denmark. BMJ Open 9(4):e025517 [Google Scholar]
  • 5.Olsen J, Melbye M, Olsen SF, Sørensen TI, Aaby P, Andersen A-MN, Taxbøl D, Hansen KD, Juhl M, Schow TB, Sørensen HT, Andresen J, Mortensen EL, Olesen AW, Søndergaard C (2001) The Danish national birth cohort—its background, structure and aim. Scand J Public Health 29(4):300–307 [DOI] [PubMed] [Google Scholar]
  • 6.Baron RM, Kenny DA (1986) The moderator-mediator variable distinction in social psychological research: conceptual, strategic, and statistical considerations. J Pers Soc Psychol 51(6):1173–1182 [DOI] [PubMed] [Google Scholar]
  • 7.Robins JM, Greenland S (1992) Identifiability and exchangeability for direct and indirect effects. Epidemiology 3(2):143–155 [DOI] [PubMed] [Google Scholar]
  • 8.Pearl J (2001) Direct and indirect effects. In Proceedings of the 17th Conference on Uncertainty in Artificial Intelligence 411–420. Morgan Kaufmann, San Francisco [Google Scholar]
  • 9.Imai K, Keele L, Tingley D (2010) A general approach to causal mediation analysis. Psychol Methods 15(4):309–334 [DOI] [PubMed] [Google Scholar]
  • 10.VanderWeele TJ, Vansteelandt S (2010) Odds ratios for mediation analysis for a dichotomous outcome. Am J Epidemiol 172(12):1339–1348 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sohn MB, Li H (2017) Compositional mediation analysis for microbiome studies. Tech. rep, 10.1101/149419 [DOI] [Google Scholar]
  • 12.Daniels MJ, Roy JA, Kim C, Hogan JW, Perri MG (2012) Bayesian inference for the causal effect of mediation. Biometrics 68(4):1028–1036 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Preacher KJ, Hayes AF (2008) Asymptotic and resampling strategies for assessing and comparing indirect effects in multiple mediator models. Behav Res Methods 40(3):879–891 [DOI] [PubMed] [Google Scholar]
  • 14.VanderWeele T, Vansteelandt S (2014) Mediation analysis with multiple mediators. Epidemiol Methods 2(1):95–115 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Chén OY, Crainiceanu C, Ogburn EL, Caffo BS, Wager TD, Lindquist MA (2015) “High-dimensional multivariate mediation with application to neuroimaging data. Tech. rep arXiv:1511.09354 [Google Scholar]
  • 16.Huang Y-T, Pan W-C (2016) Hypothesis test of mediation effect in causal mediation model with high-dimensional continuous mediators. Biometrics 72(2):402–413 [DOI] [PubMed] [Google Scholar]
  • 17.Zhao Y, Luo X (2016) Pathway Lasso: Estimate and select sparse mediation pathways with high dimensional mediators. Tech. rep, arXiv:1603.07749 [Google Scholar]
  • 18.Wang Y-B, Chen Z, Goldstein JM, Buck Louis GM, Gilman SE (2019) Bayesian regularized mediation analysis with multiple exposures. Stat Med 38(5):828–843 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Park T, Casella G (2008) The Bayesian Lasso. J Am Stat Assoc 103(482):681–686 [Google Scholar]
  • 20.Zhang Q (2019) High dimensional mediation analysis with applications to causal gene identification. Tech. rep [Google Scholar]
  • 21.Song Y, Zhou X, Zhang M, Zhao W, Liu Y, Kardia SL, Roux AVD, Needham BL, Smith JA, Mukherjee B (2020) Bayesian shrinkage estimation of high dimensional causal mediation effects in omics studies. Biometrics 76(3):700–710 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supp

RESOURCES