Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2016 Jan 31.
Published in final edited form as: Soc Sci Med. 2014 Aug 2;127:171–180. doi: 10.1016/j.socscimed.2014.08.001

What can genes tell us about the relationship between education and health?

Jason D Boardman 1, Benjamin W Domingue 2, Jonathan Daw 3
PMCID: PMC4314507  NIHMSID: NIHMS625225  PMID: 25113566

Abstract

We use genome wide data from respondents of the Health and Retirement Study (HRS) to evaluate the possibility that common genetic influences are associated with education and three health outcomes: depression, self-rated health, and body mass index. We use a total of 1.7 million single nucleotide polymorphisms obtained from the Illumina HumanOmni2.5-4v1 chip from 4,233 non-Hispanic white respondents to characterize genetic similarities among unrelated persons in the HRS. We then used the Genome Wide Complex Trait Analysis (GCTA) toolkit, to estimate univariate and bivariate heritability. We provide evidence that education (h2 = .33), BMI (h2 = .43), depression (h2 = .19), and self-rated health (h2 = .18) are all moderately heritable phenotypes. We also provide evidence that some of the correlation between depression and education as well as self-rated health and education is due to common genetic factors associated with one or both traits. We find no evidence that the correlation between education and BMI is influenced by common genetic factors.

Keywords: education, health, depression, self-rated health, BMI, genetics

Introduction

This issue of Social Science & Medicine explores the relationship between health and education with a particular emphasis on the evidence supporting or questioning the causal nature of the association. Given that genetic factors are associated with both education (Branigan et al. 2013) and health (Sullivan et al. 2000; Mosing et al. 2011; Elks et al. 2012) it is possible that some of the association between education and health is due, in part, to common genetic influences. In particular, pleiotropic effects, the effect of a single gene on multiple traits, could lead to a form of omitted variable bias in estimates of causal effects if genes are ignored in studying the relationship between education and health.

There is a substantial literature on pleiotropy (see Solovieff et al., 2013, for a review). Such effects are not rare: “A recent evaluation of genome-wide-significant single-nucleotide polymorphisms (SNPs) listed in the National Human Genome Research Institute (NHGRI) Catalogue of Published Genome-Wide Association Studies found that 4.6% of SNPs and 16.9% of genes have CP [cross-phenotype] effects” (Solovieff et al., 2013, 484). For example, the FTO gene has been linked to both BMI (Thorleifsson, et al., 2008) and melanoma (GenoMEL Consortium, 2013). While most pleiotropic effects have been established when considering disease, serotonergic and dopaminergic genes are thought to be associated with a range of outcomes more likely to be of interest to social scientists (Daw & Guo 2011; Daw et al. 2013; Guo et al. 2008a, b; Lesch et al. 1996; McHugh et al. 2010; Sabol et al. 1999). Although techniques are still evolving for the identification of pleiotropy, it is important that social scientists consider its research implications.

Previous researchers have used genetically-informed designs to examine the potential for omitted variable bias in causal analyses, but have typically employed family-based studies. For example, Fletcher and Frisvold (2009) use sibling data from the Wisconsin Longitudinal Study to examine the influence of obtaining a college degree on the use of preventative health care use. These authors examine discordant pairs and show that the college educated sibling is more likely than his or her less educated sibling to use multiple types of preventative care. Identical twin pairs are useful in this design since any within pair association between education and health cannot be due to genes. In many cases, however, including the paper by Amin, Behrman, and Kohler (2014) in this issue, researchers either fail to detect an association between education and health outcomes within identical twin pairs or the associations are dramatically reduced in magnitude (Fujiwara and Kawachi 2009; Lundborg 2013; Webbink et al. 2010; Madsen et al. 2010). This suggests that some of the link between education and health may be confounded with genetic factors related to one or both phenotypes.

However, there are serious questions regarding the generalizability of samples of identical twins (Burt and Simons 2014). Alternative methods are available that avoid this limitation. Recent advances in statistical genetics have made it possible to decompose the covariance between two traits into genetic and environmental components using genome-wide data from unrelated persons. In this paper, we review the univariate and bivariate twin models and describe previously published heritability estimates for education and health. We then describe the extension of these models to unrelated persons in which the relationship matrix for each pair of observations in the study is based on measured genetic similarity (based on molecular genetic data) rather than assumed genetic similarity (based on familial relationships). We illustrate the utility of this perspective by using genome wide data from roughly 4,000 adults in the Health and Retirement Study (HRS) for three health outcomes: self-rated health, depression, and body mass index. We also provide detail about the sensitivity of this model to detect genetic correlations and discuss the utility of this perspective in social epidemiologic inquiry.

Heritability, twins, and unrelated persons

The univariate twin model is the backbone of behavioral genetics research (Neale and Maes 1996). This model compares phenotypic correlation among identical (MZ) twins with the phenotypic correlation among fraternal (DZ) twins to provide an estimate of the proportion of total phenotypic variation that is due to genetic variation in the population. The most simple heritability estimate is computed by taking twice the difference in the correlation between MZ and DZ twin pairs. For example, if the correlation between sibling pairs is due to combination of genetic (h2) and environmental (c2) components then the correlation for identical twins is given as rmz = h2+ c2 and the correlation for fraternal twins is given as rdz = h2/2+ c2. Because the contribution of the shared environment (c2) is assumed to be the same for MZ and DZ twins, solving for h2 provides the heritability estimate described above: 2(rmz −rdz). Heritability estimates from twin-studies are broad-sense heritability estimates in that they account for all genetic effects (additive, dominant, and epistatic).

These methods have been used across a number of different studies and the results show fairly consistent evidence about the heritability of education and the three health outcomes that we evaluate here. For example, Branigan et al. (2013) report results from a meta-analysis of twin studies on the heritability of education, calculating an average of 0.40. Sullivan et al. (2000) perform a meta-analysis of family-based studies used to derive heritability estimates for major depression. They provide a point estimate of 0.37 and a 95% confidence interval of 0.31 to 0.42 for broad-sense heritability of depression in the population. Elks et al. (2012) examine 88 different twin studies for BMI and provide a median estimate of 0.75 (25th–75th CI = 0.58, 0.87). Mosing et al. (2011) provide a summary of twin studies used to describe the heritability of self-rated health, reporting a range from 0.25 to 0.64. Importantly, one of these studies used a sample of older adults (mean age = 61) and reported a heritability estimate of 0.46 for self-rated health. While it is important to consider that these estimates show variability across studies (perhaps unsurprising given the gender and cohort differences between them), the typical conclusion is that each of these three health factors and education are all thought to be moderately or highly influenced by genetic factors.

However, previous research has made it clear that identical twin pairs are more likely than dizygotic twins to have the same friends, classmates, spend more time with one another, and dress the same (Richardson 2011; Cronk et al. 2002). Because the assumption of equal environments between MZ and DZ twins undergirds this twin model, many have been skeptical of heritability estimates derived from this formulation (Horowtiz et al. 2003). Some have argued that violations of this assumption may cause an important bias in quantitative genetic estimates (Joseph 2002) but others have argued that it is not necessarily the case (Conley et al., 2013). An additional concern is whether there are limitations in generalizing from twins to a larger population (Joseph, 2002). This and other important limitations of twin and family based estimates of heritability were detailed in a recent critique of the notion of heritability within criminological research (Burt and Simons 2014). Hence, it is important to consider other approaches to estimate the contribution of genetic variation to phenotypic variation.

Visscher et al. (2006) were the first to provide evidence that comparable heritability estimates could be derived from genome-wide data among siblings without making any assumptions about the shared environment of MZ and DZ twins. Full siblings share, on average, fifty percent of their genes by descent, but this quantity varies - in their study, full sibling pairs shared between 37.4 and 61.7 percent by descent. They show that the sibling pairs who are more genetically similar are also more similar to one another with respect to height. In particular, if the squared difference of heights for siblings is regressed on πi (the pair indicator of identity by descent, or IBD) then the narrow sense heritability (variance due solely to additive genetic effects, ignoring dominant and epistatic effects), is described as –β/2σ2p where σ2p is the variance of the trait. Thus, a comparison of genetic similarity to pair similarity in height yielded a heritability of .80

More recently, this same approach has been extended to unrelated persons and has been integrated into the suite of statistical packages called Genome-wide Complex Trait Anlaysis (GCTA) (Yang et al. 2011). As with Visscher et al. (2006), the key component of the GCTA analysis is the genetic relationship matrix (GRM). In Vissher et al’s (2006) approach, the GRM entries are IBD estimates among related pairs whereas the GCTA approach focuses on identity by state (IBS) estimates among unrelated pairs of individuals. For a standardized phenotype, they show an estimate of heritability that is similar to the approach used by Vissher et al. (2006). Although the calculation of the relationship matrix differs in the two approaches, the primary parameter estimates provide the same understanding of narrow sense heritability. The GCTA methods produce a heritability estimate of height (h2=.45) that is significantly smaller than the estimate provided by the Vissher et al. (2006) approach.

These methods are useful to social epidemiologists because they provide important information about health outcomes and health behaviors that seem to be influenced by genetic factors. The first part of our analyses utilizes these approaches to retrieve genome-wide heritability estimates for education, self-rated health, depression, and body-mass index. This is one of the first papers to use data from the Health and Retirement Study to describe these parameter estimates. Importantly, our study moves beyond univariate variance decomposition to estimate bivariate models in which the covariance between two traits is modeled using similar GRM techniques. This model is an extension of the bivariate twin model which examines the covariance between the education of the twin 1 with the health of twin 2 separately for identical and fraternal twin pairs (Neale et al. 2004). If the cross-twin-cross-trait covariance is significantly higher for identical compared to fraternal twins, then it suggests that some of the population level covariance between education and health may be due to genetic factors that structure both traits. This same design can now be used with estimates of genetic similarity from unrelated person, thereby avoiding the potential limitations of twin-based modeling.

The most important parameter estimate obtained from the bivariate model is the genetic correlation coefficient (rG), which is an estimate of the additive genetic association that is common to both traits (Neale and Maes, 1996). The existence of rG would imply that education and health share a common genetic etiology. It is important to note that rG is net of the overall variation in the phenotypes due to genes. That is, two traits which are only weakly heritable may have an rG similar to two traits which are strongly heritable. Silventoinen et al. (2004) use data from nearly 7,000 twin pairs in Finland and the United States to compare the correlations between education and height across MZ and DZ pairs. They show strong evidence for genetic influences on height for Finnish (h2 ~ 0.78) and US adults (h2 = 0.74) adults and moderate evidence for educational attainment for the two groups (h2 ~ 0.46 and 0.28, respectively). The also demonstrate that the bulk of the correlation between education and height is due to environmental factors that are shared by members of a family (rC) and that there is virtually no evidence of a significant rG. The failure to find a significant rG in these studies is perhaps not surprising given the nature of the outcome (height) and the relatively low bivariate correlation between education and height in the US (red, height~ 0.10 in the US). They do find some evidence for a statistically significant rG estimate for Finnish women and they state that roughly 35% of the correlation between education and height (red, height~ 0.14) in their study may be due to common genetic influences. Vermeiren et al. (2012) examine comparable models among 388 Belgian twin pairs for five highly heritable metabolic factors including glucose, HDL-cholesterol, systolic blood pressure, triglycerides, and waist circumference (WC). Their results show that each trait has a heritability estimate ranging from a low of 0.53 (triglycerides) to a high of 0.87 (HDL-cholesterol). They find that roughly 15% of the association between education and WC is due to common genetic factors. Finally, Johnson et al. (2010) use data from the Danish Twin Registry with information on a 12 item physical health scale including activities, limitations, and self-assessed health status. They find strong evidence for a genetic correlation between education and this composite indicator of health.

Taken together, this body of work suggests three conclusions. First, educational attainment is a moderately heritable phenotype. Roughly 40% (Branigan et al., 2013) of the population variance in completed years of school is attributable to genetic variance in the population. Second, many physical and mental health outcomes of interest to social epidemiologists are also moderately heritable. Average heritability estimates for self-rated health and depression are roughly 40% and BMI’s heritability is significantly higher, at 75%. Third, increases in education are negatively associated with each of these outcomes. Results from twin/sibling studies suggest that some of this association may be due to genes that influence both traits. Here, we re-examine many of these issues using molecular genetic data instead of assumed genetic relationships using a large and representative sample of older adults in the US with detailed measures of health and genome wide data including over 1.7 million single nucleotide polymorphisms (SNPs).

Data and Methods

Phenotypic Data

This paper uses data from the Health and Retirement Survey (HRS) RAND fat files (http://www.rand.org/labor/aging/dataprod/enhanced-fat.html). Of the 9,186 individuals in the fat files who also have genetic data, we work with the subset of non-Hispanic white respondents in order to minimize the problems caused by population stratification (see next section for additional details). We also focus on the cohort born between 1930 and 1950 so that changes in health and education occurring over different generations do not bias our estimates (though we also consider results without this restriction). We thus remove 2,147 respondents who are Hispanic or non-white and 2,450 who are born outside of our cohort of interest. We then remove 105 individuals who are missing information on the variables of interest. Finally, to ensure that there are no cryptically related individuals in our sample who might bias heritability estimates, we also randomly remove one individual from each of the 251 pairs that are estimated to have genetic similarities greater than 0.025. Although this reduces our sample, this is done to ensure that the remaining individuals (N=4,233) are unlikely to share family environments (which more closely-related individuals may).

Education is known to have a genetic component in terms of both heritability (Heath et al., 1985; Branigan et al., 2013) and at the level of individual SNPs (Rietveld et al., 2013). In the HRS, education is measured as total years of education. We focus on three health-related phenotypes: BMI, depression, and self-reported health. It is crucial to note that each has been shown to be highly heritable (Haberstick et al., 2010; Johnson et al., 2002; Romeis et al., 2000), a necessary condition for rG to be meaningful. Furthermore, they are each associated with education and other measures of socioeconomic status (Hermann et al., 2011; Miech & Shanahan, 2000; Mirowsky & Ross, 2008). The descriptive statistics for all variables are shown in Table 1.

Table 1.

Descriptive Statistics for all variables used in the analyses

Mean SD r(x, education)
Years of completed education 13.32 2.49 1.00
Self-rated health 2.49 0.84 −0.32
Body Mass Index 27.72 5.04 −0.08
Depression 0.10 0.19 −0.20
Birth Year 1939 5.83 0.13
Male 0.43 0.50 0.07

Note: Data come from all waves of the Health and Retirement Study. N = 4,233.

The first health outcome is self-rated health. Self-rated health is an important predictor of morbidity and mortality (Chandola and Jenkinson, 2000; Idler and Benyamini, 1997) and its predictive validity appears to be increasing over time (Schnittker and Bacak, 2014). In the HRS, it is measured as the response to the question, “Would you say your health is excellent, very good, good, fair, or poor?” Responses were measured in the corresponding five categories, coded so that higher numerical values indicate poorer health. Consistent with typical socioeconomic gradients in health, this measure has a correlation of −0.32 with education. Body mass index (BMI) is measured as a function of body weight in kilograms and height in meters (kg/m2). Both height and weight are self-reported. BMI is only weakly correlated with education (−0.08). Finally, depression is measured using the seven-item CES-D scale, calculated as the sum of five negative affect items and two (reverse-coded) positive affect items, and is correlated at −0.20 with education. For the health outcomes, we used the mean score for each respondent over all available waves of data in order to simultaneously minimize measurement error and obtain maximum information about each respondent’s health.

Genetic Data

Genetic data for the HRS is based on DNA samples collected in two phases. The first phase was collected via buccal swabs in 2006 using the Quiagen Autopure method. The second phase used saliva samples collected in 2008 and extracted with Oragene. Genotype calls were then made based on a clustering of both data sets using the Illumina HumanOmni2.5-4v1 array. Details on this process can be found online at the HRS website: http://hrsonline.isr.umich.edu/sitedocs/genetics/HRS_QC_REPORT_MAR2012.pdf. After standard quality control procedures (such as removing SNPs that were missing in more than 5% of samples, with minor allele frequency below 1%, or failure to meet Hardy-Weinberg equilibrium - complete details are available upon request), we retained 1,707,214 SNPs for the analysis.

We conduct analyses using non-Hispanic White respondents only. This is done to minimize the difficulties associated with using cross-race data. Individuals from different races are typically genetically distinct (easily identified due to variations in allele frequencies) but these distinctions are not genetically meaningful in most cases. This is most easily observed using principal components (PCs) (Price et al., 2010). Using data from the full set of 9,186 HRS respondents, we estimated PCs and consider how they vary by self-reported race and ethnicity (shown in Figure 1). Figure 1A shows that there are two “spurs” off the cluster near the origin. The first spur running along the x-axis identifies respondents who self-identify as non-Hispanic and black (Figure 1C). Hispanics are largely grouped in the second spur (Figure 1D), although there is a large amount of variability amongst the Hispanics. This highlights the existence of very small be detectable genetic differences across racial groups, something known as population stratification (e.g., Cardon and Palmer, 2003). These subtle differences may lead to erroneous conclusions about the relative influence of genes on specific phenotypes (Hamer and Sirota, 2000) but it is critical to note that between-group genetic variation, across the entire genome, is a small fraction of total genetic variation (Long & Kittles 2003). Even though these differences are a small contribution to overall genetic variation in the population, the methods that we use rely on the measurement of genetic similarity between unrelated persons. Pairs of individuals who have the same racial or ethnic identification may appear to be far more genetically similar than pairs of individuals from different racial and ethnic groups (e.g., Domingue, et al., 2014, Figure S5). These exaggerated differences could bias subsequent estimates. For example, given the well-established health disparities in the population (Williams and Collins, 1995), social factors that cause health difference may be confounded with these small differences. Note that the estimates in Figure 1 use all respondents to estimate the PCs but we only use PCs calculated from our analytic sample of non-Hispanic white respondents in the analyses.

Figure 1.

Figure 1

First two principal components in by self-reported race/ethnicity in the HRS

Methods

The GCTA toolkit, described in detail elsewhere (Yang et al., 2011), provides a flexible and intuitive method of using cumulative information across the genome to describe the genetic similarity of all pairs of unrelated persons in a sample. This method relies upon a genetic relationship matrix which describes the genetic similarity of all pairs of respondents as a function over all available SNPs of: (1) the pairwise similarity at each SNP and (2) the allele frequency for each SNP. The resulting estimate of genetic similarity between any pair of individuals can be thought of as a weighted correlation (with weights being determined by allele frequencies) between individuals over all available SNPs. This Ajk matrix, the genetic relationship matrix (GRM), is the core of the GCTA models and is described in equation 1. Individuals j and k are characterized as genetically similar to one another when you consider the genotype of person 1 at SNP i (xi1 ∈ {0,1,2}), the genotype of person 2 at SNP i (xi2 ∈ {0,1,2}), the allele frequency of SNP i (pi), and the number of SNPs (N):

Ajk=1Ni=1N(xij-2pi)(xik-2pi)2pi(1-pi). (1)

GCTA heritability estimates measure additive genetic influences (narrow sense heritability) which is why they may be generally lower than estimates from twin studies. For example, Yang et al., 2010 estimate a heritability of 45% in a study of human height using GCTA. Twin-based studies typically find much higher estimates (e.g., a heritability of roughly 90% is reported in Macgregor et al., 2006). This is due at least in part to the fact that twin-based estimates are broad-sense heritability estimates while GCTA estimates of heritability only account for additive genetic effects (i.e., are narrow-sense heritability estimates). Yang et al. (2010) also highlight that LD structures within populations can lead to lower estimates of heritability. Thus, our estimates are likely to be lower bounds of the true genetic effects.

Like any other statistical method, properly interpreting GCTA estimates necessitates certain assumptions. One key assumption is that there are no systematic environmental differences (e.g., are more genetically similar individuals also likelier to grow up in certain types of environments) which may bias results. Some research (Conley et al., 2014) suggests that GCTA results may be robust to violations of this assumption, but more work in this area is needed. This univariate GCTA model has been recently extended to a bivariate specification in which the same genetic relationship can be used to examine covariance between two traits as a function of genetic similarity (Lee et al., 2012). The model is derived from twin-sibling approaches in which cross-twin-cross-trait covariance is compared among monozygotic and dizygotic twin pairs to see if there is evidence that the covariance is due to common genetic influences. Consider a model in which trait t is regressed on a set of fixed environmental effects (x) and normally distributed genetic effects (g):

yt=Xtbt+Ztgt+et. (2)

In this model, yt is a vector of observations for trait t (in this case trait 1 or trait 2) over all individuals, bt is a vector of fixed effects, gt captures a vector of random polygenic effects for each individual, and et are the residuals for the t-th trait. Equation 3 presents the covariance matrix between trait 1 and 2 that underlies these analyses:

V=[Z1AZ1σg12+σε12Z1AZ2σg1,g2Z2AZ1σg1,g2Z2AZ2σg22+σε22]. (3)

A is a matrix composed of the genetic similarity estimated discussed in the context of the univariate model. Z describes the incidence matrices for the effects of g for each of the t-th traits. The main diagonal describes the variance for each trait as a function of genetic (g) and environmental/residual (e) terms. The most important part of this specification for our purposes is the covariance between the genetic effects for each trait (σg1, g2) which can be standardized and expressed as a genetic correlation coefficient rG. This estimate is the primary focus of our analyses. Evidence that rG is not equal to 0 in the population suggests that there is a common genetic component underlying both traits which would complicate the claim of a causal relationship between education and health. We also explore the sensitivity of our results to inclusion of PCs in the univariate and bivariate estimates.

Results

Table 2 presents eight univariate GCTA models for the four phenotypes in our study. We present two models for each outcome that include the base model and a model in which we include controls for the top 5 PCs. Confidence intervals (95%) are shown for the heritability estimates. The key estimate in each of the models is the heritability for each trait which is simply the ratio of genetic variance to total variance. Three facts are worth noting. First, all phenotypes demonstrate moderate to large heritability estimates when PCs are not included: 0.36 for education, 0.46 for BMI, 0.35 for depression, and 0.22 for health. This education estimate is quite similar to the estimate from many twin and sibling studies (Branigan et al., 2013) and the other estimates are in line with published results for comparable measures (Haberstick et al., 2010; Johnson et al., 2002; Romeis et al., 2000). As indicated by the p-value in the bottom of each model, dropping the genetic path causes a significant loss of information for each of the four phenotypes, an additional indicator that the traits are partially explained by genotype. Second, while the heritability estimates were all reduced with controls for the top five PCs, the only phenotype to show a large reduction in magnitude was depression. Specifically, the PC controls reduced the heritability estimate for depression from .35 to .19 (a 45% reduction) but the same adjustment for education, BMI, and self-rated health led to 7, 6, 18% reductions, respectively. While not the primary aim of this paper, these results stress the importance of considering population stratification when using GCTA models to characterize quantitative genetic parameter estimates for health outcomes and health behaviors, even among relatively homogenous populations. It is also important to note that this reduced heritability of depression is roughly equivalent to the value (h2=.21) estimated by a well powered case-control design (Cross-Disorder Group of the Psychiatric Genomics Consortium 2013). Third, there are large confidence intervals for these estimates despite the relatively large sample size. For example, the estimate for education has a range from .16 to .55 which denotes a large difference in the substantive conclusions regarding the relative influence of genes on education.

Table 2.

Univariate genome wide heritability estimates for education and three health outcomes

Education Body Mass Index

Source of variance No controls 5 PCS No controls 5 PCS
 Genetic 2.219 2.056 11.561 10.859
 Environmental 3.992 4.125 13.855 14.505
 Total 6.211 6.181 25.417 25.364
Heritability 95% CI 0.357 (.163, .551) 0.333 (.135,. 529) 0.455 (.267, .643) 0.428 (.235, .620)
logL −5976.611 −5945.916 −8953.757 −8925.618
logL0 −5983.363 −5951.43 −8966.968 −8935.616
LRT 13.505 11.028 26.421 19.996
df 1 1 1 1
pr. < 0.0001 0.0004 1.00E-07 4.00E-06
Depression Self-Rated Health
Source of variance No controls 5 PCS No controls 5 PCS
 Genetic 0.012 0.007 0.153 0.124
 Environmental 0.023 0.028 0.555 0.578
 Total 0.035 0.034 0.708 0.703
Heritability 95% CI 0.348 (.164, .531) 0.192 (.015, .368) 0.216 (.025, .407) 0.177 (.001, .353)
logL 4994.264 5024.264 −1386.91 −1355.873
logL0 4985.194 5022.415 −1389.447 −1357.482
LRT 18.14 3.699 5.073 3.218
df 1 1 1 1
pr. < 1.00E-05 0.03 0.01 0.04

Note: Data come from the Health and Retirement Study; n = 4,233.

The results of the genome-wide bivariate association models, the primary focus of this paper, are presented in Table 3. The values in the third row (labeled phenotypic variance) describe the variance of each health measure and education and are similar to those presented in Table 2. The ratio of the genetic variance to the total variance still provides a heritability estimate but this estimate is now conditional upon the genetic variance of the 2nd trait. For example, the total variance for BMI is 25.345 of which 10.668 is genetic variation. This ratio (10.668/25.345) provides the heritability estimate in the table (h2 = .421) and the remaining variation is due to environmental factors. The unstandardized covariance estimates describe the extent to which the genetic variance is common between education and each health indicator. In the case of BMI, the value of the genetic covariance is very small (−.159) compared to the environmental covariance (−.788). The standardized covariance value provides the genetic correlation (rG) estimate of −.033 which is not significantly different from zero (95% CI [−0.297, 0.331]). This same conclusion can be seen in the statistically non-significant change in overall model fit without the rG estimate (−2LL = 0.032, df = 1, p < 0.40). Together, these suggest that, while both phenotypes are heritable, genes do not appear to influence the pathway linking education and BMI.

Table 3.

Bivariate genome wide covariance estimates for education and three health outcomes.

Body Mass Index Depression Self-Rated Health
Genetic variance
 Health 10.668 0.007 0.128
 Education 2.139 2.173 2.142
Cov(health, education) −0.159 −0.089 −0.477
Environmental variance
 Health 14.677 0.028 0.576
 Education 4.059 4.025 4.055
Cov(health, education) −0.788 −0.003 −0.178
Phenotypic Variance
 Health 25.345 0.034 0.704
 Education 6.197 6.198 6.198
Heritability
 Health 0.421 0.193 0.181
 Education 0.345 0.351 0.346
rG 95% CI (rg) −0.033 (−0.297, .331) −0.746 (−1.0, −0.201) −0.912 (−1.0, −0.374)
logL −14860.413 −837.535 −7089.678
logL0 (rG = 0) −14860.429 −841.035 −7094.394
LRT 0.032 6.999 9.432
df 1 1 1
pr. < 0.4 0.004 0.001

Note: Data come from the Health and Retirement Study; n = 4,233

In contrast, compare the relative magnitude of the genetic covariance estimates for depression and education (−.089) to the magnitude of the environmental covariance estimate (−.003). In this case the standardized genetic correlation coefficient is quite large (rG = −0.746) and the inclusion of this path significantly improves model fit (p<.004). The same results emerge for self-rated health in which we calculate a very large genetic correlation estimate (rG = −0.912) and the confidence interval and likelihood ratio test provide very strong evidence of statistical significance. These genetic correlation coefficients suggest that genetic factors may help explain some of the association between education and depression and education and self-rated health. However, as with the univariate results, the standard error estimates for these models are large. The substantive significance of these findings is discussed in the Conclusion section of this paper, but we first examine the robustness of these findings to a variety of assumptions.

Sensitivity Analyses

To demonstrate that our results are robust to a variety of modeling choices, we consider rG estimates from models based on several alternative set of modeling decisions. Results are shown in Table 4. We first consider how sensitive results are to varying the number of PCs included in the analyses. Crucially, the statistical significance of our results does not change as we vary the number of PCs included in our analyses from 2 to 10. The rG estimates were generally quite similar with the largest change being self-rated health/education rG estimates dropping from −0.929 (with 2 PCs) to −0.869 (with 10 PCs).

Table 4.

Results from sensitivity analyses for rG estimates.

BMI SE Depression SE SRH SE

Original Estimates −0.033 0.186 −0.746 0.278 −0.912 0.275
2 PCs −0.060 0.183 −0.764 0.281 −0.929 0.260
4 PCs −0.036 0.189 −0.781 0.286 −0.895 0.276
6 PCs −0.067 0.193 −0.728 0.288 −0.872 0.280
8 PCs −0.049 0.194 −0.758 0.292 −0.875 0.278
10 PCs −0.055 0.190 −0.757 0.295 −0.869 0.277
No DOB restrictions −0.147 0.189 −0.623 0.240 −0.635 0.188

Note: Data come from the Health and Retirement Study.

*

Analysis failed to converge.

We then consider results based on a broader pool of respondents when there are no restrictions on year of birth. These results are based on 6,395 individuals. Compared to the 4,233 individuals used in the main analyses, the expanded sample was quite similar. For example, they completed 13.30 years of education on average (compared to 13.32 in the original sample) and had a mean BMI of 27.48 (compared to 27.72 in the original sample). Results using this expanded sample were again qualitatively similar in terms of statistical significance although both the rG estimates for depression and self-rated health declined (to −0.623 and −0.635 respectively).

Finally, Webbink et al. (2010) demonstrate that there might be different associations between education and health by gender and Johnson et al. (2010) show that the genetic correlation between physical health and years of education is large and significant among women but negligible among men. To evaluate this possibility, we conducted gender-stratified analyses for the results in Table 3, however, the reduced sample sizes led to very large standard errors and fairly unreliable estimates and we were not able to make any concrete statements about the moderating role of gender. Gender differences are not the focus of this paper but we encourage future researchers to examine the mechanisms responsible for different levels of genetic correlation between education and health for men and women. In summary, our findings were not qualitatively altered by any of the modifications considered in this section.

Simulation results

Because these methods have not been widely applied, we investigated the sensitivity of the estimates to changes in several key factors by simulating data based on a random set of 500,000 SNPs from the 4,233 respondents in our analytic sample. We then generated two phenotypes based on these SNPs (using the relevant functionality from GCTA). We varied three conditions in simulation. First, we varied the heritability of the first and second traits. In doing so, we are varying the extent to which the trait is influenced by genes. Although the rG estimate should be net of the overall heritability of each trait, one might expect recovery of rG estimates to depend upon the heritability of the trait. Heritability estiamates were either 0.3, 0.6, or 0.9. Second, we varied the portion of the 500,000 SNPs that were causal (e.g., had an identified effect on the phenotype). The percentage of causal SNPs was either 1% or 0.01%. Effects were drawn from the standard normal distribution (and thus some effects are very near 0). Third, we varied the correlation of the SNP-level effects. If the SNP effects for two traits are highly correlated, then these two traits will have a high rG since the covariation between them will be genetic in origin. By controlling this facet via simulation we should be able to demonstrate direct control over the estimated rG parameters. SNP-level correlations were either 0.5 or 0.9. One limitation of our simulation is that all simulated traits have normal distributions while our empirical variables have different shapes. It is not trivial to mimic these shapes, however, since one must induce the shape via careful control of the effects for individual SNPs.

For all iterations of the simulation, results are in Table 5. Results are ordered by the value of rG estimate. Note that one combination of simulation conditions is left out since the rG estimated failed to converge. Furthermore, we also only considered those conditions in which the heritability of the first trait was less than or equal to the heritability of the second trait. Estimated rG was largely independent of heritability of the traits and the percent of causal SNPs. Essentially, there was no correlation between average true heritability across the two traits and estimated rG (r=0.05). A larger number of causal SNPs tended to produce pairs of traits with slightly higher rG, but the effect was rather weak (an IQR of 0.46 to 0.88 for 0.01% causal SNPs compared to 0.58 to 0.88 for 1% causal SNPs). As expected, the SNP level effect correlation had a huge impact on the estimated rG. When the SNP correlation was 0.5, the IQR of the estimated rG values was 0.43 to 0.60. For the 0.9 correlation, the IQR went from 0.82 to 0.93. This confirms that the rG estimate generated by GCTA accurately detected key changes in the simulation parameters.

Table 5.

Simulation Results

Iteration h2, 1 h2, 2 % Causal SNPs r (SNP- level) Phenotypic correlation rG rg (SE)
1 0.3 0.3 1.0 0.5 0.11 0.26 0.17
2 0.3 0.6 0.01 0.5 0.23 0.42 0.12
3 0.6 0.6 0.01 0.5 0.17 0.43 0.14
4 0.9 0.9 1.0 0.5 0.44 0.44 0.06
5 0.6 0.9 0.01 0.5 0.36 0.48 0.07
6 0.3 0.9 1.0 0.5 0.28 0.53 0.10
7 0.6 0.9 1.0 0.5 0.37 0.54 0.07
8 0.3 0.6 0.01 0.9 0.35 0.57 0.16
9 0.3 0.6 1.0 0.5 0.24 0.59 0.13
10 0.6 0.6 1.0 0.5 0.30 0.59 0.10
11 0.9 0.9 0.01 0.5 0.60 0.61 0.05
12 0.3 0.3 1.0 0.5 0.16 0.63 0.15
13 0.3 0.6 1.0 0.9 0.39 0.67 0.10
14 0.3 0.9 0.01 0.5 0.34 0.69 0.10
15 0.3 0.9 1.0 0.9 0.46 0.79 0.08
16 0.9 0.9 0.01 0.9 0.78 0.86 0.03
17 0.9 0.9 1.0 0.9 0.80 0.87 0.02
18 0.6 0.6 0.01 0.9 0.57 0.89 0.05
19 0.6 0.6 1.0 0.9 0.54 0.92 0.07
20 0.6 0.9 1.0 0.9 0.66 0.92 0.05
21 0.3 0.3 0.01 0.9 0.24 0.93 0.27
22 0.3 0.9 0.01 0.9 0.48 0.93 0.09
23 0.3 0.3 1.0 0.9 0.26 1.00 0.36

With Figure 2, we can compare the empirical results to those of the simulation. The results for self-rated health are similar to iterations 21–23 of the simulation (see Table 5). All of these iterations are characterized by low heritabilities and high correlations of SNP-level effects. Results for depression are similar to iterations 12 and 14 of the simulation. Both of these have a lower correlation of SNP-level effects. None of our simulation conditions produced phenotypes with observed correlations or an estimated rG as low as our empirical BMI/education result. This suggests that there is likely a weak correlation between the SNP-level effects of BMI and education.

Figure 2.

Figure 2

Simulation results with empirical results embedded.

An additional concern is the minimum rG we are able to distinguish from zero given this sample size. Given that the SEs for all three empirical rG estimates were 0.18 and 0.27, this suggests that it would be difficult to detect (in terms of statistical significance) a true rG of less than roughly 0.4. The estimated SEs in the simulation were lower than the observed values (e.g., they had a mean of 0.11). This suggests, not surprisingly, that the simulation fails to capture the full complexity inherent in real phenotypes. Only one of the simulated rG estimates was statistically non-significant. This was the lowest estimated rG in the simulation (0.26). In summary, it is likely that we only possess moderate power to detect significant rG effects below 0.4. Thus, our failure to detect a significant rG for BMI and education could be attributable to insufficient power.

Conclusion

A large and influential body of research has demonstrated that educational attainment is positively associated with better health (Adler & Rehkopf, 2008; Ross & Mirowsky, 1999). Whether measured as years of education or degree completion, higher levels of education are associated with reduced morbidities and comorbidities, delayed mortality, increased likelihood of leading a healthy lifestyle, increased access to and utilization of health care, increased sense of health-related efficacy, increased health-related knowledge, and many other factors linked to overall health status (Adler et al., 1994; Ross & Wu, 1995). Indeed, this special issue addresses many of these associations in detail. However, there has been substantial debate regarding the ordering of cause and effect. Do advances in education cause improved health, or does poor health inhibit educational attainment? Mirowsky and Ross (2003) are perhaps the most outspoken advocates of the perspective that increases in education are causally associated with increases in health. They argue that education provides problem solving skills and a sense of self-efficacy that they call “learned effectiveness” which characterizes a lifestyle in which the prioritization and knowledge of healthy choices is expected. On the other end of this continuum, some have argued that health causes education, as individuals with the poorest health are less likely than more healthy individuals to successfully accomplish on time educational goals (Haas, 2006; Palloni et al., 2009). Health shocks to children (Currie & Stabile, 2003) or parents (Boardman et al., 2012) constrain children’s success in school, which may lead to reduced likelihood of graduation because of low exam scores (Case et al., 2005).

While these two models debate the direction of causality, results from our paper (alongside the findings from previous research on twins) suggest that this framework could in some cases benefit from a consideration of the role of genes in the pathway linking education and health. We show that low education and poor mental health and poor self-rated health as they co-occur within certain individuals can be partially explained by genetic similarity. In the future, causal analyses of the relationship between education and health should consider whether more accurate estimates might be obtained if shared genes, even amongst unrelated individuals, are accounted for in the estimation of the causal relationship. Importantly, doing so properly requires additional knowledge of the nature of the pleiotropic effects in question.

To this end, it is important to highlight the distinction between “mediation pleiotropy” and “biological pleiotropy”. Assuming for the present that improved education is a cause of better health, mediation pleiotropy would imply that a genetic region affects health solely through that region’s impact on education. That is, the gene influences education directly and health only indirectly. In this scenario, a causal analysis of the effect of education on health would not suffer from omitted variable bias if it ignored genotype since the effect of genotype on health is captured by controlling for education. On the other hand, biological pleiotropy (in which a common gene is directly influencing both education and health), poses a greater threat to causal inference, since a failure to control for genotype in causal analyses will often lead to omitted variable bias. The aforementioned relationship between FTO and melanoma, for example, does not seem to be due to BMI (GenoMEL Consortium, 2013), so controlling for BMI would be insufficient in adjusting for genetic confounding in an analysis of melanoma. Distinguishing between these two types of pleiotropy is not easy and the methods demonstrated here are insufficient to do so. However, we recommend that researchers investigate the potential for genetic confounding when possible, taking advantage of recently-developed techniques for this purpose.

In considering the possibility of genetic confounding, it is helpful to consider a number of idealized scenarios (see Figure 3). Although real world relationships are likely to be more complex, these offer a starting point for thinking about how genes might bias causal estimates. Figure 3A represents a scenario in which analyses of health and education will be unbiased if genetics are ignored since the genes responsible for the two traits are disjoint (represented by the division of G into two halves)-that is, there is no pleiotropy. To the extent that our analyses based on genome-wide similarity are able to rule out pleitropic effects, our results suggest that this may be the case with BMI and education. Figures 3B and 3C represent examples of mediation pleiotropy. Such relationships have been the subject of much of the research cited earlier in this paper. Figure 3D presents the most complicated type of relationship: education and health are influenced by a common genetic source (biological pleiotropy) and also have a causal relationship amongst themselves. Given the significant rG estimates, the association between education and self-rated health or depression could be described by 3B, 3C, or 3D, but it is not yet clear which one. Distinguishing between these cases would require additional knowledge about the nature of the genetic effects on the traits that is presently lacking

Figure 3.

Figure 3

Examples of four hypothetical causal relationships between Education (E), Health (H), and Genetics (G).

Further complicating matters, genetic relationships sometimes need to be interpreted as a function of environment. For example, Johnson et al. (2010) find genetic correlations between education and self-rated health, but this correlation is stronger (rG=.75) among those with the highest levels of education compared to those with the lowest level of education (rG=.21). A lengthy consideration of the role of environment is beyond the scope of this paper, but we caution that genetic confounding might be more of a problem in some populations than others. Turning back to the present results, it is striking that the outcome typically thought to be most behaviorally influenced—BMI, which is directly influenced by both diet and activity level—shows the least evidence of genetic relationship to education. In contrast, depression and self-rated health show much stronger genetic correlation with education. Although we cannot directly test this possibility, this suggests that the mechanisms linking education to BMI are environmental in nature and unconditional on genetic factors. In other words, to the degree that education causally influences BMI, its average effect is unconditional on genetic factors, instead routing through mechanisms such as awareness of the benefits of healthful diets and activity levels and the ability to implement such a healthy lifestyle (Frohlich et al., 2001). As research on gene-environment interplay and health has become more common and prominent, numerous modes by which genetic effects may be conditional on environmental effects (and vice versa) have been described and tested (Boardman et al. 2013). However, far less attention has been given to the degree to which genomic factors may explain the relationship between an environment and a phenotype. Although gene-environment interactions are an example of such a relationship, they are but a subset of the causal configurations which could explain such an outcome. Accordingly, future researchers should pay more attention to the possibility that such relationships beyond gene-environment interactions and correlations are occurring in the explication of the social and genetic determinants of health.

Highlights.

  • Educational attainment, mental, and physical health are all moderately influenced by genotype.

  • Education and depression share common genetic influences.

  • The same is true for self-rated health but this is not the case for body mass index.

  • Genetic confounding should be considered when describing the education-health association.

Acknowledgments

This paper uses data from the Health and Retirement Study which is supported by National Institutes of Aging (U01 AG009740) and the Social Security Administration. The analyses for this paper were supported by grants from the Eunice Kennedy Shriver National Institute for Child Health and Human Development (NICHD) including R21 HD078031; R01 HD060726. Further support was provided by the NIH/NICHD funded University of Colorado Population Center (R24 HD066613).

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Contributor Information

Jason D. Boardman, University of Colorado, Boulder

Benjamin W. Domingue, University of Colorado, Boulder

Jonathan Daw, University of Alabama, Birmingham.

References

  1. Adler NE, Boyce T, Chesney MA, Cohen S, Folkman S, Kahn RL, et al. Socioeconomic-Status and Health - The Challenge of the Gradient. American Psychologist. 1994;49:15–24. doi: 10.1037//0003-066x.49.1.15. [DOI] [PubMed] [Google Scholar]
  2. Adler NE, Rehkopf DH. US disparities in health: Descriptions, causes, and mechanisms. Annual Review of Public Health. 2008;29:235–252. doi: 10.1146/annurev.publhealth.29.020907.090852. [DOI] [PubMed] [Google Scholar]
  3. Boardman JD, Alexander KB, Miech RA, MacMillan R, Shanahan MJ. The association between parent’s health and the educational attainment of their children. Social Science & Medicine. 2012;75:932–939. doi: 10.1016/j.socscimed.2012.04.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Boardman JD, Daw J, Freese J. Defining the environment in gene–environment research: lessons from social epidemiology. American Journal of Public Health. 2013;103(S1):S64–S72. doi: 10.2105/AJPH.2013.301355. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Branigan AR, McCallum KJ, Freese J. Variation in the Heritability of Educational Attainment: An International Meta-Analysis. Social Forces. 2013;92:109–140. [Google Scholar]
  6. Burt CH, Simons RL. Pulling back the curtain on heritability studies: biosocial criminology in the postgenomic era. Criminology. 2014;52(2):223–262. [Google Scholar]
  7. Case A, Fertig A, Paxson C. The Lasting Impact of Childhood Health and Circumstance. Journal of Health Economics. 2005;24:368–389. doi: 10.1016/j.jhealeco.2004.09.008. [DOI] [PubMed] [Google Scholar]
  8. Cardon LR, Palmer LJ. Population stratification and spurious allelic association. The Lancet. 2003;361(9357):598–604. doi: 10.1016/S0140-6736(03)12520-2. [DOI] [PubMed] [Google Scholar]
  9. Chandola T, Jenkinson C. Validating self-rated health in different ethnic groups. Ethnicity and Health. 2000;5(2):151–159. doi: 10.1080/713667451. [DOI] [PubMed] [Google Scholar]
  10. Conley D, Rauscher E, Dawes C, Magnusson PKE, Siegal ML. Heritability and the Equal Environments Assumption: Evidence from Multiple Samples of Misclassified Twins. Behavior Genetics. 2013;43:415–426. doi: 10.1007/s10519-013-9602-1. [DOI] [PubMed] [Google Scholar]
  11. Conley D, Siegal ML, Domingue B, Harris KM, McQueen M, Boardman J. Testing the key assumption of heritability estimates based on genome-wide genetic relatedness. Journal of human genetics. 2014 doi: 10.1038/jhg.2014.14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Cronk Nikole J, Slutske Wendy, Madden Pamela AF, Bucholz Kathleen K, Reich Wendy, Heath Andrew C. Emotional and behavioral problems among female twins: An evaluation of the equal environments assumption. Journal of the American Academy of Child & Adolescent Psychiatry. 2002;41:829–37. doi: 10.1097/00004583-200207000-00016. [DOI] [PubMed] [Google Scholar]
  13. Cross-Disorder Group of the Psychiatric Genomics Consortium. Genetic relationship between five psychiatric disorders estimated from genome-wide SNPs. Nat Genet. 2013;45(9):984–994. doi: 10.1038/ng.2711. http://dx.doi.org/10.1038/ng.2711. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Currie J, Stabile M. Socioeconomic status and child health: Why is the relationship stronger for older children? American Economic Review. 2003;93:1813–1823. doi: 10.1257/000282803322655563. [DOI] [PubMed] [Google Scholar]
  15. Daw J, Guo G. The Influence of Three Genes on Whether Adolescents Use Contraception, USA 1994–2002. Population Studies. 2011;65(3):253–271. doi: 10.1080/00324728.2011.598942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Daw J, Shanahan M, Harris KM, Smolen A, Haberstick B, Boardman JD. Genetic Sensitivity to Peer Behaviors: 5HTTLPR, Smoking, and Alcohol Consumption. Journal of Health and Social Behavior. 2013;54(1):92–108. doi: 10.1177/0022146512468591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Domingue BW, Fletcher J, Conley D, Boardman JD. Genetic and educational assortative mating among US adults. Proceedings of the National Academy of Sciences. 2014;111(22):7996–8000. doi: 10.1073/pnas.1321426111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Elks CE, den Hoed C, Zhao JH, Sharp SJ, Wareham NJ, Loos RJF, Ong KK. Frontiers in Endocrinology. Variability in the heritability of body mass index: a systematic review and meta-regression. 2012;28(3):1–16. doi: 10.3389/fendo.2012.00029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Fletcher Jason M, Frisvold David E. J Hum Cap. Higher Education and Health Investments: Does More Schooling Affect Preventive Health Care Use? 2009;3(2):144–176. doi: 10.1086/645090. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Frohlich KL, Corin E, Potvin L. A theoretical proposal for the relationship between context and disease. Sociology of Health & Illness. 2001;23:776–797. [Google Scholar]
  21. Fujiwara T, Kawachi I. Is education causally related to better health? A twin fixed-effect study in the USA. International Journal of Epidemiology. 2009;38:1310–1322. doi: 10.1093/ije/dyp226. [DOI] [PubMed] [Google Scholar]
  22. GenoMEL Consortium. A variant in FTO shows association with melanoma risk not due to BMI. Nature genetics. 2013;45(4):428–432. doi: 10.1038/ng.2571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Guo G, Roettger EM, Cai T. The integration of genetic propensities into social control models of delinquency and violence among male youths. American Sociological Review. 2008a;73:543–568. [Google Scholar]
  24. Haas SA. Health Selection and the Process of Social Stratification: The Effect of Childhood Health on Socioeconomic Attainment. Journal of Health and Social Behavior. 2006;47:339–354. doi: 10.1177/002214650604700403. [DOI] [PubMed] [Google Scholar]
  25. Haberstick BC, Lessem JM, McQueen M, Boardman JD, Hopfer CJ, Smolen A, et al. Stable genes and changing environments: body mass index across adolescence and young adulthood. Behavior Genetics. 2010;40:495–504. doi: 10.1007/s10519-009-9327-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Hamer D, Sirota L. Beware the chopsticks gene. Molecular psychiatry. 2000;5(1):11–13. doi: 10.1038/sj.mp.4000662. [DOI] [PubMed] [Google Scholar]
  27. Heath AC, Berg K, Eaves LJ, Solaas MH, Corey LA, Sundet J, Nance WE. Education policy and the heritability of educational attainment. 1985 doi: 10.1038/314734a0. [DOI] [PubMed] [Google Scholar]
  28. Hermann S, Rohrmann S, Linseisen J, May AM, Kunst A, Besson H, et al. The association of education with body mass index and waist circumference in the EPIC-PANACEA study. BMC Public Health. 2011;11 doi: 10.1186/1471-2458-11-169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Horwitz AV, Videon TM, Schmitz MF, Davis D. Rethinking Twins and Environments: Possible Social Sources for Assumed Genetic Influences in Twin Research. Journal of Health and Social Behavior. 2003;44(2):111–129. [PubMed] [Google Scholar]
  30. Idler Ellen L, Yael Benyamini. Self-Rated Health And Mortality: A Review Of Twenty-Seven Community Studies. Journal of Health and Social Behavior. 1997;38:21–37. [PubMed] [Google Scholar]
  31. Johnson W, Kyvik KO, Mortensen EL, Skytthe A, Batty GD, Deary IJ. Education reduces the effects of genetic susceptibilities to poor physical health. International Journal of Epidemiology. 2010;39:406–414. doi: 10.1093/ije/dyp314. [DOI] [PubMed] [Google Scholar]
  32. Johnson W, McGue M, Gaist D, Vaupel JW, Christensen K. Frequency and heritability of depression symptomatology in the second half of life: evidence from Danish twins over 45. Psychological Medicine. 2002;32:1175–1185. doi: 10.1017/s0033291702006207. [DOI] [PubMed] [Google Scholar]
  33. Joseph J. Twin Studies in Psychiatry and Psychology: Science or Pseudoscience? Psychiatric Quarterly. 2002;73:71–82. doi: 10.1023/a:1012896802713. [DOI] [PubMed] [Google Scholar]
  34. Lee SH, Yang J, Goddard ME, Visscher PM, Wray NR. Estimation of pleiotropy between complex diseases using SNP-derived genomic relationships and restricted maximum likelihood. Bioinformatics. 2012;28:2540–2542. doi: 10.1093/bioinformatics/bts474. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Lesch KP, Bengel D, Heils A, Sabol SZ, Greenberg BD, Petri S, Benjamin J, Muller CR, Hamer DH, Murphy DL. Association of anxiety-related traits with a polymorphism in the serotonin transporter gene regulatory region. Science. 1996;274:1527–1531. doi: 10.1126/science.274.5292.1527. [DOI] [PubMed] [Google Scholar]
  36. Long JC, Kittles RA. Human genetic diversity and the nonexistence of biological races. Human biology. 2009;81(5/6):777–798. doi: 10.3378/027.081.0621. [DOI] [PubMed] [Google Scholar]
  37. Lundborg P. The health returns to schooling-what can we learn from twins? Journal of Population Economics. 2013;26:673–701. [Google Scholar]
  38. Macgregor S, Cornes BK, Martin NG, Visscher PM. Bias, precision and heritability of self-reported and clinically measured height in Australian twins. Human genetics. 2006;120(4):571–580. doi: 10.1007/s00439-006-0240-z. [DOI] [PubMed] [Google Scholar]
  39. Madsen M, Andersen AMN, Christensen K, Andersen PK, Osler M. Does Educational Status Impact Adult Mortality in Denmark? A Twin Approach. American Journal of Epidemiology. 2010;172:225–234. doi: 10.1093/aje/kwq072. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. McHugh RK, Hofmann SG, Asnaani A, Sawyer AT, Otto MW. The serotonin transporter gene and risk for alcohol dependence: A meta-analytic review. Drug and Alcohol Dependence. 2010;108:1–6. doi: 10.1016/j.drugalcdep.2009.11.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Miech RA, Shanahan MJ. Socioeconomic status and depression over the life course. Journal of Health and Social Behavior. 2000;41:162–176. [Google Scholar]
  42. Mirowsky J, Ross CE. Education and self-rated health - Cumulative advantage and its rising importance. Research on Aging. 2008;30:93–122. [Google Scholar]
  43. Mosing MA, Verweij KJG, Medland SE, Painter J, Gordon SD, Heath AC, Madden PA, Montgomery GW, Martin NG. A Genome-wide association study of self-rated health. Twin Res Hum Genet. 2010 Aug;13(4):398–403. doi: 10.1375/twin.13.4.398. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Neale MC, Maes HH. Methodology for genetics studies of twins and families. 6. Dordrecht, The Netherlands: Kluwer; 1996. [Google Scholar]
  45. Neale MC, Boker SM, Xie G, Maes HH. Mx: statistical modeling. Richmond, VA: Virginia Institute for Psychiatric and Behavioral Genetics, Virginia Commonwealth University, Department of Psychiatry; 2004. [Google Scholar]
  46. Palloni A, Milesi C, White RG, Turner A. Early childhood health, reproduction of economic inequalities and the persistence of health and mortality differentials. Social Science & Medicine. 2009;68:1574–1582. doi: 10.1016/j.socscimed.2009.02.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Price AL, Zaitlen NA, Reich D, Patterson N. New approaches to population stratification in genome-wide association studies. Nature Reviews Genetics. 2010;11:459–463. doi: 10.1038/nrg2813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Richardson Ken. Wising up to the heritability of intelligence. GeneWatch. 2011;24:15–8. [Google Scholar]
  49. Rietveld CA, Medland SE, Derringer J, Yang J, Esko T, Martin NW, McMahon G. GWAS of 126,559 individuals identifies genetic variants associated with educational attainment. science. 2013;340(6139):1467–1471. doi: 10.1126/science.1235488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Romeis JC, Scherrer JF, Xian H, Eisen SA, Bucholz K, Health AC, et al. Heritability of self-reported health status. Health Services Research. 2000;35:995–1010. [PMC free article] [PubMed] [Google Scholar]
  51. Ross CE, Mirowsky J. Refining the association between education and health: The effects of quantity, credential, and selectivity. Demography. 1999;36:445–460. [PubMed] [Google Scholar]
  52. Ross CE, Wu CL. The Links Between Education and Health. American Sociological Review. 1995;60:719–745. [Google Scholar]
  53. Sabol SZ, Nelson ML, Fisher C, Gunzerath L, Brody CL, Hu S, Sirota LA, Marcus SE, Greenberg BD, Lucas FR, Benjamin J, Murphy DL, Hamer DH. A genetic association for cigarette smoking behavior. Health Psychology. 1999;18(1):7–13. doi: 10.1037//0278-6133.18.1.7. [DOI] [PubMed] [Google Scholar]
  54. Schnittker J, Bacak V. The Increasing Predictive Validity of Self-Rated Health. PloS one. 2014;9(1):e84933. doi: 10.1371/journal.pone.0084933. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Silventoinen K, Krueger RF, Bouchard TJ, Kaprio J, McGue M. Heritability of body height and educational attainment in an international context: Comparison of adult twins in Minnesota and Finland. American Journal of Human Biology. 2004;16:544–555. doi: 10.1002/ajhb.20060. [DOI] [PubMed] [Google Scholar]
  56. Solovieff N, Cotsapas C, Lee PH, Purcell SM, Smoller JW. Pleiotropy in complex traits: challenges and strategies. Nature Reviews Genetics. 2013;14(7):483–495. doi: 10.1038/nrg3461. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Sullivan PF, Neal MC, Kendler KS. Genetic epidemiology of major depression: review and meta-analysis. American Journal of Psychiatry. 2000;157:1552–62. doi: 10.1176/appi.ajp.157.10.1552. [DOI] [PubMed] [Google Scholar]
  58. Thorleifsson G, Walters GB, Gudbjartsson DF, Steinthorsdottir V, Sulem P, Helgadottir A, Stefansson K. Genome-wide association yields new sequence variants at seven loci that associate with measures of obesity. Nature genetics. 2008;41(1):18–24. doi: 10.1038/ng.274. [DOI] [PubMed] [Google Scholar]
  59. Vermeiren APA, Bosma H, Gielen M, Lindsey PJ, Derom C, Vlietinck R, et al. Do genetic factors contribute to the relation between education and metabolic risk factors in young adults? A twin study. European Journal of Public Health. 2012;23:986–991. doi: 10.1093/eurpub/cks167. [DOI] [PubMed] [Google Scholar]
  60. Visscher PM, et al. Assumption-free estimation of heritability from genome-wide identity-by-descent sharing between full siblings. PLoS Genet. 2006 doi: 10.1371/journal.pgen.0020041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Webbink D, Martin NG, Visscher PM. Does education reduce the probability of being overweight? Journal of Health Economics. 2010;29:29–38. doi: 10.1016/j.jhealeco.2009.11.013. [DOI] [PubMed] [Google Scholar]
  62. Williams D, Collins C. U.S. Socioeconomic and racial differences in health: patterns and explanations. Annual Review of Sociology. 1995;21:349–386. [Google Scholar]
  63. Yang J, Benyamin B, McEvoy BP, Gordon S, Henders AK, Nyholt DR, Visscher PM. Common SNPs explain a large proportion of the heritability for human height. Nature genetics. 2010;42(7):565–569. doi: 10.1038/ng.608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Yang JA, Lee SH, Goddard ME, Visscher PM. GCTA: A Tool for Genome-wide Complex Trait Analysis. American Journal of Human Genetics. 2011;88:76–82. doi: 10.1016/j.ajhg.2010.11.011. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES