Abstract
Social scientists often fit regression models with a range of covariates to infer causal effects in observational, population-based data, but the resulting estimates may be biased by unknown, unmeasured, and poorly measured confounders. Adjusting for genetic confounding using polygenic indices (PGIs) has been forwarded as one way to reduce this bias. However, whether and how relationships of interest to social scientists change when adjusting for PGIs or genetic confounding more broadly remains poorly understood. The current study sheds light on this issue by evaluating associations between years of schooling and self-rated health, body mass index, and depressive symptoms before and after adjusting for genetic confounding using data from the 2006–2012 waves of the Health and Retirement Study (n = 11,614), a nationally representative study of older U.S. adults. We adjust for genetic confounding in two ways: first by controlling for PGIs, and second by using PolygENic Genetic confoUnding INference (PENGUIN), a method based on variance component estimation. We find that controlling for PGIs modestly attenuates associations between education and each measure of health, and PENGUIN attenuates estimates further. However, a significant protective relationship between education and health remains when adjusting for genetic confounding with either method. Adjusting for genetic confounding using available methods thus does not call into question the robust relationship between education and health, underscoring the fundamental role of social and behavioral factors in shaping educational health disparities. Our findings also illustrate the limitations of adjusting for genetic confounding with PGIs specifically. In an era where PGIs are now broadly available to social scientists in population-based datasets, we urge caution when using them as controls for genetic confounding.
Highlights
-
•
We examine relationships between education and three health outcomes before and after adjusting for genetic confounding.
-
•
We consider two methods: controlling for polygenic indices (PGIs) and PolygENic Genetic confoUnding INference (PENGUIN).
-
•
Adjusting for genetic confounding with either method significantly attenuates associations between education and health.
-
•
Neither method calls into question the significant, positive relationship between education and health.
-
•
Findings underscore the important role of social and behavioral factors in structuring educational health inequalities.
1. Introduction
One of the most robust observations in the social sciences is that those with more education live longer and healthier lives than those with less (Cutler & Lleras-Muney, 2008; Link & Phelan, 1995; Zajacova & Lawrence, 2018). However, the causal effect of education on health is often contested. The relationship between education and health may be spurious, as family background, childhood health, skills, personality, and more may impact both schooling outcomes and wellbeing. Though educational gradients in health persist when adjusting for known confounders (Conti et al., 2010; Cutler & Lleras-Muney, 2008; Link et al., 2008; Montez & Hayward, 2014; Schnittker, 2005; Zheng, 2017), most studies leave one potential confounder uncontrolled: a person's DNA.
DNA influences both biological and sociobehavioral outcomes, including those related to health and education (Ge et al., 2017; Polderman et al., 2015). Furthermore, some genetic variants are associated with both health and education (Boardman et al., 2015; Bulik-Sullivan et al., 2015; Okbay et al., 2022; Wedow et al., 2018). This phenomenon, wherein the same genetic variants are related to two seemingly distinct outcomes or phenotypes, is known as pleiotropy (Solovieff et al., 2013). Pleiotropic genetic effects may confound the education-health link by shaping individual traits, behaviors, and exposures that influence both education and health.
Evidence from twin studies suggest that genetic factors do contribute to the relationship between education and health. Twin studies, which account for genetic confounding by assuming genetic similarity within twin pairs, generally return smaller estimates of education's effect on health than studies of unrelated individuals (Amin et al., 2015; Behrman et al., 2011; Böckerman & Maczulskij, 2016; Fujiwara & Kawachi, 2009; Halpern-Manners et al., 2016, 2020; Lundborg, 2013; Madsen et al., 2010, 2014) (though, see (Lundborg, Lyttkens, and Nystedt, 2016)). Yet twin studies have important limitations with respect to representativeness, external validity, precision, and bias (Boardman & Fletcher, 2015; Gilman & Loucks, 2014).
The emergence of measured genetic data in recent years has created new opportunities to understand genetic effects on education and health, and to control for genetic confounding, in population-based samples. With molecular genetic data from large samples of individuals, researchers can now estimate the genetic effects of millions of variants from across the genome on education, health, or any other outcome or phenotype using genome-wide association studies (GWAS). GWAS results, often referred to as “summary statistics”, can then be summarized in measures of genetic propensity known as polygenic indices (PGIs), polygenic scores, polygenic risk scores, or genetic risk scores (Dudbridge, 2013). In brief, a PGI is a variable that reflects the combined estimated effects of genetic variants from across a person's genome on a given phenotype.
PGIs have been forwarded as tools to control for genetic confounding in population-based research (Braudt, 2018; Cesarini & Visscher, 2017; Conley et al., 2019; Freese, 2018). PGIs have several qualities that make them attractive for this purpose. By summarizing genetic effects from all across the genome, PGIs encompass the pleiotropic effects that may confound relationships of interest. In addition, because they are based on DNA, PGIs are fixed at birth and do not change over the life course (though their predictive power can vary depending on environmental context and age); their application in regression analysis is straightforward as single genetic variables; and they are now widely available in epidemiological and population-based datasets. Important factors may, however, limit their broad utility for social science research. Most current PGIs rely on GWAS conducted in highly select datasets that have been restricted to European-ancestry respondents, raising concerns about generalizability (Martin et al., 2017; Schoeler et al., 2023). This is not merely an empirical issue; it is first and foremost an ethical one, with the potential to create inequities in future clinical and public health applications (Martin et al., 2019; Mostafavi et al., 2020). The predictive properties of PGIs also range widely, with even the most powerful PGIs explaining substantially less variation in their corresponding phenotypes than would be expected based on heritability estimates. This is because PGIs are based on estimated genetic effects and therefore contain measurement error (Pingault et al., 2021; Uddin et al., 2022; Zhao et al., 2024). Finally, because PGIs reflect estimated genetic propensities for a particular phenotype, researchers must choose which PGIs to include as covariates in their analyses, and omitted PGIs may continue to confound effect estimates (Zhao et al., 2024).
As these issues have come to light, scientists have developed alternative methods to adjust for genetic confounding. Some forward structural equation modeling approaches (Pingault et al., 2021). Others are based on variance component estimation. One such method uses GWAS summary statistics to estimate the heritability of and correlation between genetic effects on two phenotypes of interest and adjust effect estimates accordingly (Zhao et al., 2024). This method, called PolygENic Genetic confoUnding INference (PENGUIN), is “conceptually equivalent to adjusting for all genetic variants” simultaneously (Zhao et al., 2024) (p2).
While the promises and pitfalls of genetic data for social science research have been elaborated previously (Boardman & Fletcher, 2021; Cesarini & Visscher, 2017; Conley, 2016; Freese, 2018; Pingault et al., 2022), research investigating whether and how adjusting for genetic confounding alters the relationship between education and health remains incompletely understood. To shed light on this issue, the current study examines the extent to which adjusting for genetic confounding by controlling for PGIs or using PENGUIN attenuates associations between education and health in a sample of older U.S. adults of European ancestry. Specifically, we estimate associations between education and three measures of health—self-rated health, body mass index (BMI), and depressive symptoms—before and after controlling for PGIs for related health outcomes and for schooling and before and after using PENGUIN. We also compare the attenuation observed when adjusting for genetic confounding to that obtained when controlling for potential social and behavioral confounders, including measures of family background, childhood health, and early learning difficulties. Like many social scientists, our interest is not in genetic effects, per se. Instead, we are interested in the extent to which the estimated effects of education on health are attenuated when adjusting for genetic factors. We conclude with a discussion of our results in the context of the ongoing discourse surrounding the effect of education on health and in terms of the current limitations and future potential of genetic data for the social sciences.
2. Background
2.1. The link between education and health
Education is associated with better mental health, fewer functional limitations, fewer chronic conditions, reduced mortality, and more (Cutler & Lleras-Muney, 2008; Hummer & Lariscy, 2011; Zajacova & Lawrence, 2018). However, the explanation for the educational gradient in health is the subject of much debate. Education may causally impact health through diverse and dynamic mechanisms (Link & Phelan, 1995). For example, education may promote economic wellbeing, enabling access to health-enhancing goods and services. It may impart the norms, knowledge, and ability to maintain a healthy lifestyle (Mirowsky & Ross, 2003). And, those with higher education may be embedded in structural positions that confer health advantages without purposive action (Freese & Lutfey, 2011).
Alternatively, the association between education and health could be spurious, such that socially advantaged, healthier, and more conscientious kids stay in school longer and are set up for better health starting in childhood. Consider that family background is associated with education (Song & Mare, 2017) and that those from advantaged households have better health in adulthood, on average (Montez & Hayward, 2014). Poor health in childhood is also associated with educational attainment (Jackson, 2009; Palloni, 2006) as well as with poor health later in life (Haas, 2007). Similarly, early cognitive performance, non-cognitive skills, and personality predict both improved educational outcomes (Farkas, 2003; Lleras, 2008) and better health (Conti et al., 2010; Hauser & Alberto, 2011).
Studies adjusting for these and other potential confounders find that an independent, positive association between education and health remains (Conti et al., 2010; Cutler & Lleras-Muney, 2008; Link et al., 2008; Montez & Hayward, 2014; Schnittker, 2005; Zheng, 2017). However, controlled measures may not adequately reflect confounding concepts. Frequently, a key confounder is entirely unmeasured and cannot be controlled. For example, few surveys contain measures of cognitive ability, non-cognitive skills, and personality, especially in adolescence. Even when datasets include relevant measures, they may be incomplete or inaccurate. For instance, parental education may be assessed but occupation and income omitted, or childhood health may be measured with a single subjective report. Similarly, when early abilities are measured, they may be based on test scores, grades, or teacher assessments, which may fail to capture relevant talents and behaviors. Confounding concepts may continue to bias causal estimates when measured with error, and researchers may be unaware of additional confounding factors.
Studies can reduce concerns about confounding by leveraging natural experiments, such as changes in compulsory schooling laws, to estimate the effect of education on health. However, natural experiments also have limitations, including with respect to external validity and representativeness. While studies show that additional schooling spurred by policy changes improved health in the U.S., these findings are based on people whose educational attainment was influenced by the policy of interest, a group that may be highly select (Fletcher, 2015; Lleras-Muney, 2005). In addition, such analyses are often underpowered and the resulting estimates imprecise. Alternative methods of reducing confounding bias are needed for use with observational, population-based data. Adjusting for genetic confounding using PGIs has been forwarded as one potential such method.
2.2. Genetic confounding and adjustment
In observational research, genetics may confound the effect of education on health. For genetic confounding to occur, one or more genetic variants must be associated with both education and health through a phenomenon known as pleiotropy (Solovieff et al., 2013). Prior research shows that there are pleiotropic genetic effects on education and various health-related outcomes, such as general physical health, cardiometabolic and lung diseases, BMI, tobacco use, depression, and more (Boardman et al., 2015; Bulik-Sullivan et al., 2015; Okbay et al., 2022; Wedow et al., 2018).
There are two main forms of pleiotropy: biological and mediated (for a useful discussion, see Boardman et al., 2015; Wedow et al., 2018). In the case of education and health, biological pleiotropy would involve the same set of genetic variants independently influencing both education and health, perhaps indirectly via early environments, health, abilities, and traits, for example. Without the appropriate research design or statistical adjustment, biological pleiotropy could confound the estimated effect of education on health. In contrast, mediated pleiotropy through education would not confound the effect of interest. Mediated pleiotropy occurs when genetic effects on the outcome of interest (here, health) are mediated by the exposure of interest (here, education). Pleiotropic effects linking education and health may occur through some combination of biological and mediated processes. Even when a strong causal effect is suspected—such as that of education on health—biological pleiotropy could bias estimates.
For many decades, researchers concerned about genetic confounding relied heavily on studies of twins (Neale & Maes, 1996). By comparing the health outcomes of genetically identical (monozygotic) twins who were raised together but attained different levels of schooling, studies effectively hold genetic factors, as well as shared early environmental characteristics, constant (McGue et al., 2010). Twin studies examining the effect of education on health typically return smaller effect estimates than cross-sectional studies and research in unrelated samples, suggesting that genetic factors bias effect estimates upwards (Amin et al., 2015; Behrman et al., 2011; Böckerman & Maczulskij, 2016; Fujiwara & Kawachi, 2009; Halpern-Manners et al., 2016, 2020; Lundborg, 2013; Madsen et al., 2010, 2014) (though, see Lundborg et al., 2016). Several issues complicate the interpretation of twin study findings, however (Boardman & Fletcher, 2015; Gilman & Loucks, 2014). Because schooling is not randomized within twin pairs, characteristics and exposures that are unrelated to genotype, or shared environment, could continue to induce bias. Such confounders may be particularly difficult to adjust within twin pairs due to spillover effects, whereby twins seek to differentiate themselves from, or conversely, conform to, each other. Relatedly, while effect estimates are based on twins experiencing discordant exposures, it is uncommon for twins to attain different levels of schooling and twin pairs with discordant education may not be representative of all twin pairs or the broader population. Finally, as in all fixed-effect analyses, twin studies have relatively low power to detect significant effects.
Among unrelated persons, molecular genetic data is required to discern genetic effects and to adjust for genetic confounding. The collection of molecular genetic data has accelerated rapidly in recent years, and in response, genome-wide association studies (GWAS) have proliferated (McCarthy et al., 2008; Visscher et al., 2017). GWAS is a data-driven, hypothesis-free approach to the discovery of genetic associations. To conduct a GWAS, scientists estimate the effects of millions of genetic variants on an outcome, or phenotype, of interest using molecular genetic data from a very large sample of people. The effects of individual genetic variants on social and behavioral phenotypes are typically extremely small (Chabris et al., 2015). For this reason, scientists developed PGIs, which aggregate or sum the genetic effects estimated through GWAS to reflect a person's genetic propensity for a phenotype (Dudbridge, 2013). GWAS can be conducted, and PGIs constructed, for virtually any measured phenotype or outcome of interest for which there are large enough sample sizes. The current study draws on PGIs for the health outcomes studied, including self-rated health, BMI, and depressive symptoms, as well as a PGI for educational attainment.
A few important points on the interpretation of PGIs warrant mention. First, PGIs reflect genetic propensities present in a specific time, place, and population, as genetic effects are socially and environmentally contingent. (Boardman & Fletcher, 2021; Herd et al., 2019; Jencks, 1980). Consider, for example, a society experiencing long-term famine, where diets are restricted. Genetic effects on BMI might be suppressed. Relatedly, in a society that forbids women to attend school, having two X-chromosomes would completely determine the educational trajectories for half of the population. However, this effect would exist not for biological reasons, but rather because of the social environment. Second, PGIs comprise genetic effects that unfurl through some combination of biological, social, and behavioral mechanisms. For example, research shows that the relationship between a PGI for BMI and observed adiposity is mediated in part by behavioral factors (e.g., emotional eating, physical activity), personality (e.g., conscientiousness), education, and depressive symptoms (Herle et al., 2020; Stephan et al., 2020). Similarly, the relationship between a PGI for education and years of schooling is explained in part by adolescent cognitive performance (e.g., test scores), non-cognitive skills and personality traits (e.g., self-control, sociability, and openness to experience), and family SES (Belsky et al., 2016, 2018; Domingue et al., 2015; Okbay et al., 2016).
PGIs have been forwarded as potentially useful control variables for several reasons (Braudt, 2018; Cesarini & Visscher, 2017; Conley et al., 2019; Freese, 2018). Because they are based on DNA, PGIs are stable from birth and therefore causally prior to virtually all processes occurring throughout the life course, regardless of when genotyping occurs. Controlling for PGIs in regression analyses is also straightforward, and PGI variables are now widely available in epidemiological and population-based datasets. Most importantly, PGIs summarize genetic effects from all across the genome, and therefore encompass the pleiotropic genetic effects that may confound relationships of interest. One result of this is that a given PGI may be correlated not only with its corresponding phenotype, but also with numerous other traits, behaviors, and exposures, including potential confounders. Researchers could, in theory, eliminate bias by controlling for these more proximal confounders directly. In practice, however, confounding concepts are frequently poorly measured, if measured at all, in the data at hand. Adjusting for PGIs could therefore be a useful alternative or additional strategy to reduce bias attributable to confounders, both known and unknown.
Other factors may limit the utility of PGIs as control variables, however. PGIs rely on results from GWAS, which often draw on highly select samples, such as the UK Biobank (Bycroft et al., 2018) or the genetic testing company 23andMe. Genetic effects from GWAS of highly select samples may not be generalizable to the broader population (Mostafavi et al., 2020; Schoeler et al., 2023). Moreover, GWAS are often restricted to European-ancestry respondents, and resulting PGIs perform comparatively poorly in ethnoracially diverse data (Martin et al., 2017). In addition to issues stemming from non-representative GWAS, analyses seeking to adjust for genetic confounding by controlling for PGIs must consider the impact of both measurement error and model (mis)specification (Pingault et al., 2021; Uddin et al., 2022; Zhao et al., 2024). Because PGIs draw on estimated genetic effects, they measure genetic propensities imperfectly (measurement error). And, because PGIs reflect estimated genetic propensities for a particular phenotype, researchers must select which PGIs to adjust for in their analyses, and omitted PGIs may continue to bias estimates of interest (model misspecification). This is because genetic effects on a particular phenotype may evolve through multiple biological, social, and behavioral mechanisms. By aggregating genetic effects operating through these diverse mechanisms into a single continuous score, PGIs may obscure the specific mechanisms that confound the relationship of interest.
Alternative methods of adjusting for genetic confounding have recently been developed to improve upon the PGI approach, especially with respect to measurement error and model misspecification. One such method, known as PolygENic Genetic confoUnding INference (PENGUIN), relies on variance component estimation. PENGUIN is “conceptually equivalent to adjusting for all genetic variants” simultaneously and has been shown to remove considerably more confounding due to genetic factors than PGIs (Zhao et al., 2024) (p2). PENGUIN uses individual-level data on the exposure and outcome of interest in combination with GWAS summary statistics on the exposure and outcome variables to calculate the variance of the exposure, the covariance between the exposure and outcome, the heritability of the exposure, and the genetic covariance between the exposure and outcome. Using these variance components, PENGUIN adjusts the estimated effect of the exposure on the outcome for genetic confounding.
3. Study aims
In the current study, we evaluate the impact of adjusting for genetic confounding using currently available methods on the relationship between education and health in a population-based sample of older U.S. adults. Specifically, we examine associations between years of schooling and three measures of health—self-rated health, BMI, and depressive symptoms—before and after adjusting for PGIs and using PENGUIN. We also compare the attenuation observed with adjustment for genetic confounding to that observed when adjusting for known or suspected social and behavioral confounders of the education-health link. We are interested in answering the following 3 primary research questions:
-
1)
Does controlling for genetic factors significantly attenuate the association between education and health, even when adjusting for more proximal potential confounders, including demographic characteristics, family background, childhood health, and early learning difficulties?
-
2)
Does PENGUIN show greater attenuation when controlling for genetic factors than when controlling for PGIs, as the authors of PENGUIN hypothesize? If so, how do these results help us understand the limitations of using PGIs to control for genetic confounding?
-
3)
Does controlling for genetic confounding call into question the robust relationship between education and health?
4. Methods
4.1. Data
We draw individual-level data from the Health and Retirement Study (HRS). The HRS is a biennial panel survey that is funded by the National Institute on Aging (U01AG009740) and conducted by the University of Michigan (Health and Retirement Study 2022; RAND, 2022; Sonnega et al., 2014). The HRS began in 1992 with a sample of adults born between 1931 and 1941. Additional birth cohorts have since been surveyed, and more than 40,000 people have participated to date. Since 1998, the sample has been nationally representative of the U.S. population over age 50 and their spouses. Using a dataset that contains individuals above age 50 is ideal for our research questions, since most individuals have completed their education by this age, ensuring we are more accurately capturing the full effect of education, as well as any lagged effects of education that develop over time, on health. Further, we might expect more variation in health status among older individuals compared to younger individuals.
Respondents provided saliva samples for genotyping during enhanced face-to-face interviews conducted between 2006 and 2012. Nursing home residents and those responding through a proxy were ineligible for enhanced face-to-face interviews and therefore were not genotyped. A random one-half of eligible households was asked to provide saliva in 2006, and remaining households were invited in 2008. This process was repeated with new and ungenotyped respondents in 2010 and 2012. Genotyping was performed by the Center for Inherited Disease Research (Crimmins et al., 2013, 2015; Health and Retirement Study, 2021). PGIs for a range of phenotypes have since been constructed for 12,090 genotyped HRS respondents (Becker et al., 2021). PGIs are available only for respondents of European ancestries who self-reported their ethnoracial identity as non-Hispanic White. The analytic sample used for the current study is therefore also limited to non-Hispanic White respondents of European descent.
We began with data from the wave in which genotyping occurred (2006, 2008, 2010, or 2012) for the 12,090 respondents with PGIs to reduce concerns about selective mortality and attrition in the years after genotyping. We then restricted the sample to those ages 50 and older at the time of genotyping, dropping the 413 respondents below age 50. We also excluded 63 respondents who were missing information on the health outcomes of interest or who were missing education data. We multiply imputed all other variables across 20 chained imputations. Our analytic sample thus includes 11,614 respondents.
For analyses using PENGUIN, we draw GWAS summary statistics for educational attainment, self-rated health, BMI, and depression from Dr. Benjamin Neale's lab (Neale Lab, 2018). The Neale lab conducted GWAS among 361,194 European-ancestry individuals in the UK Biobank, a large-scale biomedical database and research resource containing genetic, health, social, and behavioral data from about half a million participants in the United Kingdom (Bycroft et al., 2018). We use the Neale lab summary statistics because they are freely available and because all GWAS were performed using the same methods.
4.2. Measures
4.2.1. Health outcomes
We draw on three measures of health that are commonly studied by social scientists, public health scholars, and epidemiologists, and all of which have a corresponding PGI available for use with the HRS data. They include self-rated health, BMI, and depressive symptoms. All three measures are expressed as continuous variables with higher values reflecting poorer health.
Self-rated health is from a survey question asked every year of all respondents: “Would you say your health is excellent, very good, good, fair, or poor?” Answers are recorded on a scale from 1 (Excellent) to 5 (Poor).
BMI is defined as (Centers for Disease Control and Prevention, 2022). We use measured height and weight from enhanced face-to-face interviews to calculate BMI where available. For 653 or 5.6% of respondents, we rely instead on self-reported height and weight.
Depressive symptoms are assessed with a shortened version of the Center for Epidemiologic Studies Depression Scale (CES-D) (Radloff, 1977). At each wave, respondents were asked whether they had experienced the following eight symptoms “much of the time” or more over the past week: (1) Felt depressed; (2) Felt everything was an effort; (3) Sleep was restless; (4) Felt happy (reverse coded); (5) Felt lonely; (6) Enjoyed life (reverse coded); (7) Felt sad; and (8) Couldn't get going. Depressive symptoms is an index reflecting the number of symptoms the respondent reported, ranging from 0 to 8. Our analyses include only those respondents who answered six or more items.
4.2.2. Education
We measure education with years of completed schooling. This variable is top-coded by the HRS at 17 years or more, and we collapse respondents reporting eight or fewer years of schooling into a single group (4.2%). The variable used here therefore ranges from 8 to 17.
4.2.3. PGIs and other genetic covariates
We measure genetic propensity for self-rated health, BMI, depressive symptoms, and educational attainment using PGIs in the HRS’ Polygenic Index Repository. The percentage of variation in the corresponding phenotype explained by each of these PGIs above and beyond basic demographic variables (i.e., their incremental predictive power or R2) ranges from 1.6% for depressive symptoms and 3.0% for self-rated health to 10.1% for educational attainment and 12.7% for BMI in the HRS (Becker et al., 2021; Benjamin et al., 2021). These R2 values are considered reasonably well-powered to well-powered in the literature (Plomin & Sophie von Stumm, 2022). Methodological details regarding PGI construction can be found in the Polygenic Index Repository's documentation (Becker et al., 2021; Benjamin et al., 2021). Polygenic indices were constructed using the software LDpred (Vilhjálmsson et al., 2015), applied to HapMap3 SNPs. Below, we provide a brief overview of polygenic score construction.
In short, PGIs aggregate or sum the effects on a phenotype of hundreds of thousands or even millions of single-nucleotide polymorphisms (SNPs), the type of genetic variant that is responsible for most genetic variation among humans (Dudbridge, 2013). At each SNP, a person possesses two alleles, and a person's genotype at each SNP is the number of risk or reference alleles (0, 1, or 2) found there. Their PGI is the weighted sum of their genotypes, where the weight of a particular SNP is the effect of an additional risk allele at that SNP on the outcome of interest. Equation (1) provides a standard formula for individual i's PGI,
| Equation 1 |
where is the genotype of individual i at SNP j and indicates the weight or effect size of an additional risk allele at SNP j, as estimated in an independent GWAS.
To ease interpretation of results, we standardize each PGI to have a mean of 0 and standard deviation of 1 across the analytic sample. We also re-scale the self-rated health PGI to reflect propensity for poorer self-rated health by multiplying the original standardized self-rated health PGI by −1. All health-related PGIs therefore reflect propensity for poorer outcomes, corresponding to the phenotypic measures.
In addition to the PGIs themselves, the HRS’ Polygenic Index Repository contains variables that reflect genetic ancestry. More precisely, the Repository contains variables for the top principal components (PCs) of the genetic data, or the largest dimensions of genetic variation in the sample. Statistical geneticists encourage researchers to include PCs as covariates in analyses with PGIs due to concerns that subtle ancestral variation could confound genetic effects. We therefore use the top 10 PCs as covariates in our analyses.
4.2.4. Demographic and sociobehavioral covariates
We draw on several measures known or suspected to confound the relationship between education and health as covariates, including basic demographic characteristics and measures of family background, childhood health, and early learning difficulties or academic challenges. Demographic covariates include linear and squared terms for age, year of birth, and sex/gender (female versus male). We do not control for ethnoracial identification as our analytic sample includes only respondents of European ancestries identifying as non-Hispanic White.
Measures of family background include mother's and father's years of schooling, father's occupation and employment history, perceived SES in childhood, region of birth, and rural residence in childhood. Mother's and father's years of schooling are both continuous variables ranging from 8 or fewer to 17 or more. Father's occupation is expressed in five categories: upper white collar (professional, technical, and managerial occupations); lower white collar (sales and clerical occupations or administrative support); blue collar (mechanics, construction, production, and service/labor-related occupations); farming and related occupations (e.g., fishing, forestry); and not applicable, which includes those whose father did not work or was absent or deceased. We also draw on a dichotomous variable indicating that the respondent's father was unemployed for a period of several months or more while the respondent was growing up. Perceived SES in childhood is a dichotomous variable indicating that the respondent's family was poor growing up versus about average or pretty well-off. Region of birth is coded in five categories: Northeast, Midwest, South, West, or born outside the U.S. Rural residence in childhood is a binary variable.
We measure childhood health first with a five-category variable from a question asking respondents to reflect on their health as children and report it as excellent, very good, good, fair, or poor. We also use a binary variable indicating that the respondent was “disabled for six months or more because of a health problem, that is, unable to do the usual activities of classmates or other children [their] age”. In addition, we draw on a binary measure indicating that the respondent smoked cigarettes regularly in their youth.
Finally, we draw on two binary measures that are indicative of early learning difficulties. The first is from a question asking, “In grade school or high school, did you ever have a problem in learning the usual lessons, such that you regularly attended special classes, received special training sessions, or had to attend a different school for more than six months?” The second is from a question asking, “Before you were 18 years old, did you have to do a year of school over again?”
4.3. Analysis
We begin by calculating descriptive statistics for the sample. We then estimate a series of ordinary least squares (OLS) linear regression models examining the association between years of schooling and each of our three health outcomes—self-rated health, BMI, and depressive symptoms—before and after adjusting for genetic confounding and social and behavioral covariates. We estimate three series of models: (a) models that do not adjust for genetic confounding; (b) models that adjust for genetic confounding by controlling for PGIs; and (c) models that adjust for genetic confounding using PENGUIN. Models that do not control for genetic confounding and models that control for PGIs employ HRS-constructed clustering variables to account for the complex sampling design of the HRS.
For model series (a), (b), and (c), we incorporate demographic and sociobehavioral covariates across three nested models. Thus, for each health outcome, we estimate nine models in total. Model 1A estimates the association of years of schooling (SchYrs) with the health outcome of interest (Health) when controlling for demographic characteristics, as indicated by the vector Dem. Model 1B builds on Model 1A by adding as covariates the PGI corresponding to the health outcome of interest (HealthPGI), the educational attainment PGI (EAPGI), and the top 10 PCs of the genetic data (PC). Model 1C builds on Model 1A instead by using PENGUIN to adjust for genetic confounding; details on model estimation with PENGUIN are provided below. Models 2A, 2B, and 2C parallel Models 1A, 1B, and 1C adding as covariates measures of family background, indicated by the vector Fam. Finally, Models 3A, 3B, and 3C incorporate covariates reflecting health and learning difficulties in childhood, indicated by the vector ChHlthLrn.
| Model 1A |
| Model 1B |
| Model 2A |
| Model 2B |
| Model 3A |
| Model 3B |
The analytic framework and methodological details of PENGUIN have been covered elsewhere (Zhao et al., 2024). In brief, PENGUIN uses individual-level data on the exposure and outcome variables of interest—here, educational attainment and either self-rated health, BMI, or depressive symptoms from HRS participants—in combination with GWAS summary statistics for the exposure and outcome variables to adjust effects for genetic confounding. With PENGUIN, the coefficient on years of schooling in Models 1C, 2C, and 3C is estimated with Equation (2) as
| Equation 2 |
where is the estimated covariance between years of schooling and the health outcome of interest in the HRS sample; is the estimated genetic covariance between educational attainment and the health outcome of interest based on GWAS summary statistics; is the estimated variance of the exposure from the HRS sample; and is the estimated heritability of educational attainment. The GitHub page where PENGUIN can be downloaded, along with code and a wiki for running the software and preparing data can be found at: https://github.com/qlu-lab/PENGUIN.
We are interested primarily in the magnitude and statistical significance of the coefficient on SchYrs and the extent to which it changes across models. In particular, we are interested in the proportional change or attenuation in the SchYrs coefficient when adjusting for genetic confounding with PGIs or PENGUIN. That is, we are interested in comparing estimates from Models 1B (PGIs) and 1C (PENGUIN) to estimates from Model 1A; Models 2B and 2C to Model 2A; and Models 3B and 3C to Model 3A. We calculate proportional attenuation using Equation (3). For the PGI analyses, we use Clogg tests to determine whether effect size estimates differ significantly across nested models (Clogg et al., 1995), reporting the average Clogg test p-values across the 20 imputed datasets. We are unable to conduct Clogg tests comparing results from PENGUIN.
| Equation 3 |
Note that we incorporate sets of demographic and sociobehavioral covariates across three models rather than all at once as the attenuation observed when adjusting for genetic confounding may decline when more proximal confounders are controlled directly. By estimating a series of models, we can investigate whether the attenuating impact of adjusting for genetic confounding changes when sociobehavioral covariates are also controlled.
To benchmark the level of attenuation observed when adjusting for genetic confounding, we also calculate the proportional attenuation in the coefficient on SchYrs between Models 1A and 3A—that is, when all sociobehavioral covariates are added to the baseline model. Finally, we calculate the extent to which the coefficient on SchYrs is attenuated between Models 1A and 3B and 3C—that is, when adjusting both for sociobehavioral covariates and genetic confounding.
5. Results
Table 1 presents descriptive statistics for all measures. Respondents range in age from 50 to 101, with an average of 67.4 years (SD = 10.6). Just over half are female (56.1%). On average, respondents completed 13.3 years of schooling (SD = 2.4). The average self-rated health is 2.7 (SD = 1.1) on a scale from 1 (Excellent) to 5 (Poor); average BMI is 29.3 (SD = 6.2), which is considered overweight (Centers for Disease Control and Prevention, 2022); and respondents report an average of 1.3 (SD = 1.9) depressive symptoms out of eight. All four PGIs are standardized (mean = 0, SD = 1) and range from a minimum of approximately −4 to a maximum of 4 within the sample.
Table 1.
Descriptive statistics.
| Mean (SD) or % | Min, Max | Non-Missing N | |
|---|---|---|---|
| Age in years | 67.4 (10.6) | 50, 101 | 11,614 |
| Year of birth | 1940.0 (11.4) | 1905, 1962 | 11,614 |
| Female | 56.1 | 11,614 | |
| Years of schooling | 13.3 (2.4) | 8, 17 | 11,614 |
| Self-rated healtha | 2.7 (1.1) | 1, 5 | 11,614 |
| BMI (kg/m2) | 29.3 (6.2) | 11.0, 76.5 | 11,614 |
| Depressive symptoms | 1.3 (1.9) | 0, 8 | 11,614 |
| PGI for self-rated health | 0.0 (1.0) | −4.1, 4.2 | 11,614 |
| PGI for BMI | 0.0 (1.0) | −4.2, 3.7 | 11,614 |
| PGI for depressive symptoms | 0.0 (1.0) | −3.9, 3.5 | 11,614 |
| PGI for educational attainment | 0.0 (1.0) | −3.7, 4.0 | 11,614 |
| Mother's years of schooling | 11.0 (2.5) | 8, 17 | 10,046 |
| Father's years of schooling | 10.9 (2.9) | 8, 17 | 9,711 |
| Father's occupation | 9,768 | ||
| Upper white collar | 17.5 | ||
| Lower white collar | 12.2 | ||
| Blue collar | 47.8 | ||
| Farming or forestry | 15.6 | ||
| Father did not work, absent, deceased | 6.9 | ||
| Father was unemployed | 21.7 | 10,900 | |
| Family was poor | 26.0 | 11,453 | |
| Region of birth | 11,521 | ||
| Northeast | 22.8 | ||
| Midwest | 36.2 | ||
| South | 26.9 | ||
| West | 10.3 | ||
| Outside the U.S. | 3.8 | ||
| Rural residence in childhood | 45.3 | 11,293 | |
| Health in childhood | 11,605 | ||
| Excellent | 54.7 | ||
| Very good | 25.5 | ||
| Good | 14.5 | ||
| Fair | 4.2 | ||
| Poor | 1.2 | ||
| Disabled due to poor health in childhood | 4.2 | 11,202 | |
| Smoked in childhood | 20.7 | 11,202 | |
| Had learning difficulties in childhood | 2.9 | 11,202 | |
| Repeated a grade in school | 14.0 | 10,738 |
Note. N = 11,614. Data are from the Health and Retirement Study. BMI = Body mass index; PGI = Polygenic index.
Self-rated health is coded from 1 (Excellent) to 5 (Poor).
Table 2 presents results from models of self-rated health, BMI, and depressive symptoms on years of schooling and select sets of covariates. Perhaps most importantly, additional schooling is significantly associated with better health across all nine models for all three health outcomes—even when accounting for genetic confounding with PGIs or PENGUIN. When accounting for demographic characteristics, family background, childhood health and learning difficulties, and PGIs in Model 3B, a year of schooling is associated with better self-rated health (−0.063, p < .001), lower BMI (−0.114, p < .001), and fewer depressive symptoms (−0.081, p < .001). Similarly, when using PENGUIN in Model 3C, a year of schooling is associated with better self-rated health (−0.053, p < .001), lower BMI (−0.099, p < .001), and fewer depressive symptoms (−0.077, p < .001).
Table 2.
Coefficients and standard errors from models of self-rated health, BMI, and depressive symptoms on years of schooling.
| Self-rated health |
BMI |
Depressive symptoms |
|||||||
|---|---|---|---|---|---|---|---|---|---|
| A: No adjustment for genetic confounding | B: Control for PGIs | C: Use PENGUIN | A: No adjustment for genetic confounding | B: Control for PGIs | C: Use PENGUIN | A: No adjustment for genetic confounding | B: Control for PGIs | C: Use PENGUIN | |
| Model 1: Controlling for demographics | |||||||||
| Years of schooling | −0.108 ∗∗∗ | −0.092 ∗∗∗ | −0.075 ∗∗∗ | −0.236 ∗∗∗ | −0.155 ∗∗∗ | −0.132 ∗∗∗ | −0.126 ∗∗∗ | −0.111 ∗∗∗ | −0.096 ∗∗∗ |
| (0.004) | (0.005) | (0.013) | (0.028) | (0.029) | (0.014) | (0.007) | (0.008) | (0.012) | |
| PGI for SRH, BMI, or DS | – | 0.196 ∗∗∗ | – | – | 2.116 ∗∗∗ | – | – | 0.238 ∗∗∗ | – |
| (0.009) | (0.056) | (0.014) | |||||||
| PGI for EA | – | 0.005 | – | – | 0.249 ∗∗∗ | – | – | −0.044 ∗ | – |
| (0.010) | (0.057) | (0.017) | |||||||
| Attenuation of the “Years of schooling” coefficient compared to Model 1A | – | 14.81 % | 30.56 % | – | 34.32 % | 44.07 % | – | 11.90 % | 23.81 % |
| Model 2: Controlling for demographics and family background | |||||||||
| Years of schooling | −0.087 ∗∗∗ | −0.077 ∗∗∗ | −0.064 ∗∗∗ | −0.161 ∗∗∗ | −0.109 ∗∗∗ | −0.094 ∗∗∗ | −0.111 ∗∗∗ | −0.102 ∗∗∗ | −0.092 ∗∗∗ |
| (0.005) | (0.005) | (0.012) | (0.030) | (0.030) | (0.013) | (0.007) | (0.008) | (0.012) | |
| PGI for SRH, BMI, or DS | – | 0.184 ∗∗∗ | – | – | 2.104 ∗∗∗ | – | – | 0.226 ∗∗∗ | – |
| (0.009) | (0.056) | (0.014) | |||||||
| PGI for EA | – | 0.009 | – | – | 0.279 ∗∗∗ | – | – | −0.036 ∗ | – |
| (0.010) | (0.055) | (0.017) | |||||||
| Attenuation of the “Years of schooling” coefficient compared to Model 2A | – | 11.49 % | 26.44 % | – | 32.30 % | 41.61 % | – | 8.11 % | 17.12 % |
| Model 3: Controlling for demographics, family background, and child health and learning difficulties | |||||||||
| Years of schooling | −0.071 ∗∗∗ | −0.063 ∗∗∗ | −0.053 ∗∗∗ | −0.158 ∗∗∗ | −0.114 ∗∗∗ | −0.099 ∗∗∗ | −0.088 ∗∗∗ | −0.081 ∗∗∗ | −0.077 ∗∗∗ |
| (0.006) | (0.006) | (0.013) | (0.031) | (0.030) | (0.013) | (0.008) | (0.009) | (0.014) | |
| PGI for SRH, BMI, or DS | – | 0.171 ∗∗∗ | – | – | 2.105 ∗∗∗ | – | – | 0.211 ∗∗∗ | – |
| (0.009) | (0.055) | (0.015) | |||||||
| PGI for EA | – | 0.008 | – | – | 0.272 ∗∗∗ | – | – | −0.031 | – |
| (0.010) | (0.055) | (0.018) | |||||||
| Attenuation of the “Years of schooling” coefficient compared to Model 3A | = | 11.27 % | 25.35 % | – | 27.85 % | 37.34 % | – | 7.95 % | 12.50 % |
Note. N = 11,614. Data are from the Health and Retirement Study. BMI = Body mass index; DS = Depressive symptoms; EA = Educational attainment; OLS=Ordinary Least Squares; PENGUIN = PolygENic Genetic confoUnding INference; PGI = Polygenic index; SRH = Self-rated health. Standard errors in parentheses. ∗∗∗p < .001; ∗∗p < .01; ∗p < .05 (two-tailed tests).
That said, the estimated effect of education on all three health outcomes declines in magnitude—that is, attenuates—when adjusting for genetic confounding with either PGIs or PENGUIN. This pattern is evident in Fig. 1, which plots the estimated effect of a year of schooling on each health outcome across models.
Fig. 1.
Estimated effects of a year of schooling on self-rated health, BMI, and depressive symptoms.
Note. N=11,614. Data are from the Health and Retirement Study. Coefficients and 95 % confidence intervals shown; additional results provided in Table 2. Models 1A, 2A, and 3A do not adjust for genetic confounding. Models 1B, 2B, and 3B adjust for genetic confounding by controlling for relevant PGIs and the top 10 principal components of the genetic data. Models 1C, 2C, and 3C adjust for genetic confounding using the PENGUIN method. Model 1 controls for demographics. Model 2 controls for demographics and family background. Model 3 controls for demographics, family background, and child health and learning difficulties. BMI = Body mass index; OLS = Ordinary least squares; PENGUIN = PolygENic Genetic confoUnding INference; PGI = Polygenic index.
Fig. 2 plots the proportional attenuation of the estimated effect of a year of schooling on each health outcome when adjusting for genetic confounding with PGIs or PENGUIN. Across all health outcomes, the proportional attenuation observed is much larger with PENGUIN than with the PGI approach.
Fig. 2.
Attenuation of the estimated effects of a year of schooling on self-rated health, BMI, and depressive symptoms when adjusting for genetic confounding with PGIs or PENGUIN.
Note. N=11,614. Data are from the Health and Retirement Study. Percent attenuation estimates in comparison to Models 1A, 2A, and 3A shown; additional results provided in Table 2. Models 1A, 2A, and 3A do not adjust for genetic confounding. Models 1B, 2B, and 3B adjust for genetic confounding by controlling for relevant PGIs and the top 10 principal components of the genetic data. Models 1C, 2C, and 3C adjust for genetic confounding using the PENGUIN method. Model 1 controls for demographics. Model 2 controls for demographics and family background. Model 3 controls for demographics, family background, and child health and learning difficulties. BMI = Body mass index; OLS = Ordinary least squares; PENGUIN = PolygENic Genetic confoUnding INference; PGI = Polygenic index.
Consider first results for self-rated health. In Model 1A, which controls for basic demographic characteristics only, a year of schooling is associated with a 0.108-unit (p < .001) improvement in self-rated health. When controlling for PGIs in Model 1B, the estimated effect of a year of schooling is reduced to −0.092 (p < .001), reflecting a proportional attenuation of 14.81% (Clogg p < .001). When using PENGUIN in Model 1C, the estimated coefficient on years of schooling is attenuated even further, to −0.075 (p < .001), reflecting a proportional attenuation of 30.56% relative to Model 1A. In Model 2A, which includes demographic covariates as well as measures of family background, a year of schooling is associated with a 0.087-unit (p < .001) improvement in self-rated health. Incorporating PGIs in Model 2B attenuates this estimate by 11.49% (−0.077, p < .001; Clogg p < .001), while PENGUIN attenuates the estimate by a much larger 26.44% in Model 2C (−0.064, p < .001). Finally, in Model 3A, which incorporates all sociobehavioral covariates—including demographic characteristics, family background, and childhood health and learning difficulties—a year of schooling is associated with a 0.071-unit (p < .001) improvement in self-rated health. This effect is attenuated by 11.27% (−0.063, p < .001; Clogg p < .001) when controlling for PGIs in Model 3B, versus 25.35% (−0.053, p < .001) using PENGUIN in Model 3C.
The level of attenuation observed when adjusting for genetic confounding with either the PGI approach or PENGUIN is largest for BMI. When controlling for PGIs, the estimated effect of a year of schooling on BMI is attenuated by 34.32% in Model 1B (−0.155, p < .001) versus Model 1A (−0.236, p < .001); 32.30% in Model 2B (−0.109, p < .001) versus Model 2A (−0.161, p < .001); and 27.85% in Model 3B (−0.114, p < .001) versus Model 3A (−0.158, p < .001). All differences are statistically significant (Clogg p < .05). Using PENGUIN attenuates the estimated effect of schooling on BMI even further, by 44.07% in Model 1C (−0.132, p < .001) versus Model 1A; 41.61% in Model 2C (−0.094, p < .001) versus Model 1A; and 37.34% in Model 3C (−0.099, p < .001) versus Model 1A.
For depressive symptoms, attenuation estimates are smaller. When adjusting for PGIs, the estimated effect of schooling on depressive symptoms is reduced in magnitude by 11.90% in Model 1B (−0.111, p < .001) versus Model 1A (−0.126, p < .001); 8.11% in Model 2B (−0.102, p < .001) versus Model 2A (−0.111, p < .001); and 7.95% in Model 3B (−0.081, p < .001) versus Model 3A (−0.088, p < .001). Again, all differences are statistically significant (Clogg p < .05). With PENGUIN, the estimated effect of schooling on depressive symptoms is reduced by 23.81% in Model 1C (−0.096, p < .001) versus Model 1A; 17.12% in Model 2C (−0.092, p < .001) versus Model 2A; and 12.50% in Model 3C (−0.077, p < .001) versus Model 3A.
One additional pattern stands out in Fig. 2. For all three health outcomes, the proportional attenuation in the estimated effect of education following adjustment for genetic confounding with either PGIs or PENGUIN is smaller when sociobehavioral covariates are included than when only demographics are controlled. For example, education's estimated effect on BMI is reduced in magnitude by 34.32% between Models 1A and 1B, and by 44.07% between Models 1A and 1C, which control only for demographics. Meanwhile, the coefficient on education is attenuated by a more modest 27.85% between Models 3A and 3B and by 37.34% between Models 3A and 3C, which also control for measures of family background, childhood health, and early learning difficulties. Though differences in attenuation across models are modest, results provide suggestive evidence that adjustment for genetic confounding is most useful when proximal social and behavioral confounders of the education-health link are uncontrolled.
To benchmark the attenuating impact of adjustment for genetic confounding, Fig. 3 plots the proportional attenuation of the estimated effect of a year of schooling on each health outcome between the baseline model that controls only for demographics (Model 1A) and five other models: Model 1B, which adjusts for genetic confounding with PGIs; Model 1C, which instead uses PENGUIN; Model 3A, which does not adjust for genetic confounding but rather incorporates the full range of sociobehavioral covariates; and Models 3B and 3C, which adjust for genetic confounding with PGIs or PENGUIN, respectively, and control for all sociobehavioral covariates. For self-rated health and depressive symptoms, the attenuation observed when adjusting for genetic confounding with PGIs (Model 1B) is considerably smaller than that observed when adjusting for sociobehavioral covariates (Model 3A), while the attenuation observed when using PENGUIN (Model 1C) is similar to that observed with the sociobehavioral covariates. For BMI, the attenuation observed when adjusting for genetic confounding is similar to or greater than that observed when adjusting for sociobehavioral covariates only. Thus the impact of adjusting for genetic confounding on the estimated link between education and BMI is on par with measures of family background, childhood health, and early learning difficulties, which are currently among the most recognized confounders of the education-health link. For self-rated health, BMI, and depressive symptoms, the level of attenuation is greatest when PENGUIN is used and sociobehavioral covariates are added in Model 3C relative to Model 1A, reaching 50.93% for self-rated health, 58.05% for BMI, and 38.89% for depressive symptoms.
Fig. 3.
Attenuation of the estimated effects of a year of schooling on self-rated health, BMI, and depressive symptoms when adjusting for genetic confounding with PGIs or PENGUIN, sociobehavioral covariates, or both.
Note. N=11,614. Data are from the Health and Retirement Study. Percent attenuation estimates shown in comparison to Model 1A shown; additional results provided in Table 2. Models 1A and 3A do not adjust for genetic confounding. Models 1B and 3B adjust for genetic confounding by controlling for relevant PGIs and the top 10 principal components of the genetic data. Models 1C and 3C adjust for genetic confounding using the PENGUIN method. Model 1 controls for demographics. Model 3 controls for demographics, family background, and child health and learning difficulties. BMI = Body mass index; Covars = Covariates; OLS = Ordinary least squares; PENGUIN = PolygENic Genetic confoUnding INference; PGI = Polygenic index; SB = Sociobehavioral.
6. Discussion
While the potential uses of genetic data and methods for social science research have been elaborated previously (Boardman & Fletcher, 2021; Cesarini & Visscher, 2017; Conley, 2016; Freese, 2018; Pingault et al., 2022; Zhao et al., 2024), few studies evaluate how these data and methods impact relationships of interest to social scientists, including the relationship between education and health. This study therefore estimated associations between education and three measures of health—self-rated health, BMI, and depressive symptoms—when adjusting for genetic confounding in two ways. First, we adjusted for genetic confounding by controlling for PGIs reflecting estimated genetic propensities for schooling and the health outcome of interest. Second, we used a method designed to improve upon the PGI approach using variance component estimation, known as PolygENic Genetic confoUnding Inference or PENGUIN (Zhao et al., 2024). Our analyses reveal three key findings and a fourth suggestive pattern.
First and foremost, the relationship between education and health remains statistically significant when adjusting for genetic confounding using either method. The familiar educational gradient in health persists when adjusting for education-related and health-related PGIs, such that additional years of schooling are associated with better self-rated health, lower BMI, and fewer depressive symptoms. These relationships also persist when using PENGUIN to account for genetic confounding.
That said, our second finding is that the estimated relationship between education and health is attenuated when accounting for genetic confounding using either method. The amount of attenuation observed with the PGI approach is relatively modest for self-rated health and depressive symptoms, ranging from approximately 8%–15% across models, while the estimated effect of education on BMI is attenuated more substantially, by approximately one-third, with the inclusion of PGIs. Similarly, the amount of attenuation observed when adjusting for genetic confounding with PENGUIN was largest for BMI. Indeed, for BMI, the attenuating impact of adjusting for genetic confounding was on par with or greater than that of all available social and behavioral covariates—measures of family background, childhood health, and early learning difficulties—combined.
Third, the level of attenuation observed with adjustment for genetic confounding was more modest with the PGI approach than when using PENGUIN. For self-rated health and depressive symptoms, the proportional attenuation in the estimated effect of education observed with PENGUIN was double that observed when controlling for PGIs. For BMI, the difference was smaller but still substantial, with PENGUIN attenuating the estimated effect of education by about one-third more than the PGI approach. What this suggests is that the PGI approach is less effective at controlling for genetic confounding than PENGUIN. This should come as no surprise, as PENGUIN was specifically designed to deal with some of the shortcomings of the PGI approach, including issues related to measurement error and model (mis)specification. This finding underscores a very important caveat to the PGI approach: adjusting for a select set of PGIs is unlikely to eliminate all forms of genetic confounding operating through all possible biological, social, and behavioral mechanisms. Social scientist wishing to use PGIs as controls for genetic confounding should therefore carefully consider the limitations of this approach.
Our final suggestive result is that the extent of attenuation following adjustment for genetic confounding is greatest in models that omit potentially confounding social and behavioral covariates. For example, in models controlling only for demographic characteristics, the estimated effects of education on self-rated health, BMI, and depressive symptoms are reduced by 14.8%, 34.5%, and 11.6%, respectively, with the inclusion of PGIs, whereas when accounting demographics as well as available measures of family background, childhood health, and early learning difficulties, the inclusion of PGIs attenuates these estimates by just 10.9%, 27.6%, and 7.6%, respectively. Patterns with PENGUIN are similar. This supports the idea that adjusting for genetic confounding is most useful when more proximal social and behavioral confounders are incompletely controlled.
The findings of the current study contribute to decades of research and debate over the link between education and health. Education and health demonstrate a robust, positive association (Cutler & Lleras-Muney, 2008; Zajacova & Lawrence, 2018). However, the extent to which this reflects a causal effect has long been contested, as family background, childhood health, early abilities, traits, and environments, and a multitude of other factors could confound the effect of education on health. To reduce concerns about confounding, studies have leveraged natural experiments (Fletcher, 2015; Lleras-Muney, 2005) and twins with discordant education (Amin et al., 2015; Behrman et al., 2011; Böckerman & Maczulskij, 2016; Fujiwara & Kawachi, 2009; Halpern-Manners et al., 2016, 2020; Lundborg, 2013; Lundborg et al., 2016; Madsen et al., 2010, 2014) to evaluate the link between education and health. Many such studies continue to detect a link between education and health, but the estimated effect of education is often much smaller than that obtained in observational, population-based research. Yet caution must be exercised when interpreting the results of natural experiments and twin studies, given concerns about representativeness, external validity, precision, and bias (Boardman & Fletcher, 2015; Gilman & Loucks, 2014). Many social scientists opt instead to account for confounding using direct regression adjustment. They typically find that the education-health link persists (Conti et al., 2010; Cutler & Lleras-Muney, 2008; Link et al., 2008; Montez & Hayward, 2014; Schnittker, 2005; Zheng, 2017). However, potential confounders are notoriously difficult to measure and control directly and comprehensively, and it is possible that many remain known. The current study suggests that such studies could reduce confounding further by adjusting for genetic confounding.
We find that adjusting for genetic confounding with PGIs or PENGUIN significantly attenuates or reduces the magnitude of the familiar relationship between education and health, although even when genetic confounding is controlled, additional education is significantly associated with better self-rated health, lower BMI, and fewer depressive symptoms. These findings are in line with recent work showing that long-observed relationships between birth weight and later outcomes (e.g., education, BMI) remain statistically significant when relevant PGIs are controlled (Conley et al., 2019). Current methods of adjusting for genetic confounding thus appear unlikely to force the fundamental reconsideration of established relationships between social and behavioral outcomes.
Though this study focuses on the relationship between education and health, most social and behavioral outcomes are also highly pleiotropic (Chabris et al., 2015). Large GWAS have now been conducted for a wide variety of health, social, and behavioral phenotypes, enabling the construction of relevant PGIs and the use of PENGUIN and other methods to adjust for genetic confounding (Becker et al., 2021; Neale Lab, 2018). Opportunities for eliminating bias in population-based, observational research thus abound, and they are likely to expand even further in the coming years as the collection of genetic data and development of related technologies advance. The use of PGIs, PENGUIN, and related methods could be relevant even for scholars with little interest in genetic effects, per se, as novel strategies to adjust for confounding in observational, population-based research.
The limitations of this study reflect the weaknesses of current methods of adjusting for genetic confounding as well as some of the data limitations of large, biobank-scale data. First, because PGIs measure genetic propensities with error, their use as control variables can only partially account for genetic confounding. And, because PGIs reflect additive genetic propensities for a particular phenotype, multiple PGIs may be needed to adequately control for genetic confounding, and omitted PGIs may continue to bias estimates due to model misspecification. We therefore urge social scientists to carefully consider the limitations and assumptions of using PGIs as control variables.
Next, both the PGI approach and PENGUIN rely on GWAS summary statistics, and therefore are contingent on the genetic effects present in GWAS samples. Current GWAS samples draw heavily from medical case-control studies, opt-in biobanks, and genomic testing companies. Resulting genetic effect estimates may therefore be more representative of select population subgroups, and as such, may have variable predictive capacities across the population (Mostafavi et al., 2020). Even more problematic is that GWAS are often conducted in European-ancestry samples, and resulting genetic effect estimates may not generalize to diverse populations (Martin et al., 2017). This study also is restricted to European-ancestry HRS respondents who identify as non-Hispanic White, as PGIs are largely unavailable for diverse respondents. Results of the current study thus suggest an upper bound on the impact of adjusting for genetic confounding with PGIs or PENGUIN; all else constant, the magnitude of attenuation observed may be smaller in ancestrally diverse or non-European samples. This of course dramatically reduces the usefulness of these methods in population health research, where considerations over representation and generalizability are paramount, as well as in future clinical and public health applications, which must prioritize equitable impacts and outcomes (Martin et al., 2019).
Next, our analyses with PENGUIN rely on genome-wide association studies (GWAS) conducted in the UK Biobank. We use GWAS conducted in the UK Biobank because of their large sample sizes and consistent methodological approaches. However, we acknowledge that the social context of our U.S.-based sample (HRS data) is distinct from that of the U.K.-based GWAS sample, and that social context might influence GWAS results, even within ancestry groups. As a result, GWAS conducted in one context may misrepresent socially contingent genetic effects relevant in another, leading to the potential mischaracterization of genetic confounding when applying cross-nationally.
Finally, the current study was unable to control directly for all possible social and behavioral confounders of the education-health link due to the limited availability of relevant measures in the HRS. In a way, this weakness is also a strength, as it underscores the potential utility of adjusting for genetic confounding in observational, population-based studies that lack comprehensive data on the more proximal traits, behaviors, and exposures believed to confound estimates of interest. Indeed, our results suggest that adjustment for genetic confounding attenuates the estimated effects of education on three distinct health outcomes, particularly when more proximal social and behavioral confounders are not controlled directly. Still, at this time, adjustment for genetic confounding appears unlikely to force a major reconsideration of the protective relationship between education and health. The relationship between education and health is positive, pervasive across diverse health outcomes, and robust to adjustment for genetic confounding using PGIs or alternative methods.
CRediT authorship contribution statement
Meghan Zacher: Writing – review & editing, Writing – original draft, Visualization, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Robbee Wedow: Writing – review & editing, Writing – original draft, Supervision, Methodology, Formal analysis.
Declarations
Robbee Wedow is a research fellow at AnalytiXIN, which is a consortium of health-data organizations, industry partners and university partners in Indiana primarily funded through the Lilly Endowment, IU Health and Eli Lilly and Company. Meghan Zacher declares no competing interests.
Ethical statement
Ethical review of the Add Health data was performed by the Purdue University Institutional Review Board. All other ethical guidelines were followed closely for data storage, analysis, and protection of subjects, as this data is de-identified.
Acknowledgements
This work was supported by the Population Studies and Training Center at Brown University, which receives funding from the Eunice Kennedy Shriver National Institute of Child Health and Human Development of the National Institutes of Health (grant number P2CHD041020). Meghan Zacher was supported by the National Institute on Aging of the National Institutes of Health (grant number K01AG078435). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Data availability
The authors do not have permission to share data.
References
- Amin V., Behrman J.R., Kohler H.-P. Schooling has smaller or insignificant effects on adult health in the US than suggested by cross-sectional associations: New estimates using relatively large samples of identical twins. Social Science & Medicine. 2015;127:181–189. doi: 10.1016/j.socscimed.2014.07.065. 1982. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Becker J., Burik C.A.P., Goldman G., Wang N., Jayashankar H., Bennett M., Belsky D.W., Karlsson Linnér R., Ahlskog R., Kleinman A., Hinds D.A., Caspi A., Corcoran D.L., Moffitt T.E., Poulton R., Sugden K., Williams B.S., Harris K.M., Steptoe A.…Okbay A. Resource profile and user guide of the polygenic index repository. Nature Human Behaviour. 2021:1–15. doi: 10.1038/s41562-021-01119-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Behrman J.R., Kohler H.-P., Myrup Jensen V., Pedersen D., Petersen I., Paul B., Christensen K. Does more schooling reduce hospitalization and delay mortality? New evidence based on Danish twins. Demography. 2011;48(4):1347–1375. doi: 10.1007/s13524-011-0052-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Belsky D.W., Domingue B.W., Wedow R., Arseneault L., Boardman J.D., Caspi A., Conley D., Fletcher J.M., Freese J., Herd P., Moffitt T.E., Poulton R., Sicinski K., Wertz J., Harris K.M. Genetic analysis of social-class mobility in five longitudinal studies. Proceedings of the National Academy of Sciences. 2018;115(31):E7275–E7284. doi: 10.1073/pnas.1801238115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Belsky D.W., Moffitt T.E., Corcoran D.L., Domingue B., Harrington H.L., Hogan S., Houts R., Ramrakha S., Sugden K., Williams B.S., Poulton R., Caspi A. The genetics of success: How single-nucleotide polymorphisms associated with educational attainment relate to life-course development. Psychological Science. 2016;27(7):957–972. doi: 10.1177/0956797616643070. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Benjamin D., Cesarini D., Okbay A., Turley P. Polygenic index repository user guide. (Version 1.0) https://hrsdata.isr.umich.edu/sites/default/documentation/other/User%20Guide_v1.0.pdf
- Boardman J.D., Domingue B.W., Daw J. What can genes tell us about the relationship between education and health? Social Science & Medicine. 2015;127:171–180. doi: 10.1016/j.socscimed.2014.08.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boardman J.D., Fletcher J.M. To cause or not to cause? That is the question, but identical twins might not have all of the answers. Social Science & Medicine. 2015;127:198–200. doi: 10.1016/j.socscimed.2014.10.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Boardman J.D., Fletcher J.M. Evaluating the continued integration of genetics into medical sociology. Journal of Health and Social Behavior. 2021;62(3):404–418. doi: 10.1177/00221465211032581. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Böckerman P., Maczulskij T. The education-health nexus: Fact and fiction. Social Science & Medicine. 2016;150:112–116. doi: 10.1016/j.socscimed.2015.12.036. [DOI] [PubMed] [Google Scholar]
- Braudt D.B. Sociogenomics in the 21st century: An introduction to the history and potential of genetically-informed social science. Sociology Compass. 2018;12(10) doi: 10.1111/soc4.12626. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bulik-Sullivan B., Finucane H.K., Anttila V., Gusev A., Day F.R., Po-Ru L., Consortium R.G., Genomics Consortium P., Genetic Consortium for Anorexia Nervosa of the Wellcome Trust Case Control Consortium 3. Duncan L., Perry J.R.B., Patterson N., Robinson E.B., Daly M.J., Price A.L., Neale B.M. An atlas of genetic correlations across human diseases and traits. Nature Genetics. 2015;47(11):1236–1241. doi: 10.1038/ng.3406. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K.…Marchini J. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562(7726):203–209. doi: 10.1038/s41586-018-0579-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Centers for Disease Control and Prevention Defining adult overweight and obesity. 2022. https://www.cdc.gov/obesity/basics/adult-defining.html
- Cesarini D., Visscher P.M. Genetics and educational attainment. Npj Science of Learning. 2017;2(1):4. doi: 10.1038/s41539-017-0005-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chabris C.F., Lee J.J., Cesarini D., Benjamin D.J., Laibson D.I. The fourth law of behavior genetics. Current Directions in Psychological Science. 2015;24(4):304–312. doi: 10.1177/0963721415580430. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Clogg C.C., Petkova E., Haritou A. Statistical methods for comparing regression coefficients between models. American Journal of Sociology. 1995;100(5):1261–1293. doi: 10.1086/230638. [DOI] [Google Scholar]
- Conley D. Socio-genomic research using genome-wide molecular data. Annual Review of Sociology. 2016;42(1):275–299. doi: 10.1146/annurev-soc-081715-074316. [DOI] [Google Scholar]
- Conley D., Sotoudeh R., Laidley T. Birth weight and development: Bias or heterogeneity by polygenic risk factors? Population Research and Policy Review. 2019;38(6):811–839. doi: 10.1007/s11113-019-09559-6. [DOI] [Google Scholar]
- Conti G., Heckman J., Urzua S. The education-health gradient. The American Economic Review. 2010;100(2):234–238. doi: 10.1257/aer.100.2.234. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Crimmins E., Faul J., Kim J.K., Weir D. Documentation of biomarkers in the 2010 and 2012 Health and Retirement Study. HRS documentation report. Survey Research Center, University of Michigan; Ann Arbor, MI: 2015. https://hrsdata.isr.umich.edu/sites/default/files/documentation/data-descriptions/Biomarker2010and2012_1.pdf [Google Scholar]
- Crimmins E.M., Faul J., Kim J.K., Guyer H., Langa K., Ofstedal M.B.…Weir D. Documentation of biomarkers in the 2006 and 2008 Health and Retirement Study. HRS documentation report. Survey Research Center, University of Michigan; Ann Arbor, MI: 2013. https://hrsonline.isr.umich.edu/sitedocs/userg/Biomarker2006and2008.pdf [Google Scholar]
- Cutler D., Lleras-Muney A. In: Making Americans healthier: Social and economic policy as health policy. Schoeni R.F., House J.S., Kaplan G.A., Pollack H., editors. Russell Sage Foundation; New York: 2008. Education and health: Evaluating theories and evidence; pp. 29–60. [Google Scholar]
- Domingue B.W., Belsky D.W., Conley D., Harris K.M., Boardman J.D. Polygenic influence on educational attainment: New evidence from the National Longitudinal Study of Adolescent to Adult Health. AERA Open. 2015;1(3) doi: 10.1177/2332858415599972. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dudbridge F. Power and predictive accuracy of polygenic risk scores. PLoS Genetics. 2013;9(3) doi: 10.1371/journal.pgen.1003348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Farkas G. Cognitive skills and noncognitive traits and behaviors in stratification processes. Annual Review of Sociology; Palo Alto. 2003;29:541–562. doi: 10.1146/annurev.soc.29.010202.100023. [DOI] [Google Scholar]
- Fletcher J.M. New evidence of the effects of education on health in the US: Compulsory schooling laws revisited. Social Science & Medicine. 2015;127:101–107. doi: 10.1016/j.socscimed.2014.09.052. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Freese J. The arrival of social science genomics. Contemporary Sociology. 2018;47(5):524–536. doi: 10.1177/0094306118792214a. [DOI] [Google Scholar]
- Freese J., Lutfey K. In: Handbook of the sociology of health, illness, and healing: A blueprint for the 21st century, handbooks of sociology and social research. Pescosolido B.A., Martin J.K., McLeod J.D., Rogers A., editors. Springer; New York, NY: 2011. Fundamental causality: Challenges of an animating concept for medical sociology; pp. 67–81. [Google Scholar]
- Fujiwara T., Kawachi I. Is education causally related to better health? A twin fixed-effect study in the USA. International Journal of Epidemiology. 2009;38(5):1310–1322. doi: 10.1093/ije/dyp226. [DOI] [PubMed] [Google Scholar]
- Ge T., Chen C.-Y., Neale B.M., Sabuncu M.R., Smoller J.W. Phenome-wide heritability analysis of the UK Biobank. PLoS Genetics. 2017;13(4) doi: 10.1371/journal.pgen.1006711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gilman S.E., Loucks E.B. Another casualty of sibling fixed-effects analysis of education and health: An informative null, or null information? Social Science & Medicine. 2014;118:191–193. doi: 10.1016/j.socscimed.2014.06.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haas S.A. The long-term effects of poor childhood health: An assessment and application of retrospective reports. Demography. 2007;44(1):113–135. doi: 10.1353/dem.2007.0003. [DOI] [PubMed] [Google Scholar]
- Halpern-Manners A., Helgertz J., Warren J.R., Roberts E. The effects of education on mortality: Evidence from linked U.S. census and administrative mortality data. Demography. 2020;57(4):1513–1541. doi: 10.1007/s13524-020-00892-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Halpern-Manners A., Schnabel L., Hernandez E.M., Silberg J.L., Eaves L.J. The relationship between education and mental health: New evidence from a discordant twin study. Social Forces. 2016;95(1):107–131. doi: 10.1093/sf/sow035. [DOI] [Google Scholar]
- Hauser R.M., Alberto P. Adolescent IQ and survival in the Wisconsin Longitudinal Study. Journals of Gerontology Series B: Psychological Sciences and Social Sciences. 2011;66B(Suppl 1):i91–i101. doi: 10.1093/geronb/gbr037. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Health and Retirement Study Quality control report for genotypic data. https://hrs.isr.umich.edu/sites/default/files/genetic/HRS-QC-Report-Phase-4_Nov2021_FINAL.pdf
- Health and Retirement Study . University of Michigan with funding from the National Institute on Aging; Ann Arbor, MI: 2022. RAND HRS longitudinal file 2018 (V2) public use dataset. (NIA U01AG009740) [Google Scholar]
- Herd P., Freese J., Sicinski K., Domingue B.W., Harris K.M., Wei C., Hauser R.M. Genes, gender inequality, and educational attainment. American Sociological Review. 2019;84(6):1069–1098. doi: 10.1177/0003122419886550. [DOI] [Google Scholar]
- Herle M., Smith A.D., Kininmonth A., Llewellyn C. The role of eating behaviours in genetic susceptibility to obesity. Current Obesity Reports. 2020;9(4):512–521. doi: 10.1007/s13679-020-00402-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hummer R.A., Lariscy J.T. In: International handbook of adult mortality, international handbooks of population. Rogers R.G., Crimmins E.M., editors. Springer Netherlands; Dordrecht: 2011. Educational attainment and adult mortality; pp. 241–261. [Google Scholar]
- Jackson M.I. Understanding links between adolescent health and educational attainment. Demography. 2009;46(4):671–694. doi: 10.1353/dem.0.0078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jencks C. Heredity, environment, and public policy reconsidered. American Sociological Review. 1980;45(5):723–736. doi: 10.2307/2094892. [DOI] [PubMed] [Google Scholar]
- Link B.G., Phelan J. Social conditions as fundamental causes of disease. Journal of Health and Social Behavior. 1995;35:80–94. doi: 10.2307/2626958. [DOI] [PubMed] [Google Scholar]
- Link B.G., Phelan J.C., Miech R., Westin E.L. The resources that matter: Fundamental social causes of health disparities and the challenge of intelligence. Journal of Health and Social Behavior. 2008;49(1):72–91. doi: 10.1177/002214650804900106. [DOI] [PubMed] [Google Scholar]
- Lleras C. Do skills and behaviors in high school matter? The contribution of noncognitive factors in explaining differences in educational attainment and earnings. Social Science Research. 2008;37(3):888–902. doi: 10.1016/j.ssresearch.2008.03.004. [DOI] [Google Scholar]
- Lleras-Muney A. The relationship between education and adult mortality in the United States. The Review of Economic Studies. 2005;72(1):189–221. doi: 10.1111/0034-6527.00329. [DOI] [Google Scholar]
- Lundborg P. The health returns to schooling—what can we learn from twins? Journal of Population Economics. 2013;26(2):673–701. doi: 10.1007/s00148-012-0429-5. [DOI] [Google Scholar]
- Lundborg P., Hampus Lyttkens C., Paul N. The effect of schooling on mortality: New evidence from 50,000 Swedish twins. Demography. 2016;53(4):1135–1168. doi: 10.1007/s13524-016-0489-3. [DOI] [PubMed] [Google Scholar]
- Madsen M., Andersen A.-M.N., Christensen K., Andersen P.K., Osler M. Does educational status impact adult mortality in Denmark? A twin approach. American Journal of Epidemiology. 2010;172(2):225–234. doi: 10.1093/aje/kwq072. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Madsen M., Andersen P.K., Gerster M., Andersen A.-M.N., Christensen K., Osler M. Are the educational differences in incidence of cardiovascular disease explained by underlying familial factors? A twin study. Social Science & Medicine. 2014;118:182–190. doi: 10.1016/j.socscimed.2014.04.016. 1982. [DOI] [PubMed] [Google Scholar]
- Martin A.R., Gignoux C.R., Walters R.K., Wojcik G.L., Neale B.M., Gravel S., Daly M.J., Bustamante C.D., Kenny E.E. Human demographic history impacts genetic risk prediction across diverse populations. The American Journal of Human Genetics. 2017;100(4):635–649. doi: 10.1016/j.ajhg.2017.03.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Martin A.R., Kanai M., Kamatani Y., Okada Y., Neale B.M., Daly M.J. Current clinical use of polygenic scores will risk exacerbating health disparities. Nature Genetics. 2019;51(4):584–591. doi: 10.1038/s41588-019-0379-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McCarthy M.I., Abecasis G.R., Cardon L.R., Goldstein D.B., Little J., Ioannidis J.P.A., Hirschhorn J.N. Genome-wide association studies for complex traits: Consensus, uncertainty and challenges. Nature Reviews Genetics. 2008;9(5):356–369. doi: 10.1038/nrg2344. [DOI] [PubMed] [Google Scholar]
- McGue M., Osler M., Christensen K. Causal inference and observational research: The utility of twins. Perspectives on Psychological Science: A Journal of the Association for Psychological Science. 2010;5(5):546–556. doi: 10.1177/1745691610383511. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mirowsky J., Ross C.E. Walder de Gruyter, Inc; New York, NY: 2003. Education, social status, and health. [Google Scholar]
- Montez J.K., Hayward M.D. Cumulative childhood adversity, educational attainment, and active life expectancy among U.S. adults. Demography. 2014;51(2):413–435. doi: 10.1007/s13524-013-0261-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mostafavi H., Harpak A., Agarwal I., Conley D., Pritchard J.K., Przeworski M. Variable prediction accuracy of polygenic scores within an ancestry group. eLife. 2020;9 doi: 10.7554/eLife.48376. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Neale Lab UK Biobank GWAS - Round 2. http://www.nealelab.is/uk-biobank
- Neale M.C., Maes H.M. Kluwer Academic Publishers B. V; Dordrecht, The Netherlands: 1996. Methodology for genetic studies of twins and families. [Google Scholar]
- Okbay A., Beauchamp J.P., Fontana M.A., Lee J.J., Pers T.H., Rietveld C.A., Turley P., Chen G.-B., Emilsson V., Meddens S.F.W., Oskarsson S., Pickrell J.K., Thom K., Timshel P., de Vlaming R., Abdellaoui A., Ahluwalia T.S., Bacelis J., Baumbach C.…Benjamin D.J. Genome-wide association study identifies 74 loci associated with educational attainment. Nature. 2016;533(7604):539–542. doi: 10.1038/nature17671. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Okbay A., Wu Y., Wang N., Jayashankar H., Bennett M., Nehzati S.M., Sidorenko J., Kweon H., Goldman G., Gjorgjieva T., Jiang Y., Hicks B., Tian C., Hinds D.A., Ahlskog R., Magnusson P.K.E., Oskarsson S., Hayward C., Campbell A.…Young A.I. Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nature Genetics. 2022;54(4):437–449. doi: 10.1038/s41588-022-01016-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Palloni A. Reproducing inequalities: Luck, wallets, and the enduring effects of childhood health. Demography. 2006;43(4):587–615. doi: 10.1353/dem.2006.0036. [DOI] [PubMed] [Google Scholar]
- Pingault J.-B., Allegrini A.G., Odigie T., Frach L., Baldwin J.R., Rijsdijk F., Dudbridge F. Research review: How to interpret associations between polygenic scores, environmental risks, and phenotypes. Journal of Child Psychology and Psychiatry n/a. 2022;(n/a) doi: 10.1111/jcpp.13607. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pingault, Jean-Baptiste, Rijsdijk F., Schoeler T., Choi S.W., Selzam S., Krapohl E., O'Reilly P.F., Frank D. Genetic sensitivity analysis: Adjusting for genetic confounding in epidemiological associations. PLoS Genetics. 2021;17(6) doi: 10.1371/journal.pgen.1009590. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Plomin R., Sophie von Stumm S. Polygenic scores: Prediction versus explanation. Explanations explanation. Molecular Psychiatry. 2022;27(1):49–52. doi: 10.1038/s41380-021-01348-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Polderman T.J.C., Benyamin B., de Leeuw C.A., Sullivan P.F., van Bochoven A., Visscher P.M., Posthuma D. Meta-analysis of the heritability of human traits based on fifty years of twin studies. Nature Genetics. 2015;47(7):702–709. doi: 10.1038/ng.3285. [DOI] [PubMed] [Google Scholar]
- Radloff L.S. The CES-D scale: A self-report depression scale for research in the general population. Applied Psychological Measurement. 1977;1(3):385–401. doi: 10.1177/014662167700100306. [DOI] [Google Scholar]
- RAND . RAND Center for the Study of Aging, with funding from the National Institute on Aging and the Social Security Administration; Santa Monica, CA: 2022. RAND HRS longitudinal file 2018 (V2) [Google Scholar]
- Schnittker J. Cognitive abilities and self-rated health: Is there a relationship? Is it growing? Does it explain disparities? Social Science Research. 2005;34(4):821–842. doi: 10.1016/j.ssresearch.2005.01.003. [DOI] [Google Scholar]
- Schoeler T., Speed D., Porcu E., Pirastu N., Pingault J.-B., Zoltán K. Participation bias in the UK Biobank distorts genetic associations and downstream analyses. Nature Human Behaviour. 2023;7(7):1216–1227. doi: 10.1038/s41562-023-01579-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Solovieff N., Cotsapas C., Lee P.H., Purcell S.M., Smoller J.W. Pleiotropy in complex traits: Challenges and strategies. Nature Reviews Genetics. 2013;14(7):483–495. doi: 10.1038/nrg3461. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Song X., Mare R.D. Short-term and long-term educational mobility of families: A two-sex approach. Demography. 2017;54(1):145–173. doi: 10.1007/s13524-016-0540-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sonnega A., Faul J.D., Beth Ofstedal M., Langa K.M., Phillips J.W.R., Weir D.R. Cohort profile: The Health and Retirement Study (HRS) International Journal of Epidemiology. 2014;43(2):576–585. doi: 10.1093/ije/dyu067. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stephan Y., Sutin A.R., Luchetti M., Caille P., Terracciano A. An examination of potential mediators of the relationship between polygenic scores of BMI and waist circumference and phenotypic adiposity. Psychology and Health. 2020;35(9):1151–1161. doi: 10.1080/08870446.2020.1743839. [DOI] [PubMed] [Google Scholar]
- Uddin M.J., Hjorthøj C., Ahammed T., Nordentoft M., Ekstrøm C.T. The use of polygenic risk scores as a covariate in psychological studies. Methods in Psychology. 2022;7 doi: 10.1016/j.metip.2022.100099. [DOI] [Google Scholar]
- Vilhjálmsson B.J., Yang J., Finucane H.K., Gusev A., Lindström S., Ripke S., Genovese G., et al. Modeling linkage disequilibrium increases accuracy of polygenic risk scores. The American Journal of Human Genetics. 2015;97(4):576–592. doi: 10.1016/j.ajhg.2015.09.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Visscher P.M., Wray N.R., Zhang Q., Sklar P., McCarthy M.I., Brown M.A., Yang J. 10 years of GWAS discovery: Biology, function, and translation. The American Journal of Human Genetics. 2017;101(1):5–22. doi: 10.1016/j.ajhg.2017.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wedow R., Zacher M., Huibregtse B.M., Harris K.M., Domingue B.W., Boardman J.D. Education, smoking, and cohort change: Forwarding a multidimensional theory of the environmental moderation of genetic effects. American Sociological Review. 2018;83(4):802–832. doi: 10.1177/0003122418785368. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zajacova A., Lawrence E.M. The relationship between education and health: Reducing disparities through a contextual approach. Annual Review of Public Health. 2018;39(1):273–289. doi: 10.1146/annurev-publhealth-031816-044628. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhao Z., Yang X., Dorn S., Miao J., Barcellos S.H., Fletcher J.M., Lu Q. Controlling for polygenic genetic confounding in epidemiologic association studies. Proceedings of the National Academy of Sciences. 2024;121(44) doi: 10.1073/pnas.2408715121. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zheng H. Why does college education matter? Unveiling the contributions of selection factors. Social Science Research. 2017;68:59–73. doi: 10.1016/j.ssresearch.2017.09.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The authors do not have permission to share data.



