Causality is the crown jewel of human curiosity because it allows us to make informed choices and decisions on the basis of our understanding of the factors that may improve the quality of our lives. Legend has it that Isaac Newton was inspired to discover the laws of gravity in 1687 when he saw an apple fall from a tree. Although this story is likely apocryphal, 400 years ago was a relatively short period of time compared with the hundreds of thousands or perhaps millions of years that humans have been around. This underscores the difficulty of establishing causality in research, particularly in the field of biomedicine, where it can be time-consuming and expensive when it is possible at all.
Table 1 presents a glossary of a few relevant terms and methods for ease of understanding the technical aspects.
TABLE 1.
Glossary of a few selected terms and methods.
| Term | Description |
|---|---|
| Genome-wide association studies | A panel of (hundreds of thousands or millions) genetic markers (e.g., single-nucleotide polymorphisms [SNPs]) are genotyped and tested for association with traits or disease outcomes |
| Instrumental variable | A variable with particular characteristics that are possessed by randomization in a randomized controlled trials, and it aims to control the confounding effects and filter out the causal effects |
| Linkage disequilibrium | Correlation of alleles at different loci in a population |
| Mendelian randomization | A method that uses genetic variants as instrumental variables to facilitate the causal effects of risk factors on outcomes in genetic epidemiological studies |
| Pleiotropy | Effects of 1 gene on different traits |
| SNP | Germline substitution of a single nucleotide at a specific position in the genome |
Randomized controlled trials (RCTs) are often considered the gold standard for inferring causality in medicine. For instance, Legro et al. (1) used this design to compare the effectiveness of letrozole vs. clomiphene as the first-line treatment for infertile women with polycystic ovary syndrome in achieving a live birth. Despite the ongoing debates about the technical aspects of RCTs, there is a general consensus that RCTs should be used when feasible and affordable to obtain the least biased conclusion regarding the causality of treatments. However, the feasibility of conducting an RCT can be a major problem because of the ethical and resource constraints involved.
Efforts to explore alternative designs, such as observational studies, are understandably important. Observational studies can use existing data sources such as electronic health records and publicly accessible databases, including the NIH dbGaP and UK Biobank (2). They can also collect new data, with genome-wide association studies (GWASs) being a commonly used study design to identify genetic variants for complex diseases (3). However, it is widely recognized that observational studies can establish only associations and not causal relationships.
In the example of letrozole vs. clomiphene, a wealth of data from electronic health records is available to examine the live birth outcomes of patients who received either treatment or both. Selecting a cohort of patients using the same inclusion and exclusion criteria as Legro et al. (1) is a relatively straightforward process. However, the difference between such an observational study and an RCT lies in how the patient received the treatment. In the study by Legro et al. (1), the choice of treatment was random and unknown to the patient, healthcare professionals, and investigators who may have had a vested interest in the study. Randomization in RCTs eliminates concerns about biases in the treatment allocation, which are not possible in observational studies. Although the use of electronic health records can be conducted carefully and advanced causal inference techniques are available, it is challenging to convincingly establish that the effects observed for letrozole or clomiphene are not biased by healthcare access or other factors that may affect live birth.
The desire to establish causality through observational studies coupled with the increasing number of large-scale GWASs that contain genetic markers in addition to variables such as demographic and clinical data as in typical observational studies has given rise to the concept of Mendelian randomization. The idea is to let Mendelian randomization in GWASs play the role that randomization does in RCTs so that we can assess the causal effects of certain exposures, including genetic markers.
In a letter to the editor, Katan (4) discussed the use of what is now known as Mendelian randomization to determine the causal effect of serum cholesterol on cancer by examining the relationship between apolipoprotein E isoforms and both serum cholesterol and cancer. It is noteworthy that Katan (4) did not use the term “Mendelian randomization.” Smith and Ebrahim (5) gave credit to Gray and Wheatley (6) for coining the term. It is interesting to observe that the articles by neither Katan (4) nor Gray and Wheatley (6) were published as regular original research articles.
Figure 1 illustrates the basic idea of what Mendelian randomization is and how it may or may not work. Figure 1A displays the basic design by Legro et al. (1), and we discuss the critical role of randomization in allowing them to conclude whether letrozole or clomiphene was more effective. The motivation is to accomplish the same objective with the randomization. To this end, let us put the technique of randomization in a general context of causal inference. In observational studies, we do not have the ability for randomization. To facilitate the possibility of causal inference, we need a surrogate for randomization that plays an analogous role in the observational studies. Such surrogates are referred to as instrumental variables in causal inference literature. Econometrics literature gives credit to Wright (7) for introducing this concept. For those unfamiliar with the instrumental variables, it may be easier to envision them as the randomization in RCTs.
FIGURE 1.

Correspondence between randomization in randomized controlled trials, instrumental variables in observational studies, and Mendelian randomization in genome-wide association studies. (A) Randomized controlled trial. (B) Causal inference. (C) Mendelian randomization.
In causal inference (Fig. 1B), 3 categories of variables are involved: the outcome variables (e.g., live birth [Fig. 1A and B]), the exposure variables (e.g., the treatment [Fig. 1A and B]), and the instrumental variables. In RCTs, randomization is the instrumental variable (Fig. 1A and B). Randomization is the first step to implement in the design of RCTs, so it is intuitive to understand. Instrumental variables do not present themselves to us in observational studies. We need to look for them in a given data set and verify whether they have the desired characteristics, which are described below. Such characteristics are imposed so that instrumental variables may act like randomization and facilitate the causal inference.
Specifically, an instrumental variable must meet the following 3 basic conditions: (1) Relevance: it must be associated with the exposure. In RCTs, the randomization determines the treatment assignment. It is not just associated, but it fully determines the exposure; (2) Exchangeability: no causes of an instrumental variable can affect the outcome. In RCTs, randomization is generated by computers without any knowledge whatsoever on the outcome. It is inconceivable for the process of randomization to affect the outcome. On the other hand, when the randomization is generated by an investigator who also treats the patients, it may not satisfy this condition of an instrumental variable; and (3) Exclusion restriction: the instrumental variables affect the outcome only through the exposure of interest, not otherwise directly or indirectly. This is why randomization should be blind to the investigators and patients to the extent possible without jeopardizing the safety of the participants. In open-label trials involving surgeries and behavioral interventions (8), it is extremely important to maintain the integrity of the randomization. Again, when we view an RCT as an observational study, randomization is the most obvious example of an instrumental variable. We can also view instrumental variables as being constructed to mimic the randomization so that a causal inference is possible.
Through the use of the instrumental variables, we can infer the causal relationship between the exposure and outcome (Fig. 1B). The question is: does Mendelian randomization satisfy the 3 conditions? When so, as displayed in Figure 1C, it can play the role of the randomization in RCTs.
To understand the Mendelian randomization approach, we need to revise Mendel’s first and second laws of genetic inheritance: the law of segregation and the law of independent assortment (9). The law of segregation states that at every point in the autosomal genome, an offspring randomly inherits 1 allele from their mother and 1 allele from their father. The law of independent assortment implies that these alleles will be passed to the offspring independently of each other, except in regions of the genome that are genetically linked in the DNA of the parents. Simply put, we do not inherit the entire genome from both parents but instead randomly inherit half of our genetic material from each. The fertilization process is essentially a randomization process, albeit 1 performed by 2 parents rather than a computer.
Let us suppose we want to use Mendelian randomization to evaluate the effectiveness of letrozole and clomiphene, as depicted in Figure 1. To satisfy the first condition of an instrumental variable, we need to identify a genetic marker that is associated with treatment assignment. We can examine an electronic health record database and select patients with either genotype 1 who received letrozole treatment or genotype 2 who received clomiphene treatment. However, satisfying the second condition of an instrumental variable, which requires us to ensure that any factor contributing to an offspring’s genotype has no effect on their ability to deliver a live birth, can be challenging. Furthermore, the third condition of an instrumental variable requires us to demonstrate that the effect of genotype 1 or 2 on live birth is solely because of the use of letrozole or clomiphene, which is also a difficult task.
Randomized controlled trials offer several important advantages, such as a relatively short duration between treatment assignment and outcome, close patient observation, and the ability to control treatment assignment ratios and overall sample size. However, in Mendelian randomization studies, the time from birth to the occurrence of the outcome can be lengthy and unmonitored. In addition, because humans share 99% of their genome, genotypes 1 and 2 may be imbalanced in a given population. Even for a large database such as the UK Biobank, the final cohort for a Mendelian randomization–based study may be small because of attrition throughout the process. Therefore, although Mendelian randomization may be a worthy candidate for an instrumental variable in GWASs, it may not offer definitive or persuasive evidence for a causal relationship.
Despite the limitations of Mendelian randomization, it can still offer a more robust estimate of causality than conventional observational studies. Although it may not control for confounding and reverse causation as effectively as randomization in RCTs, it can still help mitigate bias.
The Mendelian randomization approach involves 3 main steps to verify that certain genetic variants meet the criteria for instrumental variables. These genetic variants, when identified and verified, play the role of randomization as in RCTs and are used to assess the causal effects of other exposures. Step 1: identification of genetic variants that are associated with the exposures of interest. The exposures can be a treatment and/or other genetic variables for which we wish to assess the causal relation to the outcome of interest; Step 2: evaluation of the association between these genetic variants and the outcome of interest, which can be disease status or other specific variables. This is to assess the validity of the second and third conditions of the instrumental variables; and Step 3: estimation of the causal effect of the exposure on the outcome using genetic variants as instrumental variables.
Yuan et al. (10) conducted a Mendelian randomization study to examine the potential causal association between smoking, alcohol, and coffee consumption with pregnancy loss. They used a large sample of 60,565 cases with pregnancy loss and 130,687 noncases from the UK Biobank and 3,312 cases with pregnancy loss and 64,578 noncases from FinnGen. They concluded that “This study on the basis of genetic data suggests the causal potential of the association of smoking but not moderate alcohol and coffee consumption with pregnancy loss.” However, it should be noted that the investigators used a cautious language, stating that their findings suggest the “causal potential” of the association, which is not so different from any association.
On the basis of the literature, they selected 314, 84 and 12 single-nucleotide polymorphisms as instrumental variables for smoking initiation, alcohol consumption, and coffee consumption, respectively, with documented significant associations with the respective exposures (the first condition of the instrumental variables); for those associated with the same exposure, they are not in linkage disequilibrium (to minimize possible violations from the second and third conditions of the instrumental variables). Then, they conducted regression analysis with the adjustment for pleiotropy (to minimize possible violations from the third condition of the instrumental variables). The investigators concluded that “the major strength was the MR designs, which diminished residual confounding and reverse causality and, thereby, improved the causal inference in associations of smoking and alcohol and coffee consumption with pregnancy loss.” Strictly speaking, Mendelian randomization is a method of conducting causal inference in GWASs and less of a study design. The belief is that Mendelian randomization to play the role of randomization as in RCTs so that we are less concerned with potential and hidden factors that hinder our ability to identify causal effects in GWASs. However, wishful thinking does not equate the reality. Thus, it would have been useful when they provided data to support the extent to which Mendelian randomization did a reasonable job. Although the investigators took proper steps to reduce the possible violations from the conditions for the instrumental variables, it would have been useful if they provided specific data for assurance. Even for randomization in RCTs, which is the most ideal example of an instrumental variable, we cannot take it for granted that it works in practice. The comparison of variables potentially affecting the outcome is carefully examined between the intervention groups. With Mendelian randomization, using certain single-nucleotide polymorphisms as “valid instrument” variables after a few basic steps but without defining the validity and presenting the data to support the validity can lead to the misuse of the method and ultimately misleading findings.
In conclusion, Mendelian randomization is a potentially powerful approach that can improve our understanding of important relationships by exploiting genetic data. Although it is essential to follow the proper steps when using existing software for analysis, it is equally important to ensure that the obtained results are reliable and provide sufficient assurance.
Acknowledgments
Supported by NIH grants R01MH116527, R01HG010171, R01HD100336, and R01HD100369. H.Z. has nothing to disclose.
REFERENCES
- 1.Legro RS, Brzyski RG, Diamond MP, Coutifaris C, Schlaff WD, Casson P, et al. Letrozole versus clomiphene for infertility in the polycystic ovary syndrome. N Engl J Med 2014;371:119–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Liu Z, Dai W, Wang S, Yao Y, Zhang H. Deep learning identified genetic variants for COVID-19-related mortality among 28,097 affected cases in UK Biobank. Genet Epidemiol 2023;47:215–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Zhang H, Baldwin DA, Bukowski RK, Parry S, Xu Y, Song C, et al. A genome-wide association study of early spontaneous preterm delivery. Genet Epidemiol 2015;39:217–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Katan MB. Apolipoprotein E isoforms, serum cholesterol, and cancer. Lancet 1986;327:507–8. [DOI] [PubMed] [Google Scholar]
- 5.Smith GD, Ebrahim S. “Mendelian randomization”: can genetic epidemiology contribute to understanding environmental determinants of disease? Int J Epidemiol 2003;32:1–22. [DOI] [PubMed] [Google Scholar]
- 6.Gray R, Wheatley K. How to avoid bias when comparing bone marrow transplantation with chemotherapy. Bone Marrow Transplant 1991;7(Suppl 3):9–12. [PubMed] [Google Scholar]
- 7.Wright PG. The tariff on animal and vegetable oils. New York: Macmillan Co; 1928. [Google Scholar]
- 8.Legro RS, Hansen KR, Diamond MP, Steiner AZ, Coutifaris C, Cedars MI, et al. Effects of preconception lifestyle intervention in infertile women with obesity: the FIT-PLESE randomized controlled trial. PLOS Med 2022;19:e1003883. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Sanderson E, Glymour MM, Holmes MV, Kang H, Morrison J, Munafò MR, et al. Mendelian randomization. Nat Rev Methods Primers 2022;2:6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Yuan S, Liu J, Larsson SC. Smoking, alcohol and coffee consumption and pregnancy loss: a Mendelian randomization investigation. Fertil Steril 2021;116:1061–7. [DOI] [PubMed] [Google Scholar]
