Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2014 Feb 6;94(2):176–185. doi: 10.1016/j.ajhg.2013.12.010

Revisiting the Thrifty Gene Hypothesis via 65 Loci Associated with Susceptibility to Type 2 Diabetes

Qasim Ayub 1,7, Loukas Moutsianas 2,7, Yuan Chen 1, Kalliope Panoutsopoulou 1, Vincenza Colonna 1,3, Luca Pagani 1, Inga Prokopenko 2, Graham RS Ritchie 1,4, Chris Tyler-Smith 1, Mark I McCarthy 2,5,6, Eleftheria Zeggini 1, Yali Xue 1,
PMCID: PMC3928649  PMID: 24412096

Abstract

We have investigated the evidence for positive selection in samples of African, European, and East Asian ancestry at 65 loci associated with susceptibility to type 2 diabetes (T2D) previously identified through genome-wide association studies. Selection early in human evolutionary history is predicted to lead to ancestral risk alleles shared between populations, whereas late selection would result in population-specific signals at derived risk alleles. By using a wide variety of tests based on the site frequency spectrum, haplotype structure, and population differentiation, we found no global signal of enrichment for positive selection when we considered all T2D risk loci collectively. However, in a locus-by-locus analysis, we found nominal evidence for positive selection at 14 of the loci. Selection favored the protective and risk alleles in similar proportions, rather than the risk alleles specifically as predicted by the thrifty gene hypothesis, and may not be related to influence on diabetes. Overall, we conclude that past positive selection has not been a powerful influence driving the prevalence of T2D risk alleles.

Introduction

Type 2 diabetes (T2D) was responsible for more than three million deaths worldwide in 20041 and is projected to be the seventh leading cause of death by 2030.2 It shows substantial familial clustering, with an estimated lifetime risk of 38% by the age of 80 if one parent is affected and 60% by the age of 60 if both are affected.3 T2D has therefore been the focus of numerous medical-genetic studies, including candidate gene and genome-wide association studies, which have sought to identify genetic variants that influence the risk of developing T2D. In combination, these studies have identified at least 65 genomic regions associated with T2D. At most of these loci, the association signal is driven by alleles on common haplotypes, and transethnic and fine-mapping data suggest that the causal allele—which typically has yet to be identified unequivocally—is also common.4

Because T2D decreases quality-adjusted life expectancy by about 11 years,5 impairs sperm parameters,6 reduces fecundability,7 and can lead to menstrual anomalies,8 the existence of variants that have high frequency in the population but are disadvantageous in these ways creates an evolutionary puzzle. Why has natural selection not eliminated them from the population? Two classes of explanation can be considered. The risk variants might be, from an evolutionary perspective, effectively neutral, for example because their effects are all individually small, or manifest only in old age. According to this model, their high frequency would be a consequence solely of genetic drift. Alternatively, they might have some compensating advantage, now or in the past, so that their high frequency would be a consequence of positive selection. This second class of explanation has been the focus of much attention since Neel proposed that T2D was “a ‘thrifty’ genotype rendered detrimental by ‘progress’”;9,10 in other words, the genetic background that leads to disadvantageous T2D today was advantageous in past environments. We set out to investigate the genetic support for this hypothesis. In order to do so, we first need to consider in more detail the timescale of human evolution.

The human lineage (hominins) split from that of our closest living evolutionary relatives, chimpanzees and bonobos, 6–7 million years ago. Studies of hominin fossils including examination of their teeth suggest that our early ancestors were omnivorous, incorporating substantial amounts of animal as well as plant foods into their diet, especially after stone tool use became common 2.6 million years ago.11 This mixed hunter-gatherer diet was probably universal until the beginning of the Neolithic transition ∼10 thousand years ago, when domestic plants and animals began to provide alternative food sources, leading, in developed countries, to our current westernized diet. In contrast, coalescence times of segments of the autosomes are typically of the order of one million years and seldom more than two million years.12 Consequently, variants favored during hominin evolution before two million years ago would mostly be fixed in the entire human population; only variants arising in the last ∼2 million years would still be variable.

In order to formulate genetic tests of the thrifty gene hypothesis, it is therefore useful to distinguish between “thrifty early” and “thrifty late” versions. The thrifty early version seems to correspond to Neel’s view, who wrote “during the first 99 per cent or more of man’s life on earth, while he existed as a hunter and gatherer, it was often feast or famine”10 and later suggested that at this early stage “a quick insulin trigger … was an asset to our tribal, hunter-gatherer ancestors … since it should have minimized renal loss of precious glucose. Currently, however … overalimentation in the technologically advanced nations result[s] in … diabetes mellitus.”9 If selection acted for millions of years, thrifty (i.e., susceptibility) alleles should be shared by all populations; if other great ape ancestors were also thrifty, as seems likely, these thrifty alleles would also be ancestral, and common derived variants at these sites would have to arise for them to be discovered by genome-wide association studies. We consider the model in which hominids (the clade containing humans, other great apes, and the ancestors linking them) were all thrifty as the simplest version of the thrifty early hypothesis and discuss some alternative formulations of the hypothesis in the Discussion. The alternative, thrifty late interpretation would be that the thrifty genotype was favored only much later in human evolution, for example when humans were expanding into novel environments such as out of Africa ∼60 thousand years ago, into the Americas ∼15 thousand years ago, or voyaging across the Pacific after 5 thousand years ago.11 In this version, thrifty (i.e., susceptibility) alleles should be derived and probably specific to individual populations or groups of populations who experienced the particular environment.

Several previous studies have sought genetic support for the thrifty gene (implicitly the thrifty late) hypothesis. After rs7903146 in an intron of TCF7L2 (MIM 602228) was identified as a risk variant,13 a detailed study of this genomic region in 2007 identified a rapid and recent expansion of a cluster of haplotypes, designated HapA, indicative of positive selection in East Asia (but not Africa or Europe) around eight thousand years ago.14 HapA, however, carried the protective rather than the risk allele. A systematic survey of the 17 risk loci (including TCF7L2) discovered by 2009 replicated the unusual population differentiation at TCF7L2 and reported significantly extended haplotypes at rs10923931 within NOTCH2 (MIM 600275) (iHS value 2.249, 2.3 percentile) but found no overall excess of positive selection signals.15 Another study of 16 loci found unusually high levels of population differentiation in the 100 kb windows surrounding the index SNPs, with HHEX (MIM 604420) in East Asian populations being the most highly differentiated individual gene.16 A more recent investigation in 2012, and its follow-up in 2013, focused on the levels of population differentiation of 12–16 T2D risk alleles and reported a systematic pattern of increasing risk allele frequency from Africa to Europe to East Asia specific to this condition17,18 In a study comparing ten T2D variants in a herder and farmer population from Central Asia, evidence for Neolithic and post-Neolithic positive selection was found at some loci, including LEPR (MIM 601007), HHEX, and PON1 (MIM 168820), but selection was spread between the populations and acting on the protective allele.19 Thus, previous work has sought but failed to find genetic evidence for positive selection on T2D risk alleles and has therefore not provided genetic support for the thrifty gene hypothesis.

In the current study, we have reinvestigated this hypothesis. Three developments prompted us to do this. First, far more susceptibility alleles have now been identified, and we were able to investigate 65 of them, a nearly 4-fold increase over previous studies. Second, our formulation of the thrifty early and thrifty late hypotheses permits a more nuanced investigation of the selective regimes. Third, additional tests for positive selection are now possible, including ones based on multiple whole-genome sequences.20 Despite these advantages, we find little evidence to support the thrifty gene hypothesis as an overall explanation for the high frequency of these risk alleles, but we do find some evidence of positive selection influencing allele frequencies at a minority of loci.

Material and Methods

Curation of the T2D Index SNPs

The 65 T2D index SNPs were chosen according to the following criteria: all SNPs that reached genome-wide significance in any cohort and were published prior to August 2011 were included in our analysis. For SNPs found associated with T2D in European cohorts, all SNPs that reached genome-wide significance in the DIAGRAM+ study21 were included. For associations published before 2010 and subsequently replicated successfully by Voight et al.,21 lead SNPs reported in both the original and the DIAGRAM+ study were included, where the two differed. Furthermore, SNPs that reached genome-wide significance in studies by Dupuis et al.22 and Qi et al.23 were also included. For SNPs associated with T2D in Asian cohorts, signals that reached genome-wide significance in studies by Tsai et al.,24 Yamauchi et al.,25 Shu et al.,26 Unoki et al.,27 and Yasuda et al.28 were included in our analysis. The majority of the SNPs were discovered in European populations and only 13 were discovered in Asian populations. No signals were discovered in other populations prior to August 2011.

Most of our analyses are based on the 64 SNPs that lie on autosomes, excluding the SNP (rs5945326) near DUSP9 (MIM 300134) on the X chromosome. We stratified the index SNPs by allelic status of the risk allele (ancestral or derived), effect size (large, where the association was identified between 2006 and 2009, or small, where that association was identified after 2009, often by meta-analysis), and biological function (insulin resistance, reduced beta cell function, or unclassified), based on the function of the nearest gene. All the tests below were carried out on the complete set of 64 as well as the stratified subgroups.

Comparison with Archaic Humans

The allelic status of each of the 65 T2D index SNPs was determined from the EPO alignment.29 The proportion of risk alleles that correspond to the ancestral allele was determined by direct counting. Archaic human data for the allele at each T2D index SNP in the Neandertal30 and Denisovan31 genomes were obtained from the Ensembl and the high-coverage Denisovan genome,32 respectively. We used Fisher’s exact test to examine the enrichment of ancestral risk alleles in both modern humans and archaic genomes compared with matched controls.

Site Frequency Spectrum-Based Analyses

A custom pipeline has been established to investigate whether or not there is enrichment for positive selection signals based on site frequency spectrum statistics (Tajima’s D,33 Fay and Wu’s H,34 and Nielsen et al.’s CLR value35) by using the 1000 Genomes Project Phase 1 data in a chosen set of regions or genes, in comparison to a randomly selected list of control regions or genes matched for length, GC content, and average recombination rate, as described previously.36 Populations were grouped into AFR (YRI + LWK + ASW), ASN (CHB + CHS + JPT), and EUR (CEU + TSI + GBR + FIN) continental regions as the 1000 Genomes Project suggests. The pipeline has been shown to detect selection in lists containing ≥10% positively selected genes.

We compared the distribution of the statistics in nonoverlapping 10 kb windows in the regions 25 or 50 kb upstream and downstream of each index SNP with the matched controls. We divided the whole genome into similar 50 or 100 kb regions and, after excluding the index SNP regions, generated 1,000 control regions for each index SNP that were matched for GC content and recombination rate. A permutation test was used to assess whether or not the distributions of the test statistics were significantly different between the index SNP and control regions.

The nearest genes to the index SNPs were considered as candidate genes and a matched list of control genes was generated from the Ensembl Archives database (Ensembl 54: NCBI Build 36). Fifty-five genes were tagged by the 65 index SNPs, but four genes were excluded from subsequent analyses. These were DUSP9 on the X chromosome, KLF14 (MIM 609393) for which there were no recombination data, and C2CD4A (MIM 610343) and C2CD4B (MIM 610344) that were not listed in NCBI36. The 1,000 matched control sets of genes were generated and the enrichment of positive selections signals were analyzed as described previously.36

CMS Analysis

To further interrogate the T2D-associated regions for evidence of positive selection, we used a modification of the composite of multiple signals (CMS) score.37 The CMS provides an approximate posterior probability that a variant is under selection based on five tests: three measures of haplotype length (iHS, XP-EHH, ΔiHH), one of population differentiation (FST), and one of the frequency difference of the derived allele (ΔDAF). For our analysis, a missing value for any of the five scores was replaced by the mean value for that test across all sites. This was done so as to be able to include this site in the analysis, because otherwise its CMS value would have been zero. The statistic employed for all analyses presented here is the rank of each variant based on its CMS value.

We calculated the CMS value for all 3,021,594 autosomal SNPs included in the third phase of the HapMap project across the European (CEU), African (YRI), Chinese (CHB), and Japanese (JPT) populations.38 For the latter, the data were pooled (CHB+JPT). For each of the index T2D SNPs, we defined a 25 or 50 kb region on either side as a “T2D index region” and compared the highest CMS value of that region to 2,000 randomly picked regions of the same length across the genome. The latter were matched for GC content (GC% of the random regions was 98.75%–101.25% of the GC% of the T2D regions) and average recombination rate. A Mann-Whitney U test was used to compare the difference in CMS ranks between T2D regions and matched control regions, and the final p values were obtained after averaging over 10,000 iterations.

Haplotype Diversity

Nei’s haplotype diversity was calculated for regions of 10 kb surrounding the 64 autosomal index SNPs by using the 1000 Genomes Phase 1 sequencing data, grouping the populations in the same way as for the site frequency spectrum tests.39 We randomly generated 1,000 control sets for each of the 64 SNPs drawn from a subset of the 1000 Genomes Phase I data that contained only markers found in the commonly used genotyping platforms (n = 2,010,212 markers), matched by derived allele frequency (±0.001) in the population that the association was discovered in. The p value was obtained by comparing with 60,760 random genome frequency-matched control regions by a one-sided Kolmogorov-Smirnov Test.

Pairwise FST

The T2D index SNP and matched control SNP frequencies in 849 samples belonging to the AFR, ASN, and EUR continental groups were extracted from Phase 1 of the 1000 Genomes Project.39 Pairwise FST between the three groups was calculated for each of the 64 index SNPs as well as for the other biallelic autosomal SNPs available for those populations. 1,000 matched control SNPs for each of the index SNPs were chosen by matching the global derived allele frequency to ±0.001. The distribution of the 64 index SNPs as a whole or stratified by subgroup as described above was compared to its matched controls by a Mann-Whitney U test.

Identification and Analysis of Individual Genes Showing Evidence of Positive Selection

We assigned an empirical p value to each of the 64 autosomal T2D SNPs by using the most significant p value within the surrounding 100 kb region from the site frequency spectrum-based tests in the YRI, CEU, and CHB populations and the empirical p values (rank p values) obtained from its 1,000 matched controls for the pairwise FST and haplotype diversity tests. (We did not use the CMS value because this also includes haplotype-based and pairwise FST tests, and also because our CMS values are based on genotype data with complex ascertainment.) Then we calculated a combined p value by Fisher’s method,40 which is −2∑ln P (here: −2 × (ln PSFS + ln Phaplotype diversity + ln PFST1 + ln PFST2), where FST1 and FST2 are the two FST values possible for each continental comparison); it has a χ2 distribution with degrees of freedom equal to two times the number of the tests (here, eight). We used this combined p value to rank the individual genes. We set the p value cut off based on a Bonferroni correction (46 tests, 3 populations, 0.05/(463) = 3.6 × 10−4); 18 out of the 64 SNPs are in high LD with another SNP and were not considered as an independent test.

We constructed median-joining networks41 of the inferred haplotypes in high LD (r2 ≥ 0.95) with the index SNP for interesting regions. Networks were based on phased haplotypes generated from the 1000 Genomes Phase 1 low coverage resequencing data for 88 YRI, 85 CEU, and 86 CHB samples. Times to the most recent common ancestor of candidate selected haplotype clusters, and their standard deviations, were estimated by the rho statistic implemented in the network program.41

Prediction of Functional Impact of SNPs

We identified all SNPs in linkage disequilibrium with the autosomal index SNPs at r2 cutoffs of 0.4 and 0.95 in the European sample from the 1000 Genomes Project Phase 1 data.39 We annotated these variants by using custom software that identified any overlapping genes and regulatory annotations. Gene annotations were taken from GENCODE release 1542 and regulatory annotations from the ENCODE project.43 All variants were assigned a functional annotation (stop_gained, frameshift, splice_site, stop_lost, missense, UTR, ncRNA; bound_TF_motif, TFBS, DNase_footprint, DNase_peak, synonymous, intronic, or intergenic) and given a rank from 1 to 14 in this order.

For each index SNP, we used these ranks to select the highest ranked candidate functional variant from the set of linked variants. Where there were multiple candidate SNPs with the same rank, one was selected arbitrarily. We used the same software to annotate the index SNP, and we report this annotation and its associated rank in Table S1 (available online) along with the candidate functional variant.

Results

We compiled a list of 65 index SNPs that influence susceptibility to T2D: the majority of these SNPs were annotated as intronic (39) or intergenic (15), five were nonsynonymous, three lay in a 3′ UTR, and three were within 5 kb of a transcript start site (Table S1).

No Evidence to Support the Thrifty Early Hypothesis

If the thrifty early hypothesis as described above is right, we expect causative T2D risk alleles to correspond to the ancestral states and the index SNP risk alleles to be enriched for ancestral states. However, we found that only about half of the risk alleles were ancestral (36/65, 55%), not significantly different from the matched controls genome-wide. Our analysis has power to detect an enrichment of 72% ancestral states (47/65). Similarly, only 45% (28/62) of the SNPs with the highest functional ranking (see Material and Methods) in high LD with the index SNPs, i.e., possible causal alleles, were ancestral. This is not significantly different from the index SNP proportion (Fisher’s exact test p = 0.285). Therefore, the index SNPs and the candidate functional SNPs showed the same lack of enrichment for the ancestral state.

We also expect to see enrichment of ancestral status for the risk alleles in archaic humans (Neandertals and Denisovans): if thrifty variants have been selected for millions of years and become fixed, the allele carried by these hominins would be ancestral for humans. In Neandertals,30 44/65 index SNP positions were called, of which 38 were ancestral and 53% (20/38) of these were risk alleles. In Denisova,31,32 all index SNPs were called, of which 56 were ancestral and 57% (32/56) of these were risk alleles. These are again similar to controls. One SNP was derived in Denisova but ancestral in Neandertals (Table S1). A similar proportion of the Neandertal (3/6) and Denisova (4/9) genotypes were homozygous for the derived risk allele (Table S2).

No Support for the Thrifty Late Hypothesis

If the thrifty late hypothesis is true, we should be able to detect an enrichment of positive selection signals around the index SNPs, acting on the risk allele. We have applied three classes of test, based on (1) site frequency spectrum (Tajima’s D,33 Fay and Wu’s H,34 and Nielsen et al.’s CLR35); (2) haplotype characteristics (the composite of multiple signals [CMS]37 and haplotype diversity); and (3) population differentiation (pairwise FST).44 Although not entirely independent, and, strictly, detecting departures from neutrality that may or may not be best explained by positive selection, these tests should in combination capture signals of positive selection that has acted after ∼60,000 years ago when modern humans expanded out of Africa and up until the last few thousand years. Tests (1) and (2) are sensitive mainly to “classic” sweeps where selection acts on a new variant, whereas test (3) is in addition sensitive to selection on standing variation.

Site Frequency Spectrum-Based Analyses

We compared combined p values of the three site frequency spectrum-based statistics in populations of African (AFR: YRI + LWK + ASW), European (EUR: CEU + TSI + GBR + FIN), or East Asian (ASN: CHB + CHS + JPT) ancestry between the set of 64 autosomal index SNPs (windows of 50 or 100 kb centered on each SNP) and matched control regions. The distributions of the statistics for the index SNPs were not different from the matched controls (Figure 1A, Table S3). Therefore, we conclude that there is no overall enrichment of signals of positive selection detectable by these tests around the T2D index SNPs. Because strong LD in the human genome seldom extends farther than 100 kb, these tests should also detect selection on the causative variant.

Figure 1.

Figure 1

Tests for Positive Selection in the Group of 64 Autosomal T2D Susceptibility Loci

Tests were based on (A) the site frequency spectrum (Tajima’s D, Fay and Wu’s H, and Nielsen et al.’s CLR), (B) the composite of multiple signals (CMS), (C) haplotype diversity, or (D) population differentiation (pairwise FST) in a window of 25 kb on either side of the index SNP, except haplotype diversity, which used a window of 5 kb on each side. In each test, we examined either the combined set of loci (All) or subsets based on biological function (BF), the ancestral/derived status of the risk allele (RAS), or the effect size (ES), in each population or population pair. The results are summarized as the combined –log10 of the p value; the threshold for significance after Bonferroni correction is shown by the dotted line.

We also found no enrichment of positive selection signals after stratification of these SNPs by allelic status of the risk allele, effect size, and biological function, except for beta-cell function in Africa, which just exceeded the significance threshold (p = 0.0002). More detailed examination showed that this was due to the signal in Tajima’s D, and not the other two tests, which is better explained by negative selection or population substructure than positive selection (Figure 1A, Table S3). In addition, we further compared statistics for the nearest gene to each index SNP with matched control genes and observed no significant difference either (data not shown).

Haplotype-Based Analyses

We examined CMS values for the 64 autosomal index SNPs in three populations (YRI, CEU, and CHB + JPT) by using data from the third phase of the HapMap project.38 To allow for the possibility that the index SNP may not be the causal SNP, in conjunction with the suggested ability of the CMS test to pinpoint the causal SNP,37 we expanded the analysis to include the SNP with the highest CMS value in a window of 50 or 100 kb centered on each index SNP. We found no significant difference in the values compared with matched control SNPs, either overall or in the different stratification groups (Figure 1B, Table S3).

In an alternative haplotype-based approach, we examined haplotype diversity in 10 kb windows around the index SNPs compared with matched controls in populations of similar continental origin (AFR, EUR, and ASN), using the 1000 Genomes Phase 1 data set.39 We did not see overrepresentation of reduced haplotype diversity in any continental group compared with genome-wide matched controls, either overall or after stratification (Figure 1C, Table S3).

Population Differentiation-Based Analyses

Pairwise FST values44 between the three continental groups AFR, EUR, and ASN based on the 1000 Genomes Project Phase 1 data set39 were calculated. The distributions of the pairwise FST values between the three continental groups for the 64 autosomal index SNPs are not significantly different from the matched controls as a whole, nor in the stratification groups (Figure 1D, Table S3). Therefore, there is no evidence for enrichment of positive selection of the T2D SNPs based on unusual levels of population differentiation.

Overall, we did not find any enrichment for positive selection acting on the 64 T2D SNPs based on these different tests. Thus, we conclude that there is no evidence to support the thrifty late hypothesis either.

Individual Genes Showing Nominal Evidence of Positive Selection

The analyses above test whether, as a group, there was a consistent signal of selection. However, failure to detect signal across all 64 loci does not necessarily preclude signals at some of the loci when tested individually. We now discuss these individual loci.

To identify outliers in an objective way, we combined (1) the smallest p value from the site frequency spectrum-based tests in the 100 kb window surrounding each index SNP, and the empirical p values from (2) the diversity test and (3) the pairwise FST test as a single combined p value, because these p values are independent.40 Then we applied a significance cut-off of 3.6 × 10−4 to correct for multiple tests as described in the Material and Methods. We found that five index SNPs (rs340874, rs3923113, rs6780569, rs1470579, and rs1552224) at five GWAS loci (near PROX1 [MIM 601546], GRB14 [MIM 601524], UEB2E2 [MIM 602163], IGF2BP2 [MIM 608289], and ARAP1 [MIM 606646]) in African populations, three index SNPs (rs340874, rs1531343, and rs8042680) at three GWAS loci (those near PROX1, HMGA2 [MIM 600698], and PRC1 [MIM 603484]) in European populations, and nine index SNPs (rs10923931, rs11899863, rs7578597, rs3923113, rs10010131, rs1801214, rs896854, rs7903146, and rs8042680) at seven GWAS loci (NOTCH2, THADA [MIM 611800], GRB14, WFS1 [MIM 606201], TP53INP1 [MIM 606185], TCF7L2, and PRC1) in East Asian populations showed significant signals (Figure 2). Three genes showed selection signals in two populations: PROX1 in Africans and Europeans, GRB14 in Africans and Asians, and PRC1 in Europeans and Asians. Among all the selected regions, TCF7L2, NOTCH2, and THADA have been identified as showing signals of selection by previous studies.14–16,19 The other nine regions are newly identified in this study. We failed to pick up the previously reported highly differentiated signals near CDKAL1 (MIM 611259) and HHEX/IDE (MIM 146680), although these were close to the significance threshold (Figure 2).

Figure 2.

Figure 2

Tests for Positive Selection in the 64 Individual Autosomal T2D Susceptibility Loci

Tests were based on the site frequency spectrum (Tajima’s D, Fay and Wu’s H, and Nielsen et al.’s CLR), haplotype diversity, or population differentiation (pairwise FST), and each region or each index SNP was examined separately in each population. The results are summarized as the combined –log10 of the p value; the threshold for significance after Bonferroni correction is shown by the dotted line.

We constructed median-joining networks to visualize the haplotype relationships of the genomic regions in high LD (r2 ≥ 0.95) with these index SNPs and to illustrate the networks for two of these regions, those near NOTCH2 and HMGA2 (Figures 3 and 4). In the NOTCH2 network, selected in East Asians, the YRI haplotypes are diverse and scattered, whereas CHB showed massive expansion of the same haplotype, with multiple independent one- and two-step derivatives of this central haplotype also present (Figure 3), as also in CEU where the combined p value was just below the threshold value. Together, the central haplotype and its derivatives account for 95% of CHB haplotypes and 87% of CEU haplotypes. This pattern supports the hypothesis of positive selection on the central haplotype in the CHB and possibly CEU, but not the YRI, beginning sufficiently long ago for variants of the haplotype to accumulate, approximately 18 ± 5 thousand years ago.41 In the HMGA2 network, there are moderately common haplotypes in all three populations, but the haplotype with the highest population frequency occurs in the CEU (64%, compared with maximum frequencies of 43% in the CHB and 30% in the YRI), along with one- and two-step derivatives (Figure 4). This pattern similarly suggests positive selection, specific to the CEU and beginning about 30 ± 19 thousand years ago. In both cases, the likely selected haplotype carries the protective allele at the index SNP.

Figure 3.

Figure 3

Positive Selection at the NOTCH2 Locus

(A) Region on chromosome 1 that spans NOTCH2 showing GENCODE (v. 14) transcript annotation (blue, protein coding transcripts; green, non-protein coding transcripts) and variant (dbSNP 137, NHLBI Exome SNPs and Indels) tracks (red, nonsynonymous and essential splice site variants; green, synonymous; black, splice site). The NHGRI GWAS Catalog track displays the position of the index SNP (rs10923931) and a highly differentiated variant (rs835574) observed between the AFR and EUR populations in the 1000 Genomes Project. The –log10p of the combined p value (boxed area) were generated from the separate probabilities of Tajima’s D, Fay and Wu’s H, and Nielsen et al.’s CLR. The threshold, represented by the dashed line, incorporates the 5% FDR for each population. Peaks above the cut-off represent regions enriched for positive selection in the CEU and CHB+JPT.

(B) A closer look at the haplotype networks in a region that is in high LD (r2 ≥ 0.95) in CEU (gray area in A). Phased haplotypes were used to make the median joining networks. ENCODE annotation shows regions of open chromatin (Digital DNase I Hypersensitivity Clusters) and chromatin state segmentation in four cell lines (lymphoblastoid [GM12878], human umbilical vein endothelial [HUVEC], liver [HepG2], and human skeletal muscle myoblast [HSMM]). The region is associated with transcription and candidate enhancer regions are observed in HSMM. This is also supported by presence of binding sites for transcription factors (including IRF7) in its vicinity.

Figure 4.

Figure 4

Positive Selection near HMGA2

(A) Region on chromosome 12 upstream of HMGA2 showing GENCODE (v. 14) transcript annotation (blue, protein coding transcripts; green, non-protein coding transcripts) variant (dbSNP 137) and NHGRI GWAS Catalog tracks with the index SNP (rs1531343) highlighted. ENCODE annotation shows regions of open chromatin (Digital DNase I Hypersensitivity Clusters), CpG islands, and chromatin state segmentation in four cell lines (GM12878, HUVEC, HepG2, and HSMM). The region is associated with transcription and candidate enhancer and promoter regions are observed in three cell lines. deCode Recombination Maps tracks show a male specific recombination region (blue). The –log10p of the combined p value (boxed area) show peaks above the cut-off that represent regions enriched for positive selection in the CEU.

(B) A closer look at the haplotype networks in a region that is in high LD (r2 ≥ 0.95) in CEU (gray area in A). Candidate enhancer regions are observed in HUVEC and HSMM cell lines. There is a binding site for the transcription factor IRF7 in their vicinity.

All of the 14 SNPs that showed signatures of positive selection in the combined tests lay on a common haplotype in the selected population but not in other populations; this finding is expected because the tests are powered only to detect selection of a common allele, but provides some additional support for selection. In 8/14, selection was detected on the protective allele in at least one of the populations, or 8/12 if we treat the index SNPs within high LD (r2 > 0.95) as a single event (Table S1). In conclusion, these analyses confirm that there are good candidates for positive selection at a small subset of the T2D-associated loci, although even in this subset, selection is divided between the risk and protective alleles whereas it is predicted to be for the risk allele by the thrifty gene hypothesis.

Discussion

The medical impact of T2D, already a leading cause of death,2 with more than 80% of the mortality occurring in low- and middle-income countries,45 has raised interest in understanding its evolutionary-genetic basis, as well as discovering underlying pathways that may have more immediate medical relevance. Several previous studies based on small numbers of loci have unsuccessfully sought evidence of positive selection on risk alleles14–17,19 as predicted by the thrifty late interpretation of Neel’s hypothesis,9,10 instead finding that a small subset of protective alleles may have been positively selected.

Now, a greatly increased number of index SNPs, and additional resources for detecting positive selection, are available. We have therefore been able to carry out a comprehensive reinvestigation of this question. The outcome was consistent with many of the previous studies: no significant overall enrichment for positive selection on the loci as a group, but evidence for positive selection at a small number of individual loci. This proportion was similar to the level of positive selection found within a set of randomly chosen matched loci (which were chosen blind to evidence for positive selection) and therefore did not provide a signal in the group test. Selection acted on both the risk and protective alleles. We did not, however, find support for unusual levels of population differentiation at T2D index SNPs, reported previously,17,18 despite examining far more SNPs (64 versus 12–16). Of the ten SNPs analyzed in detail by Chen et al.,17 seven were included in our set and two others were in high LD with two of ours. These nine SNPs did show higher FST values than the remaining SNPs in our analyses (p = 0.046, Mann-Whitney U test), confirming that high population differentiation is a feature of these variants, but not of T2D-associated variants in general and perhaps a consequence of the particular ascertainment scheme used in these studies, which enriched for genes with large effect (p = 0.002, Fisher exact test). Although the number of T2D-associated loci is likely to increase in the future, we have examined a sufficiently large sample that the lack of support for the thrifty gene hypothesis seems unlikely to change. We therefore need to consider its implications.

Most of the analyses have been performed in the index SNP, which will usually not be the causative SNP. Would the results be very different if we were able to test the set of causative SNPs directly? This seems unlikely. The index SNP is expected to be in strong LD with the causative allele and to share properties such as location in the genome and frequency. When we attempted to increase the number of potentially causative SNPs in the set by identifying the SNP with the highest functional ranking in each window (Table S1), similar results were obtained. We have presented many of our analyses in terms of the gene nearest to the index SNP, because these names are more memorable than the SNP rs numbers. Most of the index SNPs, and probably most of the causal SNPs, will be regulatory, and thus potentially may act on a distal gene instead of the nearest one. Our analyses are based on the numbers, locations, and properties of the selection signals, except for the stratification of signals according to biological function, and so misidentification of the affected gene would have little effect on our conclusions.

The thrifty gene hypothesis can be interpreted in thrifty early and thrifty late forms, the former corresponding better to Neel’s view, but the latter being the more common interpretation by geneticists. We have suggested that according to the thrifty early hypothesis, risk variants would have been selected for during most of human evolutionary history and thus would be ancestral alleles shared by prehominins and most present-day populations. If the human-chimpanzee ancestor had not been thrifty, but early hominins had been, thrifty alleles would be human specific and appear as fixed human-chimpanzee differences. In order for such differences to be detected in association studies, common variation would need to have arisen subsequently. In this case, thrifty alleles would generally be derived with respect to other great apes, and our results also exclude this scenario. Neither the current nor previous studies have found genetic support for the thrifty late hypothesis, although further studies in populations with histories of extreme famine-abundance cycles within the last few thousand years would still be of interest.

A finding common to several studies has been the evidence for positive selection on a proportion of protective alleles. It has been suggested that the relatively low frequency of T2D in populations of European ancestry, compared to non-Europeans with similar lifestyles, might reflect selection for diabetes resistance in Europeans.46 Indeed, all three index SNPs showing positive selection in Europeans show selection on the protective allele (Table S1). Nevertheless, similar selection is seen at three of seven independent index SNPs in East Asians and two out of five in Africans. Overall, it is far from clear that these variants have been selected because of their small influence on T2D risk; selection may well have acted on other aspects of their biology.

In conclusion, we have performed the largest investigation of the genetic basis of the thrifty gene hypothesis thus far and find that it is not supported. It therefore seems that selection for thrifty alleles earlier in human history, together with a change in lifestyle, do not account for the levels of T2D susceptibility in current populations. It is more likely that T2D risk variants have been, for most of human evolutionary history, effectively neutral.

Acknowledgments

This work was supported by The Wellcome Trust (grant 098051 to Q.A., Y.C., K.P., V.C., L.P., G.R.S.R., C.T.-S., E.Z., and Y.X.; grants 098381 and WT090367MA to M.I.M.) and by Framework 7: ENGAGE (HEALTH-F4-2007- 201413 to M.I.M.).

Supplemental Data

Document S1. Tables S2 and S3
mmc1.pdf (116.1KB, pdf)
Table S1. 65 Index SNPs with Additional Information

The LD region was defined by r2 ≥ 0.95. Selected population was as identified in Figure 2. Selected haplotype frequency was identified by network analysis, including one- and two-step neighbors.

mmc2.xlsx (23.6KB, xlsx)

Web Resources

The URLs for data presented herein are as follows:

References

  • 1.World Health Organization . WHO Press; Geneva: 2009. Global Health Risks. Mortality and Burden of Disease Attributable to Selected Major Risks. [Google Scholar]
  • 2.World Health Organization . WHO Press; Geneva: 2011. Global Status Report on Noncommunicable Diseases 2010. [Google Scholar]
  • 3.Stumvoll M., Goldstein B.J., van Haeften T.W. Type 2 diabetes: principles of pathogenesis and therapy. Lancet. 2005;365:1333–1346. doi: 10.1016/S0140-6736(05)61032-X. [DOI] [PubMed] [Google Scholar]
  • 4.Maller J.B., McVean G., Byrnes J., Vukcevic D., Palin K., Su Z., Howson J.M., Auton A., Myers S., Morris A., Wellcome Trust Case Control Consortium Bayesian refinement of association signals for 14 loci in 3 common diseases. Nat. Genet. 2012;44:1294–1301. doi: 10.1038/ng.2435. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jia H., Zack M.M., Thompson W.W. The effects of diabetes, hypertension, asthma, heart disease, and stroke on quality-adjusted life expectancy. Value Health. 2013;16:140–147. doi: 10.1016/j.jval.2012.08.2208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.La Vignera S., Condorelli R., Vicari E., D’Agata R., Calogero A.E. Diabetes mellitus and sperm parameters. J. Androl. 2012;33:145–153. doi: 10.2164/jandrol.111.013193. [DOI] [PubMed] [Google Scholar]
  • 7.Whitworth K.W., Baird D.D., Stene L.C., Skjaerven R., Longnecker M.P. Fecundability among women with type 1 and type 2 diabetes in the Norwegian Mother and Child Cohort Study. Diabetologia. 2011;54:516–522. doi: 10.1007/s00125-010-2003-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Livshits A., Seidman D.S. Fertility issues in women with diabetes. Womens Health (Lond. Engl.) 2009;5:701–707. doi: 10.2217/whe.09.47. [DOI] [PubMed] [Google Scholar]
  • 9.Neel J.V. The “thrifty genotype” in 1998. Nutr. Rev. 1999;57:S2–S9. doi: 10.1111/j.1753-4887.1999.tb01782.x. [DOI] [PubMed] [Google Scholar]
  • 10.Neel J.V. Diabetes mellitus: a “thrifty” genotype rendered detrimental by “progress”? Am. J. Hum. Genet. 1962;14:353–362. [PMC free article] [PubMed] [Google Scholar]
  • 11.Jobling M.A., Hurles M.E., Tyler-Smith C. Garland Science; Abingdon: 2004. Human Evolutionary Genetics: Origins, Peoples and Disease. [Google Scholar]
  • 12.Li H., Durbin R. Inference of human population history from individual whole-genome sequences. Nature. 2011;475:493–496. doi: 10.1038/nature10231. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Grant S.F.A., Thorleifsson G., Reynisdottir I., Benediktsson R., Manolescu A., Sainz J., Helgason A., Stefansson H., Emilsson V., Helgadottir A. Variant of transcription factor 7-like 2 (TCF7L2) gene confers risk of type 2 diabetes. Nat. Genet. 2006;38:320–323. doi: 10.1038/ng1732. [DOI] [PubMed] [Google Scholar]
  • 14.Helgason A., Pálsson S., Thorleifsson G., Grant S.F., Emilsson V., Gunnarsdottir S., Adeyemo A., Chen Y., Chen G., Reynisdottir I. Refining the impact of TCF7L2 gene variants on type 2 diabetes and adaptive evolution. Nat. Genet. 2007;39:218–225. doi: 10.1038/ng1960. [DOI] [PubMed] [Google Scholar]
  • 15.Southam L., Soranzo N., Montgomery S.B., Frayling T.M., McCarthy M.I., Barroso I., Zeggini E. Is the thrifty genotype hypothesis supported by evidence based on confirmed type 2 diabetes- and obesity-susceptibility variants? Diabetologia. 2009;52:1846–1851. doi: 10.1007/s00125-009-1419-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Klimentidis Y.C., Abrams M., Wang J., Fernandez J.R., Allison D.B. Natural selection at genomic regions associated with obesity and type-2 diabetes: East Asians and sub-Saharan Africans exhibit high levels of differentiation at type-2 diabetes regions. Hum. Genet. 2011;129:407–418. doi: 10.1007/s00439-010-0935-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Chen R., Corona E., Sikora M., Dudley J.T., Morgan A.A., Moreno-Estrada A., Nilsen G.B., Ruau D., Lincoln S.E., Bustamante C.D., Butte A.J. Type 2 diabetes risk alleles demonstrate extreme directional differentiation among human populations, compared to other diseases. PLoS Genet. 2012;8:e1002621. doi: 10.1371/journal.pgen.1002621. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Corona E., Chen R., Sikora M., Morgan A.A., Patel C.J., Ramesh A., Bustamante C.D., Butte A.J. Analysis of the genetic basis of disease in the context of worldwide human relationships and migration. PLoS Genet. 2013;9:e1003447. doi: 10.1371/journal.pgen.1003447. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Ségurel L., Austerlitz F., Toupance B., Gautier M., Kelley J.L., Pasquet P., Lonjou C., Georges M., Voisin S., Cruaud C. Positive selection of protective variants for type 2 diabetes from the Neolithic onward: a case study in Central Asia. Eur. J. Hum. Genet. 2013;21:1146–1151. doi: 10.1038/ejhg.2012.295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Abecasis G.R., Altshuler D., Auton A., Brooks L.D., Durbin R.M., Gibbs R.A., Hurles M.E., McVean G.A., 1000 Genomes Project Consortium A map of human genome variation from population-scale sequencing. Nature. 2010;467:1061–1073. doi: 10.1038/nature09534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Voight B.F., Scott L.J., Steinthorsdottir V., Morris A.P., Dina C., Welch R.P., Zeggini E., Huth C., Aulchenko Y.S., Thorleifsson G., MAGIC investigators. GIANT Consortium Twelve type 2 diabetes susceptibility loci identified through large-scale association analysis. Nat. Genet. 2010;42:579–589. doi: 10.1038/ng.609. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Dupuis J., Langenberg C., Prokopenko I., Saxena R., Soranzo N., Jackson A.U., Wheeler E., Glazer N.L., Bouatia-Naji N., Gloyn A.L., DIAGRAM Consortium. GIANT Consortium. Global BPgen Consortium. Anders Hamsten on behalf of Procardis Consortium. MAGIC investigators New genetic loci implicated in fasting glucose homeostasis and their impact on type 2 diabetes risk. Nat. Genet. 2010;42:105–116. doi: 10.1038/ng.520. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Qi L., Cornelis M.C., Kraft P., Stanya K.J., Linda Kao W.H., Pankow J.S., Dupuis J., Florez J.C., Fox C.S., Paré G., Meta-Analysis of Glucose and Insulin-related traits Consortium (MAGIC) Diabetes Genetics Replication and Meta-analysis (DIAGRAM) Consortium Genetic variants at 2q24 are associated with susceptibility to type 2 diabetes. Hum. Mol. Genet. 2010;19:2706–2715. doi: 10.1093/hmg/ddq156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Tsai F.J., Yang C.F., Chen C.C., Chuang L.M., Lu C.H., Chang C.T., Wang T.Y., Chen R.H., Shiu C.F., Liu Y.M. A genome-wide association study identifies susceptibility variants for type 2 diabetes in Han Chinese. PLoS Genet. 2010;6:e1000847. doi: 10.1371/journal.pgen.1000847. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Yamauchi T., Hara K., Maeda S., Yasuda K., Takahashi A., Horikoshi M., Nakamura M., Fujita H., Grarup N., Cauchi S. A genome-wide association study in the Japanese population identifies susceptibility loci for type 2 diabetes at UBE2E2 and C2CD4A-C2CD4B. Nat. Genet. 2010;42:864–868. doi: 10.1038/ng.660. [DOI] [PubMed] [Google Scholar]
  • 26.Shu X.O., Long J., Cai Q., Qi L., Xiang Y.B., Cho Y.S., Tai E.S., Li X., Lin X., Chow W.H. Identification of new genetic risk variants for type 2 diabetes. PLoS Genet. 2010;6:e1001127. doi: 10.1371/journal.pgen.1001127. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Unoki H., Takahashi A., Kawaguchi T., Hara K., Horikoshi M., Andersen G., Ng D.P., Holmkvist J., Borch-Johnsen K., Jørgensen T. SNPs in KCNQ1 are associated with susceptibility to type 2 diabetes in East Asian and European populations. Nat. Genet. 2008;40:1098–1102. doi: 10.1038/ng.208. [DOI] [PubMed] [Google Scholar]
  • 28.Yasuda K., Miyake K., Horikawa Y., Hara K., Osawa H., Furuta H., Hirota Y., Mori H., Jonsson A., Sato Y. Variants in KCNQ1 are associated with susceptibility to type 2 diabetes mellitus. Nat. Genet. 2008;40:1092–1097. doi: 10.1038/ng.207. [DOI] [PubMed] [Google Scholar]
  • 29.Paten B., Herrero J., Beal K., Fitzgerald S., Birney E. Enredo and Pecan: genome-wide mammalian consistency-based multiple alignment with paralogs. Genome Res. 2008;18:1814–1828. doi: 10.1101/gr.076554.108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Green R.E., Krause J., Briggs A.W., Maricic T., Stenzel U., Kircher M., Patterson N., Li H., Zhai W., Fritz M.H. A draft sequence of the Neandertal genome. Science. 2010;328:710–722. doi: 10.1126/science.1188021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Reich D., Green R.E., Kircher M., Krause J., Patterson N., Durand E.Y., Viola B., Briggs A.W., Stenzel U., Johnson P.L. Genetic history of an archaic hominin group from Denisova Cave in Siberia. Nature. 2010;468:1053–1060. doi: 10.1038/nature09710. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Meyer M., Kircher M., Gansauge M.T., Li H., Racimo F., Mallick S., Schraiber J.G., Jay F., Prüfer K., de Filippo C. A high-coverage genome sequence from an archaic Denisovan individual. Science. 2012;338:222–226. doi: 10.1126/science.1224344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Tajima F. Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics. 1989;123:585–595. doi: 10.1093/genetics/123.3.585. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Fay J.C., Wu C.-I. Hitchhiking under positive Darwinian selection. Genetics. 2000;155:1405–1413. doi: 10.1093/genetics/155.3.1405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Nielsen R., Williamson S., Kim Y., Hubisz M.J., Clark A.G., Bustamante C. Genomic scans for selective sweeps using SNP data. Genome Res. 2005;15:1566–1575. doi: 10.1101/gr.4252305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Ayub Q., Yngvadottir B., Chen Y., Xue Y., Hu M., Vernes S.C., Fisher S.E., Tyler-Smith C. FOXP2 targets show evidence of positive selection in European populations. Am. J. Hum. Genet. 2013;92:696–706. doi: 10.1016/j.ajhg.2013.03.019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Grossman S.R., Shlyakhter I., Karlsson E.K., Byrne E.H., Morales S., Frieden G., Hostetter E., Angelino E., Garber M., Zuk O. A composite of multiple signals distinguishes causal variants in regions of positive selection. Science. 2010;327:883–886. doi: 10.1126/science.1183863. [DOI] [PubMed] [Google Scholar]
  • 38.Altshuler D.M., Gibbs R.A., Peltonen L., Altshuler D.M., Gibbs R.A., Peltonen L., Dermitzakis E., Schaffner S.F., Yu F., Peltonen L., International HapMap 3 Consortium Integrating common and rare genetic variation in diverse human populations. Nature. 2010;467:52–58. doi: 10.1038/nature09298. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Abecasis G.R., Auton A., Brooks L.D., DePristo M.A., Durbin R.M., Handsaker R.E., Kang H.M., Marth G.T., McVean G.A., 1000 Genomes Project Consortium An integrated map of genetic variation from 1,092 human genomes. Nature. 2012;491:56–65. doi: 10.1038/nature11632. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Fisher R.A. Twelfth Edition. Oliver and Boyd; Edinburgh: 1954. Statistical Methods for Research Workers; p. 356. [Google Scholar]
  • 41.Bandelt H.J., Forster P., Röhl A. Median-joining networks for inferring intraspecific phylogenies. Mol. Biol. Evol. 1999;16:37–48. doi: 10.1093/oxfordjournals.molbev.a026036. [DOI] [PubMed] [Google Scholar]
  • 42.Harrow J., Frankish A., Gonzalez J.M., Tapanari E., Diekhans M., Kokocinski F., Aken B.L., Barrell D., Zadissa A., Searle S. GENCODE: the reference human genome annotation for The ENCODE Project. Genome Res. 2012;22:1760–1774. doi: 10.1101/gr.135350.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Bernstein B.E., Birney E., Dunham I., Green E.D., Gunter C., Snyder M., ENCODE Project Consortium An integrated encyclopedia of DNA elements in the human genome. Nature. 2012;489:57–74. doi: 10.1038/nature11247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Weir B.S., Cockerham C.C. Estimating F-statistics for the analysis of population structure. Evolution. 1984;38:1358–1370. doi: 10.1111/j.1558-5646.1984.tb05657.x. [DOI] [PubMed] [Google Scholar]
  • 45.Mathers C.D., Loncar D. Projections of global mortality and burden of disease from 2002 to 2030. PLoS Med. 2006;3:e442. doi: 10.1371/journal.pmed.0030442. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Diamond J. The double puzzle of diabetes. Nature. 2003;423:599–602. doi: 10.1038/423599a. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Tables S2 and S3
mmc1.pdf (116.1KB, pdf)
Table S1. 65 Index SNPs with Additional Information

The LD region was defined by r2 ≥ 0.95. Selected population was as identified in Figure 2. Selected haplotype frequency was identified by network analysis, including one- and two-step neighbors.

mmc2.xlsx (23.6KB, xlsx)

Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES