Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Apr 18.
Published in final edited form as: Mol Ecol. 2013 Nov 26;22(24):6131–6148. doi: 10.1111/mec.12562

Detecting adaptive trait loci in non-model systems: divergence or admixture mapping?

Jacob E Crawford 1,*, Rasmus Nielsen 2
PMCID: PMC12007813  NIHMSID: NIHMS741499  PMID: 24128338

Abstract

Mapping adaptive trait loci (ATL) underlying ecological divergence is an essential step towards understanding the processes that generate phenotypic diversity. Technological advances have made it possible to sequence exomes in non-model systems, providing an efficient means of analyzing functional genetic variants. Divergence scans of genetic markers for outlier loci, or ‘divergence mapping’, have been used to map locally adapted genes, but this approach is likely to be underpowered when background divergence is elevated. Genotype-phenotype association tests in admixed populations, or ‘admixture mapping’, may provide a useful approach for mapping locally adapted loci when neutral divergence is high. To determine the power and limits of divergence mapping, we simulated exomes containing a single ATL across two parental populations of varying neutral divergence, estimated divergence, and quantified the power to identify the ATL. We found that divergence mapping had very high power when background FST is less than 0.2, but decreased dramatically above this level. To evaluate the utility of admixture mapping, we simulated exomes from admixed populations, then simulated phenotypes, conducted genotype-phenotype association tests, and found that even two generations of random mating after admixture could provide high mapping power in scenarios where pure divergence mapping was ineffective (FST = 0.35). Moreover, admixture mapping had high power across all levels of divergence after 20 generations since admixture. Together with high-throughput exome sequencing, admixture mapping could be used to map ATL in systems such as Heliconius butterflies or Gryllus crickets when experimental design and analytical approach are chosen accordingly.

Keywords: Adaptation, non-model systems, hybridization, admixture

INTRODUCTION

Populations colonizing novel habitats or niches often diverge phenotypically and genetically to reach new fitness optima. Understanding the mechanisms and processes that drive such ecological differentiation is a longstanding goal of evolutionary biology (Fisher 1930; Wright 1931, 1932). Examples of phenotypic divergence among populations can be found throughout the natural world. In one potent example, the mimic poison frog species Ranitomeya imitator has diverged into multiple parapatric color morphs that bear a strong resemblance to distinct toxic model species in Peru (Twomey et al. 2013). Such phenotypic divergence among allopatric or parapatric populations has also been observed in Heliconius butterflies (Turner 1981), humans (Moore 2001), Peromyscus mice (Sumner 1926), and Ensatina salamanders (Stebbins 1949; Brown 1974). Phenotypic divergence can also occur as a result of differential niche specialization in relative sympatry as observed in Rhagoletis fruitflies (Bush 1966, 1975), Acyrthosiphon pea aphids (Via 1991), and Spalax blind mole rats (Hadid et al. 2013). Pinpointing the gene(s) underlying such phenotypic divergence, or ‘adaptive trait loci’ (ATL), is an important step in understanding the genetic basis of adaptation and causal links between genotype and phenotype.

Historically, quantitative trait locus (QTL) mapping has been used to map loci underlying phenotypic variation. QTL mapping has been used to successfully map a variety of traits varying in genomic architecture in a number of both ecological and agricultural systems (Lynch & Walsh 1998; Erickson et al. 2004; Bradshaw et al. 2012; Papa et al. 2013). QTL mapping is laborious since it requires a linkage map and genotyping of hundreds of individuals from laboratory crosses. It also ultimately provides relatively low genomic resolution owing to the paucity of recombination events available in one or two generations of laboratory crosses. Although valuable information regarding genomic architecture of a trait and the approximate location of the gene(s) underlying a trait or traits can be obtained from QTL mapping, identifying specific causal variants or even genes is often impossible, especially if no a priori list of candidates fall within the QTL.

In recent years, another approach to mapping ATL coding for adaptively diverged traits has become favored involving either model-based or model-free scans for exceptionally divergent loci relative to the genome average (Storz 2005). It is now well established from both theory and empirical evidence that adaptation to novel environmental factors across heterogeneous environments is a major evolutionary force generating and maintaining phenotypic divergence (Fisher 1930; Haldane 1930; Lewontin & Krakauer 1973; Kaufman 1974; Via 1991; Yi et al. 2010). Although stochastic effects of genetic drift, variable mutation rates, or chromosomal structure can cause molecular divergence among populations to be highly heterogeneous across the genome even in the absence of natural selection (Nosil et al. 2009), genes that control ecologically diverged traits are expected to show especially high molecular differentiation among populations relative to the genome average (Fisher 1930; Haldane 1930; Lewontin & Krakauer 1973; Charlesworth et al. 1997). This expectation has motivated the relatively straightforward mapping approach of collecting allele frequency data from two phenotypically differentiated populations, scanning the genome of interest for differentiated windows or markers (hereafter referred to as ‘divergence mapping’), and examining the empirical distribution to identify outlier loci that are likely to represent locally adapted loci or loci linked to adaptive loci (Lewontin & Krakauer 1973; Akey et al. 2002; Beaumont & Balding 2004; Hohenlohe et al. 2010). Such divergence mapping has been used in an attempt to localize ATL involved in wing patterning in Heliconius butterflies (Nadeau et al. 2012), adaptation to freshwater in sticklebacks (Hohenlohe et al. 2010), host plant preference in walking-sticks (Nosil et al. 2008), and shell morphology in intertidal snails (Wilding et al. 2001), among others. The distinction between adaptive and neutral divergence should be most apparent in the early stages of population subdivision since neutral genetic differentiation is expected to be slow to develop following isolation between populations (Kimura 1957, 1962), but the point at which neutral divergence makes it difficult to identify adaptive loci has not been well established, although some comparisons have been made among model-based approaches for detecting outliers (Pérez-Figueroa et al. 2010). The ability to use signatures of excess differentiation is also complicated under more complex genetic structure models (Excoffier et al. 2009; Fourcade et al. 2013).

When background genetic divergence is relatively high, simple divergence mapping will likely be underpowered, but an extension of classical QTL mapping (Lynch & Walsh 1998; Erickson et al. 2004) may be useful under these conditions. An alternative approach exploiting recombination in natural hybrid populations admixed from divergent parental populations may provide a useful and naturally occurring alternative to laboratory crosses for mapping ATL in highly divergent systems. Since genetically divergent populations are differentiated by both adaptive and neutral loci throughout the genome, all differentially fixed loci will be highly correlated (i.e. in linkage disequilibrium) in F1 hybrids in the admixed population as they would be in lab crosses. As the admixed population persists, however, recombination will reduce linkage disequilibrium between neutral and locally adapted loci, and thus the adaptive phenotype. After some number of generations, the divergently adapted locus is expected to be the only locus with strong phenotypic association and thus could be mapped using genotype-phenotype association tests akin to those used in QTL mapping. If the hybrid population is old enough and phenotyped hybrids can be sampled in sufficient numbers, relatively simple association tests may provide some power to identify candidate adaptive loci. A similar admixture mapping approach has been used successfully to map disease alleles in human populations admixed from parental populations that differ in frequency of disease phenotypes. This method exploits natural recombination between chromosomes of different ancestry so that ancestry of relatively small admixture tracts can be correlated with disease phenotypes to identify genomic regions harboring disease alleles (Winkler et al. 2010). The situation we will consider is different from the one usually considered in human genetics, in that we will not assume the availability of a physical map of the genome. The standard Hidden Markov Model methods used in human genetics for admixture mapping (Patterson et al. 2004) are, therefore, not applicable.

In our case, we will approach the problem as a classical genetic association analysis in a natural hybrid population with linkage disequilibrium that is decreasing with time since admixture. Natural systems such as Heliconius butterflies and Ensatina salamanders are perfectly suited for this kind of approach since they are composed of many phenotypically diverged populations as well as hybrid populations (Beltrán et al. 2007; Mallet et al. 2007; Pereira & Wake 2009). Although similar approaches have been applied in candidate gene frameworks (Counterman et al. 2010; Reed et al. 2011), the utility of an exome-wide approach that queries the full set of coding sequences in a genome has not been explored in practice or in theory for non-model systems at this level of divergence, and it is not yet clear under what circumstances this approach has power.

In addition to choosing the analytical approach with the best analytical power, experimental design must also consider genetic data collection method. Full genome assembly and re-sequencing provides complete information and may be desirable when feasible, but resources are often insufficient for such endeavors and alternative reduced representation sequencing approaches may provide a more economical and efficient approach. One such strategy is transcriptome-based exome capture and sequencing. In this method, the transcriptome is sequenced and assembled de novo and used to design an exome capture array that can be used to enrich sheared genomic DNA for molecules containing exonic sequence (e.g. Bi et al. 2012). In addition to being economical and efficient, exome sequencing is attractive relative to alternative anonymous targeting strategies such as RAD-seq (Hohenlohe et al. 2010) since exome sequencing provides direct access to protein coding sequence where genetic variants that control Mendelian traits tend to lie (Stenson et al. 2009; Ng et al. 2010). Even if functional variation maps to cis-regulatory sequence as has been observed recently for major effect loci (Counterman et al. 2010; Reed et al. 2011), these sites are likely to be tightly linked to coding variants captured in exome sequencing and should thus be accessible, although exceptions to this expectation can be found (Supple et al. 2013). Exome sequencing also has the advantage of being analytically tractable when neither linkage maps or physical maps are available. However, although the complete exome can be queried efficiently using exome sequencing, a two-step approach involving follow-up studies based on laboratory crosses or molecular techniques will likely be required to verify causal variants among a list of candidates.

In this study, we sought to test the power of divergence mapping and admixture mapping to map adaptive loci under a broad range of population models varying with respect to average genetic divergence and admixture parameters. Specifically we sought to identify the evolutionary conditions under which divergence and admixture mapping have statistical power and when they do not. To do so, we conducted a simulation study where we used a combination of coalescent and forward simulations to generate population samples of diverged and admixed exomes harboring a single locally adapted locus under various population models. We then used divergence mapping and admixture mapping to determine the effect of parental population divergence and age of admixture on the power to distinguish the adapted locus from the neutral background. Our results provide some insight for when divergence mapping may be preferred as a mapping strategy and when admixture mapping may be useful. We discuss our results in the context of experimental design for mapping adaptive loci in non-model systems.

METHODS

The goal of this analysis was to compare the power of divergence-based mapping to that of genotype-phenotype association-based admixture mapping with respect to their power to detect putative adaptive loci across a range of population models. We aimed to model data one may obtain from an exon capture experiment in which the exome of population samples of individual diploids from two or more phenotypically distinct populations are sequenced and subsequently scanned for candidate adaptive loci. We approximated such datasets by simulating 499 ‘exons’ or genes evolving under a standard neutral coalescence process with recombination, and a single gene evolving under strong positive selection representing the adaptive trait locus (ATL). These datasets were then queried using a given mapping approach in order to quantify the proportion of simulations in which the ATL fell in the extreme tail of the empirical distribution.

Divergence mapping

We evaluated the power of divergence mapping in a straightforward framework in which we used coalescent simulations to simulate population samples under a model of two diverging populations and then used divergence statistics to determine the power to identify an ATL. First, we simulated the neutral exome background by simulating independent (free recombination between genes), neutrally evolving genes (Figure 1). Second, we simulated a single gene evolving under strong positive selection in just one deme, but with otherwise similar demographic parameters, conditioning on the selected site reaching fixation in the adapted population. We conducted all coalescent simulations using the coalescent simulator msms (Ewing & Hermisson 2010). Coalescent simulations were parameterized in order to generate data that corresponds to populations of small mammals such as Mus musculus house mice (Salcedo et al. 2007; Geraldes et al. 2011). For genes evolving under the neutral model, we simulated 100 chromosomes (n = 25 diploids per population) evenly distributed among two equally sized (Ne = 25,000) and reproductively isolated demes. These demes were then joined (going backward in time) after tsplit generations ranging from 500 to 100,000 (Figure 1), resulting in mean per-gene FST values per gene ranging from 0.01 to 0.64 (total of 10 divergence levels). For each gene, we simulated 2 kilobase (kb) regions with per base per generation mutation (θ) and recombination (ρ) rates of 10−8 and 10−7, respectively. To simulate an adaptive scenario where the selected sweep immediately follows a move into a novel environment, we simply added selection to the same neutral population model such that positive selection was imposed immediately following the population split (0.99×tsplit) on a new mutation (frequency of 1/2Ne). We set the population scaled fitness (2Nes or α) of the mutant homozygote (αAA) to 3000 and that of the heterozygote (αAa) to 2000 relative to the ancestral homozygote (i.e. αaa = 0). These values were arbitrarily chosen based on empirical tests in order to ensure consistent fixation of the selected site. We excluded any additional mutation at the selected site and conditioned on the selected mutation not being lost due to drift (-SFC flag in msms).

Figure 1.

Figure 1

Diagram of population model and simulation approach. Relevant model parameters are indicated. Each 2kb gene simulated independently. For divergence mapping model, n = 50 for each population (P1 and P2). For admixture mapping model, n = 250 chromosomes for each parental population (P1 and P2) and n = 500 for the admixed population. The selective sweep was always limited to P1 and tsweep was fixed at 0.99×tsplit. See Methods for details on number of replicates per model and number of models.

To measure the power to identify the ATL with divergence mapping, we measured allele frequency divergence at each SNP as well as for each gene using two statistics and quantified the proportion of iterations where the selected locus was either the most diverged in the dataset or in the top 10 of a ranked list. First, we used Weir and Cockerham’s unbiased estimator of FST (Weir & Cockerham 1984) to measure differences in allele frequency, as implemented in R by Eva Chan (www.evachan.org). We calculated FST per site as well as a multi-locus FST for each gene. For both divergence and admixture mapping (described below), we conducted the analyses in two steps. We first conducted a small, ‘chromosome-level’ scale analysis of 500 genes in order to explore a large range of conditions. We then scaled up to full exome-scale analysis with 12,500 genes, but tested only a limited subset of conditions for computation efficiency. For example, we tested the power of divergence mapping at the ‘chromosome’-level scale where the selected locus was compared to a background of 500 neutral genes, as well as at an ‘exome’-level scale that included 12,500 neutral genes. For each exome-level replicate, we calculated FST for the selected locus as well as 12,500 neutral background genes randomly selected without replacement from a total pool of 500,000 replicate simulated genes for each level of divergence.

This test was replicated 500 times for each level of divergence. We also applied a Χ2 test of homogeneity (chi.test function in R) to allele counts for each SNP, comparing observed allele counts in each population to their expectations under the null hypothesis of identical allele frequencies in the two populations. To obtain a p-value for each gene for the Χ2 test, we used Fisher’s combined p-value (Fisher 1932). Since Χ2 contingency tests have little power when cell counts are low, we restricted this analysis to SNPs with a minor allele frequency of at least 0.05 combined for the two populations. We conducted the same ranking comparison as for the FST based test, although only at the chromosome-level.

We also explored the possible effects of a population bottleneck on the power to identify adaptive loci using such a divergence mapping approach. To do so, we simply added a single population bottleneck into the neutral simulation framework while the simulations for the selected locus remained unchanged. This test was conducted using a background set of 500 neutral genes and was replicated 200 times for each bottleneck model. We tested the effects of three bottleneck models that differed in the severity of population reduction (β) including 10-fold, 100-fold, and 1,000-fold reductions. For all models, we set the duration of the bottleneck to 0.1×tsplit. In order to make these simulations comparable to the equilibrium model, we adjusted the number of generations for each bottleneck model such that the level of divergence (mean FST) at the time of sampling would be consistent with equilibrium models. We then calculated FST, ranked both SNP values and gene values, and conducted the same comparisons as above. We did not apply the contingency test here since it performed similarly or less well under no-bottleneck scenarios.

Admixture forward simulations

We modeled admixture using a combination of coalescent and forward simulations. We first simulated ‘parental’ populations with varying levels of genetic differentiation using coalescent simulations identical in structure and parameterization to those used for divergence mapping analysis except that we simulated 250 chromosomes per deme and varied divergence (mean FST) from 0.01 to 0.9. To approximate admixture among chromosomes, 499 neutral chromosomal segments (genes) and one positively selected gene were assembled into a single contiguous, linear synthetic chromosome with the selected locus falling in the center. We then simulated the population forward in time using a standard Wright-Fisher model with recombination, assuming no selection on the adapted locus in the admixed population. In generation one, 500 chromosomes were sampled with replacement from the parental populations generated from coalescent simulations to generate 250 diploid ‘genomes’ with admixture proportion equal to 0.5. We then modeled recombination as a Poisson process with a rate of λ = 1 cross-overs per individual per generation, which corresponds roughly to the rate of recombination observed per chromosome in most organisms, including laboratory crosses of Drosophila (Kulathinal et al. 2008). Cross-over events were located between genes and locations were drawn at random, without replacement, from a distribution uniform on all inter-genic segments and assuming equal distance between genes. The resulting set of recombined chromosomes was then used as input for the next generation, again sampling 500 chromosomes with replacement for each generation. We evolved 200 independent admixed ‘populations’ over 50 generations for each level of divergence between parental populations.

We then quantified the power to identify an ATL after a given number of generations since admixture. From the admixed populations simulated above, we conducted sub-sampling to emulate a sampling scheme designed to maximize the power of association tests by ensuring equal representation of each phenotype. To do so for a Mendelian dominant trait, for example, we sub-sampled 40 diploids from each replicate admixed population at each specified number of generations since admixture, conditional on the frequency of each phenotype equaling 0.5.

At multiple time-points representing a range of generations since admixture, we simulated phenotypes for each diploid, subsampled as described above, and tested each SNP for an association between diploid genotypes and phenotype. We simulated both a Mendelian, dominant trait as well as a quantitative, additive trait. To simulate a Mendelian trait, we randomly assigned chromosomes to individual diploids, assumed a dominant genetic model, and assigned phenotypes 1, 2 and 2 to genotypes aa, Aa, and AA, respectively, where A is the mutated allele. We applied two statistical methods to test for genotype-phenotype associations. First we applied the Cochran-Armitage trend test (Armitage 1955), a modified χ2 test that incorporates specific ordering of the categories, as implemented in the prop.trend.test function in R (R Development Core Team 2011) with the score parameter set to {0,1,1} to reflect dominance in the genetic model. We also used a χ2 test of homogeneity (chi.test in R) applied to the two-by-two contingency table of allele counts and phenotypes. For both statistical tests, we obtained a per-SNP p-value as well as per-gene p-values by combining per-SNP values using Fisher’s combined p-value. Since these statistical tests are underpowered when applied to rare genetic variants (e.g. Li & Leal 2008), we applied the tests only to variants segregating at a frequency of 0.2 or greater across both samples.

To obtain exome-level estimates of statistical power, we expanded the testing framework to include 25 additional samples of ‘chromosomes’ representing a total of 12,475 neutral genes in addition to the 499 included on the ‘chromosome’ harboring the selected locus. The additional ‘chromosomes’ were randomly selected without replacement from the total pool of simulated admixed populations. Exome-level analysis was conducted on only a subset of the generation times using sample sizes ranging from n=20 to n=100 and equal representation of Mendelian phenotypes. We conducted the same ranking scheme as before and quantified the proportion of iterations in which the selected locus fell within the top 10 or was top ranked.

For a limited set of population models, we also simulated an additive, quantitative trait phenotype with varying levels of narrow-sense heritability. To do so, we modeled the population phenotypic distribution in the following way. We first calculated the additive genetic variance for each sample by assuming no dominance as VA = 2pqa2, where p and q are sample allelic frequencies, and a is the allelic effect. We arbitrarily set the allelic effect to 10 and adjust the phenotypic variability (Vp) to obtain the desired narrow sense heritability, using the equation

h2=VAVP

and solving for VP. We modeled VP as a weighted mixture of genotype specific Gaussian distributions with equal variances such that

VP=σi2¯+μi2¯-μ¯2

where σi2¯ is the mean of the variances, μi2¯ is the mean of squared means, and μ̄2 is the square of the mean phenotypic value. Under the assumption of equal variances for all genotypes, and additive effects, we could then solve for the variance component σi2¯, thereby controlling the narrow sense heritability in the simulations. To test for an association between genotype and phenotype under the quantitative trait model, we first used the lm function in R to fit a linear model describing the relationship between phenotypes and genotypes. We then computed an analysis of variance for this linear model using the anova function in R and recorded the p-value for each SNP. As for the Mendelian trait tests, we also obtained a per-gene p-value using Fisher’s combined p-value. This analysis was limited to single ‘chromosome’-level represented by 499 neutral background genes.

RESULTS

Divergence mapping

We sought to identify evolutionary conditions under which divergence mapping between two phenotypically diverged populations has high analytical power to localize ATL and when it does not. To do so, we simulated exomes containing a single locally adapted locus in related populations that varied with respect to mean background FST and quantified the proportion of simulations in which the selected gene was the most diverged in the dataset or among the top 10 most diverged loci. For computational efficiency, we conducted the analysis at two scales. We conducted ‘exome’-scale analysis for a limited subset of conditions in order to obtain exome-level estimates of statistical power. We also conducted a variety of analyses under a broader range of conditions on a smaller, ‘chromosome’-level scale. In all cases, we find that the power curves are consistent with expectations in that maximal power is achieved when background genetic differentiation is low, but power is minimal when mean divergence is high and thus highly diverged loci are common (Figure 2). When divergence was measured using the FST statistic at the exome-level, we found that the selected SNP could be identified with probability > 0.9 when mean FST was less than approximately 0.1 (Figure 2). However, it remained in the top 10 most diverged in the dataset in 0.9 of the simulations until mean FST exceeded approximately 0.17. Interestingly, the power to identify the selected gene when combining data from multiple SNPs in the gene followed a curve approximately similar to that of a single SNP FST test, suggesting that the additional information added from linked SNPs only provides minimal information, but also does not increase noise sufficiently to reduce power much.

Figure 2.

Figure 2

The power of divergence mapping to recover ATL decreases with increasing background divergence. Divergence was estimated using Weir and Cockerham’s estimator of FST (Weir & Cockerham 1984) for each SNP and each gene. Orange lines indicate the proportion of simulations in which the selected SNP was the most diverged in the dataset (solid) or fell with the top 10 most diverged (dashed) as a function of mean exome-wide FST (n = 103 for each FST level). Grey lines indicate the same comparison but for the selected exon. Each of the 10 levels of divergence (FST = {0.01, 0.09, 0.16, 0.26, 0.35, 0.47, 0.51, 0.59, 0.64, 0.74}) tested was replicated with 500 simulations. Analysis was conducted at the ‘chromosome’-level (A) and ‘exome’-level (B).

We also conducted the same set of comparisons at the chromosome-level using p-values from Χ2 tests of homogeneity as the measure of divergence. Ranking divergence per SNP resulted in power curves very similar to the FST test (chromosome-level) with respect to the divergence levels at which the power decreased to less than 0.9 (Figure S1). However, when the SNP-wise p-values were combined using Fisher’s combined p-value method, power was reduced somewhat relative to both the SNP-based p-value test as well as the gene-based FST comparison. This loss in resolution stems from lower ability to distinguish between genes with a single highly significant association from those with several SNPs with slightly less extreme p-values, both of which result in similarly extreme combined p-value. Fisher’s combined p-value is likely not the most efficient way of combining information from multiple SNPs.

Thus far we have examined population models with relatively simple demographic histories, but non-equilibrium demographic effects could impact the power to identify ATL. A reduction in effective population size, following a population bottleneck, will not only increase the overall level of FST, it will also increase the variance in FST among loci. This could potentially reduce mapping power. To investigate this effect, we simulated samples under different bottleneck models and quantified the proportion of simulations in which the selected locus was identified at the chromosome-level. In these simulations, we kept the value of FST constant, to separate the effect of a bottleneck on the variance among loci from the overall effect on mean FST across the genome. We found that adding a bottleneck did not have a large impact on the power to identify adaptive loci using FST-based divergence mapping even when the bottleneck included a 1000-fold size reduction (Figure 3). This lack of effect can be best explained by the fact that while the bottleneck affects the variance among loci, it has a similar strong effect on mean FST. So for a given value of FST, the specifics of the demographic model do not seem to matter much. Although we explored only a limited number of bottleneck models, we did not find a large effect of population bottlenecks on the ability to identify ATL using divergence mapping.

Figure 3.

Figure 3

Population bottlenecks have only minimal effect on power to map adaptive loci using divergence mapping (FST). A single population bottleneck was added to parental population P1. We tested three bottleneck models that varied according to the severity of size reduction (β = fold reduction in Ne) while holding the duration of the bottleneck constant. See legend for Figure 2 for detailed description of figure lines. Each level of divergence is represented by 200 replicates of chromosome-level simulations.

Admixture mapping

In many systems, mapping ATL using classical divergence mapping is not recommended since genetic differentiation between phenotypically distinct subpopulations is elevated beyond the point where divergence mapping has power as shown above. However, admixture in secondary-contact hybrid zones may reduce correlations between ATL and loci not targeted by selection such that phenotype association tests provide a useful approach to identify the gene(s) underlying phenotypic difference(s). In this case, we specifically asked whether admixture mapping in hybrid populations provides high analytical power when background divergence is high and, if so, how many generations of recombination are needed to achieve high power. To do so, we simulated such secondary-contact hybrid populations varying broadly in the degree of differentiation among parental populations and the number of generations since admixture and tested for genotype-phenotype associations. Moreover, we structured the simulations to reflect the reduction in population size typical of most admixed populations relative to the parental populations, since admixture zones are often substantially smaller than the species or population distributions (e.g. Twomey et al. 2013). Reduced population sizes will result in increased genetic drift, which in return will increase the level of Linkage Disequilibrium (LD). High levels of LD will reduce the power to identify specific causal variants as linked variants.

To quantify the power to identify a selected locus in admixed populations, we conducted genotype-phenotype association tests in an initial test of samples of 40 individuals and 499 neutral background genes (chromosome-level). We simulated a dominant Mendelian phenotype according to genotypes at the selected site and conditioned the sample on equal frequencies of each phenotype. In the following, we consider the statistical power to be high when the proportion of simulations satisfying a given criteria exceeds 0.9 (transition to orange in Figure 4). We first tested the power to identify a selected locus using the Cochran-Armitage trend test (Armitage 1955) and found that the exact selected SNP could be identified with probability > 0.9 only when differentiation (FST) between parental populations was very low (FST = 0.01) and few generations had passed since admixture (Figure 4). Consistent with the observation of increased LD with the selected site due to drift in a small admixed population discussed below, the power to distinguish the selected site from other neutral sites decreases with the number of generations since admixture and drops below 0.9 after 30 generations. This pattern is due to the fact that sites in complete LD with the selected site, most of which are in the immediate physical region surrounding the selected SNP, show the same degree of statistical association to phenotype (Figure S2). The degree of this effect would diminish with larger admixed populations. We also considered the power to identify the selected SNP as part of a candidate list (top 10 most significant associations in the dataset) and found that statistical power was very high across a wider range of population models and increased with generations since admixture (Figure 4). At the time of admixture, the power to identify the selected SNP using admixture mapping closely resembles that of divergence mapping; power exceeds 0.9 when differentiation is low (FST < 0.26) but drops precipitously when differentiation increases above this level. After only two generations since admixture, however, the power of admixture mapping increases relative to divergence mapping such that greater than 0.9 can be achieved in systems where differentiation is as high as FST = 0.35. As recombination proceeds, mapping power reaches 0.9 at increasingly high levels of parental population differentiation such that power is greater than 0.9 at very high FST (0.74) after 50 generations of recombination (Figure 4). The power of admixture mapping to identify the causal SNP, does appear to be limited under our simulation conditions when differentiation exceeds 0.74. This limitation is partially due to the fact that we do not include recombination within genes during the admixture forward simulations prohibiting intra-gene LD from diminishing (discussed below).

Figure 4.

Figure 4

Chromosome-level admixture mapping recovers ATL for Mendelian trait from increasingly diverged populations as age of hybrid population increases. Genotype-phenotype association tests were conducted using the Cochran-Armitage trend test (Armitage 1955) at the ‘chromosome’-level with 499 neutral background genes. For each panel, columns indicate the number of generations since admixture and rows indicate the mean background neutral divergence level for that model. Colors indicate the proportion of simulations where the selected SNP (A and C) or the selected gene (B and D) showed the most significant association (A and B) to phenotype or fell within the top 10 (C and D). Note that the color gradient is non-linear and values below 0.5 have same color as 0.5.

We also tested the power of association tests when SNP-wise information is combined across SNPs into gene-wise p-values and found that overall the power to identify the selected gene showed similar patterns as the SNP-wise test, although the effect of time since admixture is more pronounced (Figure 4). If we require that the selected gene show the most significant association in the dataset, only population models with very low levels of differentiation (FST = 0.01) provided high power. The effects of drift can be seen here as well in the reduction of power as time since admixture increases beyond 30 generations. Relaxing the degree of stringency to encompass the top 10 rankings reveals a marked increase in power similar to the SNP-wise candidate list scenario but more dramatically dependent on the number of generations since admixture (Figure 4). Under this testing scenario, the power to identify the selected gene effectively saturates after 20 generations of evolution since admixture such that analytical power is high across all levels of parental differentiation.

To determine whether the specifics of the statistical test employed for the association test affect power, we also performed similar SNP-wise and combined analyses using a Χ2 allelic test of association, which should have more power in an additive model, but less power in a model with dominance. We found very similar patterns as seen with the trend test (Figure S3), suggesting that this approach is largely robust to the statistical approach employed in the association test.

Thus far we have only explored the power to conduct association tests on dominant Mendelian traits with narrow sense heritability (h2) of 1, but many traits are quantitative with non-Mendelian inheritance owing in part to environmental effects and polygenic additive genetic variance. To determine how admixture mapping performs when narrow-sense heritability (h2) is less than 1, we simulated additive quantitative traits ranging in h2 from 0.3 to 1 and conducted association tests in admixed populations as for the Mendelian trait. We first tested an experimental design where we drew a random sample of 100 individuals, selected 20 individuals from each of the phenotypic extremes, and conducted the association test. We tested the power of admixture mapping after 2, 6, and 20 generations since admixture and found that, under this sampling scheme, the power to identify the selected SNP in a candidate list of top 10 strongest associations was similar, qualitatively and quantitatively, between a Mendelian trait and quantitative traits with h2 of at least 0.5 (Figure 5). For traits with h2 equal to 0.3, analytical power was reduced somewhat, but mostly when differentiation between parental populations was high. In general, the power to identify ATL using admixture mapping in highly diverged systems was robust for both Mendelian and quantitative traits. We also tested the power of admixture mapping using a random sampling scheme where we drew a random sample of 40 individuals and expectedly found that the power curves looked very similar to those observed under the random sampling scheme (Figure S4).

Figure 5.

Figure 5

Admixture mapping recovers ATL for quantitative trait with varying levels of narrow-sense heritability. Colors indicate the trait model and narrow-sense heritability (e.g. Quant h2 = 0.5 indicates simulations for quantitative trait with narrow-sense heritability of 0.5). Y-axis indicates the proportion of simulations in which the selected SNP falls within the top 10 most significant associations with phenotype. Panels show power curves after 2, 6, and 20 generations. For each data point, 20 individuals were chosen from the phenotypic extremes of a random sample of 100 individuals for each of 100 replicate simulated populations.

The power of admixture mapping will also depend on the number of loci in the dataset as well as the number of individuals in the population sample, so we conducted association tests using larger datasets consisting of nearly 13,000 neutral background genes and a range of sample sizes to obtain estimates of power at the exome-level. We applied the Cochran-Armitage trend test to samples conditioned on a trait frequency of 0.5 simulated for a Mendelian trait and found that for FST < 0.35, even a sample of size n = 20 can provide reasonable power (Figure 6). There is, however, a clear benefit of increasing sample size, especially when differentiation between parental populations is high. Importantly, increasing the number of background neutral sites does not affect mapping power substantially, at least for n = 40. Sample size had a similar effect on the power to identify the selected gene, although the likelihood of identifying the ATL with n = 20 was slightly lower for the gene-wise test relative to the SNP-wise test (Figure 7). At the exome-level, at least 40 diploids must be included to achieve high probability of identifying the selected gene, suggesting that combining information across SNPs using Fisher’s combined p-value adds noise and contributes little to improving power particularly when sample sizes are small (i.e. n = 20). With larger sample sizes, the power to identify the selected gene increases as expected with increasing generations since admixture.

Figure 6.

Figure 6

Exome-level admixture mapping identifies adapted SNP for Mendelian trait across a range of sample sizes. The selected SNP was compared to background of SNPs within ~13,000 neutral genes. Analysis was conducted for samples of A) 20, B) 40, C) 60, and D) 100 diploids. Colors indicate the probability of the adapted SNP falling in top 10 of the ranked list. See Figure 4 legend for additional details.

Figure 7.

Figure 7

Exome-level admixture mapping identifies adapted gene for Mendelian trait across a range of sample sizes. The selected gene was compared to background of ~13,000 neutral genes. Analysis was conducted for samples of A) 20, B) 40, C) 60, and D) 100 diploids. Colors indicate the probability of the adapted SNP falling in top 10 of the ranked list. See Figure 4 legend for additional details.

DISCUSSION

Identifying adaptive loci underlying ecological divergence remains an important step in understanding the genetic basis of adaptation and evolutionary mechanisms driving phenotypic diversification. Significant progress has been made towards successful mapping efforts through improvements in analytical tools as well as technologies for genetic typing and sequencing (Erickson et al. 2004; Storz 2005). Yet a thorough understanding of the power and limits of mapping techniques under natural conditions has not been established and such efforts still lack genetic resolution. It is now possible through the use of several technological advances to collect genetic data at protein coding sequences from many individuals, thus improving our ability to identify adaptive variants or markers in LD with such variants in mapping efforts. Coupled with an assessment of analytical power of mapping techniques, mapping of adaptively diverged loci should be possible in a wide variety of systems. In this paper, we sought to understand the evolutionary conditions under which divergence and admixture mapping have ample power to identify an adaptive trait locus in an exome sequencing experimental framework such that experimental design can be tailored to specific features of any given system.

It is important to reiterate that the goal was not to evaluate specific statistics or provide a novel statistical test, but rather to quantify the effect of neutral divergence and admixture on the probability that the adaptive locus falls in the tail of an empirical distribution of a given statistic. Indeed, other methods for measuring genetic differentiation (Storz 2005) and genetic association (Balding 2006) exist and may provide improved analytical power under some circumstances. In particular, there is a large emerging literature from human genetics on methods for combining evidence in a gene from multiple SNPs that might also be applicable to exome sequencing data from non-model organisms (Potter 2006; Morgenthaler & Thilly 2007; Purcell et al. 2007; Madsen & Browning 2009; Lasky-Su et al. 2010; Price et al. 2010). Our goal was not to identify a singularly optimal method for all experimental circumstances or even to assess performance of a suite of methods. We chose several straight forward statistics for both divergence and admixture mapping in order to provide a general demonstration of how these mapping approaches can be expected to perform under various evolutionary conditions.

To assess the power of divergence and admixture mapping for mapping of adaptively diverged loci, we simulated exomes containing a single adaptive locus under a wide variety of population models and quantified the power to identify the selected locus using divergence and admixture mapping. Since most experimenters are not likely to focus exclusively on the most extreme values in a dataset, but rather on a set of candidates from the tail of the distribution, we focused on an identification scheme where we ask how often the adapted locus falls in the top 10 of the ranked list. Under this scheme and our simulation conditions, our results indicate that divergence mapping can identify a locally adapted locus with probability of 0.9 when mean background divergence (FST) is less than approximately 0.17 for samples of 50 individuals. The exact threshold value of background divergence will vary somewhat depending on whether the target is a selected SNP or a selected gene as well as features of the system such as expected heterozygosity, effective population size, and historical demography. Nonetheless, these results confirm that the power to identify a selected locus is very high when background divergence is low, as is often the case when population subdivision is recent and/or migration between subpopulations persists. Despite a number of important differences in their simulation and analytical framework, Beaumont and Balding (Beaumont & Balding 2004) found similar power to identify directionally selected loci in systems with median background FST of approximately 0.1.

Consistent with the intuition that the ability to distinguish a locus differentiated by natural selection will be diminished when highly differentiated neutral loci are common, we found the power of divergence mapping decreased from one to nearly zero over a relatively small increase in mean FST (Figure 2). Since we specifically modeled a complete selective sweep where the selected variant fixed in one population, this shift in power likely reflects the point at which new neutral mutations fix and previously segregating variants are reciprocally fixed and lost at an appreciable frequency under the combination of effective population size and mutation rate used in our simulations. Power and interpretation of divergence mapping will be further complicated if multiple loci have adaptively fixed among two populations, as may be the case in Herring for example (Lamichhaney et al. 2012), highlighting the need to use reduce linkage disequilibrium among adapted sites and explicitly associate loci with specific phenotypes.

Bottlenecks increase the variance in coalescence time, and are, therefore, expected to reduce the power to identify specific causal SNPs using allele frequency differences. We tested the effect of bottlenecks on the power of divergence mapping and found that the shift from high to low power occurred at a slightly lower mean FST level, but the effect was minimal across a range of bottleneck models for a fixed value of FST (Figure 3). Our results provide a quantitative assessment of the power of model-free empirical-distribution based divergence mapping across a range of population models that can be used to determine whether this mapping approach is appropriate for a given system when rough estimates of background divergence is available.

QTL mapping has a long history as an alternative approach for successfully mapping traits in both agricultural and evolutionary research settings (Lynch & Walsh 1998; Erickson et al. 2004), and we show here that similar logic can be applied to map traits in hybrid populations admixed from even highly diverged parental populations. We simulated populations admixed from phenotypically diverged populations, applied several association tests under a variety of phenotypic trait models, and found very promising statistical power to identify an adaptively diverged locus with data from even as few as 20 individuals. There was a clear interaction between background neutral divergence between parental populations and the number of generations needed to achieve high power. We found that, without multiple generations of recombination, statistical power was quite comparable between admixture and divergence mapping. However, admixture mapping provided high mapping power at even very high levels of background divergence as the number of generations of recombination increased (Figures 6 and 7). In general, our results point to an underappreciated opportunity to map genes underlying adaptive traits when divergence mapping is underpowered. Of course, this approach presupposes the ability to collect hybrids and quantify the adaptive phenotype in addition to genetic data collection. But if phenotypes and genetic data can be collected for a population sample of hybrids, our results provide insight into what analytical power can be expected based according to the age of the hybrid population and divergence between the parental populations.

In order to make an initial assessment of divergence and admixture, we made a number of simplifying assumptions, but we believe that these are justified and our conclusions are generally robust to violations of these assumptions. In the case of admixture mapping, we assumed a model in which the hybrid population was formed by a single pulse of admixture with no subsequent migration from parental populations. Examples of both single pulse admixtures and continuous migration can certainly be found in nature, but the major difference with respect to our analytical approach is simply the number of recombination events that have occurred since the founding of the hybrid population. The addition of ‘pure’ parental migrants to the hybrid population simply introduces chromosomes with fewer recombination events between adaptive and neutral loci. In this case, since the number of generations since admixture in our model is actually a proxy for the number of recombination events, the expected power is simply a mixture or average among several ‘generation points’ in Figures 6 and 7. For divergence mapping, including migration in the model during the divergence of the two parental populations, will only serve to increase mapping power as long as there is selection against migrant alleles at the selected locus and free introgression at neutral genomic regions.

For assessment of both divergence mapping and admixture mapping, we assumed that the populations were diverged by only a single adaptive trait and focused on monogenic trait models. For both mapping approaches, the existence of more than one locally adapted trait would not diminish statistical power dramatically. Since we are advocating a candidate list approach that implies additional experimental analysis of mapped candidates, the existence of multiple locally adapted loci in the exome simply means that the candidate list should contain additional true positive results, although perhaps the size of the list should accommodate more complex genetic architectures. In fact, association tests as employed here in the admixture mapping approach could be used to specifically test for association between diverged alleles and distinct phenotypes, provided phenotypes can be identified and quantified. Similarly, multi-trait adaptations will not affect admixture mapping as we have envisioned it here since each trait should be tested independently. Traits under polygenic control will necessarily be more difficult to map since each locus is expected to contribute small effects (Lynch & Walsh 1998; Erickson et al. 2004). Polygenic traits will be especially difficult to map using divergence mapping, since phenotypic divergence can be achieved with only minor shifts in allele frequencies under some models (Storz 2005). As such, the mapping approaches advocated here would have optimal chances of success when applied where some information regarding the genomic architecture of the trait is available, and power will be highest for Mendelian and monogenic quantitative traits, or traits for which a single locus has a strong contribution to the additive genetic variation.

The demographic model we employed to assess admixture mapping also has implications for the power of this mapping approach. We assumed that the effective population size of the admixed hybrid population would be significantly less than the size of the parental populations. This demographic model results in an increase in both genetic drift and LD in the admixed population, both of which can affect the ability to pinpoint the adapted locus. For example, we inspected patterns of LD in the admixed populations and found that the number of sites in complete LD with the selected site increases slightly among sites within and nearby (within 50 SNP window) the selected gene (Figure S5) relative to unlinked sites. This small increase reflects the fact that recombination events are rare in general and the action of genetic drift in the 50 generations of forward simulations causing some tightly linked sites to drift into LD with the selected site. When differentiation between parental populations is low, and the selected site may be one of only a small number of fixed differences between the parental populations, the addition of sites in complete LD with the selected site as genetic drift proceeds may have a significant impact on the ability to pinpoint selected sites with association tests. When differentiation between parental populations is intermediate (0.35 > FST < 0.9), however, and the parental populations are differentiated by many fixed differences including the selected site, recombination has a much larger overall effect on the data by reducing the number of sites in complete LD with the selected site, despite a very minor increase in the number of sites within and near the selected gene (Figure S5). As a result, the admixture process should increase the power of admixture mapping in the face of genetic drift. Interestingly, a slightly different pattern is observed when differentiation between parental populations is very high (e.g. FST = 0.9). At this level of differentiation, many sites throughout the exome are fixed between the parental populations at the time of admixture, and, although LD with the selected site decreases precipitously with recombination, a substantial number of sites remain in high LD with the selected site even after 50 generations (Figure S5). This pattern is due to the especially high number of segregating sites, especially within the selected gene, for the model where FST is equal to 0.9 relative other less differentiated models (1.42 times as many SNPs for FST = 0.9 relative to FST = 0.74), potentially highlighting the important role of the population-scaled mutation rate used in the simulations in determining mapping power. As recombination proceeds among all sites, these remaining sites should eventually reach linkage equilibrium with respect to the selected site, suggesting that additional generations may be required to maximize the power of admixture mapping as a function of the number of sites within close proximity to the selected site.

Although the simulation framework used here is necessarily simplistic, there are a number of natural systems that are ideally suited for exome sequencing-based mapping of adaptive loci. The utility of divergence mapping is clear for a number of systems where divergence is low. For example, one recent study applied a divergence mapping approach to data similar in structure to the exome sequencing we are advocating here to identify ATL underlying tolerance to variable levels of salinity in Atlantic herring where neutral FST = 0.022 (Lamichhaney et al. 2012). ATL controlling wing patterning have also been localized using divergence mapping in several populations of Heliconius butterflies ranging in background FST of approximately 0.1 to nearly zero (Nadeau et al. 2012). Hybrid populations admixed from highly diverged populations of Heliconius are well suited for admixture mapping. In fact, a similar approach was applied to SNPs in candidate color pattern genes in a hybrid population of Heliconius erato admixed from lowly diverged (FST < 0.05) ‘pure’ races (Counterman et al. 2010). Heliconius himera is a sister taxon to the H. erato that is differentiated phenotypically in a number of ways including wing color patterning, which is controlled by only a few major loci (Benson 1972; Barton & Hewitt 1989; Jiggins et al. 1996). Fertile hybrids between these species can be found at a series of narrow contact zones in south-western Ecuador and northern Peru (Jiggins et al. 1996; McMillan et al. 1997). Although the empirical distribution is based on only a small number of polymorphic allozymes, the distribution of FST does not seem to be unimodal, perhaps owing to the effects of both stabilizing and disruptive selection (Jiggins et al. 1997). Nonetheless, divergence is sufficiently high between these taxa (mean FST 0.35) that divergence mapping is not likely to provide useful results. According to Figures 6 and 7, ATL underlying phenotypic differences between these taxa could be reliably mapped using admixture mapping with data from as few as 40 individuals provided hybridization has been occurring for at least 10 generations. If resources allow for data collection from 40 or 60 individuals, significant mapping resolution could be achieved in F2 or F4 hybrids (Figure 6), although genotype-phenotype association failed to provide fine resolution in laboratory crosses of these species owing to a lack of recombination in the candidate region (Martin et al. 2012). A similar investment could also be leveraged to identify ecological and behavioral ATL in the secondary-contact hybrid zone between Gryllus pennsylvanicus and Gryllus firmus crickets that show intermediate (FST ≈ 0.26) levels of nuclear divergence (Broughton & Harrison 2003; Maroja et al. 2009; Andrés et al. 2013).

In this study, we undertook a simulation analysis of divergence and admixture mapping to quantify the power of these approaches for mapping ATL under a range of evolutionary conditions. Our results provide the first quantitative assessment of divergence mapping showing that this approach is very powerful for mapping ATL, but under a restricted set of scenarios. We also show for the first time that loci underlying differential adaptation among highly diverged populations can be mapped by applying simple association tests to admixed populations to a sample from a hybrid population. Together with novel exome-capture sequencing methods, our results demonstrate that ATL can be robustly mapped in many systems even when virtually no genetic resources are available.

Acknowledgments

This research was support by a Genentech Foundation Innovation Postdoctoral Fellowship in the Center for Computational Biology at The University of California, Berkeley to J.E.C and NIH Grant 2R01HG003229-09 to R.N. We are greatful to three anonymous reviewers for very insightful and helpful comments.

Footnotes

DATA ACCESSIBILITY

No data generated for this study. Custom scripts implementing parts of the analysis pipeline can be found at github.com/JEcrawford.

References

  1. Akey JM, Zhang G, Zhang K, Jin L, Shriver MD. Interrogating a high-density SNP map for signatures of natural selection. Genome research. 2002;12:1805–1814. doi: 10.1101/gr.631202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Andrés JA, Larson EL, Bogdanowicz SM, Harrison RG. Patterns of Transcriptome Divergence in the Male Accessory Gland of Two Closely Related Species of Field Crickets. Genetics. 2013;193:501–513. doi: 10.1534/genetics.112.142299. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Armitage P. Tests for Linear Trends in Proportions and Frequencies. Biometrics. 1955;11:375–386. [Google Scholar]
  4. Balding DJ. A tutorial on statistical methods for population association studies. Nature Reviews Genetics. 2006;7:781–791. doi: 10.1038/nrg1916. [DOI] [PubMed] [Google Scholar]
  5. Barton NH, Hewitt GM. Adaptation, speciation and hybrid zones. Nature. 1989;341:497–503. doi: 10.1038/341497a0. [DOI] [PubMed] [Google Scholar]
  6. Beaumont MA, Balding DJ. Identifying adaptive genetic divergence among populations from genome scans. Molecular ecology. 2004;13:969–980. doi: 10.1111/j.1365-294x.2004.02125.x. [DOI] [PubMed] [Google Scholar]
  7. Beltrán M, Jiggins CD, Brower AVZ, Bermingham E, Mallet J. Do pollen feeding, pupal-mating and larval gregariousness have a single origin in Heliconius butterflies? Inferences from multilocus DNA sequence data. Biological Journal of the Linnean Society. 2007;92:221–239. [Google Scholar]
  8. Benson WW. Strong Natural Selection in a Warning-Color Hybrid Zone. Science. 1972;176:936–939. [Google Scholar]
  9. Bi K, Vanderpool D, Singhal S, et al. Transcriptome-based exon capture enables highly cost-effective comparative genomic data collection at moderate evolutionary scales. BMC genomics. 2012;13:403. doi: 10.1186/1471-2164-13-403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Bradshaw WE, Emerson KJ, Catchen JM, Cresko WA, Holzapfel CM. Footprints in time: comparative quantitative trait loci mapping of the pitcher-plant mosquito, Wyeomyia smithii. Proceedings. Biological sciences / The Royal Society. 2012;279:4551–4558. doi: 10.1098/rspb.2012.1917. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Broughton RE, Harrison RG. Nuclear Gene Genealogies Reveal Historical, Demographic and Selective Factors Associated With Speciation in Field Crickets. Genetics. 2003;163:1389–1401. doi: 10.1093/genetics/163.4.1389. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Brown C. Hybridization among the subspecies of the plethodontid salamander Ensatina eschscholtzii. University of California Publications in Zoology. 1974;98:211–256. [Google Scholar]
  13. Bush GL. The Taxonomy, Cytology and Evolution of the Genus Rhagoletis in North America. Museum of Comparative Zoology; Cambridge, Massachusetts: 1966. [Google Scholar]
  14. Bush GL. In: Evolutionary Strategies of Parasitic Insects and Mites. Price P, editor. Plenum; New York: 1975. pp. 187–206. [Google Scholar]
  15. Charlesworth B, Nordborg M, Charlesworth D. The effects of local selection, balanced polymorphism and background selection on equilibrium patterns of genetic diversity in subdivided populations. Genetics Research. 1997;70:155–174. doi: 10.1017/s0016672397002954. [DOI] [PubMed] [Google Scholar]
  16. Counterman BA, Araujo-Perez F, Hines HM, et al. Genomic Hotspots for Adaptation: The Population Genetics of Müllerian Mimicry in Heliconius erato. PLoS Genet. 2010;6:e1000796. doi: 10.1371/journal.pgen.1000796. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Erickson DL, Fenster CB, Stenøien HK, Price D. Quantitative trait locus analyses and the study of evolutionary process. Molecular ecology. 2004;13:2505–2522. doi: 10.1111/j.1365-294X.2004.02254.x. [DOI] [PubMed] [Google Scholar]
  18. Ewing G, Hermisson J. MSMS: a coalescent simulation program including recombination, demographic structure and selection at a single locus. Bioinformatics (Oxford, England) 2010;26:2064–2065. doi: 10.1093/bioinformatics/btq322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Excoffier L, Hofer T, Foll M. Detecting loci under selection in a hierarchically structured population. Heredity. 2009;103:285–298. doi: 10.1038/hdy.2009.74. [DOI] [PubMed] [Google Scholar]
  20. Fisher RA. The genetical theory of natural selection. Clarendon Press; Oxford, UK: 1930. [Google Scholar]
  21. Fisher RA. Statistical Methods for Research Workers. Oliver and Boyd; London: 1932. [Google Scholar]
  22. Fourcade Y, Chaput-Bardy A, Secondi J, Fleurant C, Lemaire C. Is local selection so widespread in river organisms? Fractal geometry of river networks leads to high bias in outlier detection. Molecular ecology. 2013;22:2065–2073. doi: 10.1111/mec.12158. [DOI] [PubMed] [Google Scholar]
  23. Geraldes A, Basset P, Smith KL, Nachman MW. Higher differentiation among subspecies of the house mouse (Mus musculus) in genomic regions with low recombination. Molecular Ecology. 2011;20:4722–4736. doi: 10.1111/j.1365-294X.2011.05285.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Hadid Y, Tzur S, Pavlíček T, et al. Possible incipient sympatric ecological speciation in blind mole rats (Spalax) Proceedings of the National Academy of Sciences. 2013;110:2587–2592. doi: 10.1073/pnas.1222588110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Haldane JBS. A mathematical theory of natural and artificial selection. Part IV. Isolation. Proceedings of the Cambridge Philosophical Society. 1930;26:220–230. [Google Scholar]
  26. Hohenlohe PA, Bassham S, Etter PD, et al. Population Genomics of Parallel Adaptation in Threespine Stickleback using Sequenced RAD Tags. PLoS Genet. 2010;6:e1000862. doi: 10.1371/journal.pgen.1000862. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Jiggins CD, McMillan WO, King P, Mallet J. The maintenance of species differences across a Heliconius hybrid zone. Heredity. 1997;79:495–505. [Google Scholar]
  28. Jiggins CD, McMillan WO, Neukirchen W, Mallet J. What can hybrid zones tell us about speciation? The case of Heliconius erato and H. himera (Lepidoptera: Nymphalidae) Biological Journal of the Linnean Society. 1996;59:221–242. [Google Scholar]
  29. Kaufman DW. Adaptive Coloration in Peromyscus polionotus: Experimental Selection by Owls. Journal of Mammalogy. 1974;55:271–283. [Google Scholar]
  30. Kimura M. Some Problems of Stochastic Processes in Genetics. The Annals of Mathematical Statistics. 1957;28:882–901. [Google Scholar]
  31. Kimura M. On the Probability of Fixation of Mutant Genes in a Population. Genetics. 1962;47:713–719. doi: 10.1093/genetics/47.6.713. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kulathinal RJ, Bennett SM, Fitzpatrick CL, Noor MAF. Fine-scale mapping of recombination rate in Drosophila refines its correlation to diversity and divergence. Proceedings of the National Academy of Sciences. 2008;105:10051–10056. doi: 10.1073/pnas.0801848105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Lamichhaney S, Barrio AM, Rafati N, et al. Population-scale sequencing reveals genetic differentiation due to local adaptation in Atlantic herring. Proceedings of the National Academy of Sciences. 2012;109:19345–19350. doi: 10.1073/pnas.1216128109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Lasky-Su J, Murphy A, McQueen MB, Weiss S, Lange C. An omnibus test for family-based association studies with multiple SNPs and multiple phenotypes. European journal of human genetics: EJHG. 2010;18:720–725. doi: 10.1038/ejhg.2009.221. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Lewontin RC, Krakauer J. Distribution of gene frequency as a test of the theory of the selective neutrality of polymorphisms. Genetics. 1973;74:175–195. doi: 10.1093/genetics/74.1.175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Li B, Leal SM. Methods for Detecting Associations with Rare Variants for Common Diseases: Application to Analysis of Sequence Data. American Journal of Human Genetics. 2008;83:311–321. doi: 10.1016/j.ajhg.2008.06.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Lynch M, Walsh B. Genetics and Analysis of Quantitative Traits. Sinauer Associates; 1998. [Google Scholar]
  38. Madsen BE, Browning SR. A groupwise association test for rare mutations using a weighted sum statistic. PLoS genetics. 2009;5:e1000384. doi: 10.1371/journal.pgen.1000384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Mallet J, Beltrán M, Neukirchen W, Linares M. Natural hybridization in heliconiine butterflies: the species boundary as a continuum. BMC Evolutionary Biology. 2007;7:28. doi: 10.1186/1471-2148-7-28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Maroja LS, Andrés JA, Harrison RG. Genealogical Discordance and Patterns of Introgression and Selection Across a Cricket Hybrid Zone. Evolution. 2009;63:2999–3015. doi: 10.1111/j.1558-5646.2009.00767.x. [DOI] [PubMed] [Google Scholar]
  41. Martin A, Papa R, Nadeau NJ, et al. Diversification of complex butterfly wing patterns by repeated regulatory evolution of a Wnt ligand. Proceedings of the National Academy of Sciences. 2012;109:12632–12637. doi: 10.1073/pnas.1204800109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. McMillan WO, Jiggins CD, Mallet J. What initiates speciation in passion-vine butterflies? Proceedings of the National Academy of Sciences. 1997;94:8628–8633. doi: 10.1073/pnas.94.16.8628. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Moore LG. Human genetic adaptation to high altitude. High altitude medicine & biology. 2001;2:257–279. doi: 10.1089/152702901750265341. [DOI] [PubMed] [Google Scholar]
  44. Morgenthaler S, Thilly WG. A strategy to discover genes that carry multi-allelic or mono-allelic risk for common diseases: a cohort allelic sums test (CAST) Mutation research. 2007;615:28–56. doi: 10.1016/j.mrfmmm.2006.09.003. [DOI] [PubMed] [Google Scholar]
  45. Nadeau NJ, Whibley A, Jones RT, et al. Genomic islands of divergence in hybridizing Heliconius butterflies identified by large-scale targeted sequencing. Philosophical transactions of the Royal Society of London. Series B, Biological sciences. 2012;367:343–353. doi: 10.1098/rstb.2011.0198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Ng SB, Buckingham KJ, Lee C, et al. Exome sequencing identifies the cause of a Mendelian disorder. Nature genetics. 2010;42:30–35. doi: 10.1038/ng.499. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Nosil P, Egan SP, Funk DJ, Hoekstra H. Heterogeneous Genomic Differentiation Between Walking-Stick Ecotypes: “Isolation by Adaptation” and Multiple Roles for Divergent Selection. Evolution. 2008;62:316–336. doi: 10.1111/j.1558-5646.2007.00299.x. [DOI] [PubMed] [Google Scholar]
  48. Nosil P, Funk DJ, Ortiz-Barrientos D. Divergent selection and heterogeneous genomic divergence. Molecular Ecology. 2009;18:375–402. doi: 10.1111/j.1365-294X.2008.03946.x. [DOI] [PubMed] [Google Scholar]
  49. Papa R, Kapan DD, Counterman BA, et al. Multi-Allelic Major Effect Genes Interact with Minor Effect QTLs to Control Adaptive Color Pattern Variation in Heliconius erato. PLoS ONE. 2013;8:e57033. doi: 10.1371/journal.pone.0057033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Patterson N, Hattangadi N, Lane B, et al. Methods for high-density admixture mapping of disease genes. American Journal of Human Genetics. 2004;74:979–1000. doi: 10.1086/420871. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Pereira RJ, Wake DB. Genetic leakage after adaptive and nonadaptive divergence in the Ensatina eschscholtzii ring species. Evolution; international journal of organic evolution. 2009;63:2288–2301. doi: 10.1111/j.1558-5646.2009.00722.x. [DOI] [PubMed] [Google Scholar]
  52. Pérez-Figueroa A, García-Pereira MJ, Saura M, Rolán-Alvarez E, Caballero A. Comparing three different methods to detect selective loci using dominant markers. Journal of evolutionary biology. 2010;23:2267–2276. doi: 10.1111/j.1420-9101.2010.02093.x. [DOI] [PubMed] [Google Scholar]
  53. Potter DM. Omnibus permutation tests of the association of an ensemble of genetic markers with disease in case-control studies. Genetic epidemiology. 2006;30:438–446. doi: 10.1002/gepi.20155. [DOI] [PubMed] [Google Scholar]
  54. Price AL, Kryukov GV, de Bakker PIW, et al. Pooled association tests for rare variants in exon-resequencing studies. American journal of human genetics. 2010;86:832–838. doi: 10.1016/j.ajhg.2010.04.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Purcell S, Neale B, Todd-Brown K, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. American journal of human genetics. 2007;81:559–575. doi: 10.1086/519795. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. R Development Core Team. R: A language and Environment for Statistical Computing. 2011. [Google Scholar]
  57. Reed RD, Papa R, Martin A, et al. optix Drives the Repeated Convergent Evolution of Butterfly Wing Pattern Mimicry. Science. 2011;333:1137–1141. doi: 10.1126/science.1208227. [DOI] [PubMed] [Google Scholar]
  58. Salcedo T, Geraldes A, Nachman MW. Nucleotide Variation in Wild and Inbred Mice. Genetics. 2007;177:2277–2291. doi: 10.1534/genetics.107.079988. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Stebbins RC. Speciation in salamanders of the plethodontid genus Ensatina. University of California Publications in Zoology. 1949;48:377–526. [Google Scholar]
  60. Stenson PD, Mort M, Ball EV, et al. The Human Gene Mutation Database: 2008 update. Genome Medicine. 2009;1:13. doi: 10.1186/gm13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Storz JF. INVITED REVIEW: Using genome scans of DNA polymorphism to infer adaptive population divergence. Molecular Ecology. 2005;14:671–688. doi: 10.1111/j.1365-294X.2005.02437.x. [DOI] [PubMed] [Google Scholar]
  62. Sumner FB. An Analysis of Geographic Variation in Mice of the Peromyscus polionotus Group from Florida and Alabama. Journal of Mammalogy. 1926;7:149–184. [Google Scholar]
  63. Supple MA, Hines HM, Dasmahapatra KK, et al. Genomic architecture of adaptive color pattern divergence and convergence in Heliconius butterflies. Genome research. 2013;23:1248–1257. doi: 10.1101/gr.150615.112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Turner JRG. Adaptation and Evolution in Heliconius: A Defense of NeoDarwinism. Annual Review of Ecology and Systematics. 1981;12:99–121. [Google Scholar]
  65. Twomey E, Yeager J, Brown JL, et al. Phenotypic and Genetic Divergence among Poison Frog Populations in a Mimetic Radiation. PloS one. 2013;8:e55443. doi: 10.1371/journal.pone.0055443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Via S. The Genetic Structure of Host Plant Adaptation in a Spatial Patchwork: Demographic Variability among Reciprocally Transplanted Pea Aphid Clones. Evolution. 1991;45:827–852. doi: 10.1111/j.1558-5646.1991.tb04353.x. [DOI] [PubMed] [Google Scholar]
  67. Weir B, Cockerham CC. Estimating F-statistics for the analysis of population structure. Evolution. 1984:1358–1370. doi: 10.1111/j.1558-5646.1984.tb05657.x. [DOI] [PubMed] [Google Scholar]
  68. Wilding CS, Butlin RK, Grahame J. Differential gene exchange between parapatric morphs of Littorina saxatilis detected using AFLP markers. Journal of Evolutionary Biology. 2001;14:611–619. [Google Scholar]
  69. Winkler CA, Nelson GW, Smith MW. Admixture mapping comes of age. Annual Review of Genomics and Human Genetics. 2010;11:65–89. doi: 10.1146/annurev-genom-082509-141523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Wright S. Evolution in Mendelian Populations. Genetics. 1931;16:97–159. doi: 10.1093/genetics/16.2.97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Wright S. The roles of mutation, inbreeding, crossbreeding and selection in evolution. Proceedings of the sixth international congress on genetics. 1932:356–366. [Google Scholar]
  72. Yi X, Liang Y, Huerta-Sanchez E, et al. Sequencing of 50 Human Exomes Reveals Adaptation to High Altitude. Science. 2010;329:75–78. doi: 10.1126/science.1190371. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES