

MANY important questions in genetics involve looking back in time using data sampled in the present. The coalescent process describes the ancestry of a sample of genes: as lineages trace back, they coalesce at a rate inversely proportional to the effective population size. This remarkably simple approximation now dominates population genetics, both because it gives a direct intuition into the evolutionary process and because it allows efficient simulation.
The coalescent predicts the genealogical relationships between sampled genes, and it depends only on the effective population size, regardless of the detailed life history. The coalescent was first described by Kingman (1982), but was developed independently by Hudson (1983) and Tajima (1983); its influence on population genetics came primarily through Hudson. The coalescent is rooted in older concepts: Malécot’s (1948) idea of identity by descent, which is central to quantitative genetics (Kempthorne 1954), and the diffusion approximation (Kimura 1955), which depends on the same effective population size. The coalescent emerged not through radically new concepts, but, rather, because it gives a natural way to analyze the samples of DNA sequences that were just becoming available. Hudson and Kaplan (1988), with its companion (Kaplan et al. 1988), extended the coalescent to include selection and used it to interpret data on sequence variation around the alcohol dehydrogenase (Adh) locus of Drosophila melanogaster (Kreitman 1983).
Under the coalescent, lineages coalesce at a rate equal to the inverse of the effective number of genes in the population, 1/2Ne. Migration, recombination, and mutation can be included by allowing ancestral lineages to jump between locations, genetic backgrounds, or allelic states; this extension is known as the structured coalescent (Hudson 1983). Selection is much harder to incorporate because the ancestry now depends on the allelic state. However, Kaplan et al. (1988) showed that, if the selected backgrounds are taken as given, then the structured coalescent describes the genealogy of linked neutral alleles; random fluctuations in the frequency of the selected backgrounds can be treated by a diffusion that couples to the coalescent. A surprising conclusion from this method is that selection has to be extremely strong relative to random drift to distort neutral genealogies (Barton and Etheridge 2004). This is a fundamental obstacle to detecting selection from sequence data.
The fast and slow (F/S) alleles of the Adh locus of D. melanogaster differ by a single amino acid. Kreitman (1983) sequenced 11 copies of the locus and found a sharp peak of polymorphism around the amino-acid difference. This is consistent with maintenance of these alleles by long-term balancing selection. Hudson and Kaplan (1988) showed that divergence between the F and S alleles is consistent with balancing selection, albeit with a lower-than-average recombination rate. However, there was also excess variation among the S alleles, which was not expected. Even though Adh in Drosophila is one of the most intensively studied polymorphisms, we still do not know how its sequence variation has been shaped by selection (Begun et al. 1999).
Population genetics is now focused on understanding the abundance of sequence data that has recently become available. Since the first work of Kreitman (1983), the goal has been to infer the nature and strength of selection across the genome directly from the DNA sequence; “genome scans” of the kind introduced by Hudson and Kaplan (1988) are now being carried out for an enormous range of organisms. However, it is disconcerting that even the first and best-studied example is still unresolved.
Footnotes
Communicating editor: C. Gelling
ORIGINAL CITATION
The Coalescent Process in Models with Selection and Recombination
Richard R. Hudson and Norman L. Kaplan
GENETICS November 1, 1988 120: 831–840
Photo of Norman Kaplan (left) courtesy of Jotun Hein. Photo of Richard Hudson (right) courtesy of himself.
Literature Cited
- Barton N. H., Etheridge A. M., 2004. The effect of selection on genealogies. Genetics 166: 1115–1131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Begun D. J., Betancourt A. J., Langley C. H., Stephan W., 1999. Is the fast/slow allozyme variation at the Adh locus of Drosophila melanogaster an ancient balanced polymorphism? Mol. Biol. Evol. 16: 1816–1819. [DOI] [PubMed] [Google Scholar]
- Hudson R. R., 1983. Properties of a neutral allele model with intragenic recombination. Theor. Popul. Biol. 23: 183–201. [DOI] [PubMed] [Google Scholar]
- Kaplan N. L., Darden T., Hudson R. R., 1988. The coalescent process in models with selection. Genetics 120: 819–829. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kempthorne O., 1954. The correlation between relatives in a random mating population. Proc. R. Soc. Lond. Ser. B-Biol. Sci. 143: 102–113. [PubMed] [Google Scholar]
- Kimura M., 1955. Stochastic processes and distribution of gene frequencies under natural selection. Cold Spring Harb. Symp. Quant. Biol. 20: 33–55. [DOI] [PubMed] [Google Scholar]
- Kingman J. F. C., 1982. The coalescent. Stoch. Process. Their Appl. 13: 235–248. [Google Scholar]
- Kreitman M., 1983. Nucleotide polymorphisms at the alcohol dehydrogenase locus of Drosophila melanogaster. Nature 304: 412–417. [DOI] [PubMed] [Google Scholar]
- Malécot G., 1948. Les Mathématiques de l’Hérédité, Masson et Cie., Paris. [Google Scholar]
- Tajima F., 1983. Evolutionary relationship of DNA sequences in finite populations. Genetics 105: 437–460. [DOI] [PMC free article] [PubMed] [Google Scholar]
Further Reading in GENETICS
- Kingman J. F. C., 2000. Origins of the coalescent: 1974–1982. Genetics 156: 1–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nagylaki T., 1989. Gustave Malécot and the transition from classical to modern population genetics. Genetics 122: 253–268. [DOI] [PMC free article] [PubMed] [Google Scholar]
Other GENETICS Articles by R. R. Hudson and N. L. Kaplan
- Adams A. M., Hudson R. R., 2004. Maximum-likelihood estimation of demographic parameters using the frequency spectrum of unlinked single-nucleotide polymorphisms. Genetics 168: 1699–1712. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Braverman J. M., Hudson R. R., Kaplan N. L., Langley C. H., Stephan W., 1995. The hitchhiking effect on the site frequency spectrum of DNA polymorphisms. Genetics 140: 783–796. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Haubold B., Travisano M., Rainey P. B., Hudson R. R., 1998. Detecting linkage disequilibrium in bacterial populations. Genetics 150: 1341–1348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., 1982. Estimating genetic variability with restriction endonucleases. Genetics 100: 711–719. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., 1985. The sampling distribution of linkage disequilibrium under an infinite allele model without selection. Genetics 109: 611–631. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., 1992. Gene trees, species trees and the segregation of ancestral alleles. Genetics 131: 509–513. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., 2000. A new statistic for detecting genetic differentiation. Genetics 155: 2011–2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., 2001. Two-locus sampling distributions and their application. Genetics 159: 1805–1817. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Kaplan N. L., 1985. Statistical properties of the number of recombination events in the history of a sample of DNA sequences. Genetics 111: 147–164. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Kaplan N. L., 1986. On the divergence of alleles in nested subsamples from finite populations. Genetics 113: 1057–1076. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Kaplan N. L., 1995. Deleterious background selection with recombination. Genetics 141: 1605–1617. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Kreitman M., Aguadé M., 1987. A test of neutral molecular evolution based on nucleotide data. Genetics 116: 153–159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Slatkin M., Maddison W. P., 1992. Estimation of levels of gene flow from DNA sequence data. Genetics 132: 583–589. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hudson R. R., Bailey K., Skarecky D., Kwiatowski J., Ayala F. J., 1994. Evidence for positive selection in the superoxide dismutase (Sod) region of Drosophila melanogaster. Genetics 136: 1329–1340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kaplan N. L., Darden T., Hudson R. R., 1988. The coalescent process in models with selection. Genetics 120: 819–829. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kaplan N. L., Hudson R. R., Langley C. H., 1989. The “hitchhiking effect” revisited. Genetics 123: 887–899. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kreitman M., Hudson R. R., 1991. Inferring the evolutionary histories of the Adh and Adh-dup loci in Drosophila melanogaster from patterns of polymorphism and divergence. Genetics 127: 565–582. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Meyer W. K., Arbeithuber B., Ober C., Ebner T., Tiemann-Boege I., et al. , 2012. Evaluating the evidence for transmission distortion in human pedigrees. Genetics 191: 215–232. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pluzhnikov A., Di Rienzo A., Hudson R. R., 2002. Inferences about human demography based on multilocus analyses of noncoding sequences. Genetics 161: 1209–1218. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Slatkin M., Hudson R. R., 1991. Pairwise comparisons of mitochondrial DNA sequences in stable and exponentially growing populations. Genetics 129: 555–562. [DOI] [PMC free article] [PubMed] [Google Scholar]
