Skip to main content
Genetics logoLink to Genetics
. 2012 Aug;191(4):1381–1386. doi: 10.1534/genetics.112.141341

SNP-Ratio Mapping (SRM): Identifying Lethal Alleles and Mutations in Complex Genetic Backgrounds by Next-Generation Sequencing

Heike Lindner *,1, Michael T Raissig *,1, Christian Sailer *, Hiroko Shimosato-Asano *,2, Rémy Bruggmann †,, Ueli Grossniklaus *,3
PMCID: PMC3416015  PMID: 22649081

Abstract

We present a generally applicable method allowing rapid identification of causal alleles in mutagenized genomes by next-generation sequencing. Currently used approaches rely on recovering homozygotes or extensive backcrossing. In contrast, SNP-ratio mapping allows rapid cloning of lethal and/or poorly transmitted mutations and second-site modifiers, which are often in complex genetic/transgenic backgrounds.

Keywords: Arabidopsis thaliana, second-site modifier mutation, forward genetic screen, genetic model systems, lethal mutations


FORWARD genetic screens are powerful in uncovering novel gene functions in genetic model organisms. While some mutant screens can be quick to perform, the identification of the causative mutation by map-based cloning is extremely labor-intensive. Large F2 mapping populations of >1000 mutant individuals are required (Lukowitz et al. 2000; Jander et al. 2002) to fine-map a chromosomal region harboring a causative mutation. This number of mutant individuals can be difficult to obtain, especially when working with phenotypic traits that (i) are difficult to score, (ii) are weakly transmitted, or (iii) are in organisms that are hard to propagate. The recent development of next-generation sequencing (NGS) platforms has made sequencing of whole genomes quick and affordable. One application of NGS is to replace map-based cloning by the sequencing of mutagenized genomes to quickly identify causative mutations, a method successfully applied in many model organisms (Sarin et al. 2008; Smith et al. 2008; Srivatsan et al. 2008; Blumenstiel et al. 2009; Irvine et al. 2009; Schneeberger et al. 2009; Zuryn et al. 2010; Austin et al. 2011). However, current methods depend on identifying homozygous mutant individuals in an F2 mapping population after outcrossing (Schneeberger et al. 2009; Austin et al. 2011) or require several rounds of backcrossing (Zuryn et al. 2010), a time-consuming requirement not easily met in organisms with long generation times.

Here, we describe a generally applicable method, SNP-ratio mapping (SRM), which allows the rapid identification of lethal and/or poorly transmitted mutations and second-site modifiers by NGS. It is based on the distinct segregation ratio of the causative (and linked) single-nucleotide polymorphism(s) (SNPs) from that of unlinked SNPs. SRM allows the mapping of lethal mutations after only two rounds of backcrossing via NGS. After backcrossing twice to the non-mutagenized parent, any unlinked SNP created by ethyl methanesulfonate (EMS) mutagenesis segregates 1:3 in a pool of individuals. By selecting only mutant individuals in the F1 generation of the second backcross (BC2), the causative SNP is enriched and segregates 1:1 in a pool of mutant BC2 individuals (Figure 1). Thus, calculating the SNP/non-SNP segregation ratio allows the quick identification of the causative mutation. The method is applicable to any model organism and mutagen causing mostly point mutations or small indels. SRM is the method of choice when working with (i) lethal mutations, (ii) hard-to-score phenotypes, (iii) mutations with low transmission, and (iv) second-site modifiers in complex genetic/transgenic backgrounds. Here, we demonstrate the power of SRM by cloning a gametophyte lethal mutation in Arabidopsis thaliana, for which the recovery of homozygotes is not possible.

Figure 1 .

Figure 1 

SRM scheme. An EMS-treated mutant (red plant) harboring several EMS-induced SNPs throughout the genome is backcrossed to an unmutagenized wild-type parent (green plant) of the same accession. The first backcross eliminates half of the SNPs. The F1 generation is phenotyped, and a single mutant individual (red plant, BC1) is backcrossed. The F1 of the second backcross (BC2) is phenotyped, and genomic DNA of 25–50 mutant individuals is extracted and pooled. The causative SNP (red circle) is present in every mutant individual in a heterozygous state and thus segregates 1:1 in a pool of mutant individuals. Unlinked SNPs (green square) are not selected and thus segregate 1:3. After SOLiD sequencing, SNP calling, and SNP/non-SNP ratio calculation, the two different segregation ratios can be distinguished.

As proof of principle, we aimed to map the gene affected in a pollen-tube reception mutant obtained from a forward genetic screen using EMS-treated seeds of A. thaliana (Col-0 accession, Supporting Information, File S1). The turan-1 (tun-1) mutant disrupts cell–cell communication between male and female gametophytes, which is indispensable for fertilization. In flowering plants, the gametes are produced by the haploid, multicellular gametophytes. The male gametophyte (pollen tube) delivers two sperm cells to the female gametophyte (embryo sac), harboring two female gametes. Fertilization of the egg and central cell forms the embryo and the endosperm, respectively. In heterozygous tun-1 mutants, 12% (n = 1318 ovules) of the embryo sacs remain unfertilized, compared to only 1.5% (n = 1389 ovules) in the wild-type control. In tun-1 mutants, the pollen tube fails to stop growing inside the female gametophyte and does not rupture to release the sperm cells, which leads to a pollen-tube overgrowth phenotype revealed by aniline-blue staining of callose in the pollen tube’s cell wall (Figure 2 and File S1). Due to impaired fertilization and an additional effect of the mutant in the pollen, the transmission of the mutation is highly reduced, and homozygous individuals cannot be recovered. Thus, recently published methods for mutant allele identification by NGS (Schneeberger et al. 2009; Austin et al. 2011) are not applicable to mapping this gametophyte lethal mutation.

Figure 2 .

Figure 2 

Aniline-blue staining of callose in pollen tubes 2 days after pollination. The arrow indicates the place of pollen-tube arrest. (A) Fertilized wild-type ovule. (B) Ovule harboring a tun-1 embryo sac with defective pollen-tube reception. The pollen tube continues its growth and does not rupture to release the sperm cells. (C) Pollen-tube overgrowth phenotype in tun-2, an independent T-DNA line disrupting the At1g16570 gene.

To identify the TUN gene by SRM, heterozygous mutants were crossed back twice to the wild-type Col-0 parent. By selecting only mutant individuals in the F1 generation of the BC2, the causative SNP is enriched and segregates 1:1 in a pool of mutant BC2 individuals, whereas any unlinked SNP segregates 1:3 (Figure 1). We simulated a binomial distribution for a 1:1 and a 1:3 segregation to determine the optimal sample size and calculated that a 50-fold sequence coverage of the Arabidopsis genome was sufficient to distinguish a SNP segregating 1:1 from a SNP segregating 1:3 (P < 0.05, Table S1). Genomic DNA from 53 F1 individuals of the BC2 generation that displayed the mutant phenotype was pooled for sequencing (File S1). A sequencing library was prepared (File S1) and sequenced on the SOLiD 4 platform, as this method provides an incomparable sequencing accuracy optimal for SNP detection. Reads were mapped to the A. thaliana genome assembly and SNPs were called and analyzed (File S1).

We identified 2337 SNPs, of which 521 were homozygous and 1816 were heterozygous with an average sequence coverage of 57 reads (Table S2 and Table S3). The homozygous SNPs were likely due to discrepancies between our lab strain of Col-0 and the published sequence. The homozygous SNPs were discarded, since all relevant SNPs should only be heterozygous (Figure 1). Before plotting the SNP/non-SNP ratios of the heterozygous SNPs, we filtered any SNPs that showed very low or high coverage. Low-coverage SNPs could exhibit a misleading ratio due to small sample size, while very high coverage (>2× average coverage) SNPs often mapped to repetitive and/or transposable element sequences, where mapping quality is usually poor (Figure S1). Thus, we filtered out the lowest (< 19×) and the highest (> 103×) 10% quantiles, leaving 80% of the original data set. The SNP/non-SNP ratio of the remaining 1468 heterozygous SNPs was calculated and plotted against their chromosomal position (Figure 3). Any unlinked SNP should have a SNP/non-SNP ratio of ∼ 0.25, whereas a causative SNP is expected to segregate 1:1, i.e., producing a SNP/non-SNP ratio of 0.5. Furthermore, the SNPs surrounding the causative mutation should have segregation ratios > 0.25 since they have been coselected and thus cosegregate due to genetic linkage. Using this method, the causative SNP can be easily identified on the basis of the criteria that it must have a segregation ratio of ∼ 0.5, while the flanking, noncausative SNPs should cosegregate and display a ratio between 0.25 and 0.5, depending on the genetic/physical distance. The closer the flanking SNPs are, the higher this ratio will be. On the segregation ratio plot, this results in a rounded, rather flat curve, which can be visually identified without further statistical analyses (Figure 3A, red shading). Noncausative SNPs with a segregation ratio of 0.5 are likely to be surrounded by SNPs with low segregation ratios, leading to sharp drops in the SNP/non-SNP ratios of nearby SNPs (Figure 3). In our analysis of the tun-1 mutant, the only rounded peak was present in the upper arm of chromosome I (Figure 3A, red shading).

Figure 3 .

Figure 3 

SNP/non-SNP ratio plots. The SNP/non-SNP ratio of all heterozygous SNPs is calculated and plotted against the chromosomal position of the heterozygous SNPs. The red dashed line marks the SNP/non-SNP ratio at 0.5, where the causative SNP should be; the green dashed line marks the SNP/non-SNP ratio at 0.25, where all other SNPs should locate. The red shading marks the genetically linked and selected region on chromosome I with the causative SNP in At1g16570 (arrow). The gray shading marks the centromeric regions with a high SNP density, likely due to a poor mapping quality in these regions. (A–E) Chromosome I, II, III, IV, and V, respectively.

Although the causative SNP could easily be identified in our experiment, we did not want to rely on a visual identification of the rounded peak. Thus, we developed a statistical test based on the expected recombination rate of neighboring SNPs as a function of the genetic distance between the SNPs. For each SNP following the 1:1 binomial distribution (n = 118, coverage ≥ 50), we calculated the expected pattern of cosegregation with the two neighboring SNPs on each side by using the expected recombination rate according to the mean genetic distance of 1 cM/357,042 bp (File S1 and File S2). Using a χ2 goodness-of-fit test, 108 of 118 1:1 class SNPs did not lie in a linkage group (5 neighboring SNPs) that fit the expected pattern of cosegregation and therefore were discarded. Of the 10 remaining candidate SNPs, 8 reside in recombination-deficient centromeric regions. This is probably due to intrinsic problems in mapping reads to the highly repetitive centromeric sequences, leading to a high SNP density with unusual segregation ratios. Moreover, these eight linkage groups encompass 0.013 cM or less (Table S4), and the probability that 5 random SNPs lie in such close proximity is P = 1.4 × 10−8 (Poisson distribution, λ = 0.071, k = 5). Thus, any 5 SNPs that are in such close vicinity are likely of artificial nature due to mapping errors and should not be considered. In contrast, the linkage groups of the two remaining noncentromeric SNPs cover a genetic distance of 4.7 cM (ratio = 0.49) and 6.8 cM (ratio = 0.44), respectively. Both SNPs lie in the rounded peak that we visually identified on the upper arm of chromosome I and are neighbors (Figure 3, red shading).

Of the two visually and statistically identified 1:1 class SNPs, the SNP with a ratio of 0.44 was intronic whereas the SNP with a segregation ratio of 0.49 (the closest to 0.5 in the whole data set) (Figure 3A, arrow) was a nonsynonymous GC-to-AT nucleotide change, which is characteristic of most EMS-induced SNPs (Sega 1984). This nucleotide change produces a stop codon in the sixth exon of gene At1g16570, a putative UDP-glycosyltransferase superfamily protein. To demonstrate that the causative SNP was identified, the At1g16570 gene was amplified from each of the 53 DNA samples that had been pooled for sequencing (File S1). The PCR products were digested with nucleases cleaving single-base-pair mismatches in heteroduplex DNA (Till et al. 2004). Using this method, 52 samples were shown to have a SNP at the indicated position, while one sample was not cut (Figure S2). The progeny of this plant showed no phenotype, indicating that it was a sampling mistake due to wrong phenotyping in the BC2 generation. Finally, T-DNA insertion lines disrupting the identified gene At1g16570 were tested for a pollen-tube reception phenotype. The line SAIL_400_A01 (tun-2), which has an insertion in the fourth exon of At1g16570, displays the same pollen-tube overgrowth phenotype as the EMS allele tun-1 (Figure 2C), indicating that the correct gene has been identified by SRM.

In this example, no further analyses were required to identify the causative SNP. However, if genome coverage or mapping quality of the reads is lower than expected, the application of several filtering strategies could narrow down the list of potential candidate SNPs. First, selecting for exonic SNPs (nonsynonymous, synonymous) removes most of the detected SNPs (Table S2). Second, prioritizing characteristic EMS-induced SNPs (Sega 1984) should unambiguously identify the causative SNP in most cases. If not, then the rare cases where the mutation affects a regulatory region that is not exonic or an atypical EMS-induced nucleotide change have to be considered.

Interestingly, we also observed SNP ratios > 0.5. Since this should not be possible considering our genetic backcrossing strategy (Figure 1), we performed a detailed analysis of all SNPs on chromosome I with ratios > 0.5. All such SNPs display a low coverage (low-sample-size effect) or are covered by reads with low mapping quality (Figure S1). This indicates that such high ratios might be mapping artifacts in repetitive and/or transposable element regions. In addition, any SNPs found in and around the centromere display unusual segregation ratios (Figure 3, gray shading), probably representing an intrinsic problem in mapping sequence reads to highly repetitive centromeric regions.

On the whole, labor-intensive, map-based cloning has been replaced by cloning via NGS in recent years. Until now, this worked only (i) with homozygous viable mutants using the SHOREmap or similar strategies (Schneeberger et al. 2009; Austin et al. 2011) after outcrossing or (ii) for organisms with a short generation time and an easy-to-score phenotype, where multiple required backcrosses still save time (Zuryn et al. 2010). In contrast, SRM enables the mapping of zygotic or even gametophytic lethal mutations after only two rounds of backcrossing. SRM is generally applicable, but the identification of heterozygous mutant individuals may require progeny tests, e.g., scoring for the presence of aborted seeds or defective embryos among the progeny. This might involve some adaptations of the crossing scheme shown in Figure 1, including a combination of inter se crosses and backcrosses if selfing is not possible. In outcrossing species, individual males could first be crossed to siblings to identify the heterozygotes in a progeny test, as well as to wild-type females to generate the backcrossed progeny used for SRM. In fact, SRM could also be used for the cloning of the causative genes on the basis of homozygous mutants in F2 populations: the expected SNP ratios would be different (1.0 vs. 0.5), but the approach would still benefit from the small number of individuals required.

SRM is especially advantageous for (i) lethal mutations, (ii) organisms with a long generation time, (iii) hard-to-score phenotypes, and (iv) mutations with low transmission because only a small number of individuals are needed. Importantly, SRM is the method of choice for second-site modifier screens, in which a mutant with a certain phenotype is mutagenized a second time to identify novel mutant alleles that enhance or suppress this phenotype. Again, classical mapping or the SHOREmap strategy (Schneeberger et al. 2009; Austin et al. 2011), which rely on outcrossing and an F2 mapping population, require the original mutation to be present in at least two genetic backgrounds. This is possible only when another allele is available in a different accession or by outcrossing the mutation five to six times to another accession. This procedure is time-consuming and has the disadvantage that, due to a lack of recombination events close to the mutation, additional enhancer/suppressor mutations in the vicinity of the original mutation cannot be mapped. Furthermore, second-site modifier screens are often performed in complex, tailor-made backgrounds involving several mutants and/or transgenes (Page and Grossniklaus 2002). It is very hard to generate the identical genetic/transgenic constitution in two distinct accessions. By using SRM, the enhancer/suppressor mutant has to be backcrossed only to the original mutant background, no matter how complex it is.

Finally, SRM can be applied to any genetic system. In fact, we expect that SRM can also be applied in organisms without a well-annotated genome. As mentioned above, plotting the SNP/non-SNP ratio over the chromosomal positions and visually identifying flat curves indicating a region under selection can be statistically tested (File S1 and File S2) and used to identify the causative SNP. This is also possible with partly assembled and poorly annotated genomes.

In conclusion, we successfully identified a gene disrupted in a gametophyte lethal mutant in A. thaliana by SRM. The need for relatively few individuals and only two rounds of backcrosses, which are needed in any case to purify the genetic background after mutagenesis, make this a very versatile and useful method for any genetic organism.

Supplementary Material

Supporting Information

Acknowledgments

We thank A. Patrigiani for preparing the library for SOLiD sequencing and R. Schlapbach for access to the facilities of the Functional Genomics Center Zürich. We also thank S. A. Kessler and J. Jaenisch for helpful comments on the manuscript. This work was supported by grants from the European Research Council and the Swiss National Science Foundation to U.G. The costs for next-generation sequencing were covered by a project of the University Research Priority Program in Functional Genomics/System Biology of the University of Zürich.

Note added in proof: While our paper was under review, Abe et al. (2012) published a similar method based on the distinct segregation ratios of linked and unlinked SNPs to map homozygous mutants in rice. This shows that SRM can be applied to map homozygous mutants, as we suggested in our manuscript.

Footnotes

Communicating editor: C. D. Jones

Literature Cited

  1. Abe A., Kosugi S., Yoshida K., Natsume S., Takagi H., et al. , 2012.  Genome sequencing reveals agronomically important loci in rice using MutMap. Nat. Biotechnol. 30: 174–178 [DOI] [PubMed] [Google Scholar]
  2. Austin R. S., Vidaurre D., Stamatiou G., Breit R., Provart N. J., et al. , 2011.  Next-generation mapping of Arabidopsis genes. Plant J. 67: 715–725 [DOI] [PubMed] [Google Scholar]
  3. Blumenstiel J. P., Noll A. C., Griffiths J. A., Perera A. G., Walton K. N., et al. , 2009.  Identification of EMS-induced mutations in Drosophila melanogaster by whole-genome sequencing. Genetics 182: 25–32 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Irvine D. V., Goto D. B., Vaughn M. W., Nakaseko Y., McCombie W. R., et al. , 2009.  Mapping epigenetic mutations in fission yeast using whole-genome next-generation sequencing. Genome Res. 19: 1077–1083 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Jander G., Norris S. R., Rounsley S. D., Bush D. F., Levin I. M., et al. , 2002.  Arabidopsis map-based cloning in the post-genome era. Plant Physiol. 129: 440–450 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Lukowitz W., Gillmor C. S., Scheible W. R., 2000.  Positional cloning in Arabidopsis: Why it feels good to have a genome initiative working for you. Plant Physiol. 123: 795–805 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Page D., Grossniklaus U., 2002.  The art and design of genetic screens: Arabidopsis thaliana. Nat. Rev. Genet. 3: 124–136 [DOI] [PubMed] [Google Scholar]
  8. Sarin S., Prabhu S., O’Meara M. M., Pe’er I., Hobert O., 2008.  Caenorhabditis elegans mutant allele identification by whole-genome sequencing. Nat. Methods 5: 865–867 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Schneeberger K., Ossowski S., Lanz C., Juul T., Petersen A. H., et al. , 2009.  SHOREmap: simultaneous mapping and mutation identification by deep sequencing. Nat. Methods 6: 550–551 [DOI] [PubMed] [Google Scholar]
  10. Sega G. A., 1984.  A review of the genetic effects of ethyl methanesulfonate. Mutat. Res. 134: 113–142 [DOI] [PubMed] [Google Scholar]
  11. Smith D. R., Quinlan A. R., Peckham H. E., Makowsky K., Tao W., et al. , 2008.  Rapid whole-genome mutational profiling using next-generation sequencing technologies. Genome Res. 18: 1638–1642 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Srivatsan A., Han Y., Peng J., Tehranchi A. K., Gibbs R., et al. , 2008.  High-precision, whole-genome sequencing of laboratory strains facilitates genetic studies. PLoS Genet. 4: e1000139. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Till B. J., Burtner C., Comai L., Henikoff S., 2004.  Mismatch cleavage by single-strand specific nucleases. Nucleic Acids Res. 32: 2632–2641 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Zuryn S., Le Gras S., Jamet K., Jarriault S., 2010.  A strategy for direct mapping and identification of mutations by whole-genome sequencing. Genetics 186: 427–430 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting Information

Articles from Genetics are provided here courtesy of Oxford University Press

RESOURCES