Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Sep 15.
Published in final edited form as: Mutat Res Rev Mutat Res. 2023 Sep 15;792:108471. doi: 10.1016/j.mrrev.2023.108471

Next-Generation Sequencing Methodologies To Detect Low-Frequency Mutations: “Catch Me If You Can”

Vijay Menon 1, Douglas E Brash 1,2,3,*
PMCID: PMC10843083  NIHMSID: NIHMS1933139  PMID: 37716438

Abstract

Mutations, the irreversible changes in an organism’s DNA sequence, are present in tissues at a variant allele frequency (VAF) ranging from ~10−8 per bp for a founder mutation to ~10−3 for a histologically normal tissue sample containing several independent clones – compared to 1%−50% for a heterozygous tumor mutation or a polymorphism. The rarity of these events poses a challenge for accurate clinical diagnosis and prognosis, toxicology, and discovering new disease etiologies. Standard Next-Generation Sequencing (NGS) technologies report VAFs as low as 0.5% per nt, but reliably observing rarer precursor events requires additional sophistication to measure ultralow-frequency mutations. We detail the challenge; define terms used to characterize the results, which vary between laboratories and sometimes conflict between biologists and bioinformaticists; and describe recent innovations to improve standard NGS methodologies including: single-strand consensus sequence methods such as Safe-SeqS and SiMSen-Seq; tandem-strand consensus sequence methods such as o2n-Seq and SMM-Seq; and ultrasensitive parent-strand consensus sequence methods such as DuplexSeq, PacBio HiFi, SinoDuplex, OPUSeq, EcoSeq, BotSeqS, Hawk-Seq, NanoSeq, SaferSeq, and CODEC. Practical applications are also noted. Several methods quantify VAF down to 10−5 at a nt and mutation frequency (MF) in a target region down to 10−7 per nt. By expanding to >1 Mb of sites never observed twice, thus forgoing VAF, other methods quantify MF <10−9 per nt or <15 errors per haploid genome. Clonal expansion cannot be directly distinguished from independent mutations by sequencing, so it is essential for a paper to report whether its MF counted only different mutations – the minimum independent-mutation frequency MFminI – or all mutations observed including recurrences – the larger maximum independent-mutation frequency MFmaxI which may reflect clonal expansion. Ultrasensitive methods reveal that, without their use, even mutations with VAF 0.5% - 1% are usually spurious.

Keywords: Next-Generation Sequencing, Duplex Sequencing, Low-Frequency Mutations, Rare Variants, Variant Allele Frequency

1. Introduction

1.1. Detecting mutations by sequencing

Somatic mutations are present at much lower frequencies than inherited mutations (“germline mutations” in genetics). Mutation induction was therefore traditionally studied in cell culture, with mutations detected as rare colonies of cells newly-resistant to killing by a selective agent [1, 2]. Advances in genomic sequencing platforms over the past two decades have allowed rapid, accurate, and cost-effective sequencing of nucleic acids in a wide variety of samples, providing in-depth knowledge of altered nucleic acid sequence and how such a mutation’s change in genomic information content contributes to the etiology of various diseases. Moreover, Next-Generation Sequencing (NGS) has allowed detection of these changes at very low frequencies, providing a useful tool for toxicology, studies of tissue dysfunction and its occasional progression to disease, patient prognosis, and development of mutation-targeted therapies. The principal barrier has been the detection limit, set by the DNA sequencing instrument, errors introduced during PCR, and DNA damage present on one strand of the parental DNA template. This review discusses methods used to breach that barrier. The excitement generated by these advances often exceeds familiarity with underlying mutational processes and terminology, so we first summarize these in order to allow precise comparison of the rare-mutation methods and their results.

1.2. Origins of a mutant cell

Genetic mutations arise directly from deamination of 5-methylcytosine to thymine at 37 °C; from the inherent fidelity limit of DNA polymerases operating during DNA replication or repair; and from polymerase errors at sites of DNA tautomerism, cytosine deamination at body temperature, or DNA damage caused by metabolites or exposure to environmental factors. Depending on the mutagen, these mutations mainly occur at a single DNA base pair, that is, a single nucleotide (nt) on the reference strand but can be small (<< 50 bp) deletions or insertions or larger (>> 50 bp) structural variants such as amplifications, translocations, large deletions and insertions, and somatic copy number alterations [310]. Somatic mutations, which occur in non-germ cells [11, 12], are acquired during the lifetime of an individual and are implicated in proliferative diseases such as benign lesions and cancer [13, 14], aging [15, 16], and even neurodegenerative diseases including Parkinson’s disease [17], autism [18], schizophrenia [19], and Alzheimer’s disease [20]. Distinct from mutation creation is the subsequent clonal expansion of the single mutant cell into a clone [2123]. Clone expansion is commonly viewed, particularly in the oncology field, as driven by a self-sufficient “driver mutation”. But this is not the only route. First, the mutation may confer a selective advantage only in the presence of an external selection agent such as sunlight or hypoxia [2426] or an advantage only relative to adjacent cells [2729]. Or it may be a weak driver, termed a mini-driver [30]. Second, even a neutral mutation will clonally expand if it occurs in a cell that previously or subsequently acquired a driver mutation (“passenger mutation”). Third, in tissues with stochastic stem cells, any mutant will clonally expand by neutral drift, a completely random process in which some clones – mutant or not – expand while others contract and disappear (Fig. 1) [31, 32].

Figure 1:

Figure 1:

Stochastic clonal expansion (green clone and arrow) or expansion and contraction (yellow clone and arrow) in a tissue having stochastic stem cells. Arrow thickness corresponds to the size of the clone being pointed to. Computer simulation, based on data parameters measured using Confetti multi-labeled cells in mice. Adapted with permission from [32] (DOI: 10.1242/dev.060103).

Unlike detecting mutations as drug-selected clones in a cell culture dish or by immunohistochemistry [21, 25, 26], DNA sequencing cannot determine whether multiple sequence reads of a particular mutation arose from a site that was often mutated independently in many cells, i.e., has a high mutation frequency, versus arising from a mutant clone. Any claim about clone sizes or the frequency of founder mutations that is deduced from sequencing will rely on a statistical argument. Another subtlety comes in deducing a putative clone’s size from sequencing data. If a biopsy contained a million cells, 3 independent mutations in a run having 30,000x coverage depth (VAF 10−4) implies a substantial clone of 200 mutant cells.

The expansion of clones is more than a technical point: Mutagens create mutant cells in proportion to the dose, but clonal expansion is exponential over time because two daughters make four granddaughters and so forth [26]; correspondingly, tumor rates in rodents and humans are proportional to the first or second power of dose (dose1−2) but the fifth or sixth power of time (time5−6) [33].

There are subtleties to the origin of a mutant cell, even in cancer. Tumor evolution can require subclones containing more than one mutant driver gene, called “subclonal mutations” relative to the main sample. Most tumor types, and even normal tissues, also show mutations in non-driver genes at low frequencies; these are often also called subclonal mutations, although sequencing alone cannot establish their clonality or relation to the main clone [3437]. Finally, some tumors appear to be polyclonal [38]. Resolving such issues is important for early diagnosis and designing chemotherapeutic intervention.

1.3. Mutation frequencies expected

Expectations for mutation frequencies in tissue are largely based on data from traditional studies of mutations to drug-resistance: drug-resistant colonies emerging on tissue culture dishes 10–14 days after a single mutagen treatment and reflecting one generation of mutation fixation. Mutation frequency is defined as:

MF(drug-resistant colonies after mutagen treatment)/(surviving colonies after mutagen-only) (1)

Those frequencies ranged from 10−5 – 10−3 mutations per survivor per target gene for UV or polycyclic aromatic hydrocarbons in normal human cells [39, 40]. The range depends on the mutagen, dose, selective drug (usually acting on an enzyme), and number of bp in the assay’s gene target that can mutate to a scorable phenotype (small for gain of function and large for loss of function). Mutation frequencies are ~100 fold higher in repair-defective cells. Assuming that the underlying target is an exonic region of ~1 kb [41], these induced average mutation frequencies (MF) in normal cells equate to 10−8 – 10−6 mutations per nt; if mutations were concentrated at a smaller target near the enzyme’s active site [42], the MF was perhaps 10 fold higher, 10−7 – 10−5 mutations per target nt. At particular nt termed “mutation hotspots”, induced mutation frequencies are 10-fold higher than average [4347], giving a hotspot nt variant allele frequency (VAF) of 10−6 – 10−4 mutations per nt.

What is needed, then, are DNA sequencing methods that can detect an average MF of 10−7 – 10−5 mutations per nt or VAF of 10−6 – 10−4 mutations per nt.

The method’s error rate must therefore be lower than those mutation frequencies. The background error rate of standard Illumina sequencing (the technology most widely used) is VAF ~5 × 10−3 per nt (section 3), 50-fold higher than the highest expected VAF of independent hotspot mutations and at least 500-fold higher than the average mutation frequency across the gene. The background error rate of standard nanopore sequencing is MF ~15% (section 3).

Mutations observed above the 0.5% frequency range are thus a surprise and reflect additional factors at play. Candidate factors include:

  • DNA damage hyperhotspots, nt that can be 500-fold more sensitive to UV radiation [48, 49]

  • DNA repair slow spots, at which damage persists or has more time for deamination [50, 51]

  • Sequencing’s elimination of the need to express a drug-resistance phenotype [52].

  • Accumulation of in vivo exposures over years.

  • Clonal expansion, making non-independent mutations [21, 52, 53].

  • Very small biopsy sizes, which will allow a clone to occupy a majority of the sample [54].

  • Artifactual mutations introduced after DNA isolation.

The first four factors are on the order of 10–100 not 1,000, so they don’t bring the frequency above the detection limit.

Clonal expansion is an important factor. Clones in skin or esophageal epithelium range in size from ~32 cells for neutral drift over the course of a year [31, 55] to ~3000 cells for a TP53-mutant clone with growth driven by UV exposure [21, 26]. When the independence of mutations is unknown, as in sequencing, the 10–1000 fold amplification by clones could bring the range of non-hotspot VAFs to ~10−6 – 10−2 mutations per nt. Thus any mutations detected by using standard sequencing are only the largest of clones.

The technical challenge is clear. Section 4 will focus on methods that remove artifactual mutations and reach these detection limits.

1.4. Mutation detection by DNA sequencing vs drug selection assays

It is important to note key differences between mutation assays that mutagenize cells in a dish versus DNA sequencing of tissue already mutagenized in vivo..

  • Observing colonies mutagenized on a dish (or using immunohistochemistry on a tissue) identifies independent founder mutations and can therefore measure the frequency of independent mutations. DNA sequencing mixes founder mutations and their clonal descendants, so it cannot itself distinguish a nt having a high mutation frequency from one conferring a phenotype that favors clonal expansion.

  • A drug selection assay necessarily assays a gene that is expressed. DNA repair is most efficient in expressed genes [3] and the mutation frequencies are lower there [56]. Frequencies might be higher elsewhere and these will be seen with sequencing. This constraint also applies to drug selection of cells isolated from mutagenized transgenic mice.

  • A drug selection assay necessarily assays base changes that change an amino acid [52].

  • Mutation frequency using drug selection is expressed as mutations per survivor. The survival curve is measured on separate dishes, as clonal survival, i.e., ability to proliferate. Single, non-proliferating live cells are not observable. Whether the non-proliferating or truly dead cells had a higher mutation frequency has been a matter of speculation for decades. When sequencing DNA from cells in culture or in tissue, in contrast, genomes from non-proliferating or truly dead cells are typically included except for cells that floated off the culture dish.

  • Mutations in a human tissue typically reflect multiple exposures, as well as the opportunity to clonally expand over the course of years.

2. Terminology

Metrics used in NGS mutation analysis such as variant allele frequency, mutation frequency, mutation prevalence, and mutation burden are pivotal in specifying the genetic diversity within a sample and the magnitude of each variation. In reviewing the NGS methods developed to detect low frequency mutations, we encountered not only synonyms but also different uses for the same term. Often the quantity being calculated was not defined at all or the definition was ambiguous. In particular, the phrase “total number of variant bases” does not convey whether the paper is counting “different mutations”, i.e. scoring mutation types (each being the first occurrence of a particular base change at a particular nt of a particular gene), versus “all mutations”, i.e., scoring mutation instances including recurrences originating from additional cells having the same mutation. Clinical and commercial scorings lean toward the former, while mutagenesis studies lean toward the latter, without making their choice clear. We therefore define key terms below.

2.1. Variant Allele Frequency (VAF)

Variant Allele Frequency (VAF) [5759] is the most clearly defined term. Also called Variant Allele Fraction (VAF) [54, 60] or Mutant Allele Frequency (MAF) [61, 62], it is the fraction of haploid genomes in the original sample that harbor a specific mutation (a specific ‘variant allele’) at a specific nt coordinate of the reference genome:

VAF(genomes with a specific mutation at nt)/(genomes examined at nt) (2)

It has also been called Minor Allele Frequency (MAF) [63], but that term has the technical definition as the frequency of the second most common polymorphism variant in a human population. To maximize clarity we advise using “VAF”.

The units of frequency are mutations per nt or bp. The term variant allele “fraction” makes it clearer that “mutant nt per total nt” is actually a dimensionless quantity, but “frequency” is entrenched in the mutagenesis field. Unfortunately for 21st century mutagenesis conversations, the biologist considers the word “frequency” to entail “counts per x” (as in neural spikes per second) but to the bioinformaticist or statistician the word means only “counts”. When the bioinformaticist incorporates a “per x” to make what s/he calls a “rate”, the “mutation rate” achieves the meaning of the biologist’s “mutation frequency”. But to a biologist in the mutagenesis field, “mutation rate” already has a technical definition, “mutation frequency per cell generation”.

In standard sequencing, VAF is calculated as the total number of mutant reads divided by the total number of reads (the depth of coverage), at the specific genome coordinate. Using sequencing, it is unknown whether the VAF reflects many independently generated mutant cells, several small clones having the same mutation, or one large clone. Hence the quantity above could be considered a maximum independent-variant frequency, VAFmaxI. In sequencing terms:

VAF=(readcount for a specific mutation at nt)/(total readcount at nt) (3)

Because methods for detecting rare mutations use consensus among PCR progeny to discard spurious reads (section 4), the raw readcount in (3) will be replaced by the count of initial parental DNA fragments having the same mutation, each parent sequence being the consensus among its progeny family. The appropriate denominator will be replaced by the parent depth, the count of mutant parents + wildtype parents, not the count of all parental progeny.

VAF=(parent count for a specific mutation at nt)/(total parent count at nt) (4)

The genome coordinate is also referred to as the reference genome coordinate, genome position, or nucleotide position [5759]. Less precise terms are locus and genomic location, which in the pre-sequencing literature refer to entire genes. Anchoring to the reference genome, rather than the experimental genome, is important for the bioinformatics because of the possibility of deletions, insertions and, for RNA sequencing, alternate splice forms. Where clones have arisen, as in a tumor, VAF can be large and is often reported as percent [52, 64].

The theoretical detection sensitivity for VAF is limited by two factors:

  1. The number of genomes sequenced. To detect mutations at a VAF of 10−5 per nt, 105 genomes must be sequenced.

  2. The ability to distinguish distinct parental genomes. Barcodes are used to mark each parental molecule (the unique molecular identifiers, UMI, section 4) because otherwise reads could merely be PCR duplicates. UMIs must have more variants than the number of genomes being sequenced.

In practice, detection sensitivity is also limited by the accuracy of the sequencing methodology – the artifactual background mutations arising from damaged DNA templates or sequencer errors. When the number of genomes and barcodes used in an experiment are known, it is possible to decide whether a method has reached its theoretical limit or was limited by practical issues.

2.2. Mutation Frequency (MF)

Mutation Frequency (MF) is an average of mutation occurrences across many nt and without limiting to a particular base change. Before the advent of DNA sequencing, this was all one could do because the base change was unknown and the location was known only as a particular gene whose mutation conferred drug resistance. MF for mutation to drug resistance was measured as colonies on cell culture dishes [1, 2] and was calculated as in definition (1), with the units of mutations per survivor. That is, “mutations per two gene copies” in a surviving cell. Because each colony arose from an independent founder mutant cell, this MF unambiguously quantified independently generated mutations without complication from the extent of clonal expansion. This history conferred the connotation that MF measured the founder mutations created in the first cell generation after mutagen exposure [65, 66], the technical term being “mutation rate” and the units being mutations per cell generation.

A more flexible interpretation, used by the sequencing community, is that it is a measure of the frequency of mutations in the sample or biopsy at the moment, whether the mutant cells are founders or daughters and whatever number of exposures or years were required to accumulate them. The units are mutations per bp. Unknown quantities can be deduced from experimental data using MF = mutation rate (as mutations/bp per cell generation) * cell generations [65].

Crucially, extension from cell culture dishes to tissue, using transgenic mice with drug selection or using DNA sequencing, renders it unknown whether a recurrent mutation in the sample resulted from independently mutated cells or an expanding clone. Quantification of the in vivo assay in the literature has been inconsistent in treating clonal expansion. The clearest solution is to calculate both minimum-independent and maximum-independent MFs [67].

MFminI counts only the first occurrence of a particular base change at a particular nt, and is accurate if every instance of a recurrent mutation came from the same clone descended from a single founder mutant. Only different mutations are counted as independent, without regard to VAF except perhaps for a threshold in calling the mutation. This is likely an undercount, especially undercounting mutation hotspots.

MFminI(count of all different mutations in target region)/(nt examined)=(count of 1st occurrences of mutations in target region)/(nt examined)=(count of 1st occurrences of mutations in target region)/(sum of(genomes' nt examined at each target nt))=(count of 1st occurrences of mutations in target region)/(average coverage depth * target length in reference genome) (5)

MFminI is often used as an estimate of the initial, pre-expansion mutagenicity of the mutagen involved. This is a somewhat precarious number, inasmuch as a nt position’s readcount of 1 instead of 0 depends heavily on the size of the DNA sequencing run. The binomial error bar is large and most counts of 1 are “unlucky” 2s rather than very rare mutations that were “lucky” in being loaded onto the sequencer. This fact argues for using a threshold readcount of 2 or 3.

MFmaxI instead includes recurrences; it represents the case that every instance of a recurrent mutation originated from an independent founder mutation event, analogous to the original MF in cell culture. That is, it factors in each site’s VAF; indeed, MFmaxI at a single nt is VAF. A high value could be due to DNA damage or mutation hotspots, particularly for a VAF at a single nt, but for MF over a region it is likely an overcount of independent founders in vivo because of clones.

MFmaxI(total count of all occurrences of mutations in target region)/(nt examined) (6)

In methods involving consensus between PCR copies of the same parental molecule, the terms ‘mutations in target region’ and ‘nt examined’ refer to the original parent molecules, after any PCR copies have been merged into a consensus sequence for their parent.

In the literature, MFminI is reported in [59, 68]. MFmaxI is reported in [69]. As mentioned, providing the “total number of variant bases” does not eliminate the ambiguity between “number of different departures from the reference genome” (variant mutation types, MFminI) and the larger “number of variant bases in the sample, including recurrences” (variant mutation instances, MFmaxI). We recommend that papers use phrasing like “different mutations without including recurrences” or “all mutations including recurrences” to make the scoring clear, or adopt the unambiguous terms MFminI and MFmaxI. Where we cannot tell which metric was used, this review will use MF.

MFmaxI and MFminI can differ by orders of magnitude when clones are present. At most nt sites, the VAF varies by only ~10-fold, whereas clone size can easily vary from 2 to 3000 [21, 25]. This circumstance warrants close attention, as it provides an opportunity for a report to inflate the number of mutations detectable by a sequencing assay (by reporting MFmaxI) or lower a method’s detection limit or a test-compound’s mutagenicity (by reporting MFminI). Conversely, a biopsy smaller than the clone size increases the VAF, because there is less dilution by cells not mutated at that nt; this increased VAF will also increase the discrepancy between MFminI and MFmaxI.

A useful property of MF is that averaging over the large collection of target bases boosts detection sensitivity, at least for the averages needed in toxicology: The 105 cells that allow a VAF detection limit of ~10−5 per nt will allow detection of mutagenesis at an MF of 10−9 per nt when a 10 kb target is used. This extent of sequencing is feasible at current sequencing costs. Alternatively, MF provides a tissue sample-size advantage: To detect an MF of ~10−5 per nt, one can sequence 1 nt in 105 cells, 100 nt in 1000 cells, or 105 nt in 1 cell. Coverage depth is traded for coverage “width”. (The term ‘coverage breadth’ is used for the number of bp in a target that have a coverage depth meeting a depth threshold [70].

2.3. Tumor Mutation Burden (TMB) and Mutation Burden (MB)

Tumor mutation burden (TMB) is a “width” MF expressed over a region larger than “per nucleotide”, such as mutations per Mb, genome, cell, or tumor. One purpose is to generate a number greater than 1 that characterizes a typical cell in the tumor. Another purpose, particularly when a VAF threshold is stipulated for each mutation in the count, is to measure the diversity of mutations in the tumor. Diversity is a driver of tumor evolution [71] and might dictate the response to checkpoint inhibitor therapies. In a tumor each mutation is presumed to already be a large clone, so only the different mutations are counted and even the subclones are of interest. Thus TMBminI seems to be preferred in the tumor literature and is used by clinical mutation burden tests [72]. These considerations lead to the definition:

TMBminI(number of different mutations in a sample)/(target length in reference genome) (7)

This metric is straightforward in a homogeneous tumor, but a heterogenous tumor presents subtleties: If the tumor sample contains 10 subclones, the mutation burden in each may be ~1/10 of “the TMB” computed by (7). Is it the burden per cell or the 10 fold larger value per tumor that is relevant for prognosis and comparing tumor types? Should small subclones be given less prominence? A caveat is that excluding mutations having a VAF below a threshold, which in turn depends on the sequencing run size, biases the computation [72].

Definitions of TMB vary more widely than MF and are seldom clearly stated, weakening its use as a technical term. Variations on TMBminI include: total number of different somatic missense mutations per tumor, typically filtered to focus on oncogenic drivers [7375]; number of different somatic mutations per megabase of the genome [76]; and somatic nonsynonymous mutations per megabase of the “coding genome” [77, 78]. Where immunogenicity is of prime interest, TMB has been defined as TMBminI excluding oncogenic drivers [79].

TMB has also been used to indicate the fraction of tumors carrying a mutation in a particular gene [80], but this usage better falls under Prevalence (section 2.4). TMB has been considered a measure of disease or treatment prognosis [81], mutagen exposure [82], or a cell’s DNA damage response and repair pathways [8385]. This is plausible, but MFminI does not incorporate clonal expansion.

When extending the mutation burden concept to rare mutations in non-tumor tissue, the definition of MB is often unclear but appears to be analogous to (7):

MBminI(number of different mutations in a sample)/(target length in the reference genome) (8)

Yet with rare mutations, the impact of small clones and independently mutated single cells can no longer be ignored [54]. An alternative metric that represents individual contributions can therefore be defined as an average VAF across the target region [54]:

MBmaxI(sum of VAFs for each different base change at each nt of target region in sample)/(target length in reference genome) (9)

If the coverage depth is the same at every nt, this is simply:

=(sum of mutation counts)/(coverage depth * target length in reference genome)=MFmaxI

This is easier to see when VAF is considered as the dimensionless Variant Allele Fraction. When the unit of target length is Mb, the burden is in mutations per Mb.

2.4. Mutation Prevalence (MP)

In epidemiology, prevalence refers to the proportion of individuals in a population who harbor a trait or disease. By extension, mutation prevalence is the proportion of subjects [86], patients [87], tumors, or tissue samples containing a specific mutation or a collection of mutations in target regions. The proportion of genomes would be VAF or MF. The term is sometimes used for the MF in the target region when averaged across samples or individuals [65, 66]. Similarly, “recurrent” is conventionally used for mutations or sequence reads from a single sample, but it has been used to refer to a mutation located at a particular genomic position recurring across samples [88]. This review adopts the single-sample usage.

Because of the variable usage of terms, we have converted results to VAF or MF wherever possible.

3. Evolution of NGS error rates

Nucleic acid sequencing has witnessed massive advances since the development of the first sequencing methodologies in the 1970s.

First-Generation Sequencing: Developed in the 1970s, the two main sequencing methods were Sanger’s dideoxy method for terminating a DNA primer extension at one of the four nucleotides in a template [89] and the Maxam-Gilbert chemical method for cleaving a DNA strand at a subset of the four nucleotides [90]. The particular nt’s position is determined as the length of the DNA fragment upon gel electrophoresis of many molecules averaged together. Sanger’s method is the basis for most later methods and is still utilized for sequencing single DNA fragments as in molecular cloning and PCR amplicons. Because it reflects an average for thousands of genomes or PCR progeny whose signals add together, detection of variants is limited to those above ~10% of the sample. With the development of high-fidelity DNA polymerases, Sanger sequencing has a low error rate, often stated as 0.001% (1 in 100,000 nt) [91]. It is important to recognize that this number represents the polymerase’s error rate on a perfectly cloned template. Real templates contain errors that occurred during early cycles of plasmid amplification or PCR, often caused by DNA damage on the parental DNA molecule. Sequencing merely reports this pre-existing error. Those early errors occur at a level of 1 in 100–1000 nt, but will constitute 50–100% of the sequencing signal at that nt. When determining a germline sequence or even haplotypes, the problem is easily circumvented by sequencing several parental molecules separately. For identifying rare true mutations, other approaches were needed.

Second-Generation Sequencing: Termed Next-Generation Sequencing (NGS) and developed in the early 2000s, NGS revolutionized DNA sequencing by massively parallel sequencing of multiple different DNA targets arrayed on a solid surface. The target is locally amplified on the surface, after which sequencing proceeds by synthesis as in the Sanger method. For our purpose, more important than the high throughput is that NGS reports the sequence of individual molecules, providing a digital quantification that enables allele detection sensitivity below 10%. NGS error rates are the error rate on individual reads. The major NGS platforms include: Roche/454, based on pyrosequencing in which light is emitted upon nt incorporation [92, 93] and having an error rate of ~1%; Ion-Torrent Sequencing, similar to pyrosequencing but measuring hydrogen ions released during nt incorporation [94] and having an error rate of ~1.7% (with a reported drawback being inability to correctly sequence homopolymer sequences); and Illumina sequencing, which incorporates fluorescent nucleotides [95] and having an error rate of 0.1–1% that varies with nt position.

Third-Generation Sequencing: In early 2010s, Third-Generation Sequencing (TGS) technologies sequence a single single-stranded molecule passing through a microscopic pore without DNA amplification, ensuring uniform genome coverage depth without PCR bias. In Pacific Biosciences Single Molecule Real Time sequencing (SMRT) [96, 97], long double-stranded DNA fragments are ligated to two adapters that link the strands and allow a polymerase to proceed along this single strand. The 10–15 kb DNA reads facilitate studying haplotypes, poorly mappable genome regions, transcriptomes, and epigenetics. Oxford Nanopore Sequencing [98100] sequences a single strand of DNA or RNA. The error rate of Nanopore sequencing is 5–15% per nt [101]. The error rate of SMRT was initially ~15% [102], with errors being random and unbiased; its circular structure enables improvements discussed later.

Fourth-Generation Sequencing: Understanding the spatial and temporal organization of the genome within cells and tissues in situ provides a better knowledge of structural elements within cells, mechanisms of chromatin folding in different cell populations, and genome maintenance. In situ genome sequencing [103, 104], genomes in intact cells are imaged and sequenced concurrently. The error rate has not been stated; if it becomes low enough, DNA sequencing may be able to identify geographically contiguous true clones and thus study clonal evolution and tumor heterogeneity.

4. Detecting rare variants by NGS

The chief barriers to lowering the detection limit of NGS are polymerase errors during PCR, cluster amplification, or the sequencing polymerase; the presence of a DNA lesion on one strand of the parent molecule, typically an oxidized purine or deaminated cytosine, with conversion to a mutation during library preparation; or a mutation on one strand such as results directly from deamination of 5-methylcytosine. These limits have been overcome by successive innovations in error-correcting strategies that work by comparing the sequences of a family of PCR progeny strands to determine a consensus sequence. Upon solving this challenge, the detection of rare mutant cells becomes limited by the biopsy itself: the extent to which a mutant cell or small clone gets diluted by the unmutated cells in the same biopsy. Given a small enough biopsy, even standard DNA sequencing will yield VAFs that are real rather than background errors [54].

Five numbers are important in comparing consensus techniques:

  • Conversion efficiency. The percent of initial duplex fragments (after adapter ligation, PCR, and any reduced representation treatment, e.g., targeted capture or restriction digest size selection) that appear in an adequately large consensus family [105]. This can be 10% or less. The limitation is not the amount of sample that can be loaded onto the sequencer but the facts that parental duplexes drop out during processing or there may not be enough PCR progeny of a parent strand present in the sequencing run; increasing run size has little effect. Most of the sample is wasted, and increasing the biopsy size would just dilute a mutant clone.

  • Consensus family size. The number of sequence reads from progeny of the same parent duplex that are required to deduce a reliable consensus sequence. Each family represents one initial DNA duplex fragment from one genome.

  • Coverage depth. The lowest attainable VAF detection limit is 1/(coverage depth). Coverage depth refers to the number of independent consensus families, and thus parental genomes, not sequence reads. Genome-wide methods provide low coverage depth at any one nt, so cannot measure VAF and measure only MF.

  • Target size. The lowest attainable MF detection limit is 1/(coverage depth * target size).

  • MFminI or MFmaxI. For signal and background error, is the paper reporting only different mutations or all recurrences?

Few papers provide all this information or the data needed to confirm or calculate it, much less ascertain whether mutations at the sensitivity limit have been detected with 99.9% power.

It is also important to note whether a method’s reported detection limit is based on synthetic mutations constructed by serial dilution of SNPs or other known variants (providing the “internal standard” detection limit for the method, e.g., [106]; [107]; [108]), or based on the lowest number of mutations observed in untreated cells after ignoring obvious high-frequency spontaneous mutations (a “residual mutations” detection limit, e.g., [65]; [109]). The internal standard limit is often lower and it avoids complications from the DNA substrate to which it is applied, such as sporadic true mutations or clonal expansion. However by focusing on a small number of sites, it can miss the rare, randomly distributed spurious mutations present in the sample or introduced by the technique, which are of concern here. Assessing residual mutations in cells detects those substrate-based artifacts, but the measured detection limit will include both the internal standard limit of the assay and any low-frequency true sporadic mutations.

4.1. Single-strand consensus sequencing

Single-strand strategies compare the sequences of all the PCR progeny of one strand of a parent molecule to build a single-strand consensus sequence (SSCS); a lack of consensus identifies an error made during amplification and this base is not called as either normal or mutant. In principle, a modification of these methods could build a consensus from a mixture of progeny of both strands, albeit without identifying the parental strands separately. An important technical challenge applicable to any consensus method is that subsampling, due to DNA losses or using aliquots of a reaction, will introduce stochastics into the likelihood of having multiple members of the same family of progeny present in the same sequencing run.

4.1.1. Safe-SeqS

The Safe-Sequencing System or Safe-SeqS technique [65] is one of the earliest methodologies for NGS error correction using consensus between PCR progeny of the same parent. Its principles are also important for understanding all later methods. To allow grouping of one parent’s progeny, each parent molecule is marked with a unique molecular identifier (UMI), also called a tag, molecular barcode, or unique ID (Fig. 2). A 14- nt random UMI is added to one end of the progeny of (one strand of) the initial DNA molecule using amplicon-specific primers and two cycles of PCR. Such primers generate ~3×108 distinct UMIs, minimizing the likelihood that two different parent molecules acquire the same UMI (sometimes called “tag clash”). Two PCR cycles incorporate the same UMI onto top and bottom progeny strands. These uniquely tagged early-progeny strands are amplified using universal sequences present in all primers, which also contain Illumina P5/P7 sequences for sequencing. The amplification creates many daughter molecules having the same UMI sequence, thus forming a UMI family. If the mutation is real, all members of the UMI family will contain the same sequence change relative to the reference genome, creating a SSCS (termed ‘supermutant’ in the paper). If a mutation was present on only the tagged strand of the parental duplex (e.g., deaminated 5-methylcytosine), or if it occurred in the first PCR cycle (due to a polymerase error or damage on the template strand), it will be present in 100% of that strand’s progeny. Mutations arising in later amplification cycles will be present in <100% of progeny and are identifiable as spurious mutations.

Figure 2:

Figure 2:

Single-strand consensus sequence (SSCS) methods: Safe-SeqS and SiMSen-Seq. Adapted from [65, 105, 110].

Using this strategy, ~100,000 normal human cells from each of 3 unrelated individuals (~500,000 genomes total) were sequenced within a <75 bp region of the CTNNB1 gene to determine the residual mutation background VAFs. The lowest achievable VAF detection limit is thus theoretically 2×10−6 per nt and the MF lower limit is ~7×10−8. UMI family size decreased exponentially from size 2, with a mean of 68 members. The conversion efficiency of initial genomes to being a member of a UMI family was 78% for family size ≥2 and ~70% for size ≥12. The consensus sequence (SSCS) was considered a true mutation if the UMI family contained at least 2 members and >95% of members showed the same mutation.

Over 99% of UMI families were single nucleotide substitutions, with reported somatic VAFs of ~6×10−5 (Fig. 4 of ref. [65]) and an MF of 9×10−6 (Table 2 of ref. [65]). By comparison, standard sequencing gave an MF of 2×10−4, so Safe-SeqS reduced the background by ~25 fold. That is, 96% of the mutations detected by conventional sequencing were spurious. There is still room for a further 100-fold improvement, inasmuch as this paper’s coverage depth and target bases permit an MF detection limit of ~9×10−8. Alternatively, the mutations might be real; however, the individual nt with the highest background VAF in conventional sequencing were the same nt having highest background in Safe-SeqS, in all three individuals, suggesting that they were spurious and the trustable level is somewhat higher.

Figure 4:

Figure 4:

Parent-strand consensus methods: DuplexSeq. Adapted from [105, 106].

4.1.2. SiMSen-Seq

Simple, Multiplexed, PCR-based barcoding of DNA for Sensitive mutation detection using Sequencing [110, 111] is a modification of Safe-SeqS designed to prevent its large amounts of off-target amplicons or amplification failures, which the present paper found to be due to mis-priming by the random barcodes and which would prevent multiplexing. The innovation is a PCR primer in which the 12-nt barcodes are shielded by binding in a 14-nt hairpin loop that remains closed during the annealing temperature but linearizes (“open state”) at the PCR extension temperature (Fig. 2, inset). This allowed multiplexing target sites and detecting mutations at a VAF of ≤ 0.1% (1×10−3).

Briefly, genomic DNA is mechanically sheared followed by 3 cycles of PCR amplification to add barcodes onto 5–100 ng of DNA (1500–33,000 genomes); low concentrations of the primer and long annealing times are used to facilitate multiplexing by avoiding primer dimers. After terminating the reaction by protease treatment, the amplicons are amplified 18–30 cycles using standard Illumina primers for sequencing and purified using SPRI beads. The barcodes are used to create consensus read sequences for each parental DNA molecule.

From 1 to 31 genomic targets were amplified from a tumor cell line; reads were required to be 100% identical within strand families containing 10 to 20 reads or be ≥ 95% identical within strand families having > 20 reads. For 5 amplicons and a ≥30 read family cutoff, 10% of raw reads were converted to SSCSs; SSCS depth was 7,700 per ~83 bp amplicon, or 93 per bp. To determine the sensitivity, plasma DNA pooled from >10 normal individuals was spiked with ~0.6% and ~0.06% tumor DNA harboring a TP53 mutation. The VAF for the 0.06% spike-in TP53 mutation was found to be just above background.

SiMSen-Seq has been used subsequently to identify UV-induced C→T mutations at a melanoma mutation hotspot in the RPL13A promoter at TCTTCCG, at a VAF <0.5% [112]; detect mutations at RPL13A and DPH3 promoters in human keratinocytes and melanoma cells treated with UV daily for 5 or 10 weeks, at VAFs up to 2.9% [113]; and detect UV-induced mutations in post-mortem eyelid and nose, at VAFs of ~1% and 5.5% at one position of the RPL13A gene and 0.5% and 0.8% at another position, as well as at VAFs of ~2% and ~4% at one position of the DPH3 gene and ~1% and ~3% at another [114].

4.2. Tandem-replicate consensus sequencing

4.2.1. o2n-Seq

The innovation in o2n-Seq [63] is to produce 2 replicas of a parental strand in tandem within a single paired-end read (Fig. 3, left). If a variant is detected on only one of the tandem replicas, it was a PCR error. This approach generates independent copies of one strand of the parental molecule by linear amplification, before lower-fidelity PCR is used to make copies of copies that propagate previous errors. It also improves efficiency by ensuring that every insert sequenced has a SSCS consensus partner, without needing a UMI-incorporation step to identify unlinked consensus family members. The strategy was originally developed to reduce the error rate in standard NGS sequencing [115], but its rolling circle amplification introduced amplification biases. The modified version used here pushes the error rate lower by using a high-fidelity polymerase and minimizes library bias by amplifying for only four cycles. An underlying assumption is that the parent strand was perfect, with no mutation residing on just one strand (such as deaminated 5-methylcytosine) nor a DNA lesion that is always mutagenic.

Figure 3:

Figure 3:

Tandem-replicate consensus methods: o2n-Seq and SMM-Seq. Adapted from [63, 116].

Operationally, the genomic DNA is fragmented to ~100 bp, so that two tandem replicas can be sequenced within a paired-end read. After ligating the parental duplex fragment to adapters containing a unique nickable site (uracil in this case), but no UMIs, the adapter-ligated duplex is denatured and each of the single-strands is circularized. After circularization of, e.g., the parental top strand, second-strand synthesis and nicking at the unique site create two complementary nicked open-circles. Strand-displacement polymerases then extend each nicked circle to create a tandem replicate. The result is two new duplexes. One contains two independently-copied versions of the original parental top strand (which, however, might have contained a mutagenic DNA lesion); the other contains the parental top strand and a copy of a copy of that parental top strand. The “o2n”-ed DNA is used to make standard NGS libraries and the consensus sequence is determined.

Using mixtures of two E.coli strains having ~304 polymorphic sites, almost all sites at 1% synthetic VAF were detected. The background MF was 1×10−5 per bp, reducible to 10−8 if only recurrently mutated sites were counted. Most mutations were of the types caused by cytosine deamination and oxidative damage to purines. Sequencing mixtures of two polymorphic ϕX174 strains and 10,000x coverage depth, the polymorphisms were detectable at 0.1%. Samples were also taken from 1 cytologically normal and 2 tumor regions of a section of hepatocellular carcinoma. Using a 420 kb capture probe panel targeting 687 bp known to be mutated in this tumor, VAFs observed ranged from 0.3% – 4% in the normal region and 0.3% – 9% in the tumor regions.

4.2.2. SMM-Seq

Single Molecule Mutation Sequencing or SMM-Seq [116] is “single-molecule” in the same sense as o2n-Seq: sequencing (progeny of) tandem independent copies of a parental molecule strand which were generated before PCR made copies of copies that propagate previous errors. The innovation is to use using rolling circle amplification to generate multiple independent copies of the parent strand, and do so for both top and bottom strands of the parental molecule (Fig. 3, right). The resulting UMI family of combined top and bottom strand progeny establishes a consensus sequence. In addition, the complexity of the genome is reduced by restriction fragment digestion and size-selection rather than by targeted PCR primers or capture probes.

The DNA is first fragmented using a double restriction digest with AluI and MluCI, end-repaired, A-tailed, ligated to adapters, and SPRI purified. Two stem-loop adapters are ligated whose double-stranded region carries a 6-nt random UMI sequence, allowing 2×107 parental molecules to be distinguished. The adapters’ single-stranded loop contains a primer binding site that initiates rolling circle amplification on both strands of the insert, in opposite directions (“pulse RCA” in the paper). The resulting “single-stranded DNA contigs” are then amplified by PCR and sequenced. Families derived from the same parental molecule are identified by the UMI and contain reads from both strands of the parental duplex. Disagreement among family members indicates a spurious mutation. Recurrence over different families were not counted, thus quantifying MFminI.

Assuming that sequencing errors arise during PCR cycles at an error rate P(E), the paper points out that the likelihood of fortuitously creating the same error at the same position of N copies of the parent is P(E)N, much lower than the P(E)2 obtained by comparing progeny of top vs bottom strand (see section 4.3). However, when artifactual mutations arise from DNA damage on one parent strand, the tandem strategy will be similar to SafeSeq-like approaches.

The information reported does not make it possible to determine whether the restriction digest and size purification using SPRI beads confer a reduced representation of the genome sufficient to allow recurrent coverage depth and a VAF measurement, rather than constraining to MF. The reduced representation will limit the genes that are assayable.

To test the method, normal human IMR90 fibroblasts were treated with the mutagen N-ethyl-N-nitrosourea (ENU) and DNA harvested after 72 hr. Whether cells were proliferating or confluent is not stated. The conversion efficiency to a consensus sequence was 45% and the ~200 million consensus bp examined would allow an MF detection limit of 5×10−9. MF in untreated cells was 2.1×10−7 per bp and after 25 μg/mL ENU rose to 3.6×10−7, a 70% increase. At 50 μg/mL, MF was 5.4×10−7 per bp. Induced mutations were primarily T:A→A:T and T:A→C:G substitutions. If the 0 ENU control is taken as indicative of the method’s sequencing error background, DNA extracted from young and aged human liver tissues had MFminI of approximately 1×10−7 and 7×10−7 per bp, respectively, with the predominant mutation increased being T:A→C:G.

4.3. Parent-strand consensus sequencing

These strategies compare the sequences of the PCR progeny of the top strand of a parent molecule with progeny from the bottom strand; a lack of consensus identifies mutations present in just one parent strand (such as deaminated 5-methylcytosine) or mutations caused by DNA damage in just one parent strand (deaminated cytosine or oxidative damage). These events turn out to be a significant obstacle to reaching lower mutation detection limits.

4.3.1. DuplexSeq

The error-correction methods described so far efficiently reject PCR errors arising after the first PCR cycle. They are less robust for first-cycle errors and are ineffective against mutations that arise from a mutagenic DNA lesion present on one strand of the parent molecule before sequencing library preparation began. Duplex Sequencing (DuplexSeq) was developed as a solution to this first-cycle problem [106]. The innovation is to tag each strand of the parent DNA fragment with an independent UMI before PCR (Fig. 4). Progeny of the top strand are merged into a single-strand consensus family (SSCS), as in Safe-SeqS; the same is done for the progeny of the bottom strand. True mutations are located at the same base position on both SSCSs, creating a duplex consensus sequence (DCS). Spurious mutations are present in only one strand’s SSCS. In theory, Safe-SeqS would detect such errors as being present in 50% of the combined progeny, and it may catch most of them. In practice, stochastics intervenes and lower detection thresholds can be reached by distinguishing 100% from 0% rather than from 50%.

To achieve this, the insight was that the standard Illumina Y-adapter is Y-shaped because it contains different sequences on the two mismatched arms (the P5 and P7 graft sequences that hybridize to the flow cell in paired-end sequencing). If the double-strand portion of the Y-adapter contains a random UMI and each end of a DNA fragment receives a different adapter, then the progeny of the parental top strand will have as their top strand the sequence:

5’ P5 Arm-UMI1-insert-UMI2-P7 Arm 3’

whereas the progeny of the parental bottom strand will have as their top strand the sequence:

5’ P7 Arm-UMI1-insert-UMI2-P5 Arm 3’.

Operationally, the Y-adapters were each given a random 12-nt UMI, letting a pair distinguish 424 = 3×1014 parent duplexes. Two oligos, one containing the 12-nt random tag, were annealed and the complementary tag sequence on the other oligo was generated by strand extension followed by A-tailing. For generating sequencing libraries, DNA from M13mp2 phage or human brain mitochondrial DNA was mechanically sheared, end-repaired, and SPRI bead purified. ~200–500 bp DNA fragments were isolated, T-tailed, ligated to adapters, and PCR amplified. After sequencing, only families containing ≥ 3 members with ≥ 90% of these members having identical sequences at a specific position were considered to generate the SSCS for that position. The DCS conversion efficiency, the fraction of initial DNA duplex fragments that are converted into DCSs, was ~1% in these early experiments.

To determine the sensitivity of DuplexSeq, the authors first sequenced bacteriophage M13mp2, which clonal genetic assays in the literature had established to have an MFminI of 3×10−6 per base. Analyzing the data without using the strand-specific tags gave an MFminI of ~4×10−3, so the 1000-fold increase was due to spurious mutations. Applying SSCS analysis to the same dataset, in analogy to Safe-SeqS, gave MFminI ~3×10−5, 100-fold better but still 10-fold higher than the expected frequency and suggesting that 90% of the rare mutations detected by SSCS were still spurious. The base substitutions were of the type caused by cytosine deamination and oxidative damage to purines. Finally, DCS analysis gave an MFminI of 2.5×10−6, the frequency expected from genetic measurement, i.e., probably real mutations, emphasizing the importance of detecting a mutation at the same position on both DNA strands. The duplex sequencing strategy thus achieves our goal of detecting MF of 10−7 – 10−5 per nt.

Assuming that DCS eliminates all but pairs of complementary SSCS errors occurring at the same position on opposite strands, then the background error frequency is: (probability of error on one strand) * (probability of error on other strand) * (probability that both errors are complementary) = 1/3 * (SSCS error frequency)2. From the measurements above, the theoretical DuplexSeq MFminI detection limit is ~4×10−10 per nt, if there are no technical surprises blocking actual progress below 10−6.

To test the sensitivity for low-frequency VAFs, known M13mp2 variants having specific mutations were mixed in different ratios to construct internal standards. DCS analysis was performed at a coverage depth of ~20,000, which would permit a VAF sensitivity of 5×10−5 per nt. They were able to detect mutations at twice this VAF, 10−4. In a 10 kb target, this VAF sensitivity would permit an MFminI sensitivity of ~10−8 per bp. Applying the method to brain mitochondria, they found an MFminI of 3.5×10−5, with the 100-fold lower figure relative to conventional sequencing being largely due to discarding the transversion mutations at G resulting from oxidative damage rather than pre-existing mutations.

Subsequent improvements optimized the family size to 6–12, so that top and bottom strand progeny are both present with minimal sequencing and increasing DCS conversion efficiency to 10% [117]. Adding two rounds of target capture provided 97% on-target capture [118]. Capture is typically designed for a single strand, so the other strand used in the DCS comparison will actually derive from a PCR copy. Using fragmentase instead of mechanical shearing for DNA fragmentation reduces the level of artifactual mutations created at single-strand ends during the end-repair step of library preparation [105, 119]. The reagents for the DuplexSeq approach are available commercially [52] and appear to use ~100 pre-defined barcodes, allowing a VAF detection limit of ~10−5 when insert endpoints are included. It would therefore allow an MFminI of 10−9 if a 10 kb target is sequenced.)

A striking finding (Fig. 5) was that ignoring mutations with counts below the Illumina error limit of 0.1–0.5% does not guarantee quantifying true mutations: Many mutant sites with recurrence above this threshold are still spurious, rather than valid large clones.

Figure 5:

Figure 5:

Mutations in a human HRAS transgene in mice exposed to the carcinogen urethane, sequenced using standard Illumina sequencing (top) and DuplexSeq (bottom). Modified with permission from ref. [52]. The figure is copyrighted under Creative Commons Attribution License 4.0 (CC BY) (https://creativecommons.org/licenses/by/4.0/).

DuplexSeq has been used to detect low-frequency mutations in various experimental or clinical situations.

  • In cancer patients, detecting pre-treatment mutations that confer resistance to cancer therapeutics such as tyrosine kinase inhibitors might guide the choice of therapeutics [69]. Mutations were detected at VAF 8×10−5, but pre-existing mutations known to confer kinase resistance did not correlate with relapse. However, monitoring for mutations during treatment detected new kinase resistance mutations 5 months before relapse in half the patients.

  • The choice of whether to treat breast cancer by conservative lumpectomy rather than mastectomy might be guided by the absence vs presence of mutations in the uninvolved normal gland tissue; these mutations could have expanded into clones in the course of normal mammary gland development [120]. Targeted sequencing of 542 cancer-associated genes in uninvolved mammary glands of 52 breast cancer patients identified mutants in PIK3CA, TP53, AKT1, MAP3K1, CDH1, RB1, NCOR1, MED12, CBFB, TBX3, and TSHR in subsets of patients. Allelic frequencies reached 9% to 50%, suggesting clonal expansion. Of 4 patients in whom standard NGS detected PIK3CA or TP53 mutations in the primary tumor but not in uninvolved mammary gland, the duplex sequencing approach identified low-frequency PIK3CA mutations in uninvolved glands in three and low-frequency TP53 mutations in one.

  • Detecting mutations in mice treated with putative genotoxins could simplify carcinogenicity testing compared to standard plaque mutation assays and greatly speed it relative to two-year rodent tumor assays. Treating Big Blue mice with vehicle control, benzo[α]pyrene, or ENU for 4 weeks gave MFminI of 1.5×10−7, 1.1×10−6, and 1.3×10−6 per nt. The ratios of treated to vehicle were close to the MFminI per gene obtained using the standard plaque assay [52]. The method also detected carcinogen-specific mutation spectra.

  • In male germ cells, point mutations increase with paternal age. An examination of mutations in the coding region of the FGFR3 gene in sperm DNA from young (26–30 yr) and old (49–59 yr) donors identified specific sites having VAFs of 0.01–0.001% and an MFminI across the gene of ~6×10−7 [59]. The older donors showed 2–3 fold higher VAFs at some sites, and some mutations were absent in the young donors. One site identified at a VAF of 0.03% was c.1118A → G (p.Y373C), known to underlie the congenital disorder, thanatophoric dysplasia type I.

If only presence or absence of mutations is required, the capture step of DuplexSeq can use probes designed to anneal preferentially to the mutant allele, reducing the amount of sequencing by ~100 fold in a technique named MAESTRO [121].

4.3.2. PacBio HiFi

PacBio SMRT was originally a mixed-strand tandem-replicate consensus sequencing method; with both top and bottom strands sequenced in a short alternating tandem array and compared as a group to achieve a consensus sequence. In a later development, incorrect PacBio base calls were rectified by circular consensus sequencing, the repeated sequencing of the same molecule to a “family size” of 6–60 depending on insert size [122]. The greater consensus power reduced artifactual MF to ≤0.5% per nt [123] and this could be further reduced using statistical models [124]. Limiting the sequenced insert to ~1 kb, to allow more passes (which reduced errors up to ~ 10 passes), and increasing the coverage depth from the typical 5–25 up to ~10,000, allowed measurement of MFmaxI down to 10−6 per bp and VAF down to 10−4, as judged by residual mutations in untreated cells [125]. A sequencing lane contains 8 million molecules, so greater coverage depth is feasible. The tandem array also makes it possible to separately identify the two strands for parent-strand consensus, a bioinformatics addition termed HiFi sequencing. The physical linkage between top and bottom strands and the absence of PCR avoid the stochastic loss of one strand. With this approach, MF is detected at 10−7 per bp [126], as judged by residual mutations in untreated cells. We are unaware of VAF measurements. The method therefore suffices for MF measurements and would detect the VAF of medium sized clones.

4.3.3. SinoDuplex

SinoDuplex sequencing [107] is aimed at detecting mutations at a VAF of ~0.1% per nt in liquid biopsies of cell-free DNA (cfDNA) from blood plasma. The innovations are the use of error-correcting barcodes in the UMI of pre-synthesized adapters, so that sequencing errors within the UMI do not cause loss of the read or assignment to the wrong DCS family, and a Bayesian SSCS-calling approach that facilitates using family sizes as small as 2. Compared to random barcodes, pre-defined barcodes also more easily identify which PCR progeny have the same genomic coordinates due to originating from the same parental molecule. Using a small set of such barcodes additionally simplifies this identification and reduces the cost of pre-synthesis. The number of barcode combinations must exceed the desired coverage depth in order to assign PCR progeny to the correct parent, but a small number suffices for the small genome numbers in cfDNA and can be expanded using endogenous barcodes. The sequencing library is prepared using Y-shaped adapters that contain one of 16 pre-defined UMI sequences (7–8 nt long), each at least 3 edit distances from the others. That is, at least 3 independent errors are needed to convert one UMI to another. In SinoDuplex, there are thus 162 = 256 different UMI combinations possible for an insert. This allows a sequencing depth of 256 distinct parental molecules, and ~30,000 if the insert endpoints are also used to distinguish parental molecules. A depth of 30,000 would have a VAF detection limit of 3×10−5 per nt.

Targeted sequencing was carried out using three different probe panels. Using serial dilutions of reference DNA as internal standards containing a range of known polymorphic allele frequencies, the mean DCS coverage depth for ~25 ng reference samples (~8,000 genomes) was ~2,000, varying with the genomic target region. For 8,000 genomes, this is 25% DCS conversion efficiency of initial genomes to DCS families, higher than many methods and presumably resulting from the Bayesian algorithm. A DCS depth of 2,000x would allow a detection sensitivity of ~5×10−4 per nt. Correspondingly, the authors were able to detect mutations at a VAF threshold of 0.1% (1×10−3; presumably very large clones) with ~99% sensitivity and ~97% specificity. The measured allele frequency was equal to the frequency expected from the dilution, from 0.1% to 10%. Mutations present at 0.1–2% VAF were detected in cfDNA from 5 of 6 patients.

4.3.4. OPUSeq

OPUSeq, for “One-Pot dsUMI Sequencing” [127], is a modification of DuplexSeq used for detecting VAFs as low as 0.01% (1×10−4) per nt. The innovation is that, instead of incorporating UMIs into the adapters before ligation, OPUSeq ligates partially single-stranded adapters containing a 6 nt random UMI; these are then filled in in situ during the library PCR steps. This resembles the enhanced ligation method used in SaferSeq (see below) but the details differ. The one-pot approach makes it possible to incorporate dsUMIs in the same reaction as the library PCR, eliminating the need for costly pre-synthesis of adapters and reducing the number of enzymatic steps.

Two 6 nt random UMIs permit a coverage depth of 412 = ~2×107 and thus VAF detection limit of ~5×10−8. To test the actual sensitivity of OPUSeq for detecting low-frequency mutations, two DNA samples with and without a heterozygous single-nucleotide polymorphism (SNP) were mixed to produce VAFs from 0–1%. Using 1 μg of this mixed DNA (330,000 genomes, allowing a VAF detection limit of 3×10−6), capture was carried out using a panel targeting a 427 bp region of the HRAS gene and OPUSeq adapters were compared to standard commercially available adapters. Using fragmentase for fragmentation, the conversion efficiency of initial duplexes to SSCS families was ~7%, conversion of SSCSs to DCS families (“DCS recovery rate” in the paper) was 22%, and net conversion of initial duplexes to DCSs (DCS conversion efficiency) was ~2%, yielding a duplex depth of ~7400, which would allow a VAF detection limit of ~10−4 (0.01%). The loss of expected families could not be compensated for by additional sequencing, so it was due to drop-out of parental duplexes or early amplification products during processing. The SNP spike-ins were detectable above background down to 0.01%. The paper then examined an HRAS region of 427 bp to determine the average MF (termed “variant incidence”). Using fragmentase, MF was 2.6×10–5 per nt, which the authors state to be 100-fold the estimated human somatic mutation rate of 3×10–7.

Using sonication and 66,000 genomes, which would permit a VAF detection limit of ~2×10−5 and, for 427 bp, an MF of 4×10−8, they found conversion of initial duplexes to SSCS families of ~1.4%, conversion of SSCSs to DCS families again ~22%, and net DCS conversion efficiency of 0.3%, yielding a duplex depth of ~4441. However, MF background was 25-fold lower, 1×10−6 per nt, not far from somatic mutation rate. It could plausibly be the biological reality in this target region and cell line. The authors surmise that the elevated background with fragmentase was mainly due to fragmentase-mediated DNA shearing followed by end-repair of DNA ends, resulting in base changes on both strands that could not be removed computationally. The authors conclude that OPUSeq is capable of correctly calling variants down to a VAF of 0.01% and MF of 1×10−6 per nt.

4.3.5. EcoSeq

EcoSeq, for “Enzymatically Cleaved and Optimal Sequencing” [108] is a version of DuplexSeq that, like SMM-Seq, adds the innovation of enzymatic reduced representation. Instead of mechanical shearing and capture, genomic DNA is cleaved with BamH1 and size-selected to 100–700 bp. This reduces the genome size ~100-fold (30 Mb per human genome), without changing the copy number of any particular genome fragment; this allows increased coverage depth for the same amount of sequencing. The reduction is not targeted enough for measuring VAFs, so this strategy is aimed at measuring MF.

Starting with 33,000 genomes, BamHI digestion created 109 parental duplex BamHI fragments (termed “copies” in the paper). The maximum DCS conversion efficiency of 2% was obtained by using 106 of these fragments (33 genome equivalents), a randomization that will prevent VAF measurement even at the reduced representation. To circumvent the low ligation efficiency of the paper’s DuplexSeq adapters, the BamH1 cleaved site is partially filled using dATP and dGTP and then 5’-TC-tailed duplex loop adapters (termed “v1”) are ligated to the sticky ends.

Defined frequencies of rare mutations, prepared by mixing DNA with and without SNPs at ratios down to 0.01% as internal standards, were detected with 40 M reads of sequencing, at the expected average MFmaxI down to 2.4×10−8 per bp. Background mutations elsewhere were observed at ~1×10−6, which decreased to ~1×10−7 upon cloning out lineages, so they reflected mutations acquired in subclones of the cell culture. This result indicates that constructing synthetic mutation internal standards by serial dilution is a better indicator of the method’s sensitivity than detecting endogenous mutations. Next, in a stable cell line treated with the mutagen 4-nitroquinoline-1-oxide, G:C→A:T transitions and G:C→T:A transversions were detected at an MF (presumably also MFmaxI) of <10−6 per bp. Finally, analysis of normal peripheral blood cells from pediatric sarcoma patients revealed a mutation frequency of ~9×10−8 per bp in patients without chemotherapy and ~31×10−8 per bp in patients who underwent chemotherapy. Analysis of mutation accumulation before and after chemotherapy in 6 patients revealed persistence of mutations in these patients many years post-chemotherapy. A limitation of EcoSeq was an increased G+C% content in analyzed fragments as compared to the whole genome.

4.3.6. BotSeqS

The Bottleneck Sequencing System, BotSeqS [66], utilizes DuplexSeq’s strategy of separately barcoding the top and bottom strands of the parent duplex, with three innovations: a) Using the entire genome as a target rather than specific regions, thus avoiding targeted capture or PCR. b) Rendering that target feasible with affordable amounts of sequencing by introducing a 105-fold dilution step (the “bottleneck”) after ligating Illumina Y-adapters but before the library amplification, so that only a random subsample of the sheared genomes are examined; for their starting 165,000 genomes, this is the equivalent of 1.65 final genomes. Thus a nt is never seen twice and the method measures only MFminI, not VAF or MFmaxI. c) Using as the UMIs the small number of “endogenous barcodes” inherent in the range of genomic coordinates at the ends of the randomly-sheared inserts, a strategy first considered in a variation of Safe-SeqS [65]. By reducing the number of genomes being examined, dilution also increases the number of PCR cycles possible before the PCR plateau, increasing the size of a DCS family and the likelihood of containing progeny of both the top and bottom strands of the parental DNA molecule. The first read sequence (P5-insert-P7) marks progeny of the parental top strand and the second read sequence (P7-insert-P5) marks progeny of the parental bottom strand. After identifying the single-strand consensus sequence (SSCS) of the what the authors term the Watson and Crick strand duplicate-families, the two SSCSs are combined into a duplex consensus sequence, DCS.

To test the method, 44 libraries were generated from normal tissues of 34 individuals, then diluted. A sequence change was analyzed only if each parental strand had ≥2 progeny, termed Watson and Crick duplicate-families, and was considered a true mutation only if the sequence change was observed in ≥90% of each of the two duplicate-families. A mutation present on both strands of more than one parental molecule was considered clonal and the recurrences were excluded, thus quantifying MFminI. For a library starting with 165,000 haploid genomes, 105-fold dilution yields 5 million 500 bp fragments; the DCS conversion efficiency of diluted input DNA to families containing both Watson and Crick duplicates was 45%, with duplicate-family sizes of ~10. With 2×88 bp effective paired-end sequencing, the 3.5×108 DCS consensus nt sequenced permits an MF detection limit of 3×10−9 per bp. After filtering out DCSs in repetitive regions, the DCS conversion efficiency was 1.2% (equivalent to 0.2% of a haploid genome), giving ~60,000 unique molecules for final MF analysis and an MF detection limit of 1×10−7 per bp. The method’s ultimate MF detection limit was estimated by measuring the frequency of discordance between Watson and Crick duplicate-families at known SNP sites, which will be present in most libraries, then extrapolating from the square of that frequency. This extrapolated MF limit was ~3×10−12 per bp, a valid extrapolation if there are no technical surprises blocking actual progress below 10−7. The observed MFminI averaged over 25 control individuals was 5×10−7 per bp in nuclear DNA and 10−5 per bp in mitochondrial DNA. Nuclear DNA from 2 individuals having a mismatch repair deficiency found a value ~130 fold higher, 7×10−5 mutations/bp. Further, the authors report that the MFminI increases in an age-dependent fashion at a rate that varies between somatic tissues. Comparison between parental strands was important because without it nuclear mutations were ~10 fold more frequent, primarily due to G→T transversions known to often be caused during library preparation. A limitation of any method that randomly samples the genome is that it sees any particular nt only once or twice, so it cannot identify SNPs in a sample; it must rely on databases. Another limitation of random sequencing is that it needs to discard a large fraction of the data because much of the genome is not uniquely mappable and is thus not present in the reference genome. In addition, in BotSeqS the fraction of DCSs filtered out due to mapping to repetitive regions or other structural variants was 96%. This may reflect a non-random property of shearing, which the authors noted, and it is evidently important when UMIs use only endogenous barcodes.

4.3.7. Hawk-Seq

Hawk-Seq [128] uses the same method as BotSeq, considered from the point of view of starting with small amounts of genomic DNA rather than viewing this as a dilution and a bottleneck. Its goal is to achieve simplicity and high throughput for toxicology studies, so it adopts the lack of exogenous barcodes and specific genome targets and instead uses endogenous barcodes and random genome fragments. Small initial DNA amounts allow more PCR progeny from fewer genomic regions, increasing the likelihood of having PCR progeny of both top and bottoms strand of the same parental molecule loaded on the sequencer. Small DNA amounts also avoid having two parental genome fragments that have the same shear endpoints (tag clash of endogenous barcodes); this is especially relevant in small genomes like the Salmonella TA100 used in toxicity tests, which will have more copies of each genome region in the same number of ng of DNA.

A series of input DNA amounts were first tested to find the maximum input DNA amount that produced reads from both top and bottom strand progeny and maximized “sequencing efficiency”, a metric like conversion efficiency but measuring the conversion of sequence reads (rather than input DNA) to DCS families. For 50 million reads and Salmonella’s 5 MB genome, the optimum input was ~100 attomol of 350-mer fragments, or 20 pg. At higher inputs, DCSs were lost due to dropout from insufficient PCR and mutations were lost due to miscalling from tag clash. This 2×1010 bp is ~4000 Salmonella genomes and allows a theoretical MF sensitivity of 10−10 per nt. Sequencing efficiency was ~8% and only 1% of the reads were involved in a tag clash. Family size for constructing a DCS was chosen to be ≥1 read for each parental strand and in practice the average was ~1.5. The mean fraction of the genome receiving at least 1 DCS was 26% and the mean DCS depth was 1. As expected for random sampling, VAF is therefore not measurable.

The method was next used to measure mutations and mutation spectra in TA100 resulting from several mutagens, reporting MFmaxI down to 6×10−5 per bp. The method was then extended to the mouse gpt delta mutagen assay system, reporting MFmaxI down to 7×10−6 per bp. The murine starting DNA amount was not given, so the theoretically achievable sensitivity cannot be estimated; sequencing efficiency and DCS depth were also not stated. The method has subsequently used for TA100 testing of other mutagens, including alkylating agents, epoxides, aldehydes, aromatic amines, aromatic nitro compound, and polycyclic aromatic hydrocarbons [129].

4.3.8. NanoSeq

Nanorate Sequencing or NanoSeq is named after the MFminI error rate sensitivity level it achieved [130]. The method combines DuplexSeq adapters tagging top and bottom strands separately and having random UMIs, BotSeqS’s random genome sampling by dilution, and innovations to improve the quality of the DNA fragments used for the sequencing library. Its description as “single-molecule” sequencing refers to the fact that any particular DNA fragment is seen only once, so it measures MFminI not VAF or MFmaxI. In initial studies of cord blood using BotSeqS, the authors had found that their background MFminI of 2×10−7 mutations per bp came from a large excess of C→T and G→T/C point mutations near the 5′ ends of sheared DNA fragments, as well as from the complementary base change on the other strand but with base substitutions extending along the entire read. These were presumed to arise when end-repair of the sheared fragments before library preparation filled in overhangs or extended interior nicks.

NanoSeq reduces this background by using the following steps: a) Sonication is replaced by fragmenting DNA with the blunt-end restriction enzyme HpyCH4V, thereafter size selecting for 250–500 bp using SPRI beads and then dilution to achieve reduced representation. Any particular DNA fragment is thus typically encountered at most once, so only MFminI will be measured. b) The end-repair/A-tailing reaction includes dideoxyT/G/CTP, which block nick extension. c) Particularly effective, computational filters against alignment errors and ambiguous mapping are included in the bioinformatic analysis. In addition, the method employs a qPCR step to ensure an optimal UMI-duplicate family size, independent of starting DNA amount. An alternative was mechanical shearing followed by mung bean nuclease to remove 3’ and 5’ single-stranded overhangs. DCS conversion efficiency (“efficiency, E”) was optimized computationally and found to be ~12%.

These changes reduced the background 250-fold to an MFminI = < 5×10−9 per bp (<15 errors per haploid genome). Having established this level of sensitivity, the authors then measured MFminI in differentiated granulocytes versus their stem cell precursors. If (and only if) stem cells are a permanent cell type distinct from their differentiating cell progeny and protected from mutagenesis [131], one would expect the differentiated cells to have accumulated many more mutations during the ~28 cell divisions required to produce the number of differentiated cells produced in a year. Instead, the mutation rate was almost identical at ~20 per diploid genome per year; the same was seen in colonic stem cells versus differentiated colonic epithelium and in non-dividing smooth muscle cells.

4.3.9. SaferSeqS

SaferSeqS [132] returns to the PCR approach to targeting, introducing a Duplex Sequencing improvement on the PCR-based Safe-SeqS. Like DuplexSeq, it separately tags the top and bottom strands of the parental DNA, but it targets specific genomic regions by using hemi-nested PCR after adapter ligation instead of hybrid capture (Fig. 6). The method also uses a recent technique [133] for increasing the efficiency of adapter ligation by using adapter modifications that allow high adapter concentrations to be used without mis-targeting or making primer-dimers. Such intricacies are needed for PCR-targeting methods because the thermodynamics of distinguishing target from non-target puts an upper limit on the concentration of target DNA that can be used, and thus restricts the number of genomes and the mutation detection limit [134]. Sequencing reads descended from the same parent, identified using both primer-added and endogenous barcodes, are categorized into top and bottom single-strand consensus families (SSCS). If there is >80% match, they are a mutant DCS family (“duplex family”) and represent a true mutation in the parental DNA molecule.

Figure 6:

Figure 6:

Parent-strand consensus methods): SaferSeq. Adapted from [132].

The PCR approach is partly motivated by the finding that capture raises the DNA damage and mutational backgrounds and lowers the DCS conversion efficiency of parental DNA to duplex consensus families [135]. The prolonged 65°C incubation introduces DNA damage after the initial processing and the ~0.50 efficiency of capturing a parental strand leads to less than (0.50)2 for capturing both top and bottom strands from the same parent and processing both without loss from subsampling.

Operationally, circulating cell-free DNA in plasma (cfDNA, ~167 bp) or sheared DNA from leukocytes is first treated with USER enzyme to nick DNA containing uracil from in vivo spontaneous cytosine deamination, making those fragments unPCRable. DNA fragments are end-repaired by dephosphorylation and blunting and then ligated to adapters that create Y-like mismatches which tag the two strands separately. Ligation proceeds in a sequential fashion: The 3’ ends of the top and bottom strands are first ligated to the 5’ end of a long “3’-adapter” that contains a 14-nt random UID sequence (identifying 7×1016 parental duplexes) on that strand and a short oligo as the other strand to facilitate ligation. The short oligo contains a 3’-blocking group and U to make it degradable after the ligation. Ligation is followed by annealing, to the outer end of the 3’ adapter, the 3’ end of a largely mismatched “5’-adapter”. This is followed by polymerase extension to generate in situ the complementary strand of the random UMI. After ligation, the result is a double stranded UMI sequence with a Y mismatch at the ends, as in DuplexSeq. Adapter-specific primers then drive whole-genome PCR. As with DuplexSeq, progeny of the parental top strand have the order 5’ Left Arm - UMI1 - insert - UMI2 - Right Arm 3’; the bottom strand has the converse.

The second phase of library preparation supplants capture with strand-specific hemi-nested PCR of specific ~75 bp target regions to generate top- and bottom-strand PCR amplicon families. One primer is chosen from within the gene target. The other primer is a strand-specific anchor primer located in the adapter. A subtle point is that the strand-specific anchor primer will bind to both top- and bottom-strand templates, but it will extend into the insert only when its target is in the 3’ adapter; if it anneals to the same sequence at the 5’ end of progeny of the other parental strand, it extends outward and there is no PCR. A library is divided into two aliquots, each receiving primers appropriate for one of the parental strands, yielding top-strand or bottom-strand specific PCR amplification. A second PCR stage incorporates a sample barcode and graft sequence.

To test the efficiency of SaferSeqS library preparation, DNA with a known mutation was mixed with normal DNA at VAFs from 8% down to 8×10−6. All were detectable above the 0% sample and at the expected frequency, indicating a VAF sensitivity of 8×10−6; no mutations at the nt were detected in the 0% sample. Across 38 Mb sequenced in the regions, the MFmaxI was 1.6×10−7 per bp.

To benchmark clinical samples, plasma cfDNA from normal and cancer patients were mixed at different ratios down to 0.1% and 11 ng (3,700 haploid genomes if random) sequenced for each of three TP53 mutations. The fraction of on-target reads was 80% and, crucially, the median DCS conversion efficiency of initial genomes to DCS families was 89% rather than the 10–15% typical of other Duplex Sequencing methods. The required family size is not stated. Known TP53 mutations were detected at the expected VAFs down to 10−3, with no spurious mutations. Averaging over the 3.6 Mb of TP53 amplicons, the background MFmaxI was < 3×10−7 per bp, compared to MFminI of 94×10−7 for SafeSeq analysis of the same samples.

Clinical usefulness was tested in plasma obtained from 5 cancer patients who now had minimal tumor burden after treatment and whose tumor mutations had previously been identified by the original Safe-SeqS method. The earlier analysis of plasma identified 8 mutations previously identified in the tumor, at plasma VAFs of 0.1–0.1%, plus 334 different mutations not seen in the original tumor and present at plasma VAFs up to 0.01% and MFminI of 1×10−5. The latter came from 10,347 SSCS families, suggesting an average clone size of 30. SaferSeq found the same eight tumor mutations, at the same plasma VAFs, but no additional mutations. Therefore its spurious background MFminI was < 2×10−7 per bp. SaferSeqS thus provided a 100-fold improved detection sensitivity compared to Safe-SeqS.

4.3.10. CODEC

A drawback of the previous duplex sequencing approaches is that multiple progeny of each parental strand must be present in the aliquot loaded onto the sequencer. This creates two separate problems: a) A DCS family of ~8 entails 8-fold larger sequencing runs. Or for a particular run size, the output represents fewer parental genomes and thus lower coverage depth and sensitivity. b) The DCS conversion efficiency of input DNA to DCS families is often ~10–15%. This is not compensated for by larger sequencing runs, so it originates from the dropout of parental duplexes or too few PCR cycles to ensure co-occurrence of sufficient top- and bottom-progeny in the reaction. Minimizing this problem requires either reducing dilution by using fewer starting genomes, therefore sacrificing sensitivity, or increasing processing efficiencies: ligation of both ends of the parent on both strands, capture or PCR efficiency, loss avoidance during any processing and purification, and sufficient PCR amplification to exceed any dilution or aliquoting but not amplified beyond the family size needed in (a). These problems would be greatly reduced if the top and bottom strands were physically linked.

A recent advance to achieve this is CODEC or Concatenating Original Duplex for Error Correction [109]. This is a tandem-replicate method with the innovation that a tandem unit contains both a top and bottom strand sequence, allowing a DCS family size of 1 to produce the consensus. The insight is to physically link the two mismatched Y-adapters whose mismatch marks the top and bottom strands (Fig. 7). To achieve the linking, the Y-adapter arm not needed for the flow cell graft sequence is made long enough for its 3’ end to bind to a complementary sequence in the other Y-adapter. This creates a U-shaped ”adapter complex” that can ligate to both ends of the insert; it also harbors the components needed for library preparation. After the top and bottom strands of the parental duplex each receive one of the long arms. Each long arm’s 3’ end is then extended (with a parent strand behind it) across the other long arm and, by strand displacement, across the other parent strand and then across its own short arm. The result is a duplex in which each strand contains one parent strand and a copy of the other parent strand, as well as markers for each strand and the library elements needed for PCR and sequencing. Two 3-nt random UMIs distinguish 4000 different parental inserts. Using ~100 insert bases as endogenous barcodes allows ~4×105 parental molecules to be distinguished, more than adequate for studying small amounts of cfDNA, typically 20 ng (~6,000 genomes) in these experiments.

Figure 7:

Figure 7:

Parent-strand consensus methods: CODEC. Adapted from [109].

In practice, side reactions included: the insert ligating to tandem adapter complexes (adapter dimers) due to a T:T mismatch rather than the A:T pairing; ligated adapter complexes without any insert; the strand displacement reaction on one insert-adapter complex using an adjacent ligated product as a template. Even the latter spurious products were found to be usable because useful information is retrieved from one side of the insert.

The method was then tested using 6,000 genome equivalents of cfDNA (~167 bp, with no fragmentation required); this number of genomes would limit the VAF sensitivity to 2×10−4. Samples were end-repaired, A-tailed, enriched for 800 kb of cancer gene targets by hybrid capture, and then analyzed using either DuplexSeq or CODEC. Approximately 3% of parental input duplexes were converted to DCS consensus sequences (termed “recovery of molecular complexity”), giving a VAF sensitivity limit of ~0.6%. However, a single paired-end read sufficed to achieve a DCS. Averaging over many bases, the background MFs, which the authors note might be true or spurious mutations, were similar for DuplexSeq and CODEC at ~4×10−7 and ~3×10−7 per bp. A single paired-end read sufficed to achieve both a consensus and the low error rate. The authors noted that C:G→T:A error rate in a healthy individual was higher than with DuplexSeq, possibly arising in the end-repair step. The same data analyzed using SSCS gave an MF 234-fold higher than the DCS analysis.

Next the authors used CODEC after whole exome capture, where the greater complexity of the DNA fragments would greatly reduce the chance of accumulating families of DuplexSeq PCR progeny. This yielded DCSs, with an MF of 5×10−7. Progressing to further complexity, whole-genome analysis of restriction-digested sperm DNA using some of the error-reduction chemistry of NanoSeq gave DCS coverage depth of 17 and an MF of 3×10−8. By sequencing to a DCS coverage depth of 2300, it was possible to observe VAF down to 0.01% even in a whole-genome sample.

Key properties of the various methods for sequencing rare mutations are summarized in Table 1. They present different advantages depending on an experiment’s requirements for small sample size, low cost, mutation sensitivity, and whether VAF is needed or MF suffices.

Table 1.

Methods for Quantifying Rare Mutations by Sequencing and their Detection Limits.

INTRINSIC PARAMETERS ADJUSTABLE PARAMETERS RESULTS
Method Innovation Fragmentation Targeting UMI # Starting haploid Genomes (< UMIs) Family size Conversion efficiency (genomes to family member) Consensus coverage depth obtained (families per bp) VAF detection limit estimated (1/coverage) MF detection limit estimated (1/consensus bp) VAF bkgd obtained MF bkgd obtained MFminI bkgd obtained MFmaxI bkgd obtained
Single-strand Consensus
Safe-SeqS Consensus of PCR progeny of one parent strand amplicon primers 14 nt random 5,00,000 68 75% 3,74,553 3 × 10−6 9 × 10−8 0–1 × 10−4 9 × 10−6
SiMSen-Seq ”, using hairpin PCR primers to multiplex amplicon primers 12 nt random 1500–33,000 30 10% 1116 1 × 10−3 2 × 10−8 1 × 10−3
Tandem-replicate Consensus
o2n-Seq Consensus of 2 tandem replicates of target (parent strands unmatchable) shear capture none 2 39% 10,000 1 × 10−4 0.1–4% 1 × 10−5
SMM-Seq Consensus of n tandem replicates of target (parent strands matchable but mixed) restriction digest size* 2 × 6 nt random n 45% X X 5 × 10−9 X 2 × 10−7
Parent-strand Consensus
DuplexSeq [ref.107] Consensus of PCR progeny of parent top strand vs parent bottom strand shear whole M13mp2 part of mitochondria 2 × 12 nt random 6 1% 20,000 5 × 10−5 5 × 10−8 (4 × 10−10 when extrapolate from SSCS MF) 1 × 10−4 2.5 × 10−6***
[refs.108, 109] " " shear 2x capture " 6–12 1–10% 1.5 × 10−8 6 × 10−8
Commercial [ref.47] " " shear 2x capture 2 x semi-degenerate 1,65,000 7 × 10−9 1 × 10−7
Commercial (present authors) " " shear or fragmentase 2x capture 2 × 100 predefined 3,30,000 6–12 10% 33,000 3 × 10−5 3 × 10−9 3 × 10−5 7 × 10−7 (DTC )
PacBio HiFi Linked parental top and bottom strands; no PCR; consensus of n tandem replicates of top and bottom strands shear to 1 or 6 kb amplicon or none not needed 10,000 10 10,000 1 × 10−4 1 × 10−7 1 × 10−4 1 × 10−7
SinoDuplex Error-correcting UMIs, Bayesian SSCS calling shear capture 2 × 16 predefined 8000 ≥ 2 25% 2000 5 × 10−4 1 × 10−3
OPUSeq In situ UMI incorporation, simple procedure shear or fragmentase 2x capture 2 × 6 nt random 66,000 30 0.3%– 2% 4441 2 × 10−4 2 × 10−4 1 × 10−6
EcoSeq Reduced representation 100x (30 Mb genome) by restriction digest then 1000x by sampling restriction digest size 2 × 12 nt random 33,000 (predilution) 6 2% X (40) X (2.5%) 2 × 10−8 X 2.4 × 10−8
BotSeqS Reduced representation 500x (6 Mb) by 105x dilution, random genome regions, endogenous barcodes shear none* * endogenous 165,000 (predilution) ≥ 4 (10–15) 1.2% (45% pre-filter) X X 1 × 10−7 (3 ×10−9 pre-fiilter) (3 ×10−12 when extrapolate from SSCS MF) X 5 × 10−7
Hawk-Seq see BotSeq
NanoSeq Random genome regions from 500x dilution of 29% of restriction digest (2 Mb); end-repair improvements, computational filters restriction digest size* 2 × 3 nt random 17,000 ≥ 4 4% (12% pre-filters) X X X < 5 × 109
SaferSeqS Duplex PCR instead of capture amplicon amplicon 2 × 14 nt random gDNA – / cfDNA 10,000 eqvt 89% 170,000 / 6000 6 × 10−6 / 2 × 10−4 3 × 10−8 / 3 × 10−7 8 × 10−6 / 1 × 10−3 − / − − / − 1.6 × 10−7 / < 3 × 10−7
CODEC Physically link top and bottom strands cfDNA / shear / restriction capture / WES / WGS 2 × 3 nt random 6000 2 3% 180 / – / 17 & 2300 6 × 10−3 / – / 6% & 0.04% – / – / – & 0.01% 3 × 10−7 / 5 × 10−7 / 3 × 10−8

Key: – not reported; X cannot measure VAF more sensitively than standard NGS; bold italic, internal standard detection limit from serial dilutions; unbolded, lowest residual mutation frequency in cells.

*

Not expected to see a nt more than once because reduced representation is modest

* *

Not expected to see a nt more than once because representation is random

***

Same value as colony assay, so probable true mutations in standard DNA; MFminI is personal communication, J. Salk.

Commercial DNA Technical Control, 18 yo pt leukocytes; C->T at diPy is 4E-08

Note: # UMIs > # genomes to ensure that progeny of distinct parental fragments are not combined. # distinct random UMIs = ~(4EXP[2 *n]) * (endogenous barcodes)

If the # of potential endpoints of the parental fragment are used as "endogenous UMIs" having non-random sequence, the UMI count is multiplied by ~L, the average length of the parental fragments, ~100.

5. Importance of low-frequency mutation detection

The modifications and improvements introduced by various NGS methods discussed ensure low- and ultralow-frequency mutation detection at high sensitivity and specificity, as compared to conventional methods. But why is it important to detect these rare mutations and what could be the possible clinical implications? These have been extensively reviewed [105, 119] so we only highlight some of the important applications:

  1. The detection of low-frequency mutations using the aforementioned NGS methodologies provides a direct and potentially genome-wide assessment of genetic variations resulting from exposure to environmental, dietary, or therapeutic mutagens [105, 136, 137]. It can avoid systematic biases inherent to expression-dependent reporter genes [138] and has the potential for greater sensitivity. The elevated cost, primarily for library construction rather than sequencing, is offset by eliminating viral packaging and drug selection steps in transgenic mouse assays.

  2. Detecting low-frequency mutations has been instrumental in timely detection of minimal residual disease, which contains treatment-resistant tumor cells still persisting in the body and have the potential to cause cancer recurrence [139]. This information facilitates better prognosis and patient management and it suggests treatment strategies for future clinical trials.

  3. Low-frequency mutations have been implicated in the etiology of neurological disorders like epilepsy [140] and schizophrenia [19].

  4. Screening circulating cell-free fetal DNA for chromosomal abnormalities is currently carried out by the Non-Invasive Prenatal Test (NIPT), which has been shown to be quite accurate for aneuploidy. However, the NIPT test occasionally fails, primarily due to a low fraction of fetal DNA [141, 142]. Moreover, it would be ideal to routinely test for single-base changes associated with inherited disease, and to distinguish heterozygous carriers from affected homozygotes. These may be detectable with more sensitive mutation-sequencing methods, which would support a more robust screen for prospective cases of not only aneuploidies but also metabolic disorders, blood disorders, or rare diseases, without the need for invasive amniocentesis.

6. Conclusions

Several versions of the Duplex Sequencing approach of constructing parent-strand consensus sequences provide detection limits more than adequate to detect true mutations generated by typical mutagens or biological processes. Some measure both VAF and MF, while others can measure only average MF. The choice of methods depends on tradeoffs between patient biopsy size available, minimum or maximum genomic input compatible with the assay, mutation sensitivity required, need for VAF rather than MF, DNA target size, and cost. Additional considerations are the reliability of the assay and the ability to modify experimental parameters, In conveying the strengths of a new method, or the results of using one of these methods, it is important to use unambiguous terms such as VAF, MFminI, and MFmaxI, or clearly state whether: a) the scoring counted only different mutations or counted all mutation occurrences including recurrences and b) the denominator for MF counted all sequenced nt (target footprint * coverage depth) or, in a non-standard usage, counted the length of the target region in the reference genome.

In contrast, mutations detected without constructing consensus families from separately sequenced top and bottom strands of the parental duplex are typically spurious, even those with VAFs above the Illumina sequencer detection limit of ~0.1–0.5% mutations per base pair. This fact makes duplex consensus methods essential for mutation detection even when mutations are only modestly rare, such as 1% per base – 1 genome in 100.

Supplementary Material

1

ACKNOWLEDGMENTS

The research and writing were supported by National Cancer Institute grant 1R01CA240602 to D.E.B.

Abbreviations:

cfDNA

cell-free DNA in plasma

DCS

duplex consensus sequence

ENU

N-ethyl-N-nitrosourea (ENU)

MB

mutation burden

MFmaxI

maximum independent-mutation frequency

MFminI

minimum independent-mutation frequency

MP

mutation prevalence

NGS

next-generation sequencing

nt

nucleotide

SSCS

single-strand consensus sequence

TMB

tumor mutation burden

UMI

unique molecular identifier

VAF

variant allele frequency

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

CRediT AUTHORSHIP CONTRIBUTION STATEMENT

Vijay Menon: Conceptualization, Methodology, Investigation, Writing – original draft, Visualization. Douglas Brash: Conceptualization, Investigation, Writing – original draft, Writing – review and editing, Supervision.

DECLARATION OF INTERESTS

The authors declare that they have no competing financial interests.

Declaration of interests

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

REFERENCES

  • [1].DeMars R, Resistance of cultured human fibroblasts and other cells to purine and pyrimidine analogues in relation to mutagenesis detection, Mutat Res, 24 (1974) 335–364. [DOI] [PubMed] [Google Scholar]
  • [2].McCormick JJ, Maher VM, Measurement of colony- forming ability and mutagenesis in diploid human cells., in: Friedberg EC, Hanawalt PC (Eds.) DNA Repair: A Laboratory Manual of Research Procedures, Marcel Dekker, New York, 1981, pp. 501–521. [Google Scholar]
  • [3].Friedberg EC, Walker GC, Siede W, DNA Repair and Mutagenesis, ASM Press, Washington D.C., 1995. [Google Scholar]
  • [4].Debatisse M, Malfoy B, Gene amplification mechanisms, Adv Exp Med Biol, 570 (2005) 343–361. [DOI] [PubMed] [Google Scholar]
  • [5].Albertson DG, Gene amplification in cancer, Trends Genet, 22 (2006) 447–455. [DOI] [PubMed] [Google Scholar]
  • [6].Greenman C, Stephens P, Smith R, Dalgliesh GL, Hunter C, Bignell G, Davies H, Teague J, Butler A, Stevens C, Edkins S, O’Meara S, Vastrik I, Schmidt EE, Avis T, Barthorpe S, Bhamra G, Buck G, Choudhury B, Clements J, Cole J, Dicks E, Forbes S, Gray K, Halliday K, Harrison R, Hills K, Hinton J, Jenkinson A, Jones D, Menzies A, Mironenko T, Perry J, Raine K, Richardson D, Shepherd R, Small A, Tofts C, Varian J, Webb T, West S, Widaa S, Yates A, Cahill DP, Louis DN, Goldstraw P, Nicholson AG, Brasseur F, Looijenga L, Weber BL, Chiew YE, DeFazio A, Greaves MF, Green AR, Campbell P, Birney E, Easton DF, Chenevix-Trench G, Tan MH, Khoo SK, Teh BT, Yuen ST, Leung SY, Wooster R, Futreal PA, Stratton MR, Patterns of somatic mutation in human cancer genomes, Nature, 446 (2007) 153–158. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Stratton MR, Campbell PJ, Futreal PA, The cancer genome, Nature, 458 (2009) 719–724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Mahmoud M, Gobet N, Cruz-Davalos DI, Mounier N, Dessimoz C, Sedlazeck FJ, Structural variant calling: the long and the short of it, Genome Biol, 20 (2019) 246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Li Y, Roberts ND, Wala JA, Shapira O, Schumacher SE, Kumar K, Khurana E, Waszak S, Korbel JO, Haber JE, Imielinski M, Group PSVW, Weischenfeldt J, Beroukhim R, Campbell PJ, Consortium P, Patterns of somatic structural variation in human cancer genomes, Nature, 578 (2020) 112–121. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Harbers L, Agostini F, Nicos M, Poddighe D, Bienko M, Crosetto N, Somatic copy number alterations in human cancers: An analysis of publicly available data from the cancer genome atlas, Front Oncol, 11 (2021) 700568. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Garcia-Nieto PE, Morrison AJ, Fraser HB, The somatic mutation landscape of the human body, Genome Biol, 20 (2019) 298. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Vali-Pour M, Lehner B, Supek F, The impact of rare germline variants on human somatic mutation processes, Nat Commun, 13 (2022) 3724. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Watson IR, Takahashi K, Futreal PA, Chin L, Emerging patterns of somatic mutations in cancer, Nat Rev Genet, 14 (2013) 703–718. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Mustjoki S, Young NS, Somatic mutations in “Benign” disease, N Engl J Med, 384 (2021) 2039–2052. [DOI] [PubMed] [Google Scholar]
  • [15].Kennedy SR, Loeb LA, Herr AJ, Somatic mutations in aging, cancer and neurodegeneration, Mech Ageing Dev, 133 (2012) 118–126. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Risques RA, Kennedy SR, Aging and the rise of somatic cancer-associated mutations in normal tissues, PLoS Genet, 14 (2018) e1007108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Lobon I, Solis-Moruno M, Juan D, Muhaisen A, Abascal F, Esteller-Cucala P, Garcia-Perez R, Marti MJ, Tolosa E, Avila J, Rahbari R, Marques-Bonet T, Casals F, Soriano E, Somatic mutations detected in Parkinson disease could affect genes with a role in synaptic and neuronal processes, Front Aging, 3 (2022) 851039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Rodin RE, Dou Y, Kwon M, Sherman MA, D’Gama AM, Doan RN, Rento LM, Girskis KM, Bohrson CL, Kim SN, Nadig A, Luquette LJ, Gulhan DC, Network BSM, Park PJ, Walsh CA, Author Correction: The landscape of somatic mutation in cerebral cortex of autistic and neurotypical individuals revealed by ultra-deep whole-genome sequencing, Nat Neurosci, 24 (2021) 611. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Kim MH, Kim IB, Lee J, Cha DH, Park SM, Kim JH, Kim R, Park JS, An Y, Kim K, Kim S, Webster MJ, Kim S, Lee JH, Low-level brain somatic mutations are implicated in schizophrenia, Biol Psychiatry, 90 (2021) 35–46. [DOI] [PubMed] [Google Scholar]
  • [20].Miller MB, Huang AY, Kim J, Zhou Z, Kirkham SL, Maury EA, Ziegenfuss JS, Reed HC, Neil JE, Rento L, Ryu SC, Ma CC, Luquette LJ, Ames HM, Oakley DH, Frosch MP, Hyman BT, Lodato MA, Lee EA, Walsh CA, Somatic genomic changes in single Alzheimer’s disease neurons, Nature, 604 (2022) 714–722. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].Jonason AS, Kunala S, Price GJ, Restifo RJ, Spinelli HM, Persing JA, Leffell DJ, Tarone RE, Brash DE, Frequent clones of p53-mutated keratinocytes in normal human skin, Proc Natl Acad Sci U S A, 93 (1996) 14025–14029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Brash DE, Ponten J, Skin precancer, Cancer Surv, 32 (1998) 69–113. [PubMed] [Google Scholar]
  • [23].Salk JJ, Horwitz MS, Passenger mutations as a marker of clonal cell lineages in emerging neoplasia, Semin Cancer Biol, 20 (2010) 294–303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Graeber TG, Osmanian C, Jacks T, Housman DE, Koch CJ, Lowe SW, Giaccia AJ, Hypoxia-mediated selection of cells with diminished apoptotic potential in solid tumours, Nature, 379 (1996) 88–91. [DOI] [PubMed] [Google Scholar]
  • [25].Zhang W, Remenyik E, Zelterman D, Brash DE, Wikonkal NM, Escaping the stem cell compartment: sustained UVB exposure allows p53-mutant keratinocytes to colonize adjacent epidermal proliferating units without incurring additional mutations, Proc Natl Acad Sci U S A, 98 (2001) 13948–13953. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [26].Klein AM, Brash DE, Jones PH, Simons BD, Stochastic fate of p53-mutant epidermal progenitor cells is tilted toward proliferation by UV B during preneoplasia, Proc Natl Acad Sci U S A, 107 (2010) 270–275. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Alcolea MP, Greulich P, Wabik A, Frede J, Simons BD, Jones PH, Differentiation imbalance in single oesophageal progenitor cells causes clonal immortalization and field change, Nat Cell Biol, 16 (2014) 615–622. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Wong TN, Miller CA, Klco JM, Petti A, Demeter R, Helton NM, Li T, Fulton RS, Heath SE, Mardis ER, Westervelt P, DiPersio JF, Walter MJ, Welch JS, Graubert TA, Wilson RK, Ley TJ, Link DC, Rapid expansion of preexisting nonleukemic hematopoietic clones frequently follows induction therapy for de novo AML, Blood, 127 (2016) 893–897. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Baker NE, Emerging mechanisms of cell competition, Nat Rev Genet, 21 (2020) 683–697. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [30].Castro-Giner F, Ratcliffe P, Tomlinson I, The mini-driver model of polygenic cancer evolution, Nat Rev Cancer, 15 (2015) 680–685. [DOI] [PubMed] [Google Scholar]
  • [31].Clayton E, Doupe DP, Klein AM, Winton DJ, Simons BD, Jones PH, A single type of progenitor cell maintains normal epidermis, Nature, 446 (2007) 185–189. [DOI] [PubMed] [Google Scholar]
  • [32].Klein AM, Simons BD, Universal patterns of stem cell fate in cycling adult tissues, Development, 138 (2011) 3103–3111. [DOI] [PubMed] [Google Scholar]
  • [33].Brash D, Cairns J, The mysterious steps in carcinogenesis, Br J Cancer, 101 (2009) 379–380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].McFarland CD, Korolev KS, Kryukov GV, Sunyaev SR, Mirny LA, Impact of deleterious passenger mutations on cancer progression, Proc Natl Acad Sci U S A, 110 (2013) 2910–2915. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Lawrence MS, Stojanov P, Mermel CH, Robinson JT, Garraway LA, Golub TR, Meyerson M, Gabriel SB, Lander ES, Getz G, Discovery and saturation analysis of cancer genes across 21 tumour types, Nature, 505 (2014) 495–501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Lusito E, Felice B, D’Ario G, Ogier A, Montani F, Di Fiore PP, Bianchi F, Unraveling the role of low-frequency mutated genes in breast cancer, Bioinformatics, 35 (2019) 36–46. [DOI] [PubMed] [Google Scholar]
  • [37].Leal JL, John T, Immunotherapy in advanced NSCLC without driver mutations: Available therapeutic alternatives after progression and future treatment options, Clin Lung Cancer, 23 (2022) 643–658. [DOI] [PubMed] [Google Scholar]
  • [38].Parsons BL, Multiclonal tumor origin: Evidence and implications, Mutat Res Rev Mutat Res, 777 (2018) 1–18. [DOI] [PubMed] [Google Scholar]
  • [39].Maher VM, Dorney DJ, Mendrala AL, Konze-Thomas B, McCormick JJ, DNA excision-repair processes in human cells can eliminate the cytotoxic and mutagenic consequences of ultraviolet irradiation, Mutat Res, 62 (1979) 311–323. [DOI] [PubMed] [Google Scholar]
  • [40].Watanabe M, Maher VM, McCormick JJ, Excision repair of UV- or benzo[a]pyrene diol epoxide-induced lesions in xeroderma pigmentosum variant cells is ‘error free’, Mutat Res, 146 (1985) 285–294. [DOI] [PubMed] [Google Scholar]
  • [41].Yang JL, Chen RH, Maher VM, McCormick JJ, Kinds and location of mutations induced by (+/−)-7 beta,8 alpha-dihydroxy-9 alpha,10 alpha-epoxy-7,8,9,10-tetrahydrobenzo[a]pyrene in the coding region of the hypoxanthine (guanine) phosphoribosyltransferase gene in diploid human fibroblasts, Carcinogenesis, 12 (1991) 71–75. [DOI] [PubMed] [Google Scholar]
  • [42].Buschmann T, Zhang R, Brash DE, Bystrykh LV, Enhancing the detection of barcoded reads in high throughput DNA sequencing data by controlling the false discovery rate, BMC Bioinformatics, 15 (2014) 264. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [43].Bredberg A, Kraemer KH, Seidman MM, Restricted ultraviolet mutational spectrum in a shuttle vector propagated in xeroderma pigmentosum cells, Proc Natl Acad Sci U S A, 83 (1986) 8273–8277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].Brash DE, Seetharam S, Kraemer KH, Seidman MM, Bredberg A, Photoproduct frequency is not the major determinant of UV base substitution hot spots or cold spots in human cells, Proc Natl Acad Sci U S A, 84 (1987) 3782–3786. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [45].Levy DD, Groopman JD, Lim SE, Seidman MM, Kraemer KH, Sequence specificity of aflatoxin B1-induced mutations in a plasmid replicated in xeroderma pigmentosum and DNA repair proficient human cells, Cancer Res, 52 (1992) 5668–5673. [PubMed] [Google Scholar]
  • [46].Sage E, Lamolet B, Brulay E, Moustacchi E, Chteauneuf A, Drobetsky EA, Mutagenic specificity of solar UV light in nucleotide excision repair-deficient rodent cells, Proc Natl Acad Sci U S A, 93 (1996) 176–180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Ikehata H, Masuda T, Sakata H, Ono T, Analysis of mutation spectra in UVB-exposed mouse skin epidermis and dermis: frequent occurrence of C-->T transition at methylated CpG-associated dipyrimidine sites, Environ Mol Mutagen, 41 (2003) 280–292. [DOI] [PubMed] [Google Scholar]
  • [48].Premi S, Han L, Mehta S, Knight J, Zhao D, Palmatier MA, Kornacker K, Brash DE, Genomic sites hypersensitive to ultraviolet radiation, Proc Natl Acad Sci U S A, 116 (2019) 24196–24205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [49].Garcia-Ruiz A, Kornacker K, Brash DE, Cyclobutane Pyrimidine Dimer Hyperhotspots as Sensitive Indicators of Keratinocyte UV Exposure(dagger), Photochem Photobiol, 98 (2022) 987–997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Kunala S, Brash DE, Excision repair at individual bases of the Escherichia coli lacI gene: relation to mutation hot spots and transcription coupling activity, Proc Natl Acad Sci U S A, 89 (1992) 11031–11035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [51].Tornaletti S, Pfeifer GP, Slow repair of pyrimidine dimers at p53 mutation hotspots in skin cancer, Science, 263 (1994) 1436–1438. [DOI] [PubMed] [Google Scholar]
  • [52].Valentine CC 3rd, Young RR, Fielden MR, Kulkarni R, Williams LN, Li T, Minocherhomji S, Salk JJ, Direct quantification of in vivo mutagenesis and carcinogenesis using duplex sequencing, Proc Natl Acad Sci U S A, 117 (2020) 33414–33425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53].Brash D, How do mutant clones expand in normal tissue?, in: Maley C, Greaves M (Eds.) Frontiers in Cancer Research: Evolutionary Foundations, Revolutionary Directions, Springer, New York, 2016, pp. 61–98. [Google Scholar]
  • [54].Martincorena I, Roshan A, Gerstung M, Ellis P, Van Loo P, McLaren S, Wedge DC, Fullam A, Alexandrov LB, Tubio JM, Stebbings L, Menzies A, Widaa S, Stratton MR, Jones PH, Campbell PJ, Tumor evolution. High burden and pervasive positive selection of somatic mutations in normal human skin, Science, 348 (2015) 880–886. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [55].Doupe DP, Alcolea MP, Roshan A, Zhang G, Klein AM, Simons BD, Jones PH, A single progenitor population switches behavior to maintain and repair esophageal epithelium, Science, 337 (2012) 1091–1093. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [56].McGregor WG, Chen RH, Lukash L, Maher VM, McCormick JJ, Cell cycle-dependent strand bias for UV-induced mutations in the transcribed strand of excision repair-proficient human fibroblasts but not in repair-deficient cells, Mol Cell Biol, 11 (1991) 1927–1934. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [57].Guillermin Y, Lopez J, Chabane K, Hayette S, Bardel C, Salles G, Sujobert P, Huet S, What does this mutation mean? The tools and pitfalls of variant interpretation in lymphoid malignancies, Int J Mol Sci, 19 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [58].Pairawan S, Hess KR, Janku F, Sanchez NS, Mills Shaw KR, Eng C, Damodaran S, Javle M, Kaseb AO, Hong DS, Subbiah V, Fu S, Fogelman DR, Raymond VM, Lanman RB, Meric-Bernstam F, Cell-free circulating tumor DNA variant allele frequency associates with survival in metastatic cancer, Clin Cancer Res, 26 (2020) 1924–1931. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [59].Salazar R, Arbeithuber B, Ivankovic M, Heinzl M, Moura S, Hartl I, Mair T, Lahnsteiner A, Ebner T, Shebl O, Proll J, Tiemann-Boege I, Discovery of an unusually high number of de novo mutations in sperm of older men using duplex sequencing, Genome Res, 32 (2022) 499–511. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [60].Martincorena I, Fowler JC, Wabik A, Lawson ARJ, Abascal F, Hall MWJ, Cagan A, Murai K, Mahbubani K, Stratton MR, Fitzgerald RC, Handford PA, Campbell PJ, Saeb-Parsy K, Jones PH, Somatic mutant clones colonize the human esophagus with age, Science, 362 (2018) 911–917. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [61].Dienstmann R, Elez E, Argiles G, Matos I, Sanz-Garcia E, Ortiz C, Macarulla T, Capdevila J, Alsina M, Sauri T, Verdaguer H, Vilaro M, Ruiz-Pace F, Viaplana C, Garcia A, Landolfi S, Palmer HG, Nuciforo P, Rodon J, Vivancos A, Tabernero J, Analysis of mutant allele fractions in driver genes in colorectal cancer - biological and clinical insights, Mol Oncol, 11 (2017) 1263–1272. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Xia L, Li Z, Zhou B, Tian G, Zeng L, Dai H, Li X, Liu C, Lu S, Xu F, Tu X, Deng F, Xie Y, Huang W, He J, Statistical analysis of mutant allele frequency level of circulating cell-free DNA and blood cells in healthy individuals, Sci Rep, 7 (2017) 7526. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [63].Wang K, Lai S, Yang X, Zhu T, Lu X, Wu CI, Ruan J, Ultrasensitive and high-efficiency screen of de novo low-frequency mutations by o2n-seq, Nat Commun, 8 (2017) 15335. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [64].Strom SP, Current practices and guidelines for clinical next-generation sequencing oncology testing, Cancer Biol Med, 13 (2016) 3–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [65].Kinde I, Wu J, Papadopoulos N, Kinzler KW, Vogelstein B, Detection and quantification of rare mutations with massively parallel sequencing, Proc Natl Acad Sci U S A, 108 (2011) 9530–9535. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [66].Hoang ML, Kinde I, Tomasetti C, McMahon KW, Rosenquist TA, Grollman AP, Kinzler KW, Vogelstein B, Papadopoulos N, Genome-wide quantification of rare somatic mutations in normal human tissues using massively parallel sequencing, Proc Natl Acad Sci U S A, 113 (2016) 9846–9851. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [67].Dodge AE, LeBlanc DPM, Zhou G, Williams A, Meier MJ, Van P, Lo FY, Valentine CC, Salk JJ, Yauk CL, Marchetti F, Duplex sequencing provides detailed characterization of mutation frequencies and spectra in the bone marrow of MutaMouse males exposed to procarbazine hydrochloride, bioRxiv, (2023) 2023.2002.2023.529719. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [68].LeBlanc DPM, Meier M, Lo FY, Schmidt E, Valentine C 3rd, Williams A, Salk JJ, Yauk CL, Marchetti F, Duplex sequencing identifies genomic features that determine susceptibility to benzo(a)pyrene-induced in vivo mutations, BMC Genomics, 23 (2022) 542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [69].Short NJ, Kantarjian H, Kanagal-Shamanna R, Sasaki K, Ravandi F, Cortes J, Konopleva M, Issa GC, Kornblau SM, Garcia-Manero G, Garris R, Higgins J, Pratt G, Williams LN, Valentine CC 3rd, Rivera VM, Pritchard J, Salk JJ, Radich J, Jabbour E, Ultra-accurate duplex sequencing for the assessment of pretreatment ABL1 kinase domain mutations in Ph+ ALL, Blood Cancer J, 10 (2020) 61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [70].Sims D, Sudbery I, Ilott NE, Heger A, Ponting CP, Sequencing depth and coverage: key considerations in genomic analyses, Nat Rev Genet, 15 (2014) 121–132. [DOI] [PubMed] [Google Scholar]
  • [71].Merlo LM, Shah NA, Li X, Blount PL, Vaughan TL, Reid BJ, Maley CC, A comprehensive survey of clonal diversity measures in Barrett’s esophagus as biomarkers of progression to esophageal adenocarcinoma, Cancer Prev Res (Phila), 3 (2010) 1388–1397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [72].Makrooni MA, O’Sullivan B, Seoighe C, Bias and inconsistency in the estimation of tumour mutation burden, BMC Cancer, 22 (2022) 840. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [73].Carbone DP, Reck M, Paz-Ares L, Creelan B, Horn L, Steins M, Felip E, van den Heuvel MM, Ciuleanu TE, Badin F, Ready N, Hiltermann TJN, Nair S, Juergens R, Peters S, Minenza E, Wrangle JM, Rodriguez-Abreu D, Borghaei H, Blumenschein GR Jr., Villaruz LC, Havel L, Krejci J, Corral Jaime J, Chang H, Geese WJ, Bhagavatheeswaran P, Chen AC, Socinski MA, Investigators C, First-line Nivolumab in stage IV or recurrent non-small-cell lung cancer, N Engl J Med, 376 (2017) 2415–2426. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [74].Hellmann MD, Callahan MK, Awad MM, Calvo E, Ascierto PA, Atmaca A, Rizvi NA, Hirsch FR, Selvaggi G, Szustakowski JD, Sasson A, Golhar R, Vitazka P, Chang H, Geese WJ, Antonia SJ, Tumor mutational burden and efficacy of Nivolumab monotherapy and in combination with Ipilimumab in small-cell lung cancer, Cancer Cell, 33 (2018) 853–861 e854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [75].Galsky MD, Saci A, Szabo PM, Han GC, Grossfeld G, Collette S, Siefker-Radtke A, Necchi A, Sharma P, Nivolumab in patients with advanced platinum-resistant urothelial carcinoma: Efficacy, safety, and biomarker analyses with extended follow-up from CheckMate 275, Clin Cancer Res, 26 (2020) 5120–5128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [76].Johnson DB, Frampton GM, Rioth MJ, Yusko E, Xu Y, Guo X, Ennis RC, Fabrizio D, Chalmers ZR, Greenbowe J, Ali SM, Balasubramanian S, Sun JX, He Y, Frederick DT, Puzanov I, Balko JM, Cates JM, Ross JS, Sanders C, Robins H, Shyr Y, Miller VA, Stephens PJ, Sullivan RJ, Sosman JA, Lovly CM, Targeted next generation sequencing identifies markers of response to PD-1 blockade, Cancer Immunol Res, 4 (2016) 959–967. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [77].Morrison C, Pabla S, Conroy JM, Nesline MK, Glenn ST, Dressman D, Papanicolau-Sengos A, Burgher B, Andreas J, Giamo V, Qin M, Wang Y, Lenzo FL, Omilian A, Bshara W, Zibelman M, Ghatalia P, Dragnev K, Shirai K, Madden KG, Tafe LJ, Shah N, Kasuganti D, de la Cruz-Merino L, Araujo I, Saenger Y, Bogardus M, Villalona-Calero M, Diaz Z, Day R, Eisenberg M, Anderson SM, Puzanov I, Galluzzi L, Gardner M, Ernstoff MS, Predicting response to checkpoint inhibitors in melanoma beyond PD-L1 and mutational burden, J Immunother Cancer, 6 (2018) 32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [78].Hanna GJ, Lizotte P, Cavanaugh M, Kuo FC, Shivdasani P, Frieden A, Chau NG, Schoenfeld JD, Lorch JH, Uppaluri R, MacConaill LE, Haddad RI, Frameshift events predict anti-PD-1/L1 response in head and neck cancer, JCI Insight, 3 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [79].Goodman AM, Kato S, Bazhenova L, Patel SP, Frampton GM, Miller V, Stephens PJ, Daniels GA, Kurzrock R, Tumor mutational burden as an independent predictor of response to immunotherapy in diverse cancers, Mol Cancer Ther, 16 (2017) 2598–2608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [80].Krauthammer M, Kong Y, Ha BH, Evans P, Bacchiocchi A, McCusker JP, Cheng E, Davis MJ, Goh G, Choi M, Ariyan S, Narayan D, Dutton-Regester K, Capatana A, Holman EC, Bosenberg M, Sznol M, Kluger HM, Brash DE, Stern DF, Materin MA, Lo RS, Mane S, Ma S, Kidd KK, Hayward NK, Lifton RP, Schlessinger J, Boggon TJ, Halaban R, Exome sequencing identifies recurrent somatic RAC1 mutations in melanoma, Nat Genet, 44 (2012) 1006–1014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [81].Fancello L, Gandini S, Pelicci PG, Mazzarella L, Tumor mutational burden quantification from targeted gene panels: major advancements and challenges, J Immunother Cancer, 7 (2019) 183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [82].Jardim DL, Goodman A, de Melo Gagliato D, Kurzrock R, The challenges of tumor mutational burden as an immunotherapy biomarker, Cancer Cell, 39 (2021) 154–173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [83].Mei P, Freitag CE, Wei L, Zhang Y, Parwani AV, Li Z, High tumor mutation burden is associated with DNA damage repair gene mutation in breast carcinomas, Diagn Pathol, 15 (2020) 50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [84].Dai J, Jiang M, He K, Wang H, Chen P, Guo H, Zhao W, Lu H, He Y, Zhou C, DNA damage response and repair gene alterations increase tumor mutational burden and promote poor prognosis of advanced lung cancer, Front Oncol, 11 (2021) 708294. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [85].Li J, Li N, Azhar MS, Liu L, Wang L, Zhang Q, Sheng L, Wang J, Feng S, Qiu Q, Xiao Y, Analysis of mutations in DNA damage repair pathway gene in Chinese patients with hepatocellular carcinoma, Sci Rep, 12 (2022) 12330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [86].Brozek I, Cybulska C, Ratajska M, Piatkowska M, Kluska A, Balabas A, Dabrowska M, Nowakowska D, Niwinska A, Pamula-Pilat J, Tecza K, Pekala W, Rembowska J, Nowicka K, Mosor M, Januszkiewicz-Lewandowska D, Rachtan J, Grzybowska E, Nowak J, Steffen J, Limon J, Prevalence of the most frequent BRCA1 mutations in Polish population, J Appl Genet, 52 (2011) 325–330. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [87].Maxwell KN, Wubbenhorst B, D’Andrea K, Garman B, Long JM, Powers J, Rathbun K, Stopfer JE, Zhu J, Bradbury AR, Simon MS, DeMichele A, Domchek SM, Nathanson KL, Prevalence of mutations in a panel of breast cancer susceptibility genes in BRCA1/2-negative patients with early-onset breast cancer, Genet Med, 17 (2015) 630–638. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [88].Stobbe MD, Thun GA, Dieguez-Docampo A, Oliva M, Whalley JP, Raineri E, Gut IG, Recurrent somatic mutations reveal new insights into consequences of mutagenic processes in cancer, PLoS Comput Biol, 15 (2019) e1007496. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [89].Sanger F, Nicklen S, Coulson AR, DNA sequencing with chain-terminating inhibitors, Proc Natl Acad Sci U S A, 74 (1977) 5463–5467. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [90].Maxam AM, Gilbert W, A new method for sequencing DNA, Proc Natl Acad Sci U S A, 74 (1977) 560–564. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [91].Shendure J, Ji H, Next-generation DNA sequencing, Nat Biotechnol, 26 (2008) 1135–1145. [DOI] [PubMed] [Google Scholar]
  • [92].Ronaghi M, Pyrosequencing sheds light on DNA sequencing, Genome Res, 11 (2001) 3–11. [DOI] [PubMed] [Google Scholar]
  • [93].Huse SM, Huber JA, Morrison HG, Sogin ML, Welch DM, Accuracy and quality of massively parallel DNA pyrosequencing, Genome Biol, 8 (2007) R143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [94].Rothberg JM, Hinz W, Rearick TM, Schultz J, Mileski W, Davey M, Leamon JH, Johnson K, Milgrew MJ, Edwards M, Hoon J, Simons JF, Marran D, Myers JW, Davidson JF, Branting A, Nobile JR, Puc BP, Light D, Clark TA, Huber M, Branciforte JT, Stoner IB, Cawley SE, Lyons M, Fu Y, Homer N, Sedova M, Miao X, Reed B, Sabina J, Feierstein E, Schorn M, Alanjary M, Dimalanta E, Dressman D, Kasinskas R, Sokolsky T, Fidanza JA, Namsaraev E, McKernan KJ, Williams A, Roth GT, Bustillo J, An integrated semiconductor device enabling non-optical genome sequencing, Nature, 475 (2011) 348–352. [DOI] [PubMed] [Google Scholar]
  • [95].Bentley DR, Balasubramanian S, Swerdlow HP, Smith GP, Milton J, Brown CG, Hall KP, Evers DJ, Barnes CL, Bignell HR, Boutell JM, Bryant J, Carter RJ, Keira Cheetham R, Cox AJ, Ellis DJ, Flatbush MR, Gormley NA, Humphray SJ, Irving LJ, Karbelashvili MS, Kirk SM, Li H, Liu X, Maisinger KS, Murray LJ, Obradovic B, Ost T, Parkinson ML, Pratt MR, Rasolonjatovo IM, Reed MT, Rigatti R, Rodighiero C, Ross MT, Sabot A, Sankar SV, Scally A, Schroth GP, Smith ME, Smith VP, Spiridou A, Torrance PE, Tzonev SS, Vermaas EH, Walter K, Wu X, Zhang L, Alam MD, Anastasi C, Aniebo IC, Bailey DM, Bancarz IR, Banerjee S, Barbour SG, Baybayan PA, Benoit VA, Benson KF, Bevis C, Black PJ, Boodhun A, Brennan JS, Bridgham JA, Brown RC, Brown AA, Buermann DH, Bundu AA, Burrows JC, Carter NP, Castillo N, Chiara ECM, Chang S, Neil Cooley R, Crake NR, Dada OO, Diakoumakos KD, Dominguez-Fernandez B, Earnshaw DJ, Egbujor UC, Elmore DW, Etchin SS, Ewan MR, Fedurco M, Fraser LJ, Fuentes Fajardo KV, Scott Furey W, George D, Gietzen KJ, Goddard CP, Golda GS, Granieri PA, Green DE, Gustafson DL, Hansen NF, Harnish K, Haudenschild CD, Heyer NI, Hims MM, Ho JT, Horgan AM, Hoschler K, Hurwitz S, Ivanov DV, Johnson MQ, James T, Huw Jones TA, Kang GD, Kerelska TH, Kersey AD, Khrebtukova I, Kindwall AP, Kingsbury Z, Kokko-Gonzales PI, Kumar A, Laurent MA, Lawley CT, Lee SE, Lee X, Liao AK, Loch JA, Lok M, Luo S, Mammen RM, Martin JW, McCauley PG, McNitt P, Mehta P, Moon KW, Mullens JW, Newington T, Ning Z, Ling Ng B, Novo SM, O’Neill MJ, Osborne MA, Osnowski A, Ostadan O, Paraschos LL, Pickering L, Pike AC, Pike AC, Chris Pinkard D, Pliskin DP, Podhasky J, Quijano VJ, Raczy C, Rae VH, Rawlings SR, Chiva Rodriguez A, Roe PM, Rogers J, Rogert Bacigalupo MC, Romanov N, Romieu A, Roth RK, Rourke NJ, Ruediger ST, Rusman E, Sanches-Kuiper RM, Schenker MR, Seoane JM, Shaw RJ, Shiver MK, Short SW, Sizto NL, Sluis JP, Smith MA, Ernest Sohna Sohna J, Spence EJ, Stevens K, Sutton N, Szajkowski L, Tregidgo CL, Turcatti G, Vandevondele S, Verhovsky Y, Virk SM, Wakelin S, Walcott GC, Wang J, Worsley GJ, Yan J, Yau L, Zuerlein M, Rogers J, Mullikin JC, Hurles ME, McCooke NJ, West JS, Oaks FL, Lundberg PL, Klenerman D, Durbin R, Smith AJ, Accurate whole human genome sequencing using reversible terminator chemistry, Nature, 456 (2008) 53–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [96].Eid J, Fehr A, Gray J, Luong K, Lyle J, Otto G, Peluso P, Rank D, Baybayan P, Bettman B, Bibillo A, Bjornson K, Chaudhuri B, Christians F, Cicero R, Clark S, Dalal R, Dewinter A, Dixon J, Foquet M, Gaertner A, Hardenbol P, Heiner C, Hester K, Holden D, Kearns G, Kong X, Kuse R, Lacroix Y, Lin S, Lundquist P, Ma C, Marks P, Maxham M, Murphy D, Park I, Pham T, Phillips M, Roy J, Sebra R, Shen G, Sorenson J, Tomaney A, Travers K, Trulson M, Vieceli J, Wegener J, Wu D, Yang A, Zaccarin D, Zhao P, Zhong F, Korlach J, Turner S, Real-time DNA sequencing from single polymerase molecules, Science, 323 (2009) 133–138. [DOI] [PubMed] [Google Scholar]
  • [97].Travers KJ, Chin CS, Rank DR, Eid JS, Turner SW, A flexible and efficient template format for circular consensus sequencing and SNP detection, Nucleic Acids Res, 38 (2010) e159. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [98].Kasianowicz JJ, Brandin E, Branton D, Deamer DW, Characterization of individual polynucleotide molecules using a membrane channel, Proc Natl Acad Sci U S A, 93 (1996) 13770–13773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [99].Deamer D, Akeson M, Branton D, Three decades of nanopore sequencing, Nat Biotechnol, 34 (2016) 518–524. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [100].Wang Y, Zhao Y, Bollas A, Wang Y, Au KF, Nanopore sequencing technology, bioinformatics and applications, Nat Biotechnol, 39 (2021) 1348–1365. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [101].Rang FJ, Kloosterman WP, de Ridder J, From squiggle to basepair: computational approaches for improving nanopore sequencing read accuracy, Genome Biol, 19 (2018) 90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [102].Quail MA, Smith M, Coupland P, Otto TD, Harris SR, Connor TR, Bertoni A, Swerdlow HP, Gu Y, A tale of three next generation sequencing platforms: comparison of Ion Torrent, Pacific Biosciences and Illumina MiSeq sequencers, BMC Genomics, 13 (2012) 341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [103].Nguyen HQ, Chattoraj S, Castillo D, Nguyen SC, Nir G, Lioutas A, Hershberg EA, Martins NMC, Reginato PL, Hannan M, Beliveau BJ, Church GM, Daugharthy ER, Marti-Renom MA, Wu CT, 3D mapping and accelerated super-resolution imaging of the human genome using in situ sequencing, Nat Methods, 17 (2020) 822–832. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [104].Payne AC, Chiang ZD, Reginato PL, Mangiameli SM, Murray EM, Yao CC, Markoulaki S, Earl AS, Labade AS, Jaenisch R, Church GM, Boyden ES, Buenrostro JD, Chen F, In situ genome sequencing resolves DNA sequence and structure in intact biological samples, Science, 371 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [105].Salk JJ, Schmitt MW, Loeb LA, Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations, Nat Rev Genet, 19 (2018) 269–285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [106].Schmitt MW, Kennedy SR, Salk JJ, Fox EJ, Hiatt JB, Loeb LA, Detection of ultra-rare mutations by next-generation sequencing, Proc Natl Acad Sci U S A, 109 (2012) 14508–14513. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [107].Ren Y, Zhang Y, Wang D, Liu F, Fu Y, Xiang S, Su L, Li J, Dai H, Huang B, SinoDuplex: An Improved Duplex Sequencing Approach to Detect Low-frequency Variants in Plasma cfDNA Samples, Genomics Proteomics Bioinformatics, 18 (2020) 81–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [108].Ueda S, Yamashita S, Nakajima M, Kumamoto T, Ogawa C, Liu YY, Yamada H, Kubo E, Hattori N, Takeshima H, Wakabayashi M, Iida N, Shiraishi Y, Noguchi M, Sato Y, Ushijima T, A quantification method of somatic mutations in normal tissues and their accumulation in pediatric patients with chemotherapy, Proc Natl Acad Sci U S A, 119 (2022) e2123241119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [109].Bae JH, Liu R, Roberts E, Nguyen E, Tabrizi S, Rhoades J, Blewett T, Xiong K, Gydush G, Shea D, An Z, Patel S, Cheng J, Sridhar S, Liu MH, Lassen E, Skytte A-B, Grońska-Pęski M, Shoag JE, Evrony GD, Parsons HA, Mayer EL, Makrigiorgos GM, Golub TR, Adalsteinsson VA, Single duplex DNA sequencing with CODEC detects mutations with high sensitivity, Nature Genetics, (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [110].Stahlberg A, Krzyzanowski PM, Jackson JB, Egyud M, Stein L, Godfrey TE, Simple, multiplexed, PCR-based barcoding of DNA enables sensitive mutation detection in liquid biopsies using sequencing, Nucleic Acids Res, 44 (2016) e105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [111].Stahlberg A, Krzyzanowski PM, Egyud M, Filges S, Stein L, Godfrey TE, Simple multiplexed PCR-based barcoding of DNA for ultrasensitive mutation detection by next-generation sequencing, Nat Protoc, 12 (2017) 664–682. [DOI] [PubMed] [Google Scholar]
  • [112].Elliott K, Bostrom M, Filges S, Lindberg M, Van den Eynden J, Stahlberg A, Clausen AR, Larsson E, Elevated pyrimidine dimer formation at distinct genomic bases underlies promoter mutation hotspots in UV-exposed cancers, PLoS Genet, 14 (2018) e1007849. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [113].Fredriksson NJ, Elliott K, Filges S, Van den Eynden J, Stahlberg A, Larsson E, Recurrent promoter mutations in melanoma are defined by an extended context-specific mutational signature, PLoS Genet, 13 (2017) e1006773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [114].Luijts T, Elliott K, Siaw JT, Van de Velde J, Beyls E, Claeys A, Lammens T, Larsson E, Willaert W, Vral A, Van den Eynden J, A clinically annotated post-mortem approach to study multi-organ somatic mutational clonality in normal tissues, Sci Rep, 12 (2022) 10322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [115].Lou DI, Hussmann JA, McBee RM, Acevedo A, Andino R, Press WH, Sawyer SL, High-throughput DNA sequencing errors are reduced by orders of magnitude using circle sequencing, Proc Natl Acad Sci U S A, 110 (2013) 19872–19877. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [116].Maslov AY, Makhortov S, Sun S, Heid J, Dong X, Lee M, Vijg J, Single-molecule, quantitative detection of low-abundance somatic mutations by high-throughput sequencing, Sci Adv, 8 (2022) eabm3259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [117].Kennedy SR, Schmitt MW, Fox EJ, Kohrn BF, Salk JJ, Ahn EH, Prindle MJ, Kuong KJ, Shen JC, Risques RA, Loeb LA, Detecting ultralow-frequency mutations by Duplex Sequencing, Nat Protoc, 9 (2014) 2586–2606. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [118].Schmitt MW, Fox EJ, Prindle MJ, Reid-Bayliss KS, True LD, Radich JP, Loeb LA, Sequencing small genomic targets with high efficiency and extreme accuracy, Nat Methods, 12 (2015) 423–425. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [119].Momozawa Y, Mizukami K, Unique roles of rare variants in the genetics of complex diseases in humans, J Hum Genet, 66 (2021) 11–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [120].Kostecka A, Nowikiewicz T, Olszewski P, Koczkowska M, Horbacz M, Heinzl M, Andreou M, Salazar R, Mair T, Madanecki P, Gucwa M, Davies H, Skokowski J, Buckley PG, Peksa R, Srutek E, Szylberg L, Hartman J, Jankowski M, Zegarski W, Tiemann-Boege I, Dumanski JP, Piotrowski A, High prevalence of somatic PIK3CA and TP53 pathogenic variants in the normal mammary gland tissue of sporadic breast cancer patients revealed by duplex sequencing, NPJ Breast Cancer, 8 (2022) 76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [121].Gydush G, Nguyen E, Bae JH, Blewett T, Rhoades J, Reed SC, Shea D, Xiong K, Liu R, Yu F, Leong KW, Choudhury AD, Stover DG, Tolaney SM, Krop IE, Christopher Love J, Parsons HA, Mike Makrigiorgos G, Golub TR, Adalsteinsson VA, Massively parallel enrichment of low-frequency alleles enables duplex sequencing at low depth, Nat Biomed Eng, 6 (2022) 257–266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [122].Wenger AM, Peluso P, Rowell WJ, Chang PC, Hall RJ, Concepcion GT, Ebler J, Fungtammasan A, Kolesnikov A, Olson ND, Topfer A, Alonge M, Mahmoud M, Qian Y, Chin CS, Phillippy AM, Schatz MC, Myers G, DePristo MA, Ruan J, Marschall T, Sedlazeck FJ, Zook JM, Li H, Koren S, Carroll A, Rank DR, Hunkapiller MW, Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome, Nat Biotechnol, 37 (2019) 1155–1162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [123].Russo G, Patrignani A, Poveda L, Hoehn F, Scholtka B, Schlapbach R, Garvin AM, Highly sensitive, non-invasive detection of colorectal cancer mutations using single molecule, third generation sequencing, Appl Transl Genom, 7 (2015) 32–39. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [124].Callahan BJ, Wong J, Heiner C, Oh S, Theriot CM, Gulati AS, McGill SK, Dougherty MK, High-throughput amplicon sequencing of the full-length 16S rRNA gene with single-nucleotide resolution, Nucleic Acids Res, 47 (2019) e103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [125].Cartwright JF, Anderson K, Longworth J, Lobb P, James DC, Highly sensitive detection of mutations in CHO cell recombinant DNA using multi-parallel single molecule real-time DNA sequencing, Biotechnol Bioeng, 115 (2018) 1485–1498. [DOI] [PubMed] [Google Scholar]
  • [126].Miranda JA, Alund AW, Yan J, McKinzie PB, Dobrovolsky VN, Revollo JR, Genome-wide detection of ultralow-frequency substitution mutations in cultures of mouse lymphoma L5178Y cells and Caenorhabditis elegans worms by PacBio sequencing, Environ Mol Mutagen, 63 (2022) 68–75. [DOI] [PubMed] [Google Scholar]
  • [127].Alekseenko A, Wang J, Barrett D, Pelechano V, OPUSeq simplifies detection of low-frequency DNA variants and uncovers fragmentase-associated artifacts, NAR Genom Bioinform, 4 (2022) lqac048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [128].Matsumura S, Sato H, Otsubo Y, Tasaki J, Ikeda N, Morita O, Genome-wide somatic mutation analysis via Hawk-Seq reveals mutation profiles associated with chemical mutagens, Arch Toxicol, 93 (2019) 2689–2701. [DOI] [PubMed] [Google Scholar]
  • [129].Otsubo Y, Matsumura S, Ikeda N, Morita O, Hawk-Seq differentiates between various mutations in Salmonella typhimurium TA100 strain caused by exposure to Ames test-positive mutagens, Mutagenesis, 36 (2021) 245–254. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [130].Abascal F, Harvey LMR, Mitchell E, Lawson ARJ, Lensing SV, Ellis P, Russell AJC, Alcantara RE, Baez-Ortega A, Wang Y, Kwa EJ, Lee-Six H, Cagan A, Coorens THH, Chapman MS, Olafsson S, Leonard S, Jones D, Machado HE, Davies M, Obro NF, Mahubani KT, Allinson K, Gerstung M, Saeb-Parsy K, Kent DG, Laurenti E, Stratton MR, Rahbari R, Campbell PJ, Osborne RJ, Martincorena I, Somatic mutation landscapes at single-molecule resolution, Nature, 593 (2021) 405–410. [DOI] [PubMed] [Google Scholar]
  • [131].Cairns J, Mutation selection and the natural history of cancer, Nature, 255 (1975) 197–200. [DOI] [PubMed] [Google Scholar]
  • [132].Cohen JD, Douville C, Dudley JC, Mog BJ, Popoli M, Ptak J, Dobbyn L, Silliman N, Schaefer J, Tie J, Gibbs P, Tomasetti C, Papadopoulos N, Kinzler KW, Vogelstein B, Detection of low-frequency DNA variants by targeted sequencing of the Watson and Crick strands, Nat Biotechnol, 39 (2021) 1220–1227. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [133].Makarov V, Laliberte J, Enhanced adapter ligation., US patent 10,208,338B2, (2019).
  • [134].Markoulatos P, Siafakas N, Moncany M, Multiplex polymerase chain reaction: a practical approach, J Clin Lab Anal, 16 (2002) 47–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [135].Newman AM, Lovejoy AF, Klass DM, Kurtz DM, Chabon JJ, Scherer F, Stehr H, Liu CL, Bratman SV, Say C, Zhou L, Carter JN, West RB, Sledge GW, Shrager JB, Loo BW Jr., Neal JW, Wakelee HA, Diehn M, Alizadeh AA, Integrated digital error suppression for improved detection of circulating tumor DNA, Nat Biotechnol, 34 (2016) 547–555. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [136].Salk JJ, Kennedy SR, Next-Generation Genotoxicology: Using Modern Sequencing Technologies to Assess Somatic Mutagenesis and Cancer Risk, Environ Mol Mutagen, 61 (2020) 135–151. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [137].Marchetti F, Cardoso R, Chen CL, Douglas GR, Elloway J, Escobar PA, Harper T Jr., Heflich RH, Kidd D, Lynch AM, Myers MB, Parsons BL, Salk JJ, Settivari RS, Smith-Roe SL, Witt KL, Yauk C, Young RR, Zhang S, Minocherhomji S, Error-corrected next-generation sequencing to advance nonclinical genotoxicity and carcinogenicity testing, Nat Rev Drug Discov, 22 (2023) 165–166. [DOI] [PubMed] [Google Scholar]
  • [138].Maslov AY, Quispe-Tintaya W, Gorbacheva T, White RR, Vijg J, High-throughput sequencing in mutation detection: A new generation of genotoxicity tests?, Mutat Res, 776 (2015) 136–143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [139].Dai P, Wu LR, Chen SX, Wang MX, Cheng LY, Zhang JX, Hao P, Yao W, Zarka J, Issa GC, Kwong L, Zhang DY, Calibration-free NGS quantitation of mutations below 0.01% VAF, Nat Commun, 12 (2021) 6123. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [140].Kim JK, Cho J, Kim SH, Kang HC, Kim DS, Kim VN, Lee JH, Brain somatic mutations in MTOR reveal translational dysregulations underlying intractable focal epilepsy, J Clin Invest, 129 (2019) 4207–4223. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [141].Lopes JL, Lopes GS, Enninga EAL, Kearney HM, Hoppman NL, Rowsey RA, Most noninvasive prenatal screens failing due to inadequate fetal cell free DNA are negative for trisomy when repeated, Prenat Diagn, 40 (2020) 831–837. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [142].Zhang B, Zhou L, Feng C, Liu J, Yu B, More attention should be paid to pregnant women who fail non-invasive prenatal screening, Clin Biochem, 96 (2021) 33–37. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES