Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 Sep 14:2026.09.11.751022. [Version 1] doi: 10.64898/2026.09.11.751022

Estimating de novo mutation rates using parent-offspring pairs

Thuy-Trang Nguyen 1, Matthew W Hahn 1
PMCID: PMC13596197  PMID: 42779714

Abstract

Existing pedigree approaches to identifying de novo mutations (DNMs) require at least two parents and a single offspring, limiting applicability. Here, we introduce OOPS (Only One Parent Sequencing), a framework for detecting DNMs using only a single parent-offspring pair. OOPS uses short-read data from the parent and both short and long-read data from the offspring to reconstruct haplotypes in the child, one of which can then be assigned to the sequenced parent. We show that candidate de novo mutations from the assigned haplotype can be identified, allowing for estimation of the mutation rate. To demonstrate the accuracy of OOPS, we apply it to a human pedigree in which mutations have also been identified using standard trio-based approaches. OOPS achieves comparable accuracy to trio-based pipelines and recovers consistent mutation rate estimates. By removing the requirement for complete trio sequencing, OOPS expands mutation rate estimation to a wider range of settings.

Introduction

De novo mutations (DNMs) are the source of genetic variation in both evolution and disease (Veltman and Brunner 2012). The per-generation mutation rate is a fundamental parameter in molecular evolution, shaping estimates of divergence times, effective population sizes, and the genetic load carried by populations (Yoder and Tiley 2021). Accurate estimation of the mutation rate is therefore critical across disciplines.

The standard approach to DNM detection in longer-lived organisms requires sequencing a complete trio — both parents and a child — so that variants present in the child but absent from the parents can be identified as de novo (Bergeron et al. 2022). Trio-based studies have revealed key features of the mutational process, including a strong paternal bias (Kong et al. 2012; Jonsson et al. 2017; Wang et al. 2020; Wu et al. 2020; de Manuel et al. 2022; Wang et al. 2022a, 2022b; Bergeron et al. 2023; Peña-Garcia et al. 2025; Wooldridge et al. 2025), more minor maternal age-effects (Goldmann et al. 2016; Jonsson et al. 2017; Wang et al. 2025), and inter-individual variation in mutation rates (Sasani et al. 2019). Multi-generation pedigrees have further refined rate estimates, particularly since they can be used to verify Mendelian inheritance of DNMs across generations (Jonsson et al. 2017; Thomas et al. 2018; Porubsky et al. 2025).

While trio sequencing is accurate and can detect mutations transmitted by both parents, it necessarily excludes many species and study systems where two parents cannot be sampled or identified. For instance, in animals with uniparental care, one parent is typically unavailable or difficult to identify. In plants—especially wind-pollinated plants—pollen donors may be harder to identify, even when maternal and offspring tissue can be readily collected together (e.g. acorns attached to oak trees). Even in human clinical studies, one parent is sometimes unavailable or deceased. These constraints have left germline mutation rates unmeasured in many systems.

Recent advances in sequencing technology produce much longer reads, enabling haplotype-resolved genome sequencing of single individuals (Mahmoud et al. 2025). These advances allow for new approaches in mutational studies, overcoming the requirement that complete trios be sequenced. Here, we introduce OOPS (Only One Parent Sequencing), a computational method that estimates the DNM rate from a single parent–child pair. OOPS exploits the ability of long reads to phase the child's diploid genome into two haplotypes, comparing each haplotype against the single sequenced parent to identify DNM candidates. We validate the approach on a well-characterized human pedigree with high-quality DNM calls (Porubsky et al. 2025).

Novel Approaches

Mutation identification from parent-offspring pairs

OOPS requires both short-read and long-read data from the child, but only short-read data from the single parent (Fig. 1A). As direct input to the program, OOPS requires genotypes from the child and parent (as VCF files), and mapped short and long reads from the appropriate individuals (as BAM files). The pipeline proceeds in three stages (Supplementary Fig. S1).

Figure 1. Overview of haplotype assignment and de novo mutation detection in the OOPS framework.

Figure 1.

(a) In a pedigree, only one parent and a child are sequenced. Long-read sequencing of the child enables phasing , resolving the diploid genome into two distinct haplotypes (colored blocks). The available parent is sequenced with short reads, yielding unphased genotypes (hatched pattern). (b) Within a haplotype phase set (outlined by a red square in panel a), the child's two haplotypes are compared against the parenťs genotype. Here, haplotype 1 (blue) matches the parent at all 8 sites (100%), whereas haplotype 2 (orange) matches at only 4 of 8 sites (50%; mismatches marked with red crosses). This asymmetry implies that haplotype 1 is inherited from the sequenced parent, and the complementary haplotype is inherited from the absent parent, enabling parent-of-origin assignment of haplotypes and candidate de novo mutations.

In the first stage, the child’s genotype is phased into haplotype blocks using long-read sequences. The OOPS package includes WhatsHap (Martin et al. 2016) for phasing, but, if preferred, users also have the option of inputting a phased VCF for the child using other software. Most importantly, the child's diploid genome is now partitioned into phase sets — individual blocks of paired haplotypes (e.g. Fig. 1).

In the second stage, within each phase set, OOPS attempts to assign one of the two child haplotypes to the sequenced parent. As each child haplotype only has one allele, but the parental diploid genotype has two, OOPS scores as a “match” any position in the haplotype that could have come from the parent, and as a “mismatch” any position that could not have come from the parent (Fig. 1B). The expectation is that one child haplotype will match the sequenced parent at nearly every site because it is inherited from that parent; the other haplotype will show many mismatches, though the exact number and proportion depend on many factors. The OOPS default settings for numbers of matches and mismatches needed to assign haplotypes confidently are described in the Supplementary Materials.

OOPS identifies blocks showing the expected asymmetric mismatch pattern: many mismatches on one haplotype and 0 or 1 mismatches on the other. A single mismatch on the low-mismatch haplotype — a site where the child carries an allele absent from the sequenced parent, on the haplotype otherwise concordant with that parent — is flagged as a DNM candidate.

In the third stage, OOPS takes candidate DNMs through three validation steps. First, standard short-read filters are applied (Bergeron et al. 2022), including genotype quality and read-depth in both individuals, homozygosity in the parent, and allelic balance in the child (Supplementary Methods). Second, DNMs must have long-read support: the alternate (mutant) allele must be present on exactly one haplotype in the child, with a minimum number of supporting reads, and a minimum total depth of long-reads. We have found that this step eliminates many post-zygotic mutations often identified by trio-based approaches (e.g. Porubsky et al. 2025). Third, OOPS locally rephases the candidate DNM. WhatsHap is re-run in a 40 kilobase window (adjustable) around each candidate DNM to produce a new phase set, removing false positives caused by haplotype switch errors that accumulate over longer genomic distances. Candidates that pass all three filters are promoted to the final DNM call set.

Estimation of the mutation rate

The calculation of a mutation rate requires not just accurate DNM calls, but also the number of sites examined at which a mutation could have been identified—the “callable genome size”— and the false negative rate (FNR), the rate at which true mutations are missed (Bergeron et al. 2022). Using these numbers, the mutation rate within OOPS is calculated as:

μ=numberofDNMscallablegenome*1−FNR.

We define the callable genome in OOPS as the total number of bases within phased sets where one child's haplotype can be unambiguously assigned to the sequenced parent, i.e. phase sets showing the expected asymmetric mismatch pattern (with either 0 or 1 mismatch on the assigned haplotype). To estimate the false negative rate, OOPS samples one high-quality heterozygous SNPs from each assigned phase set when an eligible site is available (Supplementary Materials). The alternate allele at each sampled site is treated as a positive control mutant allele, and is subjected to the same short, long-read validation filters as the DNM candidates, except for the requirement that the sequenced parent be homozygous reference and contain no read support for the alternative allele. The FNR is the fraction of sampled control sites that failed these validation filters.

Results

OOPS accurately estimates the mutation rate

We evaluated OOPS using parent-child pairs from the human pedigree sequenced by Porubsky et al. (2025) using multiple sequencing technologies. We focus on a single child (individual NA12879) and analyzed either her mother (NA12878) or father (NA12877) as the single sequenced parent. For this child, Porubsky et al. (2025) reported 6 maternal and 37 paternal autosomal germline single-nucleotide DNMs using a trio-based method.

At 10X PacBio HiFi long-read coverage (down-sampled from the full 37X coverage dataset), WhatsHap produced 17,191 phase blocks with a median length of 42.7 kb, phasing 98.9% of heterozygous sites (Supplementary Table 1). Using the full Illumina short-read dataset for the child and mother (both at 31X coverage), OOPS recovered 2 of the 6 germline DNMs (33%), with 0 false positives (Table 1). In the child-father pair, OOPS made 14 calls: 12 matched the 37 paternal germline DNMs reported by Porubsky et al. (2025), one matched a DNM called in their dataset, but that was unassigned, and one was not in their dataset (i.e. a false positive; Table 1). The majority of the DNMs from the original study missed by OOPS occurred on phase sets for which both haplotypes had many mismatches to the sequenced parent, which appears mainly due to phase-switch errors in haplotyping (see below).

Table 1.

OOPS performance at 10X long-read coverage on mutations in offspring NA12879. Trio DNMs and rates come from Porubsky et al. (2025).

Platform Parent Trio DNMs OOPS DNMs FP Trio rate OOPS rate
PacBio HiFi Maternal 6 2 0 0.23 × 10−8 0.22 × 10−8
PacBio HiFi Paternal 37 14 1 1.39 × 10−8 1.53 × 10−8
ONT Maternal 6 1 0 0.23 × 10−8 0.25 × 10−8
ONT Paternal 37 4 0 1.39 × 10−8 1.02 × 10−8

FP = false positives. Mutation rates are expressed per bp per generation.

Despite recovering only a subset of DNMs in each parent-offspring pair, the mutation rate estimated by OOPS closely matched the trio-based estimates (Table 1). The maternal rate estimated was 0.22 × 10−8 per bp per generation (trio-based: 0.23 × 10−8) and the paternal rate was 1.53 × 10−8 (trio-based: 1.39 × 10−8), reproducing the male-biased mutation observed in humans (Kong et al. 2012).

Effect of long-read technology and read-depth

To determine the effects of long-read technology, we repeated the OOPS analyses using Oxford Nanopore (ONT) reads down-sampled to 10X coverage in the child. Again, 98.9% of the heterozygous sites were phased, but the phase-set structure is different from that obtained with PacBio: WhatsHap produced 440 phase sets with a median length of 690 kb (Supplementary Table 1). In this dataset, OOPS recovered fewer DNMs from both pairs (1 maternal, 4 paternal) and produced a slightly underestimated paternal mutation rate (Table 1). The smaller number of DNMs detected using the ONT data seems to be due to frequent phase-switch errors, with many long haplotypes showing regions with low numbers of mismatches and regions of high mismatches (Supplementary Figure S2). In these cases, OOPS cannot confidently assign haplotypes.

To determine the effects of long-read depth in the child, we analyzed PacBio HiFi datasets with 5X, 7X, 10X (the analysis above), 15X, and the full 37X coverage. Paternal DNM calls increased steadily with depth, while maternal calls plateaued (Fig. 2a; Supplementary Table 2). False positives were rare at every depth tested: one was called at 5X in the child-mother pair and one at 10X in the child-father pair, and none at any other depth. The estimated mutation rate remained stable across read-depths (Fig. 2b; Supplementary Table 2), as both the callable genome and FNR were also recalculated at each read-depth.

Fig 2.

Fig 2.

De novo mutation detection as a function of PacBio HiFi read depth. (a) Number of DNM calls of paternal (blue) and maternal (red) origin recovered by OOPS at coverage levels from 5X to 37X. (b) Estimated per-generation mutation rate at each depth, with the trio-based rates of Porubsky et al. (2025) shown as dashed lines. Bars at 10X give the 95% confidence interval across 100 independent samplings of the same data — of the number of calls in (a) and of the estimated rate in (b) (see also Supplementary Table 3).

Because there is stochasticity associated with read-sampling at every depth, some variation in coverage-specific results could arise from read sampling itself. We quantified this variability by repeating the 10X HiFi analysis 100 times, each time drawing a new set of reads and rerunning phasing and all OOPS steps. Across replicates, the mean number of calls was 13.8 for the paternal pair (range 9–19) and 2.2 for the maternal pair (range 0–4) (Fig. 2a; Supplementary Table 3). The mean mutation-rate estimates were 1.54 × 10⁻ for the paternal pair and 0.25 × 10⁻ for the maternal pair; the corresponding 95% confidence intervals are shown in Figure 2b (see also Supplementary Table 3). These replicates capture variability arising from read-sampling, phasing, and candidate filtering at 10X coverage, encompassing the range of mutation rate results captured across different read-depths (Fig. 2b).

Discussion

We demonstrate here that accurate estimation of the per-generation mutation rate is possible with sequencing data from only a parent-child pair. The key insight is that long-reads can phase a child's genome into haplotypes, one of which can often be accurately assigned to the sequenced parent. Although OOPS recovers only a minority of true DNMs — because only sites within well-phased blocks with sufficient read support are assessable — the mutation rate estimate remains accurate because the callable-genome calculation corrects for the reduced available genome (see also Figure S5 in Wang et al. 2020).

The method has several limitations. First, accuracy depends critically on phasing quality: haplotype switch errors can cause false positives by placing a true inherited variant on the wrong haplotype and can cause false negatives by making it difficult to assign haplotypes to the sequenced parent. We have found that the local rephasing step partially mitigates this issue, at least by eliminating many false positives. The same general issue is also why we have chosen not to use phased data from the sequenced parent: in addition to false positives caused by allelic gene conversion (cf. Narasimhan et al. 2017), switch errors or true recombination events in the parent will lead to more noise but not necessarily more detectable mutations. See Boukas et al. (2026) for a similar method that does use a phased parent, and Young et al. (2024) for a method using a pair of phased siblings.

As a second limitation, OOPS is currently restricted to single nucleotide variants on autosomes: indels, structural variants, multinucleotide mutations, and mutations on sex chromosomes are not detected. However, all of these restrictions can be relaxed in the future. Any type of mutation can be included, including structural variants identified using the long-reads themselves. With slight modifications, OOPS can be run on sex chromosomes, a natural application for chromosomes that are often hemizygous and that can therefore be phased even without long-reads.

Despite its limitations, OOPS fills an important gap in evolutionary genomics. Germline mutation rates have been measured using trio sequencing in a limited number of species, predominantly those that can be raised in the lab or a zoo (Bergeron et al. 2023). OOPS enables DNM-rate estimation in many non-ideal circumstances, with relatively low requirements for read-depth. As long-read sequencing costs continue to decline, this approach could dramatically expand the taxonomic breadth of mutation-rate estimates.

Supplementary Material

Supplement 1
media-1.docx (423.8KB, docx)

Acknowledgements

We thank Richard Wang, Yadira Peña-García and Jeff Rogers for valuable feedback. This work was supported by NIH grant R01-HD107120.

Data Availability

The pipeline is available at https://github.com/TrangNg-Th/OOPS-Only-One-Parent-Sequencing and is distributed as a Conda package (oops-dnm) for straightforward installation.

References

  1. Bergeron LA, Besenbacher S, Turner T, Versoza CJ, Wang RJ, Price AL, Armstrong E, Riera M, Carlson J, Chen HY, et al. The Mutationathon highlights the importance of reaching standardization in estimates of pedigree-based germline mutation rates. eLife. 2022:11:e73577. 10.7554/eLife.73577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Bergeron LA, Besenbacher S, Zheng J, Li P, Bertelsen MF, Quintard B, Hoffman JI, Li Z, St Leger J, Shao C, et al. Evolution of the germline mutation rate across vertebrates. Nature. 2023:615:285–291. 10.1038/s41586-023-05752-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Boukas L, Délot EC, Pitsava G, Lambert C, Fanslow C, Baybayan P, Belhadj S, Losic B, Harting J, Bluske K, et al. Identification of de novo variants from parent-proband duos via long-read sequencing. Am J Hum Genet. 2026:113:437–452. 10.1016/j.ajhg.2026.02.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. de Manuel M, Wu FL, Przeworski M. A paternal bias in germline mutation is widespread in amniotes and can arise independently of cell division numbers. eLife. 2022:11:e80008. 10.7554/eLife.80008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Goldmann JM, Wong WSW, Pinelli M, Farrah T, Bodian D, Stittrich AB, Glusman G, Vissers LELM, Hoischen A, Roach JC, et al. Parent-of-origin-specific signatures of de novo mutations. Nat Genet. 2016:48:935–939. 10.1038/ng.3597. [DOI] [PubMed] [Google Scholar]
  6. Jónsson H, Sulem P, Kehr B, Kristmundsdottir S, Zink F, Hjartarson E, Hardarson MT, Hjorleifsson KE, Eggertsson HP, Gudjonsson SA, et al. Parental influence on human germline de novo mutations in 1,548 trios from Iceland. Nature. 2017:549:519–522. 10.1038/nature24018. [DOI] [PubMed] [Google Scholar]
  7. Kong A, Frigge ML, Masson G, Besenbacher S, Sulem P, Magnusson G, Gudjonsson SA, Sigurdsson A, Jonasdottir A, Jonasdottir A, et al. Rate of de novo mutations and the importance of father's age to disease risk. Nature. 2012:488:471–475. 10.1038/nature11396. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Mahmoud M, Agustinho DP, Sedlazeck FJ. A Hitchhiker's Guide to long-read genomic analysis. Genome Res. 2025:35:545–558. 10.1101/gr.279975.124. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Martin M, Patterson M, Garg S, Fischer SO, Pisanti N, Klau GW, Schönhuth A, Marschall T. WhatsHap: fast and accurate read-based phasing. bioRxiv. 2016. 10.1101/085050. [DOI] [Google Scholar]
  10. Narasimhan VM, Rahbari R, Scally A, Wuster A, Mason D, Xue Y, Wright J, Trembath RC, Maher ER, van Heel DA, et al. Estimating the human mutation rate from autozygous segments reveals population differences in human mutational processes. Nat Commun. 2017:8:303. 10.1038/s41467-017-00323-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Porubsky D, Dashnow H, Sasani TA, Logsdon GA, Hallast P, Noyes MD, Kronenberg ZN, Mokveld T, Koundinya N, Nolan C, et al. Human de novo mutation rates from a four-generation pedigree reference. Nature. 2025:643:427–436. 10.1038/s41586-025-08922-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Sasani TA, Pedersen BS, Gao Z, Baird L, Przeworski M, Jorde LB, Quinlan AR. Large, three-generation human families reveal post-zygotic mosaicism and variability in germline mutation accumulation. eLife. 2019:8:e46922. 10.7554/eLife.46922. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Thomas GWC, Wang RJ, Puri A, Harris RA, Raveendran M, Hughes DST, Murali SC, Williams LE, Doddapaneni H, Muzny DM, et al. Reproductive longevity predicts mutation rates in primates. Curr Biol. 2018:28:3193–3197.e5. 10.1016/j.cub.2018.08.050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Veltman JA, Brunner HG. De novo mutations in human genetic disease. Nat Rev Genet. 2012:13:565–575. 10.1038/nrg3241. [DOI] [PubMed] [Google Scholar]
  15. Wang RJ, Thomas GWC, Raveendran M, Harris RA, Doddapaneni H, Muzny DM, Capitanio JP, Radivojac P, Rogers J, Hahn MW. Paternal age in rhesus macaques is positively associated with germline mutation accumulation but not with measures of offspring sociability. Genome Res. 2020:30:826–834. 10.1101/gr.255174.119. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Wang RJ, Peña-Garcia Y, Bibby MG, Raveendran M, Harris RA, Jansen HT, Robbins CT, Rogers J, Kelley JL, Hahn MW. Examining the effects of hibernation on germline mutation rates in grizzly bears. Genome Biol Evol. 2022a:14:evac148. 10.1093/gbe/evac148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Wang RJ, Raveendran M, Harris RA, Murphy WJ, Lyons LA, Rogers J, Hahn MW. De novo mutations in domestic cat are consistent with an effect of reproductive longevity on both the rate and spectrum of mutations. Mol Biol Evol. 2022b:39:msac147. 10.1093/molbev/msac147. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Wang RJ, Peña-García Y, Raveendran M, Harris RA, Nguyen T-T, Gingras M-C, Wu Y, Perez L, Yoder AD, Simmons JH. Unprecedented female mutation bias in the aye-aye, a highly unusual lemur from Madagascar. PLoS Biol. 2025:23:e3003015. 10.1371/journal.pbio.3003015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Wooldridge TB, Ford SM, Conwell HC, Hyde J, Harris K, Shapiro B. Direct measurement of the mutation rate and its evolutionary consequences in a critically endangered mollusk. Mol Biol Evol. 2025:42:msae266. 10.1093/molbev/msae266. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Wu FL, Strand AI, Cox LA, Ober C, Wall JD, Moorjani P, Przeworski M. A comparison of humans and baboons suggests germline mutation rates do not track cell divisions. PLoS Biol. 2020:18:e3000838. 10.1371/journal.pbio.3000838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Yoder AD, Tiley GP. The challenge and promise of estimating the de novo mutation rate from whole-genome comparisons among closely related individuals. Mol Ecol. 2021:30:6087–6100. 10.1111/mec.16007. [DOI] [PubMed] [Google Scholar]
  22. Young CL, Beichman AC, Mas Ponte D, Hemker SL, Zhu L, Kitzman JO, Shirts BH, Harris K. A maternal germline mutator phenotype in a family affected by heritable colorectal cancer. Genetics. 2024:228:iyae166. 10.1093/genetics/iyae166. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1
media-1.docx (423.8KB, docx)

Data Availability Statement

The pipeline is available at https://github.com/TrangNg-Th/OOPS-Only-One-Parent-Sequencing and is distributed as a Conda package (oops-dnm) for straightforward installation.


Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES