Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2023 Nov 1.
Published in final edited form as: Hum Mutat. 2022 Sep 18;43(11):1531–1544. doi: 10.1002/humu.24465

Long-read sequencing for molecular diagnostics in constitutional genetic disorders

Laura K Conlin 1,2, Erfan Aref-Eshghi 1, Deborah A McEldrew 1, Minjie Luo 1,2, Ramakrishnan Rajagopalan 1,2,*
PMCID: PMC9561063  NIHMSID: NIHMS1835394  PMID: 36086952

Abstract

Long-read sequencing (LRS) has been around for more than a decade, but widespread adoption of the technology has been slow due to the perceived high error rates and high sequencing cost. This is changing due to the recent advancements to produce highly accurate sequences and the reducing costs. LRS promises significant improvement over short read sequencing in four major areas: 1) better detection of structural variation 2) better resolution of highly repetitive or non-unique regions 3) accurate long-range haplotype phasing and 4) the detection of base modifications natively from the sequencing data. Several successful applications of LRS have demonstrated its ability to resolve molecular diagnoses where short-read sequencing fails to identify a cause. However, the argument for increased diagnostic yield from LRS remains to be validated. Larger cohort studies may be required to establish the realistic boundaries of LRS’s clinical utility and analytical validity, as well as the development of standards for clinical applications. We discuss the limitations of the current standard of care, and contrast with the applications and advantages of two major LRS platforms, PacBio and Oxford Nanopore, for molecular diagnostics of constitutional disorders, and present a critical argument about the potential of LRS in diagnostic settings.

Keywords: Long-read sequencing, PacBio, Oxford Nanopore, Molecular diagnostics, Constitutional disorders

Background

Over the past decade, diagnostic testing in individuals suspected of having a genetic disease has evolved from karyotype analysis and targeted gene sequencing to genome-wide approaches including chromosomal microarray analysis (CMA) and whole exome sequencing. As a result, the number of genetic diagnoses has significantly increased. Despite this, the limitations of the current standard of care tests in identifying several types of genetic variants prevent us from reaching the maximum detection rates in the affected individuals.

Limitations of the current technologies in diagnostic testing

Chromosomal microarrays (SNP arrays and array Competitive Genomic Hybridization (aCGH)) have been considered the gold standard to detect submicroscopic deletions and duplications greater than 50 kb in length for over a decade. Currently, CMA is a first-tier diagnostic test for many congenital disorders, including developmental delay, autism, and multiple congenital anomalies (Miller et al., 2010). Targeted arrays, often focusing on exonic regions, are also companions for panel-based or targeted gene testing. While CMA is the current standard of care test for detecting genome-wide copy number variants (CNVs), it has several limitations. The probes in CMA are not evenly distributed across the genome, and the number of probes in a given region will dictate the resolution of the CNVs it can detect. Since probes are spaced at intervals with no coverage in between, determining nucleotide-level breakpoints is impossible, and there is always uncertainty associated with the breakpoints of a CNV, which may be crucial for interpretation. By design, CMA can detect a variant only when there is a CNV associated with it and thus cannot detect balanced rearrangements or additional rearrangements associated with the copy number change (e.g., unbalanced translocation). While CMA can detect duplications, it cannot provide the duplicated material’s precise insertional location and orientation; CMA cannot differentiate between tandem, dispersed, and inverted duplications.

Short read Next Generation Sequencing (NGS) is the most used technology in sequencing assays since it can produce nearly error-free sequences at a very low cost. The next generation sequencers produce several copies of the same DNA fragment as short reads of length 75–250 bp which are then aligned to a reference genome using bioinformatics tools (Goodwin et al., 2016). Paired-end sequencing is a specific method that sequences the ends of DNA fragments, producing two short reads originating from the same fragment, but not sequencing the DNA in between (insert). Computational tools leverage this design using the size of the insert and orientation to detect structural variants (SVs). Several NGS applications exist depending on the genetic material sequenced (e.g., DNA, RNA) and the extent of the targeting of genomic regions by capture or amplification (e.g., exome, targeted panels). Today, various applications of NGS are widely used in diagnosing genetic disorders (Adams & Eng, 2018). Short-read sequencing is best suited for the detection of sequence variants (single nucleotide substitutions and small insertion/ deletions less than or equal to 50 base pairs in length), but it does not perform well in repetitive and low complexity regions of the genome. The ability to unambiguously map a short read in a repetitive region is limited by its length. For example, a 150 bp short read cannot be mapped uniquely to a region if the overlapping repeat element is longer than the read length (Karimzadeh et al., 2018). This limits the ability of short-read sequencing to detect variants in low complexity regions comprehensively.

More than half of the human genome is comprised of repetitive elements such as Short Tandem Repeats (STR) and interspersed repeat elements (e.g., SINE, LINE, Alu). Segmental duplications or Low Copy Repeats (LCR) are large contiguous DNA segments (10–300kb in length) repeated in the genome with greater than 90% sequence identity and occur at multiple sites across the genome (Eichler, 2001; Emanuel & Shaikh, 2001), often with various orientations and number of repeats differing between individuals. Mandelker et al. have cataloged repetitive regions in 4,773 clinically relevant genes and designated 4,264 exons in 619 genes as inaccessible by short-read sequencing (Mandelker et al., 2016). Another 7,691 exons in 1,168 genes were identified as highly challenging for short-read sequencing with a high risk for misalignments. The presence of repetitive elements affects short-read sequencing in many ways. Simple sequence repeats such as homopolymers or regions of low complexity, as well as high GC content, severely affect the sequencing efficiency and per-base sequence quality. Second, mapping short reads to repetitive regions results in a massive pileup of ambiguously aligned reads resulting in artificially inflated coverage metrics. Third, determining the exact orientation of short reads or fragments in these regions is necessary for detecting SVs from short-read paired-end sequencing, but the aligners cannot determine the origin of the fragment. Several thousands of SVs per genome can be identified from short-read sequencing data by current bioinformatics tools. These SVs include reference and mapping errors and can be enriched for false positives in complex regions of the genome. A combination of the factors mentioned above makes it challenging to detect SVs using short-read sequencing data, which can have an impact on diagnostic testing. Detection of SVs from short-read sequencing data has been a major challenge and is seen as one of the major hurdles in the comprehensive detection of pathogenic variants (Kosugi et al., 2019).

Despite these challenges, the continuously reducing cost and the advancements in bioinformatics algorithms continues to make short read sequencing a desirable technology for first tier diagnostic testing. Short read sequencing is ideal for small sequence variant (<50bp) discovery and the tools for structural variant detection are continuously improving over the years. Several specialized short-read variant callers have been developed to detect challenging variants (e.g., short tandem repeats, pseudogene discrimination). For example, ExpansionHunter was developed to detect short tandem repeat expansions from short-read sequencing data with high accuracy and it currently supports 31 known disease associated loci (Dolzhenko et al., 2019; Dolzhenko et al., 2017). Ibanez et al. (2022) used the ExpansionHunter tool to assay 13 most common repeat expansion loci associated with neurodevelopmental disorders and showed 97.3% sensitivity and 99.6% specificity compared to the standard assays for repeat expansions (Ibanez et al., 2022). Similarly, several tools have been developed to discriminate variants in specific disease-associated protein-coding genes from their pseudogenes: Gauchian (Toffoli et al., 2021) for GBA and its pseudogene GBAP, SMNCopyNumberCaller (Chen et al., 2020) for SMN1 and SMN2, and Cyrius (Chen et al., 2021) for CYP2D6 and its pseudogene CYP2D7. Continuous development of bioinformatics algorithms and tools continue to push the limits of variant detection from short-read sequencing data and maximize its clinical utility.

One application of short-read technology, exome sequencing, is a suggested first tier test for pediatric patients with congenital anomalies, developmental delay, and intellectual disability, as it can identify most of the pathogenic disease-causing variation in the uniquely mapped protein-coding regions of the genome at a fraction of the cost of sequencing the whole genome (Manickam et al., 2021). Whole-genome sequencing using short reads is evolving to become a first-tier diagnostic test due to the continuously reducing sequencing costs and its versatility in detecting small sequence variants, CNVs, and structural rearrangements in a single assay (Costain et al., 2020). The diagnostic yield of exome sequencing has been reported with a wide range (15–50%) (Manickam et al., 2021) (Sanchez-Luquez et al., 2022) (Srivastava et al., 2020) (Mellis et al., 2022) and whole-genome sequencing has been shown to provide only a modest improvement over the exome sequencing in head-to-head comparisons (Kingsmore et al., 2019) (Manickam et al., 2021). This leaves a large proportion of individuals with a clinical indication of a genetic disorder and no molecular diagnosis.

A significant proportion of the patients with a clinical indication of a genetic disease go through a long diagnostic odyssey involving several genetic tests (Wu et al., 2020), since often no single diagnostic test can detect all major classes of pathogenic variation. Failure to identify a genetic cause using the standard of care assays can be attributed to several factors, including 1) the lack of comprehensive variant detection in known disease-causing genes using current standard of care assays, 2) difficulties in interpretation of variants in yet unknown disease-causing genes or regulatory/non-coding regions, 3) the lack of ability to infer inheritance/segregation in proband-only diagnostic tests, 4) the lack of relevant tissues available for genetic testing, and 5) the possibility of non-genetic causes. In addition, due to the length of the sequencing reads, additional computation and/or testing is needed to phase the variants; often sequencing of biological parents is required to determine if variants are located on the same allele (cis) or different alleles (trans). The dependence on familial testing in diagnostics presents an ethical challenge in providing equitable care for patients without access to biological parental samples.

The current standard of care genetic testing requires multiple orthogonal assays for a comprehensive assessment of pathogenic variation in known disease-causing genes. This may be due to the limitation of any single technology to detect sequence and SVs with high confidence in the same experiment, or the complex mechanisms involved in the genetic etiology of the disease (e.g., mosaicism, uniparental disomy, genomic imprinting). For example, the current standard of care testing for hearing loss (vignette) includes a range of technologies including long-range PCR and targeted droplet digital PCR for detecting variants in specific genes with known pseudogenes. The diagnostic workflow for Beckwith-Wiedemann syndrome (BWS; MIM# 130650)may include a combination of methylation analysis of the imprinting centers in 11p15.5, CMA analysis for detection of CNVs or in the case of SNP array, copy neutral loss of heterozygosity associated with uniparental isodisomy (UPD), and sequencing assays for CDKN1C (MIM# 600856). Also, it is not uncommon to order both CMA for CNV analysis and exome sequencing for sequencing variants simultaneously. A survey of clinically ordered tests between 2016 and 2020 at the Children’s Hospital of Philadelphia revealed that 40% of the patients who were referred for exome sequencing test (prior to the implementation of CNV detection from exomes) also had a SNP array ordered. These additional tests increase the costs, complexity, and the time to achieve a definitive diagnosis.

Vignette 1: Current standard of care to perform comprehensive diagnostic testing for hearing loss

Comprehensive genetic testing for hearing loss often includes large, targeted sequencing panels for detecting single nucleotide variants and small insertion/deletions, with complimentary tests, sometimes using orthogonal methods to detect CNV/SNVs in complex regions. The comprehensive diagnostic test for hearing loss offered by the Genomic Diagnostics Lab at the Children’s Hospital of Philadelphia, the “Audiome” panel (Guan et al., 2018), uses a tiered testing protocol with an array of technologies to detect sequence and copy-number variants in genes associated with nonsyndromic hearing loss and its mimics. The current version of the Audiome involves a Tier 1 that uses fragment analysis and Sanger sequencing focusing on GJB2 (DFNB1A; MIM# 220290) and mitochondrial variants, with reflex to Tier 2 involving droplet-digital PCR and long-range PCR followed by NGS for STRC (MIM# 606440), exon-targeted custom-designed array CGH, and a virtual exome “slice” for sequence and copy number variants in 132 genes, making it one of the most complex tests offered in our diagnostic laboratory (Figure 1). Audiome is offered as a proband-only test to identify potentially pathogenic variants, with targeted parental follow-up if phasing or segregation is needed to aid in interpretation. Patients with two disease-causing variants in any of the autosomal recessive hearing loss genes are not considered to have a diagnostic result unless the phase is confirmed. This contributes to a relatively high rate of uncertainty, which can be difficult for families. While parental testing adds to the costs and complexity of the tests, it also presents an ethical challenge in providing equitable care for patients without access to parental samples. Similar panels are offered by several other academic and commercial labs including large, targeted gene panels and complementary assays to cover the spectrum of variants in these genes.

Figure 1. Technologies used in comprehensive diagnostic test for nonsyndromic hearing loss, Audiome.

Figure 1.

The comprehensive diagnostic test for Hearing Loss, Audiome, offered by the Genomic Diagnostics Lab at the Children’s Hospital of Philadelphia uses a tiered testing approach and a combination of six technologies for comprehensive variant detection (A,B,C,D,E,F). Tier 1 testing uses Sanger sequencing and fragment analysis focusing on GJB2 (DFNB1A; MIM# 220290) (A, B) and mitochondrial variants, with reflex to Tier 2 involving droplet-digital PCR (C) and long-range PCR followed by NGS (D) for STRC (MIM# 606440), a virtual exome “slice” for sequence and copy number variants (E) and exon-targeted custom-designed array CGH (F) for 132 genes.

Despite the attempt to design a comprehensive panel using multiple technologies, these tests have serious limitations in detecting all major classes of variation in known genes for hearing loss (e.g., inversions). Short-read sequencing has been well documented to have poor performance in challenging regions of the genome (e.g., repeats, segmental duplications, pseudogenes), which affects several known hearing loss genes (e.g., STRC, ESPN (MIM# 606351), OTOA (MIM# 607038)). Taken together, sequencing panels for hearing loss genes are typically limited to pre-determined targeted regions and fail to identify less common variant types in known disease-causing genes (e.g., atypical deletions, SVs, and regulatory variants). A significant proportion of patients receive a non-diagnostic result due to only a single, disease-associated variant in a known hearing loss gene; however, due to the testing limitations, a second variant may have been missed due to the variant type (e.g., SVs) or lack of coverage in the targeted assays. These limitations in the current standard of care are fundamental to the methodologies used and the types of variants that can be detected.

Long-read sequencing (LRS)

Rapidly emerging long-read sequencing (LRS) technologies can produce sequencing reads orders of magnitude longer (from 10 kilobases up to several megabases) than the current second-generation short-read sequencing which typically produce reads of length 150 – 250 base pairs. They are together called third-generation single-molecule technologies and include several emerging LRS and mapping platforms. However, for this review, we focus on two major LRS platforms, Pacific Biosciences (PacBio; Pacific Biosciences, Menlo Park, CA, USA) and Oxford Nanopore Technologies (ONT; Oxford Nanopore Technologies, Oxford, United Kingdom), and specifically discuss their potential advantages, applications in molecular diagnostics, and the limitations and challenges for applications in clinical diagnostics.

PacBio high fidelity (HiFi) sequencing starts with double-stranded high molecular weight DNA, which is ligated to create a circularized DNA. This circular DNA gets passed through several rounds of sequencing to create subreads from which a highly accurate consensus long read is generated. PacBio’s most recent sequencing platform Sequel II was released in 2019, which produces long HiFi reads of lengths 10–25 kilobases with >99.9% per-base quality (Wenger et al., 2019). Oxford Nanopore uses a motor protein to pass single-molecule DNA or RNA through a protein nanopore to measure the changes in the current produced by four nucleotides (Wang et al., 2021). Oxford Nanopore sequencing does not have a theoretical limit for maximum read length, and researchers have pushed the limits over the years to produce megabase-long reads (Payne et al., 2019). Ultra-long sequencing kits from ONT enable the consistent production of long reads with N50 > 50kb (Quick, 2018).

Research studies over the last several years have highlighted the utility of long reads in diagnosing genetic disorders where the standard of care assays fail to achieve a molecular diagnosis (Amarasinghe et al., 2020; Ardui et al., 2018; Borras et al., 2017; Melas et al., 2022; Miller et al., 2022; Miller et al., 2021). These studies fall into the two broad categories: 1) identifying pathogenic variants in regions difficult to sequence with short-reads: repeat expansions, homopolymers, protein-coding genes with pseudogenes, long insertions (>50 base pairs), and complex structural rearrangements, and 2) resolving unknown or uncertain phase in autosomal recessive conditions where short-read technology requires additional sequencing of parental samples. However, it is challenging to determine the additional diagnostic yield solely as the benefit of LRS since the standard of care genetic tests differ greatly depending on several factors and the lack of large cohorts utilizing LRS for Mendelian diagnostics.

The need for LRS for molecular diagnostics arises from the technical limitations of short-read sequencing in resolving highly repetitive and non-unique regions of the genome, and the room for additional diagnostic yield from better variant discovery in disease-causing genes. Besides the scientific rationale of better variant discovery, there is a strong operational rationale for LRS in clinical diagnostic settings as it may offer the most comprehensive single diagnostic test filling the gaps in testing using short-read whole-genome sequencing.

Identifying challenging variants using long-read sequencing

Several demonstrative applications of LRS have been published since the introduction of PacBio’s Single-Molecule Real-Time (SMRT) and Oxford Nanopore sequencing technology. Many of these include the identification of variants that are challenging to detect using the standard technologies, such as sequence repeat expansions, homologous genes, and genes with pseudogenes.

One of the earliest applications of PacBio’s SMRT sequencing technology published in 2013 demonstrated the possibility of sequencing the (CCG)n trinucleotide repeats and length heterogeneity in the FMR1 (MIM# 309550) gene which causes Fragile X syndrome (FXS; MIM# 300624) (Loomis et al., 2013). Later in 2019, Wieben et. al., used PacBio LRS to assay CAG repeats in TCF4 (MIM #602272) and demonstrated the ability to capture the heterogeneity at this locus (Wieben et al., 2019). Several publications followed suit to demonstrate the ability of LRS to assay repeat elements in clinically relevant genes such as ATXN10 (MIM# 611150) (McFarland et al., 2015; Schule et al., 2017), TAF1 (MIM# 313650) (Aneichyk et al., 2018), C9orf72 (MIM# 614260) (Ebbert et al., 2018), DMPK (MIM# 605377) (Cummings et al., 2017), SAMD12 (MIM# 618073) (Cen et al., 2019), and HTT (MIM# 613004) (Hoijer et al., 2018). More recently, Cohen et al. used PacBio sequencing in a large cohort of undiagnosed patients to uncover molecular diagnoses including a pathogenic pentamer expansion in the gene STARD7 (MIM# 616712) in a patient with global developmental delay and dystonia (Cohen et al., 2022). Stevanovski et al. used the adaptive sequencing protocol in the Oxford Nanopore platform to target 37 loci and showed the ability to detect repeat expansions accurately in 25 patients with neurological disorders. The authors demonstrate the ability of LRS to accurately detect haplotype resolved sizing of repeat alleles and methylation profiling in a single experiment (Stevanovski et al., 2022). Melas et al. used PacBio sequencing in two multi-generational pedigrees with a clinical diagnosis of synpolydactyly (SPD1; MIM# 186000) and no molecular diagnosis using short-read whole-genome sequencing (Melas et al., 2022). PacBio sequencing revealed heterozygous polyalanine tract duplications (21 and 27 basepairs in length) in both pedigrees in the HOXD13 gene. Manual inspection of short-read sequencing data showed soft-clipped reads while the GATK (v4.0.5) failed to detect the insertions. The authors highlight the low coverage in short-read whole genome sequencing data due to the high GC content and the challenge to detect variants in low complexity regions in general; however, they found that the specialized variant caller for repeat expansions ExpansionHunter identified seven of the eight variants in two families (Melas et al., 2022). The standard protocols currently in use for diagnosing nucleotide repeat expansion disorders in clinical labs involve Southern blotting or PCR-based assays which only target a single region per experiment, both of which are labor-intensive. The ideal assay for the detection of repeat expansion disorders should be able to screen multiple genomic regions simultaneously given the phenotypic overlap of many such disorders, identify the repeat counts, and detect additional clinically relevant markers including DNA methylation, repeat interruptions, and point mutations. By spanning over the length of the repeated sequence, LRS can capture the full repeat length as well as their interruptions. Also, signal kinetics from PacBio and Oxford Nanopore platforms enable the detection of methylation status of each cytosine, which has diagnostic value in several repeat expansion disorders such as Fragile X syndrome and congenital myotonic dystrophy where hypermethylation of the CG repeats are involved in the pathogenesis. For many repeat expansion diseases that can also be caused by sequence variants, the long reads enable both detection and phasing of the sequence variants with the expanded repeats, a procedure that currently requires the use of multiple parallel assays per condition and is not performed in most laboratories (Svrzikapa et al., 2020).

Homologous genes and genes with pseudogenes are the most challenging areas of molecular diagnostics currently due to limited differentiation ability. One such region is the alpha hemoglobin gene cluster on chromosome 16 which contains two homologous genes (HBA1 (MIM# 141800) and HBA2 (MIM# 141850)). Deletions within this region are the most common cause of Hb A thalassemia (MIM# 604131). These deletions are complex and can remove one or both HBA genes, or generate a hybrid functional gene. A further complication arises from the presence of common structural polymorphisms in the region, a frequent cause of allele dropout in traditional targeted assays. Similar mappability issues and common polymorphisms are seen with the beta hemoglobin (HBB (MIM# 141900)) gene cluster on chromosome 11, spans ~70kb and involves six genes. As a result, NGS with short reads is currently not used for the diagnosis of hemoglobinopathies. The current clinical testing methodology is limited to PCR-based assays, Sanger sequencing, MLPA, and targeted mutation screening. These assays target the commonly known CNVs within the HBA region, but the breakpoints, the number of genes involved, the complexity of the rearrangement, and the phasing of the variants are not determined. Long-read sequencing is shown to be capable of resolving these issues by simultaneously providing information on both HBA and HBB loci, determining CNVs and their breakpoints, identifying sequence variants, and determining the phase, all in one assay (Xu et al., 2020).

A similar approach can be applied to gene with one or more known pseudogenes. Reads produced by short-read sequencing are not specific enough to one locus, and it can be impossible to determine the origin of a variant based on short-read NGS data only. Clinical labs tend to use additional assays such as long-range PCR to target these regions, which is time-consuming and yet, not always effective in determining the origin of a variant. One such example is PKD1 (MIM# 601313) loci, which is known to cause Autosomal Dominant Polycystic Kidney Disease (PKD1; MIM# 173900). Diagnostic testing for PKD1 is complicated by the presence of six pseudogenes as well as high GC content producing low sensitivity and high false positive rate in the duplicated region of PKD1. Borras et al. used targeted long-range PCR followed by PacBio sequencing to discover pathogenic variants in PKD1 and demonstrated superior variant detection capabilities compared to the short-read sequencing methodologies (Borras et al., 2017). Another example is IKBKG (MIM# 300248), which is associated with primary immunodeficiency, and its pseudogene, IKBKGP1. Both genes are located within low copy repeats with >99% sequence identity and long-range PCR followed by sequencing is the commonly used method to differentiate the protein-coding gene from the pseudogene. A study in 2018 (Frans et al., 2018) compared the performance of targeted PacBio’s long-read sequencing with Illumina’s short-read sequencing following long-range PCR in detecting IKBKG variants. Both assays successfully unambiguously distinguished the gene from its pseudogene; however, the specificity of PacBio in calling the variants was not as high as Illumina’s short-read assay after targeted amplification. The authors indicated that PacBio’s workflow was simpler and cheaper, and the lack of computational variant callers specifically designed for long-read data was likely the main obstacle in the use of this technology (Frans et al., 2018). It is important to highlight these papers used an older PacBio RS II platform with a maximum subread length of 8kb and the PacBio’s Long Amplicon Analysis (LAA) software for data analysis which failed to produce reliable results for mapping. We believe the current generation PacBio Sequel II platform producing HiFi reads up to 20kb in length and the latest variant calling algorithms for PacBio data will resolve these regions at a much higher sensitivity and specificity.

The real power of LRS comes from identifying challenging variants occurring in complex regions of the genome. Cretu-Stancu et al. sequenced two patients with seemingly de novo chromothriptic events using Oxford Nanopore sequencing and compared it to Illumina short read sequencing (Cretu Stancu et al., 2017). They demonstrated that the Oxford Nanopore sequencing was able to identify many SVs missed by the short read sequencing and phase the extremely complex rearrangements. More recently, Hiatt et al. used PacBio sequencing to identify pathogenic structural variation including complex rearrangements. In one of the probands, they identified a likely pathogenic de novo L1-mediated insertion in CDKL5 (MIM# 300203) which was cryptic to short read sequencing (Hiatt et al., 2021). In another proband, they identified multiple de novo SVs including two insertional translocations in a chromothriptic event involving three chromosomes.

Long-range phasing using long-read sequencing

Another advantage of LRS over the short reads is the capability to infer phase with high accuracy over long genomic intervals spanning several hundred kilobases to megabase in some cases (Shafin et al., 2021). Phasing, referred to as assigning alleles to paternal or maternal chromosomes, is a crucial step for making definitive diagnosis in an individual with two variants in a gene linked to an autosomal recessive phenotype. Compound heterozygous disease-causing variants are usually the more frequent etiology for rare autosomal recessive disorders compared to homozygous variants, given the chance of two unrelated parents carrying the same rare, disease-causing variant. When two heterozygous disease-causing variants are identified in an individual with a suspected autosomal recessive disease, the next step requires determining whether they are located on the same or alternative chromosomes (cis vs. trans); if the variants are in trans, the genetic diagnosis is reached, while in cis, the cause remains inconclusive. Another situation when phasing is crucial is in a disease occurring due to mosaicism or determining the parental origin of a de novo variant. Determining the phase of these variants can provide information on the clonality of multiple mosaic variants observed in tumor tissue or help determine parent of origin for variants associated with imprinting disorders. Phasing is also used in identifying disease-risk haplotypes as well as in RNA-seq data for determining allele-specific expression.

Using short-read data, phasing can be accomplished if the two variants are located close enough to fall on the same short read or the paired reads. Alternatively, in the presence of full gene sequence data and by utilizing SNP haplotyping the phase can be computationally established (Menelaou & Marchini, 2013). This method is not used frequently in diagnostic testing since most clinical tests target exons and do not cover the polymorphic intronic regions needed for this purpose, and the targeted SNPs may not be informative to all ancestries. There are other molecular methods involving allele-specific amplification followed by sequencing or quantification to determine phase (e.g. Drop-phase); however, these methods are targeted to specific loci or an individual’s genotypes. LRS, on the other hand, provides an opportunity to perform hypothesis-free genome-wide phasing in a single experiment. The standard of care for phasing in clinical labs today is the testing of both parents to determine the inheritance of each variant and infer the phase. In cases where one of the variants has occurred de novo or is only present in gonadal mosaicism in one parent, this approach will not be fruitful. LRS is proposed to be a more effective tool to determine the phase using the occurrence of two variants in one read, which can cover the entire length of a gene, as well as using the haplotype phasing approach in larger genes. The latter tends to be more accurately performed using long reads, for which many computational tools are under development (Maestri et al., 2020).

Several research publications highlight the potential clinical utility in sequence and phase resolved SVs from LRS data. For example, Miao et al. used Oxford Nanopore sequencing on a patient with a clinical diagnosis of Glycogen Storage Disease and inconclusive prior genetic testing to achieve a definitive diagnosis (Miao et al., 2018). Exome sequencing test a homozygous single nucleotide variant, c.326G>A; p.C109Y in G6PC (MIM# 613742), which was classified as pathogenic. Sanger testing of the parents revealed that the mother was a carrier while the father did not carry this variant. Authors hypothesized the possibility of an SV and performed low coverage Oxford Nanopore sequencing to identify a 7.1kb deletion in the G6PC gene overlapping the previously identified single nucleotide variant. LRS revealed the deletion and the single nucleotide variant in the same experiment with phase resolution in the same experiment. More recently, Miller et al. used Oxford Nanopore sequencing in a cohort of nine patients with a clinical diagnosis of Werner Syndrome (WRN; MIM# 277700) (Miller et al., 2022). Werner syndrome is an autosomal recessive disorder caused by loss of function variants in the gene WRN (MIM# 604611). Standard of care testing protocol for Werner syndrome is the sequence analysis of the WRN gene followed by the targeted deletion/ duplication analysis. Sequence analysis identifies a pathogenic variant in 97% of the patients and only six copy number variants have been reported so far. A small subset of patients (9/188) in the International Registry of Werner Syndrome had a single pathogenic variant identified in WRN from prior genetic testing without a second pathogenic variant resulting in an uncertain diagnosis. Miller and colleagues used targeted Oxford Nanopore sequencing in these patients and identified a second pathogenic variant in eight of the nine patients tested. Authors highlight the advantage of phase inference from the long reads in patients without parental samples as the major advantage against a short-read sequencing approach. In this paper, we provide examples where the phase inference from LRS resolved the uncertain diagnoses from the standard of care testing in two patients with nonsyndromic bilateral sensorineural hearing loss (case vignettes).

Vignette 2 – Phasing of a single nucleotide variant and a deletion

A 7-year-old female with bilateral nonsyndromic sensorineural hearing loss was referred for genetic testing by the CHOP hearing loss clinic. Audiome Tier 1 was negative, and Tier 2 testing revealed two variants of uncertain significance (VUS) in TRIOBP (MIM# 609761), which is associated with autosomal recessive deafness 28 (DFNB28; MIM# 609823). A missense variant (chr22(GRCh38):g.37723247T>A; NM_001039141.2:c.691T>A; p.Leu231Met) variant was detected from the exome-slice panel and was not seen in any large public genomic databases (gnomAD v2.1.1). The variant was classified as a VUS and would require further evidence to determine its clinical significance. Exome-based copy number analysis (Rajagopalan et al., 2020) identified a deletion resulting in loss of exons 10 and 11 in the longest isoform of TRIOBP. While the deletion was predicted to maintain the reading frame, the exact breakpoints of this deletion cannot be determined by this analysis. Similar deletions are observed in the general population, with 16 heterozygotes and 1 homozygote in gnomAD SVs v2.1 (DEL_22_182624; chr22(GRCh38):g.37736684_37741331del; NC_000022.11:g.37736684_37741331del). The two TRIOBP variants were 13kb apart, and short-read sequencing could not resolve the phase of these variants without additional targeted Sanger sequencing for the missense variant and droplet digital PCR for the deletion in the parents.

The proband was consented for research studies using an IRB-approved research protocol (CHOP IRB# 16-013231), and a lymphoblastoid cell line (LCL) was established. High molecular weight genomic DNA from the LCL was extracted using the Qiagen Blood and Cell Culture DNA Mini Kit (Qiagen LLC, Germantown, MD, USA) and the quality was checked using the Nanodrop and quantitation was done with the Qubit 2.0 Fluorometer (Thermo-Fisher Scientific, Waltham, MA, USA). Proband-only PacBio LRS performed on a research basis at the Mt. Sinai sequencing core facility (Mount Sinai Genomics Technology Facility, Icahn School of Medicine at Mount Sinai, New York, NY, USA) using two SMRTcells targeting a genome-wide coverage of 15–20x. HiFi reads were aligned to the hg38 reference genome using pbmm2 (Pacific Biosciences, 2022a), and sequence variants (SNV/Indel) were called using DeepVariant (Google, 2022) and structural variants using pbsv (Pacific Biosciences, 2022c). Variants were phased and reads were haplotagged using the whatshap tool (Patterson et al., 2015). The whatshap tool uses aligned reads and heterozygous variants present in the sample genome to phase individual haplotypes.

PacBio LRS not only detected both variant types with the same technology, but also revealed that the two variants were in cis (Figure 2A) and therefore not diagnostic for the hearing loss in the patient. No additional diagnostic variants were detected in any of the other genes on the Audiome panel using long-read sequencing. Independent validation of the phase inheritance of these variants was not possible since the parental samples were not available.

Figure 2. Comprehensive variant detection and phase resolution without parental samples using PacBio long-read sequencing.

Figure 2.

Panel A shows the phased long-read sequencing (LRS) data for the patient who received an inconclusive result from the standard of care diagnostic testing for hearing loss (case vignette 1). Targeted exome analysis revealed two potentially pathogenic variants (SNV and a deletion) in the gene TRIOBP. The SNV and the deletion was 12 kb apart and the exome analysis was not able to resolve phase without additional sequencing of the parents. LRS detected these two variants and revealed these variants are in the same haplotype (cis) making it non-diagnostic. Panel B shows the phased LRS data for a patient who received an inconclusive result from the standard of care diagnostic testing for hearing loss (case vignette 2). Targeted long-range PCR followed by next-generation sequencing of STRC revealed a variant of uncertain significance in exon 8 and multiple likely pathogenic variants in exons 25 and 26. Variants identified in exons 25/26 of STRC are known to be present in the pseudogene STRCP1 and suggestive of a gene conversion event. Gene conversion events are challenging to resolve using short-reads due to the mappability issues in discriminating protein coding gene and its pseudogene. LRS was able to map reads unambiguously to the protein-coding gene, identified all the variants, and revealed that these variants were in different haplotypes (trans), confirming the gene conversion event to make this a diagnostic result.

Vignette 3 – Phasing of variants between a protein-coding gene and its pseudogene

A 3-year-old female with bilateral sensorineural hearing loss was referred for genetic testing, and the Audiome test revealed multiple variants, with a single heterozygous variant in multiple autosomal recessive disease genes, including a likely pathogenic variant in TRIOBP, and VUSs in ALMS1 (MIM# 606844), CDH23 (MIM# 605516), and USH2A (MIM# 608400). The long-range PCR-based short read NGS testing of STRC revealed multiple variants in exons 25 and 26 (NM_153700.2:c.[4917_4918delinsCT;4903G>T(;)5125A>G];p.[Leu1640Phe;Val1635Phe(;)Thr1709Ala]) which are known to be present in the reference sequence for the pseudogene, STRCP1. Droplet digital PCR for copy number detection revealed one copy of the probe specific for exon 26, which could be consistent with either a deletion or a drop-out event associated with a gene conversion. Gene conversion events between STRC and STRCP1 are typically difficult to resolve with short reads given the mappability issues due to high homology between the real gene and the pseudogene gene. An additional rare heterozygous missense VUS (NM_153700.2:c.2494C>T, p.Arg832Trp) in exon 8 of the protein coding gene STRC, and not known to be present in the STRCP1, was also detected. Due to the inability to phase all these variants in STRC using short-read sequencing, the patient received an inconclusive test result and further familial testing was recommended. The proband was consented for research studies using an IRB-approved research protocol (CHOP IRB# 16-013231) and proband-only PacBio LRS was performed on a research basis (methods in case vignette 1). PacBio long-read sequencing of the proband not only detected all the variants by short-read NGS (both exome and long-range PCR of STRC), but also enabled the phasing of the complex findings in STRC (Figure 2B). The five variants in exons 25 and 26 were indeed in cis and were confirmed that they were associated with a pathogenic haplotype consistent with a gene conversion event, while the missense variant in exon 8 was found to be in trans (Figure 2B). These results supported STRC as the potential underlying genetic etiology for this individual’s hearing loss using only one technology.

Applications of long-read sequencing in emerging areas of molecular diagnostics

Long-range sequencing is expected to revolutionize many areas of research in genomics, epigenomics, and transcriptomics. Among these, RNA sequencing (RNA-seq) and DNA methylation analysis are emerging as diagnostic tests for constitutional disorders (Aref-Eshghi et al., 2019; Cummings et al., 2017; Rentas et al., 2020) and have been much more widely used in cancer diagnostics (Luo et al., 2020; Sidaway, 2020).

Currently most of the RNA analysis in the diagnostic testing is performed to detect the chimeric transcripts associated with diagnostic, prognostic, and/or treatable gene fusion events associated with recurrent somatic translocations, inversion, or insertional events (Cummings et al., 2017; Lee et al., 2020; Yepez et al., 2022). RNA sequencing for constitutional disorders is currently not widely used in clinical laboratories, but many efforts are underway in the translational research setting. Most tests under development are focused on the detection of aberrant splicing events, including targeted functional analysis to help clarify a previously identified variant of uncertain significance identified from clinical DNA analysis. Transcriptome sequencing can also be used to detect differences in patterns of isoform usage and/or expression levels in relevant tissues. The detection of different isoforms from short read data requires the reconstruction of the short reads into the full transcript and is liable to miss some of the isoforms or aberrant splicing events. LRS has made it feasible to add to the accuracy of the isoform detection by sequencing the entire length of a transcript in one read. Long-read sequencing of mRNA has demonstrated the ability to detect many novel reading frames and isoforms previously unknown, in addition to allele-specific expression and variations in post-transcription RNA modifications. This knowledge will have utility in sequence variant interpretation in genetic testing (Bournazos et al., 2022).

DNA methylation alterations have long been known as an underlying etiology in diseases such as imprinting defect conditions. Recent studies are extending the role of DNA methylation in genetically unresolved individuals with developmental delays and congenital anomalies, some of whom are found to carry rare de novo promoter DNA methylation changes in genes potentially linked to their phenotypes (Barbosa et al., 2018; Garg et al., 2020). Additionally, a growing list of genetic syndromes with disease-specific genome-wide DNA methylation patterns are being utilized in diagnostic testing. The most used method in genomic DNA methylation testing is currently array hybridization or short read NGS of bisulfite converted DNA. These methods are not optimal, as they only target limited regions of the genome (in the case of arrays), or technical issues from chemical treatment interfere with interpretation in the case of bisulfite assays. Both Oxford Nanopore and PacBio can detect the methylation status of a nucleotide without any additional prior experimental requirement. By avoiding the bisulfite treatment and PCR amplification and generating long reads with higher specificity, LRS can theoretically bypass many of the above-mentioned issues resulting in uniformity in mapping, less impact of GC bias, lower read depth requirement, and higher reproducibility compared with bisulfite sequencing (Technologies, 2021). Currently, there appears to be room for improving the computational and technical capabilities of LRS in methylation calling, and this method has much promise for the use of DNA methylation analysis in molecular diagnostics. Methylation analysis will likely be routine in clinical bioinformatics workflows, and it aids in resolving uncertainty in variant classification in methylation related disorders.

Considerations in clinical implementation

Several challenges exist despite the advantages of long-read sequencing for adoption in clinical diagnostics. Production of long reads largely depends on using high-quality DNA of high molecular weight and enriched for long fragments, which may be a challenge to achieve in clinical settings. While saliva has been demonstrated as a reliable and common non-invasive source for DNA to produce good quality chromosomal microarray, short-read exome, or genome sequencing data in clinical labs, most of the library preparation protocols for long-read sequencing require high molecular weight DNA from blood or other invasive tissue collection.

Cost, accuracy, and throughput are the major factors when it comes to adopting long-read sequencing into clinical diagnostics for routine use. Any rapidly advancing technology presents a challenge for adoption in clinical labs as they typically require extensive validation before implementation as a diagnostic test. Clinical testing requires high sensitivity to ensure a diagnosis is not missed, and high specificity to minimize false positives. The limitations of the tests must be thoroughly understood and documented to be able to truly rule out a genetic disorder and/or determine the most appropriate clinical follow-up. Until recently, both Oxford Nanopore and PacBio sequencing had higher per-base error rates compared to the short read sequencing. This has changed in the last few years with the introduction of PacBio HiFi sequencing with circular consensus sequencing achieving 99.9% per-base accuracy. More recently, in early 2022, Oxford Nanopore reported 99% single molecule accuracy (modal) with the Q20+ chemistry (Kit 14) combined with the R10.4.1 flowcell (Technologies, 2022). Most recent benchmarking data for PacBio and Oxford Nanopore sequencing using the PEPPER-Margin-DeepVariant variant calling pipeline (Shafin et al., 2021) suggests a minimum of 40–50x coverage for high-quality SNV calls (99.5% recall and precision) from Oxford Nanopore data, which would require sequencing of three to four flowcells on a GridION. Producing high-quality indel calls from Oxford Nanopore data seems to be challenging as 90x coverage has achieved a recall of 60.2% and a precision of 91.3% (Shafin et al., 2021). Comparatively, 35x PacBio HiFi sequencing produced high-quality SNV and indel calls with 99.5% recall and precision. Advancements in the use of machine learning algorithms for base and variant calling for Oxford Nanopore sequencing data continue to close the gap with PacBio HiFi data. Successful phasing of the haplotypes and the length of the haplotype blocks depend on the heterozygous variants in a given genome, with long stretches of homozygous regions or genotyping errors resulting in shorter phase blocks.

In terms of cost, throughput, and speed, the two methods are very different. It takes 30 hours to sequence one PacBio SMRTcell and a maximum of 8 SMRTcells can be run in a sequential manner. A full sequencing run using PacBio Sequel II would therefore take approximately 11 days (240 hours or 10 days for sequencing and an additional day for the HiFi data production). A minimum of 15x coverage is generally recommended for reference alignment-based variant calling or 10–15x per haplotype for assembly-based variant discovery (Pacific Biosciences, 2022b). For a trio genome, it would require sequencing of two or three SMRTcells per proband and one SMRTcell for each parent. This can be several folds more expensive compared to short-read sequencing. There are several instruments currently available from Nanopore ranging from low throughput, portable MinION platform, the medium throughput, desktop sized GridION, and a high throughput benchtop device PromethION. Adaptive sequencing protocols can be used to computationally design targeted sequencing panels without any target enrichment protocol. Miller et al. (2021) used the adaptive sequencing protocol to target up to 151 Mb and demonstrated the ability to identify sequence, structural variation, methylation from the same assay (Miller et al., 2021). Real-time dynamic analysis where the sequencing, basecalling, and alignment happens instantaneously provides an opportunity to rapidly sequence in critical care settings. Recently, Gorzynski et al. used Nanopore to perform ultra-rapid genome sequencing in 12 pediatric patients in critical care setting and achieved a molecular diagnosis in five patients in less than a day with the shortest diagnosis in seven hours (Gorzynski et al., 2022). Similarly, Goenka et al. used Oxford Nanopore sequencing on two patient samples to demonstrate ultra-rapid sequencing and identified candidate variants in less than eight hours from the sample preparation (Goenka et al., 2022). Goenka et al. and Gorzynski et al. used Oxford Nanopore’s high-throughput PromethION platform and 48 flowcells in parallel to achieve the rapid data production. These demonstrations of ultra-rapid sequencing, highlight the potential for this technology where rapid testing at more cost is medically warranted.

Concluding remarks

Long-read sequencing (LRS) has the potential to solve many of the critical challenges in the comprehensive detection of pathogenic variants in patients with suspected genetic disorders. The technologies discussed here are mature and the quality of the data produced on these platforms are comparable to short reads. The cost of LRS has come down considerably over the years and yet, it is still on the higher side compared to short read sequencing. While publications from research are supportive of its clinical utility, the additional diagnostic yield solely from the LRS over short read whole-genome sequencing is not clear. LRS is desirable for clinical testing as it has the potential to become the most comprehensive diagnostic platform to date. LRS can help overcome the challenge of sequential testing paradigm, reduce the uncertainty in diagnoses by offering comprehensive variant detection and simultaneous phase inference, and reduce the time to achieve a definitive diagnosis. We expect the continuously reducing costs, increased quality, and increased throughput will accelerate LRS to become a mainstream tool in clinical settings in the next few years.

Acknowledgements

We would like to acknowledge Roberts Individualized Medical Genetics Center (RIMGC),Children’s Hospital of Philadelphia for providing the lymphoblastoid cell lines for samples presented in the case vignettes. Figures were created using BioRender.com.

Funding

This work is supported by the National Institutes of Health grant R01-HG009708.

Footnotes

Conflict of Interest Statement

The authors declare no conflict of interests.

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.

References

  1. Adams DR, & Eng CM (2018, Oct 4). Next-Generation Sequencing to Diagnose Suspected Genetic Disorders. N Engl J Med, 379(14), 1353–1362. 10.1056/NEJMra1711801 [DOI] [PubMed] [Google Scholar]
  2. Amarasinghe SL, Su S, Dong X, Zappia L, Ritchie ME, & Gouil Q (2020, Feb 7). Opportunities and challenges in long-read sequencing data analysis. Genome Biol, 21(1), 30. 10.1186/s13059-020-1935-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Aneichyk T, Hendriks WT, Yadav R, Shin D, Gao D, Vaine CA, Collins RL, Domingo A, Currall B, Stortchevoi A, Multhaupt-Buell T, Penney EB, Cruz L, Dhakal J, Brand H, Hanscom C, Antolik C, Dy M, Ragavendran A, Underwood J, Cantsilieris S, Munson KM, Eichler EE, Acuna P, Go C, Jamora RDG, Rosales RL, Church DM, Williams SR, Garcia S, Klein C, Muller U, Wilhelmsen KC, Timmers HTM, Sapir Y, Wainger BJ, Henderson D, Ito N, Weisenfeld N, Jaffe D, Sharma N, Breakefield XO, Ozelius LJ, Bragg DC, & Talkowski ME (2018, Feb 22). Dissecting the Causal Mechanism of X-Linked Dystonia-Parkinsonism by Integrating Genome and Transcriptome Assembly. Cell, 172(5), 897–909 e821. 10.1016/j.cell.2018.02.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Ardui S, Ameur A, Vermeesch JR, & Hestand MS (2018, Mar 16). Single molecule real-time (SMRT) sequencing comes of age: applications and utilities for medical diagnostics. Nucleic Acids Res, 46(5), 2159–2168. 10.1093/nar/gky066 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Aref-Eshghi E, Bend EG, Colaiacovo S, Caudle M, Chakrabarti R, Napier M, Brick L, Brady L, Carere DA, Levy MA, Kerkhof J, Stuart A, Saleh M, Beaudet AL, Li C, Kozenko M, Karp N, Prasad C, Siu VM, Tarnopolsky MA, Ainsworth PJ, Lin H, Rodenhiser DI, Krantz ID, Deardorff MA, Schwartz CE, & Sadikovic B (2019, Apr 4). Diagnostic Utility of Genome-wide DNA Methylation Testing in Genetically Unsolved Individuals with Suspected Hereditary Conditions. Am J Hum Genet, 104(4), 685–700. 10.1016/j.ajhg.2019.03.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Barbosa M, Joshi RS, Garg P, Martin-Trujillo A, Patel N, Jadhav B, Watson CT, Gibson W, Chetnik K, Tessereau C, Mei H, De Rubeis S, Reichert J, Lopes F, Vissers L, Kleefstra T, Grice DE, Edelmann L, Soares G, Maciel P, Brunner HG, Buxbaum JD, Gelb BD, & Sharp AJ (2018, May 25). Identification of rare de novo epigenetic variations in congenital disorders. Nat Commun, 9(1), 2064. 10.1038/s41467-018-04540-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Biosciences P (2022a). A minimap2 SMRT wrapper for PacBio data. Retrieved 06/17/2022 from https://github.com/PacificBiosciences/pbmm2
  8. Biosciences P (2022b). Overview of PacBio Sequel Systems Application Options and Sequencing Recommendations. Retrieved 06/17/2022 from https://www.pacb.com/wp-content/uploads/Overview-Sequel-Systems-Application-Options-and-Sequencing-Recommendations.pdf
  9. Biosciences P (2022c). PacBio structural variant (SV) calling and analysis tools. Retrieved 06/17/2022 from https://github.com/PacificBiosciences/pbsv
  10. Borras DM, Vossen R, Liem M, Buermans HPJ, Dauwerse H, van Heusden D, Gansevoort RT, den Dunnen JT, Janssen B, Peters DJM, Losekoot M, & Anvar SY (2017, Jul). Detecting PKD1 variants in polycystic kidney disease patients by single-molecule long-read sequencing. Hum Mutat, 38(7), 870–879. 10.1002/humu.23223 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Bournazos AM, Riley LG, Bommireddipalli S, Ades L, Akesson LS, Al-Shinnag M, Alexander SI, Archibald AD, Balasubramaniam S, Berman Y, Beshay V, Boggs K, Bojadzieva J, Brown NJ, Bryen SJ, Buckley MF, Chong B, Davis MR, Dawes R, Delatycki M, Donaldson L, Downie L, Edwards C, Edwards M, Engel A, Ewans LJ, Faiz F, Fennell A, Field M, Freckmann ML, Gallacher L, Gear R, Goel H, Goh S, Goodwin L, Hanna B, Harraway J, Higgins M, Ho G, Hopper BK, Horton AE, Hunter MF, Huq AJ, Josephi-Taylor S, Joshi H, Kirk E, Krzesinski E, Kumar KR, Lemckert F, Leventer RJ, Lindsey-Temple SE, Lunke S, Ma A, Macaskill S, Mallawaarachchi A, Marty M, Marum JE, McCarthy HJ, Menezes MP, McLean A, Milnes D, Mohammad S, Mowat D, Niaz A, Palmer EE, Patel C, Patel SG, Phelan D, Pinner JR, Rajagopalan S, Regan M, Rodgers J, Rodrigues M, Roxburgh RH, Sachdev R, Roscioli T, Samarasekera R, Sandaradura SA, Savva E, Schindler T, Shah M, Sinnerbrink IB, Smith JM, Smith RJ, Springer A, Stark Z, Strom SP, Sue CM, Tan K, Tan TY, Tantsis E, Tchan MC, Thompson BA, Trainer AH, van Spaendonck-Zwarts K, Walsh R, Warwick L, White S, White SM, Williams MG, Wilson MJ, Wong WK, Wright DC, Yap P, Yeung A, Young H, Jones KJ, Bennetts B, Cooper ST, & Australasian Consortium for RNAD (2022, Jan). Standardized practices for RNA diagnostics using clinically accessible specimens reclassifies 75% of putative splicing variants. Genet Med, 24(1), 130–145. 10.1016/j.gim.2021.09.001 [DOI] [PubMed] [Google Scholar]
  12. Cen Z, Chen Y, Yang D, Zhu Q, Chen S, Chen X, Wang B, Xie F, Ouyang Z, Jiang Z, Fu A, Hu B, Yin H, Qiu X, Yu F, Du X, Hao W, Liu Y, Wang H, Wang L, Yu X, Xiao Y, Liu C, Xiao J, Zhou Y, Yang W, Zhang B, & Luo W (2019, Oct). Intronic (TTTGA)n insertion in SAMD12 also causes familial cortical myoclonic tremor with epilepsy. Mov Disord, 34(10), 1571–1576. 10.1002/mds.27832 [DOI] [PubMed] [Google Scholar]
  13. Chen X, Sanchis-Juan A, French CE, Connell AJ, Delon I, Kingsbury Z, Chawla A, Halpern AL, Taft RJ, BioResource N, Bentley DR, Butchbach MER, Raymond FL, & Eberle MA (2020, May). Spinal muscular atrophy diagnosis and carrier screening from genome sequencing data. Genet Med, 22(5), 945–953. 10.1038/s41436-020-0754-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Chen X, Shen F, Gonzaludo N, Malhotra A, Rogert C, Taft RJ, Bentley DR, & Eberle MA (2021, Apr). Cyrius: accurate CYP2D6 genotyping using whole-genome sequencing data. Pharmacogenomics J, 21(2), 251–261. 10.1038/s41397-020-00205-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L, Baybayan P, Belden B, Berrios CD, Biswell RL, Buczkowicz P, Buske O, Chakraborty S, Cheung WA, Coffman KA, Cooper AM, Cross LA, Curran T, Dang TTT, Elfrink MM, Engleman KL, Fecske ED, Fieser C, Fitzgerald K, Fleming EA, Gadea RN, Gannon JL, Gelineau-Morel RN, Gibson M, Goldstein J, Grundberg E, Halpin K, Harvey BS, Heese BA, Hein W, Herd SM, Hughes SS, Ilyas M, Jacobson J, Jenkins JL, Jiang S, Johnston JJ, Keeler K, Korlach J, Kussmann J, Lambert C, Lawson C, Le Pichon JB, Leeder JS, Little VC, Louiselle DA, Lypka M, McDonald BD, Miller N, Modrcin A, Nair A, Neal SH, Oermann CM, Pacicca DM, Pawar K, Posey NL, Price N, Puckett LMB, Quezada JF, Raje N, Rowell WJ, Rush ET, Sampath V, Saunders CJ, Schwager C, Schwend RM, Shaffer E, Smail C, Soden S, Strenk ME, Sullivan BR, Sweeney BR, Tam-Williams JB, Walter AM, Welsh H, Wenger AM, Willig LK, Yan Y, Younger ST, Zhou D, Zion TN, Thiffault I, & Pastinen T (2022, Jun). Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes. Genet Med, 24(6), 1336–1348. 10.1016/j.gim.2022.02.007 [DOI] [PubMed] [Google Scholar]
  16. Costain G, Walker S, Marano M, Veenma D, Snell M, Curtis M, Luca S, Buera J, Arje D, Reuter MS, Thiruvahindrapuram B, Trost B, Sung WWL, Yuen RKC, Chitayat D, Mendoza-Londono R, Stavropoulos DJ, Scherer SW, Marshall CR, Cohn RD, Cohen E, Orkin J, Meyn MS, & Hayeems RZ (2020, Sep 1). Genome Sequencing as a Diagnostic Test in Children With Unexplained Medical Complexity. JAMA Netw Open, 3(9), e2018109. 10.1001/jamanetworkopen.2020.18109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Cretu Stancu M, van Roosmalen MJ, Renkens I, Nieboer MM, Middelkamp S, de Ligt J, Pregno G, Giachino D, Mandrile G, Espejo Valle-Inclan J, Korzelius J, de Bruijn E, Cuppen E, Talkowski ME, Marschall T, de Ridder J, & Kloosterman WP (2017, Nov 6). Mapping and phasing of structural variation in patient genomes using nanopore sequencing. Nat Commun, 8(1), 1326. 10.1038/s41467-017-01343-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Cummings BB, Marshall JL, Tukiainen T, Lek M, Donkervoort S, Foley AR, Bolduc V, Waddell LB, Sandaradura SA, O’Grady GL, Estrella E, Reddy HM, Zhao F, Weisburd B, Karczewski KJ, O’Donnell-Luria AH, Birnbaum D, Sarkozy A, Hu Y, Gonorazky H, Claeys K, Joshi H, Bournazos A, Oates EC, Ghaoui R, Davis MR, Laing NG, Topf A, Genotype-Tissue Expression C, Kang PB, Beggs AH, North KN, Straub V, Dowling JJ, Muntoni F, Clarke NF, Cooper ST, Bonnemann CG, & MacArthur DG (2017, Apr 19). Improving genetic diagnosis in Mendelian disease with transcriptome sequencing. Sci Transl Med, 9(386). 10.1126/scitranslmed.aal5209 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Dolzhenko E, Deshpande V, Schlesinger F, Krusche P, Petrovski R, Chen S, Emig-Agius D, Gross A, Narzisi G, Bowman B, Scheffler K, van Vugt J, French C, Sanchis-Juan A, Ibanez K, Tucci A, Lajoie BR, Veldink JH, Raymond FL, Taft RJ, Bentley DR, & Eberle MA (2019, Nov 1). ExpansionHunter: a sequence-graph-based tool to analyze variation in short tandem repeat regions. Bioinformatics, 35(22), 4754–4756. 10.1093/bioinformatics/btz431 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Dolzhenko E, van Vugt J, Shaw RJ, Bekritsky MA, van Blitterswijk M, Narzisi G, Ajay SS, Rajan V, Lajoie BR, Johnson NH, Kingsbury Z, Humphray SJ, Schellevis RD, Brands WJ, Baker M, Rademakers R, Kooyman M, Tazelaar GHP, van Es MA, McLaughlin R, Sproviero W, Shatunov A, Jones A, Al Khleifat A, Pittman A, Morgan S, Hardiman O, Al-Chalabi A, Shaw C, Smith B, Neo EJ, Morrison K, Shaw PJ, Reeves C, Winterkorn L, Wexler NS, Group US-VCR, Housman DE, Ng CW, Li AL, Taft RJ, van den Berg LH, Bentley DR, Veldink JH, & Eberle MA (2017, Nov). Detection of long repeat expansions from PCR-free whole-genome sequence data. Genome Res, 27(11), 1895–1903. 10.1101/gr.225672.117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Ebbert MTW, Farrugia SL, Sens JP, Jansen-West K, Gendron TF, Prudencio M, McLaughlin IJ, Bowman B, Seetin M, DeJesus-Hernandez M, Jackson J, Brown PH, Dickson DW, van Blitterswijk M, Rademakers R, Petrucelli L, & Fryer JD (2018, Aug 21). Long-read sequencing across the C9orf72 ‘GGGGCC’ repeat expansion: implications for clinical use and genetic discovery efforts in human disease. Mol Neurodegener, 13(1), 46. 10.1186/s13024-018-0274-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Eichler EE (2001, Nov). Recent duplication, domain accretion and the dynamic mutation of the human genome. Trends Genet, 17(11), 661–669. 10.1016/s0168-9525(01)02492-1 [DOI] [PubMed] [Google Scholar]
  23. Emanuel BS, & Shaikh TH (2001, Oct). Segmental duplications: an ‘expanding’ role in genomic instability and disease. Nat Rev Genet, 2(10), 791–800. 10.1038/35093500 [DOI] [PubMed] [Google Scholar]
  24. Frans G, Meert W, Van der Werff Ten Bosch J, Meyts I, Bossuyt X, Vermeesch JR, & Hestand MS (2018, Mar). Conventional and Single-Molecule Targeted Sequencing Method for Specific Variant Detection in IKBKG while Bypassing the IKBKGP1 Pseudogene. J Mol Diagn, 20(2), 195–202. 10.1016/j.jmoldx.2017.10.005 [DOI] [PubMed] [Google Scholar]
  25. Garg P, Jadhav B, Rodriguez OL, Patel N, Martin-Trujillo A, Jain M, Metsu S, Olsen H, Paten B, Ritz B, Kooy RF, Gecz J, & Sharp AJ (2020, Oct 1). A Survey of Rare Epigenetic Variation in 23,116 Human Genomes Identifies Disease-Relevant Epivariations and CGG Expansions. Am J Hum Genet, 107(4), 654–669. 10.1016/j.ajhg.2020.08.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Goenka SD, Gorzynski JE, Shafin K, Fisk DG, Pesout T, Jensen TD, Monlong J, Chang PC, Baid G, Bernstein JA, Christle JW, Dalton KP, Garalde DR, Grove ME, Guillory J, Kolesnikov A, Nattestad M, Ruzhnikov MRZ, Samadi M, Sethia A, Spiteri E, Wright CJ, Xiong K, Zhu T, Jain M, Sedlazeck FJ, Carroll A, Paten B, & Ashley EA (2022, Mar 28). Accelerated identification of disease-causing variants with ultra-rapid nanopore genome sequencing. Nat Biotechnol. 10.1038/s41587-022-01221-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Goodwin S, McPherson JD, & McCombie WR (2016, May 17). Coming of age: ten years of next-generation sequencing technologies. Nat Rev Genet, 17(6), 333–351. 10.1038/nrg.2016.49 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Google. (2022). DeepVariant. Retrieved 06/17/2022 from https://github.com/google/deepvariant
  29. Gorzynski JE, Goenka SD, Shafin K, Jensen TD, Fisk DG, Grove ME, Spiteri E, Pesout T, Monlong J, Baid G, Bernstein JA, Ceresnak S, Chang PC, Christle JW, Chubb H, Dalton KP, Dunn K, Garalde DR, Guillory J, Knowles JW, Kolesnikov A, Ma M, Moscarello T, Nattestad M, Perez M, Ruzhnikov MRZ, Samadi M, Setia A, Wright C, Wusthoff CJ, Xiong K, Zhu T, Jain M, Sedlazeck FJ, Carroll A, Paten B, & Ashley EA (2022, Feb 17). Ultrarapid Nanopore Genome Sequencing in a Critical Care Setting. N Engl J Med, 386(7), 700–702. 10.1056/NEJMc2112090 [DOI] [PubMed] [Google Scholar]
  30. Guan Q, Balciuniene J, Cao K, Fan Z, Biswas S, Wilkens A, Gallo DJ, Bedoukian E, Tarpinian J, Jayaraman P, Sarmady M, Dulik M, Santani A, Spinner N, Abou Tayoun AN, Krantz ID, Conlin LK, & Luo M (2018, Dec). AUDIOME: a tiered exome sequencing-based comprehensive gene panel for the diagnosis of heterogeneous nonsyndromic sensorineural hearing loss. Genet Med, 20(12), 1600–1608. 10.1038/gim.2018.48 [DOI] [PubMed] [Google Scholar]
  31. Hiatt SM, Lawlor JMJ, Handley LH, Ramaker RC, Rogers BB, Partridge EC, Boston LB, Williams M, Plott CB, Jenkins J, Gray DE, Holt JM, Bowling KM, Bebin EM, Grimwood J, Schmutz J, & Cooper GM (2021, Apr 8). Long-read genome sequencing for the molecular diagnosis of neurodevelopmental disorders. HGG Adv, 2(2). 10.1016/j.xhgg.2021.100023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Hoijer I, Tsai YC, Clark TA, Kotturi P, Dahl N, Stattin EL, Bondeson ML, Feuk L, Gyllensten U, & Ameur A (2018, Sep). Detailed analysis of HTT repeat elements in human blood using targeted amplification-free long-read sequencing. Hum Mutat, 39(9), 1262–1272. 10.1002/humu.23580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Ibanez K, Polke J, Hagelstrom RT, Dolzhenko E, Pasko D, Thomas ERA, Daugherty LC, Kasperaviciute D, Smith KR, Group, W. G. S. f. N. D., Deans ZC, Hill S, Fowler T, Scott RH, Hardy J, Chinnery PF, Houlden H, Rendon A, Caulfield MJ, Eberle MA, Taft RJ, Tucci A, & Genomics England Research, C. (2022, Mar). Whole genome sequencing for the diagnosis of neurological repeat expansion disorders in the UK: a retrospective diagnostic accuracy and prospective clinical validation study. Lancet Neurol, 21(3), 234–245. 10.1016/S1474-4422(21)00462-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Karimzadeh M, Ernst C, Kundaje A, & Hoffman MM (2018, Nov 16). Umap and Bismap: quantifying genome and methylome mappability. Nucleic Acids Res, 46(20), e120. 10.1093/nar/gky677 [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Kingsmore SF, Cakici JA, Clark MM, Gaughran M, Feddock M, Batalov S, Bainbridge MN, Carroll J, Caylor SA, Clarke C, Ding Y, Ellsworth K, Farnaes L, Hildreth A, Hobbs C, James K, Kint CI, Lenberg J, Nahas S, Prince L, Reyes I, Salz L, Sanford E, Schols P, Sweeney N, Tokita M, Veeraraghavan N, Watkins K, Wigby K, Wong T, Chowdhury S, Wright MS, Dimmock D, & Investigators R (2019, Oct 3). A Randomized, Controlled Trial of the Analytic and Diagnostic Performance of Singleton and Trio, Rapid Genome and Exome Sequencing in Ill Infants. Am J Hum Genet, 105(4), 719–733. 10.1016/j.ajhg.2019.08.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Kosugi S, Momozawa Y, Liu X, Terao C, Kubo M, & Kamatani Y (2019, Jun 3). Comprehensive evaluation of structural variation detection algorithms for whole genome sequencing. Genome Biol, 20(1), 117. 10.1186/s13059-019-1720-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Lee H, Huang AY, Wang LK, Yoon AJ, Renteria G, Eskin A, Signer RH, Dorrani N, Nieves-Rodriguez S, Wan J, Douine ED, Woods JD, Dell’Angelica EC, Fogel BL, Martin MG, Butte MJ, Parker NH, Wang RT, Shieh PB, Wong DA, Gallant N, Singh KE, Tavyev Asher YJ, Sinsheimer JS, Krakow D, Loo SK, Allard P, Papp JC, Undiagnosed Diseases N, Palmer CGS, Martinez-Agosto JA, & Nelson SF (2020, Mar). Diagnostic utility of transcriptome sequencing for rare Mendelian diseases. Genet Med, 22(3), 490–499. 10.1038/s41436-019-0672-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Loomis EW, Eid JS, Peluso P, Yin J, Hickey L, Rank D, McCalmon S, Hagerman RJ, Tassone F, & Hagerman PJ (2013, Jan). Sequencing the unsequenceable: expanded CGG-repeat alleles of the fragile X gene. Genome Res, 23(1), 121–128. 10.1101/gr.141705.112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Luo H, Zhao Q, Wei W, Zheng L, Yi S, Li G, Wang W, Sheng H, Pu H, Mo H, Zuo Z, Liu Z, Li C, Xie C, Zeng Z, Li W, Hao X, Liu Y, Cao S, Liu W, Gibson S, Zhang K, Xu G, & Xu RH (2020, Jan 1). Circulating tumor DNA methylation profiles enable early diagnosis, prognosis prediction, and screening for colorectal cancer. Sci Transl Med, 12(524). 10.1126/scitranslmed.aax7533 [DOI] [PubMed] [Google Scholar]
  40. Maestri S, Maturo MG, Cosentino E, Marcolungo L, Iadarola B, Fortunati E, Rossato M, & Delledonne M (2020, Dec 1). A Long-Read Sequencing Approach for Direct Haplotype Phasing in Clinical Settings. Int J Mol Sci, 21(23). 10.3390/ijms21239177 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Mandelker D, Schmidt RJ, Ankala A, McDonald Gibson K, Bowser M, Sharma H, Duffy E, Hegde M, Santani A, Lebo M, & Funke B (2016, Dec). Navigating highly homologous genes in a molecular diagnostic setting: a resource for clinical next-generation sequencing. Genet Med, 18(12), 1282–1289. 10.1038/gim.2016.58 [DOI] [PubMed] [Google Scholar]
  42. Manickam K, McClain MR, Demmer LA, Biswas S, Kearney HM, Malinowski J, Massingham LJ, Miller D, Yu TW, Hisama FM, & Directors A. B. o. (2021, Nov). Exome and genome sequencing for pediatric patients with congenital anomalies or intellectual disability: an evidence-based clinical guideline of the American College of Medical Genetics and Genomics (ACMG). Genet Med, 23(11), 2029–2037. 10.1038/s41436-021-01242-6 [DOI] [PubMed] [Google Scholar]
  43. McFarland KN, Liu J, Landrian I, Godiska R, Shanker S, Yu F, Farmerie WG, & Ashizawa T (2015). SMRT Sequencing of Long Tandem Nucleotide Repeats in SCA10 Reveals Unique Insight of Repeat Expansion Structure. PLoS One, 10(8), e0135906. 10.1371/journal.pone.0135906 [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Melas M, Kautto EA, Franklin SJ, Mori M, McBride KL, Mosher TM, Pfau RB, Hernandez-Gonzalez ME, McGrath SD, Magrini VJ, White P, Samora JB, Koboldt DC, & Wilson RK (2022, Feb). Long-read whole genome sequencing reveals HOXD13 alterations in synpolydactyly. Hum Mutat, 43(2), 189–199. 10.1002/humu.24304 [DOI] [PubMed] [Google Scholar]
  45. Mellis R, Oprych K, Scotchman E, Hill M, & Chitty LS (2022, May). Diagnostic yield of exome sequencing for prenatal diagnosis of fetal structural anomalies: A systematic review and meta-analysis. Prenat Diagn, 42(6), 662–685. 10.1002/pd.6115 [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Menelaou A, & Marchini J (2013, Jan 1). Genotype calling and phasing using next-generation sequencing reads and a haplotype scaffold. Bioinformatics, 29(1), 84–91. 10.1093/bioinformatics/bts632 [DOI] [PubMed] [Google Scholar]
  47. Miao H, Zhou J, Yang Q, Liang F, Wang D, Ma N, Gao B, Du J, Lin G, Wang K, & Zhang Q (2018). Long-read sequencing identified a causal structural variant in an exome-negative case and enabled preimplantation genetic diagnosis. Hereditas, 155, 32. 10.1186/s41065-018-0069-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Miller DE, Lee L, Galey M, Kandhaya-Pillai R, Tischkowitz M, Amalnath D, Vithlani A, Yokote K, Kato H, Maezawa Y, Takada-Watanabe A, Takemoto M, Martin GM, Eichler EE, Hisama FM, & Oshima J (2022, May 9). Targeted long-read sequencing identifies missing pathogenic variants in unsolved Werner syndrome cases. J Med Genet. 10.1136/jmedgenet-2022-108485 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Miller DE, Sulovari A, Wang T, Loucks H, Hoekzema K, Munson KM, Lewis AP, Fuerte EPA, Paschal CR, Walsh T, Thies J, Bennett JT, Glass I, Dipple KM, Patterson K, Bonkowski ES, Nelson Z, Squire A, Sikes M, Beckman E, Bennett RL, Earl D, Lee W, Allikmets R, Perlman SJ, Chow P, Hing AV, Wenger TL, Adam MP, Sun A, Lam C, Chang I, Zou X, Austin SL, Huggins E, Safi A, Iyengar AK, Reddy TE, Majoros WH, Allen AS, Crawford GE, Kishnani PS, University of Washington Center for Mendelian, G., King MC, Cherry T, Chong JX, Bamshad MJ, Nickerson DA, Mefford HC, Doherty D, & Eichler EE (2021, Aug 5). Targeted long-read sequencing identifies missing disease-causing variation. Am J Hum Genet, 108(8), 1436–1449. 10.1016/j.ajhg.2021.06.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Miller DT, Adam MP, Aradhya S, Biesecker LG, Brothman AR, Carter NP, Church DM, Crolla JA, Eichler EE, Epstein CJ, Faucett WA, Feuk L, Friedman JM, Hamosh A, Jackson L, Kaminsky EB, Kok K, Krantz ID, Kuhn RM, Lee C, Ostell JM, Rosenberg C, Scherer SW, Spinner NB, Stavropoulos DJ, Tepperberg JH, Thorland EC, Vermeesch JR, Waggoner DJ, Watson MS, Martin CL, & Ledbetter DH (2010, May 14). Consensus statement: chromosomal microarray is a first-tier clinical diagnostic test for individuals with developmental disabilities or congenital anomalies. Am J Hum Genet, 86(5), 749–764. 10.1016/j.ajhg.2010.04.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Patterson M, Marschall T, Pisanti N, van Iersel L, Stougie L, Klau GW, & Schonhuth A (2015, Jun). WhatsHap: Weighted Haplotype Assembly for Future-Generation Sequencing Reads. J Comput Biol, 22(6), 498–509. 10.1089/cmb.2014.0157 [DOI] [PubMed] [Google Scholar]
  52. Payne A, Holmes N, Rakyan V, & Loose M (2019, Jul 1). BulkVis: a graphical viewer for Oxford nanopore bulk FAST5 files. Bioinformatics, 35(13), 2193–2198. 10.1093/bioinformatics/bty841 [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Quick J (2018). Ultra-long read sequencing protocol for RAD004 V.3 Retrieved 08/27/2022 from 10.17504/protocols.io.mrxc57n [DOI]
  54. Rajagopalan R, Murrell JR, Luo M, & Conlin LK (2020, Jan 30). A highly sensitive and specific workflow for detecting rare copy-number variants from exome sequencing data. Genome Med, 12(1), 14. 10.1186/s13073-020-0712-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Rentas S, Rathi KS, Kaur M, Raman P, Krantz ID, Sarmady M, & Tayoun AA (2020, May). Diagnosing Cornelia de Lange syndrome and related neurodevelopmental disorders using RNA sequencing. Genet Med, 22(5), 927–936. 10.1038/s41436-019-0741-5 [DOI] [PubMed] [Google Scholar]
  56. Sanchez-Luquez KY, Carpena MX, Karam SM, & Tovo-Rodrigues L (2022, Jul 27). The contribution of whole-exome sequencing to intellectual disability diagnosis and knowledge of underlying molecular mechanisms: A systematic review and meta-analysis. Mutat Res Rev Mutat Res, 790, 108428. 10.1016/j.mrrev.2022.108428 [DOI] [PubMed] [Google Scholar]
  57. Schule B, McFarland KN, Lee K, Tsai YC, Nguyen KD, Sun C, Liu M, Byrne C, Gopi R, Huang N, Langston JW, Clark T, Gil FJJ, & Ashizawa T (2017). Parkinson’s disease associated with pure ATXN10 repeat expansion. NPJ Parkinsons Dis, 3, 27. 10.1038/s41531-017-0029-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Shafin K, Pesout T, Chang PC, Nattestad M, Kolesnikov A, Goel S, Baid G, Kolmogorov M, Eizenga JM, Miga KH, Carnevali P, Jain M, Carroll A, & Paten B (2021, Nov). Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat Methods, 18(11), 1322–1332. 10.1038/s41592-021-01299-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Sidaway P (2020, Apr). Diagnosis using methylation. Nat Rev Clin Oncol, 17(4), 196. 10.1038/s41571-020-0334-x [DOI] [PubMed] [Google Scholar]
  60. Srivastava S, Love-Nichols JA, Dies KA, Ledbetter DH, Martin CL, Chung WK, Firth HV, Frazier T, Hansen RL, Prock L, Brunner H, Hoang N, Scherer SW, Sahin M, Miller DT, & Group, N. D. D. E. S. R. W. (2020, Oct). Correction: Meta-analysis and multidisciplinary consensus statement: exome sequencing is a first-tier clinical diagnostic test for individuals with neurodevelopmental disorders. Genet Med, 22(10), 1731–1732. 10.1038/s41436-020-0913-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Stevanovski I, Chintalaphani SR, Gamaarachchi H, Ferguson JM, Pineda SS, Scriba CK, Tchan M, Fung V, Ng K, Cortese A, Houlden H, Dobson-Stone C, Fitzpatrick L, Halliday G, Ravenscroft G, Davis MR, Laing NG, Fellner A, Kennerson M, Kumar KR, & Deveson IW (2022, Mar 4). Comprehensive genetic diagnosis of tandem repeat expansion disorders with programmable targeted nanopore sequencing. Sci Adv, 8(9), eabm5386. 10.1126/sciadv.abm5386 [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Svrzikapa N, Longo KA, Prasad N, Boyanapalli R, Brown JM, Dorset D, Yourstone S, Powers J, Levy SE, Morris AJ, Vargeese C, & Goyal J (2020, Dec 11). Investigational Assay for Haplotype Phasing of the Huntingtin Gene. Mol Ther Methods Clin Dev, 19, 162–173. 10.1016/j.omtm.2020.09.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Technologies ON (2021). Benchmarking nanopore methylation analysis by comparison to publicly available bisulphite datasets. Retrieved 06/17/2022 from https://nanoporetech.com/sites/default/files/s3/Methylation_v2_18052021.pdf
  64. Technologies ON (2022). Q20+ Chemistry for single molecule accuracy of 99% and higher. Retrieved 06/17/2022 from https://nanoporetech.com/q20plus-chemistry
  65. Toffoli M, Chen X, Sedlazeck FJ, Lee C-Y, Mullin S, Higgins A, Koletsi S, Garcia-Segura ME, Sammler E, Scholz SW, Schapira AH, Eberle MA, & Proukakis C (2021). Comprehensive analysis of GBA using a novel algorithm for Illumina whole-genome sequence data or targeted Nanopore sequencing. medRxiv, 2021.2011.2012.21266253. 10.1101/2021.11.12.21266253 [DOI] [Google Scholar]
  66. Wang Y, Zhao Y, Bollas A, Wang Y, & Au KF (2021, Nov). Nanopore sequencing technology, bioinformatics and applications. Nat Biotechnol, 39(11), 1348–1365. 10.1038/s41587-021-01108-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Wenger AM, Peluso P, Rowell WJ, Chang PC, Hall RJ, Concepcion GT, Ebler J, Fungtammasan A, Kolesnikov A, Olson ND, Topfer A, Alonge M, Mahmoud M, Qian Y, Chin CS, Phillippy AM, Schatz MC, Myers G, DePristo MA, Ruan J, Marschall T, Sedlazeck FJ, Zook JM, Li H, Koren S, Carroll A, Rank DR, & Hunkapiller MW (2019, Oct). Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol, 37(10), 1155–1162. 10.1038/s41587-019-0217-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Wieben ED, Aleff RA, Basu S, Sarangi V, Bowman B, McLaughlin IJ, Mills JR, Butz ML, Highsmith EW, Ida CM, Ekholm JM, Baratz KH, & Fautsch MP (2019). Amplification-free long-read sequencing of TCF4 expanded trinucleotide repeats in Fuchs Endothelial Corneal Dystrophy. PLoS One, 14(7), e0219446. 10.1371/journal.pone.0219446 [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Wu AC, McMahon P, & Lu C (2020, Sep 1). Ending the Diagnostic Odyssey-Is Whole-Genome Sequencing the Answer? JAMA Pediatr, 174(9), 821–822. 10.1001/jamapediatrics.2020.1522 [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Xu L, Mao A, Liu H, Gui B, Choy KW, Huang H, Yu Q, Zhang X, Chen M, Lin N, Chen L, Han J, Wang Y, Zhang M, Li X, He D, Lin Y, Zhang J, Cram DS, & Cao H (2020, Aug). Long-Molecule Sequencing: A New Approach for Identification of Clinically Significant DNA Variants in alpha-Thalassemia and beta-Thalassemia Carriers. J Mol Diagn, 22(8), 1087–1095. 10.1016/j.jmoldx.2020.05.004 [DOI] [PubMed] [Google Scholar]
  71. Yepez VA, Gusic M, Kopajtich R, Mertes C, Smith NH, Alston CL, Ban R, Beblo S, Berutti R, Blessing H, Ciara E, Distelmaier F, Freisinger P, Haberle J, Hayflick SJ, Hempel M, Itkis YS, Kishita Y, Klopstock T, Krylova TD, Lamperti C, Lenz D, Makowski C, Mosegaard S, Muller MF, Munoz-Pujol G, Nadel A, Ohtake A, Okazaki Y, Procopio E, Schwarzmayr T, Smet J, Staufner C, Stenton SL, Strom TM, Terrile C, Tort F, Van Coster R, Vanlander A, Wagner M, Xu M, Fang F, Ghezzi D, Mayr JA, Piekutowska-Abramczuk D, Ribes A, Rotig A, Taylor RW, Wortmann SB, Murayama K, Meitinger T, Gagneur J, & Prokisch H (2022, Apr 5). Clinical implementation of RNA sequencing for Mendelian disease diagnostics. Genome Med, 14(1), 38. 10.1186/s13073-022-01019-9 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.

RESOURCES