Abstract
Rare diseases are collectively common, affecting approximately 1 in 20 individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in next-generation sequencing, development of new computational and functional genomics approaches to prioritize genes and variants and increased global sharing of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Furthermore, all data generated, currently representing over 7,500 individuals from over 3,000 families, are rapidly made available to researchers worldwide through the Analysis, Visualization and Informatics Lab-space (AnVIL) to catalyse global efforts to develop approaches for genetic diagnoses in rare diseases. Most of these families have undergone previous clinical genetic testing but remained unsolved, with most being exome-negative. Here we describe the collaborative research framework, datasets and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
The past decade has seen rapid progress in clinical genetics due to increased discovery of genes and variants involved in Mendelian diseases and ongoing advances in sequencing, variant analysis and data sharing1–5. Despite this progress, most individuals who undergo clinical genetic testing for a suspected Mendelian condition remain undiagnosed6–9. For example, in the National Human Genome Research Institute (NHGRI) Centers for Mendelian Genomics (CMG)10–12, while over 3,800 genes were implicated in Mendelian disease, only about 11,000 out of over 28,000 families received a confirmed or potential molecular diagnosis1. Thus, considerable challenges remain to increase the molecular diagnostic yield and explain currently unsolved rare genetic disorders (Box 1).
Box 1. Challenges in diagnosing rare genetic diseases.
(1). Undiscovered disease genes.
The pathogenic variant(s) may be located in a gene that is yet to be implicated in disease. Until now, over 5,000 protein-coding genes have been implicated in at least one disease, but it is estimated that still 10,000+ disease–gene relationships are undiscovered in just the remaining protein-coding genes128.
(2). Understanding non-coding variation.
Candidate pathogenic variant(s) may be located in the non-coding genome, where the mechanisms for how a variant manifests a clinical phenotype are not well understood.
(3). Current technological limitations.
The variant may be in a region that is unattainable to ascertain or difficult to detect from solely short-read, exome or genome sequencing such as long repeats, inversions and CGRs. The variant may be detectable but bioinformatic algorithms may struggle to call the variant correctly such as multi-nucleotide and mosaic variants. The variant may be detectable and called correctly but asserting its functionality or pathogenicity may require unavailable, orthogonal evidence.
(4). Complex inheritance patterns.
MPV, oligogenic, polygenic, complex compound heterozygosity, variable expressivity, incomplete penetrance, allelic heterogeneity, imprinting, epimutations, maternal effect and/or mosaicism may also be confounding a diagnosis and necessitate a broader approach to understanding the disease mechanism.
(5). Lag time in curation.
A gene–disease relationship may be published or submitted to a genetic database but has yet to be reviewed and incorporated into clinical testing. This is compounded by rapidly increasing numbers of VUSs.
(6). Costs and implementation.
The costs associated with both current and new genomic technologies, as well as the expertise required for their implementation in clinical laboratories for scaled testing, present obstacles to widespread adoption in both research and clinical settings. The need for specialized training, regulatory compliance and expensive infrastructure further limits accessibility, particularly in resource-limited settings.
(7). n = 1.
Many candidate variants and genes are n = 1 regardless of best data-sharing practices.
(8). Limitations in experimentation.
Owing to the nature of novel discovery, there is not always a functional assay available to provide orthogonal evidence for or against a candidate variant or gene. Often if a candidate does meet the inclusion criteria for an existing assay, the molecular phenotype measured in the assay may not match the potential mechanism of disease or fully recapitulate the pathophysiological impact, resulting in ambiguous results or an incorrect prediction of pathogenicity.
(9). Bias.
Databases predominantly capture genetic information from individuals of European genetic ancestry potentially propagating biases in tools and reference data for variant classification for individuals of non-European genetic ancestry112.
(10). Contribution of each challenge.
The relative contribution of each of these challenges to the overall diagnostic gap is difficult to quantify and based mostly on hypotheses, process of elimination and extrapolation. Understanding which challenges have the greatest impact—and in which contexts—remains key to prioritizing solutions effectively.
(11). High burden of proof.
Newer genomic technologies may offer advantages over short-read DNA sequencing, but effective and widespread use of new technologies requires clear guidance and broad demonstration of scaled efficacy.
In 2021, at the end of the CMG, the NHGRI launched GREGoR with five primary research sites and a data-coordinating centre to accelerate rare disease genetic research by harnessing the latest advances in sequencing including and especially genome sequencing and multi-omics; evaluating and prioritizing the use of functional genomics and new computational strategies including recent advances in artificial intelligence; translating advances into routine clinical testing; advancing data sharing to foster a quorum of evidence for discovery; and collaborating worldwide to continue discovery and reporting of genetic aetiologies for Mendelian diseases (https://gregorconsortium.org/data) (Fig. 1). Compared with other rare disease consortia worldwide, GREGoR is unique in its mission to generate and rapidly disseminate internationally a variety of genomic data for rare disease families and evaluate, validate and scale emerging technologies aimed at achieving molecular diagnoses in cases that remain unresolved despite conventional clinical genetic testing.
Fig. 1 |. Overview of GREGoR.
Strategic framework of the GREGoR consortium for accelerating genomics in rare-disease research, highlighting cross-cutting themes, systematic data generation, computational innovations and end points of success.
Evaluating emerging methods
Squeezing the exome
The most impactful approach to date for diagnosing rare diseases has been exome sequencing and periodic reanalysis of the protein-coding sequences in the human genome13–17. GREGoR has led or contributed to 83 papers studying molecular diagnoses in 365 genes with more than a third being novel disease gene discoveries or phenotypic expansions (Supplementary Table 1) and provided a variety of automated pipelines for large cohort, exome and genome reanalysis, which include phenotypic integration14–16. Reanalysis success is largely driven by new disease gene discoveries and phenotypic expansions since the original analysis15; however, new tools focused on reanalysis of well-known disease genes and loci have continued to yield diagnostic successes18.
Furthermore, GREGoR has developed new computational approaches to further increase the molecular diagnostic yield from exomes. For example, difficulty in phasing short-read sequencing can confound diagnoses of pathogenic compound heterozygous variants for recessive diseases. To overcome this, GREGoR contributed to a highly accurate method for inferring phase and has calculated and released all pairwise phasing estimates and usage guidance for rare coding variants in exomes occurring in the same gene through the Genome Aggregation Database (gnomAD)19,20. GREGoR researchers have developed tools to identify and implicate hundreds of pathogenic structural variant diagnoses from existing unsolved exomes21–25. Combined, GREGoR’s continuous efforts to extract diagnoses from existing exomes demonstrate potential for continued innovation in genomic reanalyses.
Short-read genome sequencing
GREGoR has published a framework for genetic testing when panel or exome sequencing is inconclusive26. The next step is typically short-read genome sequencing (srGS). However, most molecular diagnoses deduced by srGS are found in protein-coding genes, suggesting that they could potentially be detected by exome sequencing. By contrast, srGS offers the opportunity to profile a wide spectrum of variants within and beyond the exome, especially for structural variants, such as large deletions, duplications, insertions, inversions, translocations and complex structural variants that involve a combination of two or more variant classes27–30. To evaluate the relative utility of srGS, a large-scale study of 822 families by GREGoR researchers reported 218 patients who received a diagnosis through srGS, where 72% of variants should have been detectable by exome sequencing31. The remaining 28% of cases were explained by variants that are not readily accessible on exomes such as tandem repeat expansions, deep intronic variants, structural variants and variants in difficult-to-sequence coding regions. Overall, srGS resulted in a greater than 8% increase in diagnostic yield compared with just exome sequencing and underscored the growing support for using srGS as a first-tier test.
Large-scale srGS has also provided the opportunity to evaluate selective constraint against non-coding loci. Constraint metrics, such as LOEUF32 quantify a gene’s intolerance to loss-of-function mutations and have already proven valuable tools in gene discovery and variant classification. Identifying non-coding regions intolerant of mutation has historically proven to be more difficult due to limited sample sizes, lack of precise variant effect models and heterogeneous mutation rates influenced by broader genomic features33. A recent effort leveraged over 76,000 individuals with srGS from gnomAD and computed constraint genome-wide34. Moreover, making use of large studies of structural and copy number variants has improved measures of dosage sensitivity, including identifying triplosensitive genes28. As the sample size of reference populations increases, these metrics will be further refined, including increased resolution for the coding regions most intolerant to missense mutations35,36 and non-coding regions. GREGoR is translating knowledge of these constrained genes and regions to rare disease cases, with a focus on identifying clinically significant non-coding variants.
GREGoR is developing visualization tools for structural and copy-number variants25 by repurposing read-depth from srGS to mimic single-nucleotide polymorphism arrays to achieve resolution as low as 1 kb—beyond the 5 kb limit of the current standard using array comparative genomic hybridization. Thus, srGS could potentially serve as a cost-effective, first-line, unifying assay by simultaneously replacing both arrays and exomes and enable more accurate, nucleotide-resolution breakpoints of structural variants, which have historically been critical in deciphering mechanisms of genomic rearrangements. Most breakpoints including published structural variants, especially in highly repetitive or hypermutable regions of the genome, lack validation at the nucleotide resolution, which is relevant for genomic assembly and hypothesis-driven inferences of structural variant impact on gene expression. Furthermore, GREGoR investigators are showing local sequences surrounding candidate and pathogenic variants can offer insights into secondary structure mutagenesis and other mechanisms of genomic disorders37–39. GREGoR’s work to identify diagnoses from srGS emphasizes unrealized potential of both primary analysis and reanalysis of srGS for rare disease discoveries.
Long-read sequencing
Long-read sequencing has opened new diagnostic opportunities in rare diseases. GREGoR and others have demonstrated that targeted long-read sequencing can reveal variants, particularly structural variants spanning repetitive sequences, in both known and novel disease genes that are missed or difficult to detect by short-read sequencing40–44. Targeted sequencing panels can evolve beyond coding regions to include untranslated regions, promoters, intronic, intergenic and large expanses of non-coding regions around known disease genes. For example, GREGoR in collaboration with Twist developed the Twist Alliance Dark Genes Panel to produce phased variants across 389 medically relevant and complex autosomal genes, where short-read sequencing often fails45.
A key focus for GREGoR is comparing the molecular diagnostic yield and cost-effectiveness of short-read versus long-read sequencing. In multiple studies, long-read sequencing uncovered novel candidate variants and genes missed by short-read sequencing, including de novo, compound heterozygous, structural and epigenetic variants44,46,47. In comparison to short-read sequencing, long-read sequencing offers advantages such as better phasing, improved understanding of haplotype blocks and methylation analysis. GREGoR researchers have been leveraging these advantages by developing tools48 to use methylation data for phasing and investigating the diagnostic yield improvements from genome-wide DNA methylation arrays in relation to long-read sequencing49. Moreover, GREGoR is developing tools for improved annotation for kilobase and megabase scale variants especially using long-read sequencing with a focus on mosaic structural variants50,51. Furthermore, innovative computational tools in long-read sequencing variant calling and analysis51,52 are being developed, including de novo variant callers and long-read pipelines for mitochondrial variant calling.
Further long-read sequencing can offer multi-omic insights beyond traditional DNA sequencing and is increasingly being applied to find and understand molecular diagnoses as well as mechanisms in rare diseases. One such technology is fiber-seq, which uses long-read sequencing to simultaneously evaluate the primary DNA sequence with a nucleotide-resolution view of the surrounding chromatin architecture53,54. GREGoR is developing unifying assays looking to simultaneously assay the genome, methylome, epigenome and transcriptome to identify and understand mechanisms of rare diseases55. Such unifying assays help to explain previously elusive variants56 that may have been visible on exome or short-read genome sequencing but may not have been nominated as candidate variants. The comparison and contrast of multiple layers of -omics in a single, unifying assay mechanistically implicates the pathogenicity of these candidates, especially for non-coding variants.
Multi-omics
In rare disease diagnosis, methylome and transcriptome data have broadly demonstrated their utility by identification of outlier events in methylation57, splicing or gene expression implicating pathogenic variants. GREGoR has focused on complementing hundreds of cases with methylation and transcriptome data to facilitate development of standards and new computational methods. For example, recent activities in GREGoR have demonstrated how combined transcriptome and long-read genome analyses can aid in prioritizing structural variants when allele frequency information is limited58.
A focus of GREGoR has been creating data that enable evaluation and prioritization of multi-omic assays for rare disease diagnosis. Currently, the use of multi-omics is predominantly limited to research and limited information exists to suggest which post-genome, -omics assay would yield the most useful information. To address this challenge, GREGoR has been generating a squared-off matrix for a subset of families, for whom long-read genome, methylation, chromatin-accessibility, transcriptome, proteome and metabolome data are being collected. Complementing these data, GREGoR has been supporting development and integration of reference -omics data from the Common Fund Data Ecosystem to advance outlier detection for various -omics assays by integrating larger control datasets. Multiple efforts in GREGoR are facilitating more routine use of these data such as updates to the seqr platform to expand intake of multi-omics data types to enable routine linking of outliers in these data to underlying genomic variation59.
Reframing rare disease analysis
Reference genomes
Adoption of new reference genomes has lagged in clinical settings. Despite the publication of the GRCh38 human reference genome over a decade ago, many clinical labs to date still use GRCh37 (ref. 60). Part of the entrenchment of GRCh37 was the lag in necessary infrastructure development to support allele frequencies, in silico scores, bioinformatic tools and clinical databases on GRCh38. Thus, to date, the vast majority of known clinical disease genes and phenotypic expansions were discovered using GRCh37. GREGoR has shown the reference genome alone impacts variant calling in around 1% of the exome, with 206 genes enriched in discordant calls, including 8 known disease genes61. GREGoR showed that these discrepancies were more pronounced at the RNA level, with 1,492 genes demonstrating reference-dependent quantification and 3,377 genes exhibiting reference-exclusive expression, affecting 512 known disease genes62. GREGoR investigators have focused on fixing the GRCh38 reference63, benchmarking medically relevant genes for both GRCh37 and GRCh3864, and resolving pathogenic inversions in reference genome gaps using the telomere-to-telomere (T2T) reference genome47.
An important question for the field focuses on whether there will be development of flexible pipelines and tools capable of using the newest references, such as T2T and the pangenome65. GREGoR, in collaboration with Illumina, has benchmarked the DRAGEN pipeline, which uses graph-based alignment among many other novel features for variant calling in srGS66. Looking ahead, GREGoR is collaborating with the Human Pangenome Reference Consortium (HPRC) through methods development like the Pangenome Research Tool Kit to demonstrate accurate variant calling in regions of the genome previously too complex for accurate variant calling. Further, GREGoR investigators are using pangenome approaches to understand complex tandem repeats in known disease genes67 and exploring the infrastructure necessary for widespread adoption of the pangenome in clinical settings.
The hardest molecular diagnoses
Over time, consistent themes have emerged in the life cycle of disease gene discovery. Typically, the first and easiest candidate variants implicated are de novo and/or predicted loss-of-function (pLOF) variants that segregate with the phenotype in a pedigree. These two variant classes have logical frameworks supporting their putative mechanisms. A de novo variant, found in an affected proband but not in the unaffected parents, significantly increases the probability of being causative, especially when other de novo cases show the same clinical phenotypes68,69. pLOF variants such as nonsense single-nucleotide variants, frameshifting insertions or deletions, and splice-site altering variants imply a null effect, because many of these pLOF variants are expected to be caught by the mRNA surveillance mechanism known as nonsense-mediated decay (NMD). NMD is a highly sensitive mechanism that is present in all tissues and destroys faulty transcripts with premature stops with high efficiency and fidelity70,71. This same mechanistic reliability explains why pLOF variants are often the first variant type to be implicated in a novel gene-to-disease discovery, because clinical geneticists can reliably infer the presence of a pLOF variant will putatively lead to destruction of the faulty RNA transcript by NMD and no protein production, which is in alignment with a potential pathogenic mechanism of loss of function. Subsequently, other single-nucleotide and structural variants in the same gene or region are often implicated after a quorum of cases is established, although different types or locations of variation in the same gene can result in distinct clinical phenotypes. Furthermore, once one gene has been implicated, it serves as a seed for other genes in the same protein complex72, protein pathway38 or gene family73,74 to be implicated in the same or similar clinical phenotypes.
Currently, Online Mendelian Inheritance in Man (OMIM) and the Gene Curation Coalition (GenCC) have documented over 4,500 genes as being implicated in at least one Mendelian condition75,76. Most of these discoveries were achieved through sequencing and interpretation of primary DNA variation contextualized by a proband’s phenotypes without requiring additional -omics or integrative analyses. While many thousands more disease genes can still be discovered using these same established gene-discovery principles, GREGoR is pursuing cases that were unsolved by standard clinical genetic testing and are hypothesized to have among the rarest and most difficult to detect or interpret pathogenic variation, which often requires integration of multiple -omics to (1) discover a candidate variant refractory to traditional sequencing methods; (2) provide proper context to interpret a variant that may have been seen in the primary DNA sequence but for which there was not enough understanding to nominate the variant as a candidate; or (3) provide orthogonal validation of a candidate discovered in the primary DNA sequence but for which the interpretation was speculative at best. In Box 2, we discuss lessons learned during GREGoR, and below we discuss ten of the rarest and hardest molecular diagnoses pursued by GREGoR investigators:
Box 2. Lessons learned from GREGoR.
(1). Custom rare disease data model.
GREGoR has developed a data model emphasizing rare-disease research essentials such as accessibility, consent consistency and transparency, and use of accepted ontologies and common standards. Every variant in the GREGoR joint callset is machine-readable with both unique Global Alliance for Genomics and Health (GA4GH) Variant Representation Specification (VRS) IDs129 and ClinGen allele IDs130. The data model accommodates variants and output files from a wide variety of genomic, multi-omic, phenotypic and molecular data types, and is modularly designed to support the integration of future data types. GREGoR is also working with the GA4GH to contribute to global standards for the collection and sharing of rare disease data, including the transfer of rare disease phenotype data from electronic health records.
(2). Deep phenotyping and building a truth set.
Studying phenotypic heterogeneity in the context of genetic heterogeneity is critical to solving unsolved Mendelian disease. Assignment of Human Phenotype Ontology (HPO)131 terms is a mandatory requirement for GREGoR data collection to allow end users to link all possible genotypes to all possible phenotypes. Moreover, GREGoR is developing algorithms to quantitatively dissect genotypic heterogeneity and phenotypic heterogeneity especially for MPV39,88,132–137. Furthermore, GREGoR is building large language models for optimal phenotypic extraction from electronic health records138. GREGoR hopes to fulfil the unmet need for a truth set of deeply phenotyped rare disease cases linked to well described multimodal genotype data whereby solved cases can be used as positive controls for benchmarking new tools and focused challenges can be created to examine exome-negative, unsolved cases.
(3). Building reference databases.
GREGoR is using long-read sequencing and optical mapping from diverse individuals to create benchmarks and a database of structural variants for filtering and prioritizing candidates. Specifically, GREGoR researchers started the 1000 Genomes Project Oxford Nanopore Technologies (ONT) Sequencing Consortium139 and recently released the first 100 samples of long-read data from diverse populations. Furthermore, GREGoR investigators have established a population reference of tandem repeat expansions across ancestries from over 330,000 short-read genomes in TOPMed, UK Biobank, Estonian Biobank, 1000 Genome Project and All of Us, for interpreting repeat-driven genetic diseases101. Also GREGoR investigators have curated over 1.1 billion unique genomic variants from over 247,000 short-read genomes in All of Us v.7.1 for population allele frequency annotation and seamless integration into clinical workflows through the mass annotators of Ensembl variant effect predictor (VEP)140, dbNSFP141, Annovar142 and OpenCRAVAT143.
(4). Matchmaking beyond genes.
GREGoR is leveraging federated variant-level matchmaking with GA4GH standards through VariantMatcher144, MyGene2, Geno2MP and seqr59, with the hope of accelerating disease-causing variant discovery beyond the exome. Moreover, many clinical laboratories have undiagnosed cases potentially explained by novel gene discoveries or phenotypic expansions. GREGoR has actively engaged with clinical labs to study effective strategies for exome and genome analysis without overburdening variant analysts145. GREGoR has released recommendations for clinical labs to report variants in novel candidate genes and support follow-up investigations and data sharing, enabling broad discoveries and patient diagnoses146.
Non-coding variants occur in genomic regions that do not code for proteins, such as non-coding RNAs, promoters, enhancers and untranslated regions. Pathogenic non-coding variants can disrupt mechanisms such as gene regulation, splicing, expression, translation or RNA stability. Genome sequencing is crucial for detecting variants in non-coding regions, but techniques like chromatin immunoprecipitation with sequencing (ChIP–seq), assay for transposase-accessible chromatin with high-throughput sequencing (ATAC-seq), RNA sequencing (RNA-seq), fiber-seq and massively parallel reporter assays (MPRAs) can help to identify and provide mechanistic explanation for potential pathogenic non-coding variation. For example, non-coding variants such as deep intronic variants may be involved in alterations in splicing, whereby exons or introns may be incorrectly skipped or included during mRNA processing leading to changes in the final transcript potentially leading to clinical phenotypes77. While these variants are often visible in primary DNA sequencing, their interpretation is typically speculative without orthogonal validation such as reverse transcription quantitative PCR, RNA-seq or mini-gene splicing assays to validate whether or not aberrant splicing occurred by understanding the sequence of the processed mRNA. Similarly, long-read RNA-seq is particularly useful for detecting full-length transcripts and interpreting complex splicing patterns. Work by GREGoR and others has shown that non-coding variants can be the ‘missing variant’ in trans with a pathogenic coding variant for recessive rare diseases78,79. GREGoR is pursuing generalizable methods for understanding non-coding variant effects at scale80,81.
Non-coding genes produce molecules that perform potential regulatory, structural or catalytic roles rather than encode proteins. These include rRNA, tRNA, microRNA, long non-coding RNA, small nuclear RNA and more, and perturbations in non-coding genes may cause rare genetic diseases. A flagship example is perturbations in RNU4–2, in which cases from multiple GREGoR sites contributed to rapid discovery and progress to publication, establishing perturbations in RNU4–2 as one of the most commonly mutated causes of neurodevelopmental disorders82,83. Even though RNU4–2 is a non-coding RNA, it’s discovery timeline is analogous to the typical gene discovery life cycle. The main cases found were de novo insertions at the same site in a large cohort with phenotype-matching cases. RNU4–2 has now served as a seed for more RNU4 minor spliceosome genes being implicated in neurodevelopmental disorders84,85. Another example from GREGoR is CHASERR, a long non-coding RNA adjacent to CHD2 which was implicated in developmental and epileptic encephalopathy. The CHASERR discovery was a strong collaboration with the father of the initial proband serving as a coauthor on the manuscript, highlighting the power of patient partnerships in accelerating rare disease genomics81. Furthermore, using fiber-seq, GREGoR investigators have identified STRTS, an intergenic locus implicated in congenital hypothyroidism56. These findings illustrate the diagnostic potential of non-coding regions of the genome, which are not systematically included in standard variant analysis workflows. With thousands of non-coding transcripts still poorly understood, GREGoR continues to explore this untapped reservoir of genomic information, paving the way for novel disease-gene discoveries in the non-coding space.
Multilocus pathogenic variation (MPV) refers to the presence of multiple, independently pathogenic variants in multiple genes or loci that collectively contribute to an individual’s clinical manifestation. These variations can interact in complex ways, leading to compounded effects that may influence disease severity, onset or progression, often complicating diagnosis and treatment. GREGoR researchers have shown that as many as 5% of individuals for whom molecular diagnoses have been ascertained have MPV8, with this rate being even higher in rare disease families with parental consanguinity21,86. While individual variants comprising MPV are typically routinely diagnosed from just DNA sequencing, the interpretation of MPV typically requires deeper phenotypic analyses and represents a much broader area of gene dosage models for disease causing variation. The multiple de novo copy-number variant (MdnCNV) phenotype is a form of MPV whereby four or more independent, constitutional de novo copy-number variants arise in the same person within one generation. GREGoR researchers have shown that this ultrarare phenotype occurs in around 1 in over 12,000 individuals referred for genome-wide chromosomal microarray analysis87,88. MdnCNV typically requires integration of multiple technologies including arrays, short-read and long-read sequencing as well as quantitative phenotyping to fully characterize the genomic and clinical impact.
Complex genomic rearrangements (CGRs) are kilobase to megabase-scale structural changes in the genome involving multiple breakpoints, rearrangements and/or the integration of novel sequences in cis resulting in duplications, deletions, inversions and translocations, often affecting gene function and regulation. While CGRs have been catalogued and shown to be abundant across diverse populations, the accurate detection and assembly of CGRs, especially their breakpoints, typically requires more than short-read DNA sequencing28,29,89,90. Long-read sequencing (including adaptive sampling and ultra-long reads), optical genome mapping, chromosomal microarray analysis and linked-read sequencing can help to provide nucleotide resolution for CGRs to elucidate their mechanisms of formation to better understand how they precipitate genomic disorders and contribute to clinical variability and disease severity91–93.
Tandem repeats are sequences in which a nucleotide motif is repeated consecutively a varying number of times. These repetitive regions can be unstable, leading to expansions or contractions, which are associated with several genetic disorders94,95. Owing to the potential for multi-mapped short reads, only shorter tandem repeats have typically been well detected from short-read sequencing technologies96,97. However, long-read sequencing technologies are particularly effective at completely spanning short and longer tandem repeats, enabling accurate determination of repeat length and structure. Moreover, specialized bioinformatic tools designed for repeat analysis can help in accurately calling tandem repeats and identifying pathogenic expansions or contractions98–100. GREGoR is collating large databases and truth sets of tandem repeats across diverse populations to enable more systematic integration of tandem repeat analysis into sequencing pipelines101,102.
Mosaic variation originates from post-zygotic mutations and is considered germline if confined to the germ cells or somatic if acquired during or after the first mitotic divisions. Mosaic variation can lead to variations in phenotype, depending on the proportion and distribution of the mutant cells across tissues. Detecting mosaic variation often requires more than standard DNA sequencing due to the low variant allele fraction of mosaic alleles. Reliable detection of mosaic variation typically requires high-read-depth sequencing and orthogonal techniques such as digital droplet PCR for validation and bioinformatic discrimination between mosaic variants and sequencing artifacts. GREGoR is collaborating with the Somatic Mosaicism Across Human Tissues (SMaHT) consortium to understand pathogenic mosaic variation at scale.
Multi-nucleotide variants (MNVs) are two or more variants within the same codon on the same haplotype with over 50 MNVs per person103. Accurately identifying MNVs is bioinformatically challenging as a single MNV may be incorrectly interpreted as multiple independent variants. Missing MNVs can alter the interpretation of clinically pathogenic variation such as nonsense single-nucleotide variants that do not lead to a premature stop codon introduction104,105. Many variant effect predictors and multiplex assays of variant effects systematically score every possible MNV in a target locus. GREGoR is evaluating the use of these data towards variant of uncertain significance (VUS) reclassification and novel disease gene discovery.
NMD-escaping variants introduce premature stop codons in mRNA that evade NMD, producing truncated proteins of unknown function that may or may not manifest clinical disease. These variants are visible on primary DNA sequencing and because extensive work has been done to determine the rules of NMD escape106,107, many of these variants are already implicated in clinical disease and can be speculated on from just primary DNA sequencing70,71,108. However, newer approaches such as long read RNA-seq and proteomics are opening doors into the investigation of NMD escape alleles and their downstream mechanisms.
Incompletely penetrant variants refer to pathogenic changes in DNA that do not always result in observable clinical disease, even in individuals carrying the variant. This variability can complicate interpretation, especially in family studies where carriers appear to be unaffected. These variants often challenge diagnostic workflows because traditional penetrance assumptions do not hold, necessitating integration of orthogonal lines of evidence such as animal models and epigenetics to understand the variant pathogenicity109. These functional efforts are particularly critical for diseases where penetrance may be age-dependent, sex-influenced or modified by external factors. Furthermore, on a case-by-case analysis across all of gnomAD v4, more than 95% of incompletely penetrant pLOF variants found in severe, early-onset, highly penetrant haploinsufficient disease had explainable causes such as a downstream frame-restoring variant, predicted reinitiation by a downstream methionine, an MNV changing the interpretation of a nonsense variant to a missense or synonymous variant, or the location of the pLOF variant being in an NMD escape region110.
VUSs are most often reported in genes already established in disease pathogenesis, although, by definition, all variants found in candidate genes with insufficient evidence for disease implication are also VUSs. VUSs are accumulating rapidly over time as testing volume expands. In fact, VUSs are more often reported during panel testing compared with exome or genome sequencing due to professional practices111 and are disproportionately called in individuals from non-European ancestries112. Thus, while transitioning to consistent first-line clinical usage of exome or genome sequencing coupled with phenotypic analyses will decrease the rate of clinically reported VUSs, given most causal variants are only found in a single individual, integration of functional modelling is often required to reclassify VUSs. Multiplexed assays of variant effects (MAVEs) are high-throughput experiments orthogonal to the clinical sequencing pipeline that produce functional scores for all variant effects in a target locus and when their evidence strength is clinically calibrated and incorporated with other lines of evidence, demonstrate significant promise in massive VUS reclassification112.
Recruitment and return of results
Participation in rare-disease research can be influenced by numerous factors, including but not limited to phenotypic areas of focus or exclusion, as well as institutional, socioeconomic, geographical, linguistic, cultural, educational and insurance factors113–116. GREGoR sites, many of which are in urban centres, have implemented procedures for online enrolment, chatbots117, remote consent and offsite sample collection (including mobile phlebotomy in rural areas and for individuals with transportation challenges), and translated materials in multiple languages to improve accessibility. Attention to these details has enabled GREGoR to foster many international collaborations with local scientists to sequence and make available sequencing from thousands of individuals of non-European ancestry, actively seeking to address the disparities in genomic data availability across Middle Eastern, North African, Southeast Asian, South American and other under-represented groups such as African-American and Hispanic peoples. GREGoR is focused on solving unsolved cases, meaning exome-negative cases are prioritized and different sites will triage on a case-by-case basis for different -omic technologies including new exome and genome evaluations. While return of results varies per site and per case, GREGoR research centres have each developed rigorous standards and processes for return of high-confidence research results to participants and/or local clinicians who follow those participants. Attention is given to educating families on the distinction between ‘research results’ and ‘clinical testing results’, and centres have developed workflows to support participants in obtaining College of American Pathologists (CAP)/Clinical Laboratory Improvement Amendments (CLIA)-certified clinical laboratory confirmation of research findings when possible and appropriate.
Genomics for all
Diversifying genomics research through participant recruitment from under-represented populations is just one approach to fostering equity. GREGoR is pursuing orthogonal approaches to increase access to a molecular diagnosis by pursuing improved variant calling methods, applying multiplexed functional assays to improve interpretation of VUSs that are enriched in under-represented populations, and testing technologies and workflows that increase access to a genetic diagnosis for everyone. For example, GREGoR is collaborating with the HPRC to improve variant calling accuracy across all populations through the pangenome. GREGoR has analysed population biobanks to show a higher prevalence of VUSs and fewer ‘pathogenic’ or ‘likely pathogenic’ classifications in individuals of non-European genetic ancestry112. These disparities were alleviated by using high throughput, multiplexed functional experiments to test every possible single variant in genes of interest to resolve VUS disparities between populations. However, this study demonstrated that allele frequency and variant effect predictors contribute to unbalanced classification of variants and more work to prevent incorrect clinical variant classification is an important future priority for GREGoR. Beyond addressing disparities, the power of special and under-represented populations to make outsized contributions to rare disease genomics is unique and well established21,118. Distinct genetic phenomena—such as founder mutations in population isolates, large stretches of absence of heterozygosity in individuals from consanguineous populations, and private or ultrarare variants from individuals of ancestrally diverse backgrounds—often provide key genetic insights that lead to the discovery of disease–gene relationships and extrapolate to solving unsolved cases worldwide119–123. Finally, SeqFirst124 shows that using simple criteria to assess eligibility for rapid srGS significantly increases the proportion of non-white and Black infants who receive a precise genetic diagnosis. GREGoR and SeqFirst are conducting a comparison of long-read versus short-read sequencing in SeqFirst to better understand the relative value of these technologies within populations.
Accelerating data sharing
Data sharing is critical to advancing rare disease diagnoses125,126. GREGoR is committed to rapid release of genomic and phenotypic data to the larger research community within AnVIL and making these data findable, accessible, interoperable and reusable (FAIR)127 and machine readable. Inside AnVIL, GREGoR data are queryable through seqr, which integrates variant filtration, annotation and causal variant identification. Outside AnVIL, GREGoR has developed a public variant browser, which already includes over 95 million variants. Furthermore, de-identified phenotypes are being added to the public browser, enabling researchers to easily explore putative genotype–phenotype relationships in rare disease families. At the time of AnVIL submission and before analysis, deep phenotyping (Fig. 2a), potentially multiple orthogonal lines of -omics data (Fig. 2b,c) and pedigree data (Fig. 2d) are made available through the GREGoR data model for broader dissemination to the research community. Currently, DNA data on approximately 7,400 individuals from over 3,000 families are available with transcriptome data available for over 500 individuals and nearly 200 participants with both exome and short-read genome data (dbGaP: phs003047; Fig. 2b). Planned releases include additional short- and long-read genomes, short- and long-read RNA-seq, fiber-seq, ATAC-seq, metabolomics and proteomics. GREGoR has identified candidate discoveries and molecular diagnoses for over 400 families (Fig. 2e).
Fig. 2 |. Overview of publicly released GREGoR data.
Summary of third public data release (dbGaP: phs003047). a, The distribution of the top 30 phenotypes in GREGoR based on HPO descriptions. b, Table of numbers for probands and total individuals for each sequencing modality. c, The overlap across short-read genomes, RNA-seq and long-read genomes in data generation. d, Family structures comprising the overall cohort from a total of n = 3,610 families. e, Summary of current solved cases. Data are shared before analysis, but even the current diagnostic outcomes underscore the challenges and opportunities in resolving rare disease cases that are previously exome negative. WGS, whole-genome sequencing.
All GREGoR candidate genes are shared to Matchmaker Exchange through either GeneMatcher, seqr or MyGene2. Notably, for the 83 GREGoR publications (Supplementary Table 1) involving novel disease genes or phenotypic expansions, almost every project has been influenced by findings from connections made across or within Matchmaker Exchange nodes, whereas only 44 were supported by orthogonal functional experiments. Novel candidate genes and phenotypic expansions are curated for validity and publicly shared to the GenCC to accelerate access to early evidence of novel gene–disease relationships and aid in standardized clinical diagnostics and research. Analogously, candidate variants and molecular diagnoses are deposited in ClinVar.
Conclusion
Far from wrapping up the edges, these challenges represent a vast forefront in genomic research (Box 3), demanding both innovative methodologies and sustained collaboration to make meaningful progress. Alongside these challenges, advancements in genomic assays have required vetting at scale in individuals of diverse genetic ancestries and with diverse rare disease phenotypes. Such efforts are critical to establishing standards of when and how to use a specific approach and will guide expectations on their relative yields at scale and their adoption in clinical practice. Lastly, there is a palpable need to translate scientific discoveries into curation practices that align with formal clinical standards. To address these gaps, GREGoR provides data and infrastructure that will catalyse development and implementation of new approaches to advance genomics in rare disease by the broader community.
Box 3. The forefront of rare disease genomics.
(1). Finish the disease gene catalogue for protein-coding genes.
The effort to complete a comprehensive catalogue of genes underlying Mendelian conditions remains far from finished. While thousands of Mendelian conditions have been described, many of these conditions still lack a known genetic cause. Thousands of discoveries have been made using exome sequencing and hundreds of discoveries have been made in exome-negative cases using a variety of short-read or long-read genome sequencing, RNA-seq, novel analytical algorithms, new data sharing platforms and automated reanalysis strategies. The time is now to commit to a focused, unified, global, scaled effort to discover the remaining protein-coding disease genes and accelerate from a currently approximately 40% completed disease gene catalogue to a complete disease gene catalogue for protein-coding genes.
(2). Oligogenic and polygenic molecular diagnoses.
These involve the contribution of variants in two or more genes or loci that collectively lead to a clinical phenotype, a concept distinct from traditional monogenic inheritance patterns or MPV. These cases often present diagnostic challenges because the individual variants may not cause disease independently but act in concert to disrupt pathways or biological networks147–149. Currently, we can identify oligogenic variants through DNA sequencing, but the full scope of oligogenic and polygenic molecular diagnoses may require analysis across multiple layers of multi-omics data. For example, combinations of transcriptomic, proteomic or metabolomic alterations may converge with DNA variation to create synergistic effects that contribute to disease, representing a new frontier for rare disease genomics.
(3). Integrating large- and small-effect variants.
Currently, molecular diagnoses are considered to be complete once large-effect variant(s) are identified that can explain the majority of the primary clinical phenotype. However, non-coding modifiers and other small-effect variants can contribute to variable expressivity and incomplete penetrance. Despite their influence on the clinical presentation, these variants are not incorporated into clinical molecular diagnoses. A next frontier in rare disease genomics should include discovery of systematic approaches for identifying and integrating small-effect, modifier variants into diagnostic interpretation. Just as clinical infrastructures have been developed to support the identification and return of large-effect variants, parallel infrastructures are needed to support the clinical return of relevant small-effect variants that meaningfully modify disease presentation.
(4). Building multi-omic momentum.
Several non-coding genes have now been implicated in rare diseases, including numerous small nuclear RNA genes from the RNU family, as well as long non-coding RNAs and microRNAs, among others. These discoveries emerged from cases that remained unsolved by exome sequencing and required more comprehensive approaches such as genome sequencing or transcriptomic analysis for resolution56,81–84,150. Together, these studies demonstrate the quorum of investigations needed to build a compelling case for broader, scaled adoption of multi-omic methods—potentially as primary tools for rare disease diagnosis and discovery.
Supplementary Material
Additional information
Supplementary information The online version contains supplementary material available at https://doi.org/10.1038/s41586-025-09613-8.
Acknowledgements
We thank all of the patient participants and their families; and the expansive set of collaborators, including clinical providers, analysts and rare disease researchers. Support for title page creation and format was provided by AuthorArranger, a tool developed at the National Cancer Institute. This work was supported by the NIH NHGRI GREGoR Consortium (U01HG011758, U01HG011755, U01HG011762, U01HG011745, U01HG011744, U24HG011746).
GREGoR Partner Members
Aashish Adhikari42, Kinga M. Bujakowska43, Claudia M. B. Carvalho29, Ali Crawford44, Aimée Dudley29,45, Kelly D. Farwell Hagman46, Yang I. Li47, Jill E. Moore48, Aaron R. Quinlan49, Alex H. Wagner25,26,27, Bo Xia50 & S. Stephen Yi51,52
42Illumina Artificial Intelligence Laboratory, Illumina, Foster City, CA, USA. 43Department of Ophthalmology, Ocular Genomics Institute, Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA. 44Medical Genomics Research, Illumina, San Diego, CA, USA. 45Molecular and Cellular Biology Program, University of Washington, Seattle, WA, USA. 46Ambry Genetics, Aliso Viejo, CA, USA. 47Department of Medicine, University of Chicago, Chicago, IL, USA. 48Department of Genomics and Computational Biology, University of Massachusetts Chan Medical School, Worcester, MA, USA. 49Department of Human Genetics, University of Utah, Salt Lake City, UT, USA. 50Gene Regulation Observatory, Broad Institute of MIT and Harvard, Cambridge, MA, USA. 51Department of Neurosurgery, Baylor Research Institute, Baylor College of Medicine, Temple, TX, USA. 52Department of Medicine, School of Medicine—Temple, Baylor College of Medicine, Temple, TX, USA.
Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium
Moez Dawood1,2,3, Ben Heavner4, Marsha M. Wheeler4, Rachel A. Ungar5,6,7, Jonathan LoTempio8, Laurens Wiel5,6,9, Claudia M. B. Carvalho29, Richard A. Gibbs1,2, Casey A. Gifford5,30,31,32, Susanne May4, Danny E. Miller15,16,33,34, Heidi L. Rehm20,21, Kaitlin E. Samocha20,21, Fritz J. Sedlazeck1,2,35, Eric Vilain8, Anne O’Donnell-Luria20,21,36, Jennifer E. Posey2,37,38, Lisa H. Chadwick39, Michael J. Bamshad14,15,40, Stephen B. Montgomery5,6,41
U01HG011758
Jennifer E. Posey2,37,38, Richard A. Gibbs1,2, James R. Lupski1,2,18, Hatoon Al Ali2, Elizabeth G. Atkinson2, Sairam Behera1, Shaghayegh T. Beheshti2, Eric Boerwinkle53, Tugce Bozkurt-Yozgatli53,54, Daniel G. Calame55, Ivan Chinn56, Zeynep H. Coban-Akdemir2,53, Karen J. Coveler2, Zain Dardas2,46, Moez Dawood1,2,3, Harsha Doddapaneni1, Haowei Du2, Ruizhi Duan2, Iman Egab53, Jawid Fatih2, Mira Gandhi55, Brandon Garcia2, Nikhita Gogate2, Christopher M. Grochowski2, Jianhong Hu1, Minal Jamsandekar2, Shalini N. Jhangiani1, Angad Jolly2,57, Parneet Kaur2, Ahmed K. Saad2, Jesse M. Levine55, Richard A. Lewis2,18,58,59, Yidan Li2, Pengfei Liu2, Medhat Mahmoud1, Dana Marafi60,61,62, Tadahiro Mitani63, Chloe Munderloh2, Donna Muzny1,2, Sebastian Ochoa18, Piyush Panchal1, Shruti Pande2, Davut Pehlivan55,64,65, Archana Rai53, Edgar Andres Rivera-Munoz2, Aniko Sabo1, Evette Scott1, Fritz J. Sedlazeck1,2,35, Vernon Reid Sutton2, Kimberly Walker1, Lauren Westerfield2, Jiaoyang Xu53, Bo Yuan1,2,66 & Xinchang Zheng1
53Human Genetics Center, Department of Epidemiology, Human Genetics, and Environmental Sciences, School of Public Health, The University of Texas Health Science Center at Houston, Houston, TX, USA. 54Department of Biostatistics and Bioinformatics, Acibadem Mehmet Ali Aydinlar University, Istanbul, Turkey. 55Department of Pediatrics, Section of Pediatric Neurology and Developmental Neurosciences, Baylor College of Medicine, Houston, TX, USA. 56Department of Pediatrics, Division of Immunology, Allergy, and Retrovirology, Baylor College of Medicine, Houston, TX, USA. 57Department of Neurology, The University of Texas Dell Medical School, Austin, TX, USA. 58Department of Ophthalmology, Baylor College of Medicine, Houston, TX, USA. 59Department of Medicine, Baylor College of Medicine, Houston, TX, USA. 60Department of Pediatrics, College of Medicine, Kuwait University, Safat, Kuwait. 61Section of Child Neurology, Department of Pediatrics, Adan Hospital, Ministry of Health, Hadiya, Kuwait. 62Kuwait Medical Genetics Centre, Ministry of Health, Sulaibikhat, Kuwait. 63Department of Pediatrics, Jichi Medical University, Shimotsuke, Tochigi, Japan. 64Texas Children’s Hospital, Houston, TX, USA. 65Jan and Dan Duncan Neurological Research Institute, Texas Children’s Hospital, Houston, TX, USA. 66Baylor Genetics, Houston, TX, USA.
U01HG011755
Anne O’Donnell-Luria20,21,36, Heidi L. Rehm20,21, Michael E. Talkowski20,21,22,23,24, Siwaar Abouhala21, K. D. Ahlquist21, Mutaz Amin21, Christina Austin-Tse20,21, Samantha M. Baxter21, Benjamin Blankenmeister21, Philip M. Boone20,21,36, Harrison Brand20,21,22, Colleen Carlston21,36, Celine de Esch20,21, Stephanie DiTroia21, Michael Duyzend20,21,36, Vijay Ganesh21,67, Kiran Garimella21, Carmen Glaze21, Emily Groopman21, Sanna Gudmundsson21, Stacey Hall21, Yongqing Huang21, Julia Klugherz21, Katie Larsson21, Arthur S. Lee20,21,22, Gabrielle Lemire21, Jialan Ma21, Daniel MacArthur21,68,69, Brian Mangilog21, Daniel Marten21,36, Eva Martinez21, Olfa Messaoud21, Chloe Mighton21, Mariana Moyses21, Ashana Neale21, Emily O’Heir21, Melanie C. O’Leary21, Ikeoluwa Osei-Owusu21, Lynn Pais21, Alicia Pham21, Lindsay Romo21, Kathryn Russell21, Monica Salani20,21, Kaitlin Samocha20,21, Alba Sanchis-Juan20,21,22, Jillian Serrano21, Gulalai Shah21, Moriel Singer-Berk21, Mugdha Singh21,36, Hana Snow21, Kayla Socarras21, Sarah L. Stenton21,36, Jui-Cheng Tai20,21,22, Grace VanNoy21, Ben Weisburd21, Michael Wilson21, Monica Wojcik21,36,70, Isaac Wong21 & Rachita Yadav20,21,22
67Department of Neurology, Brigham and Women’s Hospital, Boston, MA, USA. 68Centre for Population Genomics, Garvan Institute of Medical Research, Sydney, New South Wales, Australia. 69Centre for Population Genomics, Murdoch Children’s Research Institute, Melbourne, Victoria, Australia. 70Department of Pediatrics, Division of Newborn Medicine, Boston Children’s Hospital, Boston, MA, USA.
U01HG011762
Stephen B. Montgomery5,6,41, Jonathan A. Bernstein13, Matthew T. Wheeler9, Emily Alsentzer41, Taylor M. Arriaga5, Euan A. Ashley5,9,41, Themistocles Assimes9, Gill Bejerano13,71,72, Devon Bonner13, Denver Bradley5, Jennefer Carter9, Clarisa Chavez Martinez5, Ziwei Chen71, Salil Deshpande73, Sara Emami9, Ivy Evergreen5, Casey A. Gifford5,30,31,32, Page Goddard6, John Gorzynski9, William Greenleaf5, Rodrigo Guarischi-Sousa9, Caitlin Harrington13, Sohaib Hassan74, Tanner D. Jensen5, David Jimenez-Morales9, Christopher Jin5, Aimee Juan5, Jessica Kain5, Laura Keehan13, Anshul Kundaje5,71, Soumya Kundu71, Samuel Lancaster5, Shruti Marwaha9, Dena R. Matalon13, Lauren Meador6, Hector Rodrigo Mendez9, Alexander Miller6, Matthew B. Neu30, Thuy-mi P. Nguyen6, Jonathan Nguyen6, Jeren D. Olsen6, Evin M. Padhi6, Paul Petrowski6, Astaria D. Podesta6, Elizabeth Porter13, Wanqiong Qiao6, Thomas Quertermous9, Chloe M. Reuter9, Oriane Rubio13, Stuart A. Scott6, Riya Sinha71, Kevin S. Smith6, Michael P. Snyder5, Brigitte Stark6, Suchitra Sudarshan5, Raquel L. Summers9, Christina G. Tise13, Philip Tsao9, Rachel A. Ungar5,6,7, Isabella Voutos9, Juliana M. Walrod9, Ziming Weng6, Laurens Wiel5,6,9, Frank Wong5, Yao Yang6, Jiye Yu9 & Jimmy Zhen9
71Department of Computer Science, Stanford University, Stanford, CA, USA. 72Department of Developmental Biology, School of Medicine, Stanford University, Stanford, CA, USA. 73Institute of Computational and Mathematical Engineering, Stanford University, Stanford, CA, USA. 74Department of Biomedical Data Science, School of Medicine, Stanford University, Stanford, CA, USA.
U01HG011745
Eric Vilain8, Seth Berger10,11,12, Emmanuèle C. Délot8, Miguel Almalvez75, Light Auriga8, Rebekah Barrick76, Sami Belhadj46, Krista Bluske46, Leandros Boukas11, Andrea J. Cohen10, Ya Cui77, Ivan De Dios75, Meghan Delaney78, John Harting46, Yun-Hua Hsiao46, Rachid Karam46, Charles Hadley King8, Arthur Ko11, Wei Li77, Bojan Losic46, Jonathan LoTempio8, Georgia Pitsava8 & Changrui Xiao79
75Department of Pediatrics, University of California Irvine, Irvine, CA, USA. 76Metabolic Disorders, Children’s Hospital Orange County (CHOC), Orange, CA, USA. 77Department of Biological Chemistry, School of Medicine, University of California Irvine, Irvine, CA, USA. 78Department of Pathology & Laboratory Medicine, Children’s National Hospital, Washington, DC, USA. 79Department of Neurology, School of Medicine, University of California Irvine, Irvine, CA, USA.
U01HG011744
Michael J. Bamshad14,15,40, Chia-Lin Wei16, Evan E. Eichler16,17, Jessica X. Chong14,15, Kailyn Anderson14, Peter Anderson16, Sabrina Best15,16, Elizabeth E. Blue80, Kati J. Buckingham14, Silvia Casadei15,16, Yong-Han Hank Cheng16, Colleen P. Davis16, Sophia B. Gibson16, William W. Gordon14, Jonas Gustafson14,45, William T. Harvey16, Martha Horike-Pyne80, Gail P. Jarvik16,80, Annelise Y. Mah-Som80, Colby T. Marvin14, F. Kumara Mastrorosa16, Sean R. McGee16, Heather C. Mefford81, Danny E. Miller15,16,33,34, Karynne Patterson16, Matthew Richardson16, Adriana E. Sedeño Cortés80, Joshua D. Smith16, Olivia M. Sommerland14, Lea M. Starita15,16, Andrew B. Stergachis15,16,80, Elliott G. Swanson16, Jeffrey Weiss16, Qian Yi16, Christina Zakarian16 & Miranda P. Zalusky14
80Department of Medicine, Division of Medical Genetics, University of Washington, Seattle, WA, USA. 81Cell & Molecular Biology, Center for Pediatric Neurological Disease Research, St Jude Children’s Research Hospital, Seattle, WA, USA.
U24HG011746
Susanne May4, Ali Shojaie4,19, Emily Bonkowski4, Sarah Conner4, Matthew P. Conomos4, Stephanie M. Gogarten4, Ben Heavner4, Sarah C. Nelson4, Sheryl Payne4, Jaime Prosser4, Guanghao Qi4, Adrienne M. Stilp4, Catherine C. Tong4, Marsha M. Wheeler4 & Quenna Wong4
NHGRI Program Management
Lisa H. Chadwick39, Christopher Wellington28, Sara Currin39 & Gabrielle C. Villard39
Footnotes
Competing interests R.A.G. declares that Baylor Genetics is a Baylor College of Medicine affiliate that derives revenue from genetic testing. BCM and Miraca Holdings have formed a joint venture with shared ownership and governance of Baylor Genetics, which performs clinical microarray analysis and other genomic studies (exome and genome sequencing) for patient and family care. F.J.S. has received research support from Illumina, Pacific Biosciences and Genentech. J.E.P. is an advisor to MaddieBio. S.B.M. is an advisor to MyOme, PhiTech and Valinor Therapeutics. F.J.S. and D.E.M. have received research support and/or consumables from ONT and have received travel funding to speak on behalf of ONT. D.E.M. has received travel support from Pacific Biosciences. D.E.M. is on an advisory board at ONT, a scientific advisory board at Basis Genetics and holds stock options in both MyOme and Basis Genetics. M.J.B. is the chair of the scientific advisory board of GeneDx and receives funding from the American Society of Human Genetics as the editor-in-chief of HGG Advances. J.X.C. receives funding from the American Society of Human Genetics as the Deputy Editor of HGG Advances. D.P. consults for Ionis Pharmaceuticals. H.L.R. and K.E.S. have received rare-disease research funding from Microsoft. H.L.R. has received research funding from Illumina and compensation as a past member of the scientific advisory board of Genome Medical. A.O.-L. was a paid consultant to Tome Biosciences, Ono Pharma USA, Addition Therapeutics, Congenica and receives research funding from Pacific Biosciences. E.E.E. is a scientific advisory board member of Variant Bio. M.P.S. is a cofounder and scientific advisor of Xthera, Exposomics, Filtricine, Fodsel, iollo, InVu Health, January AI, Marble Therapeutics, Mirvie, Next Thought AI, Orange Street Ventures, Personalis, Protos Biologics, Qbio, RTHM and SensOmics. M.P.S. is a scientific advisor of Abbratech, Applied Cognition, Enovone, Jupiter Therapeutics, M3 Helium, Mitrix, Neuvivo, Onza, Sigil Biosciences, Captify Inc, WndrHLTH, Yuvan Research and Ovul. A.A. and A.C. are employees and shareholders of Illumina, Inc.
A list of affiliations appears at the end of the paper.
References
- 1.Baxter SM et al. Centers for Mendelian Genomics: a decade of facilitating gene discovery. Genet. Med. 24, 784–797 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Taylor JC et al. Factors influencing success of clinical genome sequencing across a broad spectrum of disorders. Nat. Genet. 47, 717–726 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Wright CF et al. Genomic diagnosis of rare pediatric disease in the United Kingdom and Ireland. N. Engl. J. Med. 388, 1559–1571 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Rimmer A et al. Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications. Nat. Genet. 46, 912–918 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Monaco L et al. Research on rare diseases: ten years of progress and challenges at IRDiRC. Nat. Rev. Drug Discov. 21, 319–320 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Yang Y et al. Molecular findings among patients referred for clinical whole-exome sequencing. JAMA 312, 1870–1879 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Farnaes L et al. Rapid whole-genome sequencing decreases infant morbidity and cost of hospitalization. npj Genom. Med. 3, 10 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Posey JE et al. Resolution of disease phenotypes resulting from multilocus genomic variation. N. Engl. J. Med. 376, 21–31 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]; Multilocus genomic diagnoses occur in nearly 5% of solved exome cases, underscoring the complexity and need for comprehensive interpretation in rare disease phenotypes.
- 9.Turro E et al. Whole-genome sequencing of patients with rare diseases in a national health system. Nature 583, 96–102 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Bamshad MJ et al. The Centers for Mendelian Genomics: a new large-scale initiative to identify the genes underlying rare Mendelian conditions. Am. J. Med. Genet. A 158A, 1523–1525 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Posey JE et al. Insights into genetics, human biology and disease gleaned from family based genomic studies. Genet. Med. 21, 798–812 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Chong JX et al. The genetic basis of Mendelian phenotypes: discoveries, challenges, and opportunities. Am. J. Hum. Genet. 97, 199–215 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Surl D et al. Clinician-driven reanalysis of exome sequencing data from patients with inherited retinal diseases. JAMA Netw. Open 7, e2414198 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Seaby EG et al. A gene pathogenicity tool ‘GenePy’ identifies missed biallelic diagnoses in the 100,000 Genomes Project. Genet. Med. 26, 101073 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Liu P et al. Reanalysis of clinical exome sequencing data. N. Engl. J. Med. 380, 2478–2480 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]; Systematic reanalysis of clinical exome data substantially increased diagnostic yield and was driven by the newest novel disease gene discoveries.
- 16.Berger SI et al. Increased diagnostic yield from negative whole genome-slice panels using automated reanalysis. Clin. Genet. 104, 377–383 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Wenger AM, Guturu H, Bernstein JA & Bejerano G Systematic reanalysis of clinical exome data yields additional diagnoses: implications for providers. Genet. Med. 19, 209–214 (2017). [DOI] [PubMed] [Google Scholar]
- 18.Weisburd B et al. Diagnosing missed cases of spinal muscular atrophy in genome, exome, and panel sequencing data sets. Genet. Med. 27, 101336 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Guo MH et al. Inferring compound heterozygosity from large-scale exome sequencing data. Nat. Genet. 56, 152–161 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; This work presents a method to infer phasing of rare variant pairs in short-read exomes using gnomAD, enabling improved diagnosis of recessive Mendelian conditions.
- 20.Gudmundsson S et al. Variant interpretation using population databases: lessons from gnomAD. Hum. Mutat. 43, 1012–1030 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Mitani T et al. High prevalence of multilocus pathogenic variation in neurodevelopmental disorders in the Turkish population. Am. J. Hum. Genet. 108, 1981–2005 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Lemire G et al. Exome copy number variant detection, analysis, and classification in a large cohort of families with undiagnosed rare genetic disease. Am. J. Hum. Genet. 111, 863–876 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Du H et al. HMZDupFinder: a robust computational approach for detecting intragenic homozygous duplications from exome sequencing data. Nucleic Acids Res. 52, e18 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Babadi M et al. GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data. Nat. Genet. 55, 1589–1597 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Du H et al. VizCNV: An integrated platform for concurrent phased BAF and CNV analysis with trio genome sequencing data. Preprint at bioRxiv 10.1101/2024.10.27.620363 (2024). [DOI] [Google Scholar]
- 26.Wojcik MH et al. Beyond the exome: what’s next in diagnostic testing for Mendelian conditions. Am. J. Hum. Genet. 110, 1229–1248 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]; Offers a roadmap for diagnostic escalation beyond exome sequencing, including RNA-seq, genome sequencing, and long-read technologies, essential for solving unsolved Mendelian cases.
- 27.Sudmant PH et al. An integrated map of structural variation in 2,504 human genomes. Nature 526, 75–81 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Collins RL et al. A structural variation reference for medical and population genetics. Nature 581, 444–451 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Byrska-Bishop M et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Collins RL et al. Defining the diverse spectrum of inversions, complex structural variation, and chromothripsis in the morbid human genome. Genome Biol. 18, 36 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Wojcik MH et al. Genome sequencing for diagnosing rare diseases. N. Engl. J. Med. 390, 1985–1997 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; Genome sequencing provided unique diagnoses in 8% of previously unsolved cases.
- 32.Karczewski KJ et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Vitsios D, Dhindsa RS, Middleton L, Gussow AB & Petrovski S Prioritizing non-coding regions based on human genomic constraint and sequence context with deep learning. Nat. Commun. 12, 1504 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Chen S et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625, 92–100 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Zhang X et al. Genetic constraint at single amino acid resolution in protein domains improves missense variant prioritisation and gene discovery. Genome Med. 16, 88 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Chao KR et al. The landscape of regional missense mutational intolerance quantified from 125,748 exomes. Preprint at bioRxiv 10.1101/2024.04.11.588920 (2024). [DOI] [Google Scholar]
- 37.Saad AK et al. Biallelic in-frame deletion in TRAPPC4 in a family with developmental delay and cerebellar atrophy. Brain J. Neurol. 143, e83 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Dawood M et al. A biallelic frameshift indel in PPP1R35 as a cause of primary microcephaly. Am. J. Med. Genet. A 191, 794–804 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Dardas Z et al. NODAL variants are associated with a continuum of laterality defects from simple D-transposition of the great arteries to heterotaxy. Genome Med. 16, 53 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Miller DE et al. Targeted long-read sequencing identifies a retrotransposon insertion as a cause of altered GNAS Exon A/B methylation in a family with autosomal dominant pseudohypoparathyroidism type 1b (PHP1B). J. Bone Miner. Res. 37, 1711–1719 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Mori T et al. CFAP47 is implicated in X-linked polycystic kidney disease. Kidney Int. Rep. 9, 3580–3591 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Bruels CC et al. Diagnostic capabilities of nanopore long-read sequencing in muscular dystrophy. Ann. Clin. Transl. Neurol. 9, 1302–1309 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Chen X et al. Genome-wide profiling of highly similar paralogous genes using HiFi sequencing. Nat. Commun. 16, 2340 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Negi S et al. Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection. Am. J. Hum. Genet. 112, 428–449 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Mahmoud M et al. Closing the gap: solving complex medically relevant genes at scale. Preprint at medRxiv 10.1101/2024.03.14.24304179 (2024). [DOI] [Google Scholar]
- 46.Dias K-R et al. Narrowing the diagnostic gap: genomes, episignatures, long-read sequencing, and health economic analyses in an exome-negative intellectual disability cohort. Genet. Med. 26, 101076 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Bilgrav Saether K et al. Leveraging the T2T assembly to resolve rare and pathogenic inversions in reference genome gaps. Genome Res. 34, 1785–1797 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Fu Y et al. MethPhaser: methylation-based long-read haplotype phasing of human genomes. Nat. Commun. 15, 5327 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.LaFlamme CW et al. Diagnostic utility of DNA methylation analysis in genetically unsolved pediatric epilepsies and CHD2 episignature refinement. Nat. Commun. 15, 6524 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Zheng X et al. STIX: long-reads based accurate structural variation annotation at population scale. Preprint at bioRxiv 10.1101/2024.09.30.615931 (2024). [DOI] [Google Scholar]
- 51.Smolka M et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat. Biotechnol. 42, 1571–1580 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Sedlazeck FJ et al. Accurate detection of complex structural variations using single-molecule sequencing. Nat. Methods 15, 461–468 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Stergachis AB, Debo BM, Haugen E, Churchman LS & Stamatoyannopoulos JA Single-molecule regulatory architectures captured by chromatin fiber sequencing. Science 368, 1449–1454 (2020). [DOI] [PubMed] [Google Scholar]
- 54.Jha A et al. DNA-m6A calling and integrated long-read epigenetic and genetic analysis with fibertools. Genome Res. 34, 1976–1986 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Vollger MR et al. Synchronized long-read genome, methylome, epigenome and transcriptome profiling resolve a Mendelian condition. Nat. Genet. 57, 469–479 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]; By integrating long-read genome, methylome, epigenome, and transcriptome data, this study demonstrates how complex, multi-mechanism rare diseases can be mechanistically resolved in a single assay.
- 56.Grasberger H et al. STR mutations on chromosome 15q cause thyrotropin resistance by activating a primate-specific enhancer of MIR7–2/MIR1179. Nat. Genet. 56, 877–888 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Carvalho CMB et al. Interchromosomal template-switching as a novel molecular mechanism for imprinting perturbations associated with Temple syndrome. Genome Med. 11, 25 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Jensen TD et al. Integration of transcriptomics and long-read genomics prioritizes structural variants in rare disease. Genome Res. 35, 914–928 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Pais LS et al. seqr: a web-based analysis and collaboration tool for rare disease genomics. Hum. Mutat. 43, 698–707 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Lansdon LA et al. Factors affecting migration to GRCh38 in laboratories performing clinical next-generation sequencing. J. Mol. Diagn. 23, 651–657 (2021). [DOI] [PubMed] [Google Scholar]
- 61.Li H et al. Exome variant discrepancies due to reference-genome differences. Am. J. Hum. Genet. 108, 1239–1250 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]; Discrepancies between GRCh37 and GRCh38 reference genomes affect variant calling in around 200 genes including Mendelian disease genes.
- 62.Ungar RA et al. Impact of genome build on RNA-seq interpretation and diagnostics. Am. J. Hum. Genet. 111, 1282–1300 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; Genome build choice significantly alters RNA-seq interpretation across >2800 genes, directly impacting transcriptomics-guided rare disease diagnostics.
- 63.Behera S et al. FixItFelix: improving genomic analysis by fixing reference errors. Genome Biol. 24, 31 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Wagner J et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672–680 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Liao W-W et al. A draft human pangenome reference. Nature 617, 312–324 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Behera S et al. Comprehensive genome analysis and variant detection at scale using DRAGEN. Nat. Biotechnol. 43, 1177–1191 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]; The DRAGEN pipeline is a high-accuracy, scalable variant detection method that leverages multigenome mapping and machine learning to identify all major variant types—including single-nucleotide variants, copy-number variants, structural variants and short tandem repeats—across diverse populations.
- 67.Chin C-S et al. A pan-genome approach to decipher variants in the highly complex tandem repeat of LPA. Preprint at bioRxiv 10.1101/2022.06.08.495395 (2022). [DOI] [Google Scholar]
- 68.Samocha KE et al. A framework for the interpretation of de novo mutation in human disease. Nat. Genet. 46, 944–950 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]; A framework to evaluate de novo mutations improves gene discovery in rare diseases by distinguishing pathogenic mutations from background variation.
- 69.Lupski JR, Belmont JW, Boerwinkle E & Gibbs RA Clan genomics and the complex architecture of human disease. Cell 147, 32–43 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]; Describes the foundational framework for emphasizing the role of recent, rare, and private variants in disease risk and highlights the importance of family-centric rare disease analysis.
- 70.Teran NA et al. Nonsense-mediated decay is highly stable across individuals and tissues. Am. J. Hum. Genet. 108, 1401–1408 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Coban-Akdemir Z et al. Identifying genes whose mutant transcripts cause dominant disease traits by potential gain-of-function alleles. Am. J. Hum. Genet. 103, 171–187 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Valencia AM et al. Landscape of mSWI/SNF chromatin remodeling complex perturbations in neurodevelopmental disorders. Nat. Genet. 55, 1400–1412 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Paine I et al. Paralog studies augment gene discovery: DDX and DHX genes. Am. J. Hum. Genet. 105, 302–316 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Gillentine MA et al. Rare deleterious mutations of HNRNP genes result in shared neurodevelopmental disorders. Genome Med. 13, 63 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Amberger JS, Bocchini CA, Scott AF & Hamosh A OMIM.org: leveraging knowledge across phenotype-gene relationships. Nucleic Acids Res. 47, D1038–D1043 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.DiStefano MT et al. The gene curation coalition: a global effort to harmonize gene-disease evidence resources. Genet. Med. 24, 1732–1742 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Ochoa S et al. A deep intronic splice-altering AIRE variant causes APECED syndrome through antisense oligonucleotide-targetable pseudoexon inclusion. Sci. Transl. Med. 16, eadk0845 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Wu N et al. TBX6 null variants and a common hypomorphic allele in congenital scoliosis. N. Engl. J. Med. 372, 341–350 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Lord J et al. Non-coding variants are a rare cause of recessive developmental disorders in trans with coding variants. Genet. Med. 26, 101249 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Mao K et al. FOXI3 pathogenic variants cause one form of craniofacial microsomia. Nat. Commun. 14, 2026 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Ganesh VS et al. Neurodevelopmental disorder caused by deletion of CHASERR, a lncRNA gene. N. Engl. J. Med. 391, 1511–1518 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; Discovery of a neurodevelopmental disorder caused by CHASERR long non-coding deletion reveals regulatory non-coding elements as critical contributors to rare disease pathogenesis.
- 82.Greene D et al. Mutations in the U4 snRNA gene RNU4–2 cause one of the most prevalent monogenic neurodevelopmental disorders. Nat. Med. 30, 2165–2169 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Chen Y et al. De novo variants in the RNU4–2 snRNA cause a frequent neurodevelopmental syndrome. Nature 632, 832–840 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Greene D et al. Mutations in the small nuclear RNA gene RNU2–2 cause a severe neurodevelopmental disorder with prominent epilepsy. Nat. Genet. 57, 1367–1373 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Nava C et al. Dominant variants in major spliceosome U4 and U5 small nuclear RNA genes cause neurodevelopmental disorders through splicing disruption. Nat. Genet. 57, 1374–1388 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Bozkurt-Yozgatli T et al. Multilocus pathogenic variants contribute to intrafamilial clinical heterogeneity: a retrospective study of sibling pairs with neurodevelopmental disorders. BMC Med. Genom. 17, 85 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Liu P et al. An organismal CNV mutator phenotype restricted to early human development. Cell 168, 830–842 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Du H et al. The multiple de novo copy number variant (MdnCNV) phenomenon presents with peri-zygotic DNA mutational signatures and multilocus pathogenic variation. Genome Med. 14, 122 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Logsdon GA et al. Complex genetic variation in nearly complete human genomes. Nature 644, 430–441 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Ebert P et al. Haplotype-resolved diverse human genomes and integrated analysis of structural variation. Science 372, eabf7117 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Grochowski CM et al. Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci. Cell Genom. 4, 100590 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Dardas Z et al. Genomic balancing act: deciphering DNA rearrangements in the complex chromosomal aberration involving 5p15.2, 2q31.1, and 18q21.32. Eur. J. Hum. Genet. 33, 231–238 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Pehlivan D et al. Structural variant allelic heterogeneity in MECP2 duplication syndrome provides insight into clinical severity and variability of disease expression. Genome Med. 16, 146 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 94.Jakubosky D et al. Discovery and quality analysis of a comprehensive set of structural variants and short tandem repeats. Nat. Commun. 11, 2928 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Jakubosky D et al. Properties of structural variants and short tandem repeats associated with gene expression and complex traits. Nat. Commun. 11, 2927 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Dolzhenko E et al. Characterization and visualization of tandem repeats at genome scale. Nat. Biotechnol. 42, 1606–1614 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.English AC et al. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat. Biotechnol. 43, 431–442 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Dolzhenko E et al. REViewer: haplotype-resolved visualization of read alignments in and around tandem repeats. Genome Med. 14, 84 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Behera S et al. Identification of allele-specific KIV-2 repeats and impact on Lp(a) measurements for cardiovascular disease risk. BMC Med. Genom. 17, 255 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Weisburd B, Tiao G & Rehm HL Insights from a genome-wide truth set of tandem repeat variation. Preprint at bioRxiv 10.1101/2023.05.05.539588 (2023). [DOI] [Google Scholar]
- 101.Cui Y et al. A genome-wide spectrum of tandem repeat expansions in 338,963 humans. Cell 187, 2336–2341 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; This study establishes a biobank-scale, population reference of tandem repeat expansions across ancestries from short-read genome sequencing.
- 102.Weisburd B et al. Defining a tandem repeat catalog and variation clusters for genome-wide analyses and population databases. Preprint at bioRxiv 10.1101/2024.10.04.615514 (2024). [DOI] [Google Scholar]
- 103.Wang Q et al. Landscape of multi-nucleotide variants in 125,748 human exomes and 15,708 genomes. Nat. Commun. 11, 2539 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Srinivasan S et al. Misannotated multi-nucleotide variants in public cancer genomics datasets lead to inaccurate mutation calls with significant implications. Cancer Res. 81, 282–288 (2021). [DOI] [PubMed] [Google Scholar]
- 105.Campbell IM et al. Multiallelic positions in the human genome: challenges for genetic analyses. Hum. Mutat. 37, 231–234 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Lindeboom RGH, Vermeulen M, Lehner B & Supek F The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer immunotherapy. Nat. Genet. 51, 1645–1651 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Lindeboom RGH, Supek F & Lehner B The rules and impact of nonsense-mediated mRNA decay in human cancers. Nat. Genet. 48, 1112–1118 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Torene RI et al. Systematic analysis of variants escaping nonsense-mediated decay uncovers candidate Mendelian diseases. Am. J. Hum. Genet. 111, 70–81 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 109.Potter AS et al. Rare variant in MRC2 associated with familial supraventricular tachycardia and Wolff-Parkinson-White syndrome. Circ. Genomic Precis. Med. 17, e004614 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 110.Gudmundsson S et al. Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database. Preprint at bioRxiv 10.1101/2024.06.12.593113 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 111.Rehm HL et al. The landscape of reported VUS in multi-gene panel and genomic testing: time for a change. Genet. Med. 25, 100947 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]; This multi-laboratory analysis reveals the high prevalence and clinical burden of VUSs in genetic testing from panel testing and advocates for refined reporting practices to reduce VUSs.
- 112.Dawood M et al. Using multiplexed functional data to reduce variant classification inequities in underrepresented populations. Genome Med. 16, 143 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; This study defines variant classification disparities in several biobanks and demonstrates how multiplexed functional data can reduce variant classification disparities across ancestries, offering a scalable strategy to make genomic medicine more equitable.
- 113.Young JL et al. Beyond race: recruitment of diverse participants in clinical genomics research for rare disease. Front. Genet. 13, 949422 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 114.Wojcik MH et al. Rare diseases, common barriers: disparities in pediatric clinical genetics outcomes. Pediatr. Res. 93, 110–117 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Serrano JG et al. Advancing understanding of inequities in rare disease genomics. Clin. Ther. 45, 745–753 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.D’Angelo CS et al. Barriers and considerations for diagnosing rare diseases in indigenous populations. Front. Pediatr. 8, 579924 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 117.Savage SK et al. Using a chat-based informed consent tool in large-scale genomic research. J. Am. Med. Inform. Assoc. 31, 472–478 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; A chat-based consent tool successfully scaled participant enrollment for large rare disease genomics studies, reducing staff burden while maintaining participant understanding.
- 118.Abou Tayoun AN & Rehm HL Genetic variation in the Middle East—an opportunity to advance the human genetics field. Genome Med. 12, 116 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 119.AlAbdi L et al. Diagnostic implications of pitfalls in causal variant identification based on 4577 molecularly characterized families. Nat. Commun. 14, 5269 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 120.AlAbdi L et al. Beyond the exome: utility of long-read whole genome sequencing in exome-negative autosomal recessive diseases. Genome Med. 15, 114 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 121.AlAbdi L et al. Arab founder variants: Contributions to clinical genomics and precision medicine. Med 6, 100528 (2025). [DOI] [PubMed] [Google Scholar]
- 122.Sulem P et al. Identification of a large set of rare complete human knockouts. Nat. Genet. 47, 448–452 (2015). [DOI] [PubMed] [Google Scholar]
- 123.Oddsson A et al. Deficit of homozygosity among 1.52 million individuals and genetic causes of recessive lethality. Nat. Commun. 14, 3453 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 124.Wenger TL et al. SeqFirst: building equity access to a precise genetic diagnosis in critically ill newborns. Am. J. Hum. Genet. 112, 508–522 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 125.Stark Z et al. A call to action to scale up research and clinical genomic data sharing. Nat. Rev. Genet. 26, 141–147 (2025). [DOI] [PubMed] [Google Scholar]
- 126.Rehm HL Time to make rare disease diagnosis accessible to all. Nat. Med. 28, 241–242 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 127.Wilkinson MD et al. The FAIR guiding principles for scientific data management and stewardship. Sci. Data 3, 160018 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 128.Bamshad MJ, Nickerson DA & Chong JX Mendelian gene discovery: fast and furious with no end in sight. Am. J. Hum. Genet. 105, 448–455 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 129.Wagner AH et al. The GA4GH variation representation specification: a computational framework for variation representation and federated identification. Cell Genom. 1, 100027 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 130.Pawliczek P et al. ClinGen allele registry links information about genetic variants. Hum. Mutat. 39, 1690–1701 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 131.Köhler S et al. The Human Phenotype Ontology in 2021. Nucleic Acids Res. 49, D1207–D1217 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 132.Stegmann JD et al. Bi-allelic variants in CELSR3 are implicated in central nervous system and urinary tract anomalies. npj Genom. Med. 9, 18 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 133.Herman I et al. Quantitative dissection of multilocus pathogenic variation in an Egyptian infant with severe neurodevelopmental disorder resulting from multiple molecular diagnoses. Am. J. Med. Genet. A 188, 735–750 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 134.Calame DG et al. Monoallelic variation in DHX9, the gene encoding the DExH-box helicase DHX9, underlies neurodevelopment disorders and Charcot-Marie-Tooth disease. Am. J. Hum. Genet. 110, 1394–1413 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 135.Jolly A et al. Rare variant enrichment analysis supports GREB1L as a contributory driver gene in the etiology of Mayer-Rokitansky-Küster-Hauser syndrome. HGG Adv. 4, 100188 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 136.Lima AR et al. Phenotypic and mutational spectrum of ROR2-related Robinow syndrome. Hum. Mutat. 43, 900–918 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 137.Zhang C et al. Novel pathogenic variants and quantitative phenotypic analyses of Robinow syndrome: WNT signaling perturbation and phenotypic variability. HGG Adv. 3, 100074 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 138.Garcia BT et al. Improving automated deep phenotyping through large language models using retrieval-augmented generation. Genome Med. 17, 91 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 139.Gustafson JA et al. High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation. Genome Res. 34, 2061–2073 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]; High-coverage long-read ONT sequencing of 1000 Genomes Project samples enables improved detection of structural variants and epigenetic changes.
- 140.Harrison PW et al. Ensembl 2024. Nucleic Acids Res. 52, D891–D899 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 141.Liu X, Li C, Mou C, Dong Y & Tu Y dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. 12, 103 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 142.Wang K, Li M & Hakonarson H ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 143.Pagel KA et al. Integrated informatics analysis of cancer-related variants. JCO Clin. Cancer Inform. 4, 310–317 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 144.Rodrigues EDS et al. Variant-level matching for diagnosis and discovery: challenges and opportunities. Hum. Mutat. 43, 782–790 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 145.Seaby EG et al. A panel-agnostic strategy ‘HiPPo’ improves diagnostic efficiency in the UK Genomic Medicine Service. Healthcare 11, 3179 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 146.Chong JX et al. Considerations for reporting variants in novel candidate genes identified during clinical genomic testing. Genet. Med. 26, 101199 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 147.Rai A et al. Genomic rare variant mechanisms for congenital cardiac laterality defect: a digenic model approach. Am. J. Hum. Genet. 112, 1664–1680 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 148.Töpf A et al. Digenic inheritance involving a muscle-specific protein kinase and the giant titin protein causes a skeletal muscle myopathy. Nat. Genet. 56, 395–407 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 149.Gifford CA et al. Oligogenic inheritance of a human heart disease involving a genetic modifier. Science 364, 865–870 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 150.Arriaga TM et al. Transcriptome-wide outlier approach identifies individuals with minor spliceopathies. Am. J. Hum. Genet. 112, 2458–2475 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.


