Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2025 Feb 26;112(3):693–708. doi: 10.1016/j.ajhg.2025.02.001

Genome-wide prediction of dominant and recessive neurodevelopmental disorder-associated genes

Ryan S Dhindsa 1,2,3,8,, Blake A Weido 1,8, Justin S Dhindsa 2,4, Arya J Shetty 2, Chloe F Sands 1,2, Slavé Petrovski 5,6, Dimitrios Vitsios 5, Anthony W Zoghbi 1,7,∗∗
PMCID: PMC11947176  PMID: 40015282

Summary

Despite great progress, thousands of neurodevelopmental disorder (NDD) risk genes remain to be discovered. We present a computational approach that accelerates NDD risk gene identification using machine learning. First, we demonstrate that models trained solely on single-cell RNA sequencing data can robustly predict genes implicated in autism spectrum disorder (ASD), developmental and epileptic encephalopathy (DEE), and developmental delay (DD). Notably, we find differences in gene expression patterns of genes with monoallelic and bi-allelic inheritance patterns in the developing human cortex. We then integrate expression data with 300 orthogonal features, including intolerance metrics, protein-protein interaction data, and others, in a semi-supervised machine learning framework (mantis-ml) to train inheritance-specific models for these disorders. The models have high predictive power (area under the receiver operator curves [AUCs]: 0.84–0.95), and the top-ranked genes were up to 2-fold (monoallelic models) and 6-fold (bi-allelic models) more enriched for high-confidence NDD risk genes compared to genic intolerance metrics alone. Additionally, genes ranking in the top decile were 45 to 180 times more likely to have literature support than those in the bottom decile. Collectively, this work provides robust NDD risk gene predictions that can complement large-scale gene discovery efforts and underscores the importance of considering inheritance in gene risk prediction.

Keywords: genetics, machine learning, computational biology, neurodevelopmental disease, epilepsy, autism spectrum disorder, intellectual disability


Thousands of neurodevelopmental disorder (NDD)-associated genes remain undiscovered, limiting genetic diagnoses. Here, we employ inheritance-specific machine learning models, incorporating gene expression data with hundreds of genetic features to predict new NDD-associated genes. These models achieve high accuracy in predicting gene-disease associations and can thus accelerate the discovery of genetic NDDs.

Introduction

Neurodevelopmental disorders (NDDs), including autism spectrum disorder (ASD) (MIM: 209850), developmental and epileptic encephalopathy (DEE) (MIM: PS308350), and developmental delay (DD) (MIM: PS249500), are highly heritable. Researchers have made great progress in identifying hundreds of genes associated with these disorders through sequencing studies of trios, families, and case-control cohorts.1,2,3,4,5,6,7 While this growing catalog of NDD risk genes has advanced the field, our understanding remains incomplete. We use the term “risk gene” in this study to refer to such genes that harbor genetic variants associated with an increased likelihood of developing NDDs, noting that it is not the gene itself but specific mutations or variants within the gene that contribute to disease risk. Most individuals with an NDD do not receive a genetic diagnosis,8 in part because there are more NDD-associated genes to discover. In the case of ASD, only 190 of the estimated 1,000 risk genes have been confidently linked to disease,9 even as cohort sizes have grown to over 20,000 cases.6 Fully characterizing the genetic architecture of NDDs is crucial to making accurate molecular diagnoses, elucidating disease mechanisms, and developing targeted therapies but will likely require sequencing of hundreds of thousands of additional affected individuals.2

In silico approaches can help predict NDD risk genes and accelerate gene discovery. For example, we and others have shown that genes associated with severe early-onset disorders are under strong purifying selection and thus tend to be depleted of nonsynonymous variation in the general population.10,11,12,13,14 Genic intolerance metrics, which quantify the degree to which genes are intolerant to functional variation, have become a cornerstone in prioritizing NDD risk genes.1,2,6,15,16,17,18,19 However, not all intolerant genes are involved in NDDs, as any gene in which mutations reduce fecundity will be intolerant to variation (e.g., genes involved in fertility). Moreover, although population-level sequencing datasets continue to grow, intolerance metrics still suffer from a lack of power for smaller genes. Finally, perhaps the biggest current limitation is that although these scores can reliably detect purifying selection against variants with monoallelic/dominant inheritance patterns, they struggle to prioritize disease genes with bi-allelic/recessive modes of inheritance.20,21,22 Moreover, to our knowledge, there are currently no available disease-specific computational risk predictors for recessive disorders.

Other commonly used methods for predicting NDD risk genes rely on gene expression networks.23,24 However, most of these methods have been based on bulk RNA sequencing data and thus do not account for potential cell-type-specific expression patterns. Here, we hypothesized that we could bolster NDD risk gene predictions by integrating genic intolerance, bulk and single-cell RNA sequencing (scRNA-seq) data, and other orthogonal datasets in an inheritance-specific manner. First, we assess cell-type-specific expression patterns for ASD, DEE, and DD genes stratified by inheritance pattern (i.e., monoallelic versus bi-allelic). We then demonstrate that expression patterns alone can predict NDD risk genes but that these predictions significantly improve when used in combination with intolerance metrics. Finally, we use scRNA-seq data, intolerance metrics, and hundreds of other gene-level annotations in a semi-supervised machine learning approach (mantis-ml)25 to generate inheritance-specific risk gene predictions for ASD, DEE, and DD. The top risk gene predictions from these models show a striking enrichment for top genes from trio studies and large case-control analyses, expert-curated risk gene lists, and genes enriched for their related phenotype associations in published case reports and case series. We make the scores available through a public browser: https://nddgenes.com.

Methods

Seed gene list curation

A seed gene is a gene that has already been confirmed to be associated with a particular disease through previous research. Thus, these genes serve as known positive examples for model training. We used Simons Foundation Autism Research Initiative (SFARI)26 tier 1 ASD genes (n = 207), the highest confidence ranking, as the basis for our monoallelic ASD model. We then reviewed each of the tier 1 genes to ensure that they were associated with ASD through a monoallelic mechanism and removed genes that had a bi-allelic mechanism (e.g., ADSL [MIM: 608222] and ALDH5A1 [MIM: 610045]) or weak evidence of association to ASD based on the most recent large-scale studies of ASD (e.g., KATNAL2 [MIM: 614697]). After filtering, we were left with 190 monoallelic ASD seed genes.

For the DD monoallelic and bi-allelic models, we selected the 832 genes with “definitive” confidence and “brain/cognition” organ involvement from the DECIPHER Developmental Disorder Genotype-to-Phenotype Database (DDG2P).27 The DDG2P provides mechanism-of-inheritance data for each gene, and we used this information to separate the gene lists into those with monoallelic (n = 218) and bi-allelic (n = 449) inheritance patterns. For the DD monoallelic model, we combined the DDG2P27 monoallelic genes with 199 genome-wide significant genes from the largest trio exome sequencing study of DD,2 resulting in a total of 417 monoallelic seed genes for DD.

For DEE, we used a three-step selection process. First, we extracted genes from Online Mendelian Inheritance in Man (OMIM)28 specifically associated with the phenotypic series of DDE, categorizing them into monoallelic (n = 54) and bi-allelic (n = 41) inheritance patterns. Second, we performed an advanced search for “epileptic encephalopathy” and identified additional established NDD genes where at least two affected individuals had documented DEE, though these genes were not included in the official DEE phenotypic series (adding 19 monoallelic and 22 bi-allelic genes). Finally, we incorporated clinically curated monoallelic DEE genes from the most recent Epi25k study of epilepsy (adding 21 monoallelic genes).1 This comprehensive approach yielded a total of 94 monoallelic and 63 bi-allelic DEE-associated genes for model training.

If substantial evidence supported both monoallelic and bi-allelic modes of inheritance for an NDD-associated gene (e.g., ATP1A2 [MIM: 182340]), then we included it in the seed gene list for both the monoallelic and bi-allelic models. Generally, these genes can confer pathogenicity through bi-allelic loss-of-function variants and dominant-negative or gain-of-function monoallelic missense variation.

Fetal cortex scRNA-seq analysis

We downloaded scRNA expression data generated from human fetal cortical samples as described in a prior publication.29 These data, which include four samples from an 8-week span during mid-gestation, consisted of 57,868 single-cell transcriptomes. Using the same cell identity annotations from the original publication, we calculated the module score for all cells using Seurat’s AddModuleScore function with the seed gene list for each NDD as the input feature.30,31 We then Z score normalized these module scores to evaluate the relative expression of disease-associated genes between cell clusters. We also calculated the average unique molecular identifier (UMI) counts of all genes per cell type per age of tissue. These were used as features in the machine learning models.

GO enrichment analysis

We performed comprehensive Gene Ontology (GO) enrichment analysis to characterize both inheritance-specific patterns and prediction-based functional relationships across our gene sets. Our analyses utilized the gprofiler2 R package32 to query multiple functional annotation sources, including GO biological process (GO:BP), cellular component (GO:CC), molecular function (GO:MF), and human phenotype (HP) ontologies.33 Statistical significance was assessed using the default g:set counts and sizes (SCS) multiple testing correction method with a threshold of p < 0.05.

To investigate inheritance-specific enrichment patterns, we conducted two parallel GO enrichment analyses. The first analysis examined our curated list of developmental delay (DD) genes, stratified into monoallelic and bi-allelic inheritance groups. The second analysis consisted of genes from the OMIM database,28 categorized into two groups: a monoallelic group comprising genes annotated with autosomal dominant or X-linked inheritance patterns and a bi-allelic group comprising genes annotated with autosomal recessive inheritance.

To comprehensively characterize the biological and functional role of genes across the full spectrum of mantis-ml predictions, we conducted an additional stratified analysis based on prediction confidence. For each model, we ranked genes according to their mantis-ml prediction probabilities and divided them into deciles, with decile one containing genes with the highest percentile range (90th–100th). We then performed GO enrichment analysis for each decile independently, allowing us to identify biological and functional patterns associated with different levels of prediction confidence.

scRNA-seq models

To demonstrate the baseline power of scRNA-seq data as a predictor of NDD risk genes, we evaluated the performance of a random forest model for each set of curated seed genes. Due to the intrinsic imbalance between the positively labeled and unlabeled genes, machine learning models can quickly become biased toward the majority class. To address this challenge, we employed two strategies: reducing the size of the overrepresented class to prevent model overfitting and generating multiple balanced datasets to ensure robust validation across different subsets of unlabeled genes. For each inheritance-specific phenotype, we created ten balanced datasets, each containing all positively labeled genes and a random subset of unlabeled genes at a ratio of 1:1.5 positive to unlabeled genes. Next, we performed zero imputation and removed highly correlated features (Pearson’s r > 0.95). We used the scikit-learn library in Python to construct the random forest model with the default parameters and performed repeated stratified 5-fold cross-validation for 2 repeats. For each model run, we recorded the average area under the receiver operator curve (AUC) and standard deviation across all iterations. Additionally, we compared the performance of the fetal cortex and Human Protein Atlas (HPA) v.2334 scRNA-seq expression models to models trained on intolerance metrics, including missense Z (misZ), residual variation intolerance score (RVIS), loss-of-function observed/expected upper bound fraction (LOEUF), and probability of being recessive (pREC).10,13,20

The Boruta algorithm is an iterative feature selection method that uses random forests to assess the statistical significance of features. Unlike other feature selection methods where features are compared against each other, Boruta compares each feature against randomized versions of the original feature set called “shadow” features. Features achieving less significant importance than the shadow features are progressively eliminated. Eventually, a “confirmed” set of features (i.e., features that are considered predictive) are identified and ranked based on Z scores representing importance scores. To ensure a robust feature selection and reduce class imbalance, we implemented the Boruta algorithm in R across ten balanced datasets using the output mantis-ml processed feature table. Specifically, we performed the analysis across two feature sets: one including all features and another restricted to only scRNA-seq features. Feature importance scores were generated across each balanced dataset and then aggregated using the median. We repeated this step for each model using the default parameters and the corresponding processed feature table.

Mantis-ml

The mantis-ml framework has been previously described in detail.25 Briefly, mantis-ml is a semi-supervised machine learning framework for the prediction of possible disease-associated genes. Following its initial setup, disease/phenotype terms of interest provided to mantis-ml are used to automatically extract associated known disease genes and relevant features. In this manuscript, given the relative importance of using high-confidence seed genes, we elected to manually curate our seed genes. Next, mantis-ml annotates all genes in the genome with several hundred diverse features. The semi-supervised learning method employed by mantis-ml infers the risk of a gene’s association with a specific phenotype based on the similarity between the feature signature of the gene and that of the seed genes, as captured by hundreds or thousands of balanced sets comprising known and unlabeled genes. Mantis-ml then generates exome-wide gene-level risk prediction probabilities and their corresponding percentiles for the phenotype of interest.

We instituted several key improvements to the original mantis-ml framework to improve performance for NDD risk gene prediction. First, we integrated scRNA-seq data from both the human fetal cortex and HPA.34 The fetal cortex data provided average gene expression profiles for major cell types across four developmental time points, while the HPA34 data offered comprehensive whole-body, tissue, and cell-type-specific expression measures, quantified as normalized transcripts per million (nTPM). We expanded on the previous collection of gene intolerance metrics in mantis-ml by adding the gene variation intolerance rank (GeVIR),35 which performs well for smaller genes and missense intolerant genes. We also added GeVIR’s LOEUF-joined derivative ViRLoF, which has been shown to outperform LOEUF alone in prioritizing NDD risk genes. Lastly, we included GeVIR’s fold enrichment scores for autosomal dominant and recessive modes of inheritance for each gene. GO33 terms are a powerful tool for describing the relationship between a gene/gene product and its functional, molecular, and spatial properties. The original mantis-ml framework incorporated GO terms by applying a pattern search using the disease/phenotype input terms and collapsing the number of associations between a gene and the matched GO terms into a new, one-hot-encoded feature per input term. We now expand the GO feature set by also including the top 20 individual GO terms that seed genes are most enriched for compared to the rest of the exome (quantified via Fisher’s exact test), further increasing the strength of the mantis-ml feature set.

We used fixed configuration and classifier parameters for each input seed gene list and their corresponding disease/phenotype terms of interest, as described in the original mantis-ml publication (Table S1). Prior to model training and inference, mantis-ml automatically performs preprocessing and exploratory data analysis (EDA). The initial preprocessing step of mantis-ml performs feature filtering by calculating Pearson’s correlation coefficient between all features and dropping those with correlations above a defined threshold. To prevent the removal of valuable scRNA-seq features due to genes with low or non-existent expression in a given tissue of interest, we specified a high correlation threshold of 0.95 (default = 0.8). We retained the default mantis-ml parameters for the remainder of the preprocessing and EDA steps.

For the stochastic semi-supervised component of mantis-ml, we ran the random forest, support vector classifier, gradient boosting, and extreme gradient boosting (XGBoost) classifiers, with hyperparameters optimized through mantis-ml’s built-in grid search module prior to training. The mantis-ml workflow began by generating balanced datasets (M) containing a ratio of randomly selected positively labeled genes to randomly selected unlabeled genes equal to 1:1.5, with the positively labeled genes containing only 80% of known disease-associated genes. For each balanced dataset, mantis-ml performed stratified k-fold cross-validation with out-of-bag prediction using k=10 folds. After prediction probabilities are generated for the entire gene space, mantis-ml repeats this process a total of 10 times (L) by creating new balanced datasets and performing stratified k-fold cross-validation. Finally, mantis-ml generated a ranked candidate gene list by computing the mean prediction probability and corresponding percentile score for each gene.

Validation of mantis-ml using rare variant association study summary statistics

We tested whether the top predicted monoallelic mantis-ml risk genes were enriched for genes with statistical support from recent large-scale sequencing studies. We obtained summary statistics from the largest and most recently available studies of ASD and DD6 and epileptic encephalopathy.1 All three of these association studies only included dominant models. Thus, we tested for enrichment across the three relevant monoallelic mantis-ml models. Using a two-tailed Fisher’s exact test, we conducted a comprehensive sensitivity analysis by calculating the enrichment of genes at different mantis-ml prediction thresholds (90th, 95th, and 99th percentiles) within gene sets generated across a range of significant thresholds (p < 0.01, 0.001, 1 × 10−4, 1 × 10−5, and 1 × 10−6) from each of the three association studies. Fu et al.6 did not include the X chromosome in their test, so we excluded X chromosome genes in the enrichment tests for ASD and DD. We also calculated the enrichment of genes highly intolerant to loss-of-function variation (LOEUF top 10th percentile). Due to potential concerns of circularity, we repeated these enrichment tests excluding seed genes.

Validation of mantis-ml with clinically curated gene lists

We downloaded clinically curated gene lists from SFARI26 for ASD (download date: 01/18/2022) and DDG2P27 for DD (download date: 12/16/2021). SFARI26 currently provides three tiers of confidence and DDG2P27 provides five, including definitive (our seed genes), “strong,” “moderate,” “limited,” and “relevant disease and incidental finding (RD/IF)” (RD/IF). For our analysis, we only used strong and limited, as there were too few genes with moderate” and RD/IF classifications. We then plotted the distribution of monoallelic ASD mantis-ml percentiles for tier 1 (seed genes), tier 2, and tier 3 ASD genes and compared them to the distribution of the rest of the genes in the exome not included in tiers 1, 2, and 3. We repeated the same procedures for the definitive (seed genes), strong, and limited evidence genes, stratified by monoallelic and bi-allelic inheritance with the distribution of their respective mantis-ml risk percentiles.

We then evaluated the degree of enrichment of genes in the top 10th percentile of mantis-ml predictions across each tier/category using a two-tailed Fisher’s exact test. We also calculated the enrichment for two intolerance metrics: LOEUF and pREC (for bi-allelic). For each enrichment test, we compared genes within each tier/category to genes in the rest of the protein-coding genome that were not contained in any other category.

Validation of mantis-ml using an automated literature search with AMELIE

Further validation of mantis-ml results was performed using AMELIE (automatic Mendelian literature evaluation).36 Briefly, AMELIE uses natural language processing to identify manuscripts from the extant literature with a phenotype match for genes of interest. For each manuscript with a gene-phenotype match, AMELIE reports a phenotypic match score based on the strength of the match of the language in the manuscript with the HP ontology (HPO)37 input term. A match of 100% represents a perfect match for a gene and given phenotype, and lower phenotypic match percentiles are given for related descendant phenotypes in the HPO.37

For each model, we generated genome-wide AMELIE36 phenotype match scores in a two-step process. Using the default parameters, we ran AMELIE with the HPO37 terms HP: 0000729 (autistic behavior), HP: 0001250 (seizures), and HP: 0012759 (neurodevelopmental abnormality) for ASD, DEE, and DD, respectively. We repeated this process with the inheritance mode parameter set to dominant. Although AMELIE does not permit the use of “recessive” as an inheritance mode filter, it assigns both recessive and dominant scores based on the context of an article. The dominant inheritance mode instructs AMELIE to avoid returning articles for genes with higher recessive scores. Therefore, we treated the non-union of genes between the non-specified inheritance and dominant runs as our recessive set of AMELIE scores. For each set of mantis-ml-ranked predictions, we annotated genes with their corresponding phenotypic match score and removed the seed genes from the dataset. We then used Fisher’s exact test to determine the enrichment of at least one publication with a 100% phenotypic match score in each mantis-ml decile across all models. We repeated this process using the most stringent level of evidence that AMELIE allows (five or more publications with 100% phenotypic match scores) to evaluate mantis-ml’s performance with the highest confidence gene-phenotype matches.

Results

Cell-type enrichment of NDD risk genes

We examined the expression patterns of NDD risk genes using a recently published scRNA-seq atlas of the developing human cortex.29 This dataset contains 57,868 cells collected from four human fetal cortical samples spanning 8 weeks during mid-gestation, including post-conception week (PCW)16, PCW20, PCW21, and PCW24 (Figures 1A and 1B). There are 23 annotated cell types (Figure 1A), including interneurons from the medial ganglionic eminence (MGE) and central ganglionic eminence (CGE), nine different clusters of cortical excitatory neurons (glutamatergic neurons [GluNs]), precursor cells like radial glia, and other non-neuronal cell types. One of the GluN clusters corresponds to the subplate (SP), a transient cortical structure that contains some of the earliest formed neurons of the cortex (Table S2).

Figure 1.

Figure 1

Cell-type-specific expression patterns for NDD risk genes

(A) Uniform manifold approximation projection (UMAP) plot of the human fetal cortex from data generated by Trevino et al.29 The cells are colored by cell type. RG, radial glia; CycProg, cycling progenitors; tRG, truncated radial glia; mGPC, multipotent glial progenitor cell; OPC/Oligo, oligodendrocyte progenitor cell/oligodendrocyte; nIPC, neuronal intermediate progenitor cell; GluN, glutamatergic neuron; CGE IN, caudal ganglionic eminence interneuron; MGE IN, medial ganglionic eminence interneuron; SP, subplate neurons; VLMC, vascular and leptomeningeal cells; MG, microglia; Peric., pericytes.

(B) UMAP with cell types colored by age.

(C) UMAP colored by module Z scores for NDD gene sets.

(D) Distribution of module Z scores for each cell type.

Asterisks indicate Bonferroni-corrected Mann-Whitney U p < 0.05.

To test whether NDD risk genes are preferentially expressed in any of these cell types, we carefully curated genes that have been implicated in ASD, DEE, and DD (methods). We further annotated these genes as monoallelic or bi-allelic depending on the pattern of inheritance of pathogenic mutations in each gene (methods). In total, we identified 190 monoallelic ASD genes, 94 monoallelic DEE genes, and 417 monoallelic DD genes. We also identified 17 bi-allelic ASD genes, 63 bi-allelic DEE genes, and 473 bi-allelic DD genes. We excluded bi-allelic ASD genes from downstream analyses due to the relatively small size of this gene set (Table S3).

To determine whether each of these gene sets was more highly expressed in any fetal cortical cell type, we computed module Z scores as previously described (see methods).30 A positive Z score indicates that the module of genes is expressed more highly in a particular cell than in the rest of the population (Table S4). We calculated Mann-Whitney p values for each cluster by randomly sampling 400 cells from the given cluster and comparing them to 400 random cells outside of that cluster (Table S5).

Monoallelic ASD genes were most significantly enriched in several GluN clusters, particularly those corresponding to more mature neurons (GluN4–8 and SP) as well as MGE-derived interneurons (Figures 1C and 1D; Table S4). Monoallelic DEE genes showed a similar pattern to that of ASD monoallelic genes, most strongly enriched in more mature GluN neurons (GluN6–8), SP neurons, and MGE-derived interneurons. Bi-allelic DEE genes were enriched for GluN6, GluN7, GluN8, and SP neurons, but were not significantly enriched in MGE interneurons, though we note that we were less powered for this gene set given the smaller sample size compared to monoallelic genes. Monoallelic DD genes were also generally enriched in GluN neurons but were not significantly enriched in SP excitatory neurons or MGE interneurons.

Most interestingly, DD bi-allelic genes showed a strikingly different pattern from DD monoallelic genes and were preferentially expressed in more immature cell types and non-neuronal cells, such as oligodendrocyte precursor cells (OPCs), intermediate progenitor cells (IPCs), and early and transitional radial glia (Figures 1C and 1D). Altogether, these expression patterns support the notion that monoallelic ASD, DEE, and DD risk genes converge on similar cell types. However, while prior studies have suggested that DD genes are enriched for radial glia,38 we only observe a significant enrichment for bi-allelic DD genes in this cell type.

To comprehensively characterize the biological distinctions between monoallelic and bi-allelic disease genes, we performed GO enrichment analysis on both broad disease genes from OMIM28 and our curated NDD gene sets. Analysis of these genes revealed that monoallelic disease genes were significantly enriched for terms related to signal transduction, transcriptional regulation, and protein complex assembly, while bi-allelic disease genes showed enrichment for metabolic processes and organelle organization (Tables S6A and S6B). This pattern aligns with previous observations that genes associated with monoallelic diseases typically require precise control of protein dosage and often correspond to specific biological functions such as transcriptional regulation and cellular structural components, while genes associated with bi-allelic diseases frequently encode enzymes or metabolic proteins where the loss of one copy can be tolerated.39 This dichotomy was recapitulated in our NDD gene sets, where monoallelic NDD genes were enriched for terms related to nervous system development and transcriptional regulation, while bi-allelic NDD genes showed strong enrichment for metabolic processes and cellular organization (Tables S6C and S6D). These distinct functional signatures, combined with the divergent expression patterns observed in our scRNA-seq analysis, provide strong support for considering inheritance patterns when predicting NDD risk genes.

Single-cell expression data bolster NDD risk gene predictions

Motivated by their cell-type-specific expression patterns, we hypothesized that we could leverage scRNA-seq data to predict NDD risk genes stratified by inheritance pattern. In light of the importance of metabolic genes in recessive NDDs and the observed heterogeneity of expression patterns in the developing cortex, we elected to augment our scRNA-seq data with expression data across 31 tissues from the HPA.34 We trained random forest models using the scRNA-seq data for each of the NDD risk gene sets and compared their performance to models based on conventional intolerance metrics. We trained models for each disease gene list using the risk genes as the positively labeled set and a randomly selected set of genes as the negative set (1.5 times the size of the risk gene list). To ensure robust evaluation, we generated ten balanced datasets and performed 10-fold cross-validation on each dataset. We calculated the average AUCs across all folds for each disease category.

Random forest models trained purely on single-cell expression data could accurately predict NDD risk genes for each gene list (Figures 2A–2E). For monoallelic ASD, DD, and DEE, the random forest models achieved AUC statistics of 0.87, 0.82, and 0.88, respectively. The monoallelic scRNA-seq models performed nearly as well as models trained with the LOEUF score, one of the most used loss-of-function intolerance metrics (Figures 2A–2C). Interestingly, for NDD risk genes with bi-allelic patterns of inheritance, scRNA-seq models outperformed models trained on any of the intolerance metrics (Figures 2D and 2E). The expression profiles in the cell types with the highest module scores were among the most important features for each model (Figures S1–S5).

Figure 2.

Figure 2

Random forest models incorporating single-cell expression data can predict NDD genes

Mean receiver operating characteristic curves depicting the performance of random forest models trained on various feature sets, including genic intolerance metrics, whole-body scRNA-seq data from the Human Protein Atlas (HPA),34 fetal cortex scRNA-seq data, and two aggregate models: “all scRNA-seq” (combining both scRNA-seq datasets) and “all metrics” (combining all features). Results are shown for (A) monoallelic ASD genes, (B) monoallelic DD genes, (C) monoallelic DEE genes, (D) bi-allelic DD genes, and (E) bi-allelic DEE genes. TPR indicates the true positive rate, and FPR indicates the false positive rate. The numbers in parentheses in each legend represent the mean AUC and standard deviation across the 10 folds. ASD, autism spectrum disorder; DD, developmental delay; DEE, developmental and epileptic encephalopathy; LOEUF, loss-of-function observed/expected upper bound fraction; RVIS, residual variation intolerance score; pREC, probability of being recessive; misZ, missense Z score.

We next investigated whether the expression-informed models were detecting information orthogonal to intolerance. To assess this, we built random forest models that incorporated both scRNA-seq data and intolerance metrics, including LOEUF, misZ, and the RVIS.10 We also included the pREC score, a measure of genic intolerance to bi-allelic loss-of-function variants, for the bi-allelic gene sets.20 The composite models consistently outperformed all individual models for each NDD subclass, regardless of the inheritance pattern (Figure 2). Collectively, these results suggest that both scRNA-seq data and intolerance provide independent information in detecting NDD risk genes.

Incorporation of scRNA-seq data in a semi-supervised machine learning model

One major challenge in generating genome-wide disease risk predictions is that although we have a set of known risk genes for each disease, we do not know which are definitively not associated with the disease (i.e., a true negative set). To address this, we previously introduced a stochastic semi-supervised machine learning approach called mantis-ml.25 Briefly, mantis-ml takes as input a list of seed genes (the positive set) and then trains machine learning models on random balanced datasets across the protein-coding exome. It then generates final gene rankings by averaging prediction probabilities across all the iterations. Mantis-ml includes several gene-level features, including several intolerance metrics, protein-protein interaction networks, and others.

Here, we made several advances to the mantis-ml framework. Foremost, we manually curated highly confident seed gene lists for ASD, DEE, and DD (Table S3). Given the differences in intolerance and expression profiles for monoallelic and bi-allelic gene sets, we trained inheritance-specific models. In addition, we included several new features, including scRNA-seq data and a new intolerance metric, GeVIR, which was previously shown to be more sensitive for smaller genes.35 Finally, we introduced a GO feature selection strategy, in which we performed enrichment analyses on the seed gene list to determine the GOs to include as features in each model (methods).

We trained separate mantis-ml models for monoallelic ASD, monoallelic DEE, monoallelic DD, bi-allelic DEE, and bi-allelic DD using multiple machine learning approaches, including XGBoost, random forest, support vector classifier, and gradient boosting models. Although all approaches showed comparable performances (Table S7), XGBoost demonstrated the strongest performance overall with mean AUCs of 0.94, 0.95, and 0.93 for monoallelic ASD, DEE, and DD and mean AUCs of 0.84 and 0.88 for bi-allelic DEE and DD, respectively (Figure 3B; Table S7). Given their robust performance, we selected the XGBoost-derived models for downstream analyses (Tables S8–S12). To evaluate the impact of single-cell expression data, we performed a systematic comparison of mantis-ml models with and without scRNA-seq data integration. We found that incorporating fetal cortex scRNA-seq data improved model performance compared to the base mantis-ml framework, and further improvements were observed when both fetal cortex and HPA34 scRNA-seq data were included (Table S13). These improvements were most notable in the bi-allelic DD and DEE models, suggesting that tissue and cell-type-specific expression patterns provide valuable predictive power in cases where traditional genic intolerance metrics are less informative. Using the Boruta algorithm to assess feature importance, we found that constraint metrics were consistently among the top features for monoallelic models, whereas expression data and protein-protein interaction data were relatively more important in the bi-allelic models (methods; Figures S1–S5; Table S14).

Figure 3.

Figure 3

Mantis-ml XGBoost classifier performance across five neurodevelopmental disorder models

(A) Schematic of the mantis-ml framework using bi-allelic developmental delay (DD) as an example seed gene list.

(B) Score distribution of XGBoost area under the receiver operating characteristic curves (AUCs) across all five neurodevelopmental disorder risk mantis-ml models.

Mantis-ml prioritizes top genes from rare variant association studies

We sought to evaluate mantis-ml’s ability to prioritize putative NDD risk genes using results from recent large-scale exome sequencing studies of ASD, DD, and epilepsy cohorts1,6 (Tables S15–S17). To ensure a robust assessment, we conducted a sensitivity analysis across multiple thresholds, including p values ranging from <0.01 to <1 × 10−6 and scores at the top 10th, 5th, and 1st percentiles (Table S18). This analysis demonstrated that mantis-ml consistently outperformed LOEUF, with the performance gap widening at more stringent thresholds. These results indicate that mantis-ml is particularly effective at identifying high-confidence, emerging NDD risk genes. For clarity and statistical power, we focus on comparing results at p < 0.01 and the top 10th percentile, where both mantis-ml and LOEUF balanced sensitivity and specificity most effectively. Across all three dominant models, genes in the top 10th percentile of mantis-ml were highly enriched for genes with nominal evidence of increased rare variant burden in ASD, DD, and DEE (p < 0.01) (ASD odds ratio [OR] = 13.5, 95% confidence interval [CI]: [11.2, 16.1], p = 1.1 × 10−164; DD OR = 21.6, 95% CI: [18.3, 25.4], p = 1.9 × 10−300; and DEE OR = 16.1, 95% CI: [6.1, 47.1], p = 2.4 × 10−9). These enrichments remained highly significant even after removing seed genes from the evaluation (ASD OR = 9.7, 95% CI: [7.9, 11.9], p = 1.0 × 10−94; DD OR = 13.8, 95% CI: [11.3, 16.8], p = 4.9 × 10−142; and DEE OR = 12.1, 95% CI: [3.7, 42.2], p = 1.7 × 10−5).

We compared these findings to LOEUF, which was used as a gene weight in the ASD and DD burden tests.6,7 Despite this inherent advantage, nominally significant (p < 0.01) genes from the ASD, DD, and DEE studies were less strongly enriched for genes within the top 10% of LOEUF than with mantis-ml (Figure 4). For example, in the DD study (the best powered of the three studies), top mantis-ml genes (with seed genes removed) had an OR of 13.8 (95% CI: [11.3, 16.8], p = 4.9 × 10−142) compared to an OR of 9.1 (95% CI: [7.4, 11.0], p = 6.2 × 10−96) for top-ranked LOEUF genes. Finally, we compared the performance of these monoallelic-specific models to mantis-ml models that were trained on seed gene lists that were not stratified by inheritance. The inheritance-informed monoallelic models substantially outperformed the inheritance-agnostic models for both DD and DEE, with 1.8 and 2.1 times larger point estimates, respectively (Figure S6; Tables S19 and S20).

Figure 4.

Figure 4

Enrichment of mantis-ml predictions among top genes from rare variant gene-level association studies

(A–C) The enrichment of top mantis-ml predictions (≥90th percentile) and LOF-intolerant genes (measured via LOEUF) among nominally significant (p < 0.01) genes in prior gene-level association studies for ASD (n affected individuals = 20,627), DD (n affected individuals = 31,058), and DEE (n affected individuals = 1,021), respectively. Mantis-ml models for each figure represent the monoallelic model for each respective NDD. Error bars represent 95% confidence intervals. p values calculated via two-tailed Fisher’s exact test. Bonferroni-corrected p value threshold = 2.8 × 10−4 for an alpha of 0.05.

Mantis-ml risk predictions align with the degree of confidence in clinically curated gene lists

We next tested how well the mantis-ml predictions correlated with manually curated NDD risk gene lists, including those from the SFARI database26 of ASD genes and DDG2P27 (Tables S21 and S22). In both resources, each gene receives a score reflecting the strength of evidence in the published literature of a gene’s role in the disease. SFARI26 ranks genes by tier, in which tier 1 includes high-confidence genes (n = 204), tier 2 includes “strong candidate” genes (n = 208), and tier 3 includes genes with “suggestive evidence” (n = 493). The DDG2P27 resource includes definitive (n = 218), strong (n = 156), and limited(n = 63) categories for monoallelic risk genes and definitive (n = 452), strong (n = 202), and limited (n = 98) categories for bi-allelic risk genes. Genes from tier 1 and the definitive category and a subset of monoallelic genes from the strong category (n = 55) were used as seed genes for our models, providing an opportunity to test mantis-ml’s performance on the remaining gene lists (e.g., tier 2/3 and strong/limited), which mostly consist of genes that have emerged from smaller trio- and family-based sequencing studies and functional validation.

We found that the distribution of mantis-ml percentiles correlated with the levels of evidentiary support and expert curation for both ASD and DD (Figure 5). As expected, the seed genes had the highest mantis-ml percentiles (Figures 5A–5C), which were significantly higher than the remaining genes in the exome (monoallelic ASD Mann-Whitney U [MWU] p = 3.1 × 10−99, monoallelic DD MWU p = 1.0 × 10−139, and bi-allelic DD MWU p = 8.6 × 10−181). The percentile ranks of tier 2 genes and strong genes were, on average, lower than seed genes but significantly higher than the rest of the exome (Figures 5A–5C). Likewise, the percentiles of tier 3 genes and limited DD genes were still significantly higher than the remaining genes in the exome but not as enriched as the higher-confidence gene sets (Figures 5A–5C).

Figure 5.

Figure 5

Mantis-ml performance across rare variant association studies and clinically curated gene lists

(A–C) The distribution of mantis-ml risk percentiles among clinically curated gene lists from SFARI26 and DDG2P27 compared to the rest of the exome, respectively. ASD and DD seed genes were comprised of inheritance-specific SFARI tier 1 genes and DDG2P definitive genes, respectively. p values were calculated via the Mann-Whitney U test. ∗∗∗∗Bonferroni-corrected p < 1 × 10−14.

(D) Forest plots comparing the magnitude of SFARI tier 2 and tier 3 gene enrichment in top 10th percentile of monoallelic ASD mantis-ml predictions and LOEUF rankings.

(E) Enrichment of DDG2P strong and limited gene categories in the top 10th percentile of monoallelic DD mantis-ml predictions and LOEUF rankings.

(F) DDG2P gene enrichment in top 10th percentile of bi-allelic DD mantis-ml predictions and pREC scores.

ASD, autism spectrum disorder; DD, developmental delay; SFARI, Simons Foundation Autism Research Initiative; DDG2P, Developmental Disorder Genotype-Phenotype Database.

p values in (D)–(F) were calculated via two-tailed Fisher's exact test, and error bars represent 95% confidence intervals.

We next compared the enrichment of mantis-ml predictions to intolerance metrics for these expert-curated gene lists (Table S23). Consistent with the collapsing analysis enrichment tests, tier 2 and tier 3 genes were more strongly enriched for top-ranked (top 10th percentile) mantis-ml monoallelic ASD genes than the top 10th percentile of LOEUF genes (tier 2: OR = 8.7, 95% CI: [6.5, 11.6], p = 1.1 × 10−42 versus OR = 7.0, 95% CI: [5.2, 9.4], p = 3.4 × 10−34; tier 3: OR = 4.1, 95% CI: [3.3, 5.1], p = 1.4 × 10−33 versus OR = 3.4, 95% CI: [2.7, 4.2], p = 9.4 × 10−25). Likewise, mantis-ml monoallelic DD predictions were more strongly enriched among strong and limited monoallelic genes (strong: OR = 25.0, 95% CI: [16.6, 38.1], p = 5.3 × 10−58 versus OR = 13.7, 95% CI: [9.3, 20.3], p = 1.4 × 10−38; limited: OR = 10.8, 95% CI: [6.8, 17.4], p = 3.2 × 10−21 versus OR = 7.1, 95% CI: [4.3, 11.3], p = 4.1 × 10−14). Although the confidence intervals of these enrichments overlapped, the consistently higher point estimates for the top mantis-ml genes suggest that these predictions have a stronger discriminatory ability than LOEUF alone.

We observed an even more dramatic difference in enrichments among bi-allelic DD genes. We compared our bi-allelic DD mantis-ml predictions to the pREC intolerance score,20 which aims to capture the probability a gene is intolerant to bi-allelic loss of function. The top mantis-ml genes (top 10th percentile) were strongly enriched for both strong and limited bi-allelic genes. On the other hand, only the strong bi-allelic gene list was significantly enriched for genes in the top 10th percentile of pREC (OR = 2.2, 95 CI: [1.5, 3.1], p = 1.1 × 10−4 and OR = 2.0, 95% CI: [1.1, 3.4], p = 0.02, respectively). We compared the performance of the inheritance-stratified models versus the inheritance-agnostic models. For monoallelic DD, the ORs of the monoallelic DD models were 1.8 and 2.2 times larger than those of the inheritance-agnostic DD model for the strong and limited gene lists, respectively (Figure S7). Likewise, the ORs for the DD-specific models were 2.9 times larger for bi-allelic strong DD genes (Figure S7). These results suggest that the bi-allelic DD mantis-ml model could substantially help in the discovery of bi-allelic risk genes, whose discovery typically requires access to consanguineous populations or very large sample sizes (Table S24).

Mantis-ml flags genes in clinically curated databases with limited evidentiary support

The SFARI26 and DDG2P27 databases provide clinicians and researchers with broad categories of confidence for a gene’s relevance to ASD and DD, respectively. Genes within each category are considered to have the same level of evidentiary support. We sought to evaluate mantis-ml’s ability to provide a more nuanced and quantitative measure of NDD risk within these broad, manually curated categories. For each evidentiary category (i.e., tier 2/3 in SFARI26 and strong/limited in DDG2P27), we first separated genes into high (≥90th percentile) and low (<50th percentile) mantis-ml risk prediction groups. We removed any monoallelic DD seed genes that were included in the strong/limited categories. There are no ASD seed genes in tier 2/3. We then used two orthogonal validations to corroborate mantis-ml’s predictions for each gene: publications linking a given gene to either ASD or DD and statistical support (p values) from the largest ASD/DD sequencing study to date.6 To maximize statistical power, we combined genes from tiers 2 and 3 for ASD and strong/limited for DD, respectively.

We systematically assessed whether mantis-ml predictions correlated with the degree of literature support for each gene in SFARI/DDG2P26,27 using AMELIE36 (Figure S8). AMELIE is a natural language processing tool that searches all of PubMed for manuscripts that link genes to a phenotype of interest. Importantly, AMELIE can also detect whether there is language in each article that suggests a specific pattern of inheritance, which allowed us to search gene-phenotype relationships in an inheritance-specific manner. For tier 2 and 3 genes, we found that 48.7% (110 out of 226) of high mantis-ml risk genes had ≥1 publications linking them to ASD compared to 13.3% (23 out of 173) of low mantis-ml risk genes (OR 6.2, 95% CI: [3.6, 10.8], p = 2.4 × 10−14). For the strong/limited categories, 72.1% (178 out of 247) of high mantis-ml risk genes had ≥1 publication linking them to DD compared to 62.5% (30 out of 48) of low mantis-ml risk genes (OR 1.5, 95% CI: [0.8, 3.1], p = 0.2255).

We next assessed statistical human genetics evidence support from the largest and most recent sequencing study of ASD and DD6 (Figure S9). For tier 2 and 3 genes, we found that 25.8% (55 out of 213) of high mantis-ml risk genes had nominally significant p < 0.01 compared to 0% (0 out of 167) of low mantis-ml risk genes (OR ∞, 95% CI: [14.8, ∞], p = 6.8 × 10−16). Similarly, for monoallelic strong/limited categories, 47.3% (44 out of 93) of high mantis-ml risk genes versus 0% (0 out of 13) of low mantis-ml risk genes had p < 0.01 (OR ∞, 95% CI: [2.5, ∞], p = 6.1 × 10−4). Of note, X chromosome genes were not included in the p value analysis as they were not analyzed in the Fu et al.6 study.

These data demonstrate mantis-ml’s ability to flag likely false positive genes that are included in clinically curated databases such as SFARI26 and DDG2P.27 For example, CDH15 (MIM: 114019; mantis-ml 39th percentile) is a limited gene and currently has an active gene-phenotype listing in OMIM.28 However, the evidence for this association is supported only by one publication from 2008, which lists three missense variants that were purported to be associated with severe intellectual disability (MIM: 612580).40 A curation of these variants reveals that two have been reclassified as benign in ClinVar,4 and the third is present in 18 individuals in the gnomAD20 database, which is inconsistent with a pathogenic variant for severe intellectual disability. Similarly, CD96 (mantis-ml 2nd percentile; MIM: 606037) is a limited gene with an active gene-phenotype listing for C syndrome in OMIM (MIM: 211750).28 This association is only supported by one manuscript from 2007, which identified a translocation breakpoint in CD96 in an individual with C syndrome and a missense mutation (GenBank: NM_198196.3; c.839C>T [p.Thr280Met]) in CD96 in an individual with Bohring-Opitz syndrome (BOPS [MIM: 605039]).41 However, subsequent papers have largely refuted this association, including a balanced translocation disrupting CD96 without symptoms of C syndrome,42 negative mutation screening of CD96 in individuals with C syndrome,43 phenotypically normal Cd96−/− mice,44 and the presence of the p.Thr280Met missense variant in six individuals in the non-neurologic subset of gnomAD.20 These are only two of many examples of genes flagged by mantis-ml as being unlikely to be causal for NDDs.

Manually curating databases such as SFARI,26 DECIPHER,27 and OMIM28 is a time-consuming process and prone to false positives given the vast amounts of literature and human genetics evidence that need to be reviewed for thousands of genes. Given that clinicians often look to these databases when assessing the evidence for a gene’s involvement in a disease, it is critical to ensure that the genes included in these databases are of high quality. Our data show that mantis-ml can provide an automated, immediate, and inheritance-specific assessment of the evidence for each gene’s risk for NDDs that can aid clinicians and researchers who manually curate these databases.

Mantis-ml predicts gene-phenotype relationships in published literature

Before emerging as significant in large-scale sequencing studies, genes are often initially implicated in disease through case reports with supporting functional work, case series, or family-based studies. Thus, we sought to evaluate the relationship between a gene’s predicted mantis-ml risk percentile and the number of publications linked to the phenotype of interest. We used AMELIE36 to identify the number of publications linking each gene in the genome to our three phenotypes of interest (ASD, DD, and DEE) in an inheritance-specific manner.

For each mantis-ml model, we removed seed genes and binned the remaining genes into predicted mantis-ml risk deciles. Across all five models, the top mantis-ml deciles were significantly more enriched for genes with at least one publication linking the gene to the phenotype of interest when compared to the rest of the genes in the genome (Figures 6A–6E and S9; Table S25). The enrichments were even stronger when we considered the top 1st percentile (Figure 6). There was a stepwise decrease in the strength of enrichment for each successive decile. We imposed a more stringent AMELIE36 cutoff in which we tested the enrichment of genes with at least five phenotype-matching PubMed records (the maximum allowed by AMELIE) and observed even stronger enrichments among the top deciles and 1st percentile for each model (Figures 6 and S10).

Figure 6.

Figure 6

Enrichment of genes with 100% phenotype match from the published literature stratified by mantis-ml decile

We used AMELIE36 to generate gene-phenotype match scores from the literature for all genes in an inheritance specific manner for monoallelic ASD (A), monoallelic DD (B), monoallelic DEE (C), biallelic DD (D), and biallelic DEE (E). AMELIE gene-phenotype match scores range from 0% to 100%. We limited our analysis to 100% gene-phenotype matches from the literature based on the following phenotypes: HP: 0000729 (autistic behavior) for ASD, HP: 0012759 (neurodevelopmental abnormality) for DD, and HP: 0001250 (seizures) for DEE. We then plotted the enrichment of gene-phenotype matches stratified by mantis-ml prediction deciles (and top 1st percentile) with ≥1 matching publications in orange and ≥5 in blue. p values for these comparisons are available in Tables S25 and S26. Error bars represent 95% confidence intervals.

These results further support the role of mantis-ml in discriminating putative NDD risk genes. For example, in the mantis-ml bi-allelic DD model, 9.0% (163/1,818) of genes in the top decile have at least five publications linking them to bi-allelic DD in AMELIE36 versus 0.06% (1/1817) of genes in the last decile (OR = 179.2, 95% CI: [31.5, 6804.5], p = 3.3 × 10−49). Strikingly, 22.0% of genes in the top 1st percentile of risk have five or more publications linked to bi-allelic DD. Thus, these top percentile genes are more than 502 times more likely than genes in the bottom decile to have a high-confidence association with bi-allelic DD in the literature (OR = 502.3, 95% CI: [84.8, 16384.0], p = 1.3 × 10−42). The powerful discriminatory ability of mantis-ml between the top and bottom deciles of predictions is consistent across all five disease models and inheritance patterns (ASD monoallelic: OR Inf, 95% CI: [11.7, Inf], p = 8.7 × 10−14; DEE monoallelic: OR 181.3, 95% CI: [31.9, 6878.4], p = 8.2 × 10−50; DEE bi-allelic: OR 45.3, 95% CI: [15.1, 222.7], p = 6.4 × 10−35; and DD monoallelic: OR 65.8, 95% CI: [17.8, 549.5], p = 4.7 × 10−35) (Figure 6; Table S26).

Lastly, we used AMELIE36 (≥5 publications) to compare the performance of mantis-ml models trained using inheritance-specific versus inheritance-agnostic seed gene lists for DD and DEE (Table S27). The enrichments of top decile mantis-ml hits were 1.1, 1.1, 1.6, and 4.3 times greater for monoallelic DD, monoallelic DEE, bi-allelic DD, and bi-allelic DEE, respectively, for the inheritance-informed models (Figure S11). For the bi-allelic models, these enrichments were even more striking in the top percentile of mantis-ml risk, with ORs that were 5.2 and 9.6 times higher for bi-allelic DD and bi-allelic DEE, respectively (Figure S11).

GO analysis reveals functional enrichment patterns across prediction confidence levels

To better understand the biological basis of mantis-ml’s predictions, we performed GO enrichment analysis stratified by prediction confidence deciles for each model. Across all five models, we observed enrichment patterns that aligned with our understanding of the biology of NDDs. For monoallelic models (ASD, DD, and DEE), genes in the highest confidence deciles (90th–100th percentile) were strongly enriched for terms related to nervous system development, transcriptional regulation, and synaptic function (all p < 6.7 × 10−77). For bi-allelic models, the highest confidence predictions showed strong enrichment for metabolic processes and catalytic activity (all p < 6.5 × 10−122), consistent with our earlier observation that bi-allelic NDD genes are often involved in basic cellular functions and metabolism. Notably, in the bottom decile, all five models demonstrated significant enrichment for olfactory receptor activity and chemical stimulus detection (all p < 3.2 × 10−72), suggesting appropriate de-prioritization of genes less likely to be involved in neurodevelopmental processes (Tables S28A–S28E).

Discussion

While there has been great progress in identifying hundreds of genes associated with NDDs, there remain thousands of additional risk genes to be identified. Sequencing studies will require hundreds of thousands of additional participants to fully resolve the genetic architecture of NDDs.2 Here, we used the mantis-ml semi-supervised machine learning framework to provide dominant and recessive disorder-gene risk predictions across the spectrum of NDDs. We conducted multiple orthogonal validations of mantis-ml that demonstrate its ability to prioritize both monoallelic and bi-allelic risk genes for NDDs.

Our results suggest that the monoallelic ASD, DD, and DEE models outperform intolerance metrics alone in prioritizing the top results from NDD rare variant association studies. While intolerance metrics such as LOEUF, RVIS, and others have proven extremely useful in prioritizing risk genes, they are not specific to any disease. Mantis-ml leverages multiple measures of genic intolerance, bulk and scRNA-seq data, protein-protein interaction networks, and GO annotations tailored to the specific disorder and inheritance pattern of interest. We also showed that mantis-ml predictions aligned with experts’ degree of confidence in risk genes included in curated gene lists available through SFARI26 and DDG2P.27 However, mantis-ml also flagged several genes in these databases as having a low likelihood of being risk genes for ASD or DD. We showed that genes from these databases with low mantis-ml risk percentiles (<50th percentile) for ASD/DD have significantly fewer publications and weaker supporting human genetics evidence from the largest ASD/DD sequencing studies, suggesting that they are unlikely to be true risk genes. These results suggest that mantis-ml predictions can help geneticists further prioritize disease genes in clinically curated lists and that one should reconsider the evidence for those with very low mantis-ml predictions. To this point, KATNAL2, currently a tier 1 gene, was predicted by the monoallelic ASD model to have only an 8% chance of being an ASD risk gene despite being used as a seed gene in our original analysis. Indeed, a recent re-curation of the evidence for KATNAL2 as a risk gene suggests that it is unlikely to contribute to autism risk through haploinsufficiency, and it is no longer statistically significant in the most recent and largest ASD sequencing study.6,45

We foresee several clinical applications for mantis-ml. First, mantis-ml can be used in conjunction with genomic or functional evidence to accelerate gene discovery. For example, mantis-ml can provide orthogonal evidence to prioritize genes with strong human genetics evidence that do not yet meet genome-wide significance in association studies. Second, we have shown that mantis-ml can also substantially improve the reliability and confidence of manually curated disease-gene databases by flagging likely false positive genes. Third, mantis-ml can help clinicians and researchers prioritize which genes46 to nominate for deeper functional characterization using model-organism- or cell-based approaches, as has been done in the Undiagnosed Disease Network’s Model Organism Screening Center.47 Lastly, we also envision that mantis-ml could be incorporated as gene weights in gene discovery efforts to improve power. The use of genic intolerance to inform gene priors has already led to a greater than 20% increase in ASD gene discovery power,7 and mantis-ml’s outperformance of LOEUF across rare variant association studies suggests that it will provide a significant additional boost in power.

The immediate research and clinical impact of these results are significant. First, based on our validation testing, the top 1% of predicted genes from each mantis-ml model provides a high-confidence list of hundreds of likely NDD risk genes for researchers and clinicians across the NDD spectrum. For example, depending on the model, 30%–60% of these genes already have publications linking them to phenotypes of interest, a substantial enrichment compared to the rest of the genome. Moreover, the top 1% predicted risk genes are highly enriched compared to the rest of the genome for statistical associations in recent sequencing studies of ASD, DD, and DEE. Second, mantis-ml can help clinicians solve molecular diagnoses. Mantis-ml is a highly accurate NDD risk gene predictor, particularly for genes falling in the top decile of mantis-ml predictions. If a clinician or researcher is presented with an individual with two candidate variants in genes in the top and bottom deciles of mantis-ml risk, depending on the model used, they can have roughly 45–180 times more confidence that the gene in the top decile of risk will be reliably associated with the phenotype of interest. However, we note that the interpretation of the variant effect within any given remains an important challenge in clinical interpretation.

Lastly, while there are several published measures of recessive intolerance,20,21,22 to our knowledge, there are no currently available disease-specific risk predictors for recessive disorders. The discovery of novel recessive disease genes will likely require large sample sizes or access to consanguineous and founder populations, given the rarity of homozygous or compound heterozygous pathogenic variants. Until then, mantis-ml’s bi-allelic models immediately provide a high-confidence assessment of a gene’s probability of being implicated in recessive forms of epilepsy or DD, helping clinicians and researchers solve undiagnosed cases and prioritize genes for deeper functional characterization and gene-matching strategies with other clinicians and patient cohorts. Taken together, our mantis-ml NDD models provide accurate gene risk predictions across the NDD spectrum and illustrate the importance of considering inheritance patterns in generating machine learning-based gene risk predictions.

Data and code availability

The mantis-ml predictions for all five NDD models are available as supplemental information and through a publicly available browser: https://nddgenes.com. The mantis-ml NDD code is available on GitHub (https://github.com/bweido/Mantis-ml-NDD).

Acknowledgments

We thank Dr. Huda Zoghbi for useful discussions and valuable feedback. R.S.D. is supported by grants by NIH NINDS F32 NS127854, NIH DP5 OD036131, and a Longevity Impetus Grant from the Norn Group, Hevolution Foundation, and Rosenberg Foundation. A.W.Z. is supported by K23MH121669.

Author contributions

Conceptualization, R.S.D. and J.S.D.; software, B.A.W. and D.V.; validation, R.S.D., B.A.W., J.S.D., S.P., D.V., and A.W.Z.; formal analysis, R.S.D., B.A.W., J.S.D., and D.V.; investigation, R.S.D., J.S.D., B.A.W., D.V., and A.W.Z.; resources, R.S.D. and S.P.; data curation, R.S.D., J.S.D., B.A.W., D.V., and A.W.Z.; writing – original draft, R.S.D., B.A.W., J.S.D., and A.W.Z.; writing – review & editing, R.S.D., B.A.W., J.S.D., A.J.S., C.F.S., S.P., D.V., and A.W.Z.; visualization, R.S.D., B.A.W., A.J.S., and C.F.S.; supervision, R.S.D. and S.P.; funding acquisition, R.S.D., S.P., and A.W.Z.

Declaration of interests

S.P. and D.V. are current employees and/or stockholders of AstraZeneca. R.S.D. and A.W.Z. have received consulting fees from AstraZeneca.

Published: February 26, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2025.02.001.

Contributor Information

Ryan S. Dhindsa, Email: ryan.dhindsa@bcm.edu.

Anthony W. Zoghbi, Email: anthony.zoghbi@bcm.edu.

Web resources

DDG2P, https://www.deciphergenomics.org/redirect?to=https%3A%2F%2Fwww.ebi.ac.uk%2Fgene2phenotype%2Fdownloads%2FDDG2P.csv.gz

HPA v.23, https://v23.proteinatlas.org/download/rna_single_cell_type_tissue.tsv.zip

HPO, https://hpo.jax.org

OMIM, http://www.omim.org

SFARI, https://www.deciphergenomics.org/redirect?to=https%3A%2F%2Fwww.ebi.ac.uk%2Fgene2phenotype%2Fdownloads%2FDDG2P.csv.gz

Supplemental information

Document S1. Figures S1–S12
mmc1.pdf (2.5MB, pdf)
Data S1. Tables S1–S28
mmc2.xlsx (55.8MB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (6.5MB, pdf)

References

  • 1.Feng Y.-C.A., Howrigan D.P., Abbott L.E., Tashman K., Cerrato F., Singh T., Heyne H., Byrnes A., Churchhouse C., Watts N., et al. Ultra-Rare Genetic Variation in the Epilepsies: A Whole-Exome Sequencing Study of 17,606 Individuals. Am. J. Hum. Genet. 2019;105:267–282. doi: 10.1016/j.ajhg.2019.05.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Kaplanis J., Samocha K.E., Wiel L., Zhang Z., Arvai K.J., Eberhardt R.Y., Gallone G., Lelieveld S.H., Martin H.C., McRae J.F., et al. Evidence for 28 genetic disorders discovered by combining healthcare and research data. Nature. 2020;586:757–762. doi: 10.1038/s41586-020-2832-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.McRae J.F., Clayton S., Fitzgerald T.W., Kaplanis J., Prigmore E., Rajan D., Sifrim A., Aitken S., Akawi N., Alvi M., et al. Prevalence and architecture of de novo mutations in developmental disorders. Nature. 2017;542:433–438. doi: 10.1038/nature21062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Motelow J.E., Povysil G., Dhindsa R.S., Stanley K.E., Allen A.S., Feng Y.-C.A., Howrigan D.P., Abbott L.E., Tashman K., Cerrato F., et al. Sub-genic intolerance, ClinVar, and the epilepsies: A whole-exome sequencing study of 29,165 individuals. Am. J. Hum. Genet. 2021;108:965–982. doi: 10.1016/j.ajhg.2021.04.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Epi4K Consortium. Epilepsy Phenome/Genome Project. Allen A.S., Berkovic S.F., Cossette P., Delanty N., Dlugos D., Eichler E.E., Epstein M.P., Glauser T., et al. De novo mutations in epileptic encephalopathies. Nature. 2013;501:217–221. doi: 10.1038/nature12439. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Fu J.M., Satterstrom F.K., Peng M., Brand H., Collins R.L., Dong S., Klei L., Stevens C.R., Cusick C., Babadi M., et al. Rare coding variation illuminates the allelic architecture, risk genes, cellular expression patterns and phenotypic context of autism. Nat. Genet. 2022;54:1320–1331. doi: 10.1038/s41588-022-01104-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Satterstrom F.K., Kosmicki J.A., Wang J., Breen M.S., De Rubeis S., An J.-Y., Peng M., Collins R., Grove J., Klei L., et al. Large-Scale Exome Sequencing Study Implicates Both Developmental and Functional Changes in the Neurobiology of Autism. Cell. 2020;180:568–584.e23. doi: 10.1016/j.cell.2019.12.036. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Srivastava S., Love-Nichols J.A., Dies K.A., Ledbetter D.H., Martin C.L., Chung W.K., Firth H.V., Frazier T., Hansen R.L., Prock L., et al. Meta-analysis and multidisciplinary consensus statement: exome sequencing is a first-tier clinical diagnostic test for individuals with neurodevelopmental disorders. Genet. Med. 2019;21:2413–2421. doi: 10.1038/s41436-019-0554-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Iossifov I., O’Roak B.J., Sanders S.J., Ronemus M., Krumm N., Levy D., Stessman H.A., Witherspoon K.T., Vives L., Patterson K.E., et al. The contribution of de novo coding mutations to autism spectrum disorder. Nature. 2014;515:216–221. doi: 10.1038/nature13908. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Petrovski S., Wang Q., Heinzen E.L., Allen A.S., Goldstein D.B. Genic Intolerance to Functional Variation and the Interpretation of Personal Genomes. PLoS Genet. 2013;9 doi: 10.1371/journal.pgen.1003709. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Dhindsa R.S., Copeland B.R., Mustoe A.M., Goldstein D.B. Natural Selection Shapes Codon Usage in the Human Genome. Am. J. Hum. Genet. 2020;107:83–95. doi: 10.1016/j.ajhg.2020.05.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Traynelis J., Silk M., Wang Q., Berkovic S.F., Liu L., Ascher D.B., Balding D.J., Petrovski S. Optimizing genomic medicine in epilepsy through a gene-customized approach to missense variant interpretation. Genome Res. 2017;27:1715–1729. doi: 10.1101/gr.226589.117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Samocha K.E., Robinson E.B., Sanders S.J., Stevens C., Sabo A., McGrath L.M., Kosmicki J.A., Rehnström K., Mallick S., Kirby A., et al. A framework for the interpretation of de novo mutation in human disease. Nat. Genet. 2014;46:944–950. doi: 10.1038/ng.3050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Vitsios D., Dhindsa R.S., Middleton L., Gussow A.B., Petrovski S. Prioritizing non-coding regions based on human genomic constraint and sequence context with deep learning. Nat. Commun. 2021;12:1504. doi: 10.1038/s41467-021-21790-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Palmer D.S., Howrigan D.P., Chapman S.B., Adolfsson R., Bass N., Blackwood D., Boks M.P.M., Chen C.-Y., Churchhouse C., Corvin A.P., et al. Exome sequencing in bipolar disorder identifies AKAP11 as a risk gene shared with schizophrenia. Nat. Genet. 2022;54:541–547. doi: 10.1038/s41588-022-01034-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Singh T., Poterba T., Curtis D., Akil H., Al Eissa M., Barchas J.D., Bass N., Bigdeli T.B., Breen G., Bromet E.J., et al. Rare coding variants in ten genes confer substantial risk for schizophrenia. Nature. 2022;604:509–516. doi: 10.1038/s41586-022-04556-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Zoghbi A.W., Dhindsa R.S., Goldberg T.E., Mehralizade A., Motelow J.E., Wang X., Alkelai A., Harms M.B., Lieberman J.A., Markx S., Goldstein D.B. High-impact rare genetic variants in severe schizophrenia. Proc. Natl. Acad. Sci. USA. 2021;118 doi: 10.1073/pnas.2112560118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Halvorsen M., Samuels J., Wang Y., Greenberg B.D., Fyer A.J., McCracken J.T., Geller D.A., Knowles J.A., Zoghbi A.W., Pottinger T.D., et al. Exome sequencing in obsessive–compulsive disorder reveals a burden of rare damaging coding variants. Nat. Neurosci. 2021;24:1071–1076. doi: 10.1038/s41593-021-00876-8. [DOI] [PubMed] [Google Scholar]
  • 19.Satterstrom F.K., Walters R.K., Singh T., Wigdor E.M., Lescai F., Demontis D., Kosmicki J.A., Grove J., Stevens C., Bybjerg-Grauholm J., et al. Autism spectrum disorder and attention deficit hyperactivity disorder have a similar burden of rare protein-truncating variants. Nat. Neurosci. 2019;22:1961–1965. doi: 10.1038/s41593-019-0527-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Karczewski K.J., Francioli L.C., Tiao G., Cummings B.B., Alföldi J., Wang Q., Collins R.L., Laricchia K.M., Ganna A., Birnbaum D.P., et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581:434–443. doi: 10.1038/s41586-020-2308-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Balick D.J., Jordan D.M., Sunyaev S., Do R. Overcoming constraints on the detection of recessive selection in human genes from population frequency data. Am. J. Hum. Genet. 2022;109:33–49. doi: 10.1016/j.ajhg.2021.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Hsu J.S., Kwan J.S.H., Pan Z., Garcia-Barcelo M.-M., Sham P.C., Li M. Inheritance-mode specific pathogenicity prioritization (ISPP) for human protein coding genes. Bioinformatics. 2016;32:3065–3071. doi: 10.1093/bioinformatics/btw381. [DOI] [PubMed] [Google Scholar]
  • 23.Krishnan A., Zhang R., Yao V., Theesfeld C.L., Wong A.K., Tadych A., Volfovsky N., Packer A., Lash A., Troyanskaya O.G. Genome-wide prediction and functional characterization of the genetic basis of autism spectrum disorder. Nat. Neurosci. 2016;19:1454–1462. doi: 10.1038/nn.4353. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Liu L., Lei J., Sanders S.J., Willsey A.J., Kou Y., Cicek A.E., Klei L., Lu C., He X., Li M., et al. DAWN: a framework to identify autism genes and subnetworks using gene expression and genetics. Mol. Autism. 2014;5:22. doi: 10.1186/2040-2392-5-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Vitsios D., Petrovski S. Mantis-ml: Disease-Agnostic Gene Prioritization from High-Throughput Genomic Screens by Stochastic Semi-supervised Learning. Am. J. Hum. Genet. 2020;106:659–678. doi: 10.1016/j.ajhg.2020.03.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Abrahams B.S., Arking D.E., Campbell D.B., Mefford H.C., Morrow E.M., Weiss L.A., Menashe I., Wadkins T., Banerjee-Basu S., Packer A. SFARI Gene 2.0: a community-driven knowledgebase for the autism spectrum disorders (ASDs) Mol. Autism. 2013;4:36. doi: 10.1186/2040-2392-4-36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Bragin E., Chatzimichali E.A., Wright C.F., Hurles M.E., Firth H.V., Bevan A.P., Swaminathan G.J. DECIPHER: database for the interpretation of phenotype-linked plausibly pathogenic sequence and copy-number variation. Nucl. Acids Res. 2014;42:D993–D1000. doi: 10.1093/nar/gkt937. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Hamosh A., Scott A.F., Amberger J.S., Bocchini C.A., McKusick V.A. Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders. Nucleic Acids Res. 2005;33:D514–D517. doi: 10.1093/nar/gki033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Trevino A.E., Müller F., Andersen J., Sundaram L., Kathiria A., Shcherbina A., Farh K., Chang H.Y., Pașca A.M., Kundaje A., et al. Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution. Cell. 2021;184:5053–5069.e23. doi: 10.1016/j.cell.2021.07.039. [DOI] [PubMed] [Google Scholar]
  • 30.Tirosh I., Izar B., Prakadan S.M., Wadsworth M.H., Treacy D., Trombetta J.J., Rotem A., Rodman C., Lian C., Murphy G., et al. Dissecting the multicellular ecosystem of metastatic melanoma by single-cell RNA-seq. Science. 2016;352:189–196. doi: 10.1126/science.aad0501. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Stuart T., Butler A., Hoffman P., Hafemeister C., Papalexi E., Mauck W.M., Hao Y., Stoeckius M., Smibert P., Satija R. Comprehensive Integration of Single-Cell Data. Cell. 2019;177:1888–1902.e21. doi: 10.1016/j.cell.2019.05.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Kolberg L., Raudvere U., Kuzmin I., Adler P., Vilo J., Peterson H. g:Profiler—interoperable web service for functional enrichment analysis and gene identifier mapping (2023 update) Nucleic Acids Res. 2023;51:W207–W212. doi: 10.1093/nar/gkad347. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Ashburner M., Ball C.A., Blake J.A., Botstein D., Butler H., Cherry J.M., Davis A.P., Dolinski K., Dwight S.S., Eppig J.T., et al. Gene Ontology: tool for the unification of biology. Nat. Genet. 2000;25:25–29. doi: 10.1038/75556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Karlsson M., Zhang C., Méar L., Zhong W., Digre A., Katona B., Sjöstedt E., Butler L., Odeberg J., Dusart P., et al. A single–cell type transcriptomics map of human tissues. Sci. Adv. 2021;7 doi: 10.1126/sciadv.abh2169. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Abramovs N., Brass A., Tassabehji M. GeVIR is a continuous gene-level metric that uses variant distribution patterns to prioritize disease candidate genes. Nat. Genet. 2020;52:35–39. doi: 10.1038/s41588-019-0560-2. [DOI] [PubMed] [Google Scholar]
  • 36.Birgmeier J., Haeussler M., Deisseroth C.A., Steinberg E.H., Jagadeesh K.A., Ratner A.J., Guturu H., Wenger A.M., Diekhans M.E., Stenson P.D., et al. AMELIE speeds Mendelian diagnosis by matching patient phenotype and genotype to primary literature. Sci. Transl. Med. 2020;12 doi: 10.1126/scitranslmed.aau9113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Robinson P.N., Köhler S., Bauer S., Seelow D., Horn D., Mundlos S. The Human Phenotype Ontology: A Tool for Annotating and Analyzing Human Hereditary Disease. Am. J. Hum. Genet. 2008;83:610–615. doi: 10.1016/j.ajhg.2008.09.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Polioudakis D., de la Torre-Ubieta L., Langerman J., Elkins A.G., Shi X., Stein J.L., Vuong C.K., Nichterwitz S., Gevorgian M., Opland C.K., et al. A Single-Cell Transcriptomic Atlas of Human Neocortical Development during Mid-gestation. Neuron. 2019;103:785–801.e8. doi: 10.1016/j.neuron.2019.06.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Spataro N., Rodríguez J.A., Navarro A., Bosch E. Properties of human disease genes and the role of genes linked to Mendelian disorders in complex disease aetiology. Hum. Mol. Genet. 2017;26:489–500. doi: 10.1093/hmg/ddw405. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Bhalla K., Luo Y., Buchan T., Beachem M.A., Guzauskas G.F., Ladd S., Bratcher S.J., Schroer R.J., Balsamo J., DuPont B.R., et al. Alterations in CDH15 and KIRREL3 in patients with mild to severe intellectual disability. Am. J. Hum. Genet. 2008;83:703–713. doi: 10.1016/j.ajhg.2008.10.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Kaname T., Yanagi K., Chinen Y., Makita Y., Okamoto N., Maehara H., Owan I., Kanaya F., Kubota Y., Oike Y., et al. Mutations in CD96, a member of the immunoglobulin superfamily, cause a form of the C (Opitz trigonocephaly) syndrome. Am. J. Hum. Genet. 2007;81:835–841. doi: 10.1086/522014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Darlow J.M., McKay L., Dobson M.G., Barton D.E., Winship I. On the origins of renal cell carcinoma, vesicoureteric reflux and C (Opitz trigonocephaly) syndrome: A complex puzzle revealed by the sequencing of an inherited t (2; 3) translocation. Eur. J. Hum. Genet. 2013;21:145. [Google Scholar]
  • 43.Urreizti R., Roca-Ayats N., Trepat J., Garcia-Garcia F., Aleman A., Orteschi D., Marangi G., Neri G., Opitz J.M., Dopazo J., et al. Screening of CD96 and ASXL1 in 11 patients with Opitz C or Bohring–Opitz syndromes. Am. J. Med. Genet. 2016;170:24–31. doi: 10.1002/ajmg.a.37418. [DOI] [PubMed] [Google Scholar]
  • 44.Chan C.J., Martinet L., Gilfillan S., Souza-Fonseca-Guimaraes F., Chow M.T., Town L., Ritchie D.S., Colonna M., Andrews D.M., Smyth M.J. The receptors CD96 and CD226 oppose each other in the regulation of natural killer cell functions. Nat. Immunol. 2014;15:431–438. doi: 10.1038/ni.2850. [DOI] [PubMed] [Google Scholar]
  • 45.Schaaf C.P., Betancur C., Yuen R.K.C., Parr J.R., Skuse D.H., Gallagher L., Bernier R.A., Buchanan J.A., Buxbaum J.D., Chen C.-A., et al. A framework for an evidence-based gene list relevant to autism spectrum disorder. Nat. Rev. Genet. 2020;21:367–376. doi: 10.1038/s41576-020-0231-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Sobreira N., Schiettecatte F., Valle D., Hamosh A. GeneMatcher: A Matching Tool for Connecting Investigators with an Interest in the Same Gene. Hum. Mutat. 2015;36:928–930. doi: 10.1002/humu.22844. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Baldridge D., Wangler M.F., Bowman A.N., Yamamoto S., et al. Undiagnosed Diseases Network. Schedl T., Pak S.C., Postlethwait J.H., Shin J., Solnica-Krezel L. Model organisms contribute to diagnosis and discovery in the undiagnosed diseases network: current state and a future vision. Orphanet J. Rare Dis. 2021;16:206. doi: 10.1186/s13023-021-01839-9. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1–S12
mmc1.pdf (2.5MB, pdf)
Data S1. Tables S1–S28
mmc2.xlsx (55.8MB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (6.5MB, pdf)

Data Availability Statement

The mantis-ml predictions for all five NDD models are available as supplemental information and through a publicly available browser: https://nddgenes.com. The mantis-ml NDD code is available on GitHub (https://github.com/bweido/Mantis-ml-NDD).


Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES