ABSTRACT
Background
In the diagnostic process of monogenic genetic disorders, identifying pathogenic variants is a crucial step. Thanks to the widespread adoption of Next‐Generation Sequencing (NGS) technology, diagnostic efficiency has been significantly enhanced. However, with the increasing demand for diagnostic accuracy in clinical practice for monogenic genetic diseases, accurately and swiftly pinpointing pathogenic variants among numerous candidate variants remains a significant challenge. The complexity of data analysis and interpretation continues to limit both the efficiency and accuracy of diagnosis.
Methods
In this study, we have developed an innovative phenotype‐driven algorithm, geneEX. This algorithm integrates large language model technology to accurately extract phenotypes from clinical information and automatically acquire Human Phenotype Ontology (HPO) information through a semantic vector representation model, thereby identifying HPO‐associated genes. Additionally, it supports semantic matching between patients' free‐text phenotypic descriptions and disease phenotypes, further enhancing the identification of pathogenic genes. The algorithm can rank candidate causative variants, enabling rapid and precise identification of potential pathogenic variants in rare genetic disorders.
Results
geneEX demonstrates commendable performance in ranking pathogenic variants across both virtual and clinical datasets. The supplementary matching of phenotypes in free‐text form significantly enhances the precision of candidate variant prioritization for samples.
Conclusion
geneEX has achieved automated HPO acquisition through its independently developed phenotype extraction and standardization methods, thereby enabling the full‐process automated identification from clinical samples to pathogenic variants. Additionally, by integrating free‐text phenotypic descriptions with disease phenotype matching, it enhances the accuracy of pathogenic gene identification. This innovative approach significantly improves the precision and efficiency of identifying pathogenic variants in rare genetic disorders, providing robust support for the diagnosis of monogenic diseases.
Keywords: phenotype‐driven, prioritization ranking, rare disease diagnosis
geneEX has established an automated HPO acquisition system through its proprietary phenotype extraction and standardization technology framework. This comprehensive solution enables end‐to‐end automated identification from clinical specimens to pathogenic variant interpretation.

1. Introduction
Monogenic disorders are genetic diseases controlled by a single pair of alleles, following the classic Mendelian inheritance patterns (Khera et al. 2018). According to the information from the OMIM database (https://www.omim.org/), there are currently over 8000 known monogenic diseases, which are categorized into types such as autosomal dominant inheritance, autosomal recessive inheritance, and X‐linked inheritance (Amberger et al. 2019). According to the World Health Organization (WHO) (https://www.who.int/health‐topics/congenital‐anomalies#tab=tab_1), an estimated 6% of infants worldwide are born with congenital disorders, such as hereditary deafness, thalassemia, and phenylketonuria, and these conditions can severely impact critical aspects of infant development, including heart function, language ability, intellectual growth, and behavioral coordination (Church 2017; Ferraresi et al. 2023; Jiang et al. 2023; Lichter‐Konecki and Vockley 2019). Nevertheless, effective treatments and medications are still lacking for the majority of monogenic diseases (Khan et al. 2016). For these diseases, failure to diagnose and treat them promptly after birth may lead to early infant mortality (Wright et al. 2018; Yang et al. 2020).
Since 2008, the cost of whole‐exome sequencing (WES) and whole‐genome sequencing (WGS) based on next‐generation sequencing (NGS) has significantly decreased, greatly improving the diagnostic rate of monogenic genetic disorders (Goodwin et al. 2016; Schwarze et al. 2018). However, the extraction and standardization of patient phenotypic information, along with the massive amounts of data generated by next‐generation sequencing (NGS), pose significant challenges to the diagnostic process (Yeom et al. 2019). In the routine clinical diagnostic process of WES data filtering and reporting (Figure 1a), traditional methods require manual screening and matching of 150–300 variants, which is both time‐consuming and labor‐intensive. Currently, several phenotype‐driven algorithmic tools for gene and variant prioritization have been developed, including Phen‐Gen (Javed et al. 2014), PhenIX (Zemojtel et al. 2014), Exomiser (Smedley et al. 2015), Phenolyzer (Yang et al. 2015), Xrare (Li et al. 2019), LIRICAL (Robinson et al. 2020), AMELIE (Birgmeier et al. 2020), CAVaLRi (Schuetz et al. 2024), AI‐MARRVEL (Mao et al. 2024) and GENET (Liang et al. 2024), among others. Among them, GENET is an innovative algorithm based on a large language model. It utilizes tens of thousands of positive and negative cases from publicly available data as a training dataset. The algorithm employs prompts constructed based on the reasoning chains of experienced genetic disease analysts as the logical foundation for model fine‐tuning. This approach guides the base large language model to learn and master the ability to screen for pathogenic variants in rare genetic diseases. However, recent evaluation studies have shown that phenotype‐driven algorithms based on large language models (LLMs) lack accuracy, falling behind traditional bioinformatics tools. Additionally, LLMs exhibit bias, tending to predict highly cited genes more frequently (Kim et al. 2024).
FIGURE 1.

(a) Traditional method versus the geneEX algorithm sorting approach. The four steps of WES data filtering and reporting in conventional clinical diagnosis are shown. In the traditional method, manual screening and matching of 150–300 variants are required. With the introduction of the geneEX algorithm, manual review is only necessary for the top‐ranked variants. (b) The logical flowchart of the geneEX algorithm. The geneEX algorithm consists of four main components: Disease‐phenotype database construction, phenotype scoring, genotype scoring, and variant ranking. By processing the clinical information of patients and sequencing variants, genotype and phenotype scores are calculated, and the variants are ultimately ranked based on the comprehensive score.
Although the aforementioned tools have demonstrated significant value in clinical diagnostics, most of them require the conversion of patients' clinical phenotypic information into standardized phenotypic terms, such as the HPO (Köhler et al. 2019). However, the phenotypic information presented by HPO may have certain biases and still requires manual correction and supplementation by geneticists (Pearson et al. 2021). We have developed an innovative algorithm, geneEX, which rapidly narrows down the target variant range and achieves precise identification of causative variants by independently obtaining associated genes based on phenotypic text descriptions and HPO. In summary, this algorithm enhances the diagnostic efficiency for clinicians and effectively reduces clinical diagnostic costs.
2. Materials and Methods
2.1. Dataset Construction
2.1.1. Virtual Dataset
Data sources including HGMD (Stenson et al. 2014), OMIM (Amberger et al. 2009), GPCards (Li et al. 2021), and the 1000 Genomes Project (Auton et al. 2015) were selected to construct fitting and testing datasets for algorithm fitting and evaluation. Each dataset includes variant call format (VCF) files and phenotypes, as well as diagnostic variants curated by clinical experts. All variants were annotated using VEP and Annovar (McLaren et al. 2016; Wang et al. 2010).
The details of data assembly are shown in Table 1:
TABLE 1.
Dataset assembly from HGMD, OMIM, GPCards, and the 1000 Genomes Project.
| 1000Genomes | OMIM, ClinVar | HGMD | GPCards |
|---|---|---|---|
| Used for dataset assembly | OGD (7403 cases) | HGD (9335 cases) | GGD (5182 cases) |
2.1.1.1. OMIM‐1000Genomes Dataset (OGD)
Select patient variants cataloged in OMIM that are classified as Pathogenic/Likely Pathogenic (P/LP) in ClinVar.
2.1.1.2. HGMD‐1000Genomes Dataset (HGD)
Select variants classified as Disease‐causing Mutations (DM) in HGMD.
2.1.1.3. GPCards‐1000Genomes Dataset (GGD)
Select variants of the patient whose number of variants is ≤ 2.
There is no overlap among the variants across the three datasets, and all variants have corresponding phenotypic information collected.
2.1.1.4. 1000Genomes
According to the internal variant annotation pipeline, variants called from 2504 samples were annotated and filtered, ultimately retaining 163,206 variants. The retained variants were randomly assigned zygosity states and assembled into a single‐sample VCF file, with causative variants simultaneously inserted into the sample VCF file. Each sample retained approximately 150–300 variants.
2.1.1.5. Aggregate Dataset (AGD)
The three datasets mentioned above were combined to ensure the diversity of algorithm data and the effectiveness of algorithm fitting. This dataset was divided into a fitting dataset and a testing dataset, as shown in Table 2 (The information on diagnostic variants in the testing dataset (AGD‐2) is provided in Supporting Information 1).
TABLE 2.
Fitting/testing dataset assembly.
| OGD (%) | HGD (%) | GGD (%) | Total (case) a | |
|---|---|---|---|---|
| AGD‐1 (fitting) | 90 | 90 | 90 | 19,730 |
| AGD‐2 (testing) | 10 | 10 | 10 | 1168 |
Indicates that the assembled data in Table 1 were filtered due to partial failure in HPO extraction.
2.1.2. Clinical Dataset
This dataset (n = 67) was provided by the internal clinical diagnostic laboratory (The information on diagnostic variants is provided in Supporting Information 1). This study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki and was approved by the Ethics Committee of Shanghai First Maternity and Infant Hospital.
2.2. geneEX Framework Overview
The structure of the geneEX consists of four main steps. For specific details, see Figure 1b: (1) Clinical phenotype extraction and standardization processing; (2) Disease‐phenotype database construction; (3) Optimization of phenotype association scoring; (4) Genotype–phenotype comprehensive score calculation and variant ranking.
2.3. Clinical Phenotype Extraction and Standardization Processing
Clinical initial descriptions often contain noise, and manual extraction is time‐consuming and requires extensive expertise. Although existing methods, such as PhenoTagger (Luo et al. 2021) and ClinPhen (Deisseroth et al. 2019), have been developed for rapid phenotype extraction, they still have limitations in processing complex clinical texts. Compared with traditional phenotype extraction methods, large language models (LLMs) demonstrate superior semantic understanding capabilities and can more accurately handle complex clinical texts. In this study, we propose an automated phenotype extraction method based on fine‐tuning of large language models. First, we collected patients' personal information, medical history, and family history as training data. The LLM was used to assist in annotating phenotypic content in clinical texts, which was then reviewed manually to ensure accuracy. Finally, we selected the large‐scale pre‐trained model Mistral‐Nemo‐Instruct‐2407 as the base model and fine‐tuned it on the annotated clinical phenotype dataset. The fine‐tuned LLM is capable of extracting phenotypes from complex clinical texts. This method effectively distinguishes between the phenotypes of the proband and those in the family history and performs well in handling ambiguous or overlapping phenotype descriptions (see Supporting Information 2).
In addition, we employed the SimCSE framework (Gao et al. 2022) to enhance the semantic representation of domain‐specific texts in phenotypic diagnosis. We curated paired phenotypic descriptions and their synonyms from the HPO database and annotated over a thousand paired real‐world samples (i.e., the correspondence between clinical phenotypes and HPO standard phenotypes) as training data. Using the BERT pre‐trained model, we trained the model within the SimCSE framework to learn semantic vector representations of phenotypes. This approach enables the standardization of free‐text phenotypic descriptions into professional HPO terminology through vector representation and similarity calculation.
2.4. Disease Phenotype Database Construction
To facilitate the retrieval of disease‐associated phenotypic content, we constructed a disease‐phenotype library based on disease descriptions from professional disease databases and stored the phenotypes in a vectorized format. After extracting disease phenotypes, we built a dictionary linking different phenotypes to their associated diseases. On the other hand, due to the time‐consuming nature of real‐time vectorization of large volumes of text, we used a pre‐trained semantic representation model to vectorize different phenotypes in diseases and stored the results in an ordered manner in the Faiss (Jeff et al. 2017) library. This approach enables rapid retrieval of similar vectors.
2.5. Optimization of Phenotype Scoring
Traditional phenotype‐driven identification of disease‐causing genes typically involves extracting phenotypes from the proband's clinical description, standardizing them to HPO terms, and then matching these terms to disease‐causing genes. This method relies on precise matching of HPO terms to ensure accurate identification of disease‐causing genes. However, some clinical phenotypes may not directly match standard HPO terms. To address this limitation, we employed the following two strategies in our phenotype‐driven gene scoring: (1) Matching HPO terms to HPO2Gene knowledgebase (H2GKB) to identify disease‐causing genes associated with specific HPO terms. (2) Calculating semantic similarity between free‐text phenotype descriptions and disease phenotypes to identify disease‐causing genes related to the disease.
2.6. HPO Related Phenotypic Score
Phen2Gene (Zhao et al. 2020) is a gene prioritization tool based on the HPO that provides a weighted ranking of genes for each HPO term. In this study, we obtained a standardized HPO list using the automated methods described in Sections 2.1 and 2.2 and employed Phen2Gene to compute gene scores associated with the overall HPO profile.
2.7. Free Text Related Phenotypic Score
Based on the constructed disease‐phenotype library, the semantic similarity between any given phenotype and the phenotypes in a disease can be calculated. However, to more accurately represent the association between a patient's clinical information and disease descriptions, it is necessary to compute their overall correlation as integrated entities.
Assuming that the clinical information of a patient has been processed to extract and standardize N phenotype vectors, we utilized the KNN algorithm in Faiss to identify the similar K phenotypes to these N vectors within the disease‐phenotype library. The similarity calculation is given by Equation (1), denoted as , where and represent two distinct vectors. Experimental results demonstrate that when the similarity exceeds 0.75, the two vectors exhibit a strong semantic correlation.
| (1) |
In this study, we propose a two‐dimensional ranking system for associated diseases and causative genes based on phenotype similarity calculations. The first dimension is the proportion of disease‐associated phenotypes in the clinical phenotype information (Clinical Phenotype Ratio, CPR), and the second dimension is the proportion of all phenotypes contained within the disease itself (Disease Phenotype Ratio, DPR). The calculation formulas are shown in Equations (2) and (3), where represents the number of phenotypes in disease l that are related to the patient's phenotypes, represents the similarity value of the j‐th associated phenotype in disease l, and M l represents the total number of phenotypes in disease l. Based on the CPR and DPR values, diseases are ranked with priority given to those containing a higher number of associated clinical phenotypes, followed by those with higher specificity. After disease ranking is completed, causative gene scores are calculated based on the disease‐gene correspondence in OMIM.
| (2) |
| (3) |
2.8. Genotype—Phenotype Comprehensive Score Calculation and Variant Ranking
Based on the features associated with genetic variants, the pathogenic potential of variants can be assessed. Integrating the two phenotypic association scores and the genotype score of the variant, we employed a logistic regression model to rank the patient's variants. The ranking results are presented as continuous values between 0 and 1, reflecting the comprehensive score of the patient's phenotypes and the pathogenicity of the variant. The prioritization criteria, in descending order of priority, are as follows: (1) strong correlation and pathogenicity for both phenotype and variant; (2) strong correlation for phenotype and weaker pathogenicity for variant; (3) stronger pathogenicity for variant and weaker correlation for phenotype; and (4) weak correlation and pathogenicity for both phenotype and variant.
In the implementation of the algorithm, we utilized the following six categories of numerical features as input: free‐text phenotype‐associated gene scores, HPO‐associated gene scores, variant pathogenicity scores aggregated from databases such as ClinVar and HGMD, inheritance patterns, variant in silico prediction scores, and allele frequencies in populations.
Initially, these six categories of features were normalized to facilitate better model fitting. During the data preparation phase, the data were divided into training and testing sets. In the model construction process, the training set was used to fit the model parameters, while the testing set was employed to evaluate model performance.
The model expression in this study is given by:
| (4) |
Here, denotes the probability of the association between the patient's phenotype and the pathogenicity of the i‐th variant, and x represents the features. Specifically,
| (5) |
This approach integrates genotype and phenotype features through a combined feature set, enabling the ranking of disease‐associated causative variants. This process ensures the accuracy and reliability of the ranking results, thereby providing robust support for subsequent clinical diagnosis and treatment.
2.9. Comparison With Other Phenotype‐Driven Algorithmic Tools
To compare geneEX with other algorithms, we selected four open‐source phenotype‐driven variant prioritization tools based on published benchmarking studies (Jacobsen et al. 2022; Yuan et al. 2022): LIRICAL (Robinson et al. 2020), CAVaLRi (Schuetz et al. 2024), Exomiser (Smedley et al. 2015), and Xrare (Li et al. 2019). These tools were chosen as benchmarks due to their performance, utility, and data availability, enabling rigorous evaluation against other emerging tools in the field. The aforementioned algorithms utilized ClinPhen (Deisseroth et al. 2019) to uniformly extract HPO terms from patient phenotype descriptions and input them into the analysis.
3. Results
3.1. geneEX Exhibited Comparable Performance to Existing Mainstream Algorithms on the Synthetic Dataset
We compared the performance of geneEX with existing phenotype‐driven variant prioritization methods (LIRICAL, CAVaLRi, Exomiser, and Xrare) on 1168 synthetic samples, with the test dataset comprising synthetic cases from HGMD, OMIM, GPCards, and the 1000 Genomes Project. Compared to other algorithms, Exomiser (hiPHIVE) achieved the best performance in top1 (77.74%), top10 (94.61%), and top20 (96.83%). CAVaLRi performed better in top5 (92.98%), while geneEX and Exomiser (hiPHIVE) exhibited equally excellent performance in top10 (Figure 2a). Further analysis of the synthetic dataset revealed that in OGD, CAVaLRi outperformed the other four algorithms in top1 (93.33%) and top5 (97.70%), while geneEX showed superior performance in top10 (97.01%) and top20 (97.70%) (Figure 2b). In HGD, Exomiser (hiPHIVE) achieved the best performance in top1 (86.27%) and top20 (98.98%), while CAVaLRi performed excellently in top5 (97.34%) and top10 (98.36%). geneEX, although slightly inferior to these two, also demonstrated remarkable performance (Figure 2c). For GGD, Exomiser (hiPHIVE) outperformed other algorithms across top1 to top20. Overall, regardless of the source of variant data, Exomiser (hiPHIVE), CAVaLRi, and geneEX demonstrated satisfactory performance in prioritizing causative variants, with no significant differences among them. In contrast, Xrare exhibited relatively poor performance, while LIRICAL was significantly weaker than the others. Analysis of geneEX's performance revealed its superiority in OGD. The fine‐tuned large language model, with its enhanced semantic understanding capability, can more accurately process complex clinical texts. Moreover, the free‐text phenotype supplementation based on this model can compensate for the deficiencies of HPO terms. Therefore, we believe that the rich and complex patient information in OGD is the main reason for geneEX's superior performance. In summary, compared with existing phenotype‐driven variant prioritization methods, geneEX still has certain advantages, especially in handling complex clinical information samples.
FIGURE 2.

Comparison of recall rates between geneEX and other diagnostic variant prioritization algorithms on the synthetic dataset. (a) Recall rate comparison on 1168 synthetic samples derived from HGMD, OMIM, GPCards, and the 1000 Genomes Project. (b) Recall rate comparison on 435 synthetic samples from the OGD group. (c) Recall rate comparison on 488 synthetic samples from the HGD group. (d) Recall rate comparison on 245 synthetic samples from the GGD group.
3.2. Benchmark Testing Based on Real Clinical Datasets
The large language model demonstrated enhanced semantic understanding, enabling more accurate processing of complex clinical texts. Given the complexity of phenotypic descriptions in real‐world clinical settings, this study innovatively employed a fine‐tuned large language model to extract phenotypes from complex clinical texts. Subsequently, we conducted a benchmark analysis of geneEX and the aforementioned algorithms using an internal real‐world clinical dataset (n = 67). Additionally, we performed a parallel evaluation of geneEX using HPO term‐based matching (referred to as geneEX (HPO)). The results showed that geneEX ranked first in top1 (41.79%), top5 (82.09%), top10 (95.52%), and top20 (97.01%) rankings. Although geneEX (HPO) underperformed compared to geneEX, its performance was comparable to that of CAVaLRi. Notably, geneEX (HPO) outperformed Exomiser (hiPHIVE) in top5 (76.12%), top10 (88.06%), and top20 (89.55%) rankings (Figure 3). These results highlight the positive impact of incorporating free‐text phenotypic annotations in prioritizing causative variants. In summary, the benchmark results on the real‐world clinical dataset demonstrate that geneEX's standard HPO‐based phenotype–genotype matching performs comparably to leading algorithms such as CAVaLRi, Exomiser (hiPHIVE), and Xrare, meeting clinical diagnostic needs. However, the integrated phenotype matching model combining HPO terms with free‐text annotations further enhances the precision of candidate variant prioritization.
FIGURE 3.

Comparison of recall rates between geneEX and other diagnostic variant prioritization algorithms on the clinical dataset.
4. Discussion
This study proposes geneEX, a phenotype‐driven variant prioritization algorithm. The method involves: (1) collecting patient clinical information, accurately extracting clinical phenotypes using a fine‐tuned large language model, and automatically obtaining high‐quality HPO terms through a semantic vector representation model, thereby identifying HPO‐associated pathogenic genes; (2) constructing a disease phenotype library and storing phenotypes in vectorized form to enable rapid retrieval; (3) defining a formula to evaluate the semantic association between patient phenotype descriptions and disease phenotypes, thereby calculating disease‐associated gene scores and further enhancing the identification of pathogenic genes.
The results from both synthetic and clinical datasets demonstrate that geneEX exhibits performance comparable to, if not superior to, current leading algorithms. Moreover, compared to other algorithms, geneEX's unique integrated data analysis pipeline eliminates the need for manual review of input HPO terms or reliance on external tools such as ClinPhen for phenotype extraction, thereby saving valuable time for analysts. Through this innovative prioritization algorithm, geneEX significantly improves the accuracy of candidate variant prioritization, enabling analysts to more precisely identify potential causative variants. geneEX reduces the processing time for a single sample from 1–1.5 h to 0.1–0.2 h, significantly enhancing analysis efficiency. Currently, we have developed an interactive online platform, geneEX (https://geneEX.basecare.cn), based on this algorithm, facilitating variant analysis for clinicians via a web‐based interface. In conclusion, geneEX provides a more efficient option for clinicians and researchers, facilitating early clinical diagnosis and effective medical intervention.
However, this study has several limitations. Although geneEX demonstrated robust performance on both synthetic and clinical datasets, the relatively small clinical dataset (n = 67) may not fully reflect its true performance. In future work, we plan to conduct further optimization and validation on larger clinical datasets to enhance the algorithm's prioritization accuracy. For example, the DDD dataset (Fitzgerald et al. 2015) is widely recognized as the gold standard for developmental disorder research, and several leading algorithms have been benchmarked using this dataset. We intend to utilize this dataset to evaluate the performance of geneEX in future studies. Additionally, while geneEX currently supports the analysis of single nucleotide variants (SNVs) and small insertions or deletions (indels), it cannot yet analyze certain variant types, such as copy number variations (CNVs), repeat expansions, and structural variants (SVs), which are expected to account for a portion of the remaining positive cases. On the other hand, even in the most authoritative disease databases (e.g., OMIM), many rare diseases still lack sufficient phenotypic and/or genetic evidence. However, as more rare diseases and related research are published, the comprehensiveness of these databases will continue to improve, providing additional opportunities for further optimization of geneEX.
In summary, although geneEX currently has certain limitations, we are actively addressing these challenges through continuous technological innovation and collaboration. We believe that, with ongoing technological advancements and the integration of data resources, geneEX will play an increasingly important role in the diagnosis and treatment of genetic diseases.
Author Contributions
Junyu Zhang and Dongyun Liu designed the study. Yunqian Fang, Kun Dai and Xiaoxi Zhu constructed the virtual dataset. Qingqing Xu, Meiling Hou and Li Wang analyzed the data. Mei Chen, Jianfeng Wang and Jun Zhang wrote the original version of this manuscript, and Xiaoming Teng and Bo Liang reviewed and revised it. All authors approved the submission of this manuscript.
Ethics Statement
This study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki and was approved by the Ethics Committee of Shanghai First Maternity and Infant Hospital.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Table E1. The prioritization outcomes of 1168 known causative variants within the AGD‐2 (Aggregate Dataset‐2).
TABLE E2. The prioritization outcomes of 67 known causative variants within the clinical samples.
Table S1: Example of clinical phenotype extraction.
Acknowledgments
We acknowledge all participants for their support and cooperation in this study.
Zhang, J. , Liu D., Chen M., et al. 2025. “geneEX: An Integrated Phenotype‐Driven Algorithm for Rapid Identification of Causative Variants in Monogenic Disorders.” Molecular Genetics & Genomic Medicine 13, no. 9: e70139. 10.1002/mgg3.70139.
Funding: Open Research Fund of National Health Commission Key Laboratory of Birth Defects Prevention (NHCKLBDP202514).
Junyu Zhang and Dongyun Liu are contributed equally to this work.
Contributor Information
Mei Chen, Email: mei.chen@basecare.cn.
Yunqian Fang, Email: yunqian.fang@basecare.cn.
Kun Dai, Email: daikun528@163.com.
Xiaoxi Zhu, Email: xiaoxi.zhu@biolab.org.cn.
Qingqing Xu, Email: xqq_ing@163.com.
Meiling Hou, Email: meiling.hou@basecare.cn.
Li Wang, Email: lee.wang@basecare.cn.
Jianfeng Wang, Email: jeff.wang@basecare.cn.
Jun Zhang, Email: jon.zhang@basecare.cn.
Bo Liang, Email: boliang880@sjtu.edu.cn.
Xiaoming Teng, Email: tengxiaoming@hotmail.com.
Data Availability Statement
All data generated or analyzed during this study are included in this published article (and its supporting information files). If you have any further questions, data are available from the corresponding author upon reasonable request.
References
- Amberger, J. , Bocchini C. A., Scott A. F., and Hamosh A.. 2009. “Mckusick's Online Mendelian Inheritance in Man (OMIM).” Nucleic Acids Research 37: D793–D796. 10.1093/nar/gkn665. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Amberger, J. S. , Bocchini C. A., Scott A. F., and Hamosh A.. 2019. “OMIM.org: Leveraging Knowledge Across Phenotype‐Gene Relationships.” Nucleic Acids Research 47, no. D1: D1038–d1043. 10.1093/nar/gky1151. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Auton, A. , Brooks L. D., Durbin R. M., et al. 2015. “A Global Reference for Human Genetic Variation.” Nature 526, no. 7571: 68–74. 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Birgmeier, J. , Haeussler M., Deisseroth C. A., et al. 2020. “AMELIE Speeds Mendelian Diagnosis by Matching Patient Phenotype and Genotype to Primary Literature.” Science Translational Medicine 12, no. 544: eaau9113. 10.1126/scitranslmed.aau9113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Church, G. 2017. “Compelling Reasons for Repairing Human Germlines.” New England Journal of Medicine 377, no. 20: 1909–1911. 10.1056/NEJMp1710370. [DOI] [PubMed] [Google Scholar]
- Deisseroth, C. A. , Birgmeier J., Bodle E. E., et al. 2019. “ClinPhen Extracts and Prioritizes Patient Phenotypes Directly From Medical Records to Expedite Genetic Disease Diagnosis.” Genetics in Medicine 21, no. 7: 1585–1593. 10.1038/s41436-018-0381-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ferraresi, M. , Panzieri D. L., Leoni S., Cappellini M. D., Kattamis A., and Motta I.. 2023. “Therapeutic Perspective for Children and Young Adults Living With Thalassemia and Sickle Cell Disease.” European Journal of Pediatrics 182, no. 6: 2509–2519. 10.1007/s00431-023-04900-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fitzgerald, T. W. , Gerety S. S., Jones W. D., et al. 2015. “Large‐Scale Discovery of Novel Genetic Causes of Developmental Disorders.” Nature 519, no. 7542: 223–228. 10.1038/nature14135. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gao, T. , Yao X., and Chen D.. 2022. “SimCSE: Simple Contrastive Learning of Sentence Embeddings.” arXiv preprint arXiv:2104.08821. 10.48550/arXiv.2104.08821. [DOI]
- Goodwin, S. , McPherson J. D., and McCombie W. R.. 2016. “Coming of Age: Ten Years of Next‐Generation Sequencing Technologies.” Nature Reviews. Genetics 17, no. 6: 333–351. 10.1038/nrg.2016.49. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jacobsen, J. O. B. , Kelly C., Cipriani V., et al. 2022. “Phenotype‐Driven Approaches to Enhance Variant Prioritization and Diagnosis of Rare Disease.” Human Mutation 43, no. 8: 1071–1081. 10.1002/humu.24380. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Javed, A. , Agrawal S., and Ng P. C.. 2014. “Phen‐Gen: Combining Phenotype and Genotype to Analyze Rare Disorders.” Nature Methods 11, no. 9: 935–937. 10.1038/nmeth.3046. [DOI] [PubMed] [Google Scholar]
- Jeff, J. , Matthijs D., and Hervé J.. 2017. “Billion‐Scale Similarity Search With GPUs.” arXiv Preprint arXiv:1702.08734. 10.48550/arXiv.1702.08734. [DOI]
- Jiang, L. , Wang D., He Y., and Shu Y.. 2023. “Advances in Gene Therapy Hold Promise for Treating Hereditary Hearing Loss.” Molecular Therapy 31, no. 4: 934–950. 10.1016/j.ymthe.2023.02.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Khan, F. A. , Pandupuspitasari N. S., Chun‐Jie H., et al. 2016. “CRISPR/Cas9 Therapeutics: A Cure for Cancer and Other Genetic Diseases.” Oncotarget 7, no. 32: 52541–52552. 10.18632/oncotarget.9646. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Khera, A. V. , Chaffin M., Aragam K. G., et al. 2018. “Genome‐Wide Polygenic Scores for Common Diseases Identify Individuals With Risk Equivalent to Monogenic Mutations.” Nature Genetics 50, no. 9: 1219–1224. 10.1038/s41588-018-0183-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim, J. , Wang K., Weng C., and Liu C.. 2024. “Assessing the Utility of Large Language Models for Phenotype‐Driven Gene Prioritization in the Diagnosis of Rare Genetic Disease.” American Journal of Human Genetics 111, no. 10: 2190–2202. 10.1016/j.ajhg.2024.08.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Köhler, S. , Carmody L., Vasilevsky N., et al. 2019. “Expansion of the Human Phenotype Ontology (HPO) Knowledge Base and Resources.” Nucleic Acids Research 47, no. D1: D1018–D1027. 10.1093/nar/gky1105. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, B. , Wang Z., Chen Q., et al. 2021. “GPCards: An Integrated Database of Genotype‐Phenotype Correlations in Human Genetic Diseases.” Computational and Structural Biotechnology Journal 19: 1603–1611. 10.1016/j.csbj.2021.03.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Li, Q. , Zhao K., Bustamante C. D., Ma X., and Wong W. H.. 2019. “Xrare: A Machine Learning Method Jointly Modeling Phenotypes and Genetic Evidence for Rare Disease Diagnosis.” Genetics in Medicine 21, no. 9: 2126–2134. 10.1038/s41436-019-0439-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Liang, L. , Chen Y., Wang T., et al. 2024. “Genetic Transformer: An Innovative Large Language Model Driven Approach for Rapid and Accurate Identification of Causative Variants in Rare Genetic Diseases.” medRxiv, 2024.2007.2018.24310666. 10.1101/2024.07.18.24310666. [DOI]
- Lichter‐Konecki, U. , and Vockley J.. 2019. “Phenylketonuria: Current Treatments and Future Developments.” Drugs 79, no. 5: 495–500. 10.1007/s40265-019-01079-z. [DOI] [PubMed] [Google Scholar]
- Luo, L. , Yan S., Lai P. T., et al. 2021. “PhenoTagger: A Hybrid Method for Phenotype Concept Recognition Using Human Phenotype Ontology.” Bioinformatics 37, no. 13: 1884–1890. 10.1093/bioinformatics/btab019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mao, D. , Liu C., Wang L., et al. 2024. “AI‐MARRVEL – A Knowledge‐Driven AI System for Diagnosing Mendelian Disorders.” NEJM AI 1, no. 5. 10.1056/aioa2300009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McLaren, W. , Gil L., Hunt S. E., et al. 2016. “The Ensembl Variant Effect Predictor.” Genome Biology 17, no. 1: 122. 10.1186/s13059-016-0974-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Pearson, N. M. , Stolte C., Shi K., et al. 2021. “GenomeDiver: A Platform for Phenotype‐Guided Medical Genomic Diagnosis.” Genetics in Medicine 23, no. 10: 1998–2002. 10.1038/s41436-021-01219-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Robinson, P. N. , Ravanmehr V., Jacobsen J. O. B., et al. 2020. “Interpretable Clinical Genomics With a Likelihood Ratio Paradigm.” American Journal of Human Genetics 107, no. 3: 403–417. 10.1016/j.ajhg.2020.06.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schuetz, R. J. , Antoniou A. A., Lammi G. E., et al. 2024. “CAVaLRi: An Algorithm for Rapid Identification of Diagnostic Germline Variation.” Human Mutation 2024: 6411444. 10.1155/2024/6411444. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schwarze, K. , Buchanan J., Taylor J. C., and Wordsworth S.. 2018. “Are Whole‐Exome and Whole‐Genome Sequencing Approaches Cost‐Effective? A Systematic Review of the Literature.” Genetics in Medicine 20, no. 10: 1122–1130. 10.1038/gim.2017.247. [DOI] [PubMed] [Google Scholar]
- Smedley, D. , Jacobsen J. O., Jäger M., et al. 2015. “Next‐Generation Diagnostics and Disease‐Gene Discovery With the Exomiser.” Nature Protocols 10, no. 12: 2004–2015. 10.1038/nprot.2015.124. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Stenson, P. D. , Mort M., Ball E. V., Shaw K., Phillips A., and Cooper D. N.. 2014. “The Human Gene Mutation Database: Building a Comprehensive Mutation Repository for Clinical and Molecular Genetics, Diagnostic Testing and Personalized Genomic Medicine.” Human Genetics 133, no. 1: 1–9. 10.1007/s00439-013-1358-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang, K. , Li M., and Hakonarson H.. 2010. “ANNOVAR: Functional Annotation of Genetic Variants From High‐Throughput Sequencing Data.” Nucleic Acids Research 38, no. 16: e164. 10.1093/nar/gkq603. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wright, C. F. , FitzPatrick D. R., and Firth H. V.. 2018. “Paediatric Genomics: Diagnosing Rare Disease in Children.” Nature Reviews. Genetics 19, no. 5: 253–268. 10.1038/nrg.2017.116. [DOI] [PubMed] [Google Scholar]
- Yang, H. , Robinson P. N., and Wang K.. 2015. “Phenolyzer: Phenotype‐Based Prioritization of Candidate Genes for Human Diseases.” Nature Methods 12, no. 9: 841–843. 10.1038/nmeth.3484. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yang, L. , Liu X., Li Z., et al. 2020. “Genetic Aetiology of Early Infant Deaths in a Neonatal Intensive Care Unit.” Journal of Medical Genetics 57, no. 3: 169–177. 10.1136/jmedgenet-2019-106221. [DOI] [PubMed] [Google Scholar]
- Yeom, H. , Lee Y., Ryu T., et al. 2019. “Barcode‐Free Next‐Generation Sequencing Error Validation for Ultra‐Rare Variant Detection.” Nature Communications 10, no. 1: 977. 10.1038/s41467-019-08941-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Yuan, X. , Wang J., Dai B., et al. 2022. “Evaluation of Phenotype‐Driven Gene Prioritization Methods for Mendelian Diseases.” Briefings in Bioinformatics 23, no. 2: bbac019. 10.1093/bib/bbac019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zemojtel, T. , Köhler S., Mackenroth L., et al. 2014. “Effective Diagnosis of Genetic Disease by Computational Phenotype Analysis of the Disease‐Associated Genome.” Science Translational Medicine 6, no. 252: 252ra123. 10.1126/scitranslmed.3009262. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhao, M. , Havrilla J. M., Fang L., et al. 2020. “Phen2Gene: Rapid Phenotype‐Driven Gene Prioritization for Rare Diseases.” NAR Genomics and Bioinformatics 2, no. 2: lqaa032. 10.1093/nargab/lqaa032. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Table E1. The prioritization outcomes of 1168 known causative variants within the AGD‐2 (Aggregate Dataset‐2).
TABLE E2. The prioritization outcomes of 67 known causative variants within the clinical samples.
Table S1: Example of clinical phenotype extraction.
Data Availability Statement
All data generated or analyzed during this study are included in this published article (and its supporting information files). If you have any further questions, data are available from the corresponding author upon reasonable request.
