Abstract
Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant–function–disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.
Subject terms: Literature mining, Genomics, Medical genetics, Genetic databases
PubMind, an AI tool using large language models, extracts genetic variant–disease links and supporting evidence from millions of research papers, building a searchable database of ~1.3 million variants to advance precision medicine.
Introduction
The exponential growth of biomedical research has yielded a vast corpus of knowledge on human genetic variants and their disease associations. Yet much of this information remains locked in unstructured text, inaccessible to automated systems to catalog such variants. Existing resources, such as ClinVar1 and Human Gene Mutation Database (HGMD)2, attempt to catalog variants, but they both face critical limitations. ClinVar depends on voluntary submissions, leading to heterogeneity in quality, limited coverage of published literature, and frequent absence of disease or phenotype annotations—even for variants labeled Pathogenic or Likely Pathogenic (P/LP). Nevertheless, ClinVar provides great value for understanding genetic variants and has been routinely used as a benchmark dataset. Additionally, some of its limitations, such as the lack of confidence and phenotype annotations, are gradually being addressed by community efforts such as ClinVar STAR system and ClinGen consortium3, even though such efforts are still labor intensive with limited coverage of genome. On the other hand, the HGMD directly uses literature as input and provides variant-disease association through manual review and maintenance. However, the academic (publicly available) version of HGMD is infrequently updated, lags behind the professional subscription version by years (for example, the last public version update in December 2021 vs. professional version update in April 2024), and covers only ~60% of curated entries2. Moreover, HGMD records provide only variant–disease association with a reference, yet in many cases it does not specify the nature of such association (e.g., benign vs. pathogenic), and it does not provide functional/experimental evidence, which is necessary to assist in variant interpretation4.
A significant portion of variant-related knowledge, including subtle pathogenicity assertions and mechanistic insights, remains buried within full-text articles that are rarely curated into structured databases. To overcome the challenge of variant extraction from text, conventional biomedical Named Entity Recognition (NER) systems, such as tmVar3.05, PubTator36, AIONER7, GNormPlus8, and GNorm29, have improved gene, variant and disease tagging. But they are largely limited to entity-level extraction and lack the capacity to recover relational and context-dependent information, such as variant-specific pathogenicity across different disease contexts or experimental outcomes that support these claims. Rule-based heuristics further limit adaptability, requiring constant manual updates to keep pace with evolving nomenclature.
Further effort has been made through an information aggregation system, such as LitVar10, to link literature paragraphs with variant information from different sources. LitVar is a web interface which connects variant records found in literature paragraphs, extracted by tmVar3.05, with external variant database for population information and pathogenicity such as dbSNP11. LitVar greatly accelerates the variant knowledge retrieval through providing a web interface that connects different sources of information together, trying to enrich the limited functional and pathogenic information extracted by the conventional NER systems. However, the variant information in LitVar is still limited by the extraction efficiency and accuracy of the external extraction tool and relies on the existing variant records in external datasets. LitVar does not extract variant–disease–pathogenicity associations directly from text, nor do they resolve contextual semantics or experimental evidence. For example, searching for BRCA1 p.Cys61Gly (p.C61G) mutation in LitVar2 only provided pathogenicity and allele frequency info based on the matched dbSNP record11 alongside raw PubMed text snippets. To understand which paper considers this variant as pathogenic, one needs to read the paragraph from every PubMed article. Therefore, this procedure may not be optimal for quickly summarizing the mutational landscape of a gene in a specific disease.
The advent of a new generation of artificial intelligence (AI)-based methods for unstructured data, such as Large Language Models (LLMs)12–14, presents a transformative shift in this landscape. Advances in model architecture and input token capacities, exemplified by the open-source LLaMA 3 family15, have enabled LLMs to process full-text articles with nuanced contextual understanding. Unlike conventional NER pipelines that focus on isolated entities, LLMs are capable of context-aware reasoning, leveraging few-shot learning to infer complex relationships among variants, diseases, and phenotypes—even when such associations are implied rather than explicitly stated. For example, the missense variant p.F252L in IRF6 was predicted by multiple computational tools to be benign, yet functional assays in maternal-null irf6 -/- zebrafish revealed its inability to rescue the rupture phenotype16. Sophisticated LLMs can correctly infer this variant as “likely pathogenic”, prioritizing experimental evidence over misleading computational predictions. This represents a paradigm shift from syntactic extraction toward context-aware comprehension of variant knowledge.
Recent advances with domain-adapted transformer models such as BioBERT17, PubMedBERT18, PhenoBCBERT19, and ClinicalBERT20 further underscore the potential of LLMs in biomedical text mining. These models outperform traditional tools on relation extraction tasks, particularly when fine-tuned on curated corpora or refined with human-in-the-loop strategies. However, BERT-based systems remain constrained by their classification architecture: each task requires separate fine-tuning, limiting generalizability and scalability. They also lack native support for multi-turn instruction-following and show limited compositional reasoning beyond their training domain.
To address these limitations, we developed PubMind: Publication Mutation and information discovery using LLMs. PubMind represents a novel, scalable pipeline that harnesses the generative and reasoning capabilities of GPT-style models, specifically leveraging the LLaMA 3.3 architecture, for end-to-end variant extraction from both abstracts and full-text articles. Our workflow integrates a fine-tuned BERT triage stage with downstream GPT-based inference and normalization modules for variant annotation, genome coordinate resolution, and disease/phenotype mapping to standardized ontologies such as MONDO21 and HPO22. PubMind provided enriched functional and pathogenic information of variants, which is extracted directly from the literature with evidence provided by LLM reasoning. PubMind was able to provide literature-specific annotation that LitVar2 did not provide: Searching the same BRCA1 p.Cys61Gly (p.C61G) mutation, we found 79 records across multiple publications for the same variant (PubMind Variant ID: PVID117180), so that we can consolidate pathogenicity assignments, maps diseases to MONDO terms21 to standardize disease name description, and provides human interpretable reasoning (e.g., “abolishes BRCA1 interaction with BARD1”).
Through application on 41.7 million PubMed abstracts and 5.4 M full texts, we further developed PubMind-DB, a database containing ~1.3 million unique variants with rich contextual annotations, accessible via a web interface and API. By supporting complex variant types (SNVs, gene fusions, SVs, and CNVs) and offering a continuously updated, user-accessible knowledgebase, PubMind bridges the gap between unstructured biomedical text and structured genomic interpretation, advancing variant interpretation and accelerating applications in precision medicine. We note that PubMind can be applied to institutionally licensed literature resources to build private variant databases for users. Finally, PubMind empowers healthcare systems to construct secure, institution-specific variant-interpretation databases directly from clinical notes and reports, thereby facilitating the implementation of genomic medicine at scale.
Results
An overview of the PubMind workflow
We developed PubMind, a large language model (LLM)-assisted framework for Publication Mutation and information discovery, designed to extract variant–disease–pathogenicity relationships directly from biomedical literature. We also used PubMind to process 41,682,357 PubMed abstracts and 5,425,084 full-text articles from the PMC open-access subset, and transformed unstructured text into a structured, searchable knowledgebase called PubMind-DB (Fig. 1a). At its core, PubMind integrates multiple components into a scalable, multi-layer pipeline optimized for variant interpretation.
Fig. 1. Overview of the PubMind architecture.

a PubMind integrates a multi-stage workflow for large-scale extraction of variant–disease–pathogenicity associations. Publications are first processed through a fine-tuned Bidirectional Encoder Representations from Transformers (BERT) filtering module, which reduces input size by retaining only paragraphs enriched for genetic variant information. Filtered text is then passed to an instruction-tuned LLM for extraction. The inferred variant records are subsequently validated, normalized, and stored in the PubMind database (PubMind-DB), which is accessible through a web interface for querying, visualization, and data export. b Filtering module. PubMed abstracts and PMC full-text articles are segmented into sections, and a fine-tuned BERT model selects paragraphs that contain relevant gene, variant, disease, or pathogenicity content. c Inference module. Retained paragraphs are processed with structured prompts tailored to different variant types, enabling the LLM to extract standardized attributes (gene, variant representation, disease, phenotype, pathogenicity, and reasoning). d Example association. An illustration of how a paragraph yields a structured variant–disease–pathogenicity association, including variant identifiers (rsID, cDNA, protein change), mapped disease/phenotype terms, supporting textual evidence, and pathogenicity assignment.
The workflow begins with a fine-tuned BERT model that filters and prioritizes abstracts or paragraphs enriched for mentions of genes, variants, diseases, and pathogenicity (Fig. 1b). These candidate texts are then passed through a set of carefully engineered prompts, each tailored to extract different attributes for specific variant classes—single nucleotide variants (SNVs), copy number variants (CNVs), structural variants (SVs), and gene fusions. For each input, a large instruction-tuned LLM (LLaMA3.3-70B) infers the semantic relationships and extracts variant–disease–pathogenicity associations directly from the text, uncovering biological meaning that is often buried (Fig. 1c, d). Unlike traditional entity recognition methods, PubMind does not rely on manual rule sets or model retraining, allowing it to generalize across diverse literature. Extracted results are post-processed, normalized by position (coordinate) and disease/phenotype, and stored in a relational database, where they undergo validation procedures (see Methods and Supplementary Methods). Each record is further assigned a confidence score reflecting the depth of supporting evidence extracted from literature (Supplementary Table 1). As a key advantage of generative language model, PubMind can provide reasoning while extracting every association from the text, and we made this thinking process tracible and the extraction accountable by providing “LLM reasoning” information directly in the result. The curated database is then exposed through a web-based interface that supports flexible querying, visualization, and export of results, enabling users to rapidly access literature-derived variant insights. This modular architecture makes PubMind interpretable, scalable, and adaptable to new biomedical domains, providing a foundation for comprehensive literature-based variant annotation.
Benchmarking and assessment of PubMind
To rigorously assess PubMind’s performance, we first focused on the SNV extraction task as a representative benchmark. We evaluated several LLaMA and DeepSeek-R1-series models for accuracy in variant recognition (Supplementary Fig. 1). Considering the computational efficiency and memory resource requirement, LLaMA3.3-70B demonstrated the highest precision for both cDNA- and protein-level variants and was selected as the primary inference model. Next, we compared prompting strategies—including zero-shot, chain-of-thought (CoT), and few-shot prompting (Supplementary Data 1). A short system prompt with few-shot learning achieved the best balance, yielding >90% accuracy in SNV extraction while maintaining efficient GPU runtime (Supplementary Fig. 2a, b, Supplementary Table 2). Using the same prompt, PubMind also extracted disease names with 99.9% accuracy without hallucinations (Supplementary Fig. 3c, Supplementary Data 2).
We next evaluated PubMind against established mutation NER tools across multiple benchmark corpora (Supplementary Table 3). Despite being formulated as a generative extraction framework rather than a token-level tagging model, PubMind achieved consistently competitive performance for mutation recognition. On corpus-level benchmarks that combined mutation types, PubMind reached macro F1 scores of 0.750 on OSIRIS23 and 0.830 on the SETH corpus24. On DNA- and protein-specific datasets, PubMind performed particularly well after manual normalization of semantically correct variant outputs, achieving macro F1 scores of 0.804 and 0.836 on the tmVar3 DNA and protein benchmarks5, and 0.847 and 0.890 on the Wei2013 (tmVar1.0) DNA and protein benchmarks25, respectively. In the Wei2013 benchmarks, PubMind slightly surpassed tmVar3 in macro F1, highlighting that a generative framework can attain NER-level extraction accuracy while offering richer, structured outputs for downstream interpretation. Because conventional NER tools are optimized for span tagging rather than normalized structured extraction, these comparisons likely underestimate the practical utility of PubMind.
As PubMind assigns pathogenicity for every extracted variant from millions of papers, we were able to consolidate the pathogenicity assignments from multiple records and provided an aggregated pathogenicity score and assignment for each variant (see Method). Pathogenicity assignment in PubMind is inferred directly from the literature context and may not fully align with ClinVar classifications. Thus, we performed benchmarking against ClinVar to assess concordance and examine reasons when differences arise. Across all PubMind-DB SNVs, there are 1,017,540 literature-extracted variant records, which are consolidated into 579,844 unique variants. From the consolidated unique variants, 419,012 unique variants were mapped to transcripts; after accounting for multiple codon-level representations, this resulted in 916,538 transcript-mapped genomic variants (VCF format), which were annotated using ANNOVAR26. The ANNOVAR will provide ClinVar pathogenicity information and bioinformatics software predictions based on the genome coordinates of these 916,538 variants. Among these, 97,214 genomic variants (10.6%) overlapped with ClinVar records (Fig. 2a). The relatively small overlap highlighted the advantage of PubMind in cataloging literature-derived variants, while the observed concordance patterns illustrated its reliability. Agreement with ClinVar was the highest for pathogenic variants, but lower for benign classifications—consistent with known variability in annotating benign variants between different tools. To refine this comparison, we stratified concordance by ClinVar expert review status (“Star” system), which reflects confidence in classification (Fig. 2b). Agreement between PubMind and ClinVar increased with review depth, reaching 100% concordance for four-star variants. Importantly, PubMind confidence scores also tracked with ClinVar Stars, indicating that variants curated more thoroughly in ClinVar were also classified with higher confidence by PubMind. This trend underscored the robustness of PubMind’s pathogenicity assignments, while suggesting that ClinVar 0-star entries—often heterogeneous in quality—account for most discrepancies.
Fig. 2. Benchmarking PubMind pathogenicity assignments against ClinVar and bioinformatics tools.

a Concordance with ClinVar. Heatmap showing the overlap between PubMind-derived pathogenicity classifications and ClinVar labels (n = 97,214). Agreement is highest for pathogenic variants, while benign annotations show greater heterogeneity. b Stratification by ClinVar review status. Agreement between PubMind and ClinVar increases with review depth, reaching 100% concordance for four-star variants. Stacked bars indicate PubMind confidence scores, which also rise with higher ClinVar review status. c Score correlation. Comparison of PubMind-derived pathogenicity scores with the average rank score from 51 computational tools. Variant density is represented by a continuous heatmap, with darker colors indicating regions of higher concentration. Marginal histograms show the distribution of scores for PubMind (top) and computational tools (right). Variants predominantly cluster at high pathogenicity values, though PubMind additionally identifies low-scoring benign variants that are less represented by computational tools. d Tool-based categorical predictions. Average predictions from 23 categorical bioinformatics tools across different PubMind pathogenicity categories. Variants classified as pathogenic/likely pathogenic by PubMind align with higher deleterious predictions, whereas benign categories align with fewer pathogenic calls. Most tool predictions fall into the “unknown” category.
We next compared LLM-derived pathogenicity scores with consensus predictions from diverse bioinformatics tools using a density plot (Fig. 2c). This tool-based score is the average pathogenicity score (rank from 0 to 1) from 51 bioinformatics tools' predictions using dbNSFP v4.727. PubMind-derived scores ranged from 0 (benign) to 1 (pathogenic) with peaks around 0.75 and 1.0, while tool-based scores displayed a distribution skewed toward ~0.8. Both methods showed strong agreement in the high-pathogenicity range, with dense clustering in the top-right quadrant of the plot. However, PubMind additionally captured low-scoring peaks corresponding to benign calls in literature, whereas bioinformatics tools tended to yield heterogeneous or high scores. This suggests that PubMind provides contextual evidence not consistently captured by computational predictors. Finally, we compared categorical predictions from 23 bioinformatics tools in dbNSFP against PubMind classifications (Fig. 2d). As expected, variants labeled as pathogenic/likely pathogenic by PubMind showed a higher proportion of deleterious predictions, while benign/likely benign variants aligned with lower proportion of deleterious calls by bioinformatics tools. However, the majority of output from bioinformatics tools were “unknown,” reflecting the absence of confident classifications. PubMind therefore fills an important gap by extracting literature-reported pathogenicity classifications for variants that most bioinformatics tools leave unclassified, while also showing concordance with tool-based predictions when those tools can provide informative calls.
To assess PubMind’s ability to extract complete variant-level associations rather than isolated entities, we conducted a controlled case-study analysis using three PMC full-text articles published in 202628–30. Ground-truth annotations were manually curated for each paper, including gene, DNA variant, protein variant, disease, and pathogenicity labels (Supplementary Data 3 and 4). Using an identical SNV prompt, we compared paragraph-level and full-text inputs, with or without finetuned-BERT filtering. Using individual paragraph as input had no detectable effect on retrieval sensitivity, with both filtered and unfiltered paragraph settings recovering all ground-truth evidence from the source text. In contrast, full-text input had an influence on extraction behavior. Paragraph-level input identified more variant associations overall than full-text input (10 versus 3 across the three papers), capturing not only the principal reported variants but also additional variants cited elsewhere in the text. However, this increased sensitivity was accompanied by weaker global association linking, particularly in connecting cDNA and protein-level descriptions of the same variant (Supplementary Data 4, red highlights). Repeated mentions of the same variant across multiple paragraphs within a paper could introduce redundant calls and thereby inflate aggregate pathogenicity estimates. To assess this effect, we also derived a paper-level pathogenicity score by first computing pathogenicity independently for each paper and then averaging across papers. The resulting scores were highly concordant with the original paragraph-level pathogenicity scores (Pearson r = 0.997, Spearman ρ = 0.997; Supplementary Fig. 4), suggesting that redundancy from repeated mentions does not substantially alter overall pathogenicity assessment. We retained the paragraph-level score as the main metric because it was used consistently across all downstream analyses. These findings suggest that paragraph-level prompting is advantageous for sensitive association discovery, whereas full-text prompting provides stronger contextual integration for normalized variant representation.
Exploring the PubMind-DB knowledgebase
PubMind-DB encompasses four major variant classes—single nucleotide variants (SNVs), gene fusions, structural variants (SVs), and copy number variants (CNVs)—with standardized identifiers (PVID) and literature-derived annotations (Table 1). SNVs are indexed using gene names with corresponding cDNA, rsID, or amino acid changes, yielding 579,844 consolidated unique SNVs, of which 419,012 have genomic coordinates. Gene fusions are defined by their fusion partners and affected protein domains, consolidating into 85,966 fusions from 174,784 literature-derived records. SVs and CNVs are indexed by gene names, chromosomal regions, and coordinates where available, with literature-based disease and pathogenicity associations extracted for each. In total, PubMind catalogs 69,993 unique SVs consolidated from 191,643 literature-extracted SV records, and 29,457 unique CNVs consolidated from 89,953 literature-extracted CNV records, creating a unified resource across diverse variant types.
Table 1.
Overview of variant records in PubMind-DB
| Variant type | Literature-derived records | Consolidated unique variants (unique PVIDs) | Genomic/transcript-mapped variants (VCF-format) | ClinVar overlap (n, %) |
|---|---|---|---|---|
| single-nucleotide variant (SNV) | 1,017,540 | 579,844 | 916,538 | 97,214 (10.6%) |
| Gene fusion | 174,784 | 85,966 | – | – |
| Structural variant (SV) | 191,643 | 69,993 | – | – |
| Copy number variant (CNV) | 89,953 | 29,457 | – | – |
| Total | 1,473,920 | 765,260 | 916,538 | 97,214 (10.6%) |
Summary of variant classes extracted from PubMed abstracts and full-texts, after multi-stage processing (literature-derived records, consolidated unique variants, transcript/genomic mapping). ClinVar overlap is reported for transcript-mapped SNVs, based on ClinVar version number 20240914.
Figure 3 provides an overview of the pathogenicity label, confidence score, and disease associations represented in PubMind-DB. As shown in Fig. 3a, pathogenic and likely pathogenic variants dominate across all variant types, with relatively fewer benign or likely benign annotations. As expected, the number of variants decreases with increasing confidence level (Fig. 3b). To explore disease context, we identified the top 10 most frequently mentioned diseases for each variant class, stratified by pathogenicity (Fig. 3c). These distributions reveal that most variants are annotated as pathogenic or of uncertain significance (“Other”), while benign classifications are comparatively rare in the published records.
Fig. 3. Pathogenicity distribution, PubMind confidence scores, and disease associations across variant types.

a Pathogenicity classification. Distribution of PubMind-assigned pathogenicity labels for single nucleotide variants (SNVs), gene fusions, structural variants (SVs), and copy number variants (CNVs). Pathogenic and likely pathogenic annotations dominate across all variant types, whereas benign classifications are comparatively rare. b Confidence scores. Variants stratified by PubMind confidence score (0–3), reflecting evidence depth and annotation reliability. Most variants are supported at lower confidence levels, with progressively fewer variants at higher tiers. c Top 10 disease associations. The ten most frequently mentioned diseases for each variant type, with bars colored by PubMind-assigned pathogenicity category (Pathogenic/Likely Pathogenic, Benign/Likely Benign, or Other [unknown/conflicting]). The majority of top disease associations are driven by pathogenic or uncertain classifications.
We further examined the literature sources underlying PubMind annotations (Supplementary Fig. 5). Most variants were extracted from sections classified as “Other” in the XML structure, reflecting text outside standard categories such as Abstract, Introduction, or Result (Supplementary Fig. 5a). Across variant types, the temporal distribution of publications peaked around 2021, consistent with the rapid growth of sequencing and genomics literature in recent years (Supplementary Fig. 5b). Finally, we identified the top contributing journals for each variant class, with PLoS One and Scientific Reports consistently ranking among the largest sources of variant records (Supplementary Fig. 5c).
Together, these results illustrate the breadth and richness of the PubMind database, spanning millions of literature-mined variant annotations with structured pathogenicity and disease associations. By integrating multiple variant classes, a confidence scoring framework, and contextual annotations, PubMind provides a comprehensive, literature-grounded complement to human curated databases, enabling researchers to explore variant knowledge at an unprecedented scale.
Case studies: from variant reclassification to therapeutic context with PubMind-DB
To illustrate use cases of literature-driven pathogenicity annotation, we performed a case study on variants of uncertain significance (VUS) in the IRF6 gene, a key locus implicated in orofacial cleft. In prior work, Edward et al. experimentally evaluated 37 IRF6 missense mutations using a phenotype rescue assay in irf6 -/- zebrafish16. Here, we focused on 9 missense mutations reported in full-text literature16 and compared their pathogenicity assignments across multiple sources: ClinVar1, HGMD2, AlphaMissense31, an ensemble of 21 bioinformatics tools32, and PubMind (Fig. 4).
Fig. 4. Case study of IRF6 Variant of Uncertain Significance (VUS) reclassification using PubMind.

a Comparison of pathogenicity assignments. Pathogenicity calls for nine IRF6 missense variants are shown across multiple sources, including ClinVar, HGMD, AlphaMissense, an ensemble of 21 bioinformatics tools, and PubMind (single-paper and aggregated annotations). Results are benchmarked against functional rescue assays in irf6 -/- zebrafish (PMC5628943). Text colors indicate: dark blue = exact match with functional assay result; light blue = agreement without conflict (excluding exact match); dark orange = mismatch with functional assay result; light orange = no information available (N/A, variant record does not exist) or unknown (variant record exists but pathogenicity is unknown). The last two columns are the pathogenicity score calculated using the aggregated pathogenicity and the number of papers used for aggregation. b Visualization of the comparison in panel (a). Stacked bars show the number of variants with exact matches (dark blue), agreements without conflict (light blue), mismatches (dark orange), or missing information / unknown (light orange). Exact match and agreement without conflict are shown above the x-axis and are considered as concordant, while mismatch or missing information/unknown are shown below the bar and are considered not concordant. PubMind-aggregated annotations achieved complete concordance (9/9), outperforming ClinVar, HGMD, AlphaMissense, and predictions from bioinformatics tools.
HGMD provided only variant–disease associations without explicit pathogenicity calls; thus, pathogenicity was unavailable (N/A) for these entries, even though 8 of the 9 variants were indexed. For PubMind, we examined both the annotations derived from this exact zebrafish study (PMCID: PMC562894316) and the aggregated annotations across multiple publications. Among the nine variants, eight were functionally rescued in vivo, indicating a benign classification. To assess concordance between pathogenicity annotations from different resources and the functional assay results, we classified each comparison into four categories shown in Fig. 4: exact match, agreement without conflict (for example, benign vs. likely benign or benign vs. conflicting; both considered concordant), mismatch, and missing or unknown information. Exact match and agreement were considered concordant, whereas mismatch and missing/unknown annotations were considered non-concordant. Across all evaluated resources, PubMind was the only resource that achieved complete concordance (9/9, 3 exact matches and 6 agreements) when aggregating evidence across publications, correctly identifying all rescued variants as benign or non-pathogenic. Notably, when PubMind was restricted to the zebrafish-based functional study alone (PMC5628943), it produced more exact matches (5 of out 9) but assigned 4 variants as “unknown” pathogenicity, despite successfully identifying those variants in the text. Aggregating evidence across publications resolved four variants that were “unknown” when relying on a single paper (PMC5628943) alone. Together, these findings illustrate a trade-off between single-study specificity and multi-study evidence integration: single-paper extraction preserves source-specific pathogenicity assignments, whereas aggregated PubMind annotations improve classification completeness. In contrast, AlphaMissense31 correctly classified only one of the eight benign variants, and the 21-tool ensemble approach failed to classify any as benign.
Our second use case focused on PDGFRB, a receptor tyrosine kinase whose activating variants are well-established drivers of infantile myofibromatosis and other disorders, and for which targeted inhibitors such as imatinib are clinically effective33,34. In a prior collaborative study, we attempted to search published case studies on PDGFRB variants with treatment by imatinib, and manually identified seven PDGFRB variants through this exercise35. Using PubMind-DB, we systematically queried all reported PDGFRB variants and cross-referenced the “LLM reasoning” column for therapeutic context by searching for the occurrence of the keyword “imatinib” in the LLM reasoning output provided by PubMind. All four SNVs previously identified were recovered by PubMind-DB, along with additional clinically relevant variants such as Arg370Cys (PVID702484) and Asn666Ser (PVID702494). Importantly, for two recurrent SNVs (PVID702492 and PVID702578), the LLM reasoning explicitly captured therapeutic evidence, e.g., “clinical improvement when treated with imatinib, an inhibitor of several kinases” and “response to imatinib treatment”. This highlights PubMind-DB’s ability to extract not only variant–disease–pathogenicity associations, but also actionable treatment insights directly from literature.
Beyond SNVs, PubMind-DB proved especially powerful in finding gene fusions and SVs with therapeutic relevance. Across 339 PDGFRB-associated fusions found by PubMind-DB, 45 contained reasoning that directly mentioned imatinib. Examples include PDGFRB::EBF1, PDGFRB::ETV6, and COL1A1::PDGFRB, each linked by LLM reasoning to imatinib responsiveness. For instance, the entry for COL1A1::PDGFRB (PVID-GF17305) noted: “The presence of the COL1A1–PDGFRB fusion leads to activation of the PDGFRB tyrosine kinase, which can be targeted by specific inhibitors such as imatinib.” Similarly, 5 of 31 SVs in PDGFRB were annotated with imatinib in their reasoning fields. Together, these findings demonstrate how PubMind-DB can consolidate dispersed evidence into a single knowledge base, enabling rapid retrieval of both pathogenic variants and associated therapeutic opportunities.
Discussion
In this work, we address the challenge in extracting variant–disease–pathogenicity associations directly from biomedical literature by leveraging the power of large language models (LLMs). Through systematic benchmarking of models, prompting strategies, and normalization pipelines, we demonstrated that a 70B instruction-tuned model, used without finetuning, can achieve high performance in extracting genetic variants and their functional annotations. With few-shot prompting, PubMind achieves >90% accuracy in variant extraction with minimal hallucination, while capturing disease terms with 99.9% precision. Importantly, the resulting annotations show strong concordance with curated resources such as ClinVar1 with 100% agreement on ClinVar’s expert-reviewed variants (4 stars), yet provide substantial new coverage, with only ~10% overlap with ClinVar entries. By processing 41 million PubMed abstracts and 5.4 million PMC full-text articles, PubMind generated one of the most comprehensive publicly available literature-derived knowledge bases to date, spanning SNVs, gene fusions, SVs, and CNVs, each annotated with disease associations, pathogenicity, reasoning for pathogenicity, and source references.
Developing PubMind required addressing several key challenges: (1) accuracy of extraction, (2) efficiency of large-scale inference, (3) normalization of unstructured outputs, and (4) comparison to find undocumented variants relative to existing databases. Below, we discuss each of these challenges.
Accuracy
Benchmarking across multiple LLaMA15 and DeepSeek-R1 models36 confirmed that few-shot prompting and CoT reasoning provide high extraction accuracy. Cases deemed “errors” in PubMind were often formatting differences rather than true mistakes (e.g., “ΔG91” vs. “G91del”). Unlike conventional NLP approaches such as tmVar35 or PubTator36, which classify tokens, LLMs generate semantically consistent reformulations—often yielding clearer or standardized representations. Notably, even the base instruction model (LLaMA3.3-70B15) performed well without finetuning, highlighting the potential for transferability to other biomedical information extraction tasks (e.g., drug–drug interactions, protein–protein interactions).
Efficiency
Running LLM inference on millions of full texts is computationally infeasible. For example, processing just 500 PMC articles required ~3.5 hours on 4 Nvidia A100 GPUs. To overcome this, we developed and evaluated two different filtering strategies: (i) regex-based filtering to reduce input volume based on key words and predefined rules and (ii) fine-tuned BERT models to retain semantically relevant paragraphs (Supplementary Table 4). To generate the finetuning dataset for BERT-model, we used a hybrid approach, where LLMs generate training labels for a smaller BERT classifier, enabled efficient filtering while preserving high-value content. Compared with rule-based regex filtering, fine-tuned BERT retained fewer paragraphs yet yielded more extracted variants, underscoring its semantic precision and advantage in terms of turnaround time to process the entire PMC database in 5 days (Supplementary Table 5).
Normalization
Because LLM outputs are inherently unstructured, we developed a post-processing pipeline to standardize variant, disease, phenotype, and pathogenicity annotations. Variants were parsed using regex-based normalization; diseases and phenotypes were mapped to MONDO21 and HPO terms22 via PubMedBERT embeddings37; and pathogenicity assignments were consolidated across multiple records. Inspired by ClinVar1, we further implemented a tiered confidence system to score annotation reliability, enabling transparent downstream use.
Complementary role to existing databases
As a primary application of PubMind, we processed over 41 million PubMed abstracts and over 5.6 million full texts to build a variant-disease-pathogenicity knowledgebase enriched with contextual information, PubMind-DB. Comparison with ClinVar1 confirmed that PubMind-DB captures both reliable concordant variants and a large body of previously uncurated knowledge, with ~90% of variants absent from ClinVar. Agreement with ClinVar improved with review depth (reaching 100% for four-star variants), supporting the validity of PubMind’s assignments. We noticed there was high agreement with ClinVar 0-star variants, which represent the variants that do not have related literature or documents provided by the submitter. However, the majority of the agreement is contributed by the PubMind variants with a confidence score less than or equal to 1 (with 78.2% agreement). If we focus on PubMind variants with a confidence score of 2 or 3, the ClinVar 0-star variants have lower agreement (18.7%) compared to ClinVar variants with more than 1 star (24.8% in 2-stars, 30.3% in 3-stars, and 36.4% in 4-stars). The result showed a high agreement within the low-confidence portions between PubMind and ClinVar, suggesting a need for further evaluation of 0-star ClinVar variants as well as the importance of PubMind confidence in downstream use. Cross comparisons with 51 computational predictors revealed strong correlation in highly pathogenic regions, while PubMind uniquely identified benign calls aligned with functional evidence but overlooked by tools. A case study on IRF6 VUS16 highlighted PubMind’s ability to aggregate dispersed evidence across publications, achieving complete concordance with experimental zebrafish rescue assays—whereas AlphaMissense31 and ensemble predictors32 failed to identify most benign variants. The case study also underscores PubMind’s clinical relevance and ability to reclassify VUS by integrating distributed functional evidence, compared to the HGMD2 which is unable to provide the pathogenicity for these variants (or can only be considered as pathogenic association). A second case study on PDGFRB in infantile myofibromatosis33,35,38–40 further illustrated how PubMind-DB not only retrieves variant–disease associations across scattered publications but also surfaces therapeutic insights, such as imatinib response, directly from the literature. Together, these case studies demonstrate PubMind’s ability to integrate dispersed evidence across the literature to accurately reclassify variants of uncertain significance and to contextualize rare disease variants with clinically actionable insights, complementing existing databases and computational predictors.
Despite these strengths, some limitations remain. First, due to computational constraints of large models (e.g., requiring multiple H100/A100 GPUs), PubMind processes abstracts and paragraphs independently rather than entire full texts. While this reduces input size and improves efficiency, it may miss cross-paragraph associations within a single study. In a case study of three PMC full-text articles published in 2026, we evaluated the effect of paragraph-level versus full-text input on extraction performance (Supplementary Tables 4 and 5). Paragraph-level prompting improved sensitivity to localized evidence, but at the expense of cross-paragraph integration when related entities were mentioned separately in different parts of the article, for example, when a cDNA change and the corresponding protein change were not co-localized. In such cases, PubMind could produce multiple incomplete records for a single underlying variant–disease–pathogenicity association. Conversely, when applied to full-text input, the same SNV prompt showed reduced ability to recover cited variants and other non-primary findings, likely because it had originally been optimized for paragraph-level extraction. Together, these observations indicate that input preparation influences LLM-based information extraction and that additional prompt engineering may need to be adapted when moving between paragraph-scale and document-scale contexts. Second, as with any generative model, outputs may deviate from the requested format; extensive normalization mitigates this but does not eliminate rare inconsistencies (e.g., cDNA vs. protein-level misassignments). Despite this limitation, hallucinated outputs were uncommon, and extraction performance remained consistently strong. Because both the biomedical literature and underlying language models are evolving rapidly, streamlined filtering and normalization pipelines will be important for minimizing update latency and enabling the routine maintenance of periodically refreshed resources such as PubMind-DB.
In contrast to HGMD2, which provides only variant–disease associations, and ClinVar1, which depends on manual submissions and variable curation depth, PubMind offers scalable, literature-grounded extraction of explicit variant–disease–pathogenicity relationships. The pipeline can process the entire PubMed and PMC Open Access corpus in 1-2 weeks with 4 Nvidia A100 GPUs, delivering up-to-date annotations with direct links to source text. Beyond variant interpretation, the framework is generalizable to other biomedical entity–relationship extraction tasks, from pharmacogenomics to pathway annotation. Additionally, we evaluated the applicability of PubMind to in-house clinical notes for patients with known positive genetic diagnoses. In this setting, the PubMind system successfully extracted variants along with associated diseases and LLM-derived reasoning (e.g., “the gene panel confirmed the variant’s pathogenicity”). PubMind was able to efficiently filter through dozens of clinical notes for a single patient, retaining only those relevant for variant extraction. The resulting structured outputs provide a rapid, high-level index of patient data, offering a valuable foundation for clinical data management and downstream analysis.
In summary, PubMind represents a powerful and novel bridge between unstructured biomedical text and structured genomic resources. By integrating advanced LLM reasoning with efficient filtering and rigorous normalization, PubMind delivers both scale and interpretability, complementing curated databases and computational predictors. We anticipate that PubMind and PubMind-DB will serve as valuable pipelines and resources for human variant interpretation, evidence-based variant reclassification, rare disease gene discovery, and precision medicine applications.
Methods
Data sources and corpus
For the establishment of PubMind-DB, 41,682,357 PubMed Abstracts and 5,425,084 PubMed Central (PMC) full texts from OA subset have been downloaded (up to 2025.02.07) through the NCBI’s ftp depository (https://pubmed.ncbi.nlm.nih.gov/download/, https://pmc.ncbi.nlm.nih.gov/tools/ftp/). XML files were parsed using the Python package BeautifulSoup4 to extract PubMed abstracts or PMC full texts.
Input filtering for PubMed abstracts and PMC full texts
We prepared different databases for finetuning and benchmarked different BERT models for the task of input filtering (Supplementary Table 4). To filter the enormous literature input, DistilBERT model41 was finetuned using 1.5 K labelled abstracts. The finetuned DistilBERT classified each paragraph according to whether it contained variant information. The 1.5 K labeled dataset consisted of ~1000 negative examples generated using regex rules and 500 positive examples from the PubTator3.0 training corpus6. Additional BERT models have been used and tested, including GoogleBERT13, BioMedBERT42. Larger finetuning datasets of 15 K abstracts were generated and evaluated using two labeling strategies: regex-based labels (regex-label) and LLM-derived labels (LLM-label). For LLM-labels, we first ran 100,000 PubMed abstracts through the LLM; abstracts producing useful variant outputs were labeled positive, while those without were labeled negative. See Supplementary Methods for details of finetuning dataset construction.
Model selection and prompt engineering for LLM inference
To evaluate whether prompt engineering with base LLM models (without finetuning) could enable genetic variant extraction, we experimented different prompts and compared the results with benchmark dataset based on 500 PubMed abstracts with labelled variant information6. Different versions of LLaMA family15 and DeepSeek-R1 models36 were evaluated (Supplementary Fig. 1). After choosing LLaMA3.3-70B instruct model as our main model, different prompts were evaluated based on the accuracy of variant extraction and the GPU runtime (Supplementary Fig. 2). For details about all the prompts we used, please refer to Supplementary Data 1.
Variant and disease extraction benchmark
To assess reliability and accuracy, we benchmarked the LLM’s ability to extract both variants and diseases. The benchmark dataset for variant NER is the 500 PubMed abstracts with labelled variant from PubTator36. The benchmark dataset for disease NER is from NCBI disease corpus43 used in PhenoTagger44, which contains the PubMed abstract as well as the disease name found in the text. Hallucination rates were quantified by comparing all LLM outputs against these benchmark datasets (Supplementary Table 2and Supplementary Data).
Normalization and database generation
The gene names were filtered based on Ensembl v11145 to remove non-human genes from literatures. The variants (rsID, cDNA change, amino acid change) were normalized using regular expression (regex), and the transcript information and corresponding genomic coordinates for variants were extracted using PyEnsembl python package, with Ensembl v11145. The LLM output disease was normalized using MONDO human disease name21 (format-version: 1.2, data-version: releases/2025-03-04), while the LLM output phenotype was normalized using HPO term22 (format-version: 1.2, data-version: hp/releases/2025-03-03). The normalization is based on the cosine similarity >0.9 using PubMedBERT embedding for MONDO disease name and HPO term compared with LLM output disease and phenotype. The final PubMind pathogenicity has been consolidated using all records from different sources of publication. The PubMind pathogenicity score is calculated based on the equation below:
| 1 |
where P=Pathogenic, LP= Likely Pathogenic, B=Benign, LB=Likely Benign. For the paper-level pathogenicity score, the score is calculated using the same equation as above per paper, then averaged across papers to have a paper-level pathogenicity score.
Pathogenicity annotation and assessment
To annotate the variants with genome coordinates, ANNOVAR26 was used to perform variant annotation, for both gene and functional-based, population-based, and bioinformatics tools-based analyses. For gene and functional annotation, hg38 human genome from Ensembl45 (2024-05-13) was used. For population-based annotation, gnomAD46 (version 4.1 whole-genome data) and ClinVar1 (date: 20240917) database was used to cross check the population allele frequency and pathogenicity summarization. For bioinformatics pathogenicity predictions, dbNSFP version 4.7a27 was used to get pathogenicity predictions from 23 bioinformatics software, including AlphaMissense31 and MetaRNN47.
Data and website accessibility
The PubMind source code is available at GitHub (https://github.com/WGLab/PubMind). The PubMind knowledgebase is accessible via a web interface (http://PubMind.wglab.org/) and API. The web application was built with Flask (Python) and supports data queries using SQLite3.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Description of Additional Supplementary Files
Source data
Acknowledgements
We thank Wang lab members for helpful feedback and comment, and thank several users of the PubMind-DB database for insightful suggestions. We thank the IDDRC Biostatistics and Data Science core (HD105354) for consultation. We thank NCBI PubMed and PMC for providing open-access resources that enabled this study.
Author contributions
P.W. executed the project, performed data collection and curation, carried out model development and evaluation, generated all figures and tables, and drafted the initial manuscript. P.W. and K.W. discussed the specific aims of the project and revised the manuscript. K.W. initiated the idea, supervised the project, and shaped the research concept and scope. All authors reviewed and approved the final manuscript.
Peer review
Peer review information
Nature Communications thanks Matthias Heinig and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. A peer review file is available.
Funding
This project is supported by NIH grant HG013031 and the CHOP Research Institute.
Data availability
The PubMind-DB generated in this study can be accessed here: https://pubmind.wglab.org/. The raw PubMed abstracts and PubMed Central (PMC) full texts can be accessed through the open-access ftp depository (https://pubmed.ncbi.nlm.nih.gov/download/, https://pmc.ncbi.nlm.nih.gov/tools/ftp/). Source data are provided with this paper.
Code availability
PubMind is available from ref. 48 and archived on Zenodo49 (https://doi.org/10.5281/zenodo.20632115). PubMind is released under open-source license and can be accessed here: https://github.com/WGLab/PubMind.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s41467-026-76834-4.
References
- 1.Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062–D1067 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Stenson, P. D. et al. The Human Gene Mutation Database (HGMD(R)): optimizing its use in a clinical diagnostic or research setting. Hum. Genet. 139, 1197–1207 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Andersen, E. F. et al. The Clinical Genome Resource (ClinGen): Advancing Genomic Knowledge through Global Curation. Genet. Med.27, 101228 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Richards, S. et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med. 17, 405–424 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Wei, C. H., Allot, A., Riehle, K., Milosavljevic, A. & Lu, Z. Y. tmVar 3.0: an improved variant concept recognition and normalization tool. Bioinformatics38, 4449–4451 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wei, C. H. et al. PubTator 3.0: an AI-powered literature resource for unlocking biomedical knowledge. Nucleic Acids Res.52, W540–W546 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Luo, L. et al. AIONER: all-in-one scheme-based biomedical named entity recognition using deep learning. Bioinformatics39, 10.1093/bioinformatics/btad310 (2023). [DOI] [PMC free article] [PubMed]
- 8.Wei, C. H., Kao, H. Y. & Lu, Z. GNormPlus: An Integrative Approach for Tagging Genes, Gene Families, and Protein Domains. Biomed. Res Int2015, 918710 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wei, C. H., Luo, L., Islamaj, R., Lai, P. T. & Lu, Z. Y. GNorm2: an improved gene name recognition and normalization system. Bioinformatics39, 10.1093/bioinformatics/btad599 (2023). [DOI] [PMC free article] [PubMed]
- 10.Allot, A. et al. Tracking genetic variants in the biomedical literature using LitVar 2.0. Nat. Genet55, 901–903 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sherry, S. T. et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 29, 308–311 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Radford, A., Narasimhan, K., Salimans, T. & Sutskever, I. Improving language understanding by generative pre-training. Preprint at https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf (2018).
- 13.Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) 4171–4186 (2019).
- 14.Vaswani, A. et al. Attention is all you need. Advances in neural information processing systems30 (2017).
- 15.Grattafiori, A. et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024).
- 16.Li, E. B. et al. Rapid functional analysis of computationally complex rare human IRF6 gene variants using a novel zebrafish model. PLoS Genet. 13, e1007009 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lee, J. et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics36, 1234–1240 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Chakraborty, S. et al. in Proceedings of the 28th international conference on computational linguistics 669–679 (2020).
- 19.Yang, J. et al. Enhancing phenotype recognition in clinical notes using large language models: PhenoBCBERT and PhenoGPT. Patterns5, 100887 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Huang, K., Altosaar, J. & Ranganath, R. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342 (2019).
- 21.Vasilevsky, N. A. et al. Mondo: unifying diseases for the world, by the world. MedRxiv, 2022.2004. 2013.22273750 (2022).
- 22.Köhler, S. et al. Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources. Nucleic Acids Res. 47, D1018–D1027 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Furlong, L. I., Dach, H., Hofmann-Apitius, M. & Sanz, F. OSIRISv1.2: a named entity recognition system for sequence variants of genes in biomedical literature. BMC Bioinforma.9, 84 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Thomas, P., Rocktäschel, T., Hakenberg, J., Lichtblau, Y. & Leser, U. SETH detects and normalizes genetic variants in text. Bioinformatics32, 2883–2885 (2016). [DOI] [PubMed] [Google Scholar]
- 25.Wei, C. H., Harris, B. R., Kao, H. Y. & Lu, Z. tmVar: a text mining approach for extracting sequence variants in biomedical literature. Bioinformatics29, 1433–1439 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Wang, K., Li, M. & Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Liu, X., Li, C., Mou, C., Dong, Y. & Tu, Y. dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. 12, 103 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Esmaeil Lashgarian, H. et al. A novel AP4M1 variant in an iranian child with spastic paraplegia 50: a case report and molecular docking approach. Iran. J. Med Sci.51, 70–76 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Wang, W. Y., Ma, P. P., Wang, S. Y. & Wang, Y. J. Case Report: First report of a novel homozygous nonsense mutation in the CYBA gene causing chronic granulomatous disease. Front Immunol.17, 1744743 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zang, H., Yang, X., Liu, Y., Ma, C. & Yang, D. A novel de novo ATP2B1 variant causes autosomal dominant intellectual developmental disorder 66 by disrupting calcium homeostasis via impaired membrane trafficking. Exp. Biol. Med. (Maywood)251, 10834 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Cheng, J. et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science381, eadg7492 (2023). [DOI] [PubMed] [Google Scholar]
- 32.Murali, H., Wang, P., Liao, E. C. & Wang, K. Genetic variant classification by predicted protein structure: A case study on IRF6. Comput Struct. Biotechnol. J.23, 892–904 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Shahzad, F., Chappell, A. G., Purnell, C. A., Aldulescu, M. & Chamlin, S. Infantile myofibroma presenting as a large ulcerative nodule in a newborn. Case Rep. Pediatr.2019, 3476508 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Guerit, E., Arts, F., Dachy, G., Boulouadnine, B. & Demoulin, J. B. PDGF receptor mutations in human diseases. Cell Mol. Life Sci.78, 3867–3881 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Bastianelli, W. E. et al. Molecular Characterization of a ComplexPDGFRBStructural Variation in Infantile Myofibroma With Complete Response to Imatinib. JCO Precis. Oncol.10, e2500943 (2026). [DOI] [PMC free article] [PubMed]
- 36.Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature645, 633–638 (2025). [DOI] [PMC free article] [PubMed]
- 37.Gu, Y. et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans. Comput. Healthcare3, Article 2 (2021).
- 38.Braun, M. et al. Solitary cutaneous infantile myofibroma as a hallmark of myofibromatosis: Two cases and review of the literature. Pediatr. Dermatol39, 438–442 (2022). [DOI] [PubMed] [Google Scholar]
- 39.Ogita, A. & Ansai, S. I. Infantile Myofibroma: Case Report and Review of the Literature. J. Nippon Med. Sch.87, 355–358 (2021). [DOI] [PubMed] [Google Scholar]
- 40.Bastian, B. et al. (IARC, 2018).
- 41.Sanh, V., Debut, L., Chaumond, J. & Wolf, T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108 (2020).
- 42.Gu, Y. et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Trans. Comput. Healthc. (HEALTH)3, 1–23 (2021). [Google Scholar]
- 43.Dogan, R. I., Leaman, R. & Lu, Z. NCBI disease corpus: a resource for disease name recognition and concept normalization. J. Biomed. Inf.47, 1–10 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Luo, L. et al. PhenoTagger: a hybrid method for phenotype concept recognition using human phenotype ontology. Bioinformatics37, 1884–1890 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Harrison, P. W. et al. Ensembl 2024. Nucleic acids Res.52, D891–D899 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature581, 434–443 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Li, C., Zhi, D., Wang, K. & Liu, X. MetaRNN: differentiating rare pathogenic and rare benign missense SNVs and InDels using deep learning. Genome Med.14, 115 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Wang, P. & Wang, K. PubMind: literature-based genetic variant extraction and functional annotation using large language models. WGLab/PubMind 10.5281/zenodo.20632115 (2026). [DOI] [PubMed] [Google Scholar]
- 49.European Organization For Nuclear, R. & OpenAire (CERN, 2013).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Description of Additional Supplementary Files
Data Availability Statement
The PubMind-DB generated in this study can be accessed here: https://pubmind.wglab.org/. The raw PubMed abstracts and PubMed Central (PMC) full texts can be accessed through the open-access ftp depository (https://pubmed.ncbi.nlm.nih.gov/download/, https://pmc.ncbi.nlm.nih.gov/tools/ftp/). Source data are provided with this paper.
PubMind is available from ref. 48 and archived on Zenodo49 (https://doi.org/10.5281/zenodo.20632115). PubMind is released under open-source license and can be accessed here: https://github.com/WGLab/PubMind.
