Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2026 Feb 3;123(6):e2524289123. doi: 10.1073/pnas.2524289123

An interpretable molecular framework for predicting cancer driver missense mutations

Yan Yang a,b, Weikang Sun a,b, Yang Liu a,b, Jian Zhang a,b, Minghui Li a,b,1
PMCID: PMC12890849  PMID: 41632843

Significance

Missense mutations-changes that alter a single amino acid in a protein-are a leading cause of human disease, including cancer. Accurately identifying which mutations disrupt protein function or drive tumor growth remains a central challenge in genetics and oncology. In this work, we systematically analyzed over 120,000 variants to uncover the molecular features that distinguish harmful mutations from neutral ones. Based on these insights, we developed MutaPheno, a transparent and generalizable prediction framework. Trained on inherited disease variants, MutaPheno successfully identifies cancer-driving mutations with higher accuracy and robustness than existing methods. This work advances our ability to interpret genetic variation and may accelerate the discovery of therapeutic targets.

Abstract

Missense mutations play a critical role in human disease, contributing to both inherited disorders and cancer. However, accurately predicting their functional impact—particularly for cancer driver mutations—remains a major challenge due to limited validated labels and the complex molecular basis of oncogenesis. Here, we systematically characterized over 120,000 missense variants across pathogenic, benign, driver, passenger, recurrent somatic, and common population classes, using a comprehensive set of mechanistically grounded molecular features. By assessing the statistical burden of variations, we demonstrated that these features effectively discriminate among diverse variant classes and reveal a consistent enrichment of functional sites, structural integrity, and biophysical changes in pathogenic and driver mutations. Building on these insights, we developed MutaPheno, an interpretable framework for predicting the functional consequences of missense mutations. The model integrates 34 molecular-level features, encompassing structural, functional, physicochemical, and contextual descriptors, using a random forest algorithm. Trained exclusively on pathogenic and benign variants, MutaPheno achieved strong accuracy in predicting cancer driver mutations, outperforming both cancer-specific and general pathogenicity tools, while also demonstrating superior robustness when tested on unseen proteins. Our findings highlight the shared mechanisms between pathogenic and driver mutations and emphasize the role of molecular features in improving variant interpretation. MutaPheno provides a transparent and generalizable tool that can facilitate driver discovery and the development of targeted therapies.


Missense mutations—single-nucleotide substitutions that result in amino acid changes—are among the most prevalent and functionally significant genetic variants. By altering protein structure, stability, or interaction networks, these mutations can disrupt essential biological processes and contribute to disease (1, 2). When such mutations impair protein function and lead to disease, they are referred to as pathogenic missense mutations. Classic examples include the E6V substitution in the β-globin chain causing sickle cell disease (3) and missense mutations in APP, PSEN1, and PSEN2 have been shown to be causative of early-onset familial Alzheimer’s disease (4). Whereas hereditary disorders often result from germline variants, cancer is primarily driven by the accumulation of somatic mutations. Within the complex mutational landscape of tumors, only a small fraction—known as driver mutations—confer a selective growth advantage and actively promote tumorigenesis (5) while the majority are passenger mutations with no functional consequence (6). Missense mutations that promote oncogenesis are classified as cancer driver missense mutations. Representative examples of well-characterized driver mutations include the KRAS G12D substitution frequently observed in pancreatic ductal adenocarcinoma (7), the TP53 R175H hotspot mutation identified across a broad spectrum of carcinomas (8), and the BRAF V600E mutation, a key oncogenic event in melanoma (9).

Understanding and accurately characterizing the functional impact of missense variants is essential for disease diagnosis and the development of targeted therapies. While experimental assays remain the gold standard for functional validation, their high cost and low-throughput limit their scalability. Consequently, over 40 computational tools have been developed to predict the pathogenicity of missense mutations, drawing on features such as evolutionary conservation, protein structural properties, and machine learning algorithms (10, 11). However, the vast majority of these tools were originally designed for inherited diseases. Representative examples—including SIFT (12), PolyPhen-2 (13), REVEL (14), VEST4 (15), and AlphaMissense (16)—have shown strong performance in identifying pathogenic variants. In contrast, only a few tools—such as CHASM (17), ParsSNP (18), CHASMplus (19), and AI-Driver (20)—have been specifically developed for cancer driver missense mutations. Despite these advances, existing computational approaches still face two major challenges: they often operate as “black boxes” with limited mechanistic interpretability, and they exhibit suboptimal predictive accuracy and generalizability in cancer contexts—limitations further exacerbated by label scarcity and the extensive genetic heterogeneity of tumors (2123).

A promising strategy to overcome the limitations of existing variant effect predictors is to focus on molecular phenotypes—quantifiable alterations at the molecular level caused by missense mutations, such as changes in structural stability, disruption of interaction interfaces, or modification of posttranslational modification sites (24). While not all molecular phenotypes manifest as observable organismal traits, most phenotypic and fitness-related consequences ultimately stem from such molecular disruptions. Leveraging these mechanistic insights allows predictive models to achieve greater interpretability by incorporating molecular features. Over the past decade, a series of studies have explored how molecular features reveal the functional impact of missense mutations and help distinguish pathogenic, driver, and benign variants. In 2016, we showed that integrating multiple protein conformations with stability and binding affinity assessments effectively predicts the functional impact of cancer mutations in CBL (25). Iqbal et al. identified physicochemical, structural, and functional differences between pathogenic and population variants, uncovering both shared and protein class–specific signatures (26). Cheng et al. found that germline pathogenic variants are significantly enriched at protein–protein interaction interfaces (27). Laddach et al. extended this line of work by analyzing variant enrichment across structural regions and functional domains, integrating multiomics data to distinguish pathogenic, rare, and common variants (28). Most recently, Gerasimavicius et al. showed that different mutation types—loss-of-function, gain-of-function, and dominant-negative—exhibit distinct structural patterns (29). These observations suggest that pathogenic and driver mutations, despite their conceptual differences, frequently disrupt overlapping molecular features—forming a shared mechanistic layer that enables cross-context generalization.

Despite recent progress, key limitations remain. First, most studies focus on a narrow set of molecular-level features—such as interaction interfaces, posttranslational modification sites, or protein stability—to identify statistical differences across mutation classes (2530). This limits their ability to comprehensively describe mutation effects across biological contexts and hinders a full mechanistic understanding. Second, many efforts remain descriptive and fail to translate molecular features into interpretable predictive models. Although some machine learning approaches incorporate molecular features, they often depend on external scores—such as pathogenicity annotations or conservation metrics—rather than building models from intrinsic, mechanistically meaningful descriptors (26). This reliance ultimately limits both interpretability and generalizability.

To address these limitations, we developed MutaPheno, a unified and interpretable classifier that predicts the functional impact of missense mutations based on curated and computed molecular-level features. MutaPheno integrates 34 structural, functional, physicochemical, and contextual descriptors that systematically characterize pathogenic, benign, driver, passenger, recurrent somatic, and common population variants. These features capture shared and context-specific molecular signatures across variant classes. By training a random forest model exclusively on pathogenic and benign variants, MutaPheno achieves high accuracy and generalizability across both germline and somatic contexts. Notably, despite lacking driver annotations during training, it performs competitively in driver mutation prediction, underscoring the shared molecular basis of functional mutations across diseases. Our work highlights the value of biologically grounded, phenotype-driven models for robust, interpretable, and generalizable prediction of missense mutation effects, and provides a framework for future studies linking molecular mechanisms to disease phenotypes.

Methods

Data Collection and Processing for Missense Mutations.

Pathogenic and benign mutations.

Missense mutations were curated from three sources: ClinVar (31), UniProt Humsavar (32), and the dataset compiled by Sevim Bayrak et al. (33) (overview in Fig. 1A; details in SI Appendix, Table S1). From ClinVar and Humsavar, only variants annotated as “Pathogenic” or “Likely pathogenic,” and “Benign” or “Likely benign,” respectively, were retained, while entries with conflicting interpretations were excluded. From the Sevim Bayrak dataset, gain-of-function (GOF) and loss-of-function (LOF) missense mutations curated from the 2019 release of HGMD (34) were extracted. Variants labeled as Pathogenic, Likely pathogenic, GOF, or LOF were classified as pathogenic, whereas those labeled as Benign or Likely benign were classified as benign. Processed variant counts are summarized in SI Appendix, Table S2.

Fig. 1.

Multi-part figure shows missense mutation data landscape, hierarchical feature annotation, feature enrichment strategy, and model construction.

Overview of the study workflow. (A) Landscape and classification of missense mutations. Variants from major germline and somatic resources were classified as pathogenic or benign, driver or passenger, and somatic or population-derived. Somatic mutations were further stratified by recurrence, gene category (CDG, CPG, CDG∩CPG, PG), and cancer-type occurrence, while population variants were stratified by allele frequency. (B) Hierarchical feature annotation. Each mutation was annotated using 34 interpretable features spanning residue-level (n = 21), mutation-level (n = 8), and protein-level (n = 5) descriptors. (C) Feature enrichment analysis. Odds ratios (ORs) quantified the association between mutation class (positive vs. negative) and feature localization, whereas density ratios (DRs) measured absolute enrichment within feature regions normalized by residue coverage. (D) Interpretable model construction. The MutaPheno random forest classifier was trained on pathogenic and benign variants mapped to AlphaFold2 structures, integrating OR-weighted features (log10OR) with ΔΔG-based thermodynamic descriptors.

Driver and passenger mutations.

Driver and passenger mutations were collected from OncoKB (35), the Cancer Genome Interpreter (CGI) (36), and our previously published Combined Set (37). From OncoKB, variants annotated as “Oncogenic,” “Likely Oncogenic,” or “Likely Neutral” were included, while all validated oncogenic variants were retained from CGI. The Combined Set comprises experimentally characterized missense mutations across 58 cancer genes, previously classified as “Nonneutral” or “Neutral.” For this study, variants annotated as Oncogenic or Likely Oncogenic (OncoKB), oncogenic (CGI), or Nonneutral (Combined Set) were categorized as drivers, whereas Likely Neutral (OncoKB) and Neutral (Combined Set) variants were categorized as passengers (SI Appendix, Table S2).

Somatic and population mutations.

Somatic mutations were obtained from TCGA as curated by Ellrott et al. (38) and filtered following Bailey et al. (39), yielding 746,392 missense mutations across 33 cancer types (SI Appendix, Table S2). Population variants were collected from the non-TCGA subset of ExAC in the gnomAD database (40), retaining only variants passing the “PASS” filter. For comparative analyses, TCGA missense mutations were stratified by recurrence into single (n = 1) and recurrent (n > 1), while ExAC variants were grouped by allele frequency into rare (AF < 1%) and common (AF ≥ 1%). Given that recurrent somatic mutations are more likely to be drivers and common population variants are typically benign (40, 41), we focused on contrasting recurrent somatic mutations with common population variants to identify molecular features distinguishing likely drivers from frequently observed benign mutations.

Standardization of Mutations via Identifier Mapping and Sequence Alignment.

To facilitate downstream analyses, proteins and associated mutations—initially annotated using transcript-, protein-, or gene-level identifiers from multiple sources (SI Appendix, Table S2)—were standardized to UniProt canonical accessions with matched amino acid substitutions, enabling consistent sequence- and structure-level annotation. RefSeq and Ensembl transcript/protein IDs were mapped to UniProt isoform accessions via the EBI Proteins API (42), and gene symbols were converted using the UniProt ID mapping tool (43). Isoform-specific mutations were aligned to canonical sequences using Clustal Omega (v1.2.4) (44) to determine accurate residue positions. To minimize bias in downstream statistical analyses, proteins longer than 3,000 amino acids (~0.82% of the dataset) were excluded. Final counts of retained proteins and mutations are reported in SI Appendix, Table S3, with corresponding AlphaFold2-mapped counts shown in SI Appendix, Table S4.

Stratification of Recurrent Somatic Mutations by Tumor Type and Gene Category.

To enable controlled comparisons of mutation enrichment patterns by recurrence level and cancer-type specificity, recurrent missense variants were stratified in a two-step manner based on i) cancer-type specificity—mutations recurring exclusively within a single cancer type versus those observed across multiple cancer types—and ii) recurrence frequency (n = 2 vs. n > 2), yielding four subgroups: single-cancer (n = 2), single-cancer (n > 2), cross-cancer (n = 2), and cross-cancer (n > 2) (Fig. 1A). For example, single-cancer (n = 2) denotes a mutation observed twice within the same cancer type.

To explore differences across gene categories, we curated high-confidence cancer driver genes (CDG), cancer pathway genes (CPG), and noncancer pathogenic genes (PG). The CDG set was integrated from IntOGen (45), OncoVar (46), NCG (47), Cancer Gene Census (48), and ClinGen (49), yielding 1,050 unique genes. CPG were compiled from 15 cancer-related KEGG pathways in MSigDB (50) and 32 curated signaling pathways in NetSlim (51), comprising 1,042 genes. PG were obtained from ClinGen (49) by selecting genes with “Definitive” clinical validity and no cancer-related phenotype, resulting in 1,117 noncancer disease genes.

To enable mutually exclusive comparisons, 272 proteins overlapping between CDG and CPG were defined as driver–pathway genes (CDG∩CPG). After removing these, 786 CDG-specific and 782 CPG-specific proteins remained. From the PG set, genes overlapping with cancer-related categories were excluded, yielding 922 proteins exclusively associated with noncancer diseases. Recurrent mutations were mapped to one of the four gene classes—CDG, CPG, CDG∩CPG, or PG—based on their protein annotations (Fig. 1A).

Hierarchical Annotation of Missense Mutations at the Residue and Mutation Levels.

We annotated missense mutations using a hierarchical framework integrating residue-, mutation-, and protein-level features (Fig. 1B). Residue-level features capture the local structural and functional context of the mutated site, whereas mutation-level features describe the physicochemical changes introduced by amino acid substitutions. Protein-level features provide broader biological context for interpreting mutational impact. Here, we briefly summarize the first two feature types; detailed definitions, methods, and statistics are provided in the Supplementary Information.

Residue-level features:

Location classification: Residues were classified as buried (Core; rSASA < 0.2) or surface-exposed (Surface) based on relative solvent-accessible surface area (rSASA), calculated using DSSP (52).

Secondary structure: Residues were assigned to Helix, Sheet, or Loop categories based on DSSP annotations (52).

Allosteric sites (AlloSite): Combined experimentally validated sites from the Allosteric Database (ASD) (53) and predicted sites from AlloSitePro (54).

Protein–protein binding sites (PPI-BS): Integrated experimentally determined interfaces from PDB crystal structures and predicted binding sites from Interactome INSIDER (55).

Protein–ligand binding sites (PLI-BS): Derived from ligand-bound PDB structures and predicted using P2Rank (56).

Protein–nucleic acid binding sites (PNI-BS): Extracted from PDB structures and expanded using GraphBind (57) predictions for 2,047 nucleic acid-binding proteins identified from ENPD (58) and EuRBPDB (59).

Posttranslational modifications (PTM): Collected from PhosphoSitePlus (60) and dbPTM (61); residues located within 4 Å and 8 Å of PTM sites, including the modified residues themselves, were labeled PTM-4Å and PTM-8Å, respectively.

Intrinsically disordered regions (IDReg): Verified disordered segments obtained from MobiDB (62) and DisProt (63).

Short linear motifs (SLiM): Sourced from the ELM database (64).

Protein domain regions (DomReg): Domain annotations retrieved from InterPro (65); residues within annotated domains were labeled DomReg.

Physicochemical properties: The 20 standard amino acids were grouped into six categories: aliphatic (Ala, Ile, Leu, Met, Val), aromatic (Phe, Trp, Tyr), positively charged (His, Lys, Arg), negatively charged (Asp, Glu), neutral (Asn, Gln, Ser, Thr), and special (Cys, Pro, Gly).

Mutation-level features:

Stability change (ΔΔGfold): Predicted using PremPS (66), developed by our group, based on AlphaFold2 structures; positive values indicate destabilization, negative values stabilization.

Binding affinity change (ΔΔGbind): Predicted using MutaBind2 (67), developed by our group, based on experimentally resolved heterodimer structures; positive values indicate weakened binding, negative values strengthened binding.

Charge change: Amino acids assigned integer charges (+1: Lys, Arg; –1: Asp, Glu; 0: others); mutations categorized as Charge-0 (no change), Charge-1 (moderate change), or Charge-2 (strong change) based on absolute difference.

Size change: Amino acids grouped as small (0: Ala, Gly, Ser, Cys, Pro, Thr, Asp, Asn), medium (1: Val, His, Glu, Gln), or large (2: Ile, Leu, Met, Lys, Arg, Phe, Trp, Tyr); mutations classified as Size-0, Size-1, or Size-2 by size difference.

In total, we compiled 21 residue-level features: 11 structure-based (Core, Surface, Helix, Sheet, Loop, AlloSite, PPI-BS, PLI-BS, PNI-BS, PTM-4Å, PTM-8Å) and 10 sequence-based (PTM, IDReg, SLiM, DomReg, Aliphatic, Aromatic, PosCharge, NegCharge, Neutral, Special). Among the structure-based features, annotations for allosteric sites and predicted PPI-BS were obtained from external databases, while the others were computed in-house using experimentally resolved PDB structures (68) and/or AlphaFold2-predicted monomeric models (69, 70); models that were truncated or had a median pLDDT score below 70 were excluded. In addition, eight mutation-level features were defined, comprising two structure-based descriptors (ΔΔGfold and ΔΔGbind) and six sequence-based indicators capturing changes in charge and size (Charge-0/1/2 and Size-0/1/2).

Odds Ratio Analysis of Feature Associations in Positive- versus Negative-Class Mutations.

The odds ratio (OR) is a statistical metric used to quantify the relative likelihood of an event occurring in one group compared to another (26, 29). We calculated ORs to evaluate whether positive-class mutations (e.g., pathogenic, driver) are more likely than negative-class mutations (e.g., benign, passenger) to be enriched in specific categories of 21 residue-level and six mutation-level features (including charge and size changes) (Fig. 1C). ORs were calculated as

ORfeatpos/neg=Nmut.featposNmut.nonfeatposNmut.featnegNmut.nonfeatneg, [1]

where Nmut.featpos and Nmut.nonfeatpos are the counts of positive-class mutations within or outside the feature region, respectively, and Nmut.featneg and Nmut.nonfeatneg are the corresponding counts for negative-class mutations. For example, ORcorepathogenic/benign quantifies whether pathogenic mutations are more likely than benign mutations to occur in the core region. Details for this example are provided in the SI Appendix. For sequence-based features, the entire protein sequence was considered; for structure-based features, only structurally resolved regions were included. Proteins lacking relevant feature coverage were excluded.

ORs were computed using two-sided Fisher’s exact tests implemented in the SciPy (Python) (71). P-values were adjusted for multiple testing using the Benjamini–Hochberg (BH) procedure (72) to obtain q-values. The 95% CI were estimated using the Wald method (73). A feature was considered significantly associated with mutation status if OR ≠ 1, the 95% CI did not include 1, and q < 0.05. For visualization, odds ratios were log10-transformed: log10OR > 0 indicates preferential associations with positive-class mutations, whereas log10OR < 0 indicates preferential associations with negative-class mutations.

Variant Set-Specific Feature Enrichment Analysis.

While odds ratios (ORs) compare the relative prevalence of positive- versus negative-class mutations within a feature, they do not capture absolute enrichment relative to overall mutation density. To address this, we used the Density Ratio (DR) metric (28, 74), defined in Eq. 2:

DRfeatset=Nmut.featset/Nres.featsetNmut.totalset/Nres.totalset. [2]

Here, Nmut.featset and Nres.featset represent the number of mutations and residues within a given feature, respectively, for a specific dataset. Nmut.totalset and Nres.totalset are the total number of mutations and residues across all analyzed proteins. For example, DRcorepathogenic quantifies the density of pathogenic mutations in core regions relative to their overall density in the pathogenic dataset. Details for this example are provided in the SI Appendix.

To assess statistical significance, we generated a null distribution of DR values from 2,000 random simulations per dataset, randomly reassigning mutations to positions within the same protein while preserving mutation counts per protein and allowing multiple mutations at the same residue. Two-tailed P-values were obtained by comparing observed DR values to this null distribution. P-values were adjusted for multiple testing using the BH method to obtain q-values. DR values were transformed to log10DR for visualization; log10DR > 0 indicates feature enrichment. To estimate 95% CI, we performed 2,000 bootstrap resamplings with replacement at the mutation level, deriving CIs from the 2.5th and 97.5th percentiles of the bootstrap distribution.

Integrative Protein-Level Features for Mutation Impact Modeling.

Proteins were annotated using five protein-level feature categories: functional classes, subcellular localization, KEGG pathways, cancer signaling pathways, and protein domain types. Unlike residue- and mutation-level features, protein-level annotations were not analyzed individually for enrichment but were used exclusively for model construction in MutaPheno, as the large number of subclasses would yield limited mutation counts per group. A brief overview is provided below, with detailed definitions available in the Supplementary Information and Dataset S1.

Functional classes were retrieved from the PANTHER classification system (75), which assigns proteins to 24 groups based on conserved functions and evolutionary relationships.

Subcellular localization data were obtained from the Human Protein Atlas (HPA) (76), including only high-confidence annotations (Enhanced, Supported, or Approved) covering 36 compartments.

Pathway annotations were sourced from 186 KEGG pathways in MSigDB (50) and 32 curated cancer signaling pathways in NetSlim (51).

Protein domain types were annotated using InterPro (65); subdomains were consolidated into parent categories, yielding 4,608 unique domain types.

To incorporate these annotations into modeling, we computed five aggregated protein-level features—one per annotation category. For each mutation, we identified all annotation terms within a category associated with its protein and summed the log-transformed odds ratios comparing pathogenic and benign mutations to obtain the aggregated score:

log10ORprotein.featpathogenic/benign=k=1nlog10ORfeatkpathogenic/benign,

where n is the number of annotation terms in the category, and log10ORfeatkpathogenic/benign is the log-transformed odds ratio for the k-th term. The aggregated score reflects the cumulative pathogenic enrichment of a protein in each category. Additional details and illustrative examples are provided in the Supplementary Information.

Construction of the MutaPheno Model.

Training set.

We curated a high-quality dataset of pathogenic and benign missense mutations (Fig. 1A and SI Appendix, Table S3). To ensure reliable structure-based features, only variants mapped to nontruncated AlphaFold2 monomeric models with median pLDDT ≥ 70 were retained (SI Appendix, Table S4). The final training set comprised 67,791 mutations from 9,928 proteins, including 30,113 pathogenic and 37,678 benign variants.

Independent test sets.

To evaluate generalizability, we prepared two independent test sets:

  • Pathogenic mutation test set. Derived from AlphaMissense (16) and CPT-1 (77) test sets (ClinVar sourced), harmonized and filtered for canonical alignment and nonoverlap with our training data. After mapping to AlphaFold2 structures, this set included 28,717 mutations (12,539 pathogenic, 16,178 benign).

  • Driver mutation test set. Curated from the driver/passenger dataset described previously (Fig. 1A and SI Appendix, Table S4). Variants overlapping with training data were excluded, yielding 6,322 mutations (2,317 drivers, 4,005 passengers) that could be mapped to AlphaFold2 structures.

Feature engineering.

We incorporated 34 features: 21 residue-level, 8 mutation-level, and 5 aggregated protein-level annotations. For the residue-level features and six mutation-level features (charge and size changes), we calculated log10ORfeatpathogenic/benign and encoded features as binary indicators multiplied by their log10OR values, integrating feature presence and pathogenic enrichment. The two mutation-level biophysical features (folding stability and binding affinity changes) were represented by ΔΔG values from PremPS and MutaBind2; unavailable predictions were assigned a value of 0. For protein-level features, we used aggregated log10ORprotein.featpathogenic/benign scores as described earlier.

Model construction and hyperparameter tuning.

MutaPheno was constructed using a random forest (RF) classifier, with two hyperparameters optimized: the number of trees (n_estimators, 200 to 2,500 in steps of 50) and the number of features considered at each split (max_features, 1 to 34). In our previous studies (66, 67, 7880), RF consistently outperformed other conventional machine-learning methods and offers high computational efficiency, robustness to overfitting, and interpretable feature importance.

To reduce sequence redundancy and improve generalization, data were partitioned based on protein sequence similarity. Pairwise alignments were performed using MMseqs2 (81) with thresholds of 25% sequence identity and 40% alignment coverage for both query and target. An undirected graph was constructed in NetworkX (82), with proteins as nodes and edges defined by sequence similarity; clustering by connected components yielded 4,202 nonoverlapping clusters, ensuring no similarity between proteins from different clusters.

The 4,202 clusters were randomly split into training and validation sets at a 4:1 ratio, repeated 10,000 times. To account for uneven mutation distributions, partitions were ranked by deviation from the ideal 4:1 mutation ratio, and the top five were selected. These partitions preserved the original pathogenic/benign ratio (SI Appendix, Table S5A). Partition independence was further assessed using pairwise Jaccard similarity coefficients (SI Appendix, Table S5 B and C): training-set similarities ranged from 0.63 to 0.70, whereas validation-set similarities ranged from 0.05 to 0.17, indicating high distinctness.

A grid search was performed across the five partitions to identify optimal hyperparameters, using the average AUC-ROC as the evaluation metric. The final MutaPheno model was trained on the full training set using the selected parameters.

Statistical metrics and evaluation criteria.

Model performance was evaluated using multiple complementary metrics capturing different aspects of classification quality, including the area under the receiver operating characteristic curve (AUC-ROC), which measures overall discrimination across thresholds; the area under the precision–recall curve (AUC-PR), which is particularly informative for imbalanced datasets; the maximum Matthews correlation coefficient (maximum MCC), defined as the highest MCC achieved across thresholds; MCC computed at each method’s recommended or default threshold; the F1-score, the harmonic mean of precision and recall; and precision and recall individually.

Model calibration assessment.

Calibration, reflecting agreement between predicted probabilities and observed outcome frequencies, is essential for clinically interpretable risk estimation (83, 84). Global calibration was assessed using the Brier score, defined as the mean squared error between predicted probabilities and binary outcomes. Probability reliability across the prediction range was evaluated using quantile-binned calibration curves, generated by dividing predictions into ten equal-sized bins and comparing mean predicted probabilities with observed positive fractions; perfect calibration corresponds to the 45° identity line. Cox recalibration was further performed to estimate the calibration intercept (α), slope (β), and the overall miscalibration index (|β − 1 |+|α|), capturing systematic bias (ideal α = 0) and probability scaling (ideal β = 1). All calibration analyses were conducted on independent test sets for both pathogenicity and cancer driver mutation prediction tasks.

Results

To systematically investigate molecular-level characteristics of functional missense variants, we curated a comprehensive dataset spanning six major variant classes: pathogenic, benign, driver, passenger, recurrent somatic, and population variants. The dataset encompassed over 127,000 missense mutations affecting more than 15,000 proteins, each annotated with 34 molecular features describing structural and functional context, physicochemical changes, and broader protein-level properties. An overview of the study design and data composition is shown in Fig. 1 and summarized in SI Appendix, Tables S3 and S4. Mutation distributions across features are detailed in Fig. 2 and SI Appendix, Fig. S1. As expected, features such as Core and ΔΔGfold covered a large number of mutations across variant classes, reflecting their extensive representation in protein structures. In contrast, features like SLiM, PNI-BS, and ΔΔGbind mapped to far fewer mutations, likely due to their localized nature and more limited structural coverage. These differences highlight variability in statistical power, which was accounted for in subsequent enrichment analyses.

Fig. 2.

Multi-part figure shows stacked bar graphs of feature annotation coverage across mutation categories. A, B, and C show different categories.

Coverage of feature annotations across mutation categories. (A) Pathogenic vs. benign variants. (B) Driver vs. passenger variants. (C) Recurrent somatic vs. common population variants. Stacked bars (log10 scale) display the number of variants annotated for 21 residue-level and eight mutation-level features; labels indicate total counts, and color segments denote category-specific contributions. Summary statistics—including the number of variants, unique sites, and associated proteins—are provided in Dataset S2.

Feature Association and Functional Characterization.

We first assessed 21 residue-level and six mutation-level features using odds ratio (OR) analysis across three binary classification tasks: pathogenic vs. benign, driver vs. passenger, and recurrent vs. common. The two remaining mutation-level features—folding stability and binding affinity changes—were evaluated using continuous predicted ΔΔG values, as described in the section “Biophysical disruption profiles differentiate functionally relevant mutation classes”. For allosteric and binding-site features, both experimentally determined and computationally predicted annotations were available. To assess the reliability of predicted site annotations, we compared mutation association patterns (log10OR) derived from predicted sites with those from experimentally validated sites across the three tasks. Allosteric sites (SI Appendix, Fig. S2A), protein–ligand binding sites (SI Appendix, Fig. S2C), and protein–protein binding sites (SI Appendix, Fig. S2B) exhibited largely consistent and concordant association patterns, supporting the integration of both annotation types. In contrast, protein–nucleic acid binding sites (SI Appendix, Fig. S2D) showed inconsistent results, with predictions displaying an opposite association direction in the pathogenic vs. benign task and no significant signal in the other two tasks. Given these discrepancies—and to avoid introducing annotation noise—we retained only experimentally validated protein–nucleic acid binding sites in downstream analyses. OR-based analyses across tasks (Fig. 3) revealed distinct patterns of feature associations, enabling features to be grouped into four categories reflecting consistent or context-dependent associations with functional mutations.

Fig. 3.

Bar plots of log sub 10 O R for residue and mutation-level features across classification tasks. Features are grouped into four categories.

Associations of missense mutation classes with 27 molecular features. Bar plots show log10 odds ratios (ORs) for 27 residue- and mutation-level features across three classification tasks: pathogenic vs. benign, driver vs. passenger, and recurrent vs. common. Positive log10OR values indicate preferential association with the positive class (pathogenic, driver, or recurrent), whereas negative values indicate association with the corresponding negative class (benign, passenger, or common). Error bars denote 95% CI; statistical significance was assessed using two-sided Fisher’s exact tests with Benjamini–Hochberg correction (*q < 0.05; **q < 0.001; ***q < 0.0001). Features were grouped into four categories based on the direction and consistency of their associations across tasks: Category 1, features consistently associated with the positive class; Category 2, features showing context-dependent positive associations across mutation types; Category 3, features exhibiting mixed-direction associations across tasks; and Category 4, features consistently associated with the negative class. Features within each category are ordered by their mean log10OR across tasks. Underlying data are provided in Dataset S3.

Features with consistent associations with positive-class mutations.

Ten features showed significant and consistent positive associations with positive-class mutations across all tasks (log10OR > 0, q < 0.05), indicating shared molecular signatures across germline and somatic disease variants (red-labeled in Fig. 3). These features spanned functional sites (PNI-BS, PLI-BS, DomReg, AlloSite, PPI-BS, and PTM-8Å) involved in regulation, signaling, and catalysis, structural regions (Core and Sheet) linked to protein stability, and mutation-level properties capturing substantial size (Size-2) and charge (Charge-1) changes. Notably, PNI-BS remained informative despite limited coverage. Together, these patterns indicate that functional-site disruption, structural destabilization, and pronounced physicochemical perturbations represent shared mechanisms underlying disease-related mutations.

Features with context-dependent associations with positive-class mutations.

Five features showed significant positive signals (log10OR > 0, q < 0.05) in only a subset of classification tasks (orange in Fig. 3), indicating context-specific effects. Among these, PTM-4Å displayed the strongest contrast between Pathogenic vs. Benign and Driver vs. Passenger comparisons, with a more pronounced signal in somatic driver mutations. This pattern suggests that mutations proximal to posttranslational modification sites may preferentially contribute to tumorigenesis by perturbing regulatory mechanisms.

Features with mixed-direction associations across mutation types.

Six features exhibited significant but opposing patterns across classification tasks (q < 0.05; blue in Fig. 3), reflecting context-dependent effects. Notably, PTM sites and Helix showed contrasting behaviors between germline and somatic mutations: PTM sites were preferentially associated with driver mutations but depleted among pathogenic variants, consistent with a cancer-specific regulatory role, whereas Helix was observed primarily in pathogenic mutations, highlighting the importance of structural context in germline disease. Together, these opposing patterns indicate that feature effects vary by mutation class, underscoring the need to integrate diverse molecular signals when modeling mutation impact.

Features with consistent associations with negative-class mutations.

Six features showed significant and consistent negative associations across all tasks (log10OR < 0, q < 0.05; green in Fig. 3). Surface residues exhibited the strongest negative signals, consistent with the higher mutational tolerance of solvent-exposed sites. Likewise, mutations involving no change in charge (Charge-0) or size (Size-0) were preferentially associated with negative-class mutations, reflecting their limited structural and physicochemical impact. Flexible regions (Loop) and residues with aliphatic or neutral properties also showed consistent negative associations, indicating reduced functional constraint. Together, these patterns suggest that mutations carrying these features are more likely to be tolerated and functionally neutral across contexts.

Integrative Analysis of Statistical Associations and Mutation Enrichment Patterns.

To complement odds ratio (OR) analysis, we applied Density Ratio (DR) metrics to assess the absolute enrichment of mutations across residue-level features. DR quantifies mutation density within specific regions relative to the global background, helping identify areas of potential functional or structural relevance. By integrating OR and DR (Fig. 4), we identified features that are not only statistically associated with positive-class mutations but also show localized enrichment, suggesting sites under selective pressure or biological vulnerability. DR profiles for all 21 residue-level features across six mutation datasets are provided in SI Appendix, Fig. S3.

Fig. 4.

Heat map: Mutation associations for 21 residue features. Squares = log sub 10 odds ratios; Circles = log sub 10 density ratios across three tasks.

Integrative analysis of mutation–feature associations and enrichment. Shown are 21 residue-level features, as mutation-level attributes cannot be evaluated using density-based enrichment (DR) metrics. Squares denote the direction and significance of feature–mutation class associations based on log10 odds ratios (ORs) across three classification tasks: pathogenic vs. benign, driver vs. passenger, and recurrent vs. common. Red and blue squares indicate significant associations with the positive or negative mutation class, respectively, whereas gray squares denote nonsignificant associations (q ≥ 0.05). Circles indicate feature-specific mutation enrichment measured by log10 density ratios (DRs), with red and blue circles representing enrichment or depletion, respectively; circle size reflects the magnitude of |log10DR|. By jointly integrating OR- and DR-based analyses, this visualization highlights molecular features that are not only statistically associated with functional mutation classes but also exhibit localized enrichment, indicative of regions under selective pressure or functional vulnerability.

For example, nine features with consistent positive OR associations also showed significant DR enrichment in pathogenic and driver datasets (PNI-BS to Special), indicating that these regions are both functionally critical and intrinsic mutational hotspots. Although Core, Sheet, and PTM-8Å significantly distinguished recurrent from common mutations, DR analysis revealed that this distinction was mainly driven by depletion of common mutations rather than genuine enrichment of recurrent ones. This supports the view that recurrent mutations are not necessarily functional drivers and that biologically important features may not always show enrichment for such mutations. For features with mixed-direction associations, DR analysis helped clarify apparent inconsistencies. For example, PTM sites showed significant OR associations but no significant DR enrichment in pathogenic and driver mutation categories, contrasting with prior reports of PTM enrichment in disease (85). This underscores the added value of spatially extended features such as PTM-4Å and PTM-8Å, which may better capture the local structural and functional impact of PTMs.

OR and DR analyses showed that recurrent mutations exhibit association and density-based enrichment patterns distinct from those observed for driver mutations. Because background mutational processes can influence recurrence and potentially confound its interpretation as a proxy for driver status, we corrected recurrence levels using codon-level mutability and B-scores and repeated all downstream OR and DR analyses (SI Appendix, Correction for Background Mutational Processes Using Mutability and B-Scores). The resulting association and enrichment patterns remained highly consistent with the original findings, with only two features showing changes in statistical significance, whereas all major structural and functional trends were preserved. These results confirm that the observed patterns are robust and not driven by local variation in mutation rate.

Stratified Enrichment Patterns of Recurrent Mutations Improve Driver Prioritization.

To further evaluate whether enrichment patterns of carefully selected molecular features can improve driver prioritization and help identify potential driver mutations, we systematically stratified recurrent somatic mutations by cancer-type recurrence patterns and gene-level categories and performed enrichment analyses. This stratification builds on previous studies proposing that mutations recurrent across multiple cancer types are more likely pan-cancer drivers under shared selective pressures, whereas single-cancer recurrent mutations may reflect cancer-type-specific drivers or background mutational processes (39, 86, 87). Specifically, we distinguished single-cancer from cross-cancer recurrent mutations and further subdivided each group into four tiers by recurrence frequency to enable controlled comparisons and reduce confounding by recurrence level (Fig. 1A). Given the pivotal role of driver genes involved in key cancer pathways for tumorigenesis (5, 88), we also classified genes into three cancer-associated subtypes, including a category of driver–pathway genes (CDG∩CPG), and a noncancer pathogenic group. Mutation counts are summarized in SI Appendix, Table S3, with feature distributions shown in SI Appendix, Fig. S1.

Hierarchical clustering of structural and functional enrichment patterns based on log10DR profiles revealed a clear stratification of mutation categories (Fig. 5). Driver and pathogenic mutations showed highly concordant profiles, clustering tightly together with recurrent mutations in driver–pathway genes (CDG∩CPG) and highly recurrent cross-cancer mutations (n > 2), consistent with their functional importance in oncogenic processes. These mutations were strongly enriched in structurally constrained and functionally critical regions, such as ligand- and protein-binding interfaces, allosteric sites, and short linear motifs. Mutations in CDG-only genes, as well as recurrent and cross-cancer mutations occurring in two tumors (n = 2), exhibited intermediate enrichment patterns with consistently weaker magnitude, indicating that recurrence captures only partial functional relevance. In contrast, single-cancer mutations showed minimal or heterogeneous enrichment, consistent with a greater contribution from background mutational processes or tissue-specific effects. Collectively, these stratified enrichment patterns offer a refined basis for prioritizing potential functional driver mutations beyond recurrence frequency alone.

Fig. 5.

Heat map of log sub10 density ratios for residue-level features across mutation categories. Red indicates enrichment, blue indicates depletion.

Enrichment-based clustering of mutation categories using residue-level features. Heatmap showing log10 density ratios (DRs) for 21 residue-level features across 13 mutation categories, with red indicating enrichment and blue indicating depletion. Numerical log10DR values are shown in each cell. Statistical significance was assessed by comparing observed DRs to a null distribution generated from 2,000 random simulations, followed by Benjamini–Hochberg correction (*q < 0.05; **q < 0.001; ***q < 0.0001). The left dendrogram depicts hierarchical clustering of mutation categories based on their enrichment profiles, with closely clustered categories exhibiting similar feature-level signatures. Underlying data are provided in Dataset S4, and DR simulation results in Dataset S5.

Biophysical Disruption Profiles Differentiate Functionally Relevant Mutation Classes.

The two biophysical mutation-level features—folding stability and binding affinity changes—were quantified using predicted ΔΔG values. As shown in Fig. 6A, pathogenic mutations exhibited the strongest folding destabilization, significantly exceeding benign variants (Welch’s t test, q < 0.0001). Driver mutations ranked second and likewise showed greater destabilization than passengers (q < 0.0001). At the gene level, the CDG∩CPG group—ranking immediately after pathogenic and driver mutations—displayed the highest proportion of strongly destabilizing substitutions among all gene-level categories and was the only category whose mean ΔΔG value significantly exceeded that of passengers (q < 0.0001). Further stratification by cancer-type specificity and recurrence reinforced this trend: cross-cancer mutations (n > 2) showed the strongest destabilization within this four-group framework and significantly surpassed single-cancer mutations (n > 2) (q < 0.001), consistent with stronger selective constraints on pan-cancer recurrent mutations.

Fig. 6.

Two-part figure shows box plots and stacked bar graphs of predicted folding stability and protein binding affinity changes across mutation categories.

Biophysical disruption of protein folding and protein–protein interactions across mutation categories. (A, Top) Box plots of predicted folding stability changes (ΔΔGfold) across 14 mutation categories. (Bottom) Stacked bars showing the proportion of mutations classified into four stability-impact groups. (B, Top) Box plots of predicted protein–protein binding affinity changes (ΔΔGbind) for the same mutation categories. (Bottom) Stacked bars showing the corresponding distribution of binding-impact classes. In box plots, dashed red and blue lines indicate the median ΔΔG values of driver and passenger mutations, respectively, while the gray dashed line marks ΔΔG = 0. Underlying data are provided in Dataset S6 (ΔΔGfold) and Dataset S7 (ΔΔGbind).

A parallel trend was observed for binding affinity: pathogenic and driver mutations showed the strongest binding destabilization and caused significantly greater disruption of protein–protein interactions than benign and passenger variants (q < 0.05). High-recurrence cross-cancer mutations (n > 2) and the CDG∩CPG group followed, both exhibiting elevated proportions of strongly affinity-decreasing substitutions (Fig. 6B). In contrast, benign and common variants consistently displayed the weakest effects in both analyses. These findings demonstrate that folding and binding disruptions are robust indicators of mutational impact and highlight the value of ΔΔG-based features for prioritizing driver mutations in cancer (25).

MutaPheno and Its Feature-Driven Interpretability.

MutaPheno is a random forest classifier trained on 67,791 high-confidence missense variants (30,113 pathogenic; 37,678 benign), all mapped to AlphaFold2-predicted structures. It integrates 34 molecular features—21 residue-level, eight mutation-level, and five aggregated protein-level annotations—each scaled by log10OR enrichment or raw ΔΔG values. Hyperparameter tuning across five random training–validation partitions identified an optimal configuration of 2,300 trees with eight features per split. Model performance was consistently strong (AUC-ROC 0.864 ± 0.009, MCC 0.580 ± 0.020; SI Appendix, Table S6), and the final model was retrained on the full dataset with these parameters.

To evaluate the reliability of MutaPheno’s structural inputs, we compared AlphaFold2-predicted structures with experimentally resolved PDB structures, as detailed in the SI Appendix, Impact of AlphaFold2 and PDB Structural Inputs on Model Performance). AF2-predicted structures offered markedly broader proteome-wide coverage and yielded more robust and informative structural features for mutation-effect prediction.

We next assessed model-choice robustness by benchmarking multiple interpretable machine-learning algorithms (Logistic Regression, Random Forest, XGBoost, and LightGBM) using identical features and evaluation protocols. Across both pathogenicity and driver-mutation prediction tasks, tree-based models showed highly comparable performance (SI Appendix, Table S8). Random Forest achieved the most balanced performance on the pathogenic mutation test set, whereas LightGBM provided slightly stronger discrimination for driver mutations. In the absence of a clear performance advantage, model selection was guided by attribution behavior: Random Forest exhibited the most stable and evenly distributed feature-contribution profile (SI Appendix, Figs. S5 and S6) and was therefore selected as the primary model, with full benchmarking results provided in the SI Appendix, Benchmarking Interpretable Models and Comparing Attribution Behaviors.

To identify the determinants of MutaPheno predictions, we computed global SHAP values to quantify the average contribution of each feature (SI Appendix, Fig. S5). Stability change emerged as the dominant factor, followed by protein-level annotations, including KEGG pathway membership, functional class, and domain context. Among residue-level features, core/surface annotations were most influential, with additional contributions from protein–protein binding and allosteric sites. Although sparsely represented, features such as SLiMs and PNI-BS retained nonzero importance, consistent with their enrichment patterns (Fig. 3). Among mutation-level descriptors, large size changes and charge-preserving substitutions contributed most strongly. Collectively, these results indicate that MutaPheno integrates structural perturbations with multiscale biological context to achieve robust and interpretable classification.

To further illustrate how MutaPheno generates interpretable predictions, we computed SHAP values for three well-characterized driver mutations (Fig. 7). For PIK3CA-E545K, major contributors included KEGG pathway membership, predicted disruption of p110α–p85α interactions (ΔΔGbind = 1.43 kcal/mol, based on PDB 4OVU), and allosteric site involvement (Fig. 7A). Consistent with prior studies reporting that E545K weakens this inhibitory interaction, mimicking activating phosphorylation, and driving constitutive PI3Kα activation (89). KEGG pathways (e.g., ERBB, phosphatidylinositol, mTOR) are all tightly linked to PI3K/AKT signaling (Dataset S9). For TP53-R175H, the prediction was mainly driven by folding destabilization and location in the DNA-binding domain (Fig. 7B and Dataset S9), in line with experimental data showing that R175H disrupts the structural integrity of p53’s core domain, leading to loss of DNA binding (90). For BRAF-V600E, SHAP drivers included strong ΔΔGfold and enrichment in MAPK pathway annotations (Fig. 7C and Dataset S9). This matches the known phosphomimetic role of V600E, which constitutively activates the MAPK/ERK signaling cascade (9, 91, 92). These cases illustrate that MutaPheno not only delivers accurate predictions but also offers mechanistic insight, supporting its utility for hypothesis generation and experimental validation.

Fig. 7.

Three-panel bar graph shows top 10 molecular features ranked by SHAP values for driver mutations PIK 3 C A-E 545 K, T P 53-R 175 H, and BRAF-V 600 E.

Feature attribution for representative driver mutations. Bar plots show the top 10 molecular features ranked by absolute SHAP values for three canonical driver mutations: (A) PIK3CA-E545K, (B) TP53-R175H, and (C) BRAF-V600E. SHAP (SHapley Additive exPlanations) values quantify the contribution of individual molecular features to the corresponding model prediction, with larger absolute values indicating stronger influence. Underlying data are provided in Dataset S8.

Benchmarking MutaPheno against Existing Pathogenicity and Driver Mutation Prediction Tools.

We benchmarked MutaPheno on independent test sets for both pathogenicity and driver mutation prediction to assess its accuracy and generalizability. For pathogenicity prediction, we evaluated 28,717 ClinVar-derived missense variants and compared MutaPheno with 48 existing tools, including 47 models from dbNSFP4.9a (10, 11) and CPT-1 (77). After excluding methods affected by data leakage or sparse coverage, we focused on the top 10 performers and analyzed their common subset of 23,222 variants (10,485 pathogenic and 12,737 benign; the intersected pathogenic mutation test set). Performance based on continuous scores and discrete classifications is summarized in Table 1 and SI Appendix, Table S9. Among these models, AlphaMissense achieved the highest discrimination, followed by CPT-1 and VEST4, while MutaPheno ranked seventh by AUC-ROC (0.882) (SI Appendix, Fig. S7A). Using discrete outputs, MutaPheno ranked fifth by MCC (0.616). Notably, among all top-performing models, MutaPheno uniquely provides mechanistic interpretability by explicitly linking molecular features to functional impact.

Table 1.

Comparative performance of the top 10 pathogenicity prediction models on the intersected pathogenic mutation test set

Model AUC-ROC AUC-PR Maximum MCC Sensitivity at 10% FPR Brier score
AlphaMissense 0.924 0.907 0.703 0.792 0.109
CPT-1 0.923 0.908 0.691 0.774 0.134
VEST4 0.911 0.883 0.663 0.718 0.144
M-CAP 0.891 0.857 0.631 0.693 0.185
DEOGEN2 0.890 0.879 0.623 0.700 0.136
MetaSVM 0.886 0.854 0.634 0.695 0.156
MutaPheno 0.882 0.867 0.622 0.683 0.138
MVP 0.882 0.856 0.612 0.670 0.238
Eigen 0.877 0.841 0.590 0.614 0.182
MutFormer 0.873 0.844 0.593 0.609 0.193

Models are ranked in descending order of AUC-ROC. Bold values indicate the performance of the proposed method, MutaPheno. Variant-level prediction scores are provided in Dataset S9.

For driver mutation prediction, we evaluated 6,322 somatic missense variants and compared MutaPheno with both general pathogenicity predictors and established cancer-specific tools (CHASM, CHASMplus, ParsSNP, and AI-Driver). MutaPheno showed the strongest overall performance on the full dataset. For controlled comparison, we analyzed an intersected test set of 5,770 variants (2,018 drivers and 3,752 passengers; the intersected driver mutation test set) with predictions available from 12 representative models, including top pathogenicity predictors, MVP (the best-performing general tool for driver prediction), cancer-specific tools, and MutaPheno. On this set, MutaPheno achieved the highest AUC-ROC (0.843; Table 2 and SI Appendix, Fig. S7B), with ParsSNP and MVP also performing well. Although AI-Driver attained the highest maximum MCC, it showed reduced sensitivity at a 10% false-positive rate. Notably, models excelling in pathogenicity prediction showed markedly lower accuracy for driver prediction (AUC-ROC < 0.800). In classification analyses (SI Appendix, Table S10), MutaPheno achieved the second-highest MCC (0.530) and F1-score (0.707); applying ParsSNP’s recommended cutoff yielded substantially lower performance (MCC = 0.369).

Table 2.

Comparative performance of the top 12 pathogenicity and cancer-specific prediction models on the intersected driver mutation test set

Model AUC-ROC AUC-PR MaximumMCC Sensitivity at 10% FPR Brier Score
MutaPheno 0.843 0.698 0.549 0.462 0.166
ParsSNP 0.837 0.708 0.586 0.409 0.287
MVP 0.835 0.712 0.503 0.473 0.290
AI-Driver 0.835 0.631 0.641 0.289 0.233
DEOGEN2 0.794 0.656 0.434 0.424 0.222
AlphaMissense 0.780 0.638 0.428 0.403 0.252
CHASM 0.776 0.623 0.443 0.410 0.245
M-CAP 0.763 0.600 0.394 0.317 0.217
MetaSVM 0.763 0.532 0.418 0.211 0.197
CPT-1 0.762 0.635 0.418 0.447 0.181
CHASMplus 0.753 0.705 0.498 0.512 0.173
VEST4 0.692 0.505 0.302 0.217 0.305

Models are ranked in descending order of AUC-ROC. Bold values indicate the performance of the proposed method, MutaPheno. Variant-level prediction scores are provided in Dataset S9.

Importantly, MutaPheno was evaluated under a strictly independent setting, excluding all mutations overlapping with its training data. In contrast, ParsSNP, MVP, and AI-Driver were trained on datasets that partially overlapped with the test set, which may inflate their reported performance. For instance, AI-Driver’s training data included 1,665 driver and 532 passenger mutations present in the test set. ParsSNP employs a weakly supervised learning framework: during its unsupervised EM phase, it assigned driver probabilities to somatic mutations from TCGA, ICGC, and COSMIC; for example, 647 driver and 228 passenger mutations in the test set were present in the TCGA data used for label learning, and these probabilities were subsequently used as training labels.

To further evaluate MutaPheno’s predictive ability on unseen proteins, we performed leave-protein-out and leave-similar-protein-out testing. For pathogenicity prediction, excluding variants from proteins present in the training data yielded 338 variants (130 pathogenic, 208 benign), on which MutaPheno achieved an AUC-ROC of 0.841 and AUC-PR of 0.754 (SI Appendix, Table S11A). For driver prediction, fewer than 50 variants remained after leave-protein-out filtering; therefore, we adopted an alternative strategy by retraining two models with identical parameters. MutaPheno* (trained on non-overlapping proteins; 63,839 mutations) achieved an AUC-ROC of 0.844, AUC-PR of 0.710, and MCC of 0.478 (SI Appendix, Table S11B), whereas MutaPheno^ (trained on proteins sharing <25% sequence identity with the test set; 60,334 mutations) achieved an AUC-ROC of 0.841, AUC-PR of 0.699, and MCC of 0.475. These results highlight MutaPheno’s strong generalizability to unseen or distantly related proteins.

Finally, we assessed probability calibration across all benchmarked methods. For pathogenicity prediction, MutaPheno showed superior probability reliability despite moderate discrimination, ranking fourth by Brier score and achieving the lowest Cox miscalibration index (Table 1 and SI Appendix, Fig. S8). For driver prediction, MutaPheno achieved both the highest discrimination and the best calibration among all methods (Table 2 and SI Appendix, Fig. S8). Calibration curves further confirmed that its predicted probabilities closely matched observed outcome frequencies. Together, these results demonstrate that MutaPheno delivers competitive predictive accuracy while providing uniquely reliable and interpretable probability estimates, supporting its utility for translational and preclinical applications.

Discussion

In this study, we present one of the most comprehensive molecular-level examinations of missense variant effects to date, integrating over 120,000 mutations spanning six biologically distinct classes—pathogenic, benign, driver, passenger, recurrent somatic, and population variants—and annotating each with 34 curated structural, biophysical, and functional descriptors. This unified analytical framework extends prior work that typically focused on a single contrast (e.g., pathogenic vs. population variants) or a limited set of feature types. By jointly characterizing germline and somatic mutations, our work reveals a consolidated molecular landscape that clarifies both shared and context-specific determinants of functional missense variation.

Several notable findings emerge from this integrative analysis. First, features reflecting disruption of structural integrity (e.g., core, β-sheet regions, and ΔΔGfold) and functional sites (e.g., protein–protein, protein–ligand, and nucleic acid binding sites; allosteric and domain regions; ΔΔGbind) consistently distinguish pathogenic and driver mutations from benign or passenger variants. These results reinforce prior observations of functional hotspots but extend them by demonstrating that the same molecular regions are consistently vulnerable across both hereditary and somatic mutation contexts (2528). Features exhibiting mixed-direction associations across mutation types were rare; for example, only posttranslational modification (PTM) sites and helix regions showed opposite enrichment trends between pathogenic and driver mutations, underscoring the overall consistency of molecular determinants underpinning these functional classes.

Second, our stratified analysis of somatic recurrence offers a mechanistically refined framework for driver prioritization. Although recurrence is widely used as a proxy for positive selection in cancer genomics, we find that the majority of recurrent mutations do not exhibit the full spectrum of driver-like molecular signatures. In contrast, only highly recurrent cross-cancer mutations—those arising independently across diverse cancer types—show molecular profiles closely resembling established drivers, supporting the notion that pan-cancer recurrence reflects stronger positive selection (39, 86, 87). In parallel, mutations occurring in driver-pathway genes—genes that function both as cancer drivers and as core pathway regulators—exhibit similarly strong driver-like molecular signatures, reinforcing prior observations that driver genes embedded within key cancer pathways play a pivotal role in tumorigenesis (5, 88). Together, these findings establish a multilayered framework for refining driver discovery beyond recurrence thresholding alone.

Motivated by these mechanistic regularities, we developed MutaPheno, an interpretable random forest classifier that predicts both pathogenic and driver missense mutations using a unified set of 34 molecular descriptors. A central conceptual advance of MutaPheno lies in its cross-domain generalizability: although trained exclusively on germline pathogenic and benign variants, it generalizes effectively to the somatic driver–passenger task. This capability directly addresses two persistent challenges faced by existing computational approaches. First, by relying on intrinsic structural, biophysical, and functional perturbations rather than opaque latent representations, MutaPheno avoids the “black-box” nature of many current models and provides mechanistically interpretable predictions. Second, by learning transferable molecular signatures instead of task-specific statistical signals, it reduces sensitivity to label scarcity and tumor heterogeneity, enabling abundant germline annotations to inform data-scarce cancer contexts.

Benchmarking results support this mechanism-driven transferability. Whereas many pathogenicity predictors—including high-performing deep learning models such as AlphaMissense—exhibit pronounced performance declines when applied to driver-mutation tasks, MutaPheno shows only modest reductions in accuracy, highlighting the stability of mechanistic features across germline and somatic contexts and the convergent molecular signature shared by pathogenic and driver mutations. Importantly, MutaPheno also combines competitive discrimination with robust probability calibration across both tasks, enhancing its practical utility for accurate risk stratification in translational and preclinical settings.

Despite these advances, certain limitations remain, particularly the scarce coverage of several molecular features—including short linear motifs, nucleic acid-binding interfaces, and protein–protein complex structures—due to current annotation constraints. Nevertheless, OR- and DR-based enrichment analyses and feature-attribution profiles indicate that these sparsely annotated features still contribute meaningfully to distinguishing functional variants, suggesting that their biological relevance is underestimated rather than absent. As annotation resources expand and predicted protein–protein and protein–nucleic acid complex structures improve in coverage and reliability, incorporating these additional molecular layers is expected to further enhance the resolution and accuracy of mechanistic models of variant effects.

In summary, this work establishes a unified, interpretable, and transferable framework for mechanistic missense variant interpretation across germline and somatic contexts. By integrating multiscale molecular annotations, large-scale mutational genomics, and interpretable machine learning, we refine assumptions regarding recurrence, reveal convergent molecular principles underlying pathogenic and driver mutations, and introduce a feature-driven modeling paradigm that remains robust across tasks. As structural and functional annotation resources continue to grow, this framework provides a scalable foundation for next-generation precision variant prediction in both inherited disease and oncology.

Supplementary Material

Appendix 01 (PDF)

Dataset S01 (XLSX)

Dataset S02 (XLSX)

pnas.2524289123.sd02.xlsx (34.2KB, xlsx)

Dataset S03 (XLSX)

pnas.2524289123.sd03.xlsx (26.1KB, xlsx)

Dataset S04 (XLSX)

pnas.2524289123.sd04.xlsx (73.6KB, xlsx)

Dataset S05 (XLSX)

Dataset S06 (XLSX)

Dataset S07 (XLSX)

pnas.2524289123.sd07.xlsx (377.9KB, xlsx)

Dataset S08 (XLSX)

Dataset S09 (XLSX)

Acknowledgments

This work was supported by the National Natural Science Foundation of China (Grant No. 32070665), the Basic Research Program of Jiangsu Province (Grant No. BK20251896), and the Priority Academic Program Development of Jiangsu Higher Education Institutions. The funders had no role in the study design; data collection, analysis, or interpretation; the decision to publish; or the preparation of the manuscript.

Author contributions

M.L. designed research; Y.Y., W.S., Y.L., and M.L. performed research; Y.Y. and J.Z. contributed new reagents/analytic tools; Y.Y., Y.L., and M.L. analyzed data; and M.L. wrote the paper.

Competing interests

The authors declare no competing interest.

Footnotes

This article is a PNAS Direct Submission.

Data, Materials, and Software Availability

The MutaPheno program is publicly available on GitHub at https://github.com/minghuilab/MutaPheno. The feature annotations for 16 mutation datasets can also be accessed on GitHub at the same link (93), with each dataset stored in a separate sheet. Each mutation is annotated with 34 molecular features. Additionally, the feature matrices and prediction results for the training set, as well as two independent test sets used in the development and evaluation of the MutaPheno model, can also be accessed on GitHub (93).

Supporting Information

References

  • 1.Zhao F., Zheng L., Goncearenco A., Panchenko A. R., Li M., Computational approaches to prioritize cancer driver missense mutations. Int. J. Mol. Sci. 19, 2113 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Stefl S., Nishi H., Petukh M., Panchenko A. R., Alexov E., Molecular mechanisms of disease-causing missense mutations. J. Mol. Biol. 425, 3919–3936 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Bunn H. F., Pathogenesis and treatment of sickle cell disease. N. Engl. J. Med. 337, 762–769 (1997). [DOI] [PubMed] [Google Scholar]
  • 4.Hardy J., Selkoe D. J., The amyloid hypothesis of Alzheimer’s disease: Progress and problems on the road to therapeutics. Science 297, 353–356 (2002). [DOI] [PubMed] [Google Scholar]
  • 5.Vogelstein B., et al. , Cancer genome landscapes. Science 339, 1546–1558 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Stratton M. R., Campbell P. J., Futreal P. A., The cancer genome. Nature 458, 719–724 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Biankin A. V., et al. , Pancreatic cancer genomes reveal aberrations in axon guidance pathway genes. Nature 491, 399–405 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Chiang Y. T., et al. , The function of the mutant p53–R175H in cancer. Cancers (Basel) 13, 4088 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Ascierto P. A., et al. , The role of BRAF V600 mutation in melanoma. J. Transl. Med. 10, 85 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Liu X., Jian X., Boerwinkle E., dbNSFP: A lightweight database of human nonsynonymous SNPs and their functional predictions. Hum. Mutat. 32, 894–899 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Liu X., Li C., Mou C., Dong Y., Tu Y., dbNSFP v4: A comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome Med. 12, 103 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Ng P. C., Henikoff S., SIFT: Predicting amino acid changes that affect protein function. Nucleic Acids Res. 31, 3812–3814 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Adzhubei I. A., et al. , A method and server for predicting damaging missense mutations. Nat. Methods 7, 248–249 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Ioannidis N. M., et al. , REVEL: An ensemble method for predicting the pathogenicity of rare missense variants. Am. J. Hum. Genet. 99, 877–885 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Carter H., Douville C., Stenson P. D., Cooper D. N., Karchin R., Identifying Mendelian disease genes with the variant effect scoring tool. BMC Genomics 3, S3 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Cheng J., et al. , Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492 (2023). [DOI] [PubMed] [Google Scholar]
  • 17.Carter H., et al. , Cancer-specific high-throughput annotation of somatic mutations: Computational prediction of driver missense mutations. Cancer Res. 69, 6660–6667 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Kumar R. D., Swamidass S. J., Bose R., Unsupervised detection of cancer driver mutations with parsimony-guided learning. Nat. Genet. 48, 1288–1294 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Tokheim C., Karchin R., CHASMplus reveals the scope of somatic missense mutations driving human cancers. Cell Syst. 9, 9–23.e28 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wang H., et al. , AI-driver: An ensemble method for identifying driver mutations in personal cancer genomes. NAR Genom. Bioinform. 2, lqaa084 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ostroverkhova D., Sheng Y., Panchenko A., Are next-generation pathogenicity predictors applicable to cancer? J. Mol. Biol. 436, 168644 (2024). [DOI] [PubMed] [Google Scholar]
  • 22.Ostroverkhova D., Przytycka T. M., Panchenko A. R., Cancer driver mutations: Predictions and reality. Trends Mol. Med. 29, 554–566 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Chen H., et al. , Comprehensive assessment of computational algorithms in predicting cancer driver mutations. Genome Biol. 21, 43 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wierbowski S. D., Fragoza R., Liang S., Yu H., Extracting complementary insights from molecular phenotypes for prioritization of disease-associated mutations. Curr. Opin. Syst. Biol. 11, 107–116 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Li M., et al. , Balancing protein stability and activity in cancer: A new approach for identifying driver mutations affecting CBL ubiquitin ligase activation. Cancer Res. 76, 561–571 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Iqbal S., et al. , Comprehensive characterization of amino acid positions in protein structures reveals molecular effect of missense variants. Proc. Natl. Acad. Sci. U.S.A. 117, 28201–28211 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cheng F., et al. , Comprehensive characterization of protein-protein interactions perturbed by disease mutations. Nat. Genet. 53, 342–353 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Laddach A., Ng J. C. F., Fraternali F., Pathogenic missense protein variants affect different functional pathways and proteomic features than healthy population variants. PLoS Biol. 19, e3001207 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Gerasimavicius L., Livesey B. J., Marsh J. A., Loss-of-function, gain-of-function and dominant-negative mutations have profoundly different effects on protein structure. Nat. Commun. 13, 3895 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Pires D. E., Chen J., Blundell T. L., Ascher D. B., In silico functional dissection of saturation mutagenesis: Interpreting the relationship between phenotypes and changes in protein stability, interactions and activity. Sci. Rep. 6, 19848 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Landrum M. J., et al. , ClinVar: Public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res. 42, D980–D985 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.The UniProt Consortium, UniProt: The universal protein knowledgebase in 2023. Nucleic Acids Res. 51, D523–D531 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sevim Bayrak C., et al. , Identification of discriminative gene-level and protein-level features associated with pathogenic gain-of-function and loss-of-function variants. Am. J. Hum. Genet. 108, 2301–2318 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Stenson P. D., et al. , The human gene mutation database (HGMD(®)): Optimizing its use in a clinical diagnostic or research setting. Hum. Genet. 139, 1197–1207 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Chakravarty D., et al. , OncoKB: A precision oncology knowledge base. JCO Precis. Oncol. 1, 1–16 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Tamborero D., et al. , Cancer genome interpreter annotates the biological and clinical relevance of tumor alterations. Genome Med. 10, 25 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Brown A. L., Li M., Goncearenco A., Panchenko A. R., Finding driver mutations in cancer: Elucidating the role of background mutational processes. PLoS Comput. Biol. 15, e1006981 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Ellrott K., et al. , Scalable open science approach for mutation calling of tumor exomes using multiple genomic pipelines. Cell Syst. 6, 271–281.e277 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Bailey M. H., et al. , Comprehensive characterization of cancer driver genes and mutations. Cell 173, 371–385.e18 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Lek M., et al. , Analysis of protein-coding genetic variation in 60, 706 humans. Nature 536, 285–291 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Chang M. T., et al. , Identifying recurrent mutations in cancer reveals widespread lineage diversity and mutational specificity. Nat. Biotechnol. 34, 155–163 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Nightingale A., et al. , The proteins API: Accessing key integrated protein and genome information. Nucleic Acids Res. 45, W539–W544 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Zaru R., Orchard S., UniProt tools: BLAST, align, peptide search, and ID mapping. Curr. Protoc. 3, e697 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Sievers F., et al. , Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Mol. Syst. Biol. 7, 539 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Martínez-Jiménez F., et al. , A compendium of mutational cancer driver genes. Nat. Rev. Cancer 20, 555–572 (2020). [DOI] [PubMed] [Google Scholar]
  • 46.Wang T., et al. , OncoVar: An integrated database and analysis platform for oncogenic driver variants in cancers. Nucleic Acids Res. 49, D1289–D1301 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Repana D., et al. , The network of cancer genes (NCG): A comprehensive catalogue of known and candidate cancer genes from cancer sequencing screens. Genome Biol. 20, 1 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Sondka Z., et al. , The COSMIC Cancer Gene Census: Describing genetic dysfunction across all human cancers. Nat. Rev. Cancer 18, 696–705 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Rehm H. L., et al. , ClinGen–The clinical genome resource. N. Engl. J. Med. 372, 2235–2242 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Subramanian A., et al. , Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. Proc. Natl. Acad. Sci. U.S.A. 102, 15545–15550 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Raju R., et al. , NetSlim: High-confidence curated signaling maps. Database (Oxford) 2011, bar032 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Kabsch W., Sander C., Dictionary of protein secondary structure: Pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22, 2577–2637 (1983). [DOI] [PubMed] [Google Scholar]
  • 53.Liu X., et al. , Unraveling allosteric landscapes of allosterome with ASD. Nucleic Acids Res. 48, D394–D401 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Song K., et al. , Improved method for the identification and validation of allosteric sites. J. Chem. Inf. Model. 57, 2358–2363 (2017). [DOI] [PubMed] [Google Scholar]
  • 55.Meyer M. J., et al. , Interactome INSIDER: A structural interactome browser for genomic studies. Nat. Methods 15, 107–114 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Krivák R., Hoksza D., P2Rank: Machine learning based tool for rapid and accurate prediction of ligand binding sites from protein structure. J. Cheminform. 10, 39 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Xia Y., Xia C. Q., Pan X., Shen H. B., Graphbind: Protein structural context embedded rules learned by hierarchical graph neural networks for recognizing nucleic-acid-binding residues. Nucleic Acids Res. 49, e51 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Tak Leung R. W., Jiang X., Chu K. H., Qin J., ENPD–A database of eukaryotic nucleic acid binding proteins: Linking gene regulations to proteins. Nucleic Acids Res. 47, D322–D329 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Liao J. Y., et al. , EuRBPDB: A comprehensive resource for annotation, functional and oncological investigation of eukaryotic RNA binding proteins (RBPs). Nucleic Acids Res. 48, D307–D313 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Hornbeck P. V., et al. , PhosphoSitePlus, 2014: Mutations, PTMs and recalibrations. Nucleic Acids Res. 43, D512–D520 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Chung C. R., et al. , dbPTM 2025 update: Comprehensive integration of PTMs and proteomic data for advanced insights into cancer research. Nucleic Acids Res. 53, D377–D386 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Piovesan D., et al. , MOBIDB in 2025: Integrating ensemble properties and function annotations for intrinsically disordered proteins. Nucleic Acids Res. 53, D495–D503 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Aspromonte M. C., et al. , DisProt in 2024: Improving function annotation of intrinsically disordered proteins. Nucleic Acids Res. 52, D434–D441 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Kumar M., et al. , ELM-the eukaryotic linear motif resource-2024 update. Nucleic Acids Res. 52, D442–D455 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Blum M., et al. , InterPro: The protein sequence classification resource in 2025. Nucleic Acids Res. 53, D444–D456 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Chen Y., et al. , PremPS: Predicting the impact of missense mutations on protein stability. PLoS Comput. Biol. 16, e1008543 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Zhang N., et al. , MutaBind2: Predicting the impacts of single and multiple mutations on protein-protein interactions. iScience 23, 100939 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Berman H. M., et al. , The protein data bank. Nucleic Acids Res. 28, 235–242 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Jumper J., et al. , Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Varadi M., et al. , AlphaFold protein structure database in 2024: Providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 52, D368–D375 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Virtanen P., et al. , SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nat. Methods 17, 261–272 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Benjamini Y., Hochberg Y., Controlling the false discovery rate: A practical and powerful approach to multiple testing. J. R. Stat. Soc. B 57, 289–300 (1995). [Google Scholar]
  • 73.Altman D. G., Bland J. M., Interaction revisited: The difference between two estimates. BMJ 326, 219 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Deák G., Cook A. G., Missense variants reveal functional insights into the human ARID family of gene regulators. J. Mol. Biol. 434, 167529 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Thomas P. D., et al. , PANTHER: Making genome-scale phylogenetics accessible to all. Protein Sci. 31, 8–22 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Thul P. J., et al. , A subcellular map of the human proteome. Science 356, eaal3321 (2017). [DOI] [PubMed] [Google Scholar]
  • 77.Jagota M., et al. , Cross-protein transfer learning substantially improves disease variant prediction. Genome Biol. 24, 182 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Li M., Simonetti F. L., Goncearenco A., Panchenko A. R., MutaBind estimates and interprets the effects of sequence variants on protein-protein interactions. Nucleic Acids Res. 44, W494–W501 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Sun T., Chen Y., Wen Y., Zhu Z., Li M., PremPLI: A machine learning model for predicting the effects of missense mutations on protein-ligand interactions. Commun. Biol. 4, 1311 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Zheng F., Jiang X., Wen Y., Yang Y., Li M., Systematic investigation of machine learning on limited data: A study on predicting protein-protein binding strength. Comput. Struct. Biotechnol. J. 23, 460–472 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Steinegger M., Söding J., MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026–1028 (2017). [DOI] [PubMed] [Google Scholar]
  • 82.Hagberg A. A., Schult D. A., Swart P., Hagberg J., “Exploring network structure, dynamics, and function using NetworkX” in Proceedings of the Python in Science Conference, Varoquaux G., Vaught T., Millman J., Eds. (SciPy, Pasadena, CA, 2008), pp. 11–15. [Google Scholar]
  • 83.Steyerberg E. W., et al. , Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology 21, 128–138 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Steyerberg E. W., Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating (Springer, New York, 2009). [Google Scholar]
  • 85.Reimand J., Wagih O., Bader G. D., Evolutionary constraint and disease associations of post-translational modification sites in human genomes. PLoS Genet. 11, e1004919 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Priestley P., et al. , Pan-cancer whole-genome analyses of metastatic solid tumours. Nature 575, 210–216 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Hess J. M., et al. , Passenger hotspot mutations in cancer. Cancer Cell 36, 288–301.e14 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.The ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium, Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Leontiadou H., Galdadas I., Athanasiou C., Cournia Z., Insights into the mechanism of the PIK3CA E545K activating mutation using MD simulations. Sci. Rep. 8, 15544 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Chen X., et al. , Mutant p53 in cancer: From molecular mechanism to therapeutic modulation. Cell Death Dis. 13, 974 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Roviello G., et al. , Advances in anti-BRAF therapies for lung cancer. Invest. New Drugs 39, 879–890 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Kiel C., Benisty H., Lloréns-Rico V., Serrano L., The yin-yang of kinase activation and unfolding explains the peculiarity of Val600 in the activation segment of BRAF. Elife 5, e12814 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Yang Y., et al., MutaPheno: An interpretable framework for predicting pathogenic and cancer driver mutations [Data set]. GitHub. https://github.com/minghuilab/MutaPheno. Deposited 15 January 2026.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix 01 (PDF)

Dataset S01 (XLSX)

Dataset S02 (XLSX)

pnas.2524289123.sd02.xlsx (34.2KB, xlsx)

Dataset S03 (XLSX)

pnas.2524289123.sd03.xlsx (26.1KB, xlsx)

Dataset S04 (XLSX)

pnas.2524289123.sd04.xlsx (73.6KB, xlsx)

Dataset S05 (XLSX)

Dataset S06 (XLSX)

Dataset S07 (XLSX)

pnas.2524289123.sd07.xlsx (377.9KB, xlsx)

Dataset S08 (XLSX)

Dataset S09 (XLSX)

Data Availability Statement

The MutaPheno program is publicly available on GitHub at https://github.com/minghuilab/MutaPheno. The feature annotations for 16 mutation datasets can also be accessed on GitHub at the same link (93), with each dataset stored in a separate sheet. Each mutation is annotated with 34 molecular features. Additionally, the feature matrices and prediction results for the training set, as well as two independent test sets used in the development and evaluation of the MutaPheno model, can also be accessed on GitHub (93).


Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES