ABSTRACT
Interpreting the functional impact of genomic variants remains a major challenge in precision oncology, particularly in breast cancer, where many variants of unknown significance lack clear therapeutic guidance. Current annotation strategies focus on frequent driver mutations, leaving rare or understudied variants unclassified and clinically uninformative. Here, we present an Artificial Intelligence/Machine Learning (AI/ML)‐driven framework that systematically identifies variants associated with key breast cancer phenotypes, including ESR1 and EZH2 activity, by integrating genomic, transcriptomic, structural, and drug response data. Using CCLE/DepMap and TCGA datasets, we analyzed > 12,000 variants across breast cancer genomes, identifying structurally clustered mutations that share functional consequences with well‐characterized oncogenic drivers. This approach reveals that mutations in PIK3CA, TP53, and other genes strongly associate with ESR1 signaling, challenging conventional assumptions about endocrine therapy response. Additionally, EZH2‐associated variants emerge in unexpected genomic contexts, suggesting new targets for epigenetic therapies. By shifting from frequency‐based to structure‐informed classification, we expand the set of potentially actionable mutations, enabling improved patient stratification and drug repurposing strategies. This work provides a scalable, clinically relevant method to accelerate variant annotation, offering new insights into drug sensitivity and resistance mechanisms. Future validation efforts will refine these predictions and integrate clinical outcomes to guide personalized treatment strategies. Our findings highlight the transformative potential of AI/ML in redefining cancer variant interpretation, bridging the gap between genomics, functional biology, and precision medicine.
Keywords: algorithms, cancers, classification, computational biology, genomics, proteins, therapeutics
Study Highlights.
- What is the current knowledge on the topic?
-
○Genomic profiling increasingly informs breast cancer treatment decisions, but most clinical actionability is limited to frequent driver mutations (e.g., PIK3CA, TP53, ESR1). The majority of variants remain unclassified as “variants of unknown significance” (VUS), and current annotation methods are slow, relying on manual curation or low‐throughput assays that fail to capture rare mutations.
-
○
- What question did this study address?
-
○Can an AI/ML‐based framework systematically identify rare or understudied mutations that influence key breast cancer phenotypes (specifically ESR1 and EZH2 activity) by leveraging structural and functional features?
-
○
- What does this study add to our knowledge?
-
○Analyzed > 12,000 variants across breast cancer using integrated multi‐omic and structural data
-
○Identified ESR1‐associated mutation clusters, expanding the landscape of potentially actionable variants
-
○Discovered EZH2‐associated mutations, highlighting new targets for epigenetic therapies
-
○Demonstrated that 3D spatial clustering enables functional annotation of rare and uncharacterized mutations
-
○
- How might this change clinical pharmacology or translational science?
-
○Introduces a scalable, AI‐powered approach to accelerate variant annotation and functional classification.
-
○Enables more precise patient stratification based on predicted regulatory and drug response profiles
-
○Expands the scope of targeted therapy by providing a data‐driven framework for prioritizing rare mutations.
-
○
1. Introduction
Genomic profiling plays an essential role in breast cancer treatment selection, particularly in determining targeted therapies for hormone receptor‐positive tumors. Patients with estrogen receptor‐positive (ER+) breast cancer are treated with endocrine therapy [1, 2], including tamoxifen, aromatase inhibitors (e.g., letrozole, anastrozole, exemestane) and selective estrogen receptor degraders (e.g., fulvestrant), which suppress estrogen‐estrogen receptor signaling [3, 4, 5]. However, responses remain variable [6], with de novo and acquired resistance leading to many treatment failures [7]. While some ESR1 mutations (e.g., Y537S, D538G) are known to drive endocrine resistance [8], this is not the cause in most patients, and the underlying genomic factors contributing to resistance for many patients remain unclear [6]. Beyond endocrine therapy, targeted approaches such as PI3K, AKT, and mTOR inhibitors have demonstrated efficacy in patients with specific genomic alterations [9]. Yet, not all patients with these mutations respond uniformly, and many patients lack clearly actionable variants [9].
A major barrier to optimizing treatment stratification is the high number of variants of unknown significance in breast cancer [10], particularly those that may influence ESR1 signaling, epigenetic regulation, and drug response. While frequent mutations in BRCA1/2, TP53, and PIK3CA are well‐characterized [11, 12, 13, 14], many low‐frequency or previously unannotated variants remain poorly understood, leaving clinicians without clear guidance on how they influence treatment decisions [15]. Current annotation strategies rely on slow, manual curation methods or low‐throughput functional assays, which focus predominantly on frequent driver mutations. As a result, patients who lack well characterized mutations are often treated with conventional therapies [16], potentially missing opportunities for more precise, biomarker‐driven treatment selection.
To address this, we developed a scalable, high‐throughput approach to systematically classify variants based on functional impact rather than occurrence frequency. The ability to identify and prioritize mutations that drive key breast cancer phenotypes, such as estrogen signaling or epigenetic reprogramming, could significantly refine treatment strategies. By leveraging Artificial Intelligence/Machine Learning (AI/ML) and multi‐omic data resources such as CCLE/DepMap (Cancer Cell Line Encyclopedia/Dependency Map) [17, 18] and TCGA (The Cancer Genome Atlas) [19], we systematically link genomic variants to molecular phenotypes, including transcriptional activity, gene dependencies, and drug response profiles. This allows us to rank and prioritize variants based on not just their prevalence, but also their functional impact, druggability, and clinical outcome.
In this study, we apply our AI/ML‐driven framework to uncover clusters of mutations that are functionally linked to the regulatory activities of two key drivers of breast cancer progression and treatment response: ESR1 and EZH2. Rather than grouping variants based on how often they appear in patient cohorts, we map them onto the 3D structures of proteins, using AlphaFold‐predicted models [20], to evaluate their spatial relationships. This structure‐informed strategy allows us to identify clusters of mutations that converge in critical regions of the protein, even if those mutations are rare or scattered across different patients. This opens two important opportunities: (1) spatially co‐located mutations may disrupt the same functional domain, yielding similar biological effects, and (2) rare variants that fall into the same structural neighborhood as well‐characterized drivers may share pathogenic mechanisms, providing a new way to interpret their clinical relevance. By reframing how we group and prioritize mutations, we expand the scope of potentially actionable variants beyond traditional frequency‐based annotations.
Our analysis reveals that specific regions within proteins like PIK3CA and TP53 act as predictive control points for estrogen and epigenetic signaling. Using our structure‐informed AI/ML approach, we identify protein regions where mutations are highly predictive (up to 95% accuracy) of elevated ESR1 or EZH2 activity. These regions can function as clinical biomarkers of pathway activation; if a patient carries a mutation in one of these regions, we can reliably infer that their tumor exhibits high ESR1 or EZH2 activity. This framework is clinically significant because elevated ESR1 activity strongly influences response to endocrine therapies, while EZH2 activation may signal sensitivity to epigenetic treatments. By identifying these predictive protein regions, we offer a powerful new strategy for anticipating treatment response and guiding precision therapy, particularly for patients who lack well‐characterized genomic alterations. This approach also opens the door to rational design of therapies that target specific signaling states inferred directly from mutational profiles.
2. Methods
2.1. AI/ML Framework for Variant Clustering and Regulatory Activity Prediction
We developed VAMOS (Variant Annotation through Multi‐Omic Signatures) [21], a machine learning framework that integrates single‐nucleotide variant (SNV) data, RNA expression profiles, and AlphaFold‐predicted 3D protein structures [20] to link protein‐coding mutations with downstream transcriptional phenotypes (Figure 1a).
FIGURE 1.

(a) Overview of the VAMOS pipeline to identify clinically relevant variants. Predictive analytics are performed on the algorithm to determine the quality of predictions. The output of the algorithm is clinically relevant dense variant clusters. These clusters are used to prioritize treatment options for known frequent variants, and identify novel treatment options for rare and understudied variants. (b) This study looks at target pathway activity for ESR1 (encoding ER) and EZH2. ESR1 has over 200 known target genes. ER is a steroid hormone that binds chromatin to regulate gene expression in response to estrogen. ER‐activity is a primary determinant of intrinsic breast cancer biology, measured by gene expression profiles and clinical outcomes (recurrence and survival), and ER‐inhibiting drugs are first line therapy for ER‐positive tumors [22]. EZH2 has over 1000 known target genes, and is involved in epigenetic regulation primarily via chromatin remodeling, making it an exciting novel target for drug discovery [23]. Mutations in these proteins (or other proteins) can lead to alterations in these pathways, and ultimately resistance to chemotherapeutics and hormone therapies. (c) Variants from CCLE/DepMap [17, 18] with protein structural information, transcriptional information, and drug sensitivity information in the form of AUC values. 12,398 variants have all three data types available in public repositories.
To identify functionally meaningful variant clusters, we applied density‐based clustering to group spatially proximal mutations within protein structures across samples from CCLE/DepMap [17, 18] and TCGA [19]. Each cluster represents a protein region where mutations co‐localize in three‐dimensional space, regardless of their population frequency.
To assess the functional impact of these clusters, VAMOS predicts the activity of key transcriptional regulators based on the expression of curated target genes. We focused on ESR1 and EZH2, two critical regulators of breast cancer biology and therapy response [22, 24], with gene regulatory activities quantified using curated gene sets derived from literature, perturbation datasets, and co‐expression networks (see Figure 1b). We computed gene set enrichment analysis (GSEA) scores [25, 26] for ESR1 and EZH2 target gene expression, based on over 200 known targets, providing a proxy for regulator activity across 69 breast cancer cell lines and 965 tumor samples (Figure S1).
Each variant cluster was linked to transcriptional phenotypes by calculating the log‐odds that mutations within the cluster co‐occurred with elevated or depleted ESR1 or EZH2 activity. This scoring allowed us to rank clusters of mutations (and the proteins they occur in) by the strength of their association with either ESR1 or EZH2 transcriptional regulation.
2.2. Machine Learning Classification and Validation
ML models used in this study were adapted from the original VAMOS framework [21, 27]. Each variant was described by features including its Cartesian coordinates, cluster membership, and spatial density. To train the classifier, we labeled mutations as associated (1) or not associated (0) with the phenotype of interest (e.g., ESR1 activity, gene dependency), based on whether the GSEA or dependency score was ≥ 1 SD above the mean. The training dataset was balanced via oversampling.
Models were trained on one‐third of the data and tested on the remaining two‐thirds, using fivefold cross‐validation to evaluate performance. Model metrics included accuracy, precision, recall, F1‐score, AUC, and Matthews Correlation Coefficient (MCC) (Table S1 and Figure S2).
This integrated framework allowed us to identify and prioritize variant clusters most strongly predictive of ESR1 or EZH2 transcriptional activity, enabling downstream analyses of genetic dependency and drug sensitivity.
2.3. Data Retrieval
Mutation data was obtained from CCLE/DepMap v24Q4 [17, 18] and TCGA v2 [19]. Gene expression (RNA sequencing) and CRISPR‐mediated knockdown data were obtained from CCLE/DepMap. CRISPR data was obtained as a gene dependency score, where higher scores indicate increased dependency on ESR1/EZH2. Drug sensitivity data was obtained from the Cancer Therapeutics Research Portal (CTRP) [20, 28, 29].
2.4. Classification of Rare vs. Frequent Mutations
We analyzed 69 breast cancer cell lines and 965 tumor samples from the CCLE/DepMap and TCGA datasets to characterize the frequency of rare somatic mutations in our breast cancer datasets. Variants from CCLE/DepMap and TCGA were mapped to annotation databases, including OncoKB [30] and ClinVar [31]. Mutations were classified as “frequent” if they occurred in > 8% of breast cancer cell lines, and “rare” otherwise [32], and as “known drug targets” if they had a targeted therapeutic with an OncoKB evidence level of 3A or higher [30].
2.5. Gene Set Enrichment Analysis
Single‐sample Gene Set Enrichment Analysis (GSEA) [33] was performed on gene targets of interest (e.g., ESR1 and EZH2) using mutation and gene expression data. The ESR1 regulatory network was defined using an experimentally curated MSigDB [25] gene set [34], containing > 200 genes. The EZH2 regulatory network was defined using an experimentally curated MSigDB [25] gene set [26], containing > 1000 genes. Samples were ranked as “ESR1/EZH2 High” if they had GSEA scores greater than one standard deviation above the mean. Additional details about GSEA cutoffs are in the Methods S1.
2.6. Mapping to the Protein Data Bank and AlphaFold
Variants were mapped to the Protein Data Bank per residue using an adapted version of the mmtf‐Python workflow [35], and to AlphaFold v2 using a previously developed workflow [36]. Structural information was available for 17,852 variants across 8886 genes for 69 breast cancer cell lines from CCLE/DepMap, and 38,797 variants across 12,696 genes for 965 tumors in TCGA.
2.7. Statistical Analysis of Clusters
We classified mutations that strongly associated with upregulated (or downregulated) transcriptional activity of ESR1 or EZH2 as class “1” mutations. In contrast, mutations that did not associate with upregulated (or downregulated) transcriptional activities of ESR1 or EZH2 were classified as class “0.” The log odds that a cluster would contain a class “1” mutation was calculated using the formula:
where Px is the probability of a mutation being labeled as class “1” and Py is the probability of a mutation being labeled as class “0.”
We ranked mutation clusters and the proteins they occur in using a priority score. Priority scores were calculated by averaging log odds scores across all available data type combinations for each cluster.
2.8. Gene Dependency and Expression Analyses
We assessed gene dependency for ESR1 and EZH2 using DepMap CRISPR‐Cas9 loss‐of‐function screens across 69 breast cancer cell lines [17, 18], alongside matched RNA‐sequencing expression data. This analysis allowed us to determine whether specific variant clusters, particularly in genes like PIK3CA, were associated with increased dependency on ESR1 or EZH2 activity (Figures S3–S6).
2.9. Drug Sensitivity Analysis
To assess drug response differences associated with variant clusters, we analyzed Cancer Therapeutics Response Portal (CTRP) data [28, 29, 30], which includes 545 drug response profiles across 40 breast cancer cell lines. Each cell line was annotated for the presence or absence of high‐priority (class “1”) variant clusters.
For each drug, we compared the area under the curve (AUC) between cluster‐containing and non‐cluster‐containing cell lines. Statistically significant differences in drug response (p < 0.05, Wilcoxon Rank Sum test) were used to identify cluster‐associated sensitivities (Tables S2–S19; Figure 3b–d).
FIGURE 3.

(a) Three dense variant clusters in the PIK3CA protein are identified via our ML method. Cluster 1 (green) and Cluster 2 (yellow) variants are associated with upregulation of the ESR1 pathway. Cluster 3 (blue) variants are associated with upregulation of the EZH2 pathway. Variants with known clinical significance are highlighted in red. The center pie charts show the number of frequent versus rare variants and variants that are known drug targets versus variants with unknown drug interactions from OncoKB present in PIK3CA, and in each of the ML identified clusters. The likelihood that a cluster is associated with a given pathway is shown on the left. (b) A schema showing the process for finding drugs with statistically different effects on cell lines containing the ML identified clusters, shown here for Cluster 3. Cell lines are sorted into two groups: Those containing Cluster 3 variants, and those not containing Cluster 3 variants. Statistical analysis is performed to find all drugs that have lower AUC values for Cluster 3 cell lines and statistically different distributions compared to cell lines with no Cluster 3 mutations. (c) Different regions of the protein are associated with increased sensitivity towards different drugs. Cell lines with Cluster 1 variants have greater sensitivity towards AKT1 targeting and mediating drugs. Cell lines with Cluster 2 variants do not have associations with any drugs. Cell lines with Cluster 3 variants have greater sensitivity towards mTORC targeting and mediating drugs. (d) Heat maps showing the p‐values generated by the statistical analysis described in part b for Cluster 1 and Cluster 3. Additional p values from random permutation analysis are also shown. (e) These results can be used to inform precision medicine treatments for patients with mutations in the indicated clusters.
In proteins that contained both class “1” clusters and other dense clusters, we used the same test to compare sensitivity between these subgroups. To assess the likelihood of observing these results by chance, we performed permutation testing: for each protein, n random mutations (equal to the size of the class “1” cluster) were selected, and the test was repeated. Permutation trials (n = 100, 1000, 10,000, 100,000) were averaged to determine empirical significance (Tables S2–S19).
2.10. Prioritization of Clusters With Unknown or Unannotated Variants
To prioritize novel variant clusters not found in major clinical databases, we computed a log‐based priority score by averaging each cluster's log‐odds scores across all data types (transcriptional activity, dependency, and drug response) (Tables S20 and S21; Figure 5b). Based on these scores and available drug sensitivity information, proteins were categorized into three evidence tiers: High evidence, some drugs; Some evidence, potential drugs; No evidence, no drugs.
FIGURE 5.

(a) There are > 12,000 variants across > 8000 proteins total with transcriptional data, protein structure data, and drug response data. Our workflow selected 909 variants across 44 proteins with a high likelihood of association with ESR activity, and 908 variants across 38 proteins with a high likelihood of association with EZH2 activity. The breakdown of known versus unknown variants and proteins for both ESR1 and EZH2 is shown. The logodds for each cluster identified by ML is shown on the right. (b) This workflow can be used to prioritize new targets for study. Here, we show targets with a high logodds based priority score for future study. Targets are sorted into three bins. Bin 1 contains targets identified via our method that already have high evidence for impact in breast cancer, and already have known drugs. Bin 2 contains targets that have some evidence for impact in breast cancer, and have potential drugs. Bin 3 contains targets with no current evidence or drugs. (c) Clusters in DST and MGAM, both Bin 3 targets, associated with ESR1 pathway changes. (d) Clusters in KMT2C, a Bin 2 target, and TUBA1A, a Bin 3 target, associated with EZH2 pathway changes. (e) Top 5 drugs identified using our method that have higher impact on variants in ML identified DST or MGAM clusters. (f) Top drugs identified using our method that have higher impact on variants in ML identified KMT2C or TUBA1A clusters.
Most proteins with unknown variant clusters fell into the third tier, reflecting their limited prior characterization but highlighting opportunities for novel target discovery.
2.11. Identification of Statistically Significant Drug Associations
To further evaluate the biological significance of these variant clusters, we integrated functional data (including clinical annotation data, gene dependency data, and drug response datasets) to assess whether the identified clusters were associated with variants of known significance, functional dependencies, or therapeutic vulnerabilities. Specifically, we used CCLE/DepMap's RNA sequencing [17, 18] and CRISPR‐mediated gene dependency data [17, 18] for ESR1 and EZH2 across 69 breast cancer cell lines.
2.12. Code Access
All code is available at github.com/Brunk‐Lab/VAMOS_precision_oncology.
3. Results
3.1. AI‐Driven Identification of Protein Regions Predictive of ESR1 and EZH2 Network Activity
Using our AI/ML framework, VAMOS, we identified 395 variant clusters across 346 proteins that are associated with upregulated or downregulated ESR1 or EZH2 activity (Figure 1c). To estimate the activity of these regulators, we quantified expression levels of experimentally pre‐defined sets of target genes (e.g., any gene whose transcription is influenced by ESR1 or EZH2; see Methods). Variant clusters represent spatially localized regions within proteins (defined by a high density of missense mutations occurring within a 15‐Å radius) where mutations in multiple cell lines or patient samples co‐occur. While we cannot assume that these mutations are causal, their strong co‐occurrence with ESR1 or EZH2 regulatory activities suggests that these regions may serve as predictive biomarkers of these gene regulatory networks.
Across these clusters, we observed a strong association between mutation location and regulatory activity, quantified by a log‐odds score that reflects how consistently mutations in a given region coincide with elevated ESR1 or EZH2 activity. High log‐odds scores (Figure S7) indicate that most samples with mutations in that region also show increased ESR1 or EZH2 activity, making these clusters strong predictive markers of regulatory state. Based on this analysis, we identified 96 ESR1‐associated variant clusters (312 mutations) across 52 proteins and 100 EZH2‐associated clusters (634 mutations) across 41 proteins. These regions were tightly grouped in 3D protein space, with an average of four mutations within a 13–14 Å radius. Many of the top‐ranked regions fell within known cancer genes, while others highlighted previously uncharacterized candidates, suggesting new avenues for stratification and therapeutic targeting. The top 10 ESR1‐associated clusters were found in genes including PIK3CA, TP53, NF1, and MAP3K1, suggesting that mutations in these regions are highly predictive of ESR1 signaling activity, independent of their frequency across patients. Similarly, top‐ranked EZH2‐associated clusters were found in ARID1A, SMARCA4, and KMT2D, supporting their potential role in shaping epigenetic regulatory states relevant to therapy response.
These structure‐informed clusters offer a new framework for prioritizing rare or unclassified variants based on their phenotypic impact, with direct implications for drug targeting and patient stratification.
3.2. Rare Variant Annotation Through Structural Clustering
Cancer mutations follow a long‐tail distribution, with most being rare and poorly characterized [21] (Figure 2a). We hypothesized that some of these rare variants may be functionally important if they occur in close structural proximity to well‐known, disease‐relevant mutations. For instance, rare mutations located near PIK3CA mutation hotspots (like H1047R or E545K) may disrupt the same functional region and thereby influence drug response similarly (Figure 2b).
FIGURE 2.

(a) The “long tail” of distribution in cancer variants refers to a well known pattern in which a small number of genetic mutations are highly recurrent across patients, while a vast number of mutations occur infrequently. (b) Frequent and rare variants may occur in the same dense cluster on a protein. Frequent variants are well studied and may have drugs that target them specifically. It is not yet known whether rare variants in proximity of frequent variants have the same biological and clinical relevance. (c) The percentage of variants that are frequently occurring (left) or known drug targets (right) across all breast cancer cell lines and tumors. (d) The percentage of variants that are frequently occurring or in the same cluster (shown by darker red) as a frequently occurring variant (left) or in the same cluster (shown by darker green) as a known drug target variant (right) across all breast cancer cell lines and tumors. (e) The frequency of clinical occurrence for all variants in OncoKB [30] and ClinVar [31], variants present in ML identified dense clusters are highlighted in black (left). The frequency of known drug target variants for all variants in OncoKB [30] and ClinVar [31], variants present in ML identified dense clusters are highlighted in black (right). (f) The number of frequently occurring variants in each ML identified dense clusters (left). The number of known drug target variants in each ML identified dense cluster (right).
Our global analysis of 14,456 breast cancer variants revealed that the vast majority (94.9%; n = 13,724) are rare, with only 5.1% (n = 732) classified as frequent (Figure 2c). Among the rare variants, we found that 35.6% (n = 4885) cluster in three‐dimensional space with frequent, well‐characterized mutations, and 23.4% (n = 3211) fall within the same structural regions as mutations already linked to targeted therapies (Figure 2d). These findings suggest that a substantial portion of rare mutations occupy biologically meaningful regions of proteins, reinforcing the idea that spatial context can provide functional clues, even in the absence of prior annotation.
Rare variants associated with ESR1 or EZH2 activity were not concentrated in a single region or protein but instead were distributed across a wide range of clusters, spanning diverse proteins and structural contexts (Figure 2e). Many of these clusters also contained well‐characterized, frequent mutations, including those already linked to therapeutic response, suggesting that structural proximity may offer a path to infer the functional impact of rare variants. Rather than clustering independently, these rare variants often co‐occur in 3D space with mutations that have known biological or clinical significance (Figure 2f), providing a rich framework for functional extrapolation. This finding highlights the potential to reclassify uncharacterized mutations by leveraging their position within the protein structure relative to better‐understood variants, opening new avenues for biomarker discovery and drug repurposing.
3.3. PIK3CA Variant Clusters Associate With ESR1 and EZH2 Activity and Differential Drug Sensitivities
PIK3CA emerged as the top‐ranked protein in our analysis, with three distinct variant clusters showing strong associations with either ESR1 or EZH2 transcriptional activity (Figure 3a). These clusters had the highest log‐odds scores in the dataset, indicating that samples harboring mutations in these regions were significantly more likely to exhibit upregulated ESR1 or EZH2 network activity.
Despite belonging to the same protein, these clusters had functionally distinct regulatory impacts. Clusters 1 and 2 were associated with increased ESR1 activity, with average activity scores 1.2 standard deviations above the mean in cell lines carrying mutations in these regions. In contrast, Cluster 3 was uniquely linked to EZH2 upregulation, underscoring that different structural regions of PIK3CA can influence divergent transcriptional programs. These findings suggest that specific regions of PIK3CA act as regulatory control points, and that perturbations in these areas may differentially rewire downstream gene regulatory networks.
Structurally, the three clusters are located in distinct functional domains of the PIK3CA protein. Cluster 1 resides in the helical domain where the PIK3CA protein interfaces with the p85 regulatory subunit, overlapping with known ESR1‐associated mutations, including the well‐characterized E545K gain‐of‐function mutation. Cluster 2 is located within the kinase domain, also enriched for previously annotated pathogenic variants M1043I and H1047R. In contrast, Cluster 3 is located within the C2 domain but does not overlap with any previously characterized driver mutations, highlighting a novel region potentially involved in EZH2 regulation.
Functional analysis using CRISPR‐mediated gene dependency data [17, 18] revealed that cell lines with Cluster 1 mutations were markedly more dependent on ESR1, with dependency scores averaging 16.4 standard deviations above the mean (Figure S3). This suggests that mutations in this region of PIK3CA may co‐occur in cells that are functionally reliant on estrogen receptor signaling, making ESR1 a potential co‐dependency in this genomic context. From a vulnerability standpoint, these cells may also exhibit heightened sensitivity to PI3K inhibition, as the presence of Cluster 1 mutations could reflect broader network‐level dependencies involving both ESR1 and PIK3CA. In this way, mutations in Cluster 1 may serve as biomarkers for identifying tumors with dual vulnerabilities (to endocrine therapy and PI3K/AKT pathway inhibitors) informing more precise therapeutic strategies.
Among a panel of 545 therapeutics screened across cell lines (using Cancer Therapeutic Response Portal (CTRP) [28, 29, 30]; See Methods), we identified several compounds with significantly different responses between samples harboring PIK3CA cluster mutations and those without (p < 0.05, Wilcoxon rank‐sum test; Figure 3b). Notably, Cluster 3 mutations conferred increased sensitivity to mTORC inhibitors and related compounds, including Sirolimus:Bortezomib, KU‐0060648, and Tipifarnib. In contrast, Cluster 1 mutations were linked to enhanced sensitivity to AKT pathway inhibitors, such as MK‐2206, Myricetin, and 17‐AAG (Figure 3c,d).
Together, these results demonstrate that these PIK3CA‐based variant clusters can serve as predictive biomarkers of transcriptional state, oncogene dependency, and drug response. Importantly, these effects are not restricted to rare cases. Among 1034 breast cancer genomes, 14.5% of patients harbored Cluster 3 mutations, and 13% carried Cluster 1 mutations, underscoring the broad clinical relevance of these structural hotspots (Figure 3e). Cluster 1, in particular, contains the well‐known E545K gain‐of‐function mutation (present in 7.5% of tumors), providing a reference point for interpreting the functional impact of nearby, less‐characterized variants. Notably, Cluster 1 was enriched in ER+/luminal tumors, consistent with its association with ESR1 activity, while Cluster 2 showed a more heterogeneous distribution across receptor subtypes (Figure S8). These findings offer a structure‐informed framework for patient stratification, with the potential to guide more precise and effective therapies for a substantial subset of ER+ breast cancer patients each year [27].
3.4. Divergent ESR1 Activity Patterns in TP53 Variant Clusters
TP53 ranked second among all proteins identified by VAMOS based on the strength of association between its structural variant clusters and ESR1 transcriptional activity. To better understand these associations, we analyzed the distribution of TP53 variant clusters, their correlation with ESR1 regulatory states, and their functional and therapeutic implications.
We identified two distinct clusters of TP53 mutations with divergent ESR1 activity profiles (Figure 4a). Cluster 1 was associated with low ESR1 activity (ESR1_down) and was predominantly found in basal‐like tumors, while Cluster 2, linked to high ESR1 activity (ESR1_up), was enriched in luminal tumors. However, receptor status alone was insufficient to explain these differences; ESR1 transcriptional activity patterns offered additional resolution, indicating that TP53 mutation context shapes downstream regulatory networks in a manner not captured by ER status alone.
FIGURE 4.

(a) Two dense variant clusters in the TP53 protein are identified via our ML method. Cluster 1 (green) variants are associated with downregulation of the ESR1 pathway. Cluster 2 (yellow) variants are associated with upregulation of the ESR1 pathway. Variants with known clinical significance are highlighted in blue. The center pie charts show the number of frequent versus rare variants and variants that are known drug targets versus variants with unknown drug interactions from Onco KB present in TP53, and in each of the ML identified clusters. The likelihood that a cluster is associated with a given pathway is shown on the left. (b) A schema showing the process for finding drugs with statistically different effects on cell lines containing the ML identified clusters, shown here for Cluster 1. Cell lines are sorted into two groups: Those containing Cluster 1 variants, and those not containing Cluster 1 variants. Statistical analysis is performed to find all drugs that have lower AUC values for Cluster 1 cell lines and statistically different distributions compared to cell lines with no Cluster 1 mutations. (c) Different regions of the protein are associated with increased sensitivity towards different drugs. Cell lines with Cluster 1 variants have greater sensitivity towards drugs targeting or mediating DNA damage repair pathways. Cell lines with Cluster 2 variants have greater sensitivity towards ESSRA mediator drugs. (d) Heat maps showing the p‐values generated by the statistical analysis described in part b for Cluster 1 and Cluster 2. Additional p values from random permutation analysis are also shown. (e) These results can be used to inform precision medicine treatments for patients with mutations in the indicated clusters.
Functionally, Cluster 1 was composed largely of loss‐of‐function (LOF) TP53 mutations, consistent with its association with suppressed ESR1 signaling (Table S22). Across both tumor samples and cell lines, mutations within each cluster exhibited strong concordance with their ESR1 activity profiles, validating the phenotypic distinction between clusters (Figure S8).
To explore therapeutic implications, we analyzed drug response data from the Cancer Therapeutics Response Portal (CTRP). Cluster 1 mutations were associated with differential sensitivity to 13 compounds, including multiple DNA damage response and DNA repair inhibitors (Figure 4b). In contrast, Cluster 2 mutations were linked to a single compound, and notably, no drugs overlapped between the two clusters (Figure 4c,d), underscoring the functional divergence between these mutation groups.
Among 1034 breast cancer genomes, TP53 Cluster 1 and Cluster 2 mutations were found across both ER‐positive and ER‐negative tumors. However, their associated ESR1 activity states remained consistent, reinforcing the idea that structural mutation context within TP53 is a key determinant of transcriptional state. In the TCGA cohort, 13 patients carried mutations in these clusters, providing proof‐of‐concept for integrating TP53 variant cluster profiles into patient stratification strategies (Figure 4e). These findings offer a new framework for linking TP53 mutation context to regulatory phenotype and therapeutic response, with the potential to impact clinical decision‐making for thousands of patients annually [24].
3.5. Novel Variant Clusters Reveal New Candidates for Regulatory Activity and Therapeutic Targeting
Beyond well‐characterized drivers like PIK3CA and TP53, our analysis uncovered a set of 51 lesser‐known proteins harboring 887 rare or under‐characterized mutations that form distinct 3D structural clusters associated with ESR1 or EZH2 activity (Figure 5a; Tables S20 and 21). These variants showed clustering patterns similar to those seen in major oncogenes, with an average density of four variants per cluster, suggesting that even low‐frequency mutations may converge on functionally important regions.
Strikingly, over 96% of these mutations are absent from major oncogenic databases such as OncoKB [30] and ClinVar [31], and are not currently classified as clinically actionable. However, many of these clusters exhibit log‐odds scores comparable to known driver regions, implying potential functional relevance despite limited prior annotation. Several proteins in this novel cohort scored between 8 and 15 on our log‐odds–based priority scale (values typically associated with high‐confidence oncogenes) despite their low mutation frequency and lack of known drug targets. In contrast, proteins not associated with ESR1 or EZH2 activity consistently scored below 0.5, reinforcing the specificity of these associations (Figure 5b).
To assess therapeutic implications, we compared drug response profiles across breast cancer cell lines with and without mutations in each novel variant cluster. ESR1‐associated clusters were enriched for sensitivity to CDK4/6 inhibitors and PI3K pathway inhibitors, while EZH2‐associated clusters showed increased responsiveness to epigenetic therapies, including EZH2 inhibitors and BET inhibitors (Figure 5c and Tables S10–S19). These results suggest that even uncharacterized proteins carry mutational signatures predictive of drug response, offering new opportunities for repurposing or targeting.
Importantly, these findings extend beyond cell lines. In the TCGA breast cancer cohort, we identified 592 patients with mutations in top‐ranked “unknown‐variant” clusters associated with ESR1 or EZH2 activity. This represents more than 6% of the cohort, suggesting that structure‐informed variant clustering could significantly expand the pool of patients who may benefit from precision therapies targeting these undiscovered regulatory networks.
4. Discussion
Understanding how genomic variants impact cellular function remains a central challenge in precision medicine. While traditional approaches emphasize protein‐level effects of mutations, they often miss broader consequences on gene regulatory networks and therapeutic responses. Our study addresses this gap with VAMOS, an AI/ML framework that integrates 3D protein structure, gene regulatory activity, and drug response to prioritize both common and rare variants based on predicted functional relevance.
4.1. Structural Clustering Expands the Interpretive Landscape of Variants
Cancer genomes exhibit a long‐tail distribution of mutations [23], with the majority being rare and unannotated. Current annotation pipelines, largely reliant on mutation frequency, fail to capture the potential significance of these variants [31]. Our results support a structure‐guided alternative: by clustering mutations in 3D space, we can identify regions within proteins where even low‐frequency variants are likely to converge on shared biological functions. This structure‐informed clustering provides a mechanistic rationale for grouping functionally similar mutations, particularly when individual variants are too rare for direct association studies. VAMOS builds upon this principle by linking structural clusters to gene regulatory phenotypes, ESR1 and EZH2 activity, providing clear biomarkers for systems‐level impacts. As demonstrated across thousands of variants, this approach can distinguish which protein regions may act as control points for transcriptional regulation and, by extension, therapeutic response.
Protein three‐dimensional (3D) structure offers a biologically and functionally grounded axis for ranking variants for further investigation. Because protein folding brings distant sequence positions into close spatial proximity, clustering mutations based on their 3D locations enables functionally relevant aggregation, even when mutations are rare in the population. Mapping mutations onto protein structures reveals spatial patterns where rare mutations can co‐localize in critical functional domains, suggesting potential convergence on shared biological consequences. If variants disrupt a common domain or region important for protein function, their collective impact on cellular systems may be similar, providing a rationale for grouping them into functional clusters.
4.2. Functional Stratification Within Known Cancer Drivers
Beyond general principles, our framework delivers gene‐specific insights into the regulatory and therapeutic heterogeneity of mutations. In PIK3CA, we observed that mutations within different structural clusters were associated with distinct ESR1 or EZH2 activity patterns and drug sensitivities. These results reinforce the idea that even within a single gene, the location of a mutation, and not just its presence, may shape downstream effects. Notably, Cluster 1, which includes the E545K hotspot, was strongly associated with ESR1 upregulation and sensitivity to AKT‐pathway inhibitors, whereas Cluster 3 mutations predicted EZH2 activity and mTORC inhibitor sensitivity. These findings suggest that PIK3CA mutation clusters may guide dual‐targeting strategies, such as combining PI3K/AKT inhibitors with endocrine therapy in ESR1‐active tumors. The precision of this stratification, down to specific regions of the protein, offers a conceptual shift away from gene‐level generalizations towards structurally defined biomarker models.
TP53 provided a complementary case study. Though frequently mutated, TP53 variants are often grouped together in clinical annotation, despite their divergent effects on ESR1 activity and therapeutic response. Our results revealed that mutations in distinct TP53 clusters aligned with either ESR1 suppression or activation, which in turn predicted different drug vulnerabilities. This demonstrates that gene‐level annotation is insufficient to capture functionally distinct mutational subtypes, particularly when regulatory activity (not just protein stability or function) is the primary determinant of therapy response.
4.3. A Pathway to Interpret Rare, Unannotated Variants
Perhaps the most impactful contribution of this work lies in its application to understudied genes and rare variants. By validating the framework in known oncogenes, we build confidence in extending VAMOS to proteins with little or no prior clinical annotation. Indeed, our identification of 51 novel proteins with variant clusters associated with ESR1 or EZH2 activity (and corresponding drug sensitivity patterns) suggests a new class of high‐priority candidates for functional validation. While these variants are rare, their clustering near known functional domains, coupled with transcriptional and drug response signatures, supports their potential for clinical importance. Importantly, more than 6% of breast cancer patients in TCGA harbored mutations in these novel clusters, indicating that this strategy could meaningfully expand the population eligible for precision therapies. This is a significant step towards democratizing precision oncology, allowing rarer genotypes to be evaluated systematically rather than excluded due to sample size constraints.
4.4. A New Direction for Functional Variant Annotation
Collectively, our findings propose a new direction for variant interpretation: one rooted in structural context, regulatory impact, and therapeutic consequence. Rather than rely solely on variant frequency, we demonstrate that phenotype‐informed, structure‐aware clustering can illuminate hidden drivers of gene regulatory activity, inform mechanistic hypotheses, and guide more precise treatment decisions.
While this study focused on ESR1 and EZH2 activity in breast cancer, the framework is generalizable to other cancers, regulators, and phenotypes. Future work will focus on expanding this approach to additional transcriptional networks and integrating proteomic and epigenetic datasets to further refine predictions. Ultimately, VAMOS offers a scalable and clinically relevant tool to bridge variant impacts and regulatory network behavior and to unlock the therapeutic potential of the genomic “dark matter” that traditional annotation approaches leave behind.
5. Limitations and Future Directions
While our framework advances variant annotation through structural and functional integration, several important limitations should be acknowledged. First, our analysis was limited to coding mutations mapped to proteins with available three‐dimensional structural models. Many regions of proteins, particularly disordered or flexible domains, remain unresolved by experimental methods and cannot yet be reliably predicted by computational approaches such as AlphaFold. As a result, mutations in these structurally uncharacterized regions were excluded from clustering and analysis.
Second, although our method links variant clusters to transcriptional phenotypes, these associations are inherently correlative. The framework identifies mutations that co‐occur with specific molecular states but does not establish direct causality. Thus, while these variants may serve as effective biomarkers for pathway activity or drug sensitivity, they may not be mechanistic drivers of the observed phenotypes. This distinction is critical when interpreting the biological significance of novel variant clusters.
Several future directions can extend and strengthen this framework. Experimental validation of AI‐predicted variant clusters will be critical for establishing functional relevance. Constructs carrying specific variants from prioritized clusters can be introduced into cell lines, followed by assays measuring changes in transcriptional activity, pathway engagement, and drug sensitivity. Testing variant‐phenotype associations in clinically relevant models, including patient‐derived xenografts and organoid systems, will provide additional evidence for causal relationships between mutations and therapeutic responses.
In parallel, clinical validation will require linking variant clusters to patient outcome data. Retrospective analysis of existing clinical cohorts can assess whether mutations within prioritized clusters correlate with treatment response, resistance, or survival. Incorporating patient‐derived genomic and transcriptomic profiles into the framework will further refine prioritization and help identify clinically actionable mutations. Additional refinements to the framework, such as integrating protein–ligand and protein–protein interaction data, could enhance functional annotations by capturing how mutations disrupt key molecular interactions. Together, these directions offer a roadmap for transforming rare variant interpretation, bridging computational predictions with experimental validation, and advancing the application of structure‐informed precision medicine.
Author Contributions
E.B., P.M.S., and K.S. wrote the manuscript; E.B., P.M.S., and K.S. designed the research; K.S. and Y.W. performed the research; K.S. analyzed the data.
Conflicts of Interest
The authors declare no conflicts of interest.
Supporting information
Data S1: Supporting Information.
Acknowledgments
The authors thank the UNC Longleaf Computer Cluster and Informatics team for their support in code optimization. This work was supported by funding from the Computational Medicine Pilot Award (no reference number) and the PhRMA Foundation Drug Discovery Faculty Starter Grant (no reference number).
Shukla K., Wang Y., Spanheimer P. M., and Brunk E., “ AI‐Driven Variant Annotation for Precision Oncology in Breast Cancer,” Clinical and Translational Science 18, no. 10 (2025): e70350, 10.1111/cts.70350.
Funding: The authors received no specific funding for this work.
This paper was selected as a 2025 PhRMA Foundation Trainee Challenge Award winner.
References
- 1. Johnston S. R. D. and Dowsett M., “Aromatase Inhibitors for Breast Cancer: Lessons From the Laboratory,” Nature Reviews. Cancer 3, no. 11 (2003): 821–831. [DOI] [PubMed] [Google Scholar]
- 2. Tyson J. J., Baumann W. T., Chen C., et al., “Dynamic Modeling of Estrogen Signaling and Cell Fate in Breast Cancer Cells,” Nature Reviews. Cancer 11, no. 7 (2011): 523–532. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Kim C., Tang G., Pogue‐Geile K. L., et al., “Estrogen Receptor (ESR1) mRNA Expression and Benefit From Tamoxifen in the Treatment and Prevention of Estrogen Receptor–Positive Breast Cancer,” Journal of Clinical Oncology 29, no. 31 (2011): 4160–4167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Patel H. K. and Bihani T., “Selective Estrogen Receptor Modulators (SERMs) and Selective Estrogen Receptor Degraders (SERDs) in Cancer Treatment,” Pharmacology & Therapeutics 186 (2018): 1–24. [DOI] [PubMed] [Google Scholar]
- 5. Hernando C., Ortega‐Morillo B., Tapia M., et al., “Oral Selective Estrogen Receptor Degraders (SERDs) as a Novel Breast Cancer Therapy: Present and Future From a Clinical Perspective,” International Journal of Molecular Sciences 22, no. 15 (2021): 7812. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Lloyd M. R., Wander S. A., Hamilton E., Razavi P., and Bardia A., “Next‐Generation Selective Estrogen Receptor Degraders and Other Novel Endocrine Therapies for Management of Metastatic Hormone Receptor‐Positive Breast Cancer: Current and Emerging Role,” Therapeutic Advances in Medical Oncology 30, no. 14 (2022): 17588359221113694. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Reinert T., Saad E. D., Barrios C. H., and Bines J., “Clinical Implications of ESR1 Mutations in Hormone Receptor‐Positive Advanced Breast Cancer,” Frontiers in Oncology 7 (2017): 26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Fanning S. W., Mayne C. G., Dharmarajan V., et al., “Estrogen Receptor Alpha Somatic Mutations Y537S and D538G Confer Breast Cancer Endocrine Resistance by Stabilizing the Activating Function‐2 Binding Conformation,” eLife 5 (2016): e12792. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Cerma K., Piacentini F., Moscetti L., et al., “Targeting PI3K/AKT/mTOR Pathway in Breast Cancer: From Biology to Clinical Challenges,” Biomedicine 11, no. 1 (2023): 109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Federici G. and Soddu S., “Variants of Uncertain Significance in the Era of High‐Throughput Genome Sequencing: A Lesson From Breast and Ovary Cancers,” Journal of Experimental & Clinical Cancer Research 39 (2020): 46. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Sluiter M. D. and van Rensburg E. J., “Large Genomic Rearrangements of the BRCA1 and BRCA2 Genes: Review of the Literature and Report of a Novel BRCA1 Mutation,” Breast Cancer Research and Treatment 125, no. 2 (2011): 325–349. [DOI] [PubMed] [Google Scholar]
- 12. van den Broek A. J., Schmidt M. K., van't Veer L. J., Tollenaar R. A. E. M., and van Leeuwen F. E., “Worse Breast Cancer Prognosis of BRCA1/BRCA2 Mutation Carriers: What's the Evidence? A Systematic Review With Meta‐Analysis,” PLoS One 10, no. 3 (2015): e0120189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Olivier M., Hollstein M., and Hainaut P., “TP53 Mutations in Human Cancers: Origins, Consequences, and Clinical Use,” Cold Spring Harbor Perspectives in Biology 2, no. 1 (2010): a001008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Samuels Y. and Velculescu V. E., “Oncogenic Mutations of PIK3CA in Human Cancers,” Cell Cycle 3, no. 10 (2004): 1221–1224. [DOI] [PubMed] [Google Scholar]
- 15. Lee S., Abecasis G. R., Boehnke M., and Lin X., “Rare‐Variant Association Analysis: Study Designs and Statistical Tests,” American Journal of Human Genetics 95, no. 1 (2014): 5–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Vivekanandhan S., Bahr D., Kothari A., Ashary M. A., Baksh M., and Gabriel E., “Immunotherapies in Rare Cancers,” Molecular Cancer 22, no. 1 (2023): 23. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Tsherniak A., Vazquez F., Montgomery P. G., et al., “Defining a Cancer Dependency Map,” Cell 170, no. 3 (2017): 564–576.e16. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Ghandi M., Huang F. W., Jané‐Valbuena J., et al., “Next‐Generation Characterization of the Cancer Cell Line Encyclopedia,” Nature 569, no. 7757 (2019): 503–508. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Weinstein J. N., Collisson E. A., Mills G. B., et al., “The Cancer Genome Atlas Pan‐Cancer Analysis Project,” Nature Genetics 45, no. 10 (2013): 1113–1120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Seashore‐Ludlow B., Rees M. G., Cheah J. H., et al., “Harnessing Connectivity in a Large‐Scale Small‐Molecule Sensitivity Dataset,” Cancer Discovery 5, no. 11 (2015): 1210–1223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Qing T., Mohsen H., Rozenblit M., et al., “Abstract PS19‐05: Functional Importance of Long‐Tail Mutations in Breast Cancer,” Cancer Research 81, no. 4_Supplement (2021): PS19‐05. [Google Scholar]
- 22. Dustin D., Gu G., and Fuqua S. A. W., “ESR1 Mutations in Breast Cancer,” Cancer 125, no. 21 (2019): 3714–3728. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Kleer C. G., Cao Q., Varambally S., et al., “EZH2 Is a Marker of Aggressive Breast Cancer and Promotes Neoplastic Transformation of Breast Epithelial Cells,” Proceedings of the National Academy of Sciences 100, no. 20 (2003): 11606–11611. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Hurson A. N., Abubakar M., Hamilton A. M., et al., “Prognostic Significance of RNA‐Based TP53 Pathway Function Among Estrogen Receptor Positive and Negative Breast Cancer Cases,” NPJ Breast Cancer 8, no. 1 (2022): 1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Liberzon A., Birger C., Thorvaldsdóttir H., Ghandi M., Mesirov J. P., and Tamayo P., “The Molecular Signatures Database (MSigDB) Hallmark Gene Set Collection,” Cell Systems 1, no. 6 (2015): 417–425. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Nuytten M., Beke L., Van Eynde A., et al., “The Transcriptional Repressor NIPP1 Is an Essential Player in EZH2‐Mediated Gene Silencing,” Oncogene 27, no. 10 (2008): 1449–1460. [DOI] [PubMed] [Google Scholar]
- 27. Siegel R. L., Giaquinto A. N., and Jemal A., “Cancer Statistics, 2024,” CA: A Cancer Journal for Clinicians 74, no. 1 (2024): 12–49. [DOI] [PubMed] [Google Scholar]
- 28. Rees M. G., Seashore‐Ludlow B., Cheah J. H., et al., “Correlating Chemical Sensitivity and Basal Gene Expression Reveals Mechanism of Action,” Nature Chemical Biology 12, no. 2 (2016): 109–116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Basu A., Bodycombe N. E., Cheah J. H., et al., “An Interactive Resource to Identify Cancer Genetic and Lineage Dependencies Targeted by Small Molecules,” Cell 154, no. 5 (2013): 1151–1161. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Suehnholz S. P., Nissan M. H., Zhang H., et al., “Quantifying the Expanding Landscape of Clinical Actionability for Patients With Cancer,” Cancer Discovery 14, no. 1 (2024): 49–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Landrum M. J., Lee J. M., Riley G. R., et al., “ClinVar: Public Archive of Relationships Among Sequence Variation and Human Phenotype,” Nucleic Acids Research 42 (2014): D980–D985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Ma J., Fu Y., Tu Y., et al., “Mutation Allele Frequency Threshold Does Not Affect Prognostic Analysis Using Next‐Generation Sequencing in Oral Squamous Cell Carcinoma,” BMC Cancer 18 (2018): 758. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Subramanian A., Tamayo P., Mootha V. K., et al., “Gene Set Enrichment Analysis: A Knowledge‐Based Approach for Interpreting Genome‐Wide Expression Profiles,” Proceedings of the National Academy of Sciences 102, no. 43 (2005): 15545–15550. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Gozgit J. M., Pentecost B. T., Marconi S. A., Ricketts‐Loriaux R. S. J., Otis C. N., and Arcaro K. F., “PLD1 Is Overexpressed in an ER‐Negative MCF‐7 Cell Line Variant and a Subset of Phospho‐Akt‐Negative Breast Carcinomas,” British Journal of Cancer 97, no. 6 (2007): 809–817. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Bradley A. R., Rose A. S., Pavelka A., et al., “MMTF—An Efficient File Format for the Transmission, Visualization, and Analysis of Macromolecular Structures,” PLoS Computational Biology 13, no. 6 (2017): e1005575. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Shukla K., Idanwekhai K., Naradikian M., Ting S., Schoenberger S. P., and Brunk E., “Machine Learning of Three‐Dimensional Protein Structures to Predict the Functional Impacts of Genome Variation,” Journal of Chemical Information and Modeling 64, no. 13 (2024): 5328–5343. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data S1: Supporting Information.
