Abstract
Breast cancer is the most prevalent and lethal form of cancer being the utmost common medical concern of women. Breast cancer etiology implicates numerous cellular protein receptors such as estrogen receptors (ER), progesterone receptors (PR), and human epidermal growth factor/receptor 2 (HER2) which turn on oncogenic cascade often attributed to certain genetic variations. Breast Cancer is thus classified into ER + /-, PR + /-, HER2 ± and Triple Negative types. This study seeks to build upon our current knowledge of HER2 + and TNBC BC types to discover novel patterns for diagnosis and prognosis. The study exploits wealth of HER2 + and TNBC transcriptome (RNA Seq) data to elucidate the key hub genes, their associated networks, pathways, stage-wise expression profile, role in prognosis and survival expectancy, and regulatory transcription factors. The study also employs machine learning models including support vector machine (SVM), XGBoost, Random Forest, k nearest neighbor (kNN), Naïve Bayes and Voting Classifier to distinguish between HER2 + and TNBC transcriptomes which is a key variable for early detection and choice of therapeutic alternatives. RNA Seq datasets consisting of 49 HER2 + and 44 TNBC breast tumor samples were retrieved and pre-processed. Differentially Expressed Genes (DEGs) along with their logFC and p-values were fetched. The KEGG (Kyoto Encyclopedia of Genes and Genomes) and GO (Gene Ontology) analyses of DEGs were conducted on DAVID (the Database for Annotation, Visualization and Integrated Discovery) and interaction network was constructed through Cytoscape. Ten hub genes were obtained based on maximum clique centrality (MCC), maximum neighborhood component (MNC), degree, closeness and betweenness using cytoHubba which included ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5). These hub genes were found to be associated with cell proliferation, invasion and migration. Transcription factors and association of the expression profile of these hub genes with survival expectancy was also determined. Among the ML models, SVM stood out, exhibiting classification success between HER2 + and TNBC transcriptomes with an accuracy of 90%. The findings of this study can therefore effectively aid in tracing the initial prognosis of BC and identify biomarkers for the personalized prevention, prediction, diagnosis, and treatment of BC.
Keywords: Breast cancer (BC), Triple negative breast cancer (TNBC), Human epidermal growth factor 2 (HER2 +), Differentially expressed genes (DEGs’), Pathway analysis, Network analysis, Cytoscape, CytoHubba, Hub genes, Machine learning, Gene ontology, Expression analysis, GEPIA, UALCAN, Survival analysis, NetworkAnalyst and TIMER
Subject terms: Cancer, Breast cancer
Introduction
In the emerging world of science where many diseases are being cured day by day, cancer remains as a leading cause of death. Cancers can be both fatal (pancreatic cancer, breast cancer or lung cancer) and curable (thyroid cancer or Hodgkin lymphoma)1. Breast cancer is frequently diagnosed in females, but rarely observed in males with < 1% occurrence2. In terms of new cases, breast cancer has surpassed lung cancer in global prevalence based on statistics from 20203. The incidence rate of BC in women older than 50 years is approximately 75% while 5% in women younger than 40 years4. BC is categorized in molecular subtypes: ER + , ER- (Estrogen Receptor + /-); HER2 + , HER2-(Human Epithelial growth factor Receptor 2 + /-), PR + , PR- (Progesterone Receptor + /−) and TNBC (Triple Negative Breast Cancer) and all these receptors hinge on genetic alterations5. TNBC is an aggressive molecular subtype that is defined by the absence of expression of ER, PR and HER2 + which accounts for approximately 15% to 20% of BC5–10. Almost 25% of breast cancers are known to overexpress the epithelial growth factor receptor 2 (HER2), a tyrosine kinase that belongs to family of EGFRs (an epithelial growth factor receptor)11,12. A high rate of recurrence and mortality is associated with an overexpression of HER2. In addition, HER2 + breast cancers have an increased potential for metastasizing to the brain11,13.
The current breast cancer treatment strategies are predominantly directed towards ER + /- and PR + /, but HER2 and TNBC patients are not able to benefit from endocrine therapy or targeted therapy because of the lack of knowledge of targets, poor prognoses, high recurrence, aggressive metastasis and high mortality rates14–16. The molecular and genetic profiles of HER2 + and TNBC are distinct from other subtypes of breast cancers. The aforementioned unique characteristics associated to HER2 + and TNBC underscore the importance of identifying new targets which could provide more insight into the underlying biology of these subtypes17,18.
Thus, it is imperative to identify key genes and pathways associated with HER2 + and TNBC to uncover their pathogenesis and provide a direction for new therapeutic strategies. This study focuses on the identification of novel biomarkers their functional enrichment and network analysis associated with the pathogenesis and prognosis of BC employing hierarchical analysis of RNA Seq datasets of the HER2 + and TNBC affected patents. RNA Seq datasets containing HER2 + and TNBC tumor samples were obtained from ArrayExpress database (https://www.ebi.ac.uk/biostudies/arrayexpress) and analyzed using the Galaxy server (https://usegalaxy.eu/)19 , followed by identification of the DEGs.
David (https://david.ncifcrf.gov/) is a suit of multiple tools to discover the functional classification, biochemical pathways and conserved protein domain architectures20,21. Cytoscape is an open-source biological application designed for an easy and powerful analysis, data integration and visualization of biological networks22–25. Survival and expression analysis of hub genes were performed using Kaplan–Meier plotter (https://kmplot.com/analysis/)26, GEPIA (http://gepia.cancer-pku.cn/)27 and UALCAN (https://ualcan.path.uab.edu/)28.
Machine learning models are widely employed in genomics research for the analysis of microarray and next generation data to decipher complex biological patterns29–31. Previously, supervised and unsupervised learning techniques have been used for microarray data analysis32,33. This study seeks to integrate conventional NGS analysis methods with machine learning techniques for the analysis of RNA Seq data of HER2 + and TNBC breast cancer types. We employed supervised learning classifiers such as SVM34,35, XGBoost36, Random Forest37, Naïve Bayesian classifier38–40, kNN38,40 and voting classifier41 to assess the association of gene expression data with breast cancer with a demonstrated accuracy of 90%.
Methodology
Dataset collection and preprocessing
RNA Seq datasets with accession numbers: E-GEOD-45419, E-GEOD-52194 and E-GEOD-68086, consisting of 49 HER2 + and 44 TNBC breast tumor samples were retrieved from ArrayExpress database (https://www.ebi.ac.uk/biostudies/arrayexpress) (Table 1). The data was preprocessed at Galaxy server (https://usegalaxy.eu/) using computational pipeline. FASTQC was used to check the quality of raw sequence data which was generated from high throughput RNA sequencing (RNA Seq). The poor-quality reads were trimmed and the data was aligned to the human genome (hg38) using HISAT2 to obtain the exact matches.
Table 1.
Detail of RNA Seq datasets with accession numbers: E-GEOD-45419, E-GEOD-52194 and E-GEOD-68086, consisting of 44 TNBC and 49 HER2 + breast tumor samples.
| No | Datasets | ArrayExpress Accession No | ENA Accession No | TNBC Samples | HER2 + Samples |
|---|---|---|---|---|---|
| 1 | An Integrated Model of the Transcriptome Landscape of HER-2 positive Breast Cancer | E-GEOD-45419 | ENA-SRP042620 | 8 | 8 |
| 2 | mRNA-sequencing of breast cancer subtypes and normal tissue | E-GEOD-52194 | ENA-SRP032789 | 30 | 30 |
| 3 | RNA-seq of tumor-educated platelets enables blood-based pan-cancer, multiclass and molecular pathway cancer diagnostics | E-GEOD-68086 | ENA- SRP057500 | 6 | 11 |
Identification of DEGs
A comprehensive analysis was performed to identify significant differentially expressed genes (DEGs) in HER2 + and TNBC samples using DESeq2. The DEGs with p-value < 0.05 and log2FC < 1 or log2FC > 1 was considered statistically significant.
Functional annotation of DEGs
The Database for Annotation, Visualization and Integrated Discovery (DAVID) (https://david.ncifcrf.gov/)42 and Enrichr (https://amp.pharm.mssm.edu/Enrichr/) was used for functional annotation including biological process (BP), cellular components (CC), molecular functions (MF) and the pathways of the genes, the criteria was set at p-value < 0.0543–46.
Interaction network construction and detection of hub genes
STRING (The Search Tool for the Retrieval of Interacting Genes) (https://string-db.org/)47,48 was used to determine the interaction within the DEGs (n = 140). Based on STRING annotations, the networks were generated using Cytoscape to find genetic interactions, physical interaction, co-expression and colocalization. Network topology parameters such as MCC, MNC, degree, closeness and betweenness were calculated to explore the key hub genes in an interactome network using cytoHubba plug-in49. The top 10 genes detected by these parameters were: ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5).
Prognostic value verification of hub genes
The prognostic values of ten hub genes were assessed by Kaplan–Meier Plotter (https://kmplot.com/analysis/ ) to analyze the correlation between the hub genes and survival time of patients50 by calculating a hazard ratio (HR) and p-value < 0.05 was considered statistically significant.
GEPIA analysis
Expression level between BC and control samples (|Log2FC| Cutoff = 1, and p-value cutoff = 0.05) and correlation of hub genes with TFs was performed using GEPIA (http://gepia.cancer-pku.cn/detail.php) for the analysis of RNA-seq data51. To indicate statistical significance, we assessed the predictive value of all key hub genes throughout the TCGA dataset using the GEPIA default parameters.
Expression analysis of hub genes
An in-depth analysis of hub gene expression from breast cancer was conducted using UALCAN (https://ualcan.path.uab.edu/index.html)28 with a cutoff p-value < 0.05.
Transcription factor target gene and expression correlation analyses
TF of hub genes were identified using NetworkAnalyst (http://www.networkanalyst.ca)52, the Genotype-Tissue Expression Project (GTEx) database and breast tissue were selected to filter out genes by tissue specificity. The transcription-regulated network was visualized using the Cytoscape and expression correlation analysis based on TCGA samples were conducted in GEPIA (http://gepia.cancer-pku.cn/index.html)53.
BC immune infiltration cells
TIMER (https://cistrome.shinyapps.io/timer/)54 was used to estimate the infiltration of the six types of immune cells in BC.
Model prediction and performance estimation
Six different models including SVM, XGBoost, Random Forest, kNN, Naïve Bayes and Voting Classifier were used for classification. The best model was evaluated on the basis of performance indicators including accuracy, sensitivity, specificity, Positive Predictive Value (PPV) and Negative Predictive Value (NPV)55,56. The performance indicators were computed as follows:
where, TP = true positive, TN = true negative, FN = false negative, FP = false positive.
Results
Identification of DEGs
Differentially expressed genes analysis was performed using DESeq2 on RNA Seq datasets (E-GEOD-45419, E-GEOD-52194 and E-GEOD-68086). DESeq2 generated principal component analysis plot (PCA), histogram, MA plot and dispersion graph for each dataset (Fig. 1, 2 and 3). A Venn diagram was generated based on |log2(FC)|≥ 1 and corrected p-value < 0.05 resulting in 2396 DEGs using Venny tool (https://csbg.cnb.csic.es/BioinfoGP/venny.html) shown in Fig. 4. Further analysis of these DEGs revealed that in E-GEOD-45419 were 1586 upregulated genes and 1090 downregulated genes, E-GEOD-45419 have 12,163 upregulated genes and 2045 downregulated genes whereas in E-GEOD-68086 1408 were upregulated genes and 26 downregulated genes consistently observed and their volcano plots and boxplots were presented in Fig. 5.
Fig. 1.
The PC plot (A), Dispersion estimates (B), histogram (C) and MA plot (D) were obtained from DESeq2 for E-GEOD-45419. A PC plot shows the estimation of variance-mean dependence of two groups: HER2 + and TN. The blue color signifies TN samples whereas red color represents HER2 + samples. B Plot display dispersion estimation of the mean of the normalized counts. C Histogram is a graphical representation of samples frequency and p value by displaying the DEG’s which are grouped into bins on the basis of genes frequency. D MA plot shows the samples of HER2 + and TN by transforming the data by using log ratio and mean average. The red color shows the dispersion of differentially expressed genes while grey color shows no variation. Significant genes were observed in the range 2 to -2. The upregulated genes were observed in range 2 to 4 while down regulated genes lie in range between − 2 to -6 represented by red dots.
Fig. 2.
The PC plot (A), Dispersion estimates (B), histogram (C) and MA plot (D) were obtained from DESeq2 for E-GEOD-52194. A PC plot shows the estimation of variance-mean dependence of two groups: HER2 + and TN. The blue color signifies TN samples whereas red color represents HER2 + samples. B Plot display dispersion estimation of the mean of the normalized counts. C Histogram is a graphical representation of samples frequency and p value by displaying the DEG’s which are grouped into bins on the basis of genes frequency. D MA plot shows the samples of HER2 + and TN by transforming the data by using log ratio and mean average. The red color shows the dispersion of differentially expressed genes while grey color shows no variation. Significant genes were observed in the range 2 to -2. The upregulated genes were observed in range 2 to 6 while down regulated genes lie in range between − 2 to -4 represented by red dots.
Fig. 3.
The PC plot (A), Dispersion estimates (B), histogram (C) and MA plot (D) were obtained from DESeq2 for E-GEOD-68086. A PC plot shows the estimation of variance-mean dependence of two groups: HER2 + and TN. The blue color signifies TN samples whereas red color represents HER2 + samples. B Plot display dispersion estimation of the mean of the normalized counts. C Histogram is a graphical representation of samples frequency and p value by displaying the DEG’s which are grouped into bins on the basis of genes frequency. D MA plot shows the samples of HER2 + and TN by transforming the data by using log ratio and mean average. The red color shows the dispersion of differentially expressed genes while grey color shows no variation. No significant genes were observed in the range 2 to -2. The upregulated genes were observed in range + 4 to + 6 represented by red dots.
Fig. 4.

Venn diagram of 2396 overlapping DEGs from the E-GEOD-45419, E-GEOD-52194 and E-GEOD-68086 datasets. DEGs, differentially expressed genes.
Fig. 5.
Volcano plot and boxplots of DEGs for each dataset were generated. The volcano plot A of E-GEOD-45419, C of E-GEOD-52194 and E of E-GEOD-68086 highlights the differential expressed genes between two conditions in which genes having significant up-regulation are shown with red color while the down-regulated genes are shown with blue color that are separated on the basis of fold-change thresholds. The non-significant ones are represented by grey color which are differentiated statistically on the basis of adjusted p-value threshold. The boxplot B of E-GEOD-45419, D of E-GEOD-52194 and F of E-GEOD-68086 illustrates the distribution of log2 fold changes (log2(FC)) for gene categories: up-regulated, down-regulated, and not significant. The up-regulated genes (Red) show primarily positive log2(FC) values with a narrow range showing highly upregulated outliers while Down-regulated genes (Blue) have mainly negative log2(FC) values with a slighter spread and some outliers showing extreme down-regulation. The Not significant genes (grey) show log2(FC) values are centered around zero with minimal change in expression and a wide range of variability. This boxplot highlights distribution and changeability of expression patterns across significant and non-significant gene categories.
Pathway enrichment analysis
Enrichr and David was used to identify the KEGG pathways in which differentially expressed genes are involved and the results showed that the DEGs were enriched in ubiquitin mediated proteolysis, protein processing in endoplasmic reticulum, ATP-dependent chromatin remodeling, p53 signaling pathway, proteoglycans in cancer, cycle, endocytosis, cellular senescence, adherens junction and FoxO signaling pathway (Table 2, Fig. 6 and Fig. 7).
Table 2.
KEGG pathways of differentially expressed genes obtained by using Enrichr.
| Term | P-value | Gene count |
|---|---|---|
| hsa04120: Ubiquitin mediated proteolysis | 1.28E-06 | 30 |
| hsa04141: Protein processing in endoplasmic reticulum | 4.91E-05 | 30 |
| hsa03082: ATP-dependent chromatin remodeling | 0.001738318 | 20 |
| hsa04115: p53 signaling pathway | 0.001798866 | 15 |
| hsa05205: Proteoglycans in cancer | 0.002411603 | 29 |
| hsa04110: Cell cycle | 0.002690893 | 24 |
| hsa04144: Endocytosis | 0.004263587 | 33 |
| hsa04218: Cellular senescence | 0.005255122 | 23 |
| hsa04520: Adherens junction | 0.005448527 | 16 |
| hsa04068: FoxO signaling pathway | 0.006854397 | 20 |
| hsa05212: Pancreatic cancer | 0.036456418 | 12 |
| hsa04512: ECM-receptor interaction | 0.04388441 | 13 |
| hsa04010: MAPK signaling pathway | 0.04499527 | 33 |
Fig. 6.
The pathway enrichment category plots classify pathways in six functional groups which highlights the distribution of gene numbers associated with class/ category. The highest gene enrichment is showcased in metabolism. The other categories like genetic and environmental information processing shows active involvement in genetic maintenance and cellular communication while categories including human diseases and organismal systems pin points genes associated with disease mechanisms and functions. This distribution highlights the biological processes most impacted.
Fig. 7.
KEGG pathways of DEGs. The legend at the right side in which dot’s color indicates the p-value and its size indicates the number of enriched target genes in the pathway.
Gene ontology enrichment analysis of DEGs
To investigate the functional roles of DEGs, GO terms including biological processes (BP), molecular functions (MF) and cellular components (CC) were analyzed at David (p-value < 0.05; FDR < 0.05). The results depict that the DEGs were involved in biological processes including DNA damage response, ubiquitin-dependent protein catabolic process, protein localization, chromatin remodeling, protein polyubiquitination (Fig. 8 and Table 3). The molecular functions of DEGs were concentrated in protein binding, RNA binding, ubiquitin-protein transferase activity, ATP binding, small GTPase binding (Fig. 8, Table 3). The cellular components in which DEGs were enriched in Nucleoplasm, cytosol, nucleus, nuclear speck, cytoplasm (Fig. 8, Table 3).
Fig. 8.
This bar plot illustrates the results of Gene ontology enrichment analysis including three categories biological processes, cellular components and molecular functions, each bar shows genes associated with GO terms. The biological functions are shown with red bars, cellular components with blue and molecular functions with green bars. This analysis reveals the cellular roles and functional diversity of analyzed genes providing insights in the biological significance.
Table 3.
Biological Processes, Cellular Components and Molecular Functions of DEGs.
| Category | Term | Description | Count |
|---|---|---|---|
| BP | GO:0,006,974 | DNA damage response | 51 |
| BP | GO:0,006,511 | ubiquitin-dependent protein catabolic process | 42 |
| BP | GO:0,008,104 | protein localization | 31 |
| BP | GO:0,006,338 | chromatin remodeling | 51 |
| BP | GO:0,000,209 | protein polyubiquitination | 32 |
| CC | GO:0,005,654 | Nucleoplasm | 475 |
| CC | GO:0,005,829 | Cytosol | 594 |
| CC | GO:0,005,634 | Nucleus | 616 |
| CC | GO:0,016,607 | nuclear speck | 72 |
| CC | GO:0,005,737 | Cytoplasm | 528 |
| MF | GO:0,005,515 | protein binding | 1225 |
| MF | GO:0,003,723 | RNA binding | 196 |
| MF | GO:0,004,842 | ubiquitin-protein transferase activity | 51 |
| MF | GO:0,005,524 | ATP binding | 172 |
| MF | GO:0,031,267 | small GTPase binding | 46 |
Network analysis
To gain a knowledgeable insight network of 2396 DEGs was constructed using STRING database and Cytoscape (Fig. 9 and 10). Moreover, MCC, MNC, degree, closeness and betweenness was used to identify hub genes using cytoHubba with default parameters in Cytoscape (Fig. 11). The top ten hub genes obtained are ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5) (Fig. 12).
Fig. 9.
The network was constructed using STRING database and visualized by Cytoscape of DEGs. The extensive network includes nodes and edges. The edges represent different interactions between genes on the basis of co-expression, pathways, shared protein domains, physical interactions predicted, co-localization and genetic interactions which are represented by different colors: Physical interactions- > light pink, Co-expression- > light purple, Genetic interactions- > green, Pathways- > light green and shared protein domains- > pink.
Fig. 10.
Subnetwork of DEGs that includes nodes and edges. The network also includes the ten hub genes, ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5) identified in this study. The edges represent different interactions between genes on the basis of co-expression, pathways, shared protein domains, physical interactions predicted, co-localization and genetic interactions which are represented by different colors: Physical interactions- > light pink, Co-expression- > light purple, Genetic interactions- > green, Pathways- > light green and shared protein domains- > pink.
Fig. 11.

A Venn diagram of 10 overlapping hub genes obtained from cytoHubba using MCC, MNC, degree, closeness and betweenness algorithms.
Fig. 12.

Hub genes network visualization.
Survival analysis of hub genes
The Kaplan–Meier survival analysis highlights the prognostic significance of gene expression levels in overall survival outcomes of ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5) (Fig. 13). High expression of ESR1 (HR = 0.72, P = 0.0079), KRAS (HR = 0.73, P = 0.0052), ATM (HR = 0.78, P = 0.031), MDM2 (HR = 0.7, P = 0.002), and TP53 (HR = 0.57, P = 5.2e-07) is associated with significantly better survival, suggesting these genes as favorable prognostic markers. Conversely, high expression of GAPDH (HR = 1.3, P = 0.028) and SIRT1 (HR = 1.8, P = 4e-07) correlates with worse survival outcomes, indicating their potential as unfavorable prognostic markers. In contrast, genes such as ACTB (HR = 1.2, P = 0.11) and HNRNPK (HR = 1.15, P = 0.23) show no statistically significant associations.
Fig. 13.
Prognostic value of nine ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53 in BC patients.
Gene expression profiling of key hub genes
GEPIA was employed to analyze the ten hub genes, their overall expression level as compared to normal tissues. The box plot (Fig. 14) of the ten hub genes illustrates that the genes were significantly overexpressed in breast cancer as compared to normal breast tissue. The genes namely ACTB, ATM, ESR1, HNRNPK, KRAS, MDM2 and TP53 were found up regulated while SIRT1 and H3F3C (H3-5) were found to be non-significantly expressed in breast cancer.
Fig. 14.
Comparisons of the expression of the ten hub genes between breast cancer and normal breast tissues in TCGA and GTEx based on GEPIA. The Y axis represents the log2 (TPM + 1) for gene expression. The Grey boxplot indicates the normal tissues while the red boxplot shows the BC tissues. TPM: transcripts per kilobase million.
Expression level of hub genes with respect to BC tumor stages, menopause status and BC-subclasses
The gene expression analysis from UALCAN highlights elevated expression of identified hub genes, including ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5) in breast cancer (BRCA) compared to normal tissue. These hub genes play crucial roles in tumor progression and their expression patterns vary across different tumor stages, breast cancer subtypes and Menopause status were shown in Fig. 15, 16 and 17.
Fig. 15.
The expression profile of Hub genes at different stages of breast cancer.
Fig. 16.
Expression of hub genes based on breast cancer subclasses.
Fig. 17.
Expression of hub genes based on Menopause status.
Transcription factor
The transcription-regulated network was predicted by the NetworkAnalyst that can regulate the expression of hub genes. The transcription network included 102 nodes and 119 edges. Results showed that TP53 had 51 connections, ESR1 had 29, MDM2 had 23 connections followed by SIRT1, ATM, KRAS and GAPDH with 11, 8, 1 and 3 connections respectively (Fig. 18). GEPIA was employed for correlation expression analysis between transcription factor (MYC, MYCN, DNMT1, RELA, CEBPA, CEBPD and HIF1A) and hub genes (ATM, KRAS, SIRT1, TP53, ESR1 and GAPDH) as shown in Fig. 19.
Fig. 18.
Network of transcription factors and hub genes. The red nodes represent hub genes, the diamond shaped blue nodes represent transcription factors that are associated with hub genes.
Fig. 19.
Association of transcription factors and hub genes. The scatter plots showing correlations between transcription factors (x-axis, given as log2 TPM) and hub genes (y-axis, likely also log2 TPM). Each subplot has a correlation coefficient (R) and p-value displayed, which indicate the strength and significance of the correlation.
BC immune infiltration cells
The abundance of six types of infiltrating immune cells, including B Cells, CD8+ T Cells, CD4+ T Cells, macrophages, neutrophils, and dendritic cells, was estimated to predict the influence of immune cell infiltration on the progression and clinical outcomes of BC (Fig. 20). The results revealed that the expression of candidate hub genes was significantly correlated with the infiltration levels of CD8+ T cells, B cells, neutrophils, dendritic cells, macrophages, and CD4+ T cells (P < 0.05).
Fig. 20.
The correlations between the copy number (deep deletion, arm-level deletion, diploid/normal, arm-level gain, and high amplification) of hub genes expression and tumor immune cell infiltration levels in BC tissues. CD8 + , CD8 + T cells, CD4 + , CD4 + T cells. p-value Significant Codes: 0 ≤ *** < 0.001 ≤ ** < 0.01 ≤ * < 0.05 ≤ . < 0.1.
Classification models
In this study, six machine-learning algorithms including Support Vector Machine, XGBoost, Random Forest, k-nearest neighbor, Naïve Bayes and voting classifier were employed for classification between RNA Seq datasets HER2 + and TNBC affected patients. The comparison among the classifiers shows that the best accuracy of 90% was obtained from SVM classifier (Table 4).
Table 4.
Performance of classification models.
| Classifier | Sensitivity | Specificity | Accuracy | Precision | Accuracy Percentage |
|---|---|---|---|---|---|
| SVM | 1.00 | 80 | 0.90 | 0.83 | 90% |
| XGBoost | 0.60 | 1.00 | 0.78 | 0.67 | 77.78% |
| Random Forest | 0.80 | 1.00 | 0.89 | 0.80 | 88.89% |
| Naive Byes | 0.600 | 1.00 | 0.7778 | 0.7143 | 78% |
| kNN | 0.40 | 1.00 | 0.67 | 1.00 | 66.67% |
| Voting Classifier | 0.60 | 1.00 | 0.78 | 0.57 | 77.78% |
Discussion
The Breast cancer (BC) is one of the most prevalent and life-threatening malignancies affecting women worldwide. Currently it is the major cause of cancer related deaths in women, due to which the death rate has augmented drastically. Molecular studies of breast cancer have revealed that it is a heterogeneous disease which has several receptor types including: ER + , ER-, HER2 + , HER2-, PR + and PR-. All these receptors hinge on genetic variations. The most aggressive BC subtypes are HER2 + and TNBC, which have a high recurrence rate and a worse prognosis. The cancers caused by overexpression of HER2 + are responsible for twenty percent of breast cancer cases57. Triple Negative is the tumor caused when there is no variation in expression of ER, PR and HER2. In comparison to HER2, Triple negative causes about fifteen percent of the breast cancers, being more amongst younger women58,59. The objective of the present study was to investigate the molecular underpinnings and the aggressive nature of BC by identifying significant hub genes. RNA sequencing has revolutionized the study of transcriptome with the help of high-throughput sequencing technologies which can be exploited to figure out the genetic alterations which lead to disease. Zhai et al. (2019) identified differentially expressed genes (DEGs) to discriminate between transcriptome samples from TNBC and non-TNBC individuals through analyses on multiple datasets from GEO and TCGA databases; the study successfully highlighted key genes associated with overall survival (OS) and relapse-free survival (RFS)60. The protein–protein interaction (PPI) networks enabled the identification of ten hub genes with prognostic significance, contributing to risk stratification and potential therapeutic targeting in TNBC namely EGFR, KRT16, RET, SOX10, PDZK1, XBP1, TFF3, PTGER3, NME5, and IL6ST. Similarly, Chen et al. (2020) focused on KEGG-expressed genes and pathways in TNBC, identifying 2998 DEGs and highlighting pathways such as “Pathways in cancer” and “Systemic lupus erythematosus”61. Their study emphasized the role of top hub genes, including EGFR, in TNBC prognosis.
Further expanding on TNBC research, Wang et al. (2019) integrated Weighted Gene Co-expression Network Analysis (WGCNA) with DEG analysis to explore prognostic biomarkers in TNBC62. Using data from GSE76275, they identified SIDT1, ANKRD30A, GPR160, and CA12 as key genes linked to RFS, with SIDT1 demonstrating a suppressive role in TNBC cell proliferation and tumor growth. This study underscored SIDT1’s potential as a prognostic biomarker and therapeutic target in TNBC.
While these studies provide valuable insights into TNBC-related DEGs and key prognostic biomarkers, our study offers a broader and more integrative approach by focusing on both TNBC and HER2 + breast cancer subtypes. We utilized transcriptomic samples combined with advanced machine learning models, including Support Vector Machine (SVM), XGBoost, Random Forest, k-Nearest Neighbors (kNN), Naïve Bayes, and a Voting Classifier, to enhance subtype classification and predictive accuracy. Our bioinformatics pipeline incorporated KEGG and GO enrichment analyses of DEGs and employed five algorithms—MCC, MNC, Degree, Betweenness, and Closeness—for robust hub gene identification.
Additionally, we applied tools such as Kaplan–Meier plotter, GEPIA, UALCAN, NetworkAnalyst, and network analysis for comprehensive validation, survival assessment, and interaction mapping. Unlike previous studies focused solely on TNBC, our research extends the analysis to HER2 + subtypes, integrating machine learning and multi-dimensional bioinformatics approaches to uncover novel biomarkers and potential therapeutic targets across both subtypes, thus contributing to more personalized treatment strategies.
RNA Seq datasets containing HER2 + and TNBC tumor samples were obtained from ArrayExpress database. The data was preprocessed, aligned and differentially expressed genes were obtained. ML algorithms including SVM, XGBoost, Random Forest, Naïve Bayesian classifier, kNN and voting classifier were employed to classify between the HER2 + and TNBC transcriptomes. These approaches were evaluated by means of several measures such as accuracy, sensitivity, specificity, positive predictive value and negative predictive value.
David functional annotation tool and Enrichr were used to obtain significant GO terms and pathways associated to DEGs. The genes were enriched in pathways such as ubiquitin mediated proteolysis, protein processing in endoplasmic reticulum, ATP-dependent chromatin remodeling, p53 signaling pathway, proteoglycans in cancer, cycle, endocytosis, cellular senescence, adherens junction, FoxO signaling pathway and MAPK signaling pathways.
The network analysis of genes was done at STRING and Cytoscape. Ten hub genes (ACTB, ATM, ESR1, GAPDH, HNRNPK, KRAS, MDM2, SIRT1, TP53, and H3F3C (H3-5) based on five diverse and highly cited algorithms MCC, MNC, Degree, Betweenness and Closeness were identified which can serve as prognostic biomarkers. Kaplan–Meier method was used to delineated the expression profiles of these hub genes in various substages of BC. Additionally, analyses at UALCAN and GEPIA of aforementioned hub genes also exhibited variable expression levels associated with each of individual substages of BC. Therefore, these hub genes have been found to exhibit significant differential expression during various substages of BC vis a vis healthy individual, profiling the variability of expression along the disease progression. Transcription factors and hub genes regulatory network was also constructed to explore the molecular mechanism of BC progression and associated TFs were identified which regulate the expression of these hub genes.
The cytoskeletal protein ACTB plays a crucial role in cell migration, division, embryonic development, wound healing, immunity, and gene expression63,64. Recent evidence reveals that this gene plays a significant role in various diseases, particularly cancer65,66. Tumor growth and metastasis are governed by the cytoskeleton and actin microfilaments67. In addition, ATM is a regulatory gene which plays a critical role in the cell cycle and DNA repair68. ATM gene mutations result in Ataxia Telangiectasia (A-T), which is a rare but serious autosomal recessive disorder that causes ionizing radiation sensitivity, cerebellar neurodegeneration, immunodeficiency, and an increased risk of cancer69,70. In the presence of an activated ATM, several downstream targets can be phosphorylated, such as p53, chek2 and BRCA1, which will stop the cell cycle, repair DNA or trigger apoptosis, leading to an increase in cancer incidence due to insufficient cellular repair71–73.
The human ESR1 gene is a steroid hormone receptor gene. The expression and transcription of proteins are significantly affected by several SNPs in ESR1, resulting in an increased risk of breast cancer74,75. Moreover, an enzyme called glyceraldehyde-3-phosphate dehydrogenase (GAPDH) is responsible for the reversible conversion of glyceraldehyde-3-phosphate (G-3-P) into 1,3-diphosphoglycerate76,77. Apart from its glycolytic function, GAPDH is also involved in several cellular functions, such as nuclear tRNA export, DNA replication and repair, endocytosis, exocytosis, cytoskeletal organization, iron metabolism, and carcinogenesis78.
Heterogeneous nuclear ribonucleoprotein K (HNRNPK) is a key member of the HNRNP family that is widely expressed in human cells in the nucleus, cytoplasm, mitochondria, and cell membrane79,80. They form ribonucleoprotein complexes that are involved in gene expression, signal transduction, DNA repair, and telomere synthesis81. HNRNPK regulates gene expression regulation, cell signaling pathways, and cell signal82. Additionally, KRAS is an oncogene in the rat sarcoma viral family (RAS), which also includes Harvey and neuroblastoma rat sarcoma viral oncogenes (HRAS, NRAS)83,84. It encodes two highly related protein isoforms, KRAS-4B and KRAS-4A, which contain 188 and 189 amino acids, respectively. KRAS is generally known as KRAS-4B because cells have high levels of mRNA encoding KRAS-4B85,86.
MDM2 (Murine double minute 2) produces oncoproteins that play a crucial role in that regulatory process. MDM2 binds to p53 and prevents its tumor suppressor activities and promotes its degradation. These two proteins thus form an autoregulatory feedback loop in which p53 positively regulates MDM2 levels and MDM2 negatively regulates p53 levels and activity87.Moreover, SIRT1, a member of the sirtuin family, has gained significant attention for its diverse roles in physiological and pathological processes, including lifespan regulation, neurodegeneration, age-related disorders, obesity, heart disease, inflammation, and cancer88–91. Notably, SIRT1 has been implicated in promoting tumor development and progression.
TP53, located on chromosome 17p13, spans 20 kilobases and consists of 11 exons92. It encodes a primarily nuclear phosphoprotein of 53 kDa and belongs to a highly conserved gene family, including P63 and TP7393,94. Unlike most tumor suppressors, over 75% of TP53 mutations are missense, leading to single amino acid substitutions95. Lastly, H3F3C (H3-5) encodes the histone H3 variant H3.5, which appears to play a crucial role in the complex dynamics of various cell types96. However, its regulation and specific functions remain largely unexplored. The key hub genes identified are associated with tumorigenesis and progression of breast cancer, they may serve as biomarkers for diagnosis and therapeutic targets. Further studies in pre-clinical and prospective models are required to validate these findings.
Conclusion
In summary, our comprehensive analysis of HER2 + and TNBC transcriptomes identified key hub genes and ML classifiers which can help identifying the genes with unique expression profiles which can serve as prognostic biomarkers for diagnosis. These genes can provide knowledgeable insight in comprehending the pathogenesis, progression and prognosis of breast cancer. SVM was found to serve as best classifier between samples of HER2 + and TNBC affected patients.
Acknowledgements
Not Applicable
Abbreviations
- BC
Breast Cancer
- TNBC
Triple Negative Breast Cancer
- HER2 +
Human Epidermal Growth Factor 2
- ER
Estrogen Receptor
- PR
Progesterone Receptor
- EGFR
Epithelial Growth Factor Receptor
- ML
Machine Learning
- DEGs
Differentially Expressed Genes
- GO
Gene Ontology
- SVM
Support Vector Machine
- kNN
K-nearest Neighbor
- DAVID
Database for Annotation, Visualization and Integrated Discovery
- BP
Biological Processes
- MF
Molecular Functions
- CC
Cellular Components
- STRING
Search Tool for the Retrieval of Interacting Genes
- MCC
Maximum Clique Centrality
- MNC
Maximum Network Clique
- HR
Hazard Ratio
- PCA
Principal Component Analysis
- GEPIA
Gene Expression ProfilingIinteractive Analysis
- TIMER
Tumor Immune Estimation Resource
- KEGG
Kyoto Encyclopedia of Genes and Genomes
Author contributions
MZ and AS contributed equally in the study. MS and LZ performed formal analysis. FS, FSA and SAR assisted in writeup manuscript submission and formal analysis. IA, ZM and WB.A assisted in writing—review & editing. ARS conceptualized and supervised the study. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported by the Researchers Supporting Project Number (RSP2025R332), King Saud University, Riyadh, Saudi Arabia.
Data availability
Yes. The data is publicly available in both repositories i.e. ArrayExpress and ENA. The accession numbers for both repositories are as below: Accession Numbers of datasets in ArrayExpress are listed below: 1.E-GEOD-45419 2.E-GEOD-52194 3.E-GEOD-68086 Accession Numbers of datasets in ENA are listed below: 1.ENA-SRP042620 2.ENA-SRP032789 3.ENA- SRP057500.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Contributor Information
Wadi B. Alonazi, Email: waalonazi@ksu.edu.sa
Abdul Rauf Siddiqi, Email: araufsiddiqi@comsats.edu.pk.
References
- 1.Mattiuzzi, C. & Lippi, G. Current cancer epidemiology. J. Epidemiol. Glob. Health9(4), 217 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Heber, D. Nutritional oncology (Elsevier, 2011). [Google Scholar]
- 3.Sung, H. et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J. Clin.71(3), 209–249 (2021). [DOI] [PubMed] [Google Scholar]
- 4.DeSantis, C., Ma, J., Bryan, L. & Jemal, A. Breast cancer statistics, 2013. Cancer J. Clin.64(1), 52–62 (2014). [DOI] [PubMed] [Google Scholar]
- 5.Lee, S. K. et al. Distinguishing low-risk luminal a breast cancer subtypes with Ki-67 and p53 is more predictive of long-term survival. PLoS ONE10(8), e0124658 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Naorem, L. D., Muthaiyan, M. & Venkatesan, A. Integrated network analysis and machine learning approach for the identification of key genes of triple-negative breast cancer. J. Cell. Biochem.120(4), 6154–6167 (2019). [DOI] [PubMed] [Google Scholar]
- 7.Criscitiello, C., Azim, H. Jr., Schouten, P., Linn, S. & Sotiriou, C. Understanding the biology of triple-negative breast cancer. Ann. Oncol.23, 13–18 (2012). [DOI] [PubMed] [Google Scholar]
- 8.Lehmann, B. D. & Pietenpol, J. A. Identification and use of biomarkers in treatment strategies for triple-negative breast cancer subtypes. J. Pathol.232(2), 142–150 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bae, S. Y. et al. Poor prognosis of single hormone receptor-positive breast cancer: similar outcome as triple-negative breast cancer. BMC Cancer15, 1–9 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xu, J. et al. Delving into the heterogeneity of different breast cancer subtypes and the prognostic models utilizing scRNA-Seq and bulk RNA-Seq. Int. J. Mol. Sci.23(17), 9936 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Patel, A., Unni, N. & Peng, Y. The changing paradigm for the treatment of HER2-positive breast cancer. Cancers12(8), 2081 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Dean-Colomb, W. & Esteva, F. J. Her2-positive breast cancer: herceptin and beyond. Eur. J. Cancer44(18), 2806–2812 (2008). [DOI] [PubMed] [Google Scholar]
- 13.Mounsey, L. A. et al. Changing natural history of HER2–Positive breast cancer metastatic to the brain in the era of new targeted therapies. Clin. Breast Cancer18(1), 29–37 (2018). [DOI] [PubMed] [Google Scholar]
- 14.Mao, J. et al. Single hormone receptor-positive metaplastic breast cancer: similar outcome as triple-negative subtype. Front. Endocrinol.12, 628939 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Van Asten, K. et al. Prognostic value of the progesterone receptor by subtype in patients with estrogen receptor-positive, HER-2 negative breast cancer. Oncologist24(2), 165–171 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Higgins, M. J. & Baselga, J. Targeted therapies for breast cancer. J. Clin. Investig.121(10), 3797–3803 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Gorshein, E., Klein, P., Boolbol, S. K. & Shao, T. Clinical significance of HER2-positive and triple-negative status in small (≤ 1 cm) node-negative breast cancer. Clin. Breast Cancer14(5), 309–314 (2014). [DOI] [PubMed] [Google Scholar]
- 18.Wei, Z. et al. Analysis of differentially expressed proteins between HER2 positive and triple negative breast cancer and their prognostic significance. Ann. Diagn. Pathol.55, 151834 (2021). [DOI] [PubMed] [Google Scholar]
- 19.Jalili, V. et al. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2020 update. Nucleic Acids Res.48(W1), W395–W402 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Dennis, G. et al. DAVID: database for annotation, visualization, and integrated discovery. Genome Biol.4(9), 1–11 (2003). [PubMed] [Google Scholar]
- 21.Huang, D. W., Sherman, B. T. & Lempicki, R. A. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nat. Protoc.4(1), 44–57 (2009). [DOI] [PubMed] [Google Scholar]
- 22.Cline, M. S. et al. Integration of biological networks and gene expression data using Cytoscape. Nat. Protoc.2(10), 2366–2382 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Shannon, P. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res.13(11), 2498–2504 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Otasek, D., Morris, J. H., Bouças, J., Pico, A. R. & Demchak, B. Cytoscape automation: empowering workflow-based network analysis. Genome Biol.20(1), 1–15 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.M. D. Mares-Quiñones, E. Galán-Vásquez, E. Pérez-Rueda, D. G. Pérez-Ishiwara, M. O. Medel-Flores, and M. d. C. Gómez-García, "Identification of modules and key genes associated with breast cancer subtypes through network analysis," Scientific Reports, vol. 14, no. 1, p. 12350, 2024. [DOI] [PMC free article] [PubMed]
- 26.B. Győrffy, "Integrated analysis of public datasets for the discovery and validation of survival-associated genes in solid tumors," The Innovation, vol. 5, no. 3, 2024. [DOI] [PMC free article] [PubMed]
- 27.Tang, Z. et al. GEPIA: a web server for cancer and normal gene expression profiling and interactive analyses. Nucleic Acids Res.45(W1), W98–W102 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Chandrashekar, D. S. et al. UALCAN: An update to the integrated cancer data analysis platform. Neoplasia25, 18–27 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.C. Liu, D. Che, X. Liu, and Y. Song, "Applications of machine learning in genomics and systems biology," ed: Hindawi, 2013. [DOI] [PMC free article] [PubMed]
- 30.Neelima, E. & Babu, M. A comparative study of machine learning classifiers over gene expressions towards cardio vascular diseases prediction. International Journal of Computational Intelligence Research13(3), 403–424 (2017). [Google Scholar]
- 31.Tarca, A. L., Carey, V. J., Chen, X.-W., Romero, R. & Drăghici, S. Machine learning and its applications to biology. PLoS computational biology3(6), e116 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Libbrecht, M. W. & Noble, W. S. Machine learning applications in genetics and genomics. Nat. Rev. Genet.16(6), 321–332 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Vandesompele, J. et al. Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes. Gen. Biol.3(7), 1–12 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.V. Jakkula, "Tutorial on support vector machine (svm)," School of EECS, Washington State University, vol. 37, no. 2.5, p. 3, 2006.
- 35.Zararsız, G. et al. A comprehensive simulation study on classification of RNA-Seq data. PloS one12(8), e0182507 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Liu, P. et al. Optimizing survival analysis of XGBoost for ties to predict disease progression of breast cancer. IEEE Trans. Biomed. Eng.68(1), 148–160 (2020). [DOI] [PubMed] [Google Scholar]
- 37.B. Dai, R.-C. Chen, S.-Z. Zhu, and W.-W. Zhang, "Using random forest algorithm for breast cancer diagnosis," in 2018 International symposium on computer, consumer and control (IS3C), 2018: IEEE, pp. 449–452.
- 38.Tomar, D. & Agarwal, S. A survey on Data Mining approaches for Healthcare. Int. J. Bio-Sci. Bio-Technol.5(5), 241–266 (2013). [Google Scholar]
- 39.M. J. Islam, Q. J. Wu, M. Ahmadi, and M. A. Sid-Ahmed, "Investigating the performance of naive-bayes classifiers and k-nearest neighbor classifiers," in 2007 international conference on convergence information technology (ICCIT 2007), 2007: IEEE, pp. 1541–1546.
- 40.A. Jabeen, N. Ahmad, and K. Raza, "Machine learning-based state-of-the-art methods for the classification of rna-seq data," Classification in BioApps: Automation of Decision Making, pp. 133–172, 2018.
- 41.U. K. Kumar, M. S. Nikhil, and K. Sumangali, "Prediction of breast cancer using voting classifier technique," in 2017 IEEE international conference on smart technologies and management for computing, communication, controls, energy and materials (ICSTM), 2017: IEEE, pp. 108–114.
- 42.Huang, D. W. et al. DAVID Bioinformatics Resources: expanded annotation database and novel algorithms to better extract biology from large gene lists. Nucleic Acids Research35(suppl_2), W169–W175. 10.1093/nar/gkm415 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Kuleshov, M. V. et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic acids research44(W1), W90–W97 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Xie, Z. et al. Gene set knowledge discovery with Enrichr. Current Protocols1(3), e90 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Chen, E. Y. et al. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinf.14(1), 1–14 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Kanehisa, M., Furumichi, M., Sato, Y., Matsuura, Y. & Ishiguro-Watanabe, M. KEGG: biological systems database as a model of the real world. Nucl. Acids Res.53(D1), D672–D677 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.D. Szklarczyk et al., "The STRING database in 2017: quality-controlled protein–protein association networks, made broadly accessible," Nucleic acids research, p. gkw937, 2016. [DOI] [PMC free article] [PubMed]
- 48.Szklarczyk, D. et al. STRING v10: protein–protein interaction networks, integrated over the tree of life. Nucl. Acids Res.43(D1), D447–D452 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Chin, C.-H. et al. cytoHubba: identifying hub objects and sub-networks from complex interactome. BMC Syst. Biol.8(4), 1–7 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Chen, J. et al. KEGG-expressed genes and pathways in triple negative breast cancer: Protocol for a systematic review and data mining. Medicine99(18), e19986 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Salavaty, A., Rezvani, Z. & Najafi, A. Survival analysis and functional annotation of long non-coding RNAs in lung adenocarcinoma. J. Cell. Mol. Med.23(8), 5600–5617 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Zhou, G. et al. NetworkAnalyst 3.0: a visual analytics platform for comprehensive gene expression profiling and meta-analysis. Nucl. Acids Res.47(W1), W234–W241 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Tang, Z., Kang, B., Li, C., Chen, T. & Zhang, Z. GEPIA2: an enhanced web server for large-scale expression profiling and interactive analysis. Nucl. Acids Res.47(W1), W556–W560 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Yang, X. et al. Potential value of PRKDC as a therapeutic target and prognostic biomarker in pan-cancer. Medicine101(27), e29628 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Wu, J. & Hicks, C. Breast cancer type classification using machine learning. J. Personal. Med.11(2), 61 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.D. M. Powers, "Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation, arXiv preprint arXiv:2010.16061 2020.
- 57.Zubair, M., Wang, S. & Ali, N. Advanced approaches to breast cancer classification and diagnosis. Front. Pharmacol.11, 632079 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Foulkes, W. D., Smith, I. E. & Reis-Filho, J. S. Triple-negative breast cancer. New Engl. J. Med.363(20), 1938–1948 (2010). [DOI] [PubMed] [Google Scholar]
- 59.Kumar, P. & Aggarwal, R. An overview of triple-negative breast cancer. Arch. of Gynecol. Obstetr.293, 247–269 (2016). [DOI] [PubMed] [Google Scholar]
- 60.Zhai, Q., Li, H., Sun, L., Yuan, Y. & Wang, X. Identification of differentially expressed genes between triple and non-triple-negative breast cancer using bioinformatics analysis. Breast Cancer26, 784–791 (2019). [DOI] [PubMed] [Google Scholar]
- 61.Chen, J. et al. KEGG-expressed genes and pathways in triple negative breast cancer: Protocol for a systematic review and data mining. Medicine99(18), e19986 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Y. Wang et al., "Integrated bioinformatics data analysis reveals prognostic significance of SIDT1 in triple-negative breast cancer," OncoTargets and therapy, pp. 8401–8410, 2019. [DOI] [PMC free article] [PubMed]
- 63.Guo, C., Liu, S., Wang, J., Sun, M.-Z. & Greenaway, F. T. ACTB in cancer. Clin. Chimica Acta417, 39–44 (2013). [DOI] [PubMed] [Google Scholar]
- 64.Bunnell, T. M., Burbach, B. J., Shimizu, Y. & Ervasti, J. M. β-Actin specifically controls cell growth, migration, and the G-actin pool. Mol. Biol. Cell22(21), 4047–4058 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Yang, H., Zhang, L. & Liu, S. Determination of reference genes for ovine pulmonary adenocarcinoma infected lung tissues using RNA-seq transcriptome profiling. J. Virolog. Methods284, 113923 (2020). [DOI] [PubMed] [Google Scholar]
- 66.Herath, S. et al. Selection and validation of reference genes for normalisation of gene expression in ischaemic and toxicological studies in kidney disease. PLoS One15(5), e0233109 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Sadek, K. H., Cagampang, F. R., Bruce, K. D., Macklon, N. & Cheong, Y. Variation in stability of housekeeping genes in healthy and adhesion-related mesothelium. Fertil. Steril.98(4), 1023–1027 (2012). [DOI] [PubMed] [Google Scholar]
- 68.Bogdanova, N. et al. A nonsense mutation (E1978X) in the ATM gene is associated with breast cancer. Breast Cancer Res. Treat.118, 207–211 (2009). [DOI] [PubMed] [Google Scholar]
- 69.Swift, M., Reitnauer, P. J., Morrell, D. & Chase, C. L. Breast and other cancers in families with ataxia-telangiectasia. New Engl. J. Med.316(21), 1289–1294 (1987). [DOI] [PubMed] [Google Scholar]
- 70.Chen, J., Birkholtz, G. G., Lindblom, P., Rubio, C. & Lindblom, A. The role of ataxia-telangiectasia heterozygotes in familial breast cancer. Cancer Res.58(7), 1376–1379 (1998). [PubMed] [Google Scholar]
- 71.Bernstein, J. et al. Population-based estimates of breast cancer risks associated with ATM gene variants c. 7271T> G and c. 1066–6T> G (IVS10–6T> G) from the Breast Cancer Family Registry. Human Mutation27(11), 1122–1128 (2006). [DOI] [PubMed] [Google Scholar]
- 72.Abraham, R. T. PI 3-kinase related kinases:‘big’players in stress-induced signaling pathways. DNA Repair3(8–9), 883–887 (2004). [DOI] [PubMed] [Google Scholar]
- 73.Dombernowsky, S. L. et al. Risk of cancer by ATM missense mutations in the general population. J. Clin. Oncol.26(18), 3057–3062 (2008). [DOI] [PubMed] [Google Scholar]
- 74.Li, T. et al. A meta-analysis of the association between ESR1 genetic variants and the risk of breast cancer. PloS one11(4), e0153314 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Lipphardt, M. F., Deryal, M., Ong, M. F., Schmidt, W. & Mahlknecht, U. ESR1 single nucleotide polymorphisms predict breast cancer susceptibility in the central European Caucasian population. Int. J. Clin. Exp. Med.6(4), 282 (2013). [PMC free article] [PubMed] [Google Scholar]
- 76.Zhang, J.-Y. et al. Critical protein GAPDH and its regulatory mechanisms in cancer cells. Cancer Biol. Med.12(1), 10 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Colell Riera, A., Green, D. R. & Ricci, J.-E. Novel roles for GAPDH in cell death and carcinogenesis. Cell Death Differ16(12), 1573-1581A (2009). [DOI] [PubMed] [Google Scholar]
- 78.Sheokand, N. et al. Moonlighting cell-surface GAPDH recruits apotransferrin to effect iron egress from mammalian cells. J. Cell Sci.127(19), 4279–4291 (2014). [DOI] [PubMed] [Google Scholar]
- 79.Bomsztyk, K., Denisenko, O. & Ostrowski, J. hnRNP K: one protein multiple processes. Bioessays26(6), 629–638 (2004). [DOI] [PubMed] [Google Scholar]
- 80.Wang, Z. et al. The emerging roles of hnRNPK. J. Cel. Physiol.235(3), 1995–2008 (2020). [DOI] [PubMed] [Google Scholar]
- 81.Piñol-Roma, S. & Dreyfuss, G. hnRNP proteins: localization and transport between the nucleus and the cytoplasm. Trends Cell Biol.3(5), 151–155 (1993). [DOI] [PubMed] [Google Scholar]
- 82.Geuens, T., Bouhy, D. & Timmerman, V. The hnRNP family: insights into their role in health and disease. Human Gen.135, 851–867 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Huang, L., Guo, Z., Wang, F. & Fu, L. KRAS mutation: from undruggable to druggable in cancer. Signal Transd. Target. Ther.6(1), 386 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Barbacid, M. Ras genes. Ann. Rev. Biochem.56(1), 779–827 (1987). [DOI] [PubMed] [Google Scholar]
- 85.Plowman, S. et al. K-ras 4A and 4B are co-expressed widely in human tissues, and their ratio is altered in sporadic colorectal cancer. J. Exp. Clin. Cancer Res. CR25(2), 259–267 (2006). [PubMed] [Google Scholar]
- 86.Karnoub, A. E. & Weinberg, R. A. Ras oncogenes: split personalities. Nat. Rev. Mol. Cell Biol.9(7), 517–531 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Freedman, D., Wu, L. & Levine, A. Functions of the MDM2 oncoprotein. Cell. Mol. Life Sci. CMLS55, 96–107 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Zhang, J. et al. The type III histone deacetylase Sirt1 is essential for maintenance of T cell tolerance in mice. J. Clin. Invest.119(10), 3048–3058 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Lin, Z. et al. USP22 antagonizes p53 transcriptional activation by deubiquitinating Sirt1 to suppress cell apoptosis and is required for mouse embryonic development. Mol. Cell46(4), 484–494 (2012). [DOI] [PubMed] [Google Scholar]
- 90.Wu, Y. et al. Resveratrol-activated AMPK/SIRT1/autophagy in cellular models of Parkinson’s disease. Neurosignals19(3), 163–174 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Chen, W. Y. et al. Tumor suppressor HIC1 directly regulates SIRT1 to modulate p53-dependent DNA-damage responses. Cell123(3), 437–448 (2005). [DOI] [PubMed] [Google Scholar]
- 92.Bénard, J., Douc-Rasy, S. & Ahomadegbe, J. C. TP53 family members and human cancers. Human Mutation21(3), 182–191 (2003). [DOI] [PubMed] [Google Scholar]
- 93.Levrero, M. et al. The p53/p63/p73 family of transcription factors: overlapping and distinct functions. J. Cell Sci.113(10), 1661–1670 (2000). [DOI] [PubMed] [Google Scholar]
- 94.Guimaraes, D. & Hainaut, P. TP53: a key gene in human cancer. Biochimie84(1), 83–93 (2002). [DOI] [PubMed] [Google Scholar]
- 95.Hainaut, P. & Hollstein, M. p53 and human cancer: the first ten thousand mutations. Adv. Cancer Res.77, 81–137 (1999). [DOI] [PubMed] [Google Scholar]
- 96.Schenk, R., Jenke, A., Zilbauer, M., Wirth, S. & Postberg, J. H3. 5 is a novel hominid-specific histone H3 variant that is specifically expressed in the seminiferous tubules of human testes. Chromosoma120, 275–285 (2011). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Yes. The data is publicly available in both repositories i.e. ArrayExpress and ENA. The accession numbers for both repositories are as below: Accession Numbers of datasets in ArrayExpress are listed below: 1.E-GEOD-45419 2.E-GEOD-52194 3.E-GEOD-68086 Accession Numbers of datasets in ENA are listed below: 1.ENA-SRP042620 2.ENA-SRP032789 3.ENA- SRP057500.

















