Skip to main content
Cancer Biomarkers: Section A of Disease Markers logoLink to Cancer Biomarkers: Section A of Disease Markers
. 2026 May 28;43:18758592261454886. doi: 10.1177/18758592261454886

Stage-specific biomarkers in triple-negative breast cancer: Preliminary findings from bioinformatics approaches

Faeze Shahraki 1, Azadeh Meshkini 1,, Elham Nazari 2
PMCID: PMC13219939  PMID: 42206911

Abstract

Background

Integrated bioinformatics approaches were used to identify stage-specific candidate genes and potential drug targets in triple-negative breast cancer (TNBC).

Methods

Microarray (164 early-stage, 33 advanced-stage, and 53 normal samples) and RNA-seq (113 normal, 163 early-stage, and 30 advanced-stage TNBC samples) datasets were analyzed. Differentially expressed genes (DEGs) were identified, followed by co-expression analysis using Weighted Gene Co-expression Network Analysis (WGCNA) and protein–protein interaction analysis using the STRING database. miRNA co-regulation was evaluated using multiMiR and TCGA correlation analyses. Candidate genes were validated using UALCAN and immunohistochemistry data. Molecular docking assessed potential therapeutic agents.

Results

Novel stage-specific candidate biomarkers were identified, including DNAJC6, SKP2, MOCOS, and NCAPD2 in early-stage TNBC, and F11R, FOXO6, PPP4C, and TMEM51 in advanced-stage TNBC. UALCAN analysis confirmed the dysregulation of these genes across 23 additional malignancies. STRING-based network analysis revealed stage-specific protein–protein interactions, including SKP2–SKP1 in early-stage and F11R–TJP1 in advanced-stage TNBC. miRNA co-regulation distinguished early-stage TNBC through PI3K–AKT–related pathways and advanced-stage TNBC through tumor progression–associated pathways. Docking-based drug repurposing highlighted conventional agents (e.g., doxorubicin) and potential novel candidates (e.g., sunitinib).

Conclusion

This study identifies novel stage-specific gene candidates and suggests repurposable drugs for TNBC, supporting progression-specific targeted therapeutic strategies.

Keywords: Triple-Negative breast cancer (TNBC), cancer staging, microarray gene expression profiles, drug repurposing, bioinformatics

Introduction

Triple-negative breast cancer (TNBC) is an aggressive, heterogeneous breast cancer subtype. Its lack of targeted therapies poses significant diagnostic and therapeutic challenges. 1 Unlike other subtypes, TNBC lacks the three key receptors—Estrogen Receptor (ER), Progesterone Receptor (PR), and Human Epidermal Growth Factor Receptor-2 (HER2)—that are commonly targeted in breast cancer treatments. 2 Consequently, unlike other breast cancer subtypes, TNBC is insensitive to hormonal and HER2-targeted therapies. This leaves patients with limited treatment options, primarily involving chemotherapy, which is often associated with severe side effects and variable efficacy.3,4 Clinically, TNBC exhibits a more aggressive behavior, a higher propensity for metastasis, and a poorer prognosis compared to other breast cancer subtypes.5,6 The five-year survival rate for TNBC patients is approximately 77%, which is significantly lower than that observed in hormone receptor–positive breast cancers.7,8 Therefore, identifying reliable prognostic factors is critical for improving personalized treatment strategies, survival outcomes, and patient quality of life. 9

Determining the stage of TNBC is critical for guiding treatment decisions and improving patient outcomes. 10 Cancer staging provides essential information regarding tumor size, the extent of lymph node involvement, and the presence of distant metastases. Early-stage TNBC (stages I and II) is generally associated with smaller tumors and limited lymph node involvement. Patients at these stages often benefit from surgery, followed by chemotherapy and radiation therapy.11,12 In contrast, advanced-stage TNBC (stages III and IV) is characterized by larger tumors, extensive lymph node involvement, and distant metastases. These cases typically require more aggressive treatment strategies, including combination chemotherapy, targeted therapies, and immunotherapy.3,13 Accurate staging enables clinicians to tailor treatment strategies, maximize therapeutic efficacy, and minimize unnecessary toxicity.

Despite advances in understanding TNBC, a critical need remains to identify biomarkers that inform stage-specific therapeutic strategies. Biomarkers are essential for predicting prognosis, guiding treatment decisions, and monitoring disease progression. Recent studies, summarized in Table SI1, have revealed several potential biomarkers and explored molecular markers associated with TNBC prognosis, including CD3D, CCNB1, and MKI67.14,15 Other promising markers, such as TMEM100, have been identified, with evidence suggesting that its regulation affects breast cancer progression and chemosensitivity; however, its precise prognostic value and mechanistic role remain unclear and require further investigation. 16 However, many of these studies have not addressed the stage-specific relevance of these biomarkers. For example, a recent study using weighted gene co-expression network analysis (WGCNA) identified prognostic transcription factors in TNBC but focused on generic biomarker discovery rather than a stage-stratified approach. 17 Collectively, these limitations highlight a significant gap in understanding stage-specific TNBC progression and therapeutic opportunities.

The molecular mechanisms driving stage-specific progression remain poorly characterized. This knowledge gap is compounded by the historical treatment of TNBC as a monolithic disease and a lack of stratified clinical data. As a result, stage-specific therapeutic strategies remain underexplored.

To address this gap, we leveraged advanced transcriptomic analysis by integrating data from the Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA). Our goal was to identify distinct molecular signatures, uncover stage-specific biomarkers, gain insights into progression mechanisms, and identify therapeutic targets for each TNBC stage. Therefore, this study provides a foundational resource for personalized TNBC medicine. By pinpointing candidate targets, we aimed to address the unique challenges of this aggressive subtype and support the development of precise, stage-specific interventions. To achieve this goal, we conducted WGCNA integrated with differential gene expression analysis. This approach aligns with the evolution of network-based methods for biomarker discovery, such as FUNMarker, which integrates multiple gene relationships to identify robust prognostic markers. 18 Central to these methods is the analysis of gene co-expression networks. WGCNA is a powerful biological method that analyzes gene expression data to identify clusters (modules) of highly correlated genes and associates them with phenotypic traits. 17

In this study, WGCNA was used to pinpoint novel stage-specific candidate biomarkers for each TNBC stage. Pan-cancer dysregulation patterns of these genes were then assessed using the University of Alabama at Birmingham Cancer Data Analysis portal (UALCAN), and the literature was curated to evaluate their associations with tumor aggressiveness. Immunohistochemistry (IHC) data from the Human Protein Atlas (HPA) were analyzed to validate TNBC-specific expression profiles. The physical and genetic interactions of these genes were examined using the STRING database. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses, were also performed. Stage-specific regulatory networks, such as miRNA–mRNA interactions, were systematically characterized with integrated bioinformatics tools to clarify their roles in early- and advanced-stage TNBC. Using these findings, we explored repurposable and conventional drug molecules for each stage via the Comparative Toxicogenomics Database (CTD). Finally, molecular docking was conducted to evaluate interactions between candidate drugs and proteins encoded by critical genes. This approach provides a comprehensive view of stage-specific therapeutic opportunities in TNBC. The workflow of the study is illustrated in Figure 1.

Figure 1.

Figure 1.

Workflow of this study. (QC: quality control, KGs: key genes).

Materials and methods

Data collection

Gene expression data were sourced from public repositories to leverage both platform homogeneity and independent validation. Three datasets were obtained from the GEO database, all derived from the same platform (GPL570) Affymetrix Human Genome U133 Plus 2.0 Array, to facilitate effective data integration. Using datasets from the same platform ensures greater data homogeneity, reduces batch effects, and minimizes technical variability, thereby improving the reliability of the analysis.19,20

A total of 197 TNBC samples were analyzed, comprising 164 early-stage and 33 advanced-stage samples, along with 53 normal samples. These were sourced from three omics datasets with accession numbers GSE76275, 21 GSE29044, 22 and GSE42568 23 (Table 1). Additionally, RNA sequencing (RNA-Seq) data from TCGA were utilized, encompassing 193 TNBC samples with defined staging status and 113 normal tissue samples. Among the TNBC samples, 163 were categorized as early-stage and 30 as advanced-stage.

Table 1.

List of datasets with brief descriptions that are used in this study.

Accession number Sample size Probe Platform Country
TNBC Normal
GSE76275 197 _ 54675 GPL570AffymetrixHumanGenomeU133Plus2.0Array USA
GSE29044 _ 36 54675 GPL570AffymetrixHumanGenomeU133Plus2.0Array Saudi Arabia
GSE42568 _ 17 54675 GPL570AffymetrixHumanGenomeU133Plus2.0Array Ireland
TCGA data 193 113 60660 NA NA

Staging was determined based on the TNM classification system developed by the American Joint Committee on Cancer (AJCC), which considers tumor size (T), lymph node involvement (N), and metastasis (M). 24 Detailed criteria for staging TNBC patients are outlined in Table SI2, specifying the parameters used for classification. This structured approach enables a thorough comparison of gene expression profiles between early-stage and advanced-stage TNBC, enhancing our understanding of the molecular characteristics associated with different disease stages.

Data preprocessing and identification of differentially expressed genes (DEGs)

Differential expression analysis utilized platform-specific methods: DESeq2 with median-of-ratios normalization and negative binomial modeling for RNA-seq count data, and limma with Robust Multi-Array Average (RMA) normalization, linear modeling, and empirical Bayes moderation for microarray continuous data. Raw data from three datasets available in the GEO database were downloaded and combined for analysis (Table 1). RMA normalization, which includes background correction, normalization, and median polish summarization, was applied using the Affy package in R software.25,26 Normalizing the datasets collectively ensured consistent calculation of parameters like mean and standard deviation for z-scores, effectively minimizing batch effects and reducing variability. Joint normalization (instead of processing them independently) also enhanced comparability, improving the reliability of interpretations and the overall robustness of the results. Following data merging and normalization, a Principal Component Analysis (PCA) plot was generated as a quality control measure to identify low-quality samples or outliers following the simultaneous normalization of the three datasets. The analysis revealed three outlier samples (GSM1045194_N5_14_12_04.CEL, GSM1045192_N4_14_1204.CEL, and GSM1045193_N5_15_12_04.CEL), which were subsequently removed to enhance data quality for further analyses. The PCA plot also confirmed two distinct clusters corresponding to the biological conditions of TNBC and normal samples, with no evidence of a batch effect based on the original dataset source (Fig. SI1).

Probe IDs were converted to gene symbols using the biomaRt package (v-2.58.2) in R, and the datasets were merged. Genes with zero expression and duplicates were removed to clean the data. DEGs were identified between early-stage TNBC and normal samples, as well as advanced-stage TNBC and normal samples, using the Linear Models for Microarray Analysis (limma) statistical method implemented in the limma package (v-3.58.1) in R. 27 The analysis followed the guidelines from the Linear Models for Microarray and RNA-Seq Data User's Guide by Gordon K. Smyth (last revised April 22, 2023) and focused on single-channel experimental designs and Affymetrix arrays. 28 The threshold for identifying up- and down-regulated DEGs in the combined data has been considered as follows;

DEG={DEG(upregulated)ifadj.pvalue0.05andLog2FC+1.0DEG(downregulated)ifadj.pvalue<0.05andLog2FC<1.0

Where adj. p-value is the adjusted p-value, and Log2 FC is Log fold change.

RNA-seq data from TCGA were analyzed for differential expression using the DESeq2 package under the same conditions and thresholds applied to the microarray data.

An online tool (http://bioinformatics.psb.ugent.be/webtools/Venn/) was used to create Venn diagrams for the DEGs obtained from the GEO and TCGA databases. DEG analysis was performed separately between early-stage and normal samples, as well as advanced-stage and normal samples. Common genes between the two DEG sets were removed to identify genes specifically involved in early stages and those uniquely associated with advanced stages.

Construction of co-expression network and identification of key modules

WGCNA was specifically selected for its ability to infer shared regulatory relationships, such as microRNA co-regulation, from correlation patterns in expression data. This capability is fundamental to the goal of identifying coregulated biomarker candidates.

Data normalization and transformation

WGCNA was employed to identify gene modules strongly associated with the external traits of the samples. 29 To ensure the reliability of WGCNA, normalization and transformation were applied to DEGs involved in early stages and those uniquely associated with advanced stages, which were identified from microarray and RNA-seq datasets, respectively. Normalization scaled the data to a comparable range, and transformation stabilized variance, making the data more suitable for correlation analyses such as WGCNA.2932

For the microarray data, Robust Multi-array Average (RMA) normalization was applied, followed by log2 transformation function to stabilize variance and prepare the data for downstream analysis. For the RNA-seq data, Transcripts Per Million (TPM) normalization, downloaded from TCGA, was utilized, followed by the Variance Stabilizing Transformation (VST) function. Genes with low expression levels and low variance were filtered out to include only informative genes in the analysis. These preprocessing steps were essential for building a WGCNA network free from variance and bias.

Validation of stage-specific key genes in TNBC

To move beyond in silico mRNA validation and obtain direct protein-level evidence, IHC data from HPA (https://www.proteinatlas.org) were analyzed to validate the expression of stage-specific key genes in TNBC (early-stage: I–II; advanced-stage: III–IV) versus normal breast tissue. 33

Pan-Cancer evaluation of key genes

Six TNBC stage-specific genes (three early-stage [I–II] and three advanced-stage [III–IV]) were evaluated across 23 cancer types through differential mRNA expression analysis using UALCAN (https://ualcan.path.uab.edu/) (thresholds: |log2FC| > 1, adjusted p < 0.05), supplemented with supporting evidence from NCBI Gene. This systematic validation preserved the stage-specific context of each gene in TNBC while assessing their relevance across multiple cancer types.

Interaction analysis: physical and genetic perspectives

Data on protein–protein interactions (PPIs) and genetic interactions were obtained from the STRING database (https://string-db.org). Genetic interactions, which indicate how genes influence each other's expression, were used to map gene relationships, analyze pathways, and investigate complex traits. Physical interactions, representing direct protein associations, were used to elucidate protein mechanisms, metabolic pathways, and potential drug targets.

KEGG and go functional enrichment analysis

KEGG and GO functional enrichment analyses were performed using the STRING database to determine the role of critical genes involved in the early and advanced stages of TNBC in biological pathways.

Co-Regulation by miRNAs of key TNBC-associated genes in early and advanced stages

To investigate the regulatory roles of miRNAs in TNBC progression, a systematic analysis was conducted using the multiMiR package in R, which integrates multiple miRNA-target interaction databases, including TarBase, miRTarBase, and TargetScan. miRNAs co-regulating three key genes in the early stage and three in the advanced stage of TNBC were identified. miRNA–gene interactions (experimental and predicted) were first extracted from multiMiR. miRNAs simultaneously regulating key genes in each stage were then identified, and overlapping miRNAs between early and advanced stages were analyzed to determine dual regulatory roles. Stage-specific miRNAs were classified into early-stage-specific (regulating only early-stage genes) and advanced-stage-specific (regulating only advanced-stage genes) groups. Pairwise Pearson correlation analysis was performed to assess potential co-expression relationships between the 44 early-stage and 20 advanced-stage miRNAs, using miRNA expression data from TCGA processed with the stringr package in R.

Drug repurposing

The proteins encoded by key genes at each stage of TNBC were considered target proteins. Drug agents associated with these key genes were identified using the CTD (http://ctdbase.org/). 34 Drugs selected for further analysis were either FDA-approved or in clinical trials, as determined from DrugBank (https://www.drugbank.com), and included both conventional and repurposed agents with the potential to downregulate the expression of the target proteins.

Molecular docking calculations

Molecular docking studies of repurposed drugs (3D structures obtained from PubChem) with proteins expressed from key genes in each stage of TNBC were conducted using AutoDock Vina. The 3D structures of all proteins were obtained from AlphaFold (https://alphafold.ebi.ac.uk). The structures of repurposed drugs were optimized for geometry, charges, and atomic energy using VEGA ZZ software. Polar hydrogens were added to the proteins, and Gasteiger charges were calculated to prepare them for docking. A grid with a spacing of 0.375 Å was constructed around a specific domain region of the proteins. The flexible drug structures were then docked to the rigid proteins. The docking scores from AutoDock Vina were used to rank the conformations based on their binding free energy (ΔG, kcal/mol) and the number of hydrogen bonds. The best pose was visualized using PyMOL software.

Results

Identification of DEGs

Both microarray and RNA-seq data were analyzed across different stages of TNBC. For the microarray data, a total of 425 DEGs were identified in 164 early-stage TNBC samples compared to 50 normal samples. In the advanced-stage analysis, 567 DEGs were identified between 33 advanced-stage TNBC samples and 50 normal samples. To distinguish genes specifically involved in early and advanced stages, 1248 common genes between the two DEG sets were removed (Figure 2A).

Figure 2.

Figure 2.

Venn diagrams illustrating the distribution of DEGs between the early and advanced stages of TNBC, constructed using two distinct datasets. For the GEO dataset (A), 425 DEGs were identified in early-stage TNBC, while 567 DEGs were found in advanced-stage TNBC. Analysis of the TCGA dataset (B) revealed 2601 DEGs in early-stage TNBC and 1662 DEGs in advanced-stage TNBC.

Similarly, the RNA-seq analysis identified 2601 DEGs in 163 early-stage TNBC samples compared to normal samples. For advanced-stage TNBC, 1662 DEGs were identified between 30 advanced-stage TNBC samples and 113 normal samples. Additionally, 4806 common genes between the two DEG sets were removed to isolate stage-specific genes (Figure 2B).

Weighted co-expression network construction and gene modules identification

WGCNA was used to identify key modules significantly associated with both early and advanced stages of TNBC across microarray and RNA-seq datasets (Table 2). To pinpoint the most biologically relevant gene sets, for each developmental stage, the modules exhibiting the highest correlation and strongest statistical significance with the stage were selected as the top correlated modules for further analysis. These key modules were analyzed to uncover distinct gene co-expression patterns associated with the developmental stages of TNBC, offering insights into the underlying molecular mechanisms and potential biomarkers for early and advanced stages of the disease.

Table 2.

Key modules were identified that exhibited significant correlations with distinct stages of TNBC progression.

Stage Dataset Module Correlation with TNBC stage
Early Stage RNA-seq Brown Module with the highest correlation with the stage trait
Advanced Stage RNA-seq Blue Module with the highest correlation with the stage trait
Early Stage Microarray Blue Module with the highest correlation with the stage trait
Advanced Stage Microarray Brown Module with the highest correlation with the stage trait

To ensure accurate adjacency calculations, an appropriate soft-thresholding power (β) was determined. The lowest power at which the scale-free topology fit index curve flattened out at a high value (R2 > 0.85) was selected, resulting in an unsigned network structure. A dendrogram, presented in Figure 3, was constructed to visually represent the clustering of genes based on their topological overlap. This analysis categorized genes into distinct modules, each assigned a specific color for ease of identification. The dendrogram, constructed using a topological overlap matrix (TOM), quantified connectivity among genes by measuring shared neighbors, thereby highlighting strong correlations while minimizing noise from weaker associations.

Figure 3.

Figure 3.

Module-trait association heatmap for early and advanced cancer stages with common gene identification. This figure illustrates module-trait associations in early and advanced cancer stages using heatmaps derived from two distinct data types. Panel A represents RNA-seq data, with the upper section depicting early-stage cancer and the lower section showing advanced-stage cancer. Panel B displays microarray data, similarly organized with early-stage cancer in the upper section and advanced-stage cancer in the lower section. The heatmaps’ color intensity indicates the strength of correlations between gene modules and clinical traits, with accompanying p-values denoting statistical significance. These heatmaps display both correlation coefficients and p-values, enabling a detailed assessment of the relationships. Notably, four common genes, represented by circular nodes, were identified within key modules at each stage across both datasets, potentially signifying robust biomarkers or therapeutic targets consistent across data types and cancer progression stages.

Additionally, Figure 3 includes a module-trait association heatmap that displays the correlation coefficients and p-values for various traits associated with early and advanced stages of cancer in the TCGA and microarray datasets. This heatmap facilitates a quick visual assessment of the relationships between gene modules and clinical traits, emphasizing significant correlations that may indicate potential biomarkers or therapeutic targets.

Within these key modules, four common genes were identified at each stage across the two datasets. Module membership (MM) and gene significance (GS) values for these common genes are provided in Table SI3. The genes identified for the early stage included DNAJC6, SKP2, MOCOS, and NCAPD2, while F11R, PPP4C, FOXO6, and TMEM51 were identified for the advanced stage. These genes were subsequently analyzed further to investigate their roles and potential as biomarkers in TNBC progression.

Stage-Specific key gene dysregulation in TNBC: A validation study

IHC staining data retrieved from the HPA confirmed significant dysregulation in the expression of key stage-specific genes in TNBC (Figures 4 and 5). Compared to normal breast tissue, early-stage (I-II) tumors exhibited upregulation of SKP2, MOCOS, and NCAPD2 (Figure 4), while advanced-stage (III-IV) tumors showed upregulation of F11R, TMEM51, and PPP4C protein expression (Figure 5). These findings align with transcriptomic data and support the functional relevance of these genes in TNBC progression.

Figure 4.

Figure 4.

Validation of early-stage TNBC key gene (SKP2, MOCOS, NCAPD2) expression through IHC from the HPA database, showing protein-level dysregulation compared to normal breast tissue.

Figure 5.

Figure 5.

Validation of advanced-stage TNBC key gene (F11R, TMEM51, PPP4C) expression through IHC from the HPA, demonstrating protein-level alterations versus normal tissue.

Pan-Cancer validation

UALCAN Pan-Cancer analysis confirmed widespread dysregulation of these key genes across malignancies. Differential mRNA expression analysis (|log2FC| > 1, adj. p < 0.05) across 23 cancer types revealed consistent upregulation of both early-stage genes (SKP2, MOCOS, NCAPD2) and advanced-stage genes (F11R, TMEM51, PPP4C) (Figure 6). The overexpression of these genes in diverse cancers underscores broader oncogenic relevance.

Figure 6.

Figure 6.

Differential expression (|log2FC| > 1, adj. p < 0.05) across 23 cancers confirmed consistent upregulation of early-stage (SKP2, MOCOS, NCAPD2) and advanced-stage (F11R, TMEM51, PPP4C) genes, supporting their pan-cancer relevance. Cancer abbreviations: BLCA: Bladder Urothelial Carcinoma; BRCA: Breast Invasive Carcinoma; CESC: Cervical Squamous Cell Carcinoma; CHOL: Cholangiocarcinoma; COAD: Colon Adenocarcinoma; ESCA: Esophageal Carcinoma; GBM: Glioblastoma Multiforme; HNSC: Head and Neck Squamous Cell Carcinoma; KICH: Kidney Chromophobe; KIRC: Kidney Renal Clear Cell Carcinoma; KIRP: Kidney Renal Papillary Cell Carcinoma; LIHC: Liver Hepatocellular Carcinoma; LUAD: Lung Adenocarcinoma; LUSC: Lung Squamous Cell Carcinoma; PAAD: Pancreatic Adenocarcinoma; PRAD: Prostate Adenocarcinoma; READ: Rectum Adenocarcinoma; SARC: Sarcoma; SKCM: Skin Cutaneous Melanoma; STAD: Stomach Adenocarcinoma; THCA: Thyroid Carcinoma; THYM: Thymoma; UCEC: Uterine Corpus Endometrial Carcinoma.

Literature validation further clarified the roles of these genes in cancer progression. DNAJC6 drives hepatocellular carcinoma via epithelial-to-mesenchymal transition (EMT), 35 while SKP2 overexpression in glioblastomas and melanoma correlates with poor survival.36,37 NCAPD2 is associated with TNBC aggressiveness and lymph node metastasis,38,39 and F11R/JAM-A promotes tumor aggression, with elevated expression in glioblastomas and links to cancer stemness.4044 FOXO6 accelerates progression in colorectal and gastric cancers,45,46 and PPP4C enhances lung cancer proliferation and poor outcomes. 47 Supporting their potential clinical relevance, MOCOS is a prognostic marker in renal, pancreatic, and lung cancers, and TMEM51 in renal cancer, according to the HPA. Collectively, these findings suggest that all identified genes are pivotal pan-cancer biomarkers with potential as therapeutic targets.

Physical and genetic interaction analysis with molecular mechanisms and pathway insights in early and advanced triple-negative breast cancer

To investigate the early molecular events in TNBC, gene interactions were analyzed using the STRING database, focusing on physical and genetic interactions of key genes involved in early TNBC progression. STRING provides a confidence score based on evidence types, including experimental data and co-expression, enabling the identification of biologically relevant relationships (Figure 7).

Figure 7.

Figure 7.

Physical and genetic interactions of genes (co-expressed genes) involved in the early stage of TNBC.

Our analysis identified several high-confidence physical interactions critical to TNBC progression (Figure 8, Table 3, and Fig. SI2-4). DNAJC6, involved in endocytosis, interacts with CLTC (score: 0.983), a gene known for its role in vesicle trafficking and cellular signaling. This interaction may regulate receptor recycling in TNBC cells, contributing to altered signaling pathways. Similarly, SKP2 interacts strongly with SKP1 (score: 0.999), a component of the SCF ubiquitin ligase complex. SKP2's role in degrading p27Kip1 and promoting the G1-to-S phase transition underscores its potential as a driver of unchecked cell proliferation in TNBC. Furthermore, the interaction between MOCOS and ISCU (score: 0.955) suggests involvement in metabolic adaptations, specifically within the trans-sulfuration pathway, which is often upregulated in cancers to counter oxidative stress. 48 According to information from the HPA, the MOCOS gene functions as a prognostic marker in both pancreatic and renal cancers. Lastly, NCAPD2 protein, a key player in chromatin condensation, forms a robust interaction with SMC4 (score: 0.993), highlighting the importance of proper chromosomal segregation and genomic stability during TNBC cell division.

Figure 8.

Figure 8.

KEGG pathway network illustrating the involvement of genes associated with the early stage of TNBC.

Table 3.

Physical and genetic interactions of genes (co-expressed genes) are involved in the early stage of TNBC.

Stage Gene Physical interaction Score Co-expressed genes Score Function
Early
DNAJC6 CLTC 0.983 SH3GL2 - Endophilin-A1 0.175 Endocytosis
SKP2 SKP1 0.999 CCNA2 - Cyclin-A2 0.331 Cell cycle
MOCOS ISCU 0.955 CTH 0.092 trans-sulfuration pathway
NCAPD2 SMC4 0.993 SMC4 0.882 Chromatin condensation

In advanced-stage TNBC (Figures 9 and 10, Table 4 and Fig. SI5-SI7), distinct gene interactions emerge. F11R-TJP1 (score: 0.948) regulates the integrity of tight junctions, a key process in metastasis. FOXO6-MYC (score: 0.561) links transcriptional regulation to proto-oncogenic signaling, contributing to TNBC's aggressive nature. PPP4C-PPP4R2 (score: 0.994) governs mitotic progression by regulating key mitotic processes, including chromosome segregation and spindle assembly. Additionally, the PPP4C protein plays a crucial role in DNA damage repair pathways, such as the dephosphorylation of H2AX, which is critical for maintaining genomic stability. In cancer, this dual role in mitosis and DNA repair may contribute to tumor cell survival and proliferation, particularly under the genomic instability characteristic of TNBC. TMEM51 interacts directly with TMEM86A (score: 0.062) and is implicated in signal transduction. Although the precise role of TMEM51 in cancer remains unclear, its involvement in cellular differentiation and signaling pathways suggests a potential contribution to tumor progression. In normal cells, TMEM51 is associated with cell signaling and differentiation, whereas in cancer cells, it may modulate pathways that promote cell survival, metastasis, and therapeutic resistance.49,50

Figure 9.

Figure 9.

Physical and genetic interactions of genes (co-expressed genes) involved in the advanced stage of TNBC.

Figure 10.

Figure 10.

KEGG pathway network illustrating the involvement of genes associated with the advanced stage of TNBC.

Table 4.

Physical and genetic interactions of genes (co-expressed genes) involved in the advanced stage of TNBC.

Stage Gene Physical interaction Score Coexpressed genes Score Function
Advanced
F11R TJP1 0.948 OCLN – Occludin 0.231 Regulation of the tight junction
FOXO6 MYC 0.561 MYC 0.078 Proto-oncogene
TMEM51 - - TMEM86A 0.062 Signal Transduction
PPP4C PPP4R2 0.994 (PPP4R2) 0.163 Mitotic progression

The identified interactions reveal critical transitions in TNBC progression. Early-stage genes emphasize cell cycle regulation, oxidative stress management, and chromosomal stability. In contrast, advanced-stage genes reflect changes in cell adhesion, oncogenic signaling, and mitotic regulation. These findings provide a foundation for developing targeted interventions specific to each TNBC stage.

Dynamic shifts in molecular processes are highlighted through the comparison of interactions across TNBC stages. These stage-specific insights may be used to guide the development of tailored therapeutic strategies targeting distinct pathways in TNBC progression.

Dual-Role miRNAs and differential co-regulation networks define stage-specific regulation in TNBC

Our analysis revealed distinct miRNA regulatory networks associated with early and advanced TNBC stages. We identified numerous stage-specific miRNAs, with a subset of 13 demonstrating a dual regulatory role by influencing genes in both stages, which may indicate their involvement in TNBC transition mechanisms (see Table SI4 for the complete list).

To explore potential cooperative interactions among the identified miRNAs, a pairwise correlation analysis was performed (Figure 11A and B). The results revealed higher pairwise correlations between miRNAs, suggesting functional synergy. Among the advanced-stage miRNAs, a notable co-expression pattern was observed between hsa-let-7b and hsa-let-7c (r = 0.89, p < 0.01), indicating potential synergistic regulatory effects. Similarly, within the early-stage miRNAs, strong co-expression was detected between hsa-mir-29a and hsa-mir-29c (r = 0.92, p < 0.01), suggesting coordinated gene regulation. These findings highlight stage-specific miRNA regulatory networks in TNBC, offering insights into potential biomarkers for early detection and advanced-stage therapeutic intervention.

Figure 11.

Figure 11.

Pairwise correlation analysis of miRNA co-expression was performed using Pearson's correlation coefficient (|r| > 0.8, p < 0.01), with results visualized in stage-specific heatmaps (Panel A: advanced TNBC; Panel B: early-stage TNBC). The right panels display miRNA expression dynamics across TNBC progression stages, analyzed using [RPM normalization], with significant dysregulation defined as [|log2FC| > 1, FDR-adjusted p < 0.05]. Panels C and D illustrate miRNA expression dynamics during TNBC progression, analyzed via RPM normalization. Significantly dysregulated miRNAs (|log2FC| > 1, FDR-adjusted p < 0.05) include hsa-let-7b and hsa-let-7c (Panel C) and hsa-mir-29a and hsa-mir-29c (Panel D). These miRNAs exhibit strong stage-specific expression patterns and high pairwise correlation.

Further analysis of miRNA expression dynamics across TNBC stages revealed significant downregulation of hsa-let-7b and hsa-let-7c (Figure 11C) as well as hsa-mir-29a and hsa-mir-29c (Figure 11D). Functional enrichment analysis using the DIANA TOOLS (http://diana.imis.athena-innovation.gr/DianaTools/index.php) demonstrated that the pairwise miRNAs associated with early-stage TNBC were predominantly involved in ECM-receptor interactions and PI3K-AKT signaling pathways. In contrast, miRNAs linked to advanced-stage TNBC were enriched in pathways related to cancer progression, apoptosis, and ribosome biogenesis. These results underscore distinct mechanistic roles of stage-specific miRNAs in TNBC pathogenesis.

Stage-Specific therapies and drug repurposing in TNBC

Following the identification of genes associated with early and advanced stages of TNBC, their relevance to conventional and repurposing drugs was analyzed using the CTD. Drug selection on the CTD was based on the downregulation of the regarded gene by the drug, and only FDA-approved drugs or those currently in clinical trials were considered.

In early-stage TNBC, the primary treatment goal is curative intent, combining surgery, chemotherapy, and radiation. Chemotherapeutic agents, such as anthracyclines, taxanes, and platinum-based drugs, play a central role in targeting cancer cells. These therapies may be administered as neoadjuvant (before surgery) or adjuvant (after surgery) approaches. Additionally, immunotherapy with atezolizumab and nab-paclitaxel may be considered for early-stage PD-L1-positive TNBC. Our findings identified several conventional drugs that are known to interact with genes implicated in early-stage TNBC (Table 5). For example, doxorubicin, which targets DNAJC6, 51 and cisplatin, targeting SKP2, 52 are already key components of standard chemotherapeutic regimens. NCAPD2, targeted by doxorubicin 51 and palbociclib, 53 further highlights the role of drugs that interfere with cell cycle progression. Additionally, JQ1, which has the ability to downregulate MOCOS, 54 adds another layer of insight into potential therapeutic strategies. These gene-drug interactions are consistent with the known mechanisms of current chemotherapies and suggest how gene-specific insights could potentially align with clinical strategies.

Table 5.

Conventional and repurposed drugs that exhibit genetic (indirect) interactions with genes involved in various stages of TNBC.

Stage Gene Conventional drugs Repurposing drugs Current clinical use
Early  
DNAJC6 Doxorubicin
  • Cyclosporine

  • Sunitinib

  • anti-rheumatoid arthritis

  • Anti-cancer

SKP2
  • Cisplatin

  • Afuresertib

  • Doxorubicin

  • Estradiol

  • Methotrexate

  • Raloxifene Hydrochloride

  • Arsenic trioxide

  • Azathioprine

  • Bortezomib

  • Calcitriol

  • Cannabidiol

  • Leflunomide

  • Sunitinib

  • Valproic Acid

  • Anti-cancer

  • Anti-inflammatory

  • Anti-cancer

  • Anti-hyperparathyroidism

  • Anti-seizures

  • Anti-rheumatoid arthritis

  • Anti-cancer

  • Anti-seizures

MOCOS JQ1
  • Isotretinoin

  • Lucanthone

  • Anti-seizures

  • Radiation sensitizer

NCAPD2
  • Doxorubicin

  • Palbociclib

  • Calcitriol

  • Cyclosporine

  • Dasatinib

  • Ivermectin

  • Valproic Acid

  • Leflunomide

  • Sunitinib

  • Anti-hyperparathyroidism

  • Anti-hyperparathyroidism

  • Anti-cancer

  • Anti-worm infection

  • Anti-seizures

  • Anti-rheumatoid arthritis

  • Anti-cancer

Advanced  
F11R Fulvestrant
  • Vorinostat

  • Anti-cancer

FOXO6 - -
TMEM51 -
  • Belinostat

  • Temozolomide

  • Anti-cancer

  • Anti-cancer

PPP4C -
  • Aspirin

  • Tretinoin

  • Anti-inflammatory

  • anti-acne, Anti-cancer

In the advanced stage, treatment focuses on controlling disease progression, improving survival, and maintaining quality of life. Systemic therapies, including chemotherapy (e.g., taxanes, gemcitabine, platinum-based drugs), immunotherapy (atezolizumab or pembrolizumab for PD-L1-positive cases), PARP inhibitors (for BRCA-mutated TNBC), and antibody-drug conjugates (e.g., sacituzumab govitecan), form the backbone of advanced-stage TNBC management.

In advanced-stage TNBC, genes such as F11R, FOXO6, TMEM51, and PPP4C are implicated. Conventional therapies have been identified for some of these genes, such as fulvestrant for F11R, 55 while others (FOXO6, TMEM51, and PPP4C) lack targeted conventional drugs, underscoring the need for further research (Table 5).

Repurposing drugs initially developed for other diseases represents a promising and innovative approach that warrants further investigation for TNBC treatment (Table 5). For early-stage TNBC, agents such as cyclosporine, sunitinib, arsenic trioxide, and isotretinoin demonstrated the potential to modulate key genes (DNAJC6, SKP2, MOCOS, and NCAPD2), targeting cancer-related pathways through mechanisms like angiogenesis inhibition, DNA damage induction, and epigenetic regulation. Note that while cyclosporine is primarily used as an immunosuppressant for preventing organ transplant rejection and treating autoimmune diseases, there is emerging research suggesting its potential use in cancer therapy, particularly in leukemia and lymphoma. 56 Valproic acid, originally used as an anticonvulsant and mood stabilizer, has been repurposed in cancer therapy, particularly for gliomas, breast cancer, and hematologic malignancies. As a histone deacetylase inhibitor, it has shown potential in overcoming drug resistance in cancer cells.57,58

For advanced-stage TNBC, drugs such as vorinostat, belinostat, and temozolomide were identified for genes like F11R and TMEM51. Vorinostat, belinostat, and temozolomide, identified as potential treatments for advanced-stage TNBC, have demonstrated efficacy in other cancers and may target key pathways in TNBC.59,60 Vorinostat and belinostat, both histone deacetylase (HDAC) inhibitors, have been approved for T-cell lymphomas and act by modulating epigenetic regulation, inducing apoptosis, and enhancing chemotherapy sensitivity. Temozolomide, an alkylating agent used in glioblastoma, induces DNA damage in rapidly dividing cells, making it effective against cancers with DNA repair deficiencies, such as TNBC. Notably, prior studies have shown that these drugs can reduce the expression of F11R and TMEM51, genes implicated in metastasis and tumor progression, further supporting their potential utility in TNBC management. 61 The predicted modulation of PPP4C by aspirin and tretinoin highlights the versatility of repurposed drugs in targeting inflammation and cancer pathways. Aspirin, in particular, shows promise for advanced-stage cancer treatment due to its anti-inflammatory and anti-metastatic effects, as demonstrated in other cancer types and supported by emerging studies in TNBC.62,63

Our findings highlight the potential utility of stage-specific treatment strategies in TNBC. Conventional drugs remain the foundation of therapy, particularly in the early stages, while the proposed repurposing drugs represent a complementary approach that could potentially be used to address treatment resistance and metastasis, especially in advanced stages. The identification of genes lacking conventional drug targets highlights areas for future research and drug development, paving the way for personalized medicine approaches in TNBC.

Domain-Specific docking analysis of repurposed drugs targeting early and advanced TNBC proteins

Docking analysis was conducted to evaluate direct interactions between repurposed drugs and proteins expressed from genes implicated in early and advanced stages of TNBC (Table 6 and Figure 12). For early-stage TNBC, MOCOS protein exhibited high affinity with lucanthone (−7.8 kcal/mol) and isotretinoin (−7.1 kcal/mol), potentially targeting its PLP-dependent transferase and MOSC domains. These interactions may disrupt sulfur transfer and molybdenum cofactor-related enzymatic activity, potentially altering metabolic pathways critical for cancer cell survival and proliferation, thereby impairing tumor progression. SKP2 protein demonstrated strong interactions with azathioprine (−7.6 kcal/mol) and leflunomide (−7.3 kcal/mol) at its leucine-rich repeat domain, suggesting a possible capacity to influence protein ubiquitination. Meanwhile, NCAPD2 protein demonstrated strong affinities with dasatinib (−7.2 kcal/mol) and leflunomide (−7.1 kcal/mol) at the HEAT repeat and C-terminal domains, highlighting their potential roles in disrupting cell cycle regulation and chromosomal dynamics essential for cancer cell proliferation.

Table 6.

Assessment of the direct interactions between repurposed drugs and their respective targets through docking analysis.

Stage Gene UniProt Repurposed drug Protein domain: docking score (Affinity (kcal/mol)
Early DNAJC6 P60510 Sunitinib Clathrin Domain: −6.8
SKP2 Q13309 Azathioprine F-box Domain: −6.0
Leucine-Rich Repeat Domain: −7.6
Calcitriol F-box Domain: −6.6
Leucine-Rich Repeat Domain: −6.0
Cannabidiol F-box Domain: −6.3
Leucine-Rich Repeat Domain: −6.5
Leflunomide F-box Domain: −6.9
Leucine-Rich Repeat Domain: −7.3
Sunitinib F-box Domain: −6.6
Leucine-Rich Repeat Domain: −6.7
Valproic Acid F-box Domain: −4.2
Leucine-Rich Repeat Domain: −3.9
MOCOS Q96EN8 Isotretinoin PLP-Dependent Transferase Domaina: −7.1
MOSC Domain: −6.2
Lucanthone PLP-Dependent Transferase Domain: −7.8
MOSC Domain: −7.5
NCAPD2 Q15021 Calcitriol HEAT Repeat Domainb: −6.2
C-terminal Domain: −6.9
Dasatinib HEAT Repeat Domain: −7.2
C-terminal Domain: −7.2
Valproic Acid HEAT Repeat Domains: −3.9
C-terminal Domain: −4.2
Leflunomide HEAT Repeat Domains: −6.5
C-terminal Domain: −7.1
Sunitinib HEAT Repeat Domains: −6.8
C-terminal Domain: −6.6
Advanced F11R Q9Y624 Vorinostat Immunoglobulin-like Domains: −5.0
FOXO6 A8MYZ6 - -
TMEM51 Q9NW97 Belinostat Extracellular Domain: −5.7
Temozolomide Extracellular Domain: −4.8
PPP4C P60510 Aspirin Catalytic domain: −5.7
a

PLP-Dependent Transferase: Pyridoxal Phosphate-Dependent Transferase.

b

HEAT Repeat Domain: Huntingtin, Elongation Factor 3 (EF3), Protein Phosphatase 2A (PP2A), and Target of Rapamycin 1 (TOR1).

Figure 12.

Figure 12.

Docking complexes of repurposed drugs with proteins involved in early and advanced stages of TNBC. (A) Docking complexes of Lucanthone with the MOCOS protein: (a) at different domains; (b) with the PLP-dependent transferase domain; (c) with the MOSC domain. (B) Docking complexes of Azathioprine with the SKP2 protein: (a) at different domains; (b) with the Leucine-rich domain; (c) with the F-box domain. (C) Docking complex of Vorinostat with (a) the F11R protein; (b) hydrogen bonds between Vorinostat and the immunoglobulin-like domain of the F11R protein.

For advanced-stage TNBC, F11R protein exhibited moderate interaction with vorinostat (−5.0 kcal/mol) at its immunoglobulin-like domains, and TMEM51 protein showed an affinity with belinostat (−6.5 kcal/mol) at its extracellular domain, indicating their potential in regulating adhesion and signaling. Aspirin was bound moderately to PPP4C protein (−5.7 kcal/mol) in its catalytic domain, suggesting a role in modulating inflammation and mitotic processes.

These findings position lucanthone, leflunomide, and azathioprine as promising candidates for further experimental validation in TNBC treatment while emphasizing the importance of domain-specific drug targeting for therapeutic efficacy.

Discussion

In this study, we identified four genes specifically associated with early-stage TNBC and four additional genes linked to advanced-stage TNBC. These stage-specific markers were found to play a crucial role in the molecular progression of TNBC. Their identification emphasized the need to move beyond a one-size-fits-all understanding of this aggressive cancer. Taken together, our findings supported the established paradigm that recognizes distinct biological subtypes within TNBC, which critically influence potential therapeutic strategies and patient outcomes.

The prognostic significance of the genes identified in our study was supported by their known roles in other malignancies. For example, DNAJC6, identified here for its relevance to TNBC outcomes, is associated with poor prognosis in hepatocellular carcinoma and shows a positive correlation with lymph node metastasis and reduced survival in TNBC patients, highlighting its potential as a potent oncogene.35,38 Similarly, SKP2, also identified in our analysis, is a well-characterized promoter of tumorigenesis; its overexpression drives proliferation and is linked to poor survival across multiple cancers. 64 For the less-studied genes, our study provided novel biological insights. MOCOS, previously designated as a prognostic marker in renal, pancreatic, and lung cancers, is a key regulator of purine metabolism and redox homeostasis through activation of xanthine dehydrogenase. Its dysregulation may contribute to cancer progression by disrupting metabolic energetics and promoting a pro-oxidant state that fuels genomic instability. Likewise, TMEM51, a prognostic marker in renal cancer, is implicated in cell adhesion and glycosylation. Its dysregulation could likely facilitate metastasis—a hallmark of advanced disease—by enhancing EMT and immune evasion. Collectively, these biologically grounded observations strengthen their credibility as compelling candidates for further functional validation.

This stage-specific molecular architecture also extended to our advanced-stage genes. F11R (JAM-A) has been shown to promote tumor progression and poor prognosis in multiple myeloma, while FOXO6 overexpression is linked to aggressive tumor behavior in gastric cancer.42,46 Additionally, the upregulation of PPP4C in lung cancer and its correlation with poor prognosis further validated our approach to identifying high-impact targets. 47

The translational potential of these findings was significant. By leveraging comprehensive drug databases, we identified both repurposed drugs and novel therapeutic agents that target these novel stage-specific candidate biomarkers. This aligned with the progress in immunotherapy for advanced TNBC, where the combination of a PD-1 inhibitor and GP regimen (gemcitabine and cisplatin) has been shown to provide precise therapeutic efficacy. This combination has facilitated the restoration of patients’ immune function, reduced tumor marker levels, suppressed tumor angiogenesis, and improved short-term prognosis with a favorable safety profile. 65 The efficacy and targeted delivery of these therapeutic candidates could be further enhanced by leveraging recent advances in biomaterials and nanoplatforms. For instance, smart hydrogels designed using machine learning offer a new platform for precise, stimuli-responsive drug release in cancer therapy. 66 This strategy highlights a promising path for developing personalized treatment regimens tailored to the specific genetic makeup and stage of a patient's TNBC, as well as the distinct signaling pathways active at each phase of progression. Our research substantiated the growing consensus that TNBC's heterogeneity demands a tailored therapeutic approach, and we provided an actionable framework by pinpointing bioactive small molecules aligned with TNBC staging.

Limitations of the study

While our study provides valuable insights into stage-specific molecular profiles in TNBC, several limitations must be acknowledged. The primary constraint is the inherent scarcity of advanced-stage TNBC samples in public repositories, which directly leads to an unbalanced distribution of samples across disease stages. This imbalance, a common feature of databases like TCGA and GEO, arises because it is more clinically challenging to obtain biopsies from patients in advanced stages. We recognize that this uneven sample size can affect the statistical power for analyses concerning the group with fewer samples (the advanced stage). Furthermore, the necessary integration of multiple datasets to achieve a sufficient cohort size introduces methodological challenges, including the risk of creating technical batch effects and the suboptimal practice of normalizing datasets individually rather than as a collective entity.

Additionally, an important limitation of our work is the absence of in vitro or functional experimental validation. As our findings are derived solely from transcriptomic analyses, further laboratory-based experiments will be required to verify the biological relevance and mechanistic roles of the identified genes.

To mitigate these concerns and ensure the robustness of our findings, we employed a multifaceted strategy: First, we integrated multiple cohorts from a unified platform (GPL-570) and applied collective normalization to minimize technical batch effects. Second, and most critically, our analytical design identified stage-specific dysregulation through independent comparisons of each stage against a common normal tissue control group, rather than a direct head-to-head comparison between the highly unbalanced stage groups. This approach, supplemented by stringent statistical filters and subsequent validation in an independent TCGA cohort, strengthens our confidence that the identified molecular signatures are robust and biologically meaningful.

Therefore, while we believe our methodology provides reliable results, we emphasize that these findings require validation in more balanced datasets or larger prospective studies. Future work with access to larger, well-balanced cohorts will be essential to further validate and refine these stage-specific biomarkers and to fully elucidate the molecular progression of TNBC.

Future direction

Building upon this work, future research should focus on several key areas. First, the identified biomarkers and their associated pathways require validation in prospective, large-scale clinical cohorts to confirm their diagnostic and prognostic utility across diverse populations. Furthermore, to translate these molecular insights into a clinical setting, innovative diagnostic methods should be explored. For instance, building upon the precedent of using serum Raman spectroscopy combined with a convolutional neural network for the rapid diagnosis of TNBC, 67 we propose developing a similar fast, low-cost, and non-invasive screening method. Second, a priority is the functional validation of key targets from our model, specifically SKP2, F11R, and TMEM51. Experimental investigations using in vitro and in vivo models are needed to elucidate the mechanistic roles of these and other understudied genes, and to evaluate the therapeutic efficacy and safety of the proposed drug candidates.

To enhance the clinical translatability of our findings, a promising strategy would be to integrate this molecular signature with established and readily available clinical metrics. For instance, combining our genetic panel with preoperative inflammatory and nutritional indices—such as the pan-immuno-inflammatory value (PIV) and albumin-to-globulin ratio (AGR), which are powerful prognostic tools in other cancers could yield a more robust and holistic model for patient risk stratification. 68

Furthermore, integrating multi-omics data (e.g., proteomics, metabolomics) may uncover additional layers of TNBC heterogeneity and identify novel therapeutic targets. Finally, exploration of combination therapies targeting multiple stage-specific pathways identified in this study could be a crucial strategy to overcome resistance and improve treatment outcomes.

Conclusion

This study provides stage-specific molecular insights into TNBC progression by identifying distinct biomarkers for early vs. advanced stages and their potential therapeutic targets. Using advanced transcriptomic analyses, including WGCNA, we uncovered distinct stage-specific gene modules. Early-stage TNBC was characterized by key drivers in cell cycle and DNA replication, such as DNAJC6, SKP2, MOCOS, and NCAPD2, while advanced-stage TNBC was associated with metastasis and immune evasion genes, including F11R, PPP4C, FOXO6, and TMEM51. Notably, UALCAN pan-cancer analysis revealed the consistent dysregulation of these genes across 23 cancer types. Further validation using literature and IHC data from the HPA confirmed their clinical relevance to tumor aggressiveness and prognosis. Moreover, dual-role miRNAs and stage-specific co-regulation networks were identified, with early-stage miRNAs enriched in PI3K-AKT signaling and advanced-stage miRNAs linked to tumor progression pathways. By integrating these findings with drug repurposing and molecular docking studies, we identified potential therapeutic options for each stage of TNBC.

The stage-stratified approach developed here, identifying stage-unique biomarkers and repurposing existing drugs, provides a transferable framework for other aggressive cancers lacking targeted therapies. This methodology bridges computational discovery with clinical applicability, supporting the broader shift in oncology toward molecularly-guided, stage-aware treatment strategies.

This study addresses a critical gap in TNBC research by emphasizing the stage-specific nature of its molecular underpinnings and therapeutic needs. Ultimately, this work contributes to the advancement of personalized medicine in breast cancer care, with the potential to improve survival rates and quality of life for TNBC patients.

Supplemental Material

sj-docx-1-cbm-10.1177_18758592261454886 - Supplemental material for Stage-specific biomarkers in triple-negative breast cancer: Preliminary findings from bioinformatics approaches

Supplemental material, sj-docx-1-cbm-10.1177_18758592261454886 for Stage-specific biomarkers in triple-negative breast cancer: Preliminary findings from bioinformatics approaches by Faeze Shahraki, Azadeh Meshkini and Elham Nazari in Cancer Biomarkers

Acknowledgments

The authors greatly appreciate the financial support by the Research Council of Ferdowsi University of Mashhad.

Footnotes

ORCID iD: Azadeh Meshkini https://orcid.org/0000-0002-5804-7827

Ethical approval and consent to participate: Ethical approval was not required for the present study because the analysis was performed exclusively on previously published, de-identified genomic data obtained from public databases (GEO and TCGA). This research, therefore, did not involve human participants or animals directly. Our study complies with the relevant national guidelines and institutional policies that exempt such secondary data analysis from ethics committee review.

Consent for publication: All authors consent for the manuscript to be published.

Authors’ contributions: FS: Methodology, Validation, Formal analysis, Investigation, Writing – original draft, AM: Project administration, Funding acquisition, Conceptualization, Resources, Methodology, Validation, Formal analysis, Writing - Review & Editing, EN: Investigation, Review & Editing.

Funding: This work was supported by Ferdowsi University of Mashhad, grant number: 3/62137.

The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Data availability statement: The data supporting the findings of this study are available from the corresponding author upon reasonable request.

Availability of supporting data: Gene expression data supporting this study's findings are publicly available in the Gene Expression Omnibus (GEO) under the accession numbers GSE76275, GSE29044, and GSE42568 for microarray data. Additionally, data from The Cancer Genome Atlas (TCGA) can be accessed through the Genomic Data Commons (GDC) portal (https://portal.gdc.cancer.gov/) under project code TCGA-BRCA.

Supplemental material: Supplemental material for this article is available online.

References

  • 1.Khan G, Hussain MS, Ahmad S, et al. Metabolomics as a tool for understanding and treating triple-negative breast cancer. Naunyn Schmiedebergs Arch Pharmacol 2025; 398: 13351–13370. [DOI] [PubMed] [Google Scholar]
  • 2.Griffiths CL, Olin JL. Triple negative breast cancer: a brief review of its characteristics and treatment options. J Pharm Pract 2012; 25: 319–323. [DOI] [PubMed] [Google Scholar]
  • 3.Obidiro O, Battogtokh G, Akala EO. Triple Negative Breast Cancer Treatment Options and Limitations: Future Outlook. Pharmaceutics 2023; 15: 1796. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Mahmoud R, Ordóñez-Morán P, Allegrucci C. Challenges for Triple Negative Breast Cancer Treatment: Defeating Heterogeneity and Cancer Stemness. Cancers (Basel) 2022; 14: 4280. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Chen Y, Anwar M, Wang X, et al. Integrative transcriptomic and single-cell analysis reveals IL27RA as a key immune regulator and therapeutic indicator in breast cancer. Discov Oncol 2025; 16: 977. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Qiu J, Xue X, Hu C, et al. Comparison of clinicopathological features and prognosis in triple-negative and non-triple negative breast cancer. J Cancer 2016; 7: 167–173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Hsu J-Y, Chang C-J, Cheng J-S. Survival, treatment regimens and medical costs of women newly diagnosed with metastatic triple-negative breast cancer. Sci Rep 2022; 12: 729. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wu Q, Siddharth S, Sharma D. Triple Negative Breast Cancer: A Mountain Yet to Be Scaled Despite the Triumphs. Cancers (Basel) 2021; 13: 3697. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Chen Y, Zhang B, Wang X, et al. Prognostic value of preoperative modified Glasgow prognostic score in predicting overall survival in breast cancer patients: a retrospective cohort study. Oncol Lett 2025; 29: 180. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Zhang Z, Zhang R, Li D. Molecular biology mechanisms and emerging therapeutics of triple-negative breast cancer. Biologics 2023; 17: 113–128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Lee J. Current Treatment Landscape for Early Triple-Negative Breast Cancer (TNBC). J Clin Med 2023; 12: 1524–1539. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Brierley J. The evolving TNM cancer staging system: an essential component of cancer care. C Can Med Assoc J = J L’Association Medicale Can 2006; 174: 155–156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Gospodarowicz MK, Miller D, Groome PA, et al. The process for continuous improvement of the TNM classification. Cancer 2004; 100: 1–5. [DOI] [PubMed] [Google Scholar]
  • 14.Li L, Huang H, Zhu M, et al. Identification of hub genes and pathways of triple negative breast cancer by expression profiles analysis. Cancer Manag Res 2021; 13: 2095–2104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Wei L-M, Li X-Y, Wang Z-M, et al. Identification of hub genes in triple-negative breast cancer by integrated bioinformatics analysis. Gland Surg 2021; 10: 799–806. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Liu X, Zhang G, Zhao L. Detection of transmembrane protein 100 in breast cancer: correlation with malignant progression and chemosensitivity. Cytojournal 2024; 21: 65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang H, Hao R, Liu W, et al. Identification of transcription factors associated with the disease-free survival of triple-negative breast cancer through weighted gene co-expression network analysis. Cytojournal 2024; 21: 71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Li X, Xiang J, Wang J, et al. FUNMarker: fusion network-based method to identify prognostic and heterogeneous breast cancer biomarkers. IEEE/ACM Trans Comput Biol Bioinforma 2021; 18: 2483–2491. [DOI] [PubMed] [Google Scholar]
  • 19.Hamid JS, Hu P, Roslin NM, et al. Data integration in genetics and genomics: methods and challenges. Hum Genomics Proteomics 2009; 2009: 869093. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhang S, Shao J, Yu D, et al. Matchmixer: a cross-platform normalization method for gene expression data integration. Bioinformatics 2020; 36: 2486–2491. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Burstein MD, Tsimelzon A, Poage GM, et al. Comprehensive genomic analysis identifies novel subtypes and targets of triple-negative breast cancer. Clin Cancer Res an Off J Am Assoc Cancer Res 2015; 21: 1688–1698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Colak D, Nofal A, Albakheet A, et al. Age-specific gene expression signatures for breast tumors and cross-species conserved potential cancer progression markers in young women. PLoS One 2013; 8: e63204. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Clarke C, Madden SF, Doolan P, et al. Correlating transcriptional networks to breast cancer survival: a large-scale coexpression analysis. Carcinogenesis 2013; 34: 2300–2308. [DOI] [PubMed] [Google Scholar]
  • 24.Smith A, Cavalli C, Harling L, et al. Impact of the TNM staging system for thymoma, Mediastinum; Vol 5 (December 25, 2021) Mediastinum. (2021). [DOI] [PMC free article] [PubMed]
  • 25.Irizarry RA, Hobbs B, Collin F, et al. Exploration, normalization, and summaries of high density oligonucleotide array probe level data. Biostatistics 2003; 4: 249–264. [DOI] [PubMed] [Google Scholar]
  • 26.Kenn M, Cacsire Castillo-Tong D, Singer CF, et al. Microarray normalization revisited for reproducible breast cancer biomarkers. Biomed Res Int 2020; 2020: 1363827. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ritchie ME, Phipson B, Wu D, et al. Limma powers differential expression analyses for RNA-Sequencing and microarray studies. Nucleic Acids Res 2015; 43: e47–e47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Smyth G, Thorne N, Wettenhall J. limma: Linear Models for Microarray Data User’s Guide. Bioinforma Comput Biol Solut Using R Bioconductor 2011. [Google Scholar]
  • 29.Langfelder P, Horvath S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics 2008; 9: 559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Rau A, Maugis-Rabusseau C. Transformation and model choice for RNA-Seq co-expression analysis. Brief Bioinform 2017; 19: 425–436. [DOI] [PubMed] [Google Scholar]
  • 31.Livesey M, Rossouw SC, Blignaut R, et al. Transforming RNA-Seq gene expression to track cancer progression in the multi-stage early to advanced-stage cancer development. PLoS One 2023; 18: e0284458. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Doerge RW. Bioinformatics and computational biology solutions using R and bioconductor edited by Gentleman, R., Carey, V., Huber, W., Irizarry, R., and Dudoit, S. Biometrics 2006; 62: 1270–1271. [Google Scholar]
  • 33.Li B, Lin R, Hua Y, et al. Single-cell RNA sequencing reveals TMEM71 as an immunomodulatory biomarker predicting immune checkpoint blockade response in breast cancer. Discov Oncol 2025; 16: 1256. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Davis AP, Grondin CJ, Johnson RJ, et al. Comparative toxicogenomics database (CTD): update 2021. Nucleic Acids Res 2021; 49: D1138–D1143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Yang T, Li X-N, Li X-G, et al. DNAJC6 Promotes hepatocellular carcinoma progression through induction of epithelial–mesenchymal transition. Biochem Biophys Res Commun 2014; 455: 298–304. [DOI] [PubMed] [Google Scholar]
  • 36.Saigusa K, Hashimoto N, Tsuda H, et al. Overexpressed Skp2 within 5p amplification detected by array-based comparative genomic hybridization is associated with poor prognosis of glioblastomas. Cancer Sci 2005; 96: 676–683. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Li Q, Murphy M, Ross J, et al. Skp2 and p27Kip1 expression in melanocytic nevi and melanoma: an inverse relationship. J Cutan Pathol 2004; 31: 633–642. [DOI] [PubMed] [Google Scholar]
  • 38.Zhang Y, Liu F, Zhang C, et al. Non-SMC condensin I Complex subunit D2 is a prognostic factor in triple-negative breast cancer for the ability to promote cell cycle and enhance invasion. Am J Pathol 2020; 190: 37–47. [DOI] [PubMed] [Google Scholar]
  • 39.Dong X, Liu T, Li Z, et al. Non-SMC condensin I complex subunit D2 (NCAPD2) reveals its prognostic and immunologic features in human cancers. Aging (Albany NY) 2023; 15: 7237–7257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Goetsch L, Haeuw J-F, Beau-Larvor C, et al. A novel role for junctional adhesion molecule-A in tumor proliferation: modulation by an anti-JAM-A monoclonal antibody. Int J Cancer 2013; 132: 1463–1474. [DOI] [PubMed] [Google Scholar]
  • 41.Leech AO, Vellanki SH, Rutherford EJ, et al. Cleavage of the extracellular domain of junctional adhesion molecule-A is associated with resistance to anti-HER2 therapies in breast cancer settings. Breast Cancer Res 2018; 20: 140. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Solimando AG, Brandl A, Mattenheimer K, et al. JAM-A as a prognostic factor and new therapeutic target in multiple myeloma. Leukemia 2018; 32: 736–743. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Magara K, Takasawa A, Osanai M, et al. Elevated expression of JAM-A promotes neoplastic properties of lung adenocarcinoma. Cancer Sci 2017; 108: 2306–2314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Rosager AM, Sørensen MD, Dahlrot RH, et al. Expression and prognostic value of JAM-A in gliomas. J Neurooncol 2017; 135: 107–117. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Li Q, Tang H, Hu F, et al. Silencing of FOXO6 inhibits the proliferation, invasion, and glycolysis in colorectal cancer cells. J Cell Biochem 2019; 120: 3853–3860. [DOI] [PubMed] [Google Scholar]
  • 46.Wang J-H, Tang H-S, Li X-S, et al. Elevated FOXO6 expression correlates with progression and prognosis in gastric cancer. Oncotarget 2017; 8: 31682–31691. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Wang B, Zhu X, Pan L, et al. PP4C Facilitates lung cancer proliferation and inhibits apoptosis via activating MAPK/ERK pathway. Pathol - Res Pract 2020; 216: 152910. [DOI] [PubMed] [Google Scholar]
  • 48.Rosado JO, Salvador M, Bonatto D. Importance of the trans-sulfuration pathway in cancer prevention and promotion. Mol Cell Biochem 2007; 301: 1–12. [DOI] [PubMed] [Google Scholar]
  • 49.Ye C, Ren S, Sadula A, et al. The expression characteristics of transmembrane protein genes in pancreatic ductal adenocarcinoma through comprehensive analysis of bulk and single-cell RNA sequence. Front Oncol 2023; 13: 1047377. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Hui L, Wang J, Zhang J, et al. lncRNA TMEM51-AS1 and RUSC1-AS1 function as ceRNAs for induction of laryngeal squamous cell carcinoma and prediction of prognosis. PeerJ 2019; 7: e7456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Verheijen M, Schrooders Y, Gmuender H, et al. Bringing in vitro analysis closer to in vivo: studying doxorubicin toxicity and associated mechanisms in 3D human microtissues with PBPK-based dose modelling. Toxicol Lett 2018; 294: 184–192. [DOI] [PubMed] [Google Scholar]
  • 52.Uddin S, Ahmed M, Hussain AR, et al. Bortezomib-mediated expression of p27Kip1 through S-phase kinase protein 2 degradation in epithelial ovarian cancer. Lab Investig 2009; 89: 1115–1127. [DOI] [PubMed] [Google Scholar]
  • 53.Liu F, Korc M. Cdk4/6 inhibition induces epithelial-mesenchymal transition and enhances invasiveness in pancreatic cancer cells. Mol Cancer Ther 2012; 11: 2138–2148. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.McCleland ML, Mesh K, Lorenzana E, et al. CCAT1 Is an enhancer-templated RNA that predicts BET sensitivity in colorectal cancer. J Clin Invest 2016; 126: 639–652. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Awada Z, Nasr R, Akika R, et al. DNA methylome-wide alterations associated with estrogen receptor-dependent effects of bisphenols in breast cancer. Clin Epigenetics 2019; 11: 138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Wright SJ, Keating MJ. Cyclosporine A in chronic lymphocytic leukemia: dual anti-leukemic and immunosuppressive role? Leuk Lymphoma 1995; 20: 131–136. [DOI] [PubMed] [Google Scholar]
  • 57.Brodie SA, Brandes JC. Could valproic acid be an effective anticancer agent? The evidence so far. Expert Rev Anticancer Ther 2014; 14: 1097–1100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Han W, Guan W. Valproic acid: a promising therapeutic agent in glioma treatment. Front Oncol 2021; 11: 687362. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Wawruszak A, Borkiewicz L, Okon E, et al. Vorinostat (SAHA) and Breast Cancer: An Overview. Cancers (Basel) 2021; 13: 4700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.El Omari N, Bakrim S, Khalid A, et al. Anticancer clinical efficiency and stochastic mechanisms of belinostat. Biomed Pharmacother 2023; 165: 115212. [DOI] [PubMed] [Google Scholar]
  • 61.Shinde V, Hoelting L, Srinivasan SP, et al. Definition of transcriptome-based indices for quantitative characterization of chemically disturbed stem cell development: introduction of the STOP-Tox(ukn) and STOP-Tox(ukk) tests. Arch Toxicol 2017; 91: 839–864. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Shiao J, Thomas KM, Rahimi AS, et al. Aspirin/antiplatelet agent use improves disease-free survival and reduces the risk of distant metastases in stage II and III triple-negative breast cancer patients. Breast Cancer Res Treat 2017; 161: 463–471. [DOI] [PubMed] [Google Scholar]
  • 63.Ma J, Fan Z, Tang Q, et al. Aspirin attenuates YAP and β-catenin expression by promoting β-TrCP to overcome docetaxel and vinorelbine resistance in triple-negative breast cancer. Cell Death Dis 2020; 11: 530. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Cai Z, Moten A, Peng D, et al. The Skp2 pathway: a critical target for cancer therapy. Semin Cancer Biol 2020; 67: 16–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Kaiyuan JQ, Zhu Q and Lv Q. Efficacy and prognosis analysis of PD-1 inhibitor combined with GP regimen in the treatment of patients with advanced triple-negative breast cancer. IJP 2024; 20: 1030–1039. [Google Scholar]
  • 66.Guo L, Fu Z, Li H, et al. Smart hydrogel: a new platform for cancer therapy. Adv Colloid Interface Sci 2025; 340: 103470. [DOI] [PubMed] [Google Scholar]
  • 67.Zeng Q, Chen C, Chen C, et al. Serum Raman spectroscopy combined with convolutional neural network for rapid diagnosis of HER2-positive and triple-negative breast cancer. Spectrochim Acta Part A Mol Biomol Spectrosc 2023; 286: 122000. [DOI] [PubMed] [Google Scholar]
  • 68.Li K, Chen Y, Zhang Z, et al. Preoperative pan-immuno-inflammatory values and albumin-to-globulin ratio predict the prognosis of stage I-III colorectal cancer. Sci Rep 2025; 15: 11517. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

sj-docx-1-cbm-10.1177_18758592261454886 - Supplemental material for Stage-specific biomarkers in triple-negative breast cancer: Preliminary findings from bioinformatics approaches

Supplemental material, sj-docx-1-cbm-10.1177_18758592261454886 for Stage-specific biomarkers in triple-negative breast cancer: Preliminary findings from bioinformatics approaches by Faeze Shahraki, Azadeh Meshkini and Elham Nazari in Cancer Biomarkers


Articles from Cancer Biomarkers: Section A of Disease Markers are provided here courtesy of SAGE Publications

RESOURCES