Abstract
Background
Luminal breast cancer has a high incidence and significant heterogeneity. Exosomes, as key mediators of intercellular communication, are involved in the tumor process, but the related genes in Luminal breast cancer regarding prognosis and mechanism remain unclear.
Methods
This study integrated a total of 1251 Luminal breast cancer samples from the TCGA and GEO databases. Differential expression analysis and Cox regression were used to screen out prognostic-related exosomal genes, and a multivariate Cox risk prognostic model was constructed. Kaplan-Meier survival analysis and time-dependent ROC curves were used to evaluate the model performance. Further immunological analysis, gene set enrichment analysis (GSEA), transcription factor/ceRNA regulatory network construction, and drug sensitivity analysis were conducted. Results: A total of 4 prognostic-related exosomal genes (LCN2, RRAS2, LRG1, LTF) were identified. Based on this, a risk model was constructed that can effectively distinguish high-risk and low-risk patients, and the overall survival of the high-risk group was significantly shorter. Immunological analysis showed that the high-risk group had increased M2 Macrophage infiltration and accompanied by immune function inhibition. GSEA suggested that the high-risk group was enriched in tumor progression-related pathways such as DNA replication and homologous recombination. Drug sensitivity analysis indicated that the high-risk group showed higher sensitivity to PI3K inhibitors and chemotherapy drugs.
Conclusion
LCN2, RRAS2, LRG1, and LTF can be used as prognostic markers for Luminal breast cancer, related to immune suppression and tumor progression, and the high-risk group is more sensitive to PI3K inhibitors and chemotherapy drugs, providing potential targets for precision treatment.
Keywords: Luminal breast cancer, Exosomes, Bioinformatics, Immune infiltration, Drug sensitivity
Introduction
Global cancer statistics from 2022 indicate that approximately 20 million individuals were newly diagnosed with cancer worldwide. Among these cases, female breast cancer (BC) represented 11.6% of total cancer incidences and was responsible for 6.9% of global cancer-related mortality [1]. A marked increase in both incidence and mortality rates has been observed for female BC, underscoring its considerable threat to women’s health [2]. Historically, BC has been stratified into four principal molecular subtypes: luminal (estrogen receptor-positive, ER+), basal-like, HER2-overexpressing, and normal-like [3]. The luminal subtype, which constitutes more than 60% of all breast cancers, is the most prevalent [4]. Relative to other BC subtypes, luminal tumors generally demonstrate heightened sensitivity to endocrine therapy and are associated with a more favorable prognosis. Nevertheless, achieving optimal clinical outcomes continues to depend on tailored, multimodal treatment strategies [5].
The development of BC involves a combination of genetic, hormonal, and lifestyle-related factors [6]. Current diagnostic standards depend heavily on imaging and histopathological examination. Although these conventional approaches provide essential diagnostic information, they are limited in early detection and are inherently invasive. In contrast, liquid biopsy offers a minimally invasive alternative with high sensitivity, enabling rapid diagnosis through the analysis of readily accessible biofluids such as blood, urine, or saliva. These bodily fluids are rich in extracellular vesicles (EVs), including exosomes derived from the endoplasmic reticulum and microvesicles formed by budding from the plasma membrane. Such vesicles carry a diverse array of biomolecules—such as lipids, proteins, and nucleic acids—and play a key role in mediating intercellular communication [7]. Significant heterogeneity exists among breast cancer molecular subtypes, leading to pronounced differences in prognosis and treatment response. In many cases, therapeutic efficacy is limited by the emergence of drug resistance [6, 8]. For Luminal BC, endocrine therapy—which inhibits estrogen-driven proliferation—plays a central role and has substantially improved clinical outcomes [9]. Despite advances combining surgery, chemoradiotherapy, and endocrine therapy, a subset of Luminal BC patients still experience disease recurrence during or after treatment. Thus, the identification of more reliable prognostic biomarkers and novel therapeutic targets remains essential to further improve outcomes in this population.
Exosomes are cell-derived vesicles characterized by a double-layered membrane and a diameter ranging from approximately 30 to 200 nanometers. They are abundantly present in diverse bodily fluids and tissues, encapsulating a variety of molecular cargoes such as nucleic acids, proteins, and lipids. Through this capacity, exosomes serve as critical mediators of intercellular communication and play an essential role in various physiological processes [10]. In oncology, exosomes are increasingly recognized for their involvement in tumor initiation, progression, and metastasis. Studies suggest that exosomes can shape the tumor microenvironment, thereby influencing cancer development, treatment response, and disease outcome [11]. Hannafon et al. reported that [12] plasma extracellular vesicles containing miR-1246 and miR-21 may serve as potential biomarkers for BC, with combined detection offering higher diagnostic accuracy than either marker alone. Despite these advances, the prognostic relevance of extracellular vesicle–associated genes in luminal BC and their underlying biological mechanisms remain poorly understood.
Recent advances in high-throughput sequencing, big data analytics, and artificial intelligence have driven substantial progress in bioinformatics. Bioinformatics has become essential for deciphering the complex molecular mechanisms underlying BC, particularly through the integration of multi-omics data—including genomics, transcriptomics, proteomics, and metabolomics. The application of bioinformatics tools facilitates the identification of potential diagnostic and prognostic biomarkers and helps elucidate processes involved in tumor initiation, progression, and metastasis, thereby supporting more precise BC management [13, 14]. Despite these contributions, current bioinformatics-based studies in BC face several limitations, such as inadequate sample sizes, reliance on single data sources, and a lack of external validation, which collectively constrain the robustness and generalizability of their findings.
By integrating transcriptomic datasets from the TCGA and GEO databases, this study comprehensively characterized the expression landscape of exosome-related genes in luminal breast cancer and evaluated their prognostic relevance. Differentially expressed exosome-associated genes with prognostic significance were identified through combined differential expression analysis and univariate Cox regression, and subsequently used to construct and validate a risk prediction model. Beyond survival analysis, we further investigated the relationships between these candidate genes and key biological features, including immune-related activity, tumor mutational burden, and therapeutic response. Moreover, ceRNA interactions and transcription factor–mediated regulatory networks were constructed to elucidate the underlying biological mechanisms. This multi-level investigation into the molecular mechanisms of Luminal BC aims to provide novel strategies for precision prognosis and treatment in Luminal BC.
Materials and methods
Raw data download
Transcriptomic data, single nucleotide variation (SNV) data, and corresponding clinical information were retrieved from The Cancer Genome Atlas (TCGA) database [15]. Patients with luminal breast cancer—defined as estrogen receptor (ER)-positive, progesterone receptor (PR)-positive, and human epidermal growth factor receptor 2 (HER2)-negative—were selected according to established molecular subtyping criteria. This yielded 518 tumor samples, together with 113 normal breast tissue samples for control comparisons. For external validation, the corresponding cohort was acquired from the Gene Expression Omnibus (GEO) database [16]. Inclusion criteria are patients with Luminal breast cancer who have complete overall survival information and clearly defined molecular subtype information. A total of 4 GEO datasets were ultimately included, comprising GSE17705 [17] (298 cases), GSE21653 [18, 19] (122 cases), GSE25066 [20–22] (236 cases), and GSE69031 [23] (77 cases), totaling 733 Luminal breast cancer patients for subsequent validation analysis. The exosome-related gene set was derived from the GeneCards database [24] (https://www.genecards.org/). A keyword search identified 1,143 exosome-related genes, which were used as a candidate gene set for subsequent bioinformatics analysis.
Raw data processing and cleaning
We preprocessed TCGA breast cancer transcriptome expression data and extracted exosome-related genes for subsequent differential expression analysis. Only retain gene expression information with explicit gene annotations, and perform gene-level screening of the expression matrix based on the exosome-related gene set obtained from the GeneCards database. The expression data from the GEO dataset are processed uniformly within the R programming environment. We first standardized the gene names in the expression matrices across all datasets and removed duplicates, retaining only genes common to all datasets for subsequent merged analysis. Subsequently, merge the expression matrices from different GEO datasets and set batch information based on the data source. The merged expression data underwent batch correction using the ComBat method from the sva package. We compared the distribution of samples before and after batch correction using principal component analysis to evaluate the effectiveness of batch effect correction. The expression matrix calibrated by batches is used for subsequent survival analysis and model construction.
Differential gene expression analysis
Differential expression analysis was performed using standardized breast cancer transcriptomic data from TCGA. The expression matrix was first processed at the gene level to remove duplicate entries and filter out low-expression genes; only those with an average expression value greater than 0 were retained for downstream analysis. Counts per million (CPM) were then calculated for each gene using the edgeR package, and genes with an average CPM > 1 were kept. The filtered dataset was normalized and log-transformed. Limma-based differential expression analysis was carried out. Samples were grouped into normal breast tissue (n = 113) and luminal BC (n = 518) based on their origin. A linear model was fitted and empirical Bayes moderation was applied. P-values were corrected for multiple testing using the false discovery rate (FDR) method. Genes with |log₂ fold change| > 1 and an adjusted P-value < 0.05 were defined as differentially expressed. Results were visualized using a volcano plot and a heatmap, with the heatmap displaying the most significantly altered genes.
Screening of prognosis-associated genes, model building, and SHAP interpretation analysis
Overall survival served as the endpoint for screening differentially expressed genes (DEGs) within the TCGA cohort. Each DEG was subjected to univariate Cox proportional hazards regression analysis (via the survival package) to evaluate its association with prognosis. DEGs exhibiting a significant association (P < 0.05) in the univariate analysis were retained and subsequently incorporated into a multivariate Cox model to construct the final prognostic signature. To interpret the contribution of each gene in the multivariate model, we employed SHAP (Shapley Additive Explanations) analysis [25]. The SHAP value for each gene was calculated to estimate its relative importance and directional effect on the model output. Genes were then ranked according to their mean absolute SHAP value. Overall feature contributions were visualized using bar and beeswarm plots, while individual sample-level risk contributions were displayed using waterfall plots. Using the linear predictor derived from the multivariate Cox model, we calculated a risk score for each patient. Based on the median risk score, patients were then stratified into high- and low-risk groups. This stratification was subsequently evaluated using Kaplan–Meier survival analysis in both the TCGA training cohort and the GEO validation cohort to compare overall survival between the two groups, with statistical significance assessed by the log-rank test.
Dimension reduction visualization for risk stratification and time-dependent ROC evaluation
Model performance was further assessed in the TCGA cohort based on the calculated risk scores. To explore global expression differences of the signature genes between the high-risk and low-risk groups, the corresponding gene expression matrix was extracted from the risk score dataset. Principal component analysis (PCA) was conducted using the prcomp function with data scaling enabled. Samples were subsequently projected onto a two-dimensional principal component space and colored according to risk classification, with confidence ellipses applied to depict the distribution patterns within each group. Additionally, time-dependent receiver operating characteristic (ROC) analysis was conducted using the timeROC package to assess the predictive performance of the risk score. Using overall survival time and status as endpoint variables and the risk score as a continuous predictor, we assessed the model’s discriminative ability at 1-, 3-, and 5-year follow-up intervals by generating ROC curves and calculating the corresponding area under the curve (AUC) values.
Analysis of immune infiltration and immune function
Immune infiltration within the TCGA cohort was assessed by estimating the relative proportions of 22 immune cell types using the CIBERSORT algorithm (1,000 permutations) [26]. Subsequent analyses included only samples with a CIBERSORT output P-value < 0.05. To compare immune infiltration between risk groups, the risk stratification file was matched to the CIBERSORT results, including tumor samples only. The immune cell composition of high- and low-risk groups was visualized using stacked bar plots. Immune cell fractions were then converted to long format, and differences between risk groups were compared via boxplots, with significance tested by between-group comparison and annotated with appropriate symbols. We computed immune-related functional scores using ssGSEA (GSVA package) and compared them between risk groups within tumor samples via boxplots. We then assessed the correlation between the expression of model genes and the abundance of infiltrating immune cells, using samples for which both gene expression profiles and corresponding immune cell estimates derived from CIBERSORT were available. Spearman rank correlation was applied to calculate correlation coefficients and P-values, and statistically significant relationships were visualized using a correlation-coupling plot generated with the linkET package, highlighting interaction patterns between signature genes and immune cell populations.
Tumor mutation burden calculation and correlation analysis
Using the Mutation Annotation Format (MAF) file from TCGA, we calculated the tumor mutational burden (TMB). Variants lacking amino acid alterations were excluded, and mutations annotated as “Silent” or “Splice_Region” were removed to retain only non-synonymous events. For each sample, the total number of retained mutations was quantified and normalized to an assumed exonic coverage of 38 Mb, yielding TMB values expressed as mutations per megabase. In parallel, a gene-level mutation matrix was generated to summarize mutation frequencies across samples. TMB values were subsequently integrated with TCGA-derived risk stratification data to facilitate group-wise comparison. Following log₂(TMB + 1) transformation, differences in TMB between the high-risk and low-risk groups were visualized using boxplots and assessed for statistical significance. The relationship between tumor mutational burden (TMB) and patient survival was also examined in the TCGA cohort. An optimal TMB cutoff was established using the maximally selected rank statistics method implemented in the survminer package, which was then used to classify patients into high- and low-TMB groups. We performed Kaplan–Meier survival analysis and log-rank testing to compare overall survival across TMB groups. In parallel, we integrated TMB with the risk model to establish a four-group composite stratification, thereby assessing their combined prognostic effect.
GSEA enrichment analysis
Following alignment of expression data with the established TCGA risk stratification, mean expression levels for each gene were calculated separately within the high- and low-risk groups. Subsequent gene set enrichment analysis (GSEA) was then performed to compare pathway activities between these groups. We then computed log2(meanHigh) − log2(meanLow) as a gene ranking metric and constructed a pre-ranked gene list in descending order. Enrichment analysis employed KEGG pathway gene sets (c2.cp.kegg; MSigDB official website: https://www.gsea-msigdb.org/gsea/msigdb/) from the MSigDB database, implemented using GSEA within the clusterProfiler package, and outputted enrichment results. Pathway enrichment direction was determined based on normalized enrichment scores (NES). Pathways significantly enriched in the high-risk group (NES > 0) and low-risk group (NES < 0) were separately identified. Corresponding GSEA plots were generated to visualize enrichment results.
ceRNA network construction and visualization
We constructed a ceRNA network based on the core genes of the prognostic model. Firstly, we collected the miRNA target gene prediction results from four databases (miRanda, miRDB, TargetScan, and miRWalk) (each read as miRNA–mRNA pairs), and used the core gene list as the input to extract the candidate miRNAs corresponding to each gene. To enhance reliability, we only retained the miRNA–mRNA relationships that were simultaneously supported by all four databases. In the candidate screening, we standardized the prediction scores from different databases within the database and summarized them, sorted the candidate miRNAs by the comprehensive score, and retained the top 6 miRNAs for each gene for subsequent analysis to generate the miRNA–mRNA relationship file (gene miRNA_refined). Subsequently, we introduced miRNA–lncRNA interaction information from the spongeScan database and took the intersection with the aforementioned screened miRNAs to obtain miRNA–lncRNA pairs. Finally, a ceRNA network was constructed by integrating miRNA–mRNA and miRNA–lncRNA interactions, with node attributes (mRNA, miRNA, lncRNA) annotated accordingly. The resulting network was visualized and its topology displayed using Cytoscape software (v3.10.1, https://cytoscape.org/).
Analysis of transcription factor regulatory networks
We constructed a transcription factor (TF) regulatory network for the core genes of our model based on the TRRUST database [27]. TRRUST is a manually curated database of human transcription factor–target gene regulatory relationships (official website: https://www.grnpedia.org/trrust/). We downloaded the TRRUST human transcription regulation raw data, extracted the TF and target gene information, and imported the model core gene list. By screening the regulatory relationships between target genes and core genes in the TRRUST database, we obtained TF–Target pairs associated with core genes. After deduping and organizing the results, we constructed a transcription factor regulatory network. The resulting TF–gene regulatory relationships were utilized for network construction and topological visualization. Both the transcription factor regulatory network and the previously mentioned ceRNA network were visualized using Cytoscape software.
Drug sensitivity prediction analysis
We employed the oncoPredict package to predict drug sensitivity in TCGA breast cancer tumor samples, focusing on evaluating potential response differences to commonly used chemotherapy agents, endocrine therapies, and targeted therapies. First, the tumor sample transcriptome expression matrix underwent gene deduplication and low-expression filtering (retaining genes with average expression values greater than 0.5). Normal samples were excluded, and only tumor samples were included for predictive analysis. The drug response training data is sourced from the GDSC2 training set within the Genomics of Drug Sensitivity in Cancer (GDSC) database [28](official website: https://www.cancerrxgene.org/). We read the GDSC2 cell line expression profiles and corresponding drug response data, then use the calcPhenotype function in oncoPredict to transfer the training set model to the TCGA cohort, predicting the sensitivity phenotype of each sample to different drugs. During prediction, an empirical Bayesian approach was employed for batch correction (batchCorrect = eb). Drug response data underwent power transformation processing, while low-varying genes were removed (removeLowVaryingGenes = 0.2) to enhance the model’s robustness against interference. After obtaining predicted drug sensitivities for each sample, we compared sensitivities across risk groups and visualized the differences using boxplots.
Statistical analysis
We performed all analyses in R (v4.5.0). We assessed differential expression with edgeR and limma, applying cutoffs of |log₂FC| > 1 and FDR < 0.05. Prognostic genes screened by univariate Cox regression were used to build a multivariate Cox model. We evaluated the model using PCA and time-dependent ROC curves (AUC). Gene contributions were interpreted via SHAP values (DALEX package, kernel SHAP method), and prognostic utility was validated with Kaplan–Meier analysis and the log-rank test. TMB was calculated based on MAF files. Immune cell proportions were estimated using CIBERSORT, with intergroup differences assessed via the Wilcoxon rank-sum test. Regulatory networks comprising competing endogenous RNAs (ceRNAs) and transcription factors were constructed from the core genes. Gene set enrichment analysis (GSEA) was conducted using the clusterProfiler package, and drug sensitivity was predicted with oncoPredict utilizing data from the GDSC2 database. A P-value < 0.05 was considered statistically significant.
Results
Data integration and batch correction effect evaluation
We integrated and analyzed breast cancer transcriptomic data from the TCGA database alongside validation datasets obtained from GEO. Data integration was followed by batch-effect correction using the ComBat algorithm. The success of this normalization was evidenced by principal component analysis (PCA), which demonstrated improved coherence across the four GEO datasets post-correction (Fig. 1A, B).
Fig. 1.
Effect of batch effect correction. A Sample distribution before batch correction; B Sample distribution after batch correction
Acquisition of exosome-related differentially expressed genes (DEGs)
Differential expression analysis was conducted on standardized TCGA-BRCA transcriptome data, comparing 518 luminal breast cancer tissues with 113 normal breast tissues. This analysis identified 106 exosome-related DEGs, consisting of 34 upregulated and 72 downregulated genes (|log₂FC| > 1, FDR < 0.05) (Fig. 2A, B).
Fig. 2.
Identification results of exosome-associated DEGs. A Heatmap of exosome-associated DEGs; B Volcano plot of exosome-associated DEGs
Screening of prognosis-related exosomal DEGs and SHAP interpretation analysis
To identify prognostic exosomal DEGs, univariate Cox regression was applied to the TCGA cohort. This analysis yielded four significant genes (LCN2, RRAS2, LRG1, and LTF), and their hazard ratios were presented in a forest plot (Fig. 3A).
Fig. 3.
Forest plot and SHAP interpretation analysis results of exosomal DEGs associated with prognosis. A Forest plot of univariate Cox proportional hazards regression analysis; B SHAP summary bar chart; C SHAP swarm plot; D SHAP waterfall plot
The SHAP summary plot (Fig. 3B) displayed the mean absolute SHAP values for the four prognostic genes: LCN2 (0.294), RRAS2 (0.283), LRG1 (0.158), and LTF (0.088), with LCN2 contributing most and LTF least to the model output. Consistent negative SHAP values for high expression of these genes were observed in the swarm plot (Fig. 3C), indicating that elevated expression correlates with an increased risk prediction, thus functioning as negative prognostic predictors.
The SHAP waterfall plot (Fig. 3D) further illustrates the contribution of individual genes to the model’s prediction for representative samples. With a baseline model output E[f(x)] of -5.6, the expression of LCN2 (4.85) corresponds to a decrease of 0.327 in the predicted value. In comparison, LRG1 (expression = 4.76) contributes a slight increase of 0.07, whereas LTF (expression = 6.23) and RRAS2 (expression = 4.13) are associated with decreases of 0.0537 and 0.0254, respectively. These combined effects produced a final prediction of f(x) = -5.94. Since this value falls below the prognostic classification threshold (< -5.6), the sample was classified as having a poor prognosis.
Risk model validation and performance evaluation
Kaplan–Meier analysis revealed significantly shorter overall survival in the high-risk group compared to the low-risk group, both in the TCGA training cohort (p < 0.001; Fig. 4A) and the GEO validation cohort (p = 0.033; Fig. 4B). Principal component analysis (PCA) of the model genes in the TCGA cohort showed clear separation between risk groups (Fig. 4C), indicating distinct expression patterns. Time-dependent ROC analysis demonstrated the model’s robust and consistent predictive performance, with AUC values of 0.886, 0.691, and 0.742 at 1, 3, and 5 years, respectively (Fig. 4D).
Fig. 4.
Kaplan–Meier survival analysis, PCA, and time-dependent ROC evaluation results. A Kaplan–Meier survival analysis in the TCGA training cohort; B Kaplan–Meier survival analysis in the GEO external validation cohort; C Principal Component Analysis (PCA); D Time-dependent ROC analysis
Risk model immune landscape and its correlation with prognostic-related exosomal DEGs
We performed immune profiling to compare the high- and low-risk groups defined by the exosomal DEG-based model. This included CIBERSORT analysis of 22 immune cell types (such as B cells, T cells, NK cells, macrophages, dendritic cells) for infiltration abundance, and ssGSEA for immune functional evaluation (Fig. 5A). Comparative analysis revealed distinct immune infiltration patterns between the risk groups (Fig. 5B). Specifically, the high-risk group showed significantly elevated M2 macrophage infiltration and reduced levels of T follicular helper cells and activated dendritic cells relative to the low-risk group.
Fig. 5.
Immunological landscape across risk groups and correlation analysis with prognosis-related exosomal DEGs. A Percentage of immune cell infiltration across risk groups; B Differences in immune cell infiltration across risk groups; C Differences in immune function across risk groups; D Correlation between prognosis-related exosomal DEGs and immune cell infiltration (*, p < 0.05; **, p < 0.001; p < 0.01; ***, p < 0.001)
Further evaluation of immune-related functional activity revealed broad suppression of immune functions in the high-risk group (Fig. 5C). Specifically, multiple immune functional signatures, including antigen-presenting cell (APC) co-inhibition and co-stimulation, chemokine receptor activity, immune checkpoint signaling, dendritic cell–related functions, HLA expression, inflammatory and para-inflammatory responses, as well as effector immune cell signatures such as NK cells, neutrophils, mast cells, and various T cell subsets, were significantly downregulated relative to the low-risk group. Collectively, these alterations suggest an immunosuppressive tumor microenvironment in high-risk luminal breast cancer. Impaired immune surveillance and antitumor immune activity may facilitate tumor progression and immune evasion, thereby contributing to unfavorable clinical outcomes in this patient population.
Spearman correlation analysis was conducted to examine the associations between the expression of the four model genes (LCN2, RRAS2, LRG1, and LTF) and immune cell infiltration levels (Fig. 5D). The results indicated that RRAS2 expression was positively correlated with resting CD4 memory T cells, resting dendritic cells, and activated dendritic cells, while showing a negative correlation with M2 macrophages. LRG1 was positively associated with naïve B cells, resting dendritic cells, activated dendritic cells, and neutrophils, while negatively associated with M2 macrophages. LTF was positively correlated with naïve B cells, monocytes, resting dendritic cells, and resting mast cells, while showing negative correlations with M0 and M2 macrophages. In contrast, LCN2 showed positive correlations with naïve B cells, T follicular helper cells, and activated dendritic cells, and negative correlations with memory B cells and M2 macrophages.
Tumor mutation burden (TMB) differential analysis and prognostic value assessment
Tumor mutational burden (TMB) showed no significant difference between the high- and low-risk groups (Fig. 6A). Consistently, Kaplan–Meier analysis revealed no significant survival difference between patients with high versus low TMB (P = 0.071; Fig. 6B). Furthermore, a composite stratification combining TMB status and risk groups into four categories revealed no significant overall survival difference among them (P = 0.167; Fig. 6C).
Fig. 6.
TMB differential analysis and Kaplan–Meier survival analysis. A TMB differences across risk groups; B Kaplan–Meier survival analysis by TMB stratification; C Kaplan–Meier survival analysis with combined TMB and risk score stratification
Gene set enrichment analysis (GSEA) based on risk stratification
We performed GSEA on the TCGA cohort to assess pathway differences between risk groups. In the high-risk group, pathways related to tumor progression and proliferation—such as DNA replication, ganglioside biosynthesis, and homologous recombination—were markedly enriched (Fig. 7A). Conversely, the low-risk group exhibited significant enrichment in immune- and metabolism-related pathways, including the complement and coagulation cascade, cytokine–cytokine receptor interaction, and galactose metabolism (Fig. 7B). These enrichment patterns suggest that low-risk tumors reside in a more immunologically active microenvironment, while high-risk tumors are characterized by enhanced proliferative and DNA repair capacities—hallmarks of increased tumor aggressiveness and poorer clinical outcomes.
Fig. 7.
GSEA results for different risk groups. A GSEA for the high-risk group; B GSEA for the low-risk group
Construction and analysis of ceRNA networks and transcription factor (TF) regulatory networks
The ceRNA network (Fig. 8A) was constructed by integrating prediction results from four databases: miRanda, miRDB, TargetScan, and miRWalk. This network comprised 60 nodes, including 4 mRNA nodes, 9 miRNA nodes, and 47 lncRNA nodes. Within this network, each prognostic-related exosomal DEG (core gene) was connected to multiple lncRNAs via miRNAs, forming a complex regulatory interaction. The transcription factor (TF) regulatory network for the model’s core genes (Fig. 8B) revealed that the LCN2 gene was regulated by multiple TFs, including STAT1, RELA, NFKB1, and NR3C2. The core gene LTF was regulated by SP1 and CEBPE, while RRAS2 was regulated by BTF3. These TFs may influence core gene expression directly or indirectly, potentially affecting related biological processes and disease progression.
Fig. 8.
ceRNA network and transcription factor regulatory network of core genes. A ceRNA network associated with prognostic core genes. Red rectangles denote core genes, green triangles indicate miRNAs, and blue ovals represent lncRNAs. Edges represent predicted interactions between miRNAs and mRNAs or lncRNAs. B Transcription factor regulatory network of core genes. Red ellipses represent core genes, blue diamonds indicate transcription factors, and edges denote predicted regulatory relationships
Risk-based drug sensitivity analysis
Drug sensitivity analysis using the oncoPredict package revealed significant differences in sensitivity to multiple drugs between the high-risk group and the low-risk group. Specifically, the high-risk group exhibited significantly higher sensitivity to the PI3K inhibitors Alpelisib (p = 2.1e-06) (Fig. 9A), Buparlisib (p = 4.4e-07) (Fig. 9B), and Taselisib (p = 3.5e-06) (Fig. 9H) compared to the low-risk group. While the high-risk group was more sensitive to docetaxel (p = 0.00027) (Fig. 9C), epirubicin (p = 0.00013) (Fig. 9D), and paclitaxel (p = 0.015) (Fig. 9E) than the low-risk group, no significant intergroup differences were observed for the CDK4/6 inhibitors palbociclib (p = 0.05) (Fig. 9F) and ribociclib (p = 0.89) (Fig. 9G). No significant difference in tamoxifen sensitivity was observed between the risk groups (p = 0.67) (Fig. 9I). These findings indicate that the risk model may not reliably predict efficacy for all therapeutic classes, such as CDK4/6 inhibitors and endocrine therapy.
Fig. 9.
Drug sensitivity analysis by risk stratification. A Alpelisib; B Buparlisib; C Docetaxel; D Epirubicin; E Paclitaxel; F Palbociclib; G Ribociclib; H Taselisib; I Tamoxifen
Discussion
Luminal breast cancer (BC) is the most common subtype of BC, exhibiting distinct biological characteristics and clinical manifestations. Although treatment for Luminal BC has advanced significantly, recurrence and drug resistance remain major clinical challenges, substantially limiting long-term patient survival and quality of life [29]. In recent years, exosomes have garnered increasing research interest due to their elucidated roles in tumor invasion, metastasis, immune evasion, and the development of therapy resistance [30–32]. Exosomes serve as key mediators of intercellular communication within the tumor microenvironment, contributing to cancer cell invasion and metastasis [33]. In addition, they can influence the activity of immune cells, thereby enhancing the tumor’s capacity for immune evasion and enabling it to escape immune surveillance [34]. Furthermore, exosome-mediated regulatory functions are implicated in resistance to chemotherapeutic agents [35].
We identified 106 exosome-related differentially expressed genes from the TCGA transcriptome. Using univariate Cox regression, we further selected four with strong prognostic relevance: LCN2, RRAS2, LRG1, and LTF. SHAP-based model interpretation provided additional insight into the relative contribution of each gene, revealing LCN2 as the most influential feature in driving risk prediction. Importantly, elevated expression of all four prognostic exosome-associated genes was consistently linked to negative SHAP values, suggesting an association with adverse clinical outcomes. Survival analysis revealed that patients in the high-risk group had significantly poorer overall survival compared to those in the low-risk group—a result independently validated in an external GEO cohort, underscoring the robustness and generalizability of the proposed risk model. Evidence from a recent systematic review shows that exosome-derived markers are frequently upregulated in breast cancer and linked to adverse clinical features, encompassing chemotherapy resistance, survival, recurrence, and metastasis [36]. Similarly, an independent study focusing on triple-negative breast cancer demonstrated a strong association between exosome-related gene signatures and poor prognosis [37]. Collectively, these findings align with the present results and underscore the importance of exosome-related genes in prognostic stratification across breast cancer subtypes.
In recent years, the integration of bioinformatics analysis with experimental validation has provided crucial support for exploring the molecular mechanisms of breast cancer and identifying potential biomarkers [38–40]. In this study, bioinformatics analysis revealed that LCN2, a prognosis-related differentially expressed gene (DEG) among exosomal transcripts, contributed most significantly to model prediction, and its high expression was closely associated with poor prognosis in patients with Luminal-type breast cancer. Existing literature supports a tumor-promoting role for LCN2 in breast cancer. Studies have shown [41] that LCN2 drives disease progression by promoting angiogenesis and enhancing tumor cell migration and invasion. Proteomic analyses further suggest that LCN2 may influence cell cycle regulation and genomic stability by modulating DNA repair-related proteins [42]. Moreover, LCN2 expression has been linked to PD-L1, indicating a possible role in immune checkpoint regulation [43]. While these findings provide important clues for understanding the involvement of LCN2 in Luminal-type breast cancer, the underlying molecular mechanisms require further validation through in vitro cellular assays and in vivo animal models to establish its potential as a therapeutic target and biomarker.
Analyses of immune infiltration and function demonstrated that high-risk tumors were characterized by a significant accumulation of M2 macrophages and concurrent broad impairment of immune-related functions relative to the low-risk group. Significant correlations were observed between the four prognostic exosomal DEGs and diverse immune cell infiltration, highlighting their potential role in shaping an immunosuppressive microenvironment associated with high-risk disease. Physiologically, macrophages exist predominantly in an M0 state. Upon specific signals from the microenvironment, they undergo functional polarization, shifting toward either the pro-inflammatory M1 or the immunosuppressive M2 phenotype [44]. Macrophage polarization is increasingly recognized as a key driver of breast cancer initiation and progression Tumor cells often exploit this process, promoting the polarization of macrophages toward an immunosuppressive M2 phenotype. This shift fosters a tumor-permissive microenvironment and contributes to disease advancement [45]. Emerging evidence indicates that cancer-derived exosomes modulate macrophage behavior. One mechanism involves the transfer of exosomal miR-138-5p to tumor-associated macrophages, which suppresses M1 polarization and enhances M2 polarization, in turn promoting breast cancer development [46]. Collectively, these findings suggest that alterations in exosome-associated gene expression and macrophage polarization status may provide valuable insight into tumor immune regulation and has practical utility for prognostic stratification in breast cancer.
Gene set enrichment analysis suggested a phenotype characterized by heightened proliferation and DNA repair in high-risk tumors. This molecular profile correlates with clinical aggressiveness and adverse prognosis, potentially underlying treatment resistance and tumor recurrence. Consistent with these observations, integrative analyses of large-scale gene expression datasets have demonstrated that pathways related to DNA repair and cell cycle regulation play central roles in breast cancer progression [47]. Through the construction of competing endogenous RNA and transcription factor regulatory networks, the present study further delineated the intricate transcriptional and post-transcriptional regulatory landscape governing the core prognostic genes. Dissecting these regulatory interactions to identify key transcription factors and non-coding RNAs may provide mechanistic insight into their involvement in tumor proliferation, apoptotic resistance, and immune escape [48]. Exosomes contribute to breast cancer pathogenesis from early malignant transformation to mediating intercellular crosstalk within the tumor microenvironment, ultimately driving immune evasion and enhancing invasive and metastatic potential [49].
Drug sensitivity analysis revealed that high-risk tumors exhibited heightened sensitivity to PI3K inhibitors and several chemotherapeutic agents compared to low-risk tumors. However, no significant differences in the predicted efficacy of CDK4/6 inhibitors or the endocrine agent tamoxifen were observed between the two risk strata. Such differential drug sensitivity is likely attributable to distinct molecular characteristics and signaling pathway activities underlying each risk group. The greater response of high-risk tumors to PI3K inhibitors may be attributable to enhanced PI3K/AKT/mTOR signaling, a pathway which governs key tumorigenic processes including proliferation, survival, and metabolic reprogramming [50]. Chemotherapeutic agents exert their antitumor effects primarily by inducing apoptosis or suppressing cell proliferation. Accordingly, the increased proliferative capacity and DNA repair activity observed in high-risk tumors may partly explain their relatively greater responsiveness to chemotherapy. CDK4/6 inhibitor efficacy depends critically on RB pathway function. The absence of significant intergroup differences in sensitivity could be attributed to preserved RB pathway activity in both risk groups [51]. Tamoxifen, a selective estrogen receptor modulator widely used in luminal breast cancer, likewise showed no marked efficacy difference between risk groups, potentially due to relatively stable estrogen receptor expression levels in both groups [52]. The drug sensitivity analysis in this study was based on pharmacogenomic data derived from cancer cell lines in the GDSC database. Such in-silico predictions do not fully account for in vivo factors including drug metabolism, stromal interactions, and immune microenvironment influences, and therefore should be interpreted cautiously. Nevertheless, the increased predicted sensitivity to PI3K inhibitors in the high-risk group is biologically plausible. Aberrant activation of the PI3K/AKT pathway is common in hormone receptor–positive breast cancer, frequently driven by PIK3CA mutations or PTEN loss. Clinical studies have demonstrated that PI3K pathway alterations can confer sensitivity to PI3K inhibitors such as alpelisib in ER-positive disease. It is therefore possible that the transcriptional profile captured by our risk model reflects enhanced PI3K pathway activity, which may partly explain the observed difference in predicted drug response. Further validation in patient-derived models and prospective clinical cohorts will be necessary to determine whether the proposed risk signature can reliably guide therapeutic decision-making.
Compared with previous studies that were often constrained by single datasets or limited analytical strategies, the present study adopted a more comprehensive and integrative analytical framework. By leveraging large-scale multi-omics data, we identified key exosome-related genes and established a prognostic model, followed by systematic downstream analyses encompassing transcriptomic profiling, somatic mutation patterns, ceRNA interactions, transcription factor regulation, and drug sensitivity prediction. Together, this multidimensional framework, combined with its large sample size and external validation, enhances the robustness and generalizability of the findings, offering a more comprehensive view of the biological and clinical relevance of exosome-related signatures in luminal breast cancer.
It should be noted that this study has several limitations. First, all analyses were based on publicly available datasets, and experimental validation was not performed, which may limit the biological confirmation of the findings. Additionally, the absence of dimensionality reduction techniques such as LASSO regression, as well as the lack of integration with clinical models, compromises the robustness of the predictive model. In addition, tumor mutational burden analysis did not reveal a significant association with risk stratification, nor were survival differences observed between high- and low-TMB groups. This result may be partly explained by the exclusive focus on luminal breast cancer, as the prognostic and predictive relevance of TMB is known to vary across molecular subtypes, potentially reducing its discriminative value in this specific setting. Furthermore, although in silico drug sensitivity analyses suggested that high-risk tumors may exhibit increased responsiveness to certain therapeutic agents, these predictions have not been validated in clinical trials, and their translational relevance therefore requires cautious interpretation. Future studies incorporating larger and more diverse patient cohorts, additional breast cancer subtypes, employing more advanced machine learning approaches, and prospective clinical data will be essential to validate these observations, strengthen the robustness of the conclusions, and facilitate clinical application.
Conclusion
Through analysis of RNA-seq data from Luminal breast cancer patients in the TCGA and GEO databases, this study identified four exosome-related prognostic differentially expressed genes (LCN2, RRAS2, LRG1, and LTF). A risk-stratification model built on these genes effectively distinguished patients into high- and low-risk groups, which exhibited marked differences in immune microenvironment composition and key cellular pathway activities. Drug-response profiling further suggested that high-risk patients may display increased sensitivity to PI3K inhibitors and selected chemotherapeutic agents. Together, these results provide mechanistic insights and highlight potential therapeutic targets to advance precision management of Luminal breast cancer.
Acknowledgements
Not applicable.
Author contributions
(I) Conception: J Huang and X Tang; (II) Administrative support: Z Luo; (III) Collection and analysis of data: Q Su, D Fang and J Wang; (IV) Data visualization: Q Su; (V) Manuscript final approval: All authors.
Funding
This research was supported by the Science and Technology Plan Projects in Baise (BK20243434).
Data availability
All data used in this study are publicly available. The TCGA-BRCA dataset was obtained from The Cancer Genome Atlas (TCGA) database via the Genomic Data Commons (GDC) portal at https://portal.gdc.cancer.gov/projects/TCGA-BRCA. The Gene Expression Omnibus (GEO) datasets used for external validation include GSE17705 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE17705), GSE21653 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE21653), GSE25066 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE25066), and GSE69031 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE69031). All datasets are publicly accessible without restriction.
Declarations
Ethics approval and consent to participate
This study was conducted entirely based on publicly available datasets (TCGA and GEO). All data were obtained from open-access databases, and no human participants or animals were directly involved. Therefore, ethical approval was not required. Not applicable. The study used publicly available anonymized data.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Jian Huang, Xiuling Tang and Qiyuan Su have contributed equally to this work.
References
- 1.Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–63. [DOI] [PubMed] [Google Scholar]
- 2.Xia C, Dong X, Li H, Cao M, Sun D, He S, et al. Cancer statistics in China and United States, 2022: profiles, trends, and determinants. Chin Med J (Engl). 2022;135(5):584–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Perou CM, Sørlie T, Eisen MB, van de Rijn M, Jeffrey SS, Rees CA, et al. Molecular portraits of human breast tumours. Nature. 2000;406(6797):747–52. [DOI] [PubMed] [Google Scholar]
- 4.Wang H, Liu XY, Jiang YZ, Shao ZM. [Challenges and countermeasures in the treatment of luminal breast cancer]. Zhonghua Zhong Liu Za Zhi. 2020;42(3):192–6. [DOI] [PubMed] [Google Scholar]
- 5.Sarhangi N, Hajjari S, Heydari SF, Ganjizadeh M, Rouhollah F, Hasanzad M. Breast cancer in the era of precision medicine. Mol Biol Rep. 2022;49(10):10023–37. [DOI] [PubMed] [Google Scholar]
- 6.Xiong X, Zheng LW, Ding Y, Chen YF, Cai YW, Wang LP, et al. Breast cancer: pathogenesis and treatments. Signal Transduct Target Ther. 2025;10(1):49. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.De Rubis G, Rajeev Krishnan S, Bebawy M. Liquid Biopsies in Cancer Diagnosis, Monitoring, and Prognosis. Trends Pharmacol Sci. 2019;40(3):172–86. [DOI] [PubMed] [Google Scholar]
- 8.Fang D, Li Y, Li Y, Chen Y, Huang Q, Luo Z, et al. Identification of immune-related biomarkers for predicting neoadjuvant chemotherapy sensitivity in HER2 negative breast cancer via bioinformatics analysis. Gland Surg. 2022;11(6):1026–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Reinbolt RE, Mangini N, Hill JL, Levine LB, Dempsey JL, Singaravelu J, et al. Endocrine therapy in breast cancer: the neoadjuvant, adjuvant, and metastatic approach. Semin Oncol Nurs. 2015;31(2):146–55. [DOI] [PubMed] [Google Scholar]
- 10.Shao H, Im H, Castro CM, Breakefield X, Weissleder R, Lee H. New Technologies for Analysis of Extracellular Vesicles. Chem Rev. 2018;118(4):1917–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Mortezaee K. Exosomes in bridging macrophage-fibroblast polarity and cancer stemness. Med Oncol. 2025;42(6):216. [DOI] [PubMed] [Google Scholar]
- 12.Hannafon BN, Trigoso YD, Calloway CL, Zhao YD, Lum DH, Welm AL, et al. Plasma exosome microRNAs are indicative of breast cancer. Breast Cancer Res. 2016;18(1):90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Veneziano M, Savini I, Cortellesi E, Gasperi V, Gambacurta A, Catani MV. Bioinformatics Strategies in Breast Cancer Research. Biomolecules. 2025;15(10). [DOI] [PMC free article] [PubMed]
- 14.Jing Y, Wang Y, Li Y, Huang X, Wang J, Yelihamu D, et al. Diagnostics and immunological function of CENPN in human tumors: from pan-cancer analysis to validation in breast cancer. Transl Cancer Res. 2025;14(2):881–906. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Tomczak K, Czerwińska P, Wiznerowicz M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp Oncol (Pozn). 2015;19(1a):A68–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Clough E, Barrett T, Wilhite SE, Ledoux P, Evangelista C, Kim IF, et al. NCBI GEO: archive for gene expression and epigenomics data sets: 23-year update. Nucleic Acids Res. 2024;52(D1):D138–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Symmans WF, Hatzis C, Sotiriou C, Andre F, Peintinger F, Regitnig P, et al. Genomic index of sensitivity to endocrine therapy for breast cancer. J Clin Oncol. 2010;28(27):4111–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Sabatier R, Finetti P, Cervera N, Lambaudie E, Esterni B, Mamessier E, et al. A gene expression signature identifies two prognostic subgroups of basal breast cancer. Breast Cancer Res Treat. 2011;126(2):407–20. [DOI] [PubMed] [Google Scholar]
- 19.Sabatier R, Finetti P, Adelaide J, Guille A, Borg JP, Chaffanet M, et al. Down-regulation of ECRG4, a candidate tumor suppressor gene, in human breast cancer. PLoS ONE. 2011;6(11):e27656. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hatzis C, Pusztai L, Valero V, Booser DJ, Esserman L, Lluch A, et al. A genomic predictor of response and survival following taxane-anthracycline chemotherapy for invasive breast cancer. JAMA. 2011;305(18):1873–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Itoh M, Iwamoto T, Matsuoka J, Nogami T, Motoki T, Shien T, et al. Estrogen receptor (ER) mRNA expression and molecular subtype distribution in ER-negative/progesterone receptor-positive breast cancers. Breast Cancer Res Treat. 2014;143(2):403–9. [DOI] [PubMed] [Google Scholar]
- 22.Baldasici O, Balacescu L, Cruceriu D, Roman A, Lisencu C, Fetica B et al. Circulating Small EVs miRNAs as Predictors of Pathological Response to Neo-Adjuvant Therapy in Breast Cancer Patients. Int J Mol Sci. 2022;23(20). [DOI] [PMC free article] [PubMed]
- 23.Chin K, DeVries S, Fridlyand J, Spellman PT, Roydasgupta R, Kuo WL, et al. Genomic and transcriptional aberrations linked to breast cancer pathophysiologies. Cancer Cell. 2006;10(6):529–41. [DOI] [PubMed] [Google Scholar]
- 24.Stelzer G, Rosen N, Plaschkes I, Zimmerman S, Twik M, Fishilevich S et al. The GeneCards Suite: From Gene Data Mining to Disease Genome Sequence Analyses. Curr Protoc Bioinformatics. 2016;54:1.30.1-1.3. [DOI] [PubMed]
- 25.Ponce-Bobadilla AV, Schmitt V, Maier CS, Mensing S, Stodtmann S. Practical guide to SHAP analysis: Explaining supervised machine learning model predictions in drug development. Clin Transl Sci. 2024;17(11):e70056. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Newman AM, Liu CL, Green MR, Gentles AJ, Feng W, Xu Y, et al. Robust enumeration of cell subsets from tissue expression profiles. Nat Methods. 2015;12(5):453–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Han H, Shim H, Shin D, Shim JE, Ko Y, Shin J, et al. TRRUST: a reference database of human transcriptional regulatory interactions. Sci Rep. 2015;5:11432. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Iorio F, Knijnenburg TA, Vis DJ, Bignell GR, Menden MP, Schubert M, et al. A Landscape of Pharmacogenomic Interactions in Cancer. Cell. 2016;166(3):740–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Lutz C, Messal HA, Vareslija D, Prekovic S. The complex landscape of luminal breast cancer. Endocr Relat Cancer. 2025;32(1). [DOI] [PubMed]
- 30.Li I, Nabet BY. Exosomes in the tumor microenvironment as mediators of cancer therapy resistance. Mol Cancer. 2019;18(1):32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Li C, Teixeira AF, Zhu HJ, Ten Dijke P. Cancer associated-fibroblast-derived exosomes in cancer progression. Mol Cancer. 2021;20(1):154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Krylova SV, Feng D. The Machinery of Exosomes: Biogenesis, Release, and Uptake. Int J Mol Sci. 2023;24(2). [DOI] [PMC free article] [PubMed]
- 33.Liu N, Wu T, Han G, Chen M. Exosome-mediated ferroptosis in the tumor microenvironment: from molecular mechanisms to clinical application. Cell Death Discov. 2025;11(1):221. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Wang X, Luo G, Zhang K, Cao J, Huang C, Jiang T, et al. Hypoxic Tumor-Derived Exosomal miR-301a Mediates M2 Macrophage Polarization via PTEN/PI3Kγ to Promote Pancreatic Cancer Metastasis. Cancer Res. 2018;78(16):4586–98. [DOI] [PubMed] [Google Scholar]
- 35.Qu X, Liu B, Wang L, Liu L, Zhao W, Liu C, et al. Loss of cancer-associated fibroblast-derived exosomal DACT3-AS1 promotes malignant transformation and ferroptosis-mediated oxaliplatin resistance in gastric cancer. Drug Resist Updat. 2023;68:100936. [DOI] [PubMed] [Google Scholar]
- 36.Wang M, Ji S, Shao G, Zhang J, Zhao K, Wang Z, et al. Effect of exosome biomarkers for diagnosis and prognosis of breast cancer patients. Clin Transl Oncol. 2018;20(7):906–11. [DOI] [PubMed] [Google Scholar]
- 37.Yang Q, Cai X, Qian Y, Yang T, Xia W, Zhong M, et al. Implication of an exosome-based gene signature for estimating clinical outcomes of triple-negative breast cancer and assisting in individualized therapy. Sci Rep. 2025;15(1):37774. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Wang J, Wang Y, Ma H, Li Y, Hou J, Li J, et al. CLIC6’s role in cancer: from broad analysis to breast cancer validation. Front Oncol. 2025;15:1667589. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Elihamu D, Li Y, Wang Y, Cui H, Xing Y, Peng H, et al. CORO1A: a pan-cancer prognosis, diagnostic and immune biomarker based on breast cancer validation. Front Oncol. 2025;15:1670526. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Aimaiti X, Wang Y, Ismtula D, Li Y, Ma H, Wang J, et al. Bystin is a Prognosis and Immune Biomarker: From Pan-Cancer Analysis to Validation in Breast Cancer. Breast Cancer (Dove Med Press). 2025;17:755–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Keller DT, Weiskirchen R, Schröder-Lange SK. Lipocalin-2 in Triple-Negative Breast Cancer: A Review of Its Pathophysiological Role in the Metastatic Cascade. Int J Mol Sci. 2025;26(22). [DOI] [PMC free article] [PubMed]
- 42.Villodre ES, Hu X, Larson R, Finetti P, Gomez K, Balema W, et al. Lipocalin 2 promotes inflammatory breast cancer tumorigenesis and skin invasion. Mol Oncol. 2021;15(10):2752–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Ekemen S, Bilir E, Soultan HEA, Zafar S, Demir F, Tabandeh B, et al. The Programmed Cell Death Ligand 1 and Lipocalin 2 Expressions in Primary Breast Cancer and Their Associations with Molecular Subtypes and Prognostic Factors. Breast Cancer (Dove Med Press). 2024;16:1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Huang G, Yu Y, Su H, Gan H, Chu L. Integrating RNA-seq and scRNA-seq to explore the prognostic features and immune landscape of exosome-related genes in breast cancer metastasis. Ann Med. 2025;57(1):2447917. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Ma C, He D, Tian P, Wang Y, He Y, Wu Q et al. miR-182 targeting reprograms tumor-associated macrophages and limits breast cancer progression. Proc Natl Acad Sci U S A. 2022;119(6). [DOI] [PMC free article] [PubMed]
- 46.Xun J, Du L, Gao R, Shen L, Wang D, Kang L, et al. Cancer-derived exosomal miR-138-5p modulates polarization of tumor-associated macrophages through inhibition of KDM6B. Theranostics. 2021;11(14):6847–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Li WX, He K, Tang L, Dai SX, Li GH, Lv WW, et al. Comprehensive tissue-specific gene set enrichment analysis and transcription factor analysis of breast cancer by integrating 14 gene expression datasets. Oncotarget. 2017;8(4):6775–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Yin J, Lin C, Jiang M, Tang X, Xie D, Chen J, et al. ISG20L2, LSM4, MRPL3 are four novel hub genes and may serve as diagnostic and prognostic markers in breast cancer. Sci Rep. 2021;11(1):15610. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Jia Y, Chen Y, Wang Q, Jayasinghe U, Luo X, Wei Q, et al. Exosome: emerging biomarker in breast cancer. Oncotarget. 2017;8(25):41717–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Nixon MJ, Formisano L, Mayer IA, Estrada MV, González-Ericsson PI, Isakoff SJ, et al. PIK3CA and MAP3K1 alterations imply luminal A status and are associated with clinical benefit from pan-PI3K inhibitor buparlisib and letrozole in ER+ metastatic breast cancer. NPJ Breast Cancer. 2019;5:31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Shanabag A, Armand J, Son E, Yang HW. Targeting CDK4/6 in breast cancer. Exp Mol Med. 2025;57(2):312–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Tang J, Cui Q, Zhang D, Liao X, Zhu J, Wu G. An estrogen receptor (ER)-related signature in predicting prognosis of ER-positive breast cancer following endocrine treatment. J Cell Mol Med. 2019;23(8):4980–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All data used in this study are publicly available. The TCGA-BRCA dataset was obtained from The Cancer Genome Atlas (TCGA) database via the Genomic Data Commons (GDC) portal at https://portal.gdc.cancer.gov/projects/TCGA-BRCA. The Gene Expression Omnibus (GEO) datasets used for external validation include GSE17705 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE17705), GSE21653 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE21653), GSE25066 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE25066), and GSE69031 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE69031). All datasets are publicly accessible without restriction.









