Abstract
Background
Breast cancer metastasis remains a major clinical challenge due to its complex molecular mechanisms, highlighting the need to identify key regulatory factors.
Methods
Multiple breast cancer metastasis datasets were integrated to comprehensively identify metastasis-related genes. Differential expression analysis, weighted gene co-expression network analysis (WGCNA), and five machine learning algorithms were employed to screen for key candidate genes. Immune cell infiltration was assessed using the CIBERSORT algorithm, and miRNA expression profiles associated with lymph node metastasis were analyzed to construct potential regulatory networks. The biological functions of target genes were further verified through in vitro functional experiments.
Results
PAICS was identified as a metastasis-associated gene and was associated with patient survival in TCGA-BRCA. Its expression correlated with the infiltration of multiple immune cell subsets. Four metastasis-related miRNAs were involved in potential regulatory networks. Functional assays demonstrated that PAICS knockdown suppressed migration and invasion while increasing E-cadherin and decreasing N-cadherin and Vimentin expression.
Conclusions
PAICS acts as a biomarker that promotes breast cancer metastasis by modulating epithelial–mesenchymal transition and the immune microenvironment, offering potential for clinical translation.
Supplementary Information
The online version contains supplementary material available at 10.1007/s12672-026-04968-4.
Keywords: Breast cancer metastasis, PAICS, Machine learning, Immune infiltration analysis, MiRNA regulatory network
Introduction
Across the globe, breast cancer is the most widespread malignant tumor, while also functioning as a key cause of cancer-related fatalities [1]. A recent global analysis estimated that approximately 2.3 million new cases were diagnosed in 2022 across 185 countries, representing 11.7% of all female cancers, and causing around 6.9% of global cancer deaths [2]. In China, recent cancer statistics estimates indicate that there were approximately 369,769 new cases of female breast cancer in 2024, making breast cancer the most common cancer among Chinese women [3]. Breast cancer, owing to its high incidence and mortality, remains a major threat to women’s health. Combined treatments, including surgery, radiotherapy, and chemotherapy, have markedly improved the prognosis of patients with early-stage disease. However, the 5-year survival rate of patients with metastatic breast cancer remains below 30% [4, 5]. Therefore, revealing the driving mechanisms of breast cancer metastasis and developing accurate prognostic tools are essential for improving the early diagnosis rate and treatment effect of metastatic breast cancer.
Breast cancer metastasis is a complex malignant process in which primary tumor cells disseminate and establish secondary tumors through a series of sequential biological events. These include local invasion, intravasation, survival in the circulatory system, colonization of distant organs such as the bone, liver, lung, brain, and lymph nodes, and subsequent adaptation to the local microenvironment [5]. Bone, liver, lung, brain, and lymph nodes are among the most common locations for metastasis [6–10]. Current investigations into breast cancer are primarily directed toward the epithelial-mesenchymal transition (EMT) [11], and they also focus on advances in chemotherapy, targeted therapy, and the development of therapeutic drugs, such as trastuzumab, tamoxifen, bevacizumab [12–14]. Recent studies have shown that inhibiting the early steps of EMT in breast cancer can prevent breast cancer cell metastasis by suppressing key signaling pathways such as TGF-β/SMAD3 and PI3K/AKT, thereby markedly reducing the migratory and invasive capabilities of cancer cells and blocking tumor dissemination [15, 16]. Although these treatment strategies have achieved substantial progress in breast cancer metastasis, their clinical application still faces multiple challenges, as drug resistance and adverse effects limit long-term efficacy, and tumor heterogeneity along with the complexity of the microenvironment further contribute to insufficient response rates to single-agent therapies [17].
Therefore, identifying biomarkers associated with breast cancer metastasis is of great clinical significance for early diagnosis, prognostic assessment, and the development of targeted therapies. This study aimed to systematically identify key genes involved in breast cancer metastasis by integrating bioinformatics and machine learning approaches. Core candidate genes were screened using weighted gene co-expression network analysis (WGCNA) and multiple machine learning algorithms, followed by assessment of their relationships with immune cell infiltration. To elucidate potential upstream regulatory mechanisms, miRNA expression profiles associated with lymph node metastasis were analyzed. The biological roles of the identified genes were further validated in vitro by evaluating their effects on cell migration, invasion, and the expression of EMT markers. This multi-level analytical framework provides new insights into the molecular mechanisms driving breast cancer metastasis and establishes a theoretical foundation for improving prognostic assessment and developing targeted therapeutic strategies.
Materials and methods
Data sources
This study utilized multiple publicly available datasets. The GSE191230 dataset from the gene expression omnibus (GEO) included 20 samples, comprising 13 primary breast tumors and 7 distant metastasis tumors. In addition, the cancer genome atlas breast invasive carcinoma (TCGA-BRCA) dataset, encompassing 1,035 breast cancer samples with complete clinical follow-up data, was included. To identify metastasis-related genes, the GeneCards database was searched with the exact keyword “breast cancer metastasis”, and protein-coding genes with a relevance score > 15 were selected for subsequent analyses.
Data preprocessing
The GSE191230 dataset counts were analyzed using the R package DESeq2 (v1.50.2). Normalized counts were extracted using DESeq2’s counts function and transformed by Log2(count + 1) for downstream analyses. For the TCGA-BRCA data were log2-transformed using log2(TPM + 1) to approximate a normal distribution, followed by batch effect correction using the ComBat function from the sva package. This preprocessing workflow generated normalized gene expression matrices suitable for downstream statistical and integrative analyses.
Differential gene expression analysis
Differentially expressed genes (DEGs) associated with breast cancer metastasis were identified in the GSE191230 dataset by comparing metastatic samples to non-metastatic controls. Expression differences were assessed using an independent sample t-test, with variance homogeneity confirmed via Levene’s test. To improve stability for high-dimensional data, empirical Bayes moderation from the limma package (R version 4.3.1) was applied. P-values were adjusted for multiple testing using the Benjamini-Hochberg method to control the false discovery rate (FDR). DEGs were defined as genes with log2|Fold Change| (log2|FC|) > 0.75 and FDR < 0.05.
Functional enrichment analysis
Using the R package clusterProfiler (Version 4.10.0), we conducted gene ontology (GO) enrichment analysis for upregulated genes, including biological processes (BP), cellular components (CC), and molecular functions (MF). Kyoto encyclopedia of genes and genomes (KEGG) pathway analysis was performed to identify involved metabolic and signaling pathways. Enrichment significance was determined using the hypergeometric test, and multiple testing was controlled by the Benjamini–Hochberg procedure, with results considered significant at adj.p < 0.05 and FDR < 0.25.
WGCNA
The WGCNA R package was used to build a co-expression network from the most variably expressed genes. Soft thresholding powers from 1 to 20 were tested to identify the optimal value for network construction. The resulting correlation matrix was converted into an adjacency matrix and subsequently into a topological overlap matrix (TOM). Modules were defined by average linkage hierarchical clustering of the TOM, with a minimum size of 100 genes, and highly similar modules were merged. Pearson correlation analysis was then used to evaluate associations between the merged modules and disease status, and the modules showing the strongest positive and negative correlations were designated as key modules.
Candidate gene integration
To narrow down key genes, upregulated DEGs, genes from metastasis-related WGCNA modules, and GeneCards-derived metastasis-related genes were intersected using the VennDiagram R package. The overlapping genes were defined as candidate key genes.
Feature selection and machine learning modeling
To identify metastasis-associated genes, we constructed a binary classification framework distinguishing metastatic and non-metastatic samples based on candidate gene expression profiles. Five machine-learning algorithms were applied, including adaptive boosting (AdaBoost), least absolute shrinkage and selection operator (LASSO) regression, support vector machine–recursive feature elimination (SVM-RFE), random forest, and eXtreme gradient boosting (XGBoost), implemented using the adabag, glmnet, e1071, kernlab, caret, randomForest, and XGBoost packages. To reduce overfitting and information leakage, repeated nested five-fold cross-validation was performed, with feature selection and model tuning conducted exclusively within the training folds. Hyperparameters were optimized using cross-validation: the optimal lambda for LASSO was selected based on minimal cross-validation error; XGBoost parameters were tuned via grid search to maximize AUC; the number of trees in random forest was selected by minimizing cross-validation error; Boruta analysis was conducted with a fixed number of trees to ensure stability of feature importance estimation. Genes consistently prioritized across multiple algorithms were defined as hub genes associated with breast cancer metastasis.
Survival analysis and clinical association analysis
Expression data and clinical information from the TCGA-BRCA dataset were analyzed to explore the potential clinical relevance of the five core genes. Univariate Cox regression and Kaplan–Meier survival analyses were performed in an exploratory manner to assess associations between gene expression and patient outcomes. Kaplan–Meier survival curves were generated to visualize differences between high- and low-expression groups, with cutoffs determined using the surv_cutpoint function from the R survival package (default method). Additionally, the Wilcoxon test was used to examine differential expression of core genes across breast cancer metastasis subtypes—T (primary tumor size and invasion), N (regional lymph node involvement), and M (distant metastasis)—and box plots were generated to visualize these differences.
Immune infiltration and correlation analysis
Immune cell composition was estimated using the CIBERSORT algorithm, which applies linear support vector regression to deconvolute bulk transcriptomic data [18]. Immune infiltration analysis was conducted on the GSE191230 RNA-seq dataset using FPKM-normalized gene expression values as input. The normalized expression matrix was uploaded to the CIBERSORT web portal with 1,000 permutations to estimate the relative proportions of 22 immune cell subsets based on the LM22 signature matrix. All samples were included for exploratory analyses without filtering based on deconvolution p-values. The relative proportions of immune cell types were visualized using stacked bar plots, and differences between groups were illustrated using violin plots generated with the ggplot2 package in R. To investigate the association between PAICS expression and immune infiltration, correlation heatmaps were constructed using the corrplot package. Samples were stratified into high- and low-PAICS expression groups according to the median PAICS expression level, and differences in immune cell proportions were compared between groups. Correlations between PAICS expression and immune checkpoint-related genes were also evaluated. Additionally, immune infiltration scores were assessed using estimation-based approaches to provide a complementary overview of the tumor immune microenvironment. All immune-related analyses were conducted for exploratory and hypothesis-generating purposes.
MiRNA regulatory network construction
The miRNA expression dataset related to lymph node metastasis in breast cancer (GSE148846) was obtained from the GEO database. Differential expression analysis was performed to identify significantly dysregulated miRNAs, with thresholds set at log2|FC| > 2 and P < 0.05. The potential target genes of these miRNAs were predicted using four databases: TargetScan, miRDB, miRWalk, and NPInter. The intersecting genes among these databases were selected, and a miRNA–target gene regulatory network was constructed using Cytoscape (v3.10.2) to elucidate the regulatory mechanisms of miRNAs in breast cancer lymphatic metastasis.
Cell culture
MCF-7 (ATCC HTB-22), MDA-MB-231 (ATCC HTB-26), and MCF-10 A (ATCC CRL-10317) cell lines were obtained from Genetimes ExCell Technology. MCF-7 and MDA-MB-231 cells were maintained in a 1:1 mixture of Ham’s F12 medium and Dulbecco’s Modified Eagle’s Medium (both from Genetimes ExCell Technology) supplemented with 2 mM L-glutamine. Cells were cultured at 37 °C in a humidified incubator with 5% CO₂ and passaged at 80% confluence using 0.25% trypsin–0.02% EDTA (T1300, Salabao, Beijing). All reagents used for the culture of MCF-10 A cells were obtained from Thermo Fisher Scientific.MCF-10 A cells were cultured in DMEM/F12 medium (12634010) supplemented with 5% (v/v) horse serum (16050122), 20 ng/mL EGF (PHG0311L), 10 mg/mL insulin (41400045), 0.5 mg/mL cholera toxin (C34775), 50 µM human keratinocyte growth supplement (S0015), and 1% (v/v) penicillin–streptomycin (15140122). Cells were maintained at 37 °C with 5% CO₂, with medium replaced every 3 days. When confluent, cells were passaged with 0.05% trypsin–0.01% EDTA.
Cell transfection
For establishment of PAICS-knockdown models, lentiviral vectors encoding specific shRNAs (sh-PAICS#1, sh-PAICS#2) and a non-targeting control (sh-NC) were acquired from GenePharma (Suzhou, China). MDA-MB-231 cells were seeded in six-well plates until reaching 50% confluence, followed by infection with the respective lentiviruses for 24 h. Stable cell lines were subsequently selected through continuous culture with 2 µg/ml puromycin (Sigma, USA) over a two-week period. For rescue experiments, transient transfection of relevant plasmids into the knockdown and control cells was carried out employing Lipofectamine 2000 reagent (Invitrogen, Carlsbad, USA) in compliance with the manufacturer’s protocol. The knockdown efficiency was validated by reverse transcription quantitative real-time polymerase chain reaction (RT-qPCR).
RT-qPCR
Total RNA was extracted from cells using TriQuick Reagent (Solarbio Bioscience & Technology, Beijing, R1100). During the extraction process, a chloroform substitute (ECOTOP, Guangzhou, ES-8522-100mL) was used for phase separation, and RNA was precipitated with isopropanol (Baishi Chemical Industry Co., Ltd., Tianjin). The resulting RNA was dissolved in DEPC-treated water for subsequent experiments. First-strand cDNA synthesis was performed using the SureScript™ First-Strand cDNA Synthesis Kit (GeneCopeia, USA, QP056T), and the synthesized cDNA was stored at − 20 °C until use. Quantitative PCR was conducted using 2×SYBR Green qPCR Master Mix (None ROX) (Servicebio, Wuhan, G3320-05) with gene-specific primers. The qPCR amplification was performed on the SLAN-96S RT-qPCR System (Shanghai Hongshi Medical Technology Co., Ltd., Shanghai). Relative mRNA expression levels were normalized to GAPDH using the 2−ΔΔCt method. Primer sequences used were as follows: PAICS forward: 5’-CAGGAGATCGAAGCCAACA-3’; PAICS reverse: 5’-AGCCCATCAACACTACAACC-3’. GAPDH forward: 5’-TCCAAAATCAAGTGGGGCGA-3’; GAPDH reverse: 5’-TGATGACCCTTTTGGCTCCC-3’.
Western blot
Cells were washed with PBS (C3580-0500, Shanghai Xiaopeng Biotechnology) and lysed on a shaker at 4 °C for 30 min using 1× RIPA buffer (P0100, Wuhan Bosch De Biotechnology). Protein lysates were collected by centrifugation, and concentrations were determined using the Bradford assay with breast cancerA standards. Proteins were separated by SDS-PAGE (AR0138, Wuhan Bosch De Biotechnology) using 10% separating gel and 5% stacking gel, then transferred to PVDF membranes (IPVH00010, Millipore, USA) at 350 mA for 2 h at 4 °C. Membranes were blocked with 5% skimmed milk at room temperature for 2 h and incubated overnight at 4 °C with anti-PAICS (1:3000; 12967-1-AP, Wuhan Sanying), anti-E-cadherin (1:20000; 20874-1-AP, Proteintech), anti-Vimentin (1:20000; 10366-1-AP, Proteintech), anti-N-cadherin (1:5000; 22018-1-AP, Proteintech) primary antibody. After three washes with TBST, membranes were incubated with HRP-conjugated secondary antibody (ab7090, Abcam, UK) at 1:8000 dilution for 1.5 h at room temperature. Signals were detected using Western Lightning™ Chemiluminescence Reagent (NEL10300EA, PerkinElmer, USA) and captured with a Tanon 4800 imaging system. Relative protein expression was quantified using ImageJ (version 1.8.0), with GAPDH serving as the internal reference.
Transwell assay
Using 0.02% EDTA and 0.25% trypsin, cells were digested; the reaction was neutralized with culture medium, and the cells were subsequently counted. Matrigel was diluted and coated (50 µL per chamber) over chambers pre-coated with 10 µL 0.5 mg/mL fibronectin and air-dried under sterile conditions. Approximately 1 × 105 cells suspended in 200 µL serum-free medium were added to the upper chamber, while the lower chamber contained medium supplemented with 20% FBS to serve as a chemoattractant. Following 24 h of incubation at 37 °C, non-invading cells were removed. Invading cells on the lower membrane surface were fixed with a methanol-acetic acid mixture (3:1) for 30 min, stained with crystal violet for 15 min, washed, mounted, and imaged under a microscope at three random fields.
Statistical analysis
All statistical analyses and graphing were performed using R (version 4.2.2) and GraphPad Prism 9.0. Values foreach sample were from at least 3 biologically independent experiments with at least 3 technical replicates. All data in this study are displayed as the mean + standard deviation (SD). For comparisons between two groups, independent samples t-tests were used when data were normally distributed, and Mann–Whitney U tests were applied otherwise. For comparisons among three or more groups, one-way ANOVA followed by Tukey’s post-hoc test was used for normally distributed data, while the Kruskal–Wallis test followed by Dunn’s post-hoc test was applied for non-normal data. P < 0.05 considered statistically significant. *P < 0.05, **P < 0.01, ***P < 0.001.
Results
Screening and functional enrichment analysis of DEGs
Using the breast cancer metastasis-related dataset GSE191230 from the GEO database, a total of 919 upregulated genes were identified based on the thresholds of |log₂FC| > 0.75 and FDR < 0.05 (Fig. 1A). Heatmap visualization of the top 100 differentially expressed genes revealed distinct transcriptional profiles between metastatic and non-metastatic breast cancer samples (Fig. 1B).
Fig. 1.
Identification and functional annotation of differentially expressed genes in breast cancer metastasis. A Volcano plot of differentially expressed genes. Red dots represent upregulated genes; green dots represent downregulated genes. B Heatmap of the top 100 differentially expressed genes in primary and metastatic breast tumors. C Bar chart of GO enrichment analysis showing the ten most significantly enriched terms in biological process, cellular component, and molecular function categories. D Bar chart of KEGG pathway enrichment analysis showing the twenty most significantly enriched pathways
GO enrichment analysis indicated that these upregulated genes were predominantly involved in BP related to immune cell adhesion, activation, and differentiation. In terms of CC, they were mainly localized to the outer plasma membrane, extracellular matrix, and membrane rafts, reflecting their involvement in cell–cell interaction and signal transduction. MF analysis showed enrichment in integrin binding, cytokine binding, and receptor activity (Fig. 1C). KEGG pathway enrichment analysis further demonstrated that these genes participated in several major biological pathways. Immune-related pathways included cell adhesion molecules, cytokine–cytokine receptor interaction, T-helper cell differentiation, as well as NF-κB and TNF signaling. In addition, viral infection pathways such as human T-cell leukemia virus, human papillomavirus, and Epstein–Barr virus infection were enriched. Pathways associated with cell survival and adhesion, including the PI3K–Akt signaling pathway and focal adhesion, were also significantly represented (Fig. 1D).
WGCNA identified key modules and candidate genes
A weighted gene co-expression network was constructed based on the DEGs. Outlier samples were first removed to ensure data reliability (Fig. 2A). To determine the optimal soft threshold, the scale-free topology was evaluated under a range of power values. A soft threshold power of β = 6 (corresponding to a scale-free topology fit index of R² = 0.9) was selected to construct a network that met the scale-free distribution criterion (Fig. 2B). Using dynamic tree cutting, genes were hierarchically clustered into 12 distinct co-expression modules (Fig. 2C). Correlation analysis between module eigengenes and clinical traits revealed that the green and purple modules were most strongly associated with distant metastatic breast cancer (cor = 0.51, P = 0.03), containing 436 and 283 genes, respectively (Fig. 2D). To identify metastasis-related hub genes, we intersected genes from these two modules with DEGs and breast cancer metastasis-associated genes from the GeneCards database. This integrative analysis yielded 28 overlapping candidate genes for subsequent functional validation (Fig. 2E).
Fig. 2.
WGCNA of differentially expressed genes in metastatic breast cancer. A Sample clustering tree after removal of outliers. B Selection of the soft threshold power. The scale-free topology fit index (left) and mean connectivity (right) were plotted against various soft threshold powers. A power of β = 6 (red dashed line) was chosen to satisfy the scale-free topology criterion (R² = 0.9). C Clustering dendrogram showing 12 gene co-expression modules, each represented by a distinct color. D Heatmap of module–trait relationships. Each cell displays the correlation coefficient and p-value; color intensity reflects the strength and direction of the association (red for positive, blue for negative). E Venn diagram showing the intersection of DEGs, GeneCards breast cancer metastasis-related genes, and genes from the green and purple modules
Screening of core cenes based on machine learning
LASSO regression, SVM-RFE, random forest, XGBoost and AdaBoost were used to screen the key genes related to breast cancer metastasis. LASSO regression identified six potential genes (Fig. 3A, B), while SVM-RFE analysis yielded eleven candidate genes (Fig. 3C). The random forest algorithm ranked gene importance based on the MeanDecreaseGini index, retaining eight genes with values greater than 0.5 (Fig. 3D, E). XGBoost further identified six hub genes (Fig. 3F), and AdaBoost analysis highlighted thirteen genes (Fig. 3G). By intersecting the results from all five algorithms, five core genes—CCNB1, E2F1, NACC1, PAICS, and FEN1—were consistently identified as key regulators associated with breast cancer metastasis (Fig. 3H).
Fig. 3.
Identification of core genes associated with breast cancer metastasis using five machine learning algorithms. A Regression coefficient path plot and (B) cross-validation curves of the LASSO logistic regression model. C Gene selection using the SVM-RFE algorithm. (D) Relationship between the number of decision trees and error rate, and (E) gene importance ranking based on the Gini coefficient in the random forest classifier. (F) Feature importance ranking in the XGBoost model. (G) Feature importance ranking in the AdaBoost model. (H) Venn diagram showing the intersection of genes identified by the five algorithms
Clinical relevance of core genes
To explore the potential clinical relevance of the five core genes (CCNB1, E2F1, NACC1, PAICS, and FEN1), univariate Cox regression analysis was performed using TCGA data. NACC1 and PAICS were identified as genes associated with patient survival. High NACC1 expression was associated with reduced risk (HR = 0.61, p = 0.003), whereas elevated PAICS expression was associated with poorer survival (HR = 1.453, p = 0.023) (Fig. 4A). An exploratory five-gene risk score was constructed and analyzed in relation to clinical variables, including age and tumor stage. Univariate Cox regression showed that riskScore, age, tumor stage, and N and M stages were associated with survival outcomes (Fig. 4B). Kaplan–Meier survival analyses further indicated that high NACC1 expression was associated with longer overall survival (p = 0.000655) (Fig. 4C), whereas high PAICS expression was associated with shorter survival (p = 0.00323) (Fig. 4D).
Fig. 4.
Survival association analysis of core genes. A Forest plot showing univariate Cox regression analysis of the five core genes. B Forest plots of univariate Cox regression analyses of riskScore and clinical variables. C, D Kaplan–Meier survival curves based on (C) NACC1 and (D) PAICS expression
Analysis of immune infiltration in metastatic breast cancer and immune-related study of PAICS
Using the CIBERSORT algorithm, we analyzed immune cell infiltration patterns in breast cancer samples from the GSE191 dataset. The results revealed heterogeneous distributions of 22 immune cell subsets within the tumor microenvironment (TME) (Fig. 5A). Correlation analysis further indicated significant interactions among these immune cell populations (Fig. 5B). Comparing primary and metastatic breast tumors, we observed that resting mast cells were significantly more abundant in metastatic tumors (P = 0.031), whereas eosinophils were enriched in primary tumors (P = 0.011) (Fig. 5C). Furthermore, correlation analysis between PAICS expression and immune cells demonstrated significant positive associations with macrophage M0 (cor = 0.51, P = 0.022) and resting NK cells (cor = 0.50, P = 0.025), and a negative correlation with memory B cells (cor = −0.50, P = 0.025) (Figs. 5D, E).
Fig. 5.
Immune microenvironment analysis in breast cancer and correlation with PAICS expression. A Proportions of 22 immune cell subsets. B Heatmap showing correlations among immune cell infiltration levels. C Differences in immune cell infiltration between primary and metastatic tumors. D Bubble plot of correlations between PAICS and immune cell subsets. E Scatter plots illustrating correlations between PAICS and 3 immune cells
Construction of miRNA regulatory network and functional enrichment of target genes in breast cancer lymph node metastasis
Using the GSE148846 dataset from the GEO database, which profiles miRNA expression in breast cancer lymph node metastasis, differential expression analysis identified four significantly upregulated miRNAs: miR-142-3p, miR-142-5p, miR-150-5p, and miR-203a-3p. Target prediction and intersection analysis revealed 113 common target genes, which were subsequently used to construct a regulatory network connecting these four miRNAs with their targets (Fig. 6A).
Fig. 6.
MiRNA-gene regulatory network construction and functional enrichment analysis. A Schematic diagram of miRNA-Gene regulatory network. B GO enrichment analysis of target genes. C KEGG pathway enrichment analysis
GO enrichment analysis indicated that these target genes were primarily involved in BP such as regulation of vascular development, mRNA localization, positive regulation of endothelial cell migration, angiogenesis, and stress fiber assembly. About CC, enrichment was observed in cytoplasmic ribonucleoprotein granules, stress granules, and cytosolic solute regions. MF analysis highlighted significant involvement in SMAD binding, miRNA binding, and GTPase regulatory and activation activities (Fig. 6B). KEGG pathway analysis demonstrated that the target genes were significantly enriched in diverse pathways, including cancer-related pathways (renal cell carcinoma, colorectal cancer), major signaling pathways (Wnt, TGF-β, FoxO, Hippo), viral infection processes (HIV-1 life cycle), metabolic pathways (glucose signaling, AGE-RAGE signaling in diabetic complications), and pathways associated with cellular structure and invasion, such as regulation of the actin cytoskeleton, focal adhesion, and bacterial invasion of epithelial cells (Fig. 6C).
Knockdown of PAICS inhibits breast cancer cell migration and modulates EMT
PAICS expression in breast cancer was assessed by measuring mRNA and protein levels in normal breast epithelial cells (MCF-10 A) and breast cancer cell lines (MDA-MB-231, MCF-7). PAICS mRNA was significantly elevated in MDA-MB-231 cells compared with MCF-10 A (P < 0.01) (Fig. 7A). Consistently, Western blot analysis revealed higher PAICS protein expression in both MCF-7 (P < 0.01) and MDA-MB-231 (P < 0.001), with the highest levels observed in MDA-MB-231 cells (Figs. 7B, C). Given its elevated expression in highly invasive breast cancer cells, we hypothesized that PAICS contributes to tumor cell motility and invasiveness. To test this, PAICS was knocked down in MDA-MB-231 cells, and cell migration and invasion were assessed using wound-healing and Transwell assays, respectively. Compared with sh-NC controls, wound closure was reduced by 55.52% and 52.99% in sh-PAICS #1 and sh-PAICS #2 cells, respectively (P < 0.001) (Figs. 7D, E), indicating impaired migratory capacity. In addition, Transwell assays demonstrated that PAICS knockdown significantly suppressed cell invasion (P < 0.001) (Figs. 7F, G). Since EMT is a key regulator of tumor cell migration, we evaluated EMT-related markers. Western blot analysis showed that PAICS knockdown markedly increased the epithelial marker E-cadherin (P < 0.05) and decreased mesenchymal markers N-cadherin (P < 0.01) and Vimentin (P < 0.001) (Figs. 7H, I). These findings indicate that PAICS promotes breast cancer cell migration, at least in part, by modulating the EMT process.
Fig. 7.
Expression and functional analysis of PAICS in breast cancer cells. A Relative PAICS mRNA levels across breast cancer cell lines. B Representative Western blot bands for PAICS protein. C Semi-quantitative analysis of PAICS protein expression. D Representative images from scratch assay following PAICS knockdown (Scale bar: 200 μm). E Quantification of scratch assay wound closure. F Representative Transwell assay images post-PAICS knockdown (Scale bar: 100 μm). G Quantification of migrating cells in Transwell assay. H Western blot bands for EMT-related markers after PAICS knockdown. I Semi-quantitative analysis of EMT marker protein expression. Data are presented as mean ± SD from three independent experiments. (n = 3 biological replicates); *p < 0.05, **p < 0.01, ***p < 0.001
Discussion
Through integrative multi-omics analysis and experimental validation, this study systematically highlighted the critical role of PAICS in breast cancer metastasis. Our results demonstrate that PAICS serves not only as an independent prognostic factor but also actively promotes breast cancer cell migration by modulating the EMT. These findings offer novel insights into the molecular mechanisms underlying breast cancer metastasis.
As a key enzyme in the de novo purine synthesis pathway, PAICS is not only aberrantly overexpressed in breast cancer but also demonstrates significant prognostic value across multiple tumor types. In colorectal cancer, PAICS is highly expressed in over 70% of cases, with elevated levels correlating with reduced five-year survival. This association remains significant after adjustment for conventional clinicopathological factors such as age, sex, and tumor stage [19]. Similarly, in lung adenocarcinoma, high PAICS expression is linked to poorer survival, and multivariate Cox analysis confirms its role as an independent prognostic factor [20, 21]. Functional studies, including CRISPR-Cas9–based knockdown, have shown that targeting PAICS markedly inhibits breast cancer cell proliferation and induces cell cycle arrest, highlighting its involvement in tumor progression [22]. Consistent with these observations, our study demonstrates that PAICS is highly expressed in breast cancer and is associated with patient survival in TCGA. However, these findings should be interpreted as exploratory, and further validation in independent clinical cohorts is required to determine whether PAICS provides prognostic information beyond established clinicopathologic factors. Although PAICS may represent a metabolically relevant biomarker and potential therapeutic target, its clinical utility remains to be established through prospective studies.
We confirmed that PAICS knockdown upregulates the epithelial marker E-cadherin and downregulates the mesenchymal markers N-cadherin and vimentin, indicating that PAICS may enhance breast cancer cell migration and invasion by promoting EMT. A study indicated that PAICS and its catalytic product SAICAR have been shown to activate PKM2, ERK1/2, and FAK signaling pathways, thereby promoting EMT and tumor cell invasion and metastasis [23]. Evidence from multiple tumor types further supports the role of PAICS in EMT. In breast cancer, miR-4731-5p suppresses glycolytic activity, migration, and EMT phenotypes by targeting PAICS and inhibiting PAICS-induced FAK phosphorylation [24]. In colorectal cancer, PAICS knockdown restores E-cadherin expression and reduces metastasis to bone and liver [25]. Similarly, in pancreatic cancer, elevated PAICS expression enhances invasive potential, whereas its knockdown inhibits invasion and increases E-cadherin levels [26]. Together, these studies suggest that PAICS not only supports purine synthesis in tumor cells but also promotes migration and invasion by triggering EMT-related mechanisms. Our findings align with these observations and further indicate that PAICS may act as a critical metabolic–signaling hub mediating metastasis in breast cancer. Future studies should investigate whether PAICS regulates classical EMT transcription factors such as Snail and ZEB1, adhesion molecules including integrin α6/β1, and how its EMT-promoting effects vary across breast cancer subtypes (ER/PR-positive, HER2-positive, triple-negative).
Immune infiltration analysis revealed that PAICS expression is positively correlated with macrophage M0 and resting NK cells, but negatively correlated with memory B cells, indicating a potential role for PAICS in shaping the immune cell landscape of the TME. Recent studies have highlighted that highly active metabolic pathways, including de novo purine synthesis, can remodel the TME by influencing immune cell recruitment, regulating immunosuppressive gene expression, and altering cellular metabolic states [27, 28]. For instance, upregulation of the PAICS-driven purine biosynthesis pathway in tumor cells not only promotes proliferation but may also modulate the extracellular matrix and immune cell behavior through metabolite secretion [21, 29]. Notably, a key product of purine metabolism, adenosine, accumulates in the TME and exerts profound immunoregulatory effects. Adenosine, generated primarily from extracellular ATP hydrolysis, is elevated in various solid tumors and interacts with A2A and A2B receptors on immune cells to suppress antitumor immune responses and facilitate immune evasion. It can inhibit the activation and cytokine production of effector T cells and NK cells, while promoting the function of immunosuppressive cells, such as regulatory T cells and M2-polarized macrophages, thereby contributing to a cold tumor immune phenotype [30]. Moreover, adenosine signaling has been identified as a key regulator of tumor-associated macrophage (TAM) function, promoting immunosuppressive polarization and further reducing the efficacy of immunotherapy [31]. These findings suggest that metabolic alterations, such as PAICS overexpression, may reshape the TME by increasing the accumulation of immunosuppressive metabolites like adenosine, thereby directly influencing immune cell composition and functional states. In pancreatic cancer, a metabolic prognostic index revealed that tumors with enhanced purine metabolism exhibited reduced CD8⁺ T cell infiltration and increased expression of immune checkpoints, linking metabolic reprogramming to immune evasion [32]. These observations suggest that the association of PAICS with increased macrophage M0 and decreased memory B cells in breast cancer may reflect metabolic alterations that influence immune composition and function. This underscores the need to evaluate whether PAICS expression correlates with immunotherapy response, macrophage polarization (M1/M2), NK cell activation, or immune checkpoint expression. Moreover, PAICS may represent a promising target for immunometabolic interventions or for identifying patient subgroups with distinct immune infiltration profiles, warranting further investigation.
Several limitations of this study should be acknowledged. (1) Our analysis relied primarily on public databases and in vitro experiments. Although cross-validation across multiple datasets strengthened the reliability of our findings, the absence of in vivo animal models and prospective clinical cohort validation limits a comprehensive understanding of the tumor microenvironment. (2) Further experimental work is required to elucidate the molecular mechanisms by which PAICS influences breast cancer metastasis, particularly the upstream and downstream signaling pathways. Notably, in the present study, we only performed PAICS knockdown experiments to assess its effects on migration, invasion, and EMT; overexpression experiments were not conducted. Future studies should include PAICS overexpression models to comprehensively evaluate its functional role and confirm dose-dependent effects on breast cancer aggressiveness. (3) Breast cancer exhibits substantial heterogeneity, and variations in genetic background and tumor microenvironment among different molecular subtypes may impact the function of PAICS and its associated miRNA networks, which were not systematically evaluated in this study. To address these limitations, future research will focus on developing in vivo models of breast cancer metastasis, validating PAICS biological function in an in vivo context, and employing multi-omics approaches to dissect the key signaling axes through which PAICS regulates EMT. Additionally, the universality and subtype specificity of PAICS will be assessed to inform the development of precision therapeutic strategies tailored to molecular subtypes.
In conclusion, this study systematically identified PAICS as a key regulator of breast cancer metastasis and demonstrated that it promotes tumor cell migration and invasion by inducing EMT. These findings enhance our understanding of breast cancer metastatic mechanisms and provide a foundation for the development of targeted therapies against PAICS, highlighting its potential clinical relevance.
Conclusion
This study identified PAICS as a key driver of breast cancer metastasis via multi-omics integration, machine learning, and in vitro validation. High PAICS expression correlates with poor prognosis and distant metastasis risk. Mechanistically, it promotes tumor cell migration/invasion by regulating EMT, modulates NK cell and M0 macrophage infiltration to affect the tumor microenvironment, and is regulated by miRNAs involved in angiogenesis/cell migration. In vitro, PAICS knockdown inhibits cell migration/invasion and reverses EMT markers. PAICS holds promise as a diagnostic biomarker and therapeutic target for breast cancer metastasis, with further in vivo studies needed to advance its clinical translation.
Supplementary Information
Acknowledgements
Not applicable.
Author contributions
Jian GUO was responsible for conceptualization, formal analysis, investigation, and writing the original draft; Lichao CEN contributed to methodology, software, and validation; Xinye QIAN handled data curation, visualization, and resources; and Shengban YOU, as the corresponding author, provided supervision, project administration, and writing review and editing. Jian GUO and Shengban YOU acquired fundings. All authors have read and approved the final version of the manuscript.
Funding Information
This project is funded by Traditional Chinese Medicine Research Project of Zhejiang Provincial (grant no.2023ZL654) and Natural Science Funding of Wenzhou Science and Technology Bureau (grant no. Y20220081).
Data availability
All data generated or analyzed in this study are available from the corresponding author (Shengban You) upon reasonable request. Public datasets used in this work include the Gene Expression Omnibus (GEO; accession numbers GSE191230, GSE125989 and GSE100534), the Cancer Genome Atlas Breast Invasive Carcinoma (TCGA-BRCA), and the GeneCards Human Gene Database (GeneCards). These data can be found here: GEO (https://www.ncbi.nlm.nih.gov/geo/), TCGA-BRCA (https://www.cancerimagingarchive.net/collection/tcga-brca/), and GeneCards (https://www.genecards.org/).
Declarations
Competing interests
The authors declare no competing interests.
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
8.References
- 1.Katsura C, et al. Breast cancer: presentation, investigation and management. Br J Hosp Med (Lond). 2022;83(2):1–7. [DOI] [PubMed] [Google Scholar]
- 2.Kim J, et al. Global patterns and trends in breast cancer incidence and mortality across 185 countries. Nat Med. 2025;31(4):1154–62. [DOI] [PubMed] [Google Scholar]
- 3.Wu Y, et al. Comparative analysis of cancer statistics in China and the United States in 2024. Chin Med J (Engl). 2024;137(24):3093–100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Riggio AI, Varley KE, Welm AL. The lingering mysteries of metastatic recurrence in breast cancer. Br J Cancer. 2021;124(1):13–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Park M et al. Breast cancer metastasis: mechanisms and therapeutic implications. Int J Mol Sci. 2022;23(12). [DOI] [PMC free article] [PubMed]
- 6.Li X, et al. CST6 protein and peptides inhibit breast cancer bone metastasis by suppressing CTSB activity and osteoclastogenesis. Theranostics. 2021;11(20):9821–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Medeiros B, Allan AL. Molecular mechanisms of breast cancer metastasis to the lung: clinical and experimental perspectives. Int J Mol Sci. 2019. 10.3390/ijms20092272. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Rashid NS, et al. Breast cancer liver metastasis: current and future treatment approaches. Clin Exp Metastasis. 2021;38(3):263–77. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ivanova M, et al. Breast cancer with brain metastasis: molecular insights and clinical management. Genes. 2023. 10.3390/genes14061160. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wu T, et al. Axillary lymph node metastasis in breast cancer: from historical axillary surgery to updated advances in the preoperative diagnosis and axillary management. BMC Surg. 2025;25(1):81. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Lüönd F, et al. Distinct contributions of partial and full EMT to breast cancer malignancy. Dev Cell. 2021;56(23):3203–e322111. [DOI] [PubMed] [Google Scholar]
- 12.Mosele F, Maio MD. Trastuzumab deruxtecan for breast cancer: do patients experience a comprehensive benefit? Ann Oncol. 2023;34(7):567–8. [DOI] [PubMed] [Google Scholar]
- 13.Jordan VC. Tamoxifen treatment for breast cancer: concept to gold standard. Oncol (Williston Park). 1997;11(2 Suppl 1):7–13. [PubMed] [Google Scholar]
- 14.Garcia J, et al. Bevacizumab (Avastin®) in cancer treatment: a review of 15 years of clinical experience and future outlook. Cancer Treat Rev. 2020;86:102017. [DOI] [PubMed] [Google Scholar]
- 15.Wang H, et al. Cisplatin prevents breast cancer metastasis through blocking early EMT and retards cancer growth together with paclitaxel. Theranostics. 2021;11(5):2442–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Zhu K, et al. PI3K/AKT/mTOR-targeted therapy for breast cancer. Cells. 2022. 10.3390/cells11162508.36611931 [Google Scholar]
- 17.Tabor S, et al. How to predict metastasis in luminal breast cancer? Current solutions and future prospects. Int J Mol Sci. 2020. 10.3390/ijms21218415. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lian G, et al. Bioinformatics analysis of the immune cell infiltration characteristics and correlation with crucial diagnostic markers in pulmonary arterial hypertension. BMC Pulm Med. 2023;23(1):300. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Agarwal S, et al. PAICS, a purine nucleotide metabolic enzyme, is involved in tumor growth and the metastasis of colorectal cancer. Cancers (Basel). 2020. 10.3390/cancers12040772. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Goswami MT, et al. Role and regulation of coordinately expressed de novo purine biosynthetic enzymes PPAT and PAICS in lung cancer. Oncotarget. 2015;6(27):23445–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Liu J, et al. PAICS-driven purine biosynthesis and its prognostic implications in lung adenocarcinoma: a novel risk stratification model and therapeutic insights. Curr Issues Mol Biol. 2025. 10.3390/cimb47050366. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Hany D, Vafeiadou V, Picard D. CRISPR-Cas9 screen reveals a role of purine synthesis for estrogen receptor α activity and tamoxifen resistance of breast cancer cells. Sci Adv. 2023;9(19):eadd3685. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Huo A, Xiong X. PAICS as a potential target for cancer therapy linking purine biosynthesis to cancer progression. Life Sci. 2023;331:122070. [DOI] [PubMed] [Google Scholar]
- 24.Lang L, et al. Tumor suppressive role of microRNA-4731-5p in breast cancer through reduction of PAICS-induced FAK phosphorylation. Cell Death Discov. 2022;8(1):154. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Agarwal S, et al. PAICS, a purine nucleotide metabolic enzyme, is involved in tumor growth and the metastasis of colorectal cancer. Cancers. 2020. 10.3390/cancers12040772. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Agarwal S, et al. PAICS, a De Novo Purine Biosynthetic Enzyme, Is Overexpressed in Pancreatic Cancer and Is Involved in Its Progression. Transl Oncol. 2020;13(7):100776. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Zhang F, et al. Cancer associated fibroblasts and metabolic reprogramming: unraveling the intricate crosstalk in tumor evolution. J Hematol Oncol. 2024;17(1):80. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Xia L, et al. The cancer metabolic reprogramming and immune response. Mol Cancer. 2021;20(1):28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Chakravarthi BV, et al. Expression and Role of PAICS, a De Novo Purine Biosynthetic Gene in Prostate Cancer. Prostate. 2017;77(1):10–21. [DOI] [PubMed] [Google Scholar]
- 30.Han Y, et al. Unlocking the adenosine receptor mechanism of the tumour immune microenvironment. Front Immunol. 2024;15:1434118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Yang L, Zhang Y, Yang L. Adenosine signaling in tumor-associated macrophages and targeting adenosine signaling for cancer therapy. Cancer Biol Med. 2024;21(11):995–1011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Song W, et al. Metabolic reprogramming shapes the immune microenvironment in pancreatic adenocarcinoma: prognostic implications and therapeutic targets. Front Immunol. 2025;16:1555287. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data generated or analyzed in this study are available from the corresponding author (Shengban You) upon reasonable request. Public datasets used in this work include the Gene Expression Omnibus (GEO; accession numbers GSE191230, GSE125989 and GSE100534), the Cancer Genome Atlas Breast Invasive Carcinoma (TCGA-BRCA), and the GeneCards Human Gene Database (GeneCards). These data can be found here: GEO (https://www.ncbi.nlm.nih.gov/geo/), TCGA-BRCA (https://www.cancerimagingarchive.net/collection/tcga-brca/), and GeneCards (https://www.genecards.org/).







