Skip to main content
Cancer Innovation logoLink to Cancer Innovation
. 2025 Jan 7;4(1):e142. doi: 10.1002/cai2.142

Identification of Significant Single‐Nucleotide Polymorphisms Associated With Breast Cancer Recurrence and Metastasis Using GWAS

Shujuan Sun 1, Sha Yin 1, Jie Huang 1, Dongdong Zhou 1, Qiaorui Tan 1, Xiaochu Man 1, Wen Wang 1, Jiale Zhang 1, Huihui Li 1,
PMCID: PMC11705446  PMID: 39777116

Abstract

Background

Identification of risk genes and loci associated with the recurrence and metastasis of breast cancer (BC) is of utmost importance. Genome‐wide association studies (GWASs) represent valuable tools for identifying the disease risk associated with a given single‐nucleotide polymorphism (SNP); they offer significant insights into the disease progression mechanism by analyzing SNP information of the entire genome. Though GWAS has already identified several genetic susceptibility SNPs for BC, their significance in the recurrence and metastasis of this cancer remains unclear. Here, we used a GWAS approach to identify SNPs specifically associated with the risk of BC recurrence and metastasis.

Methods

This study adopted a two‐stage GWAS approach. In the first stage, 97 pairs of BC patients with or without recurrence and metastasis, treated at the Shandong Cancer Hospital and Institute from November 2013 to April 2014, were identified using propensity score matching. DNA extracted from the patient peripheral blood was then subjected to Illumina ASA chip analysis for genome‐wide SNP detection. In the second stage, the findings were verified in a validation set of 854 BC patients recruited at the same hospital from May 2014 to June 2015. SNP genotyping was performed using time‐of‐flight mass spectrometry. The SNP loci and their corresponding genes and pathways were analyzed using the DAVID (https://david.ncifcrf.gov/) online enrichment analysis tool.

Results

Based on the GWAS results, 191 SNP‐related genes significantly associated with BC recurrence and metastasis were identified as expression quantitivative trait loci (p < 0.001). Functional and pathway enrichment analyses subsequently revealed the potential involvement of glutamatergic synaptic transmission, calcium signaling, and insulin secretion pathways in BC recurrence and metastasis. Based on genotype correlation and database expression levels, rs10108514, rs12920540, rs4273077, and rs4730155 were found to be significantly associated with the risk of BC recurrence and metastasis.

Conclusion

Our study suggests that the SNPs rs10108514, rs12920540, rs4273077, and rs4730155 are correlated with the risk of BC recurrence and metastasis, potentially by being implicated in glutamatergic synaptic transmission, calcium signaling, and insulin secretion pathways.

Keywords: breast cancer, metastasis, recurrence, single‐nucleotide polymorphisms


Our study employed a two‐stage genome‐wide association study strategy to identify specific single‐nucleotide polymorphisms (SNPs) that are uniquely associated with the risk of breast cancer (BC) recurrence and metastasis. In the Screening phase, 191 SNPs were significantly associated with BC (p < 0.001). Kyoto Encyclopedia of Genes and Genomes enrichment analysis indicated potential roles for glutamatergic synaptic transmission, calcium signaling, and insulin secretion in BC recurrence and metastasis. Verification phase results confirmed that the SNPs rs10108514, rs12920540, rs4273077, and rs4730155 are associated with the risk of BC recurrence and metastasis.

graphic file with name CAI2-4-e142-g005.jpg


Abbreviations

BC

breast cancer

CASC16

cancer susceptibility candidate 16

CI

confidence interval

DFS

disease‐free survival

EAS

East Asian Population

GWAS

genome‐wide association study

HER2

human epidermal growth factor receptor 2

HR

hazard ratio

IQR

interquartile range

MAF

minor allele frequency

NAMPT

nicotinamide phosphoribosyl transferase

PVT1

plasmacytoma variant translocation 1

SNP

single‐nucleotide polymorphism

1. Introduction

Breast cancer (BC) is the most common malignancy affecting women worldwide, with an increasing incidence and a low average age of onset [1]. Recent advances in early screening and treatment strategies have significantly improved the outcomes for BC patients. However, a significant proportion (20%–30%) of patients with early‐stage BC still face the challenge of disease progression to advanced stages, characterized by local recurrence or distant metastasis [2, 3, 4]. Despite improved accessibility to various treatment options, based on cancer stage and molecular profiles, some patients with BC are still at a significant risk of disease progression [5]. Hence, the primary objective of current treatment strategies for BC is to prevent recurrence and metastasis by ensuring the early identification of patients at high risk of BC progression. Although some progress in delineating the molecular mechanisms underlying the recurrence and metastasis of BC has been made [6], the heterogeneity of this disease limits the generalizability of these findings. Thus, the development of individualized treatment approaches is becoming increasingly important.

Predicting tumor recurrence and metastasis for the clinical monitoring and treatment of patients with BC is challenging [7]. For instance, an individual's genetic characteristics have a substantial impact on BC tumor initiation and progression [8]. Genome‐wide association studies (GWASs) enable the exploration of the relationship between genes and diseases or traits by comparing large‐scale genetic variations. GWAS are highly effective at identifying genes that render individuals more susceptible to diseases such as BC. Single‐nucleotide polymorphisms (SNPs), the most common form of genetic variation in the human genome, can affect gene expression or protein function, thus contributing to the development of diseases, including cancer [9]. Numerous preliminary studies [10, 11, 12] have identified SNPs that are associated with BC risk and prognosis. For example, Miedl and colleagues [13] reported that the SNPs targeting estrogen receptor 1, specifically rs2046210 and rs9383590, were correlated with BC risk. Furthermore, Cui et al. [14] demonstrated that the SNP rs2071095 might influence the susceptibility of Chinese women to BC by affecting the expression of the H19 gene.

Though this type of research has identified certain SNPs that are associated with BC outcomes, the use of candidate SNP research strategies is limited by the restricted number of genetic variations and the reliance on existing knowledge of cancer biology. Conversely, GWAS offers comprehensive gene coverage, enabling a complete genomic investigation, which can reveal greater numbers of disease or phenotype risk loci [15]. This allows for precise genome mapping and in‐depth investigations of diseases or phenotypes. The identified risk SNPs, together with other prognostic factors, may provide valuable information to enable the effective prediction, prevention, and treatment of BC.

Here, we performed GWAS of patients with and without BC recurrence by adopting a multi‐stage, case–control study design and using Illumina ASA chip technology. After identifying SNPs associated with the risk of BC recurrence, we selected a subset of these loci for further validation in an expanded BC cohort to identify the true SNP variants associated with BC recurrence. This research aims to provide reliable molecular markers that can be used to predict the rates of BC progression in clinical practice and offer theoretical insights into the molecular mechanisms underlying BC recurrence and metastasis.

2. Methods

2.1. Study Design

This study used a two‐phase GWAS research strategy (Figure 1). Using propensity score matching, 97 pairs of BC patients, treated at the Shandong Cancer Hospital and Institute between November 2013 and April 2014, with or without recurrence and metastasis at the first stage, were selected for this study. The DNA was extracted from the patient peripheral blood samples and used for genome‐wide detection on Illumina ASA chips. The SNP loci associated with BC recurrence and metastasis were identified using a logistic regression analysis of GWAS data. Between May 2014 and June 2015, a total of 854 BC patients from the same hospital were enrolled in the second phase of the study. SNP genotyping was performed using time‐of‐flight mass spectrometry to confirm the association with BC recurrence and metastasis.

Figure 1.

Figure 1

Study design flow diagram. BC, breast cancer; DFS, disease‐free survival; EAS, East Asian population; GWAS, genome‐wide association studies; MAF, minor allele frequency; SNP, single‐nucleotide polymorphism.

2.2. Study Subjects

The study population consisted of patients with BC who visited the Shandong Cancer Hospital and Institute between November 1, 2013, and June 30, 2015. The inclusion criteria were as follows: (1) women with primary BC, who had a histopathologically confirmed diagnosis of BC, (2) local surgical treatment of the breast, and (3) blood and DNA samples that met the quality control standards. The exclusion criteria were as follows: (1) primary stage IV BC, (2) double primary BC, (3) a history of other malignancies, (4) incomplete clinicopathological data, (5) < 1 month of disease‐free survival (DFS), or (6) unknown loss to follow‐up. Based on these criteria, 1048 female BC patients were enrolled in the study. None of these individuals had any familial connection to the Han Chinese ethnic group. The Ethics Committee of Shandong Cancer Hospital and Institute approved this study (approval number: SDTHEC2022009018). Patients signed a consent form, and 2 mL of peripheral blood was collected from the elbow vein. The samples were stored at −80°C until use.

2.3. DNA Extraction and Illumina ASA Microarray Whole‐Genome Assay

Peripheral venous blood (2 mL) was collected from patients into EDTA‐containing tubes. DNA was extracted from the blood samples using a DNA extraction kit and stored at −80°C. The concentration and purity of the DNA were determined, and then each sample was subjected to experimental quality control using the Illumina ASA high‐throughput genotyping chip. The samples were chosen based on a detection rate >90%, a minor allele frequency (MAF) > 0.01 at each locus, and a Hardy–Weinberg equilibrium p value < 0.000001.

2.4. SNP Selection and Validation

After analyzing the GWAS data, SNP loci that met the following criteria: p value < 10−3, loci located on autosomes, and an MAF > 0.05 in the East Asian population (EAS) were selected. The dbSNP database (https://www.ncbi.nlm.nih.gov/) was searched to find the genes associated with the SNP loci mentioned above and their positions on the DNA strands of a given gene. A search was performed in PubMed (http://www.ncbi.nlm.nih.gov/pubmed). The three‐dimensional (3D) SNP database (http://cbportal.org/3dsnp/) was searched to determine the current status of the SNP loci, focusing specifically on malignant tumors and breast tumors, and retrieved the chromatin spatial interactions of SNP loci. The priority was to identify SNP loci in genes related to tumor evolution, which were present in gene‐coding regions, and that may interact with chromatin spatial structures in the noncoding regions. This led to the selection of 21 SNPs and their inclusion in the subsequent validation phase. Primer sequences were designed and synthesized based on the SNP loci identified through screening. The SNP genotypes were detected using time‐of‐flight mass spectrometry.

2.5. Bioinformatics Analysis

This study analyzed gene expression data obtained from both BC and normal tissues. A combination of two databases, namely, The Cancer Genome Atlas (TCGA) (https://portal.gdc.cancer.gov/) and the Genotype‐Tissue Expression (GTEx) database (https://commonfund.nih.gov/GTEx/), were used to examine differential gene expression. Based on the results of SNP loci in the 3D SNP database (http://cbportal.org/3dsnp/), information on the regulation mediated by the noncoding SNP loci and chromatin loops was obtained. This approach provided a new perspective for studying the influence of SNPs on gene regulation. Furthermore, the DAVID online enrichment analysis website (https://david.ncifcrf.gov/) was used to investigate the enrichment of relevant pathways, enabling an exploration of correlations between SNPs and enriched pathways.

2.6. Statistical Analysis

The statistical analysis was performed using SPSS 22.0 software. Study subject data with a normal distribution were conveyed as the mean ± standard deviation, whereas the median with interquartile range (IQR) was used to describe data with a skewed distribution. The t test was used to determine the statistical significance of differences between two groups with normal distributions, whereas the Wilcoxon rank‐sum test was used to compare groups whose distributions were not normal. Statistical data were expressed as the number of patients and percentages, whereas the chi‐square test or Fisher's exact test was used to compare the two groups. The Kaplan–Meier method was used to calculate survival time and plot survival curves, whereas the log‐rank test was employed to compare the differences in DFS of patients bearing different SNP loci. Cox regression analysis was used to adjust for patient age, hormone receptor status, HER‐2 status, and TNM staging. The hazard ratio (HR) and 95% confidence intervals (CIs) for recurrent metastasis were then calculated. Statistical tests were performed as two‐sided probability tests, with a p value < 0.05 considered as a measure of statistically significant differences.

3. RESULTS

3.1. Basic Information About the Study Subjects

The study participants were divided into the screening and validation cohorts (Table 1). The screening cohort consisted of 194 participants, with a median age of 46 (IQR = 40–51) years. The validation cohort included 854 participants, with a median age of 47.5 (IQR = 41–55) years. There were no statistically significant differences between the two groups regarding BC laterality, pathological type, histological grading, TNM staging, and family history (p > 0.05). The predominant pathological BC type observed was invasive ductal carcinoma, comprising 80% of cases. Approximately 35% of cases were considered as Grade III based on the histological grading, whereas 45% were classified as Grades I–II. Regarding TNM staging, Stage II was the most prevalent, representing approximately 41% of cases, followed by Stage III (~35%) and Stage I (~24%). Approximately 25% of the patients had a family history (immediate relatives) of cancer.

Table 1.

Patient characteristics.

Variable Screening stage (n = 194, %) Validation stage (n = 854) p value
Age, median (inner distance) 46 (40–51) 47.5 (41–55) 0.013
Laterality 0.925
Left 100 (51.55) 437 (51.17)
Right 94 (48.45) 417 (48.83)
Pathological type 0.063
Infiltrating ductal carcinoma 170 (87.63) 701 (82.08)
Others 24 (12.37) 153 (17.92)
Histological grading 0.475
Grades I–II 90 (46.39) 384 (44.96)
Grade III 71 (36.60) 292 (34.19)
Unknown 33 (17.01) 178 (20.84)
Breast subtype 0.002
Luminal A 47 (24.23) 206 (24.12)
Luminal B HER2‐negative 51 (26.29) 252 (29.51)
Luminal B HER2‐positive 15 (7.73) 122 (14.29)
HER2+ 39 (20.10) 102 (11.94)
Triple‐negative 38 (19.59) 128 (14.99)
Unknown 4 (2.06) 44 (5.15)
TNM staging 0.956
Phase I 45 (23.20) 206 (24.12)
Phase II 81 (41.75) 356 (41.69)
Phase III 68 (35.05) 292 (34.19)
Family history 0.355
Yes 52 (26.80) 202 (23.65)
No 142 (73.20) 652 (76.35)

Note: Tumor‐node‐metastasis (TNM) staging was defined according to the 2017 American Joint Commission for Cancer (AJCC) (8th edition) guidelines; Molecular staging was defined according to the 2021 NCCN guidelines.

Abbreviation: HER2, human epidermal growth factor receptor 2.

3.2. Results of the Genome‐Wide Association Analysis

After processing the raw Illumina ASA chip data of 194 patients with BC, we conducted a systematic quality control assessment of the genotyping data, which confirmed that the results passed the quality control criteria. Population stratification issues in the samples were identified and adjusted using principal component analysis. After removing outliers, the genetic variation between the samples was found to be minimal, indicating a good level of consistency (Figure 2a–c). Subsequently, an association analysis was performed on the 510,152 high‐quality genotyped loci. An additive model was used to investigate the association between the genetic loci and the recurrence and metastasis of BC. Covariates such as patient age, hormone receptor status, HER‐2 status, and TNM stage were adjusted using logistic regression to identify SNPs associated with BC recurrence and metastasis. A Manhattan plot was used to illustrate the distribution of genetic variants associated with BC recurrence and metastasis across chromosomes and the whole genome (Figure 2d). Furthermore, the Quantile–Quantile (QQ) plot was used to demonstrate the differences between the observed and predicted values of the genetic variants (Figure 2e).

Figure 2.

Figure 2

Population hierarchical assessment and GWAS visualizations for breast cancer (BC) recurrence and metastasis. (a–c) Sample population hierarchical assessment chart. (d) Manhattan plot for the GWAS of breast cancer recurrence and metastasis. (e) QQ map for the GWAS of BC recurrence and metastasis. GWAS, Genome‐Wide Association Study.

In this study, 20,132 SNP sites associated with BC recurrence and metastasis were identified based on the additive model (p < 0.05); of these, 3368 SNP sites had a p < 0.01 and 191 SNP sites had a p < 0.001. The most significant SNP site had a p value of 1.88 × 10−6; however, no SNP sites in this study met the significance threshold of the statistical analysis (p < 5.00 × 10−8). Based on the results of the GWAS analysis, we followed the SNP screening strategy described in the Methods section. This resulted in the identification of 21 SNP loci (Tables 2 and 3), which could be subjected to subsequent validation.

Table 2.

Primer sequences for the 21 SNPs identified in this study.

Loci Sequence 1 (5′–3′) Sequence 2 (5′–3′) Amplification length (bp)
rs1003533 TGGATGGAGAGTTGAATGCCAGCCAC TGGATGGAACACATGTGATTTACACC 139
rs10108514 TGGATGATTAGCTTGAGTGCCTGGTG TGGATGGACAGTGAGGGCCAGATTC 119
rs1043996 TGGATGGGGTCACAGTCATTGATGTC TGGATGATGGTATCTGCACCAACCTG 117
rs10752609 TGGATGGCAGCTTCAGTTGTCTGGTG TGGATGTGCAGTGTTCAGTCCCTCAG 125
rs1109866 TGGATGTTCTTCTTCACGCTCGTGCC TGGATGCCAAGACAGCGAATCAGCAC 134
rs111770568 TGGATGCTCGACAGAATGGTACCATC TGGATGCAAGAGCAGTCCAAGAGTAT 121
rs1164760 TGGATGAAAACCTCAGCGAATGCACC TGGATGTCTAGGTCTTTCCACTGTCC 96
rs12920540 TGGATGCGCTTCCATAAAGTGGTGAG TGGATGAGGGATTATATGGCCTGCTC 110
rs1799801 TGGATGTTATACTTCTCTGACTCGGG TGGATGGAGCTGAAACAAAGCAAGCC 112
rs2227251 TGGATGTGTGTCTGCCGGGATGATG TGGATGTGAAGACTGAGTCCAGGTGC 153
rs249820 TGGATGAAACCAGCTCCCTTAAGTGC TGGATGCTTCATGGTGGGCATTAAGG 91
rs2561530 TGGATGGAATTAGTAACCTGGCTGTC TGGATGAGTTGCCCACCAGCAATTTC 100
rs3746069 TGGATGTGCTTCAGACACTGCCGTAG TGGATGACAGGGACACAGCACCCAA 113
rs3814811 TGGATGCCTGCCCCCTTAAAACAGAG TGGATGGGATCTTGATGCCAAGTGTG 107
rs4273077 TGGATGTTCACAGTGCCAGCGGATTC TGGATGTCACCATGGCTACAGGTTTC 116
rs4730155 TGGATGTTGAAATCGAGCCAAGATCC TGGATGTCATACTTGATAGTGTTCGG 105
rs74860409 TGGATGTTATCCACTCCCATTTCAAG TGGATGCTTTAGTCTCCCCACCATTC 111
rs7719624 TGGATGGTTGGCTTCTGAGTTCCATC TGGATGTCCTCAATGAGATCCTGGTG 100
rs8736 TGGATGTCCCATGGCTTCCATCTGAG TGGATGGATTTACACACGGTGACCTG 114
rs921943 TGGATGCCCGTGTAGAGCTTCTAACT TGGATGTCTTCCAAAAGAGACCCTGC 128
rs9830253 TGGATGAAAGCAGCTGTTAACCTCCG TGGATGGGGTGGGATGCTATCTTTTC 118

Note: The amplification length encompasses the sequence tagged and the longest SNP sequence.

Abbreviation: SNP, single‐nucleotide polymorphism.

Table 3.

Basic information for the 21 verified SNPs.

Gene SNP ID Chromosome location Alleles gene MAF‐EAS OR (95% CI)
IRF1‐AS1 rs1003533 chr5:131755650 C>T 0.331 2.459 (1.484–4.075)
PVT1 rs10108514 chr8:128822708 A>G 0.309 2.825 (1.667–4.788)
NOTCH3 rsl043996 chr19:15295133 G>A 0.593 2.549 (1.483–4.379)
KCNN3 rsl0752609 chrl:154791127 G>A 0.800 3.429 (1.679–7.004)
ABCB6 rsl109866 chr2:220083279 C>T 0.138 0.208 (0.094–0.459)
POLN rsll1770568 chr4:2167331 G>A 0.083 3.426 (1.702–6.897)
CHST11 rs1164760 chr12:105063253 C>T 0.280 0.340 (0.181–0.642)
CASC16 rs12920540 chr16:52625612 C>A 0.430 2.472 (1.469–4.159)
ERCC4 rsl799801 chr16:14041957 T>C 0.259 0.334 (0.175–0.637)
CTDSP1 rs2227251 chr2:219266422 C>A 0.071 4.09 (1.853–9.030)
LINCO2453 rs249820 chr12:98897472 A>G 0.105 10.47 (3.157–34.690)
COQ8B, ITPKC rs2561530 chr19:41224313 G>A 0.714 2.516 (1.456–4.346)
GNA11 rs3746069 chr19:3114863 C>T 0.252 3.309 (1.751–6.253)
NYNRIN rs3814811 chr14:24887461 T>C 0.784 0.207 (0.096–0.445)
TNFRSF13B rs4273077 chr17:16849138 A>G 0.479 2.788 (1.560–4.983)
NAMPT rs4730155 chr7:105924142 T>C 0.876 0.182 (0.067–0.491)
ZBTB38 rs74860409 chr3:141144140 C>T 0.082 5.194 (2.066–13.05)
TGFBI rs7719624 chr5:135377565 C>T 0.581 2.593 (1.509–4.457)
MBOAT7, TMC4 rs8736 chr19:54677188 T>C 0.218 2.894 (1.550–5.404)
DMGDH rs921943 chr5:78316475 C>T 0.139 0.288 (0.145–0.573)
COL6A6 rs9830253 chr3:130284283 G>A 0.444 3.186 (1.834–5.535)

Abbreviations: CI, confidence interval; EAS, East Asian Population; MAF, minor allele frequency; OR, odds ratio; SNP, single‐nucleotide polymorphism.

3.3. Association Between the Genotypes of rs10108514, rs12920540, rs4273077, and rs4730155 and the DFS of Patients With BC

The genotypes of the 21 SNP loci mentioned above and their corresponding recurrence and metastasis events were examined using the log‐rank test in the verified population. The survival curves revealed that, among the different genotypes of the four significant SNPs loci (p < 0.05), only rs10108514 (A>G), rs12920540 (C>A), rs4273077 (A>G), and rs4730155 (T>C) significantly impacted DFS (Figure 3a–d). The identified SNPs were stratified to evaluate their specific impact on tumor metastasis or recurrence. However, no significant SNPs associated with tumor metastasis were observed. Nonetheless, we did identify three SNPs, namely, rs1799801 (T>C), rs2227251 (C>A), and rs4730155 (T>C), which were significantly associated with tumor recurrence (Supporting Information S1: Figure S1).

Figure 3.

Figure 3

Log‐rank and Cox regression tests of differential DFS in breast cancer (BC) patients by genotype for four SNPs. Log‐rank test of DFS in BC patients with different genotypes of rs10108514 (a), rs12920540 (b), rs4273077 (c), or rs4730155 (d). (e) The Cox regression test was used to compare the contributions of the different genotypes of rs10108514, rs12920540, rs4273077, or rs4730155 to the risk of BC recurrence and metastasis. CI, confidence interval; DFS, disease‐free survival; SNP, single‐nucleotide polymorphism.

Subsequent Cox regression analysis further delineated the differences in the risk of recurrent metastasis across these genotypes (Figure 3e). Compared with homozygous AA, carriers of rs10108514 (AG/GG) had a 27% reduced risk of BC recurrence and metastasis (HR = 0.73, 95% CI = 0.54–0.98), consistent with the GWAS result (OR = 2.825, 95% CI = 1.667–4.788) for the A allele being a risk allele. Compared with homozygous CC, carriers of rs12920540 (CA/AA) had a 52% increased risk of BC recurrence and metastasis (HR = 1.52, 95% CI = 1.06–2.18), consistent with the GWAS result (OR = 2.472, 95% CI = 1.469–4.159) for the C allele being a risk allele. Carriers of rs4273077 (AG/GG) had a 26% reduced risk of BC recurrence and metastasis compared with homozygous AA (HR = 0.74, 95% CI = 0.54–1.01), consistent with the GWAS result for the A allele being a risk allele (OR = 2.788, 95% CI = 1.56–4.983). Conversely, carriers of rs4730155 (CT/TT) had a 35% reduced risk of BC recurrence and metastasis compared with homozygous CC (HR = 0.65, 95% CI = 0.45–0.93), consistent with the GWAS result for the T allele being a protective allele (OR = 0.182, 95% CI = 0.067–0.491). Of note, all four SNPs are located in intronic regions of the gene.

3.4. 3D SNP Spatial Role Information

We used the 3D SNP database (http://cbportal.org/3dsnp) to determine the distribution and spatial interactions of the four SNPs (rs10108514, rs12920540, rs4273077, and rs4730155) in the genome. Additionally, Circos plots were used to visually represent the chromosomal interactions among the noncoding mutations, distal regulatory elements, and promoters. The expression quantitative trait loci analysis revealed that rs10108514 was significantly associated with plasmacytoma variant translocation 1 (PVT1) expression in the thyroid. The rs10108514 is located in the intron region of the PVT1 gene (on chromosome 8q24) [16], which encodes a long noncoding RNA [16, 17]. The functional annotation of the rs10108514 mutation revealed its primary involvement in three functional categories: transcription factor binding sites (TFBS, score: 64.66), enhancers (score: 50.57), and promoters (score: 20.95). Specifically, in pancreatic PANC‐1 cells, rs10108514 resided at the binding site of POLR2A with high DNA accessibility (852/1000). In addition, it resided at the binding sites of EP300, ESR1, and FOXA1 with high DNA accessibility (265/1000) in BC T‐47D cells. In vascular tissue HUVEC cells, the binding sites for JUN, MYC, and FOXA1 exhibited high DNA accessibility (223/1000) (Figure 4a). The total score of the rs12920540 mutation was derived from its enhancer (0.25) function. The gene located 2 kb upstream or downstream of the SNP was identified as cancer susceptibility candidate 16 (CASC16) (Figure 4b). The total score of the rs4273077 mutation was derived from its role as a TFBS (17.49) and an enhancer (5.00). The gene located 2 kb upstream or downstream of rs4273077 was identified as tumor necrosis factor receptor superfamily member 13B (TNFRSF13B). In the TFBS fraction, rs4273077 was located at the binding site of CHD2, CTCF, RAD21, YY1, and ZNF143 in H1‐hESC cells of esophageal squamous epithelial tissue with DNA accessibility (310/1000) (Figure 4c). Finally, the rs4730155 mutation predominantly scored in the promoter category (100.00) and TFBS (40.93). The gene located 2 kb upstream or downstream of this SNP was identified as nicotinamide phosphoribosyl transferase (NAMPT). In the expression quantitative trait loci fraction, rs4730155 was significantly associated with the expression level of NAMPT in cells that have been transformed into fibroblasts and left ventricular cells. In the TFBS fraction, rs4730155 was located at the binding site of POLR2A in esophageal squamous epithelial tissue H1‐hESC cells with high DNA accessibility (652/1000) (Figure 4d). Overall, these findings provide valuable insights into the functional roles of these mutations and their association with gene expression and DNA binding.

Figure 4.

Figure 4

Circos diagrams visualizing chromatin interactions and SNP annotations for four SNPs. Circos diagrams for rs10108514 (a), rs12920540 (b), rs4273077 (c), and rs4730155 (d). Chromatin, annotated genes, histones (depicted in red), transcription factors (depicted in blue), current and associated SNPs, and 3D chromatin interactions (in that order) can be seen moving from the outer to the inner rings of the Circos diagram. SNP, single‐nucleotide polymorphism.

3.5. Analysis of Differential Gene Expression

To understand the expression of the four SNPs in breast tumor tissues compared with normal or paracancerous tissues, the following SNPs were investigated: PVT1: rs10108514 (A>G), CASC16: rs12920540 (C>A), TNFRSF13B: rs4273077 (A>G), and NAMPT: rs4730155 (T>C). For this study, we obtained paired breast invasive carcinoma and nearby tissue samples from the BRCA project in TCGA. We also acquired normal tissue data from the GTEx database for comparison. The results showed a significant increase in the expression of the PVT1 gene corresponding to rs10108514 in breast tumor tissues when compared with normal tissues and paired adjacent tissues (p < 0.001) (Figure 5a). Similarly, the expression of the CASC16 gene corresponding to rs12920540 was substantially increased in breast tumor tissues compared with normal tissues and paired adjacent tissues (p < 0.001) (Figure 5b). However, the expression of the TNFRSF13B gene corresponding to rs4273077 was not significantly different among BC tumor tissues, normal tissues, or paired adjacent tissues (p ≥ 0.05) (Figure 5c). Finally, the expression of the NAMPT gene corresponding to rs4730155 was significantly higher in breast tumor tissues compared with normal tissues and paired adjacent tissues (p < 0.05) (Figure 5d). This study demonstrated distinct gene expression patterns in breast tumor tissues compared with normal and paracancerous tissues for the investigated SNPs.

Figure 5.

Figure 5

Differential expression of PVT1 (a), CASC16 (b), TNFRSF13B (c), and NAMPT (d) genes. ***, p < 0.001; *, p < 0.05; ns, p ≥ 0.05.

3.6. Significant Association Between High TNFRSF13B Expression and BC Prognosis

To investigate the association between the genes corresponding to the four identified SNP loci and BC prognosis, we next obtained survival and prognosis information from TCGA. We categorized gene expression into high and low based on a 50% cutoff. The results revealed that high TNFRSF13B expression was significantly associated with a more favorable prognosis, as indicated by the prolonged overall survival and progress‐free interval of BC patients (p < 0.05). Moreover, patients with high TNFRSF13B expression tended to have a longer disease‐specific survival than those with low TNFRSF13B expression (p = 0.077). Conversely, PVT1, CASC16, and NAMPT expression did not significantly impact BC prognosis (Figure 6a–c).

Figure 6.

Figure 6

Functional enrichment analysis and survival analysis of differential expression of the TNFRSF13B gene. (a) OS log‐rank test for the differential expression of the TNFRSF13B gene. (b) DSS log‐rank test for the differential expression of the TNFRSF13B gene. (c) PFI log‐rank test for the differential expression of the TNFRSF13B gene. (d) GO/KEGG enrichment analysis of genes corresponding to differential SNPs in GWAS analysis results. DSS, disease‐specific survival; GO, gene ontology; HR, hazard ratio; KEGG, Kyoto Encyclopedia of Genes and Genomes; OS, overall survival; PFI, progress‐free interval; SNP, single‐nucleotide polymorphism.

3.7. GO/KEGG Enrichment Analysis

Using the online DAVID tool (https://david.ncifcrf.gov/), we performed gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses on genes corresponding to SNP loci substantially associated with BC recurrence and metastasis (p < 0.001). The results revealed that the enriched target genes were involved in biological processes such as glutamatergic synaptic transmission and KEGG signaling pathways such as the calcium signaling system and the insulin secretion route (Figure 6d).

4. Discussion

BC is the most common malignant tumor affecting women worldwide [1]. The long‐term survival of BC patients is hampered by a high risk of recurrence and metastasis. Thus, accurately predicting the risk of recurrence and metastasis for BC is crucial [18]. BC arises due to a series of genetic, epigenetic, and phenotypic changes. Genetic polymorphisms associated with BC risk have been identified in genes involved in multiple biological pathways [19]. These genetic polymorphisms further contribute to variations in disease susceptibility and severity among individuals [20]. In this study, we conducted unbiased screening of candidate SNPs through GWAS analysis and validated them using a larger clinical cohort. During the validation phase, four SNP loci, namely, PVT1: rs10108514 (A>G), CASC16: rs12920540 (C>A), TNFRSF13B: rs4273077 (A>G), and NAMPT: rs4730155 (T>C), were significantly associated with the risk of BC recurrence and metastasis. The 3D SNP analysis provides a new perspective on the spatial and functional relationships among SNPs within the genome. This method enabled us to determine the potential regulatory roles of candidate SNPs in gene expression and cellular processes associated with BC progression by identifying specific DNA binding sites and transcription factor interactions for the four SNPs.

This study revealed a significant association between the rs10108514 (A>G) mutation in PVT1 and an increased risk of BC recurrence and metastasis, confirming the role of this gene in BC progression [21]. Previous research has shown that SNP variants of PVT1 can induce functional changes that have been linked to the development of various malignant tumors, including BC [22]. The second SNP identified, rs12920540, was located in the intron region of the CASC16 (LOC643714) gene. Our findings also align with previous studies indicating the significant association between this gene and susceptibility to BC. The genomic region of TOX3/LOC643714 has been extensively studied and linked to an increased risk of BC [23]. Recent research has shown that genetic variants (including the common rs3803662 and rs4784227 genetic variants) at multiple loci on TOX3/LOC643714 are associated with an increased risk of BC [24].

Here, we found a correlation between the rs4273077, located in the intronic region of the TNFRSF13B gene, and the risk of BC recurrence and metastasis. Moreover, we identified a significant association between high TNFRSF13B expression and a favorable BC prognosis. In a previous report, TNFRSF13B was proposed as a predictive marker for BC progression and a potential therapeutic target for triple‐negative BC [25]. Thus, our findings further support the potential role of TNFRSF13B in BC progression. However, the underlying mechanisms and functional implications of the rs4273077 genetic variant in the development and metastasis of BC remain unclear. In addition, we discovered a significant association between the rs4730155 locus in the intronic region of NAMPT and susceptibility to BC recurrence and metastasis. The functional implications of NAMPT in tumor biology are well‐established, including its involvement in DNA repair, metastasis, angiogenesis, immune regulation, and drug resistance [26]. Nevertheless, there is a lack of research examining the relationship between NAMPT polymorphisms and the development and progression of BC. To the best of our knowledge, the present study was the first to establish a connection between the rs4730155 locus in NAMPT and the susceptibility to BC recurrence and metastasis. Further research should clarify how the rs4730155 locus affects NAMPT function and BC progression.

Our GWAS identified genes significantly associated with BC recurrence and metastasis based on the presence of specific SNPs. The GO and KEGG enrichment analyses revealed that the genes targeted by these SNPs were involved in biological processes related to glutamatergic synaptic transmission and the calcium and insulin signaling pathways. These findings are consistent with previous research, emphasizing the importance of these pathways in BC metastasis. Glutamatergic synaptic transmission is crucial for normal brain development and function [27]. Moreover, multiple studies have demonstrated that BC brain metastasis is caused by the formation of pseudo‐triple synapses between cancer cells and glutamatergic neurons [28]. The calcium signaling pathway regulates fundamental cellular processes such as cell proliferation, survival, apoptosis, and immunity; it is also associated with the development of various diseases, including cancer and autoimmune disorders [29]. Similarly, components of the insulin signaling pathway, such as PI3K and MAPK, are pivotal in the regulation of cell survival, growth, proliferation, and differentiation. Thus, the disruption of the insulin signaling pathway is also associated with cancer progression and metastasis [30, 31]. Hence, the present study provides novel evidence for the involvement of glutamatergic synaptic transmission, calcium signaling, and insulin signaling pathways in BC recurrence and metastasis. Understanding the mechanisms of these processes may have important implications for the development of targeted therapies and intervention strategies for BC.

This study identified four risk SNPs (rs10108514, rs12920540, rs4273077, and rs4730155) using GWAS, thus, providing valuable insights into the genetic variants associated with BC recurrence and metastasis. Our research lays the groundwork for further investigation and offers direction for future studies. However, there are several limitations to this study, including the small sample size, which may have produced false‐negative or false‐positive results. Moreover, genetic variations among different individuals may affect our predictions of SNP function. Environmental and genetic factors interact to influence an individual's phenotype. Given that our study focused solely on genetic factors while ignoring environmental influences, it may offer a simplistic view of SNP function in BC. Additionally, we acknowledge that we determined the functions of the identified SNPs through 3D SNP analysis alone, without subsequent experimental validation. In future studies, we will address these limitations by increasing the sample size, considering the role of genomic heterogeneity, and accounting for the contribution of environmental factors. We believe that this approach will help us better understand the relationship between human genetics and disease.

5. Conclusion

In this study, we used a two‐stage GWAS approach to demonstrate that four SNPs loci, namely, PVT1: rs10108514 (A>G), CASC16: rs12920540 (C>A), TNFRSF13B: rs4273077 (A>G), and NAMPT: rs4730155 (T>C), were associated with the risk of BC recurrence and metastasis. TCGA data were used to reveal a substantial correlation between TNFRSF13B expression and BC prognosis. The enrichment analysis of GWAS results revealed that the calcium signaling and insulin secretion pathways were potentially significant signaling pathways in BC recurrence and metastasis. Our study provides insights into the genetic basis of BC recurrence and metastasis, which may advance future research efforts and facilitate the development of new treatment strategies.

Author Contributions

Shujuan Sun: writing – original draft (equal). Sha Yin: writing – original draft (equal). Jie Huang: formal analysis (equal), writing – review and editing (equal). Dongdong Zhou: formal analysis (equal). Qiaorui Tan: data curation (equal). Xiaochu Man: data curation (equal). Wen Wang: Data curation (equal). Jiale Zhang: data curation (equal). Huihui Li: conceptualization (equal).

Ethics Statement

This study was approved by the Ethics Committee of Shandong Cancer Hospital and Institute (SDTHEC2022009018).

Consent

All patients provided written informed consent at the time of entering this study.

Conflicts of Interest

The authors declare no conflicts of interest.

Supporting information

Figure S1: The log‐rank test was used to determine the association between breast cancer recurrence and the different genotypes of rs2227251 (a), rs4730155 (b), or rs1799801 (c).

CAI2-4-e142-s001.tif (362KB, tif)

Acknowledgments

The authors are grateful to all our team members and the funding bodies who financially supported this study.

Shujuan Sun, Sha Yin, and Jie Huang contributed equally to this study.

Data Availability Statement

The authors have nothing to report.

References

  • 1. Siegel R. L., Giaquinto A. N., and Jemal A., “Cancer Statistics, 2024,” CA: A Cancer Journal for Clinicians 74, no. 1 (2024): 12–49, 10.3322/caac.21820. [DOI] [PubMed] [Google Scholar]
  • 2. Malmgren J. A., Mayer M., Atwood M. K., and Kaplan H. G., “Differential Presentation and Survival of de novo and Recurrent Metastatic Breast Cancer Over Time: 1990‐2010,” Breast Cancer Research and Treatment 167, no. 2 (2018): 579–590, 10.1007/s10549-017-4529-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Mariotto A. B., Etzioni R., Hurlbert M., Penberthy L., and Mayer M., “Estimation of the Number of Women Living With Metastatic Breast Cancer in the United States,” Cancer Epidemiology, Biomarkers & Prevention 26, no. 6 (2017): 809–815, 10.1158/1055-9965.EPI-16-0889. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Gennari A., Conte P., Rosso R., Orlandini C., and Bruzzi P., “Survival of Metastatic Breast Carcinoma Patients Over a 20‐Year Period: A Retrospective Analysis Based on Individual Patient Data From Six Consecutive Studies,” Cancer 104, no. 8 (2005): 1742–1750, 10.1002/cncr.21359. [DOI] [PubMed] [Google Scholar]
  • 5. Yan J., Liu Z., Du S., Li J., Ma L., and Li L., “Diagnosis and Treatment of Breast Cancer in the Precision Medicine Era,” Methods in Molecular Biology 2204 (2020): 53–61, 10.1007/978-1-0716-0904-0_5. [DOI] [PubMed] [Google Scholar]
  • 6. Si H., Esquivel M., Mendoza Mendoza E., and Roarty K., “The Covert Symphony: Cellular and Molecular Accomplices in Breast Cancer Metastasis,” Frontiers in Cell and Developmental Biology 11 (2023): 1221784, 10.3389/fcell.2023.1221784. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Kashyap D., Pal D., Sharma R., et al., “Global Increase in Breast Cancer Incidence: Risk Factors and Preventive Measures,” BioMed Research International 2022 (2022): 9605439, 10.1155/2022/9605439. [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  • 8. Smolarz B., Nowak A. Z., and Romanowicz H., “Breast Cancer‐Epidemiology, Classification, Pathogenesis and Treatment (Review of Literature),” Cancers 14, no. 10 (2022): 2569, 10.3390/cancers14102569. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Pérez‐Losada J., Castellanos‐Martín A., and Mao J. H., “Cancer Evolution and Individual Susceptibility,” Integrative Biology 3, no. 4 (2011): 316–328, 10.1039/C0IB00094A. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Chen Y., Shi C., and Guo Q., “TNRC9 rs12443621 and FGFR2 rs2981582 Polymorphisms and Breast Cancer Risk,” World journal of surgical oncology 14, no. 1 (2016): 50, 10.1186/s12957-016-0795-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Gudmundsdottir E. T., Barkardottir R. B., Arason A., et al., “The Risk Allele of SNP rs3803662 and the mRNA Level of its Closest Genes TOX3 and LOC643714 Predict Adverse Outcome for Breast Cancer Patients,” BMC Cancer 12 (2012): 621, 10.1186/1471-2407-12-621. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Bouhniz O. E., Zaied S., Naija L., et al., “Association Between HER2 and IL‐6 Genes Polymorphisms and Clinicopathological Characteristics of Breast Cancer: Significant Role of Genetic Variability in Specific Breast Cancer Subtype,” Clinical and Experimental Medicine 20, no. 3 (2020): 427–436, 10.1007/s10238-020-00632-5. [DOI] [PubMed] [Google Scholar]
  • 13. Miedl H., Oswald D., Haslinger I., et al., “Association of the Estrogen Receptor 1 Polymorphisms rs2046210 and rs9383590 With the Risk, Age at Onset and Prognosis of Breast Cancer,” Cells 12, no. 4 (2023): 515, 10.3390/cells12040515. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Cui P., Zhao Y., Chu X., et al., “SNP rs2071095 in LincRNA H19 Is Associated With Breast Cancer Risk,” Breast Cancer Research and Treatment 171, no. 1 (2018): 161–171, 10.1007/s10549-018-4814-y. [DOI] [PubMed] [Google Scholar]
  • 15. Couzin J. and Kaiser J., “Closing the Net on Common Disease Genes,” Science 316, no. 5826 (2007): 820–822, 10.1126/science.316.5826.820. [DOI] [PubMed] [Google Scholar]
  • 16. Shtivelman E., Henglein B., Groitl P., Lipp M., and Bishop J. M., “Identification of a Human Transcription Unit Affected by the Variant Chromosomal Translocations 2;8 and 8;22 of Burkitt Lymphoma,” Proceedings of the National Academy of Sciences USA 86, no. 9 (1989): 3257–3260, 10.1073/pnas.86.9.3257. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Graham M. and Adams J. M., “Chromosome 8 Breakpoint Far 3′ of the C‐Myc Oncogene in a Burkitt's Lymphoma 2;8 Variant Translocation Is Equivalent to the Murine pvt‐1 Locus,” EMBO Journal 5, no. 11 (1986): 2845–2851, 10.1002/j.1460-2075.1986.tb04578.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Kim M. Y., “Breast Cancer Metastasis,” Experimental Medicine and Biology 1187 (2021): 183–204, 10.1007/978-981-32-9620-6_9. [DOI] [PubMed] [Google Scholar]
  • 19. Coughlin S. S., “Epidemiology of Breast Cancer in Women,” Advances in Experimental Medicine and Biology 1152 (2019): 9–29, 10.1007/978-3-030-20301-6_2. [DOI] [PubMed] [Google Scholar]
  • 20. Malins D. C. and Haimanot R., “Major Alterations in the Nucleotide Structure of DNA in Cancer of the Female Breast,” Cancer Research 51, no. 19 (1991): 5430–5432. [PubMed] [Google Scholar]
  • 21. Zhang Z., Zhu Z., Zhang B., et al., “Frequent Mutation of rs13281615 and Its Association With PVT1 Expression and Cell Proliferation in Breast Cancer,” Journal of Genetics and Genomics 41, no. 4 (2014): 187–195, 10.1016/j.jgg.2014.03.006. [DOI] [PubMed] [Google Scholar]
  • 22. Lin H. Y., Callan C. Y., Fang Z., Tung H. Y., and Park J. Y., “Interactions of PVT1 and CASC11 on Prostate Cancer Risk in African Americans,” Cancer Epidemiology, Biomarkers & Prevention 28, no. 6 (2019): 1067–1075, 10.1158/1055-9965. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Xu W., Zhong Y., Yang H., Gong Y., Dao J., and Bao L., “Association Between the rs4784227‐CASC16 Polymorphism and the Risk of Breast Cancer: A Meta‐Analysis,” Medicine 101, no. 34 (2022): e30218, 10.1097/MD.0000000000030218. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Zuo X., Wang H., Mi Y., et al., “The Association of CASC16 Variants With Breast Cancer Risk in a Northwest Chinese Female Population,” Molecular Medicine 26, no. 1 (2020): 11, 10.1186/s10020-020-0137-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Hinterleitner C., Zhou Y., Tandler C., et al., “Platelet‐Expressed TNFRSF13B (TACI) Predicts Breast Cancer Progression,” Frontiers in Oncology 11 (2021): 642170, 10.3389/fonc.2021.642170. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Gasparrini M. and Audrito V., “NAMPT: A Critical Driver and Therapeutic Target for Cancer,” International Journal of Biochemistry & Cell Biology 145 (2022): 106189, 10.1016/j.biocel.2022.106189. [DOI] [PubMed] [Google Scholar]
  • 27. Martynyuk A. E., Glushakov A. V., Sumners C., Laipis P. J., Dennis D. M., and Seubert C. N., “Impaired Glutamatergic Synaptic Transmission in the PKU Brain,” Molecular Genetics and Metabolism 86, no. Suppl 1 (2005): 34–42, 10.1016/j.ymgme.2005.06.014. [DOI] [PubMed] [Google Scholar]
  • 28. Zeng Q., Michael I. P., Zhang P., et al., “Synaptic Proximity Enables NMDAR Signalling to Promote Brain Metastasis,” Nature 573, no. 7775 (2019): 526–531, 10.1038/s41586-019-1576-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Panda S., Chatterjee O., Roy L., and Chatterjee S., “Targeting Ca2+ signaling: A New Arsenal Against Cancer,” Drug Discovery Today 27, no. 3 (2022): 923–934, 10.1016/j.drudis.2021.11.012. [DOI] [PubMed] [Google Scholar]
  • 30. Yee L. D., Mortimer J. E., Natarajan R., Dietze E. C., and Seewaldt V. L., “Metabolic Health, Insulin, and Breast Cancer: Why Oncologists Should Care about Insulin,” Frontiers in Endocrinology 11 (2020): 58, 10.3389/fendo.2020.00058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Haeusler R. A., McGraw T. E., and Accili D., “Biochemical and Cellular Properties of Insulin Receptor Signalling,” Nature Reviews Molecular Cell Biology 19, no. 1 (2018): 31–44, 10.1038/nrm.2017.89. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1: The log‐rank test was used to determine the association between breast cancer recurrence and the different genotypes of rs2227251 (a), rs4730155 (b), or rs1799801 (c).

CAI2-4-e142-s001.tif (362KB, tif)

Data Availability Statement

The authors have nothing to report.


Articles from Cancer Innovation are provided here courtesy of John Wiley & Sons Ltd. on behalf of Tsinghua University Press

RESOURCES