Skip to main content
Breast Cancer : Targets and Therapy logoLink to Breast Cancer : Targets and Therapy
. 2026 Aug 27;18:611734. doi: 10.2147/BCTT.S611734

MST1 and CPNE1 Emerge as Candidate Targets in Breast Cancer: Integration of Mendelian Randomization, Multi-omics Validation, and Functional Interrogation

Meng Jiang 1,*, Qilong Wang 2,*, Hong Xu 1,✉
PMCID: PMC13528792  PMID: 42677314

Abstract

Background

Breast cancer remains a leading cause of female mortality worldwide, with therapeutic benefit limited by tumor heterogeneity and drug resistance. Identification of novel therapeutic targets through integration of genetic causality inference and functional validation is urgently needed.

Methods

Plasma protein quantitative trait loci (pQTL) from the Fenland cohort (~10,700 individuals) and the FinnGen R10 SomaScan subset (n = 828) were integrated with breast cancer genome-wide association study data from the Breast Cancer Association Consortium (122,977 cases and 105,974 controls). Causal protein-disease relationships were inferred using summary-data-based Mendelian randomization (SMR), with colocalization and HEIDI tests. Multi-level validation was performed using TCGA-BRCA transcriptomic data. Functional validation included CCK-8, EdU, wound healing, Transwell invasion, and flow cytometry apoptosis analysis evaluated CPNE1 knocdown and/or overexpression of the MST1 gene which encodes macrophage-simulating protein (MSP).

Results

SMR identified 23 proteins in Fenland and 10 in FinnGen at FDR < 0.05. MST1/MSP and CPNE1 were supported in both datasets with PP.H4 ≥ 0.80 and non-significant HEIDI tests. Genetically predicted circulating MSP was positively associated with breast cancer risk, whereas circulating CPNE1 showed an inverse association. TCGA-BRCA showed higher CPNE1 mRNA in tumors (P = 2.57×10⁻25) and lower MST1 mRNA (P = 3.04×10⁻23). CPNE1 expression was highest in triple-negative breast cancer and correlated positively with clinical stage (ρ = 0.211), whereas MST1 expression was lowest in triple-negative breast cancer and correlated inversely with stage (ρ = −0.164). In T47D and MDA-MB-231 cells, CPNE1 knockdown and MST1 overexpression each reduced proliferation, migration, and invasion and increased apoptosis; the combined group showed greater changes than either single intervention.

Conclusion

This study prioritizes the MST1 gene, which encodes MSP, and CPNE1 as candidate proteins for further investigation in breast cancer. However, the circulating-protein associations, tumor-mRNA patterns, and cell-autonomous perturbations represent distinct and directionally discordant biological contexts. The findings therefore support context-dependent candidate roles and justify mechanistic, in vivo, and formal interaction studies, but do not yet establish therapeutic efficacy or synergy.

Keywords: breast cancer, CPNE1, MST1, Mendelian randomization, therapeutic target, SMR

Introduction

Breast cancer (BC) remains the most common malignancy among women worldwide. In 2022, approximately 2.3 million new cases and 670,000 deaths were reported globally, and in some countries the annual incidence increased by up to 5%, underscoring its growing global health burden.1 BC is characterized by pronounced heterogeneity, with molecular subtypes including luminal, HER2-positive, and triple-negative breast cancer (TNBC) differing substantially in pathogenesis, therapeutic responsiveness, and prognosis. High-risk subtypes such as TNBC continue to exhibit poor five-year survival rates below 40%.2 Moreover, epidemiological evidence indicates a gradual shift toward earlier age at onset, with increasing proportions of young patients in certain regions, emphasizing the importance of early detection and subtype-specific management.3

Recent advances in proteomics offer a molecular framework for dissecting tumor heterogeneity, and BC cohorts have been profiled extensively by mass spectrometry. Conventional observational studies using laser capture microdissection to procure lesion-specific material, followed by protein quantification via label-free mass spectrometry or liquid chromatography-selected reaction monitoring, have nominated candidate biomarkers.4,5 Yet such associations often cannot establish causality, constraining clinical translation. To overcome this limitation, plasma proteomics has been integrated with Mendelian randomization and colocalization to prioritize putatively causal proteins and strengthen mechanistic inference.6

Mendelian randomization (MR), which uses germline variants as instrumental variables, has emerged as a robust approach for protein-disease causal inference by mitigating confounding and reverse causation.7–9 With the expansion of multi-omics resources, summary-data MR (SMR) combined with the HEIDI test has been widely employed to integrate protein quantitative trait loci (pQTL) with genome-wide association studies (GWAS).10,11 SMR identifies proteins whose genetically predicted levels are associated with disease, whereas the HEIDI test distinguishes true colocalization from linkage disequilibrium-driven signals.12 Despite these advances, a translational gap persists: few BC studies jointly apply multi-cohort pQTL integration, convergent causal frameworks, tissue-level expression validation, and wet-lab functional interrogation within one pipeline.13–16

A fundamental consideration in pQTL-based MR studies is the distinction between circulating protein abundance and tissue-level gene expression. Plasma proteins integrate systemic production, secretion, processing, distribution, and clearance, whereas tumor transcriptomes capture local gene expression; apparently discordant directions can therefore reflect different biological compartments.17,18 In this study, MST1 refers to the official human gene encoding macrophage-stimulating protein (MSP), also known as hepatocyte growth factor-like protein (HGFL).

To address the translational gap, we implemented a comprehensive framework integrating SMR, HEIDI testing, and Bayesian colocalization with multi-level validation. Following genetic nomination from plasma pQTL-GWAS integration, protein-protein interaction (PPI) networks were constructed and Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed.19,20 Phenome-wide association studies (PheWAS) were conducted to survey broader phenotype associations of the genetic instruments.21,22, Tissue-level expression analyses using TCGA-BRCA data examined differential expression, molecular subtype specificity, clinical stage correlations, and relationships with established proliferation markers. Finally, in vitro functional assays evaluated single and combined perturbations of the prioritized candidates. Together, this framework was used to assess MST1/MSP and CPNE1 across genetic association, tissue expression, and cellular perturbation, while recognizing that these evidence layers represent distinct biological contexts.

Methods and Materials

Data Sources

Protein quantitative trait locus (pQTL) data were obtained from the Fenland study, a University of Cambridge-led cohort of approximately 10,700 adults of European ancestry without chronic disease at baseline. Plasma concentrations of approximately 4,000 proteins were measured using the SomaScan high-throughput proteomic platform. A comprehensive pQTL analysis integrating genotype and proteomic data was reported by Pietzner et al,23 identifying 10,674 associations across 3,892 plasma proteins. FinnGen is a nationwide genotype-health-register resource.24 The publicly available FinnGen R10 SomaScan pQTL summary statistics used in this study were derived from 828 unrelated participants and included 7,596 protein assays.

Breast cancer GWAS summary statistics were obtained from the Breast Cancer Association Consortium (BCAC; IEU identifier: ieu-a-1126), comprising 122,977 BC cases and 105,974 controls with genome-wide data on approximately 15 million common variants.25,26

SMR Analysis and Colocalization

Associations between genetically predicted plasma protein levels and BC risk were evaluated using the summary-data-based Mendelian randomization (SMR) framework (v1.3.1). SMR P values were adjusted separately within each pQTL dataset using the Benjamini-Hochberg procedure, with FDR < 0.05 used as the screening threshold. Signals passing the FDR threshold were further examined using the heterogeneity in dependent instruments (HEIDI) test; P_HEIDI > 0.05 was interpreted as no evidence of heterogeneity attributable to linkage, rather than as proof of causality. Two-trait Bayesian colocalization was performed using the R package coloc (v5.2.3) within ±100 kb windows, with PP.H4 > 0.80 prespecified as strong evidence of colocalization.27

Pathway Enrichment and PheWAS

Protein-protein interaction networks were constructed using STRING (v12.0; https://string-db.org/) under a minimum interaction score threshold of 0.15 and visualized using Cytoscape software (v3.10.2; https://cytoscape.org/).28 GO and KEGG enrichment analyses were performed using the R package clusterProfiler (v4.16.0) with Benjamini-Hochberg correction (FDR < 0.05).29 PheWAS was applied to test associations between prioritized pQTLs and ICD-10-coded phenotypes using the IEU Open GWAS project (https://gwas.mrcieu.ac.uk/).20,22 P values from the PheWAS analyses were adjusted using the Benjamini-Hochberg procedure, with FDR < 0.05 used as the reporting threshold.

TCGA-BRCA Expression Analysis

Gene expression data from The Cancer Genome Atlas Breast Invasive Carcinoma cohort were downloaded from the UCSC Xena platform (https://xenabrowser.net/). Differential expression analysis compared CPNE1 and MST1 mRNA levels between tumor and paired normal tissue using 100 matched sample pairs with paired Wilcoxon signed-rank test. Expression patterns were stratified by molecular subtype (Luminal A, Luminal B, HER2-enriched, and TNBC) based on PAM50 classification, and differences among subtypes were assessed using Kruskal–Wallis test. Correlations between gene expression and clinical stage were assessed using Spearman rank correlation. Relationships between candidate genes and proliferation markers (MKI67 and PCNA) were evaluated using Spearman correlation analysis. Subtype stratification, clinical stage correlation, and proliferation marker analyses were performed using the full TCGA-BRCA cohort (n = 1,098). All statistical analyses were performed using R software (v4.2.0).

Cell Culture

Five cell lines were used in this study, including four human breast cancer cell lines representing different molecular subtypes (MDA-MB-231, triple-negative; MCF-7, Luminal A; T47D, Luminal A; SK-BR-3, HER2-positive) and one non-tumorigenic mammary epithelial cell line (MCF-10A) as normal control. All cell lines were purchased from Procell Life Science and Technology Co., Ltd (Wuhan, China). MDA-MB-231, MCF-7, T47D, and SK-BR-3 cells were cultured in high-glucose Dulbecco’s Modified Eagle’s Medium (DMEM; Procell) supplemented with 10% fetal bovine serum (FBS; Procell) and 1% penicillin-streptomycin (100 U/mL penicillin and 100 μg/mL streptomycin; Procell). MCF-10A cells were maintained in DMEM/F12 medium (Procell) supplemented with 5% horse serum (Gibco, USA), 20 ng/mL epidermal growth factor (Peprotech, USA), 0.5 μg/mL hydrocortisone (Sigma-Aldrich, USA), 100 ng/mL cholera toxin (Sigma-Aldrich), 10 μg/mL insulin (Sigma-Aldrich), and 1% penicillin-streptomycin. All cells were maintained at 37°C in a humidified atmosphere containing 5% CO2 and were routinely tested for mycoplasma contamination.

RNA Extraction and Quantitative Real-Time PCR

Total RNA was extracted from cells using TRIzol reagent (Invitrogen, Carlsbad, CA, USA) according to the manufacturer’s instructions. RNA concentration and purity were assessed using a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific, Waltham, MA, USA). Complementary DNA (cDNA) was synthesized from 1 μg of total RNA using the PrimeScript RT Reagent Kit with gDNA Eraser (Takara Bio, Kusatsu, Japan). Quantitative real-time PCR (qRT-PCR) was performed using TB Green Premix Ex Taq II (Takara Bio) on an ABI 7500 Real-Time PCR System (Applied Biosystems, Foster City, CA, USA). The primer sequences were as follows: CPNE1 forward 5’-TGAAGAGGAGGAGGAGGAGG-3’ and reverse 5’-CTCCAGCTTCTCCTTGGTGA-3’; MST1 forward’5′-GACCAGCCGCCATCAATC-3′ and reverse 5′-CTTGGAACGCCGCTGATC-3′; GAPDH forward 5’-GAAGGTGAAGGTCGGAGTC-3’ and reverse 5’-GAAGATGGTGATGGGATTTC-3’. The thermal cycling conditions were as follows: initial denaturation at 95°C for 30 seconds, followed by 40 cycles of denaturation at 95°C for 5 seconds and annealing/extension at 60°C for 34 seconds. Melting curve analysis was performed to confirm amplification specificity. Relative mRNA expression levels were calculated using the 2−ΔΔCt method with GAPDH as the internal reference gene. All reactions were performed in triplicate.

Western Blotting

Total protein was extracted from cells using RIPA lysis buffer (Beyotime Biotechnology, Shanghai, China) supplemented with phenylmethylsulfonyl fluoride (PMSF; Beyotime) and phosphatase inhibitor cocktail (Beyotime). Protein concentrations were determined using the BCA Protein Assay Kit (Beyotime). Equal amounts of protein (30 μg per lane) were separated by 10% sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) and transferred to polyvinylidene difluoride (PVDF) membranes (Millipore, Burlington, MA, USA). Membranes were blocked with 5% non-fat dry milk in Tris-buffered saline containing 0.1% Tween-20 (TBST) for 1 hour at room temperature. Membranes were then incubated overnight at 4°C with primary antibodies against CPNE1 (1:1,000, Cat. No. 12215-1-AP, Proteintech, Wuhan, China), MSP (the protein encoded by MST1; clone OTI4B12; 1:2,000; Cat. No. CF808369; Invitrogen/Thermo Fisher Scientific), and GAPDH (1:5,000, Cat. No. 60004-1-Ig, Proteintech). After washing three times with TBST, membranes were incubated with horseradish peroxidase (HRP)-conjugated secondary antibodies (1:5000, Proteintech) for 1 hour at room temperature. Protein bands were visualized using an enhanced chemiluminescence (ECL) detection kit (Beyotime) and imaged using a ChemiDoc XRS+ System (Bio-Rad, Hercules, CA, USA). Band intensities were quantified using ImageJ software (v1.53k, National Institutes of Health, Bethesda, MD, USA).

siRNA Transfection

Three small interfering RNA (siRNA) sequences targeting CPNE1 and a negative control siRNA (siNC) were designed and synthesized by GenePharma Co., Ltd (Shanghai, China). The siRNA sequences were as follows: siCPNE1-1 5’-GCUGGAAACUUCAGCAAAUTT-3’; siCPNE1-2 5’-GGACAAGAUCAACGAGAAUTT-3’; siCPNE1-3 5’-GCAGCUACAUCAAGACCAATT-3’. Cells were seeded in 6-well plates at a density of 2×105 cells per well and cultured until reaching 60–70% confluence. Transfection was performed using Lipofectamine 3000 (Invitrogen) according to the manufacturer’s protocol. Briefly, siRNA (50 nM final concentration) was diluted in Opti-MEM reduced serum medium (Gibco) and mixed with Lipofectamine 3000. The siRNA-lipid complexes were added to cells and incubated for 6 hours, after which the medium was replaced with complete growth medium. Knockdown efficiency was assessed by qRT-PCR and Western blotting at 48 hours post-transfection. The siRNA with the highest knockdown efficiency was selected for subsequent functional experiments.

Plasmid Transfection

The full-length human MST1 coding sequence (encoding MSP/HGFL) was cloned into the pCMV-3×Flag expression vector by Genechem Co., Ltd (Shanghai, China). The empty pCMV-3×Flag vector served as the negative control for overexpression (OE-NC). Cells were seeded in 6-well plates and cultured until reaching 70–80% confluence. Transfection was performed using Lipofectamine 3000 according to the manufacturer’s instructions. Cells were transfected with 2.5 μg of MST1 overexpression plasmid (MST1-OE) or empty vector control (OE-NC) per well. Overexpression efficiency was confirmed by qRT-PCR and Western blotting at 48 hours post-transfection.

For functional experiments, five experimental groups were established: Control (untreated), NC (siNC plus OE-NC), siCPNE1 (siCPNE1 plus OE-NC), MST1-OE (siNC plus MST1-OE plasmid), and Combined (siCPNE1 plus MST1-OE plasmid). Cells were co-transfected with the corresponding siRNA and plasmid simultaneously using Lipofectamine 3000.

Cell Proliferation Assay

Cell proliferation was measured using the Cell Counting Kit-8 (CCK-8; Biosharp, Hefei, China) according to the manufacturer’s instructions. Cells were seeded in 96-well plates at a density of 3 × 103 cells per well in 100 μL of complete medium. At designated time points (0, 24, 48, 72, and 96 hours), 10 μL of CCK-8 solution was added to each well and incubated for 2 hours at 37°C. Absorbance was measured at 450 nm using a Multiskan FC microplate reader (Thermo Fisher Scientific). Each experiment was performed with six replicates per group and repeated three times independently.

EdU Incorporation Assay

Cell proliferation was further assessed using the EdU Cell Proliferation Kit with Alexa Fluor 594 (RiboBio, Guangzhou, China). Cells were seeded in 96-well plates and treated as indicated. After 48 hours, cells were incubated with 50 μM 5-ethynyl-2’-deoxyuridine (EdU) for 2 hours at 37°C. Cells were then fixed with 4% paraformaldehyde for 30 minutes, permeabilized with 0.5% Triton X-100 for 10 minutes, and stained with Apollo reaction cocktail for 30 minutes in the dark. Nuclei were counterstained with Hoechst 33342 for 30 minutes. Images were captured using a fluorescence microscope (Olympus IX73, Tokyo, Japan) and the percentage of EdU-positive cells was calculated from at least five random fields per well using ImageJ software.

Wound Healing Assay

Cell migration was assessed using the wound healing assay. Cells were seeded in 6-well plates and cultured to 90–95% confluence. A sterile 200 μL pipette tip was used to create a straight scratch across the cell monolayer. Cells were washed twice with phosphate-buffered saline (PBS) to remove cell debris and then cultured in serum-free medium to minimize proliferation effects. Images were captured at 0 and 48 hours post-scratch using an inverted microscope (Olympus CKX53). Wound closure percentage was calculated using ImageJ software as follows: wound closure (%) = (wound area at 0 h − wound area at 48 h)/wound area at 0 h × 100%.

Transwell Invasion Assay

Cell invasion was evaluated using 24-well Transwell chambers with 8 μm pore size polycarbonate membranes (Corning, Corning, NY, USA). The upper surface of the membrane was pre-coated with Matrigel (BD Biosciences, San Jose, CA, USA) diluted 1:8 in serum-free medium. Cells (5 × 104) suspended in 200 μL of serum-free medium were seeded into the upper chamber, while 600 μL of medium containing 10% FBS was added to the lower chamber as a chemoattractant. After 48 hours of incubation at 37°C, non-invaded cells on the upper surface of the membrane were removed using cotton swabs. Invaded cells on the lower surface were fixed with 4% paraformaldehyde for 15 minutes and stained with 0.1% crystal violet for 20 minutes. Invaded cells were counted under a light microscope (Olympus CX23) in five random fields per chamber.

Flow Cytometry Apoptosis Analysis

Cell apoptosis was detected using the Annexin V-FITC/PI Apoptosis Detection Kit (BD Biosciences) according to the manufacturer’s protocol. Cells were harvested by trypsinization without EDTA, washed twice with cold PBS, and resuspended in 1× binding buffer at a concentration of 1 × 106 cells/mL. Cell suspension (100 μL) was incubated with 5 μL of Annexin V-FITC and 5 μL of propidium iodide (PI) for 15 minutes at room temperature in the dark. After adding 400 μL of binding buffer, samples were analyzed within 1 hour using a FACSCalibur flow cytometer (BD Biosciences). Data were analyzed using FlowJo software (v10.8.1, BD Biosciences). The apoptosis rate was calculated as the sum of early apoptotic (Annexin V+/PI−) and late apoptotic (Annexin V+/PI+) cells.

Statistical Analysis

-Statistical analyses were performed using R v4.2.0 and GraphPad Prism v9.0. Data are presented as mean ± SD from three independent experiments unless otherwise stated; technical replicates were averaged before analysis. Two-group comparisons used two-sided unpaired Student’s t-tests. For the CCK-8 assay, comparisons among groups at each time point were analyzed by one-way ANOVA followed by Tukey’s post hoc test. Other multi-group cell experiments were analyzed by one-way ANOVA followed by Tukey’s post hoc test. TCGA paired, subtype, and correlation analyses used the tests specified above. A two-sided P < 0.05 was considered statistically significant. Comparisons of the Combined group with each single-intervention group quantify greater combined effects but were not interpreted as formal synergy tests.

Results

SMR and Colocalization Analyses Prioritize MST1/MSP and CPNE1 as candidate therapeutic targets

SMR analysis identified 23 proteins associated with BC in the Fenland dataset and 10 proteins in the FinnGen dataset at FDR < 0.05 (Figure 1A, B and Supplementary Table 1). Among these, genetically predicted circulating MST1/MSP consistently showed positive associations with BC risk in both cohorts (Fenland: OR = 1.02, 95% CI 1.01–1.03, P_FDR = 8.33×10−3; FinnGen: OR = 1.03, 95% CI 1.01–1.04, P_FDR = 0.013), whereas CPNE1 showed inverse associations (Fenland: OR = 0.97, 95% CI 0.95–0.98, P_FDR = 8.33×10−3; FinnGen: OR = 0.97, 95% CI 0.95–0.98, P_FDR = 0.013) (Figure 1C and D).

Figure 1.

Five plots of SMR proteins linked to breast cancer risk: volcano, Manhattan, circos patterns. The figure consists of five panels labeled A to E, illustrating SMR analysis results for breast cancer risk. Panel A: A volcano plot with x-axis ′betaSMA′ (effect size) and y-axis ′-log10(pvaladj)′ (adjusted P-value). Proteins like TLR1 show strong positive associations. Panel B: Another volcano plot for the FinnGen cohort, highlighting GSTM1 and MST1. Panel C: A Manhattan plot for the Fenland cohort with x-axis ′Chr′ (chromosome) and y-axis ′-log10(p)′ (raw P-value), showing peaks for SEMA4A and ANXA4. Panel D: A Manhattan plot for the FinnGen cohort, highlighting GSTM1 and CPNE1. Panel E: A circos plot showing chromosomal distribution of candidate proteins, with outer tracks labeling gene names and inner rings representing genomic features. The legend indicates significance levels from 1 to 13. Color intensity and point size in plots indicate significance and effect size, respectively. Dashed lines represent significance thresholds.

SMR analysis identifies candidate proteins associated with breast cancer risk. (A) Volcano plot displaying SMR-identified proteins based on SMR effect size (beta_SMR) and adjusted P-value (−log10 scale). Red dots represent proteins positively associated with breast cancer risk (eg, TLR1), with color intensity indicating −log10(P_adj). (B) Volcano plot showing SMR-identified proteins in the FinnGen cohort, with GSTM1 and MST1 highlighted. (C) Manhattan plot of SMR association results in the Fenland cohort, plotting −log10 (P) against chromosomal position; key proteins (eg, SEMA4A, ANXA4) are labeled. (D) Manhattan plot of SMR association results in the FinnGen cohort, highlighting GSTM1 and CPNE1. (E) Circos plot illustrating chromosomal distribution of candidate proteins, with outer tracks labeling gene names and inner rings representing genomic features. MST1 in the figure denotes the pQTL annotation for MSP. SMR P-values were adjusted using the Benjamini-Hochberg method; proteins with FDR < 0.05 and P_HEIDI > 0.05 were considered significant.

HEIDI testing indicated no evidence of LD-induced heterogeneity (P_HEIDI > 0.05) for 18 of 23 proteins in Fenland and 9 of 10 proteins in FinnGen (Supplementary Table 1).

Proteins were stratified into three evidence tiers based on colocalization (PP.H4) and HEIDI results (Figure 1E). Tier 1 consisted of MST1/MSP and CPNE1, each demonstrating strong colocalization in both datasets (PP.H4 ≥ 0.80) and no LD-induced heterogeneity (P_HEIDI > 0.05). Tier 2 included A4GALT, BTN3A3, and NAGLU with partial replication. Tier 3 included proteins with single-cohort evidence only. Considering all criteria, MST1/MSP and CPNE1 showed the strongest genetic support as candidate proteins for further investigation.

MST1/MSP- and CPNE1-associated interaction networks show distinct enrichment profiles

PPI network analysis showed that MST1/MSP was connected to extracellular and plasma-protein partners, including HGFAC, PLG, F2, SERPINC1, VTN, albumin, and apolipoproteins (Figure 2A). GO enrichment highlighted wound healing, blood coagulation, and hemostasis (Figure 2B), while KEGG analysis showed enrichment in complement and coagulation cascades, PI3K-AKT, MAPK, Rap1, FoxO, and other signaling pathways (Figure 2C). CPNE1 was positioned within an actin-cytoskeleton network (Figure 2D), with GO terms dominated by actin organization and DNA-damage checkpoint regulation (Figure 2E), and KEGG enrichment in DNA replication and actin cytoskeleton regulation (Figure 2F).

Figure 2.

Six panels show PPI networks and enrichment analyses for MST1 and CPNE1-related genes. The image A shows a PPI network centered on MST1, with nodes labeled APOH, FGA, SERPINF2, APOA1, AHS3, PROC, SERPINC1, F2, LPA, VTN, PLG, ALB, APOA2, HGFAC and AMBP. The image B shows a bar plot of GO biological process enrichment for MST1-related genes, listing pathways like wound healing, blood coagulation and regulation of response to wounding. The image C shows a bubble plot of KEGG pathway enrichment for MST1-related genes, including PI3K-Akt signaling and MAPK signaling pathways. The image D shows a PPI network centered on CPNE1, with nodes labeled ACTB, AKT1, ACTR1B, CORO1B, ANXA1, CDC42, ACTG1, CAPZB, YWHAG, TPM3, CAPZA2, PFN1, POLM, WDR1 and CAPZ. The image E shows a bar plot of GO biological process enrichment for CPNE1-related genes, including actin filament organization and regulation of DNA damage checkpoint. The image F shows a bubble plot of KEGG pathway enrichment for CPNE1-related genes, including Salmonella infection and DNA replication.

Functional enrichment and protein-protein interaction (PPI) network analysis of candidate targets. (A) PPI network centered on MST1/MSP; node size and color denote interaction degree. (B) Bar plot of Gene Ontology (GO) biological process enrichment for MST1/MSP-related genes, showing top pathways including wound healing and blood coagulation. (C) Bubble plot of Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment for MST1/MSP-related genes; dot size represents gene count, and color indicates Benjamini-Hochberg adjusted P-value. (D) PPI network centered on CPNE1, highlighting its interacting partners. (E) Bar plot of GO biological process enrichment for CPNE1-related genes, including actin filament organization. (F) Bubble plot of KEGG pathway enrichment for CPNE1-related genes, including the Salmonella infection pathway. PPI networks were constructed using STRING (v12.0) with a minimum interaction score of 0.15 and visualized using Cytoscape (v3.10.2). Enrichment analyses were performed using clusterProfiler (v4.16.0) with FDR < 0.05.

PheWAS Reveals Exploratory Phenotype Associations

PheWAS revealed that the MST1/MSP pQTL instrument was positively associated with gastrointestinal inflammatory phenotypes including ulcerative colitis (OR = 1.000555, P_FDR = 1.18×10−4) and inversely associated with cardiometabolic phenotypes including hypertension (OR = 0.997607, P_FDR = 0.014) and myocardial infarction (OR = 0.999548, P_FDR = 0.015) (Supplementary Table 2). CPNE1 showed positive association with breast cancer (C50.9, OR = 1.001872, P_FDR = 1.08×10−4) and inverse association with facial skin malignancy (C44.3, OR = 0.998569, P_FDR = 7.34×10−3) (Supplementary Table 3). Given the small effect estimates and broad phenotype screen, these associations should be interpreted as exploratory and not as direct evidence of therapeutic efficacy or safety.

CPNE1 is Upregulated and MST1 is Downregulated in Breast Cancer Tissues

Analysis of TCGA-BRCA paired samples (n = 100) revealed that CPNE1 mRNA was significantly elevated in tumor compared to paired normal tissue (P = 2.57×10−25), while MST1 showed significantly reduced expression in tumors (P = 3.04×10−23) (Figure 3A and B). When stratified by molecular subtype, CPNE1 expression was significantly higher in TNBC compared to other subtypes (Kruskal–Wallis P = 4.19×10−20), while MST1 was significantly lower in TNBC (P = 2.39×10−4) (Figure 3C and D). Clinical stage correlation analysis revealed that CPNE1 positively correlated with stage (Spearman ρ = 0.211, P = 1.57×10−11), while MST1 showed significant inverse correlation (ρ = −0.164, P = 1.77×10−7) (Figure 3E and F).

Figure 3.

Graphs of CPNE1 & MST1 mRNA in breast cancer by sample type, subtype, stage and proliferation markers. Box plots compare CPNE1 and MST1 mRNA expression log2(TPM+1) between Normal and Tumor samples using 100 paired TCGA-BRCA samples. CPNE1 has a higher median in Tumor (P = 2.57e-25), while MST1 is lower in Tumor (P = 3.04e-23). Violin plots show expression across four subtypes: Luminal A, Luminal B, HER2+ and TNBC. CPNE1 peaks in TNBC (Kruskal-Wallis P = 4.19e-20) and MST1 is lowest in TNBC (P = 2.39e-4). Another set of violin plots shows expression across cancer stages I-IV. CPNE1 rises with stage (Spearman rho = 0.211, P = 1.57e-11), while MST1 declines (rho = -0.164, P = 1.77e-7). Scatter plots display CPNE1 expression against MKI67, PCNA and MST1. Spearman correlations: CPNE1 vs MKI67 = 0.821, CPNE1 vs PCNA = 0.802 and CPNE1 vs MST1 = -0.072. Distinct markers differentiate datasets.

Expression profiles of CPNE1 and MST1 in breast cancer clinical cohorts. (A) Box plot comparing CPNE1 mRNA expression between tumor and adjacent normal tissues in TCGA-BRCA paired samples (n = 100 pairs; paired Wilcoxon signed-rank test, P = 2.57×10−25). (B) Box plot of MST1 mRNA expression in tumor vs paired normal tissues (n = 100 pairs; paired Wilcoxon signed-rank test, P = 3.04×10−23). (C) Violin plot showing CPNE1 expression across breast cancer molecular subtypes (Luminal A, Luminal B, HER2-enriched, and TNBC) based on PAM50 classification (Kruskal–Wallis test, P = 4.19×10−20). (D) Violin plot of MST1 expression across molecular subtypes (Kruskal–Wallis test, P = 2.39×10−4). (E) Violin plot of CPNE1 expression across clinical stages (Spearman ρ = 0.211, P = 1.57×10−11). (F) Violin plot of MST1 expression across clinical stages (Spearman ρ = −0.164, P = 1.77×10−7). (G) Scatter plots showing Spearman correlations between CPNE1 and proliferation markers MKI67 (ρ = 0.821, P < 2.2×10−16), PCNA (ρ = 0.802, P < 2.2×10−16), and MST1 (ρ = −0.072, P = 2.35×10−2). Subtype, stage, and correlation analyses were performed using the full TCGA-BRCA cohort (n = 1,098).

CPNE1 showed strong positive correlations with proliferation markers MKI67 (ρ = 0.821, P < 2.2×10−16) and PCNA (ρ = 0.802, P < 2.2×10−16), while CPNE1 and MST1 showed modest inverse correlation (ρ = −0.072, P = 2.35×10−2) (Figure 3G). These results characterize the tissue-level expression patterns of CPNE1 and MST1 in breast cancer.

CPNE1 and MST1 Show Reciprocal Expression Patterns in Breast Cancer Cell Lines

To validate the expression patterns of CPNE1 and MST1 in breast cancer cell lines, we examined their mRNA and protein levels across five cell lines: four breast cancer cell lines representing different molecular subtypes (MDA-MB-231, triple-negative; MCF-7, Luminal A; T47D, Luminal A; SK-BR-3, HER2-positive) and one non-tumorigenic mammary epithelial cell line (MCF-10A) serving as normal control. qPCR and Western blot analyses confirmed that CPNE1 mRNA and protein were significantly upregulated in all four BC cell lines compared to MCF-10A, with the highest expression observed in MDA-MB-231.Conversely, MST1 mRNA and MSP protein were significantly downregulated in all four BC cell lines, with the lowest expression in MDA-MB-231 (Figure 4A–E). This expression pattern is consistent with the TCGA-BRCA finding that TNBC displays the highest CPNE1 and lowest MST1 expression among molecular subtypes. Based on these expression profiles, two cell lines were selected for subsequent functional studies: MDA-MB-231, which exhibited the most pronounced differential expression pattern (highest CPNE1 and lowest MST1) and represents the triple-negative subtype with the poorest clinical prognosis and greatest unmet therapeutic need; and T47D, representing the luminal subtype, to validate the generalizability of functional effects across molecular subtypes.

Figure 4.

Bar charts and blots showing CPNE1 and MST1 expression in breast cancer cell lines and transfections. Image A: CPNE1 mRNA expression increases across cell lines, peaking in MDA-MB-231. Image B: MST1 mRNA expression decreases, lowest in MDA-MB-231. Image C: Western blot shows stronger CPNE1 and weaker MST1 in cancer lines; GAPDH is the control. Image D: CPNE1 protein expression highest in MDA-MB-231. Image E: MST1 protein expression lowest in MDA-MB-231. Image F: CPNE1 mRNA reduced after siRNA knockdown. Image G: Weaker CPNE1 bands post-knockdown. Image H: CPNE1 protein reduced post-knockdown. Image I: MST1 mRNA increases after overexpression. Image J: Stronger MST1 bands post-overexpression. Image K: MST1 protein increases post-overexpression. Bars are color-coded, error bars show mean ± SD, significance marked by asterisks.

Validation of CPNE1 and MST1 expression in breast cancer cell lines and transfection efficiency. (A) CPNE1 mRNA expression in MCF-10A (normal mammary epithelial), MDA-MB-231 (triple-negative), MCF-7 (Luminal A), T47D (Luminal A), and SK-BR-3 (HER2-positive) cells measured by qRT-PCR. (B) MST1 mRNA expression in the same cell lines by qRT-PCR. (C) Representative Western blot showing CPNE1 and MSPprotein expression across cell lines; GAPDH served as a loading control. (D) Densitometric quantification of CPNE1 protein expression from (C), normalized to GAPDH. (E) Densitometric quantification of MSPprotein expression from (C), normalized to GAPDH. (F) CPNE1 mRNA expression after transfection with three independent siRNA sequences (siCPNE1-1, siCPNE1-2, siCPNE1-3) or negative control siRNA (siNC) in MDA-MB-231 and T47D cells. (G) Representative Western blot confirming CPNE1 knockdown. (H) Densitometric quantification of CPNE1 protein expression after knockdown. (I) MST1 mRNA expression after transfection with MST1 overexpression plasmid (MST1-OE) or empty vector control (OE-NC). (J) Representative Western blot confirming increased MSP protein after MST1 overexpression. (K) Densitometric quantification of MSP protein expression after overexpression. Data are presented as mean ± SD from three independent experiments. For panels (A)–(E) ***P < 0.001 vs MCF-10A group. For panels (F)–(H) ***P < 0.001 vs siNC group. For panels (I)–(K) **P < 0.01, ***P < 0.001 vs OE-NC group.

Among the three siRNA sequences tested, siCPNE1-2 achieved the highest knockdown efficiency and was selected for subsequent experiments (Figure 4F–H). MST1 overexpression plasmid transfection significantly increased MST1 mRNA and MSP protein levels (Figure 4I–K).

Combined CPNE1 Knockdown and MST1 Overexpression Produces Greater Phenotypic Changes Than Either Intervention Alone

To evaluate the functional roles of CPNE1 and MST1 in breast cancer cells, five experimental groups were established: Control (untreated), NC (siNC plus OE-NC), siCPNE1 (siCPNE1 plus OE-NC), MST1-OE (siNC plus MST1-OE plasmid), and Combined (siCPNE1 plus MST1-OE plasmid).

In T47D cells, CCK-8 assays demonstrated that both siCPNE1 and MST1-OE groups exhibited significantly reduced proliferation compared to the NC group, while the Combined group showed the most pronounced inhibition of cell growth (Figure 5A). EdU incorporation assays confirmed these findings, with the Combined group displaying the lowest percentage of EdU-positive cells (Figure 5B and C). Wound healing assays revealed that siCPNE1, MST1-OE, and Combined groups all exhibited significantly reduced migration capacity compared to NC, with the Combined group showing the greatest impairment in wound closure (Figure 5D–E). Transwell invasion assays demonstrated similar patterns, with the Combined group exhibiting the fewest invaded cells (Figure 5F and G). Flow cytometry analysis showed that apoptosis rates were significantly elevated in siCPNE1 and MST1-OE groups compared to NC, and the Combined group displayed the highest apoptosis rate among all treatment groups (Figure 5H and I).

Figure 5.

An infographic of T47D assays comparing Control, NC, siCPNE1, MST1-OE and Combined outcomes. A scientific infographic details T47D cell phenotypes in five groups: Control, NC, siCPNE1, MST1-OE and Combined. A CCK-8 graph shows cell viability over time, with Control and NC highest at 96h, siCPNE1 and MST1-OE lower and Combined lowest. EdU assay micrographs are labeled DAPI, EdU and Merge. A bar chart displays the percentage of EdU-positive cells. Wound healing assay micrographs are shown at 0h and 24h. Another bar chart illustrates migration rates. Transwell invasion assay micrographs are provided, with a bar chart showing migration rates. Annexin V-FITC vs PI flow cytometry plots reveal quadrant values: Control 1.27, 1.20, 1.48, 96.0; NC 1.54, 2.79, 1.41, 94.3; siCPNE1 0.76, 9.68, 4.58, 85.0; MST1-OE 0.62, 8.07, 5.75, 85.6; Combined 1.21, 14.2, 13.2, 71.4. A bar chart shows apoptosis rates. siCPNE1, MST1-OE and Combined exhibit lower viability, EdU, migration and Transwell migration than NC, with higher apoptosis, especially in Combined.

Effects of CPNE1 knockdown and MST1 overexpression on T47D cell phenotypes. Five experimental groups were established: Control (untreated), NC (siNC + OE-NC), siCPNE1 (siCPNE1 + OE-NC), MST1-OE (siNC + MST1-OE plasmid), and Combined (siCPNE1 + MST1-OE plasmid). (A) Cell viability assessed by CCK-8 assay at 0, 24, 48, 72, and 96 hours post-transfection, showing reduced proliferation in siCPNE1, MST1-OE, and Combined groups compared to NC. (B) Representative fluorescence images of EdU incorporation (red) with Hoechst 33342 nuclear counterstaining (blue) at 48 hours. Scale bar: 100 μm. (C) Quantification of EdU-positive cell percentage. (D) Representative images of wound healing assay at 0 and 48 hours post-scratch. Scale bar: 200 μm. (E) Quantification of wound closure rate at 48 hours. (F) Representative images of Transwell invasion assay (crystal violet staining) at 48 hours. Scale bar: 100 μm. (G) Quantification of invaded cell numbers per field. (H) Representative flow cytometry dot plots of Annexin V-FITC/PI staining showing apoptotic cell populations. (I) Quantification of total apoptosis rate (early + late apoptosis). Data are presented as mean ± SD from three independent experiments (n = 3). For panel A, groups were compared separately at each time point using one-way ANOVA followed by Tukey’s post hoc test. For panels C, E, G, and I, statistical analysis was performed using one-way ANOVA followed by Tukey’s post hoc test. *P < 0.05, **P < 0.01, ***P < 0.001 vs NC group; #P < 0.05, ##P < 0.01, ###P < 0.001, Combined vs siCPNE1 group; &P < 0.05, &&P < 0.01, &&&P < 0.001, Combined vs MST1-OE group.

Similar results were observed in MDA-MB-231 cells (Figure 6A–I). Across both cell lines, the Combined group showed greater phenotypic changes than either single-intervention group in the reported comparisons. These findings indicate an enhanced combined phenotype under the tested conditions but do not establish formal synergy or therapeutic efficacy.

Figure 6.

A multi-panel infographic of MDA-MB-231 assays for Control, NC, siCPNE1, MST1-OE, Combined. The images illustrate assays on MDA-MB-231 cells. Image A shows a graph of cell viability, with Control and NC higher than siCPNE1, MST1-OE and Combined. Image B′s EdU assay micrographs reveal reduced signals in siCPNE1, MST1-OE and Combined versus Control and NC. Image C′s bar chart shows fewer EdU-positive cells in siCPNE1, MST1-OE and Combined. Image D′s wound healing micrographs indicate more gap closure in Control and NC. Image E′s bar chart shows higher migration rates for Control and NC. Image F′s Transwell micrographs have fewer stained cells in siCPNE1, MST1-OE and Combined. Image G′s bar chart shows higher Transwell migration rates for Control and NC. Image H′s flow cytometry plots reveal increased apoptosis in siCPNE1, MST1-OE and Combined. Image I′s bar chart shows higher apoptosis rates in siCPNE1, MST1-OE and Combined. Significance is marked on bars.

Effects of CPNE1 knockdown and MST1 overexpression on MDA-MB-231 cell phenotypes. Five experimental groups were established as described in Figure 5. (A) Cell viability assessed by CCK-8 assay at 0, 24, 48, 72, and 96 hours post-transfection, showing reduced proliferation in siCPNE1, MST1-OE, and Combined groups compared to NC. (B) Representative fluorescence images of EdU incorporation (red) with Hoechst 33342 nuclear counterstaining (blue) at 48 hours. Scale bar: 100 μm. (C) Quantification of EdU-positive cell percentage. (D) Representative images of wound healing assay at 0 and 48 hours post-scratch. Scale bar: 200 μm. (E) Quantification of wound closure rate at 48 hours. (F) Representative images of Transwell invasion assay (crystal violet staining) at 48 hours. Scale bar: 100 μm. (G) Quantification of invaded cell numbers per field. (H) Representative flow cytometry dot plots of Annexin V-FITC/PI staining showing apoptotic cell populations. (I) Quantification of total apoptosis rate (early + late apoptosis). Data are presented as mean ± SD from three independent experiments (n = 3). For panel A, groups were compared separately at each time point using one-way ANOVA followed by Tukey’s post hoc test. For panels C, E, G, and I, statistical analysis was performed using one-way ANOVA followed by Tukey’s post hoc test. *P < 0.05, **P < 0.01, ***P < 0.001 vs NC group; #P < 0.05, ##P < 0.01, ###P < 0.001, Combined vs siCPNE1 group; &P < 0.05, &&P < 0.01, &&&P < 0.001, Combined vs MST1-OE group.

Discussion

This study presents an integrated framework combining Mendelian randomization, tissue-level expression analysis, and comprehensive functional characterization to investigate MST1 and CPNE1 as candidate therapeutic targets in breast cancer. Through evidence spanning genetic association and colocalization, TCGA-based expression analysis, and cellular functional interrogation, our findings support further investigation of these candidates, while the compartment-dependent results preclude a simple therapeutic interpretation.

The replication of MST1/MSP and CPNE1 associations across two independent pQTL datasets (Fenland and FinnGen), combined with strong colocalization signals (PP.H4 ≥ 0.80) and non-significant HEIDI tests, strengthens genetic support for these associations. Traditional biomarker discovery approaches, while generating numerous candidate proteins, are often limited by confounding and reverse causation, leading to poor reproducibility and translational failure.30 The genetic approach reduces some confounding and reverse-causation concerns by using germline variants as instrumental variables, but it does not by itself prove causality or therapeutic direction.

A notable aspect of our findings is the apparent discordance between plasma pQTL-based association estimates and tissue-level expression patterns. Our SMR analysis showed that genetically predicted higher circulating MST1/MSP was associated with higher BC risk, whereas higher circulating CPNE1 was inversely associated with BC risk. Intriguingly, TCGA-BRCA analysis revealed that at the tissue level, CPNE1 is markedly upregulated (P = 2.57×10−25) while MST1 mRNA is significantly downregulated (P = 3.04×10−23) in tumors. This observation, rather than representing a contradiction, reflects the complex biology underlying plasma protein levels versus tissue expression. Plasma protein concentrations represent an integration of multiple biological processes including tissue production rates, secretion dynamics, systemic clearance, and distribution volumes, which may differ substantially from local tumor microenvironment expression.17,18 For CPNE1, the inverse circulating-protein association and higher tumor expression may reflect distinct systemic and tumor-local biology, although the underlying mechanisms remain unclear. Cancer cells may hijack the CPNE1-dependent signaling machinery to activate PI3K-AKT pathways and drive proliferation, independent of systemic protein dynamics.31 MST1 encodes the secreted protein MSP/HGFL.32 MSP signals primarily through the RON/MST1R receptor, and its biological output depends on ligand processing, receptor availability, and cellular context.33,34 Accordingly, the present evidence should be interpreted as context-dependent rather than as a directionally convergent target-validation chain.

The tissue-level analysis in TCGA-BRCA provides biological insights into the relevance of these candidates. The observation that CPNE1 expression is highest in TNBC—the most aggressive BC subtype with limited therapeutic options—and positively correlated with clinical stage (ρ = 0.211, P = 1.57×10−11) suggests that CPNE1 may be particularly relevant for patients with advanced or aggressive disease. The strong correlations between CPNE1 and established proliferation markers MKI67 (ρ = 0.821) and PCNA (ρ = 0.802) further support an association between CPNE1 expression and proliferative activity. These findings are consistent with recent reports demonstrating that CPNE1 upregulation promotes tumorigenesis and radioresistance in TNBC through activation of the AKT signaling pathway,35 and that CPNE1 enhances cancer progression in HER2-positive and Luminal BC subtypes.36 The Ca2⁺-dependent membrane-binding properties of CPNE1, mediated by its tandem C2 domains, position it as a scaffold protein linking receptor trafficking and membrane dynamics to downstream kinase activation,37 providing a mechanistic basis for its oncogenic functions. Notably, the expression pattern observed in breast cancer cell lines—with MDA-MB-231 (TNBC) exhibiting the highest CPNE1 and lowest MST1 mRNA and MSP protein levels—recapitulates the TCGA tissue-level findings, reinforcing the relevance of these candidates, particularly in the triple-negative subtype.

The MST1/MSP findings require especially cautious interpretation. In the present experiments, MST1 overexpression reduced proliferation, migration, and invasion and increased apoptosis. By contrast, prior breast cancer studies often report tumor-promoting MSP-RON activity: MSP-pathway activation was associated with metastasis and poor prognosis.38 MSP stimulated invasion in RON-positive MDA-MB-231 cells39; in murine models, global HGFL/MSP loss impaired mammary tumor growth and metastasis and altered immune-cell recruitment32; and HGFL-dependent RON activation promoted PCNA-linked proliferation.40 This directional discrepancy is biologically important. Our experiments measured MST1 mRNA and cell-associated MSP after plasmid overexpression, but did not quantify extracellular pro-MSP, proteolytically activated MSP, or RON signaling. Differences in precursor processing, secretion, receptor abundance or isoforms, and cell state could contribute, but remain untested. The data therefore support a context-dependent MST1/MSP phenotype but do not establish the underlying signaling mechanism.

A notable finding of this study is the observation of enhanced cellular phenotypic effects from combined CPNE1 knockdown and MST1 overexpression. The combined intervention produced significantly greater effects than either single intervention across all phenotypic readouts. This provides evidence of an enhanced combined phenotype under the tested conditions, but it does not establish molecular convergence or synergy. Formal interaction assessment requires an appropriate dose-response matrix and a prespecified interaction model, such as Bliss independence or another response-surface-based approach.41,42 Because MSP processing/RON signaling and the relevant CPNE1 downstream pathway were not measured, no shared signaling mechanism is claimed. The combination data instead motivate future interaction and rescue experiments.

The exploratory PheWAS analysis identified phenotype associations that may inform follow-up studies. The MST1/MSP pQTL instrument showed positive associations with gastrointestinal inflammatory phenotypes (ulcerative colitis, rectal polyps) and inverse associations with cardiometabolic phenotypes (hypertension, myocardial infarction). Because the effect estimates were close to unity and the analyses used inherited proxies rather than a defined intervention, these associations should not be interpreted as predicted toxicity, cardiovascular benefit, or treatment safety.

Several limitations should be acknowledged. First, our genetic analyses are based predominantly on European-ancestry populations, and validation in diverse populations with ancestry-aware fine-mapping is needed to assess generalizability. Second, pQTL instruments reflect plasma protein abundance, which as discussed may differ substantially from tumor microenvironment regulation; integration with tumor tissue QTL resources and single-cell data would enhance resolution. Third, the cell experiments measured MST1 mRNA and cell-associated MSP but did not quantify secreted pro-MSP, activated MSP, or RON/MST1R signaling. Fourth, while our in vitro studies demonstrate cellular effects, in vivo validation is required to determine whether these findings extend to a more physiologically relevant context. Fifth, the precise molecular mechanisms underlying the combined CPNE1 knockdown and MST1 overexpression effects warrant further investigation through pathway-rescue experiments, phosphoproteomic profiling, and analysis of downstream transcriptional programs. Sixth, although the combined intervention produced significantly greater cellular phenotypic effects than either single intervention alone, formal synergy assessment using established methods such as the Bliss independence model or combination index analysis would be needed to distinguish true synergistic from additive interactions. Finally, development of pharmacological modulators specifically targeting CPNE1 or modulating tumor-cell MST1 remains an important translational challenge.

Conclusions

Integrated genetic, transcriptomic, and cellular analyses prioritize CPNE1 and the MST1 gene (encoding MSP/HGFL) for further investigation in breast cancer. CPNE1 knockdown and MST1 overexpression each produced anti-tumor phenotypes in two cell lines, and the combined intervention produced greater changes than either single intervention under the tested conditions. However, the positive association of genetically predicted circulating MSP with breast cancer risk, lower tumor MST1 mRNA, and the cell-autonomous overexpression results represent distinct biological contexts. The findings therefore support context-dependent candidate roles and justify mechanistic and in vivo follow-up, but do not yet establish either protein, the direction of its modulation, or their combination as a therapeutic strategy.

Funding Statement

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Data Sharing Statement

The datasets analyzed during the current study are publicly available. The GWAS and pQTL data are publicly accessible through the IEU OpenGWAS project (https://gwas.mrcieu.ac.uk/), FinnGen (https://www.finngen.fi/en), and published literature, while gene expression data are available from the UCSC Xena platform (https://xenabrowser.net/). Specifically, breast cancer GWAS summary statistics were obtained from the Breast Cancer Association Consortium via the IEU OpenGWAS project (identifier: ieu-a-1126);25,26 plasma pQTL data were obtained from the Fenland study as reported by Pietzner et al23 and from the FinnGen DF10 public release;24 and TCGA-BRCA gene expression data were downloaded from the UCSC Xena platform. Additional experimental data and analysis codes that support the findings of this study are available from the corresponding author upon reasonable request.

Ethics Approval and Consent to Participate

This study utilized publicly available GWAS summary statistics and gene expression data from the Breast Cancer Association Consortium, Fenland study, FinnGen study, and TCGA-BRCA cohort. All these studies received ethical approval from their respective institutional review boards, with informed consent obtained from the participants. The data from the Fenland study are publicly accessible, and our use of this data complies with all relevant guidelines. The experimental validation using commercially available cell lines does not require ethical approval, as no human subjects, tissues, or animals were involved. This study was reviewed by the Medical Ethics Committee of the First Affiliated Hospital of Soochow University, which determined that it qualifies for exemption from ethical review (granted 26 June 2026), as the analyses relied exclusively on publicly available, de-identified datasets and commercially available cell lines and did not involve patient-derived clinical samples, individual-level patient data, human tissue, or animals. This exemption is consistent with the Ethical Review Measures for Life Sciences and Medicine Research Involving Humans (2023, China).

Author Contributions

Meng Jiang: Conceptualization, Methodology, Formal analysis, Writing – original draft. Qilong Wang: Methodology, Investigation, Data curation, Formal analysis, Writing – original draft. Hong Xu: Conceptualization, Supervision, Project administration, Writing – review & editing. All authors made a significant contribution to the work reported, whether that is in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; have agreed on the journal to which the article has been submitted; and agree to be accountable for all aspects of the work.

Disclosure

Meng Jiang and Qilong Wang are co-first authors for this study. The authors declare that there are no conflicts of interest regarding this study.

References

  • 1.Bray F, Laversanne M, Sung H. et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74:229–18. doi: 10.3322/caac.21834 [DOI] [PubMed] [Google Scholar]
  • 2.Caswell-Jin JL, Sun LP, Munoz D, et al. Analysis of Breast Cancer Mortality in the US-1975 to 2019. Jama. 2024;331:233–241. doi: 10.1001/jama.2023.25881 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Giaquinto AN, Sung H, Newman LA, et al. Breast cancer statistics 2024. CA Cancer J Clin. 2024;74(6):477–495. doi: 10.3322/caac.21863 [DOI] [PubMed] [Google Scholar]
  • 4.Peila R, Rohan TE. Circulating levels of biomarkers and risk of ductal carcinoma in situ of the breast in the UK Biobank study. Int J Cancer. 2024;154(7):1191–1203. doi: 10.1002/ijc.34795 [DOI] [PubMed] [Google Scholar]
  • 5.Yang Y, Zhao D, Luo J, et al. Quantitative Site-Specific Glycoproteomics Reveals Glyco-Signatures for Breast Cancer Diagnosis. Anal Chem. 2025;97(1):114–121. doi: 10.1021/acs.analchem.4c03069 [DOI] [PubMed] [Google Scholar]
  • 6.Song J, Yang H. Identifying new biomarkers and potential therapeutic targets for breast cancer through the integration of human plasma proteomics: a Mendelian randomization study and colocalization analysis. Front Endocrinol. 2024;15:1449668. doi: 10.3389/fendo.2024.1449668 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Burgess S, Scott RA, Timpson NJ, et al. Using published data in Mendelian randomization: a blueprint for efficient identification of causal risk factors. Eur J Epidemiol. 2015;30(7):543–552. doi: 10.1007/s10654-015-0011-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Davey Smith G, Hemani G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Hum Mol Genet. 2014;23(R1):R89–98. doi: 10.1093/hmg/ddu328 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Lawlor DA, Harbord RM, Sterne JAC, et al. Mendelian randomization: using genes as instruments for making causal inferences in epidemiology. Stat Med. 2008;27(8):1133–1163. doi: 10.1002/sim.3034 [DOI] [PubMed] [Google Scholar]
  • 10.Zhu Z, Zhang F, Hu H, et al. Integration of summary data from GWAS and eQTL studies predicts complex trait gene targets. Nat Genet. 2016;48(5):481–487. doi: 10.1038/ng.3538 [DOI] [PubMed] [Google Scholar]
  • 11.Xu F, Yu EY, Cai X, et al. Genome-wide genotype-serum proteome mapping provides insights into the cross-ancestry differences in cardiometabolic disease susceptibility. Nat Commun. 2023;14(1):896. doi: 10.1038/s41467-023-36491-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Suhre K, Arnold M, Bhagwat AM, et al. Connecting genetic risk to disease end points through the human blood plasma proteome. Nat Commun. 2017;8(1):14357. doi: 10.1038/ncomms14357 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Cano-Gamez E, Trynka G. From GWAS to Function: using Functional Genomics to Identify the Mechanisms Underlying Complex Diseases. Front Genet. 2020;11:424. doi: 10.3389/fgene.2020.00424 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Zuber V, Grinberg NF, Gill D, et al. Combining evidence from Mendelian randomization and colocalization: review and comparison of approaches. Am J Hum Genet. 2022;109(5):767–782. doi: 10.1016/j.ajhg.2022.04.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Li Y, Dou Y, Da Veiga Leprevost F, et al. Proteogenomic data and resources for pan-cancer analysis. Cancer Cell. 2023;41(8):1397–1406. doi: 10.1016/j.ccell.2023.06.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Sun BB, Chiou J, Traylor M, et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature. 2023;622(7982):329–338. doi: 10.1038/s41586-023-06592-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Pietzner M, Wheeler E, Carrasco-Zanini J, et al. Synergistic insights into human health from aptamer- and antibody-based proteomic profiling. Nat Commun. 2021;12(1):6822. doi: 10.1038/s41467-021-27164-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Ferkingstad E, Sulem P, Atlason BA, et al. Large-scale integration of the plasma proteome with genetics and disease. Nat Genet. 2021;53(12):1712–1721. doi: 10.1038/s41588-021-00978-w [DOI] [PubMed] [Google Scholar]
  • 19.Gene Ontology Consortium. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Research. 2021;49(D1):D325–D334. doi: 10.1093/nar/gkaa1113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Kanehisa M, Furumichi M, Sato Y. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 2023;51(D1):D587–d592. doi: 10.1093/nar/gkac963 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Denny JC, Bastarache L, Roden DM. Phenome-Wide Association Studies as aTool to Advance Precision Medicine. Annu Rev Genomics Hum Genet. 2016;17(D1):353–373. doi: 10.1146/annurev-genom-090314-024956 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Wu P, Gifford A, Meng X, et al. Mapping ICD-10 and ICD-10-CM Codes to Phecodes: workflow Development and Initial Evaluation. JMIR Med Inform. 2019;7(4):e14325. doi: 10.2196/14325 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Pietzner M, Wheeler E, Carrasco-Zanini J, et al. Mapping the proteo-genomic convergence of human diseases. Science. 2021;374(6569):eabj1541. doi: 10.1126/science.abj1541 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Kurki MI, Karjalainen J, Palta P, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613(7944):508–518. doi: 10.1038/s41586-022-05473-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Lyon MS, Andrews SJ, Elsworth B, et al. The variant call format provides efficient and robust storage of GWAS summary statistics. Genome Biol. 2021;22(1):32. doi: 10.1186/s13059-020-02248-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Michailidou K, Lindström S, Dennis J, et al. Association analysis identifies 65 new breast cancer risk loci. Nature. 2017;551(7678):92–94. doi: 10.1038/nature24284 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Foley CN, Staley JR, Breen PG, et al. A fast and efficient colocalization algorithm for identifying shared genetic risk factors across multiple traits. Nat Commun. 2021;12(1):764. doi: 10.1038/s41467-020-20885-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Szklarczyk D, Kirsch R, Koutrouli M, et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51:D638–d646. doi: 10.1093/nar/gkac1000 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Xu S, Hu E, Cai Y, et al. Using clusterProfiler to characterize multiomics data. Nat Protoc. 2024;19(11):3292–3320. doi: 10.1038/s41596-024-01020-z [DOI] [PubMed] [Google Scholar]
  • 30.Schmidt AF, Finan C, Gordillo-Marañón M, et al. Genetic drug target validation using Mendelian randomisation. Nat Commun. 2020;11(1):3255. doi: 10.1038/s41467-020-16969-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Cao J, Cao R, Liu Y, et al. CPNE1 mediates glycolysis and metastasis of breast cancer through activation of PI3K/AKT/HIF-1α signaling. Pathol Res Pract. 2023;248:154634. doi: 10.1016/j.prp.2023.154634 [DOI] [PubMed] [Google Scholar]
  • 32.Hunt BG, Jones A, Lester C, Davis JC, Benight NM and Waltz SE. RON (MST1R) and HGFL (MST1) Co-Overexpression Supports Breast Tumorigenesis through Autocrine and Paracrine Cellular Crosstalk. Cancers. 2022;14(10):2493. doi: 10.3390/cancers14102493 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Huang L, Fang X, Shi D, et al. MSP-RON Pathway: potential Regulator of Inflammation and Innate Immunity. Front Immunol. 2020;11:569082. doi: 10.3389/fimmu.2020.569082 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Cazes A, Childers BG, Esparza E, et al. The MST1R/RON Tyrosine Kinase in Cancer: oncogenic Functions and Therapeutic Strategies. Cancers. 2022;14(8):2037. doi: 10.3390/cancers14082037 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Shao Z, Ma X, Zhang Y, et al. CPNE1 predicts poor prognosis and promotes tumorigenesis and radioresistance via the AKT singling pathway in triple-negative breast cancer. Mol Carcinog. 2020;59:533–544. doi: 10.1002/mc.23177 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Choi HY, Lee HJ, Moon KM, et al. Up-regulation of CPNE1 Appears to Enhance Cancer Progression in HER2-positive and Luminal A Breast Cancer Cells. Anticancer Res. 2022;42(7):3445–3452. doi: 10.21873/anticanres.15831 [DOI] [PubMed] [Google Scholar]
  • 37.Tang H, Pang P, Qin Z, et al. The CPNE Family and Their Role in Cancers. Front Genet. 2021;12:689097. doi: 10.3389/fgene.2021.689097 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Welm AL, Sneddon JB, Taylor C, et al. The macrophage-stimulating protein pathway promotes metastasis in a mouse model for breast cancer and predicts poor prognosis in humans. Proc. Natl. Acad. Sci. U.S.A. 2007;104(18):7570–7575. doi: 10.1073/pnas.0702095104 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Thangasamy A, Rogge J and Ammanamanchi S. Regulation of RON Tyrosine Kinase-mediated Invasion of Breast Cancer Cells. Journal of Biological Chemistry. 2008;283(9):5335–5343. doi: 10.1074/jbc.M706957200 [DOI] [PubMed] [Google Scholar]
  • 40.Zhao H, Chen M, Lo Y, et al.The Ron receptor tyrosine kinase activates c-Abl to promote cell proliferation through tyrosine phosphorylation of PCNA in breast cancer. Oncogene. 2014;33(11):1429–1437. doi: 10.1038/onc.2013.84 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Boumendil L, Fontaine M, Lévy V, et al. Drug combinations screening using a Bayesian ranking approach based on dose–response models. Biometrical J. 2024;66(1):e2200332. doi: 10.1002/bimj.202200332 [DOI] [PubMed] [Google Scholar]
  • 42.Xu S, Esmaeili S, Cardozo-Ojeda EF, et al. Two-way pharmacodynamic modeling of drug combinations and its application to pairs of repurposed Ebola and SARS-CoV-2 agents. Antimicrob Agents Chemother. 2024;68(4):e0101523. doi: 10.1128/aac.01015-23 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The datasets analyzed during the current study are publicly available. The GWAS and pQTL data are publicly accessible through the IEU OpenGWAS project (https://gwas.mrcieu.ac.uk/), FinnGen (https://www.finngen.fi/en), and published literature, while gene expression data are available from the UCSC Xena platform (https://xenabrowser.net/). Specifically, breast cancer GWAS summary statistics were obtained from the Breast Cancer Association Consortium via the IEU OpenGWAS project (identifier: ieu-a-1126);25,26 plasma pQTL data were obtained from the Fenland study as reported by Pietzner et al23 and from the FinnGen DF10 public release;24 and TCGA-BRCA gene expression data were downloaded from the UCSC Xena platform. Additional experimental data and analysis codes that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Breast Cancer : Targets and Therapy are provided here courtesy of Dove Press

RESOURCES