Skip to main content
Medicine logoLink to Medicine
. 2026 May 22;105(21):e48953. doi: 10.1097/MD.0000000000048953

Causal effects of plasma protein–protein ratios on lung cancer subtypes: A proteome-wide Mendelian randomization and mediation analysis

Xin Geng a, Jingyuan Xue a, Kaisen Liu a, Zipei Song a, Yuheng Wang b, Xincen Cao a, Zhihua Li a, Liang Chen a,*
PMCID: PMC13200959  PMID: 42175454

Abstract

Lung cancer remains the leading cause of cancer-related mortality worldwide, with major subtypes including lung adenocarcinoma, lung squamous cell carcinoma (LUSC), and small cell lung cancer, yet the causal role of plasma protein–protein ratios in subtype-specific development remains poorly understood. We performed a bidirectional, two-sample Mendelian randomization (MR) analysis to assess the causal effects of plasma protein–protein ratios on lung cancer. Genetic instruments for protein–protein ratios were derived from a large-scale proteomic genome-wide association study, and lung cancer summary statistics were obtained from the FinnGen and Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium databases. Our study identified 17 plasma protein–protein ratios associated with lung cancer across both databases, including 5 for overall lung cancer (LC), 4 for lung adenocarcinoma, 5 for LUSC, and 4 for small cell lung cancer. Notably, a higher melanoma inhibitory activity (MIA)/DAN family BMP antagonist ratio was associated with increased risk for both LC and LUSC. When analyzing the constituent proteins of these ratios individually, only the causal association between MIA and LC remained significant. Subsequent 2-step MR analysis revealed that the MIA/DAN family BMP antagonist ratio mediated the effect of body mass index on LUSC development, with consistent mediation proportions across both databases (13.46%, 95% confidence interval: 3.14%–23.77% in FinnGen; 9.26%, 95% confidence interval: 0.05%–18.47% in Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium). Additional exploratory mediation pathways were identified for smoking initiation and body mass index; however, these findings should be interpreted with caution, given the assumptions underlying 2-step MR and the exploratory nature of the analysis. Multiple sensitivity analyses confirmed the robustness of these findings. This MR study offers novel proteomic insights into lung cancer pathogenesis and identifies candidate protein–protein ratios that warrant further functional investigation to assess their therapeutic potential.

Keywords: lung cancer, mediation analysis, Mendelian randomization, MIA/NBL1, protein–protein ratios

1. Introduction

Lung cancer remains one of the most prevalent and lethal malignancies worldwide, ranking as the leading cause of cancer-related mortality with an estimated 2.5 million new cases and 1.8 million deaths in 2022.[1] The most prevalent histological subtype is lung adenocarcinoma (LUAD, 45%), followed by lung squamous cell carcinoma (LUSC, 20%), small cell lung cancer (SCLC, 11%), and large cell lung carcinoma (7%).[2] Despite therapeutic advances, the 5-year relative survival for lung cancer in the United States remains low at about 23%, partly because many cases are diagnosed at advanced stages.[3] The pathogenesis of lung cancer involves a complex interplay of genetic susceptibility, environmental exposures such as air pollution and occupational hazards, and lifestyle factors, most notably tobacco use.[4] Within this intricate network, proteins play a pivotal role in oncogenesis.[5,6] Revolutionary advancements in proteomics have paved the way for large-scale circulating proteomic studies, offering deeper insights into the molecular basis of cancer while enabling the discovery of novel biomarkers and therapeutic targets. Recent observational studies have revealed associations between plasma proteins and lung cancer risk, underscoring the proteome’s influence on tumor development and offering clues to underlying pathways.[7,8]

Compared with conventional analyses focused on protein quantitative trait loci (pQTLs), ratio quantitative trait loci (rQTLs) offer enhanced insights by revealing the presence of shared genetic variance or nongenetic regulatory processes between proteins that may better reflect biological pathways relevant to disease.[9] However, existing research remains largely limited to a handful of studies examining specific protein ratios for prognostic evaluation, leaving a substantial knowledge gap regarding comprehensive proteomic ratio profiling and their etiological significance.[10,11] Moreover, traditional observational studies are inherently constrained by potential residual confounding and reverse causation, making it challenging to establish definitive causal relationships.

Mendelian randomization (MR) emerges as a powerful epidemiological approach that harnesses genetic variants as instrumental variables (IVs) to infer causal relationships between exposures and outcomes.[12] According to Mendelian genetics, alleles undergo random allocation during meiosis at conception. This process of natural randomization effectively minimizes confounding and avoids reverse causation.[13] Therefore, this approach enables robust causal inference when randomized controlled trials are not feasible, addressing fundamental limitations inherent in observational research.

In this study, we hypothesized that genetically predicted plasma protein–protein ratios exert causal effects on lung cancer and offer distinct etiologic insights beyond individual proteins. Accordingly, our primary analysis employed a bidirectional, two-sample MR to systematically investigate the causal effects of 2821 protein–protein ratios on overall lung cancer (LC) and its major histological subtypes across discovery and replication datasets. As a secondary analysis, we performed single-protein MR and colocalization analyses for the identified ratios to determine whether the observed associations were driven by the shared variance of the ratios or by their individual constituent proteins. Finally, considering the well-established associations of smoking and body mass index (BMI) with lung cancer,[14–16] we conducted an exploratory 2-step MR analysis to assess whether the identified protein–protein ratios act as mediators in these causal pathways. Our findings aim to provide novel mechanistic insights into the proteomic landscape of lung cancer causation, highlighting candidate biological pathways that warrant further functional validation and clinical investigation. We present this article in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology using-MR reporting checklist.

2. Methods

2.1. Study design

As depicted in Figure 1, we applied a bidirectional two-sample univariable MR (UVMR) to examine causal associations between 2821 plasma protein–protein ratios and lung cancer, including LC and its major subtypes: LUAD, LUSC, and SCLC. Our MR analysis relies on 3 core assumptions[17]: genetic variants are strongly associated with the exposure (relevance); genetic variants are independent of confounders (independence); and genetic variants affect the outcome only through the exposure pathway (exclusion restriction). The FinnGen database served as the discovery dataset, with the Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium (TRICL-ILCCO) used for replication validation. To isolate genuine causal effects, we performed a union analysis of results from both databases, incorporating Bayesian weighted MR (BWMR) and comprehensive sensitivity analyses. Intersection analysis then identified protein–protein ratios with robust and consistent associations across both databases. Subsequently, we extracted plasma proteins corresponding to these ratios for MR and colocalization analyses to examine their causal relationships with lung cancer. Reverse MR was performed to address potential reverse causation. Finally, we conducted a 2-step MR analysis to evaluate the mediating role of the identified protein–protein ratios in the causal pathways from smoking initiation and BMI to lung cancer.

Figure 1.

Figure 1.

Overview of the Mendelian randomization (MR) study design. BMI = body mass index, BWMR = Bayesian weighted Mendelian randomization, COPD = chronic obstructive pulmonary disease, GORD = gastroesophageal reflux disease, IV = instrumental variable, IVW = inverse-variance weighted, LC = overall lung cancer, LD = linkage disequilibrium, LUAD = lung adenocarcinoma, LUSC = lung squamous cell carcinoma, MAF = minor allele frequency, MR-PRESSO = Mendelian Randomization Pleiotropy Residual Sum and Outlier, PLIN3 = perilipin 3, pQTL = protein quantitative trait locus, rQTL = ratio quantitative trait locus, SCLC = small cell lung cancer, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium, TYMP = thymidine, UVMR = univariable Mendelian randomization.

2.2. Data sources

Genetic data on rQTLs were obtained from a genome-wide association study (GWAS) by Suhre et al, which employed Olink proteomics data for 1463 proteins in over 54,000 UK Biobank participants and identified 4248 associations across 2821 unique protein pairs.[9] We sourced pQTL data from the UK Biobank Pharma Proteomics Project reported by Sun et al.[18] Summary statistics for smoking initiation came from the GWAS and Sequencing Consortium of Alcohol and Nicotine use, with 311,629 cases and 321,173 controls.[19] BMI genetic associations were drawn from the Genetic Investigation of Anthropometric Traits consortium, involving 681,275 individuals.[20] For lung cancer outcomes, GWAS summary statistics were acquired from 2 large databases. The primary set originated from the FinnGen consortium (release R10), comprising 6340 cases and 314,193 controls for LC, along with subtypes: LUAD (1590 cases), LUSC (1510 cases), and SCLC (717 cases), each with 314,193 controls.[21,22] Replication relied on the TRICL-ILCCO meta-analysis, including 29,266 cases and 56,450 controls for LC, with subtypes: LUAD (11,273 cases, 55,483 controls), LUSC (7426 cases, 55,627 controls), and SCLC (2664 cases, 21,444 controls).[23] All analyses utilized publicly available GWAS summary statistics from European-ancestry populations, with details summarized in Table 1.

Table 1.

Data sources for the Mendelian randomization analysis in the study.

Phenotype Data source nCase nControl Phenotypic code Ancestry PMID
Exposure (two-sample MR)
 rQTLs Suhre et al[9] 54,306 – GCST90313126 to GCST90315946 European 38412862
 pQTLs Sun et al[18] 54,306 – – European 37794186
Exposure (two-step MR)
 Smoking initiation GSCAN 311,629 321,173 ieu-b-4877 European 30643251
 BMI GIANT 681,275 – ieu-b-40 European 30124842
Outcome (discovery)
 Overall lung cancer FinnGen 6340 314,193 C3_BRONCHUS_LUNG_EXALLC European –
 Lung adenocarcinoma FinnGen 1590 314,193 C3_NSCLC_ADENO_EXALLC European –
 Squamous cell lung carcinoma FinnGen 1510 314,193 C3_NSCLC_SQUAM_EXALLC European –
 Small cell lung cancer FinnGen 717 314,193 C3_SCLC_EXALLC European –
Outcome (replication)
 Overall lung cancer TRICL-ILCCO 29,266 56,450 GCST004748 European 28604730
 Lung adenocarcinoma TRICL-ILCCO 11,273 55,483 GCST004744 European 28604730
 Squamous cell lung carcinoma TRICL-ILCCO 7426 55,627 GCST004750 European 28604730
 Small cell lung cancer TRICL-ILCCO 2664 21,444 GCST004746 European 28604730

BMI = body mass index, GIANT = Genetic Investigation of Anthropometric Traits, GSCAN = GWAS and Sequencing Consortium of Alcohol and Nicotine use, MR = Mendelian randomization, NSCLC = non-small cell lung cancer, pQTLs = protein quantitative trait loci, rQTLs = ratio quantitative trait loci, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium.

In accordance with Strengthening the Reporting of Observational Studies in Epidemiology using-MR guidelines, we assessed potential sample overlap. The exposure GWAS for both protein–protein ratios and single proteins were derived exclusively from the UK Biobank.[9,18] The lung cancer outcome data were obtained from FinnGen, an independent Finnish cohort,[21] and the TRICL-ILCCO consortium, an international meta-analysis independent of the UK Biobank.[23] Therefore, no sample overlap exists between the exposure and outcome datasets in the primary analyses. For the subsequent mediation analyses, the GWAS datasets for smoking initiation and BMI both incorporate UK Biobank participants,[19,20] introducing potential overlap with the exposure data. However, previous methodological research demonstrates that two-sample MR remains highly robust to sample overlap when genetic instruments are strong.[24] Because we strictly retained instruments with F-statistics exceeding 10, the risk of overlap-induced weak instrument bias is minimal.

2.3. IV selection

Genetic instruments were identified through a stepwise process. First, we applied a significance threshold of P < 1 × 10−5 for rQTLs to ensure sufficient single-nucleotide polymorphisms (SNPs),[25] while using the genome-wide threshold (P < 5 × 10−8) for pQTLs. Next, linkage disequilibrium clumping was performed based on the European 1000 Genomes reference panel, selecting independent variants with r2 < 0.001 within 10,000 kb. To ensure accurate causal estimates, we harmonized alleles and excluded palindromic SNPs with ambiguous strand alignment. We also removed variants with minor allele frequency <0.01 to maintain statistical power. For phenotype scanning, we eliminated variants showing pleiotropic associations with secondary traits (including smoking,[26] BMI,[16] chronic obstructive pulmonary disease,[27] gastroesophageal reflux disease,[28] and educational attainment[29]) or those directly associated with the outcomes, by querying the GWAS Catalog (https://www.ebi.ac.uk/gwas/) at P < 1 × 10−5. Finally, we evaluated instrument strength using F-statistics and retained only SNPs with F > 10 to minimize weak instrument bias.[30]

2.4. UVMR and BWMR analysis

We utilized the random-effects inverse-variance weighted (IVW) method as the main approach for causal estimates in both FinnGen and TRICL-ILCCO databases.[31] The IVW method assumes all genetic variants are valid instruments, which provides optimal statistical power by calculating a weighted mean of individual variant effects.[32] To enhance robustness of the IVW results, we conducted supplementary MR analyses, including MR-Egger, weighted median, weighted mode, and simple mode. In the union analysis of findings from both databases, we further employed BWMR for additional validation. This approach refines causal estimates by effectively managing potential biases from weak instruments and identifying pleiotropic outliers.[33] Results were reported as odds ratios (ORs) with 95% confidence intervals (CIs) per 1-standard deviation increase in the genetically predicted exposure.

2.5. Multiple-testing correction and candidate selection

To account for multiple hypothesis testing across the 2821 protein–protein ratios, we applied the Benjamini–Hochberg false discovery rate (FDR) correction. However, because strict multiplicity control in large-scale proteomics can substantially increase false-negative rates and restrict discovery power, we adopted stringent cross-database replication as our primary empirical safeguard against false positives. According to our study design, candidate selection was executed through a rigorous 3-step statistical framework. First, initial screening within each database required nominal statistical significance in the IVW method (unadjusted P < .05); consistent effect directions across all applied MR methods; and no evidence of horizontal pleiotropy. Second, candidates passing these criteria were pooled into union sets and subjected to reevaluation, where BWMR was required to validate the IVW estimates, and MR Pleiotropy Residual Sum and Outlier (MR-PRESSO), alongside Radial MR, were applied to confirm the absence of pleiotropic outliers. Finally, intersection analysis was performed to strictly retain only those ratios demonstrating concordant and significant associations across both the FinnGen and TRICL-ILCCO databases.

2.6. Bayesian colocalization analysis

For plasma proteins identified as associated with lung cancer in the preceding MR analysis, we further performed colocalization analysis using the “coloc” R package.[34] This method evaluates the posterior probability of a shared causal variant between plasma protein and lung cancer, resulting in 5 mutually exclusive hypotheses: H0, no association with either trait; H1, association with plasma protein only; H2, association with lung cancer only; H3, association with both traits driven by distinct causal variants; and H4, association with both traits mediated by a shared causal variant. SNPs within a 1000 kb window (±500 kb) centered on the lead variant were extracted for colocalization analysis. We defined strong evidence of colocalization as a posterior probability for H4 >0.80 and moderate evidence as 0.5 < posterior probability for H4 ≤ 0.80.[35]

2.7. Reverse MR analysis

In this study, reverse MR analysis was performed, treating lung cancer as the exposure and its associated protein–protein ratios or proteins as the outcomes. To prevent erroneous interpretation of the results, any associations showing evidence of reverse causation were excluded from further analysis.

2.8. Mediation analysis

To elucidate the potential mediating role of the identified protein–protein ratios in the causal pathway linking smoking initiation and BMI to lung cancer, we employed a 2-step MR analysis. This approach yields unbiased estimates of mediation when the genetic instruments for the exposure and mediator are strong and satisfy the exclusion restriction criterion.[36] To strengthen the plausibility of these conditions, we performed a comprehensive phenotype scan during IV selection, removing variants with pleiotropic associations to secondary traits, including smoking and BMI (all at P < 1 × 10−5). Given that all utilized instruments possessed F-statistics exceeding 10, and multiple sensitivity analyses showed no evidence of horizontal pleiotropy, 2-step MR is considered methodologically appropriate for estimating mediation.[36] In the first step, a UVMR analysis was conducted to evaluate the effect (β1) of exp on med, and in the second step, we derived the effect (β2) of the corresponding med on out from our previous bidirectional MR screening.[36] The SNPs employed in the second step were rigorously examined to be nonoverlapping with those used in the first step.[37] The total effect (β) of the exp on out was decomposed into direct and indirect effects. We quantified the indirect effect using the product of coefficients method (β1 × β2) and calculated the mediation proportion by dividing the indirect effect by the total effect.[36] The 95% CIs were estimated using the delta method. Mediators exhibiting directionally consistent indirect and total effects were subsequently classified into 2 evidence levels. Potential evidence of mediation required a complete triangular causal relationship with significant effects across all 3 pathways. This was upgraded to strong evidence when the 95% CI for the indirect effect excluded 0. We performed mediation analyses in both FinnGen and TRICL-ILCCO databases to assess the robustness and consistency of the observed effects.

2.9. Sensitivity analysis

We employed a suite of sensitivity analyses to detect and mitigate potential biases from pleiotropy, heterogeneity, and the undue influence of individual SNPs. Horizontal pleiotropy arises when IVs affect the outcome through pathways independent of the exposure, thereby biasing causal inferences. To address this concern, we first performed MR-Egger regression to evaluate horizontal pleiotropy, where the intercept term quantifies the average pleiotropic effect across SNPs.[38] A nonzero intercept (P < .05) indicates the presence of horizontal pleiotropy. To further probe for pleiotropic effects, the MR-PRESSO method was also applied. This method identifies and excludes influential outliers to ensure the reliability of the causal estimates.[39] Specifically, if the global test P-value remained below .05 after MR-PRESSO correction, we implemented Radial MR analysis in both IVW and MR-Egger estimates.[40] This approach utilizes weighted residuals from the Radial plot to identify influential variants potentially harboring pleiotropic effects.[41] Outliers were sequentially removed until the global test indicated no significant pleiotropy (P > .05). The refined estimates were then verified for consistency with the primary analysis. Additionally, the Cochran Q test was employed to evaluate heterogeneity among the IVs.[42] A P-value < .05 was considered indicative of significant heterogeneity, which refers to substantial variation in causal effect estimates across different IVs. Furthermore, to ensure that our results were not driven by any single genetic variant, we conducted a leave-one-out sensitivity analysis by iteratively excluding each SNP to evaluate whether the overall MR estimates were disproportionately influenced by any individual variant.[43]

All MR analyses were conducted using R software (version 4.3.2; R Foundation for Statistical Computing, Vienna, Austria) with several packages, including “TwoSampleMR” (version 0.6.0),[44] “MendelianRandomization” (version 0.10.0),[45] “BWMR” (version 0.1.1),[33] “MR-PRESSO” (version 1.0),[39] and “RadialMR” (version 1.1).[41]

3. Results

3.1. Genetic instruments

Following the predefined criteria for IV selection, we identified genetic instruments for each exposure trait, with full details provided in Tables S1 and S2, Supplemental Digital Content. The number of SNPs for rQTLs ranged from 6 to 97 (median = 30), while pQTLs yielded 5 to 22 SNPs (median = 11). Additionally, 217 and 458 independent variants were selected for smoking initiation and BMI, respectively. When lung cancer was treated as the exposure, the SNP counts varied from 64 to 153 (median = 122.5). The proportion of variance in the exposures explained by individual variants spanned 0.00003 to 0.331. All instruments demonstrated strong statistical power, with F-statistics exceeding 10, thereby minimizing the risk of weak instrument bias in subsequent analyses.

3.2. Causal effects of plasma protein–protein ratios on lung cancer

We performed systematic UVMR analyses to investigate the causal effects of 2821 protein–protein ratios on lung cancer and its major subtypes, in both the FinnGen and TRICL-ILCCO databases. The complete multistep analytical workflow, detailing the progressive filtering steps and the precise counts of retained candidate ratios at each stage, is explicitly illustrated in Figure 2. After applying the Benjamini–Hochberg FDR correction, no protein–protein ratios reached the strict statistical significance threshold of FDR-adjusted q < 0.05 (Table S3, Supplemental Digital Content). Nevertheless, several candidate ratios demonstrated unadjusted low P-values that warrant further investigation.[46] Therefore, to balance discovery power with false-positive control, we selected candidate protein–protein ratios based on the following strict criteria: nominal statistical significance in the IVW method (unadjusted P < .05); consistent effect directions across all MR methods; and no evidence of horizontal pleiotropy as assessed by MR-Egger regression. This initial screening yielded 135, 114, 117, and 127 protein–protein ratios associated with LC, LUAD, LUSC, and SCLC, respectively, in the FinnGen database, while the TRICL-ILCCO database showed 123, 135, 123, and 100 corresponding protein–protein ratios (Tables S3 and S4, Supplemental Digital Content). To enhance robustness, we pooled candidate protein–protein ratios from both databases, generating union sets comprising 251 protein–protein ratios for LC, 243 for LUAD, 234 for LUSC, and 223 for SCLC. We then reevaluated these associations by incorporating BWMR to validate IVW estimates and by conducting sensitivity analyses (MR-PRESSO and Radial MR) to further assess horizontal pleiotropy. These analyses refined the sets of protein–protein ratios for FinnGen (LC: 135; LUAD: 114; LUSC: 115; SCLC: 126) and TRICL-ILCCO (LC: 121; LUAD: 128; LUSC: 123; SCLC: 100) (Fig. S1 and Table S5, Supplemental Digital Content). An intersection analysis identified protein–protein ratios with nominally significant, directionally consistent effects across both databases, yielding 5 for LC (after excluding 1 ratio due to reverse causation, as detailed below), 4 for LUAD, 5 for LUSC, and 4 for SCLC (Fig. 3A–H). As shown in Figure 3I, each 1-standard deviation increase in genetically predicted protein–protein ratio levels was associated with varying lung cancer risks across subtypes. In FinnGen, higher genetically predicted levels of 3 protein–protein ratios were associated with an increased LC risk: melanoma inhibitory activity (MIA)/DAN family BMP antagonist (NBL1; OR 1.167, 95% CI: 1.070–1.274, P = .0005), HEXIM P-TEFb complex subunit 1/interaction protein for cytohesin exchange factors 1 (OR 1.161, 95% CI: 1.035–1.303, P = .011), and secretogranin II/transmembrane serine protease 15 (OR 1.140, 95% CI: 1.026–1.266, P = .014), while elevated levels of hippocalcin-like 1/S100 calcium-binding protein A4 (HPCAL1/S100A4; OR 0.725, 95% CI: 0.590–0.892, P = .002) and glyoxalase domain containing 4/S100A4 (OR 0.860, 95% CI: 0.751–0.985, P = .030) were found to exert a protective effect. For LUAD, 4 protein–protein ratios showed robust associations with increased risk: adhesion G protein-coupled receptor E2/CD58 (OR 1.347, 95% CI: 1.111–1.632, P = .002), angiopoietin-like 3/TFPI (OR 1.360, 95% CI: 1.149–1.609, P = .0003), fatty acid-binding protein 5 (FABP5)/NSFL1 cofactor p47 (OR 1.535, 95% CI: 1.090–2.160, P = .014), and glycoprotein 2/E-selectin (SELE) (OR 1.223, 95% CI: 1.018–1.471, P = .032). For LUSC, butyrophilin subfamily 2 member A1/sialic acid-binding Ig-like lectin 1 (BTN2A1/SIGLEC1; OR 0.850, 95% CI: 0.732–0.987, P = .032), and layilin/UL16 binding protein 2 (LAYN/ULBP2; OR 0.836, 95% CI: 0.726–0.962, P = .013) were inversely associated with risk, whereas contactin 5/roundabout guidance receptor 2 (OR 1.219, 95% CI: 1.030–1.443, P = .022), cellular repressor of E1A stimulated genes 1/N ‐ acylethanolamine acid amidase (OR 1.312, 95% CI: 1.113–1.546, P = .001), and MIA/NBL1 (OR 1.269, 95% CI: 1.072–1.501, P = .006) were linked to a higher risk. For SCLC, 3 protein–protein ratios were positively associated with risk: C-X-C motif chemokine ligand 5 (CXCL5)/C-X-C motif chemokine ligand 8 (OR 1.269, 95% CI: 1.006–1.602, P = .045), methionine sulfoxide reductase A/serpin family B member 1 (SERPINB1) (OR 1.494, 95% CI: 1.075–2.076, P = .017), and NSFL1 cofactor p47/SERPINB1 (OR 1.985, 95% CI: 1.091–3.614, P = .025), while SERPINB1/stress induced phosphoprotein 1 was associated with a reduced risk (OR 0.579, 95% CI: 0.353–0.949, P = .030). Replication in TRICL-ILCCO showed directionally consistent associations, accompanied by a general attenuation of effect magnitudes. Additionally, sensitivity analyses indicated no evidence of pleiotropy or heterogeneity among the identified protein–protein ratios (Fig. S2 and Table S6, Supplemental Digital Content). In reverse MR, we found that LC was associated with lower perilipin 3/thymidine levels in FinnGen (IVW: β = −0.037, P = .009) (Tables S7 and S8, Supplemental Digital Content). To mitigate potential reverse causation, perilipin 3/thymidine was excluded from the final sets of identified ratios and subsequent analyses.

Figure 2.

Figure 2.

Flow diagram of the multistep analytical workflow. Numbers represent counts of protein–protein ratios retained at each filtering stage. * For LC, 1 ratio (PLIN3/TYMP) was excluded after reverse MR analysis, yielding 5 final identified rQTLs. LC = overall lung cancer, LUAD = lung adenocarcinoma, LUSC = lung squamous cell carcinoma, PLIN3 = perilipin 3, rQTL = ratio quantitative trait locus, SCLC = small cell lung cancer, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium, TYMP = thymidine.

Figure 3.

Figure 3.

Causal effects of identified protein–protein ratios on lung cancer. (A–H) Volcano plots of protein–protein ratios with significant and consistent causal effects in both the FinnGen (A–D) and TRICL‑ILCCO (E–H) databases. (I) Forest plot of inverse-variance weighted (IVW) estimates for the associations between the identified protein–protein ratios and lung cancer subtypes. BTN2A1 = butyrophilin subfamily 2 member A1, CI = confidence interval, CNTN5 = contactin 5, CXCL5 = C-X-C motif chemokine ligand 5, CXCL8 = C-X-C motif chemokine ligand 8, GLOD4 = glyoxalase domain containing 4, GP2 = glycoprotein 2, HPCAL1 = hippocalcin-like 1, LAYN = layilin, LC = overall lung cancer, LUAD = lung adenocarcinoma, LUSC = lung squamous cell carcinoma, NBL1 = DAN family BMP antagonist, Not sig = not significant, OR = odds ratio, ROBO2 = roundabout guidance receptor 2, S100A4 = S100 calcium-binding protein A4, SCG2 = secretogranin II, SCLC = small cell lung cancer, SELE = E-selectin, SERPINB1 = serpin family B member 1, SIGLEC1 = sialic acid-binding Ig-like lectin 1, SNP = single-nucleotide polymorphism, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium, ULBP2 = UL16 binding protein 2.

3.3. Causal effects of plasma proteins on lung cancer

Genetic associations with protein–protein ratios can uncover biologically relevant links between proteins by capturing their shared genetic and nongenetic variance, which may be masked at the single-protein level. We identified 30 unique plasma proteins from 17 protein–protein ratios and performed MR analyses linking each protein to the lung cancer subtypes associated with its corresponding ratio. To further validate whether protein–protein ratios enhance signal strength by integrating interactions between proteins, we assessed the differential causal effects of protein–protein ratios versus single-protein levels on lung cancer. The results revealed 6 significant associations in FinnGen (LC: 2; LUAD: 3; LUSC: 1; SCLC: 0) and 7 in TRICL-ILCCO (LC: 3; LUAD: 1; LUSC: 2; SCLC: 1) (Fig. 4 and Table S9, Supplemental Digital Content). Subsequent colocalization analysis provided strong evidence for the MIA–LC association in FinnGen and the BTN2A1–LUSC association in TRICL-ILCCO (Table S10, Supplemental Digital Content). Notably, only the causal effect of MIA on LC showed directional consistency and significance across both databases. In contrast, the remaining 29 proteins did not exhibit robust, independent associations across both cohorts. Furthermore, no evidence of reverse causation was observed, as detailed in Table S11, Supplemental Digital Content. These findings passed checks for pleiotropy and heterogeneity (Fig. S3 and Tables S12 and S13, Supplemental Digital Content).

Figure 4.

Figure 4.

Heatmap of causal effects of individual plasma proteins on lung cancer. Nominal significance is marked with * for P < .05. BTN2A1 = butyrophilin subfamily 2 member A1, BWMR = Bayesian weighted Mendelian randomization, CNTN5 = contactin 5, CXCL5 = C-X-C motif chemokine ligand 5, CXCL8 = C-X-C motif chemokine ligand 8, GLOD4 = glyoxalase domain containing 4, GP2 = glycoprotein 2, HPCAL1 = hippocalcin-like 1, IVW = inverse-variance weighted, LAYN = layilin, LC = overall lung cancer, LUAD = lung adenocarcinoma, LUSC = lung squamous cell carcinoma, MIA = melanoma inhibitory activity, NBL1 = DAN family BMP antagonist, Not sig = not significant, ROBO2 = roundabout guidance receptor 2, S100A4 = S100 calcium-binding protein A4, SCG2 = secretogranin II, SCLC = small cell lung cancer, SELE = E-selectin, SERPINB1 = serpin family B member 1, SIGLEC1 = sialic acid-binding Ig-like lectin 1, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium, ULBP2 = UL16 binding protein 2.

3.4. Mediation analyses of potential protein–protein ratios

Given the established association of smoking and BMI with lung cancer, we investigated whether the identified protein–protein ratios acted as mediators in these causal pathways within the FinnGen and TRICL-ILCCO databases. Following UVMR analyses to establish associations (Tables S14 and S15, Supplemental Digital Content) and subsequent directional screening of effects (β, β1, and β2), we identified the same mediation pathways in both databases, though with varying evidence strength (Fig. 5A). In the FinnGen database, we found strong evidence for 2 mediation pathways: the MIA/NBL1 ratio mediated the effect of BMI on LC and on LUSC, accounting for 17.42% (95% CI: 2.60%–32.25%) and 13.46% (95% CI: 3.14%–23.77%) of the total effect, respectively. Four additional pathways showed potential evidence: smoking initiation influencing LC via glyoxalase domain containing 4/S100A4 and HPCAL1/S100A4, smoking initiation influencing LUSC via LAYN/ULBP2, and BMI influencing LUSC via contactin 5/roundabout guidance receptor 2, with mediation proportions ranging from 6.64% to 16.43% (Fig. 5B and Table S16, Supplemental Digital Content). In the TRICL-ILCCO database, the BMI–LUSC pathway mediated by MIA/NBL1 maintained strong evidence, accounting for 9.26% (95% CI: 0.05%–18.47%) of the total effect, while the remaining 5 pathways demonstrated potential evidence, with proportions ranging from 3.90% to 13.05% (Fig. 5C and Table S16, Supplemental Digital Content). Overall, only the BMI–MIA/NBL1–LUSC pathway demonstrated consistent, strong evidence across both databases; therefore, the remaining pathways are strictly classified as exploratory and should be interpreted with caution. Moreover, no horizontal pleiotropy or heterogeneity was detected across all MR mediation analyses in both databases, underscoring the robustness of these findings (Tables S17 and S18, Supplemental Digital Content).

Figure 5.

Figure 5.

Mediating roles of protein–protein ratios in the effects of smoking initiation and BMI on lung cancer. (A) Sankey diagram illustrating identified mediation pathways. Symbols denote evidence strength: * and † indicate strong evidence in FinnGen and TRICL-ILCCO, respectively, while # represents potential evidence. (B, C) Forest plots showing indirect effects and mediated proportions in the FinnGen (B) and TRICL-ILCCO (C) databases. Pathways demonstrating strong evidence across datasets are highlighted in red text. BMI = body mass index, CI = confidence interval, CNTN5 = contactin 5, GLOD4 = glyoxalase domain containing 4, HPCAL1 = hippocalcin-like 1, LAYN = layilin, LC = overall lung cancer, LUSC = lung squamous cell carcinoma, MIA = melanoma inhibitory activity, NBL1 = DAN family BMP antagonist, ROBO2 = roundabout guidance receptor 2, S100A4 = S100 calcium-binding protein A4, TRICL-ILCCO = Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium, ULBP2 = UL16 binding protein 2.

4. Discussion

Distinct pathogenic pathways and corresponding therapeutic strategies for lung cancer subtypes rest on substantial genetic heterogeneity.[23,47] However, this heterogeneity has not been extensively explored at the proteome level, particularly through subtype-specific plasma protein–protein ratios. To bridge this gap, our MR study provides, to our knowledge, the first systematic evaluation of the causal effects of 2821 plasma protein–protein ratios on lung cancer and its subtypes. Through discovery and replication in 2 independent databases, we identified 17 unique plasma protein–protein ratios robustly associated with lung cancer subtypes (LC: 5; LUAD: 4; LUSC: 5; SCLC: 4). Among the 30 unique proteins constituting these ratios, only MIA showed a causal association with LC. Moreover, 2-step MR indicated that several identified ratios partially mediated the effects of smoking initiation and BMI on lung cancer risk. Notably, the BMI–MIA/NBL1–LUSC pathway stood out as the sole mediation pathway with consistently strong evidence across both databases, while the remaining pathways should be considered exploratory.

In the LC analysis, 2 protein ratios involving S100A4 were associated with a reduced risk. This inverse association is particularly intriguing given the established pro-tumorigenic role of S100A4. As a key participant in the epithelial–mesenchymal transition, S100A4 is linked to metastasis in various malignancies.[48–50] Clinically, S100A4 overexpression is associated with tumor progression and poor prognosis in non-small cell lung cancer (NSCLC).[51] Mechanistically, S100A4 promotes lung tumor proliferation by inhibiting autophagy via the β-catenin pathway and enhances metastatic growth in response to the expulsion of nuclei from apoptotic tumor cells.[52,53] Additionally, exploratory mediation analysis suggested that these 2 ratios may mediate the effect of smoking initiation on lung cancer risk. Notably, the reported elevation of S100A4 in NSCLC with increased smoking pack-years provides a plausible biological link to this mediation finding.[54] HPCAL1 appears to play a dual role: it can drive NSCLC growth by reinforcing lactate dehydrogenase A activation,[55] yet also acts as a specific autophagy receptor targeting cadherin 2 to amplify ferroptosis, thereby exhibiting tumor-suppressive potential.[56] In parallel, our study revealed that the MIA/NBL1 ratio increased the risk for both LC and LUSC. Notably, this ratio served as the only mediator with consistent strong evidence across both databases, mediating the effect of BMI on LUSC development. MIA, a small secreted protein, is known to facilitate cancer migration and invasion.[57,58] Its misexpression in embryonic lung epithelium implies a possible fetal oncogenic role.[59] Consistent with our results, MIA is highly expressed in the peripheral blood and tumor tissues of lung cancer patients, where it promotes lung cancer progression by inhibiting tumor cell apoptosis.[60] These established functions underscore the potential of the MIA/NBL1 ratio as a candidate for future biomarker research, pending rigorous functional validation.

Our analysis further elucidated several causal risk factors for LUAD. Adhesion G protein-coupled receptor E2 has been identified as an oncoprotein primarily enriched in LUAD. It reportedly counteracts tumor suppression in NSCLC, potentially through epithelial–mesenchymal transition-driven stemness activation.[61] Angiogenesis is a critical hallmark of advanced cancer progression.[62] Angiopoietin-like 3 is a novel, liver-specific secreted angiogenic factor.[63] It promotes tumor-associated angiogenesis in LUAD through its interaction with Y-box binding protein 1, making it a candidate for further investigation in anti-angiogenic research.[64] Alterations in tumor metabolism also play a crucial role. High expression of FABP5, a lipid-chaperone protein found in macrophages, is linked to a poor prognosis in LUAD patients.[65] This association may be attributed to complement component 1q+ tumor-associated macrophages in malignant pleural effusion, which enhance FABP5-mediated fatty acid metabolism and subsequently promote tumor immunosuppression.[66] Alongside metabolic alterations, SELE plays a prominent role in metastasis. It shapes the tumor microenvironment by recruiting and activating monocytes, a process that facilitates the extravasation of lung tumor cells.[67] Notably, a clinical study found significantly elevated serum SELE levels in patients with metastatic SCLC or NSCLC, highlighting its relevance as a candidate for future prognostic research in lung cancer.[68]

Our results indicate that both BTN2A1/SIGLEC1 and LAYN/ULBP2 are associated with a reduced LUSC risk. However, the constituent proteins of the former and latter ratios exert opposing immunological functions. On one hand, as an activating immune checkpoint for Vγ9Vδ2 T cells, BTN2A1 has been reported to enhance their cytotoxicity against tumors and induce pyroptosis in malignant cells.[69,70] This aligns with our results and is supported by colocalization analysis. Similarly, SIGLEC1 may act as a co-stimulatory molecule for cytotoxic T cells, exerting antitumor properties.[71] On the other hand, the transmembrane protein LAYN has been shown to inhibit the killing effect of tumor-infiltrating cluster of differentiation 8 positive T cells. Thereby, its downregulation can bolster antitumor immunity in lung cancer.[72] Furthermore, serum levels of ULBP2 are markedly elevated in patients with LUSC and directly impair the cytolytic activity of peripheral blood mononuclear cells.[73] Taken together, this apparent contradiction implies that the biological effect of these protein ratios is governed by a complex regulatory network, in which the relative abundance of the constituent proteins supersedes their individual functional roles.

In SCLC, neutrophil-mediated inflammation appears to be an important driver of tumor progression. Our study found that SERPINB1 was involved in 3 protein ratios significantly associated with SCLC. As a potent inhibitor of neutrophil elastase and neutrophil extracellular traps, SERPINB1 has been demonstrated to suppress the invasion and motility of lung and breast cancer cells.[74] In addition, its decreased expression is regarded as a novel biomarker for prostate cancer progression.[75] Conversely, circulating tumor cells in advanced SCLC recruit macrophages, which elevate the expression of CXCL5 and C-X-C motif chemokine ligand 8. These chemokines then attract pro-angiogenic neutrophils that facilitate tumor migration and invasion.[76] Additionally, CXCL5-mediated neutrophil accumulation further impairs the function of antitumor cluster of differentiation 8 positive T cells, contributing to immune evasion.[77]

Crucially, a clear distinction must be made between ratio-level associations and single-protein corroboration. As noted, our individual MR analysis identified MIA as the only constituent protein exhibiting significant and directionally consistent causal effects across both databases. Consequently, the MIA/NBL1 ratio possesses a distinct, higher level of biological anchoring, as its causal signal is independently supported by the MIA protein itself. Conversely, the remaining ratio-based associations lacked consistent independent effects from their single constituents. This divergence underscores the unique biological significance of protein–protein ratios.; it suggests that, for these specific pathways, the causal mechanism is likely driven by the imbalance in their relative abundance, which captures shared variations masked in single-protein analyses. By distinguishing these levels of evidence, our study provides a nuanced layer of biological insight, acknowledging that findings with single-protein corroboration currently carry stronger mechanistic weight. Furthermore, the results of our 2-step mediation analysis provide novel insights into how smoking and BMI influence lung cancer risk. Notably, while the BMI–MIA/NBL1–LUSC pathway demonstrated robust evidence, the remaining mediation pathways should be viewed as exploratory.

This study offers several strengths. First, we utilized genetic summary data from extensive, large-scale GWAS studies designed to minimize potential sample overlap. Second, we employed stringent criteria to select genetic instruments, thereby mitigating potential confounding and weak instrument bias. Third, to establish credible and consistent causal inferences, we performed tens of thousands of association estimates across 2 independent databases, covering multiple lung cancer subtypes. Finally, the robustness and reliability of our results were systematically confirmed through comprehensive sensitivity analyses.

Several limitations of this study should be acknowledged. First, reliance on summary-level data constrained our capacity to conduct stratified analyses or to explore individual-level interactions. Second, restriction to a European-ancestry population limits the generalizability of the findings to other ethnic groups. Third, although certain protein–protein ratios appeared subtype-specific, their potential relevance to LC risk cannot be entirely excluded given shared genetic architecture. Fourth, while we performed FDR correction across the 2821 plasma protein–protein ratios, no associations survived this stringent threshold. As highlighted by a recent proteomic MR study on lung cancer subtypes,[22] applying strict multiple-testing correction across thousands of plasma proteins can substantially increase type II errors and inherently restrict discovery power.[78] Therefore, our primary findings are based on unadjusted P-values, with stringent cross-database replication serving as the principal safeguard against false positives. Fifth, we must acknowledge the limited corroboration at the single-protein level. With the exception of MIA, the identified ratio-based associations lacked consistent independent effects from their individual constituent proteins, suggesting these findings predominantly capture relative abundance imbalances rather than the efficacy of single targets. Sixth, the mediation analyses are fundamentally exploratory. Given that several indirect-effect CIs bordered on the null or crossed 0 during replication, these pathways must be strictly interpreted as hypothesis-generating. While we used 2-step MR for this exploratory analysis, multivariable MR could provide complementary evidence by simultaneously adjusting for correlated mediators. Future studies with dedicated confirmatory designs may employ multivariable MR to further validate these pathways. Seventh, although we rigorously aligned the phenotypic definitions and diagnostic codes across databases (Table 1), inherent differences in study design, ascertainment methods, and patient enrollment between the FinnGen database and the TRICL-ILCCO consortium may still introduce latent outcome-definition mismatch and heterogeneity, particularly for LC. Finally, these genetic inferences require validation in experimental systems and large, multiethnic population-based studies to clarify underlying biological mechanisms.

5. Conclusion

This proteomic MR study identified 17 plasma protein–protein ratios associated with lung cancer risk. Subsequent mediation analyses demonstrated that the BMI–MIA/NBL1–LUSC pathway was the sole mediation pathway with consistent strong evidence across both databases, while the remaining pathways mediating the effects of smoking and BMI should be considered exploratory. Collectively, these findings provide novel proteomic insights into the causal pathways underlying lung cancer development, highlighting genetically supported candidate pathways that warrant further functional validation and clinical investigation.

Acknowledgments

Our sincere gratitude extends to all the researchers and participants of the contributing studies, whose efforts made this research possible. We are also grateful to UKB-PPP, GWAS and Sequencing Consortium of Alcohol and Nicotine use, Genetic Investigation of Anthropometric Traits, FinnGen, TRICL, ILCCO, and the GWAS Catalog for sharing the GWAS data. Additionally, we acknowledge the developers of the R packages that were used in our analysis.

Author contributions

Conceptualization: Xin Geng.

Methodology: Xin Geng, Zhihua Li.

Data curation: Xin Geng, Jingyuan Xue, Zipei Song.

Formal analysis: Xin Geng, Kaisen Liu, Yuheng Wang.

Investigation: Xin Geng, Xincen Cao.

Project administration: Liang Chen, Zhihua Li.

Resources: Liang Chen, Xin Geng, Zhihua Li.

Supervision: Liang Chen.

Visualization: Xin Geng, Yuheng Wang.

Validation: Jingyuan Xue, Kaisen Liu, Zipei Song.

Writing – review & editing: Liang Chen.

Writing – original draft: Xin Geng, Zhihua Li.

medi-105-e48953-s006.xlsx (374.9KB, xlsx)
medi-105-e48953-s007.xlsx (33.9KB, xlsx)
medi-105-e48953-s008.xlsx (55.1KB, xlsx)
medi-105-e48953-s009.xlsx (11.5KB, xlsx)
medi-105-e48953-s010.xlsx (20.1KB, xlsx)
medi-105-e48953-s012.xlsx (23.6KB, xlsx)
medi-105-e48953-s013.xlsx (34.9KB, xlsx)
medi-105-e48953-s014.xlsx (10.6KB, xlsx)
medi-105-e48953-s015.xlsx (16.8KB, xlsx)
medi-105-e48953-s018.xlsx (19.1KB, xlsx)
medi-105-e48953-s019.xlsx (14.2KB, xlsx)
medi-105-e48953-s020.xlsx (22.3KB, xlsx)
medi-105-e48953-s021.xlsx (15.3KB, xlsx)

Abbreviations:

BMI
body mass index
BTN2A1
butyrophilin subfamily 2 member A1
BWMR
Bayesian weighted Mendelian randomization
CI
confidence interval
CXCL5
C-X-C motif chemokine ligand 5
FABP5
fatty acid-binding protein 5
FDR
false discovery rate
GWAS
genome-wide association studies
HPCAL1
hippocalcin-like 1
IV
instrumental variable
IVW
inverse-variance weighted
LAYN
layilin
LC
overall lung cancer
LUAD
lung adenocarcinoma
LUSC
lung squamous cell carcinoma
MIA
melanoma inhibitory activity
MR
Mendelian randomization
MR-PRESSO
MR pleiotropy residual sum and outlier
NBL1
DAN family BMP antagonist
NSCLC
non-small cell lung cancer
OR
odds ratio
pQTL
protein quantitative trait locus
rQTL
ratio quantitative trait locus
S100A4
S100 calcium-binding protein A4
SCLC
small cell lung cancer
SELE
E-selectin
SERPINB1
serpin family B member 1
SIGLEC1
sialic acid-binding Ig-like lectin 1
SNP
single-nucleotide polymorphism
TRICL-ILCCO
Transdisciplinary Research in Cancer of the Lung and International Lung Cancer Consortium
ULBP2
UL16 binding protein 2
UVMR
univariable Mendelian randomization

This work was supported by the Major Program of Science and Technology Foundation of Jiangsu Province (BE2023832 and BE2018746) and the Specialized Diseases Clinical Research Fund of Jiangsu Province Hospital (DL202402).

This study was based exclusively on publicly available summary-level data. Therefore, no additional ethical approval or informed consent was required.

The authors have no conflicts of interest to disclose.

All data generated or analyzed during this study are included in this published article (and its supplementary information files).

Supplemental Digital Content is available in the online version of this article (http://dx.doi.org/10.1097/MD.0000000000048953).

How to cite this article: Geng X, Xue J, Liu K, Song Z, Wang Y, Cao X, Li Z, Chen L. Causal effects of plasma protein–protein ratios on lung cancer subtypes: A proteome-wide Mendelian randomization and mediation analysis. Medicine 2026;105:21(e48953).

XG, JX, and KL contributed to this article equally.

Contributor Information

Xin Geng, Email: xingeng1108@163.com.

Jingyuan Xue, Email: jyxue2002@gmail.com.

Kaisen Liu, Email: kaisenliu@stu.njmu.edu.cn.

Zipei Song, Email: songzipei1003@163.com.

Yuheng Wang, Email: 460629480@qq.com.

Xincen Cao, Email: 1346068275@qq.com.

Zhihua Li, Email: lizhihua_njmu@126.com.

References

  • [1].Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74:229–63. [DOI] [PubMed] [Google Scholar]
  • [2].Zhang Y, Vaccarella S, Morgan E, et al. Global variations in lung cancer incidence by histological subtype in 2020: a population-based study. Lancet Oncol. 2023;24:1206–18. [DOI] [PubMed] [Google Scholar]
  • [3].Siegel RL, Miller KD, Wagle NS, Jemal A. Cancer statistics, 2023. CA Cancer J Clin. 2023;73:17–48. [DOI] [PubMed] [Google Scholar]
  • [4].Malhotra J, Malvezzi M, Negri E, La Vecchia C, Boffetta P. Risk factors for lung cancer worldwide. Eur Respir J. 2016;48:889–902. [DOI] [PubMed] [Google Scholar]
  • [5].Yanovich-Arad G, Geiger T. Across the globe: proteogenomic landscapes of lung cancer. Cell. 2020;182:9–11. [DOI] [PubMed] [Google Scholar]
  • [6].Gillette MA, Satpathy S, Cao S, et al. Proteogenomic characterization reveals therapeutic vulnerabilities in lung adenocarcinoma. Cell. 2020;182:200–25.e35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Davies MPA, Sato T, Ashoor H, et al. Plasma protein biomarkers for early prediction of lung cancer. EBioMedicine. 2023;93:104686. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Zhang M, Li H, Ma S, et al. Serum proteome profiling reveals HGFA as a candidate biomarker for pulmonary arterial hypertension. Respir Res. 2024;25:418. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Suhre K. Genetic associations with ratios between protein levels detect new pQTLs and reveal protein-protein interactions. Cell Genom. 2024;4:100506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Yamauchi Y, Safi S, Muley T, et al. C-reactive protein-albumin ratio is an independent prognostic predictor of tumor recurrence in stage IIIA-N2 lung adenocarcinoma patients. Lung Cancer. 2017;114:62–7. [DOI] [PubMed] [Google Scholar]
  • [11].Jones JM, McGonigle NC, McAnespie M, Cran GW, Graham AN. Plasma fibrinogen and serum C-reactive protein are associated with non-small cell lung cancer. Lung Cancer. 2006;53:97–101. [DOI] [PubMed] [Google Scholar]
  • [12].Bowden J, Holmes MV. Meta-analysis and Mendelian randomization: a review. Res Synth Methods. 2019;10:486–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Lawlor DA, Harbord RM, Sterne JA, Timpson N, Davey Smith G. Mendelian randomization: using genes as instruments for making causal inferences in epidemiology. Stat Med. 2008;27:1133–63. [DOI] [PubMed] [Google Scholar]
  • [14].Jeon J, Holford TR, Levy DT, et al. Smoking and lung cancer mortality in the United States from 2015 to 2065: a comparative modeling approach. Ann Intern Med. 2018;169:684–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Barbi J, Patnaik SK, Pabla S, et al. Visceral obesity promotes lung cancer progression-toward resolution of the obesity paradox in lung cancer. J Thorac Oncol. 2021;16:1333–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Gao C, Patel CJ, Michailidou K, et al. Mendelian randomization study of adiposity-related traits and risk of breast, ovarian, prostate, lung and colorectal cancer. Int J Epidemiol. 2016;45:896–908. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Didelez V, Sheehan N. Mendelian randomization as an instrumental variable approach to causal inference. Stat Methods Med Res. 2007;16:309–30. [DOI] [PubMed] [Google Scholar]
  • [18].Sun BB, Chiou J, Traylor M, et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature. 2023;622:329–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Liu M, Jiang Y, Wedow R, et al. Association studies of up to 1.2 million individuals yield new insights into the genetic etiology of tobacco and alcohol use. Nat Genet. 2019;51:237–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [20].Yengo L, Sidorenko J, Kemper KE, et al. Meta-analysis of genome-wide association studies for height and body mass index in approximately 700000 individuals of European ancestry. Hum Mol Genet. 2018;27:3641–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].Kurki MI, Karjalainen J, Palta P, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613:508–18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Lyu Z, Si G, Xing M, et al. Circulating proteins associated with histological subtypes of lung cancer from genetic and population-based perspectives. PLoS Genet. 2025;21:e1011821. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].McKay JD, Hung RJ, Han Y, et al. Large-scale association analysis identifies new lung cancer susceptibility loci and heterogeneity in genetic susceptibility across histological subtypes. Nat Genet. 2017;49:1126–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Burgess S, Davies NM, Thompson SG. Bias due to participant overlap in two-sample Mendelian randomization. Genet Epidemiol. 2016;40:597–608. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Sanna S, van Zuydam NR, Mahajan A, et al. Causal relationships among the gut microbiome, short-chain fatty acids and metabolic diseases. Nat Genet. 2019;51:600–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [26].Larsson SC, Carter P, Kar S, et al. Smoking, alcohol consumption, and cancer: a Mendelian randomisation study in UK Biobank and international genetic consortia participants. PLoS Med. 2020;17:e1003178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Ding T, Lin Q, Qu X. From chronic obstructive pulmonary disease (COPD) to lung cancer: a Mendelian randomization study revealing mediation pathways through plasma metabolomics, proteomics, and immunophenotyping. Discov Oncol. 2025;16:629. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Liu Y, Lai H, Zhang R, Xia L, Liu L. Causal relationship between gastro-esophageal reflux disease and risk of lung cancer: insights from multivariable Mendelian randomization and mediation analysis. Int J Epidemiol. 2023;52:1435–47. [DOI] [PubMed] [Google Scholar]
  • [29].Zhou H, Zhang Y, Liu J, et al. Education and lung cancer: a Mendelian randomization study. Int J Epidemiol. 2019;48:743–50. [DOI] [PubMed] [Google Scholar]
  • [30].Pierce BL, Ahsan H, Vanderweele TJ. Power and instrument strength requirements for Mendelian randomization studies using multiple genetic variants. Int J Epidemiol. 2011;40:740–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [31].Lin Z, Deng Y, Pan W. Combining the strengths of inverse-variance weighting and Egger regression in Mendelian randomization using a mixture of regressions model. PLoS Genet. 2021;17:e1009922. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Burgess S, Butterworth A, Thompson SG. Mendelian randomization analysis with multiple genetic variants using summarized data. Genet Epidemiol. 2013;37:658–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Zhao J, Ming J, Hu X, Chen G, Liu J, Yang C. Bayesian weighted Mendelian randomization for causal inference based on summary statistics. Bioinformatics. 2020;36:1501–8. [DOI] [PubMed] [Google Scholar]
  • [34].Giambartolomei C, Vukcevic D, Schadt EE, et al. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 2014;10:e1004383. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Chen J, Xu F, Ruan X, et al. Therapeutic targets for inflammatory bowel disease: proteome-wide Mendelian randomization and colocalization analyses. EBioMedicine. 2023;89:104494. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Carter AR, Sanderson E, Hammerton G, et al. Mendelian randomisation for mediation analysis: current methods and challenges for implementation. Eur J Epidemiol. 2021;36:465–78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Jie J, Gong Y, Hu H, Liu S. The role of cerebrospinal fluid metabolites in mediating the impact of lipids on late-onset Alzheimer’s disease: a two-step Mendelian randomization analysis. J Transl Med. 2024;22:1077. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Bowden J, Davey Smith G, Burgess S. Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. Int J Epidemiol. 2015;44:512–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [39].Verbanck M, Chen CY, Neale B, Do R. Detection of widespread horizontal pleiotropy in causal relationships inferred from Mendelian randomization between complex traits and diseases. Nat Genet. 2018;50:693–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Qiu Y, Li C, Huang Y, et al. Exploring the causal associations of micronutrients on urate levels and the risk of gout: a Mendelian randomization study. Clin Nutr. 2024;43:1001–12. [DOI] [PubMed] [Google Scholar]
  • [41].Bowden J, Spiller W, Del Greco MF, et al. Improving the visualization, interpretation and analysis of two-sample summary data Mendelian randomization via the Radial plot and Radial regression. Int J Epidemiol. 2018;47:1264– 78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Cohen JF, Chalumeau M, Cohen R, Korevaar DA, Khoshnood B, Bossuyt PMM. Cochran’s Q test was useful to assess heterogeneity in likelihood ratios in studies of diagnostic accuracy. J Clin Epidemiol. 2015;68:299–306. [DOI] [PubMed] [Google Scholar]
  • [43].Burgess S, Bowden J, Fall T, Ingelsson E, Thompson SG. Sensitivity analyses for robust causal inference from Mendelian randomization analyses with multiple genetic variants. Epidemiology. 2017;28:30–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [44].Hemani G, Zheng J, Elsworth B, et al. The MR-Base platform supports systematic causal inference across the human phenome. Elife. 2018;7:e34408. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [45].Yavorska OO, Burgess S. MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data. Int J Epidemiol. 2017;46:1734–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [46].Ma Z, Zhao M, Zhao H, Qu N. Causal role of immune cells in generalized anxiety disorder: Mendelian randomization study. Front Immunol. 2023;14:1338083. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [47].Long E, Patel H, Golden A, et al. High-throughput characterization of functional variants highlights heterogeneity and polygenicity underlying lung cancer susceptibility. Am J Hum Genet. 2024;111:1405–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [48].Boye K, Maelandsmo GM. S100A4 and metastasis: a small actor playing many roles. Am J Pathol. 2010;176:528–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [49].Ning Q, Li F, Wang L, et al. S100A4 amplifies TGF-beta-induced epithelial-mesenchymal transition in a pleural mesothelial cell line. J Investig Med. 2018;66:334–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [50].Hua T, Liu S, Xin X, et al. S100A4 promotes endometrial cancer progress through epithelial-mesenchymal transition regulation. Oncol Rep. 2016;35:3419–26. [DOI] [PubMed] [Google Scholar]
  • [51].Zhang J, Gu Y, Liu X, Rao X, Huang G, Ouyang Y. Clinicopathological and prognostic value of S100A4 expression in non-small cell lung cancer: a meta-analysis. Biosci Rep. 2020;40:BSR20201710. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [52].Hou S, Tian T, Qi D, et al. S100A4 promotes lung tumor development through beta-catenin pathway-mediated autophagy inhibition. Cell Death Dis. 2018;9:277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [53].Park WY, Gray JM, Holewinski RJ, et al. Apoptosis-induced nuclear expulsion in tumor cells drives S100a4-mediated metastatic outgrowth through the RAGE pathway. Nat Cancer. 2023;4:419–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [54].Kongsgaard A, Lund-Iversen M, Berge G, et al. Expression of S100A4, ephrin-A1 and osteopontin in non-small cell lung cancer. BMC Cancer. 2012;12:333. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [55].Wang X, Xie X, Zhang Y, et al. Hippocalcin-like 1 is a key regulator of LDHA activation that promotes the growth of non-small cell lung carcinoma. Cell Oncol (Dordr). 2022;45:179–91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [56].Chen X, Song X, Li J, et al. Identification of HPCAL1 as a specific autophagy receptor involved in ferroptosis. Autophagy. 2023;19:54–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [57].Poser I, Tatzel J, Kuphal S, Bosserhoff AK. Functional role of MIA in melanocytes and early development of melanoma. Oncogene. 2004;23:6115–24. [DOI] [PubMed] [Google Scholar]
  • [58].El Fitori J, Kleeff J, Giese NA, et al. Melanoma inhibitory activity (MIA) increases the invasiveness of pancreatic cancer cells. Cancer Cell Int. 2005;5:3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [59].Lin S, Ikegami M, Xu Y, Bosserhoff A-K, Malkinson AM, Shannon JM. Misexpression of MIA disrupts lung morphogenesis and causes neonatal death. Dev Biol. 2008;316:441–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [60].Gu QH, Li D, Xie ZH, Shen QB. The clinical significance of MIA gene in tumorigenesis of lung cancer. Neoplasma. 2020;67:660–7. [DOI] [PubMed] [Google Scholar]
  • [61].Feliciano A, Garcia-Mayea Y, Jubierre L, et al. miR-99a reveals two novel oncogenic proteins E2F2 and EMR2 and represses stemness in lung cancer. Cell Death Dis. 2017;8:e3141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [62].Al-Abd AM, Alamoudi AJ, Abdel-Naim AB, Neamatallah TA, Ashour OM. Anti-angiogenic agents for the treatment of solid tumors: potential pathways, therapy and current strategies - a review. J Adv Res. 2017;8:591–605. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [63].Camenisch G, Pisabarro MT, Sherman D, et al. ANGPTL3 stimulates endothelial cell adhesion and migration via integrin alpha vbeta 3 and induces blood vessel formation in vivo. J Biol Chem. 2002;277:17281–90. [DOI] [PubMed] [Google Scholar]
  • [64].Cong Z, Diao Y, Li X, et al. Long non-coding RNA linc00665 interacts with YB-1 and promotes angiogenesis in lung adenocarcinoma. Biochem Biophys Res Commun. 2020;527:545–52. [DOI] [PubMed] [Google Scholar]
  • [65].Garcia KA, Costa ML, Lacunza E, Martinez ME, Corsico B, Scaglia N. Fatty acid binding protein 5 regulates lipogenesis and tumor growth in lung adenocarcinoma. Life Sci. 2022;301:120621. [DOI] [PubMed] [Google Scholar]
  • [66].Zhang S, Peng W, Wang H, et al. C1q(+) tumor-associated macrophages contribute to immunosuppression through fatty acid metabolic reprogramming in malignant pleural effusion. J Immunother Cancer. 2023;11:e007441. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [67].Hauselmann I, Roblek M, Protsyuk D, et al. Monocyte induction of E-selectin-mediated endothelial activation releases VE-cadherin junctions to promote tumor cell extravasation in the metastasis cascade. Cancer Res. 2016;76:5302–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [68].Gogali A, Charalabopoulos K, Zampira I, et al. Soluble adhesion molecules E-cadherin, intercellular adhesion molecule-1, and E-selectin as lung cancer biomarkers. Chest. 2010;138:1173–9. [DOI] [PubMed] [Google Scholar]
  • [69].Cano CE, Pasero C, De Gassart A, et al. BTN2A1, an immune checkpoint targeting Vgamma9Vdelta2 T cell cytotoxicity against malignant cells. Cell Rep. 2021;36:109359. [DOI] [PubMed] [Google Scholar]
  • [70].Le Floch AC, Imbert C, Boucherit N, et al. Targeting BTN2A1 enhances Vgamma9Vdelta2 T-Cell effector functions and triggers tumor cell pyroptosis. Cancer Immunol Res. 2024;12:1677–90. [DOI] [PubMed] [Google Scholar]
  • [71].Zhang Y, Li JQ, Jiang ZZ, Li L, Wu Y, Zheng L. CD169 identifies an anti-tumour macrophage subpopulation in human hepatocellular carcinoma. J Pathol. 2016;239:231–41. [DOI] [PubMed] [Google Scholar]
  • [72].Yang B, Deng B, Jiao XD, et al. Low-dose anti-VEGFR2 therapy promotes anti-tumor immunity in lung adenocarcinoma by down-regulating the expression of layilin on tumor-infiltrating CD8(+)T cells. Cell Oncol (Dordr). 2022;45:1297–309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [73].Yamaguchi K, Chikumi H, Shimizu A, et al. Diagnostic and prognostic impact of serum-soluble UL16-binding protein 2 in lung cancer patients. Cancer Sci. 2012;103:1405–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [74].Chou RH, Wen HC, Liang WG, et al. Suppression of the invasion and migration of cancer cells by SERPINB family genes and their derived peptides. Oncol Rep. 2012;27:238–45. [DOI] [PubMed] [Google Scholar]
  • [75].Lerman I, Ma X, Seger C, et al. Epigenetic suppression of SERPINB1 promotes inflammation-mediated prostate cancer progression. Mol Cancer Res. 2019;17:845–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [76].Hamilton G, Rath B, Klameth L, Hochmair MJ. Small cell lung cancer: recruitment of macrophages by circulating tumor cells. Oncoimmunology. 2016;5:e1093277. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [77].Simoncello F, Piperno GM, Caronni N, et al. CXCL5-mediated accumulation of mature neutrophils in lung cancer tissues impairs the differentiation program of anticancer CD8 T cells and limits the efficacy of checkpoint inhibitors. Oncoimmunology. 2022;11:2059876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [78].Mazidi M, Wright N, Yao P, et al. Plasma proteomics to identify drug targets for ischemic heart disease. J Am Coll Cardiol. 2023;82:1906–20. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

medi-105-e48953-s006.xlsx (374.9KB, xlsx)
medi-105-e48953-s007.xlsx (33.9KB, xlsx)
medi-105-e48953-s008.xlsx (55.1KB, xlsx)
medi-105-e48953-s009.xlsx (11.5KB, xlsx)
medi-105-e48953-s010.xlsx (20.1KB, xlsx)
medi-105-e48953-s012.xlsx (23.6KB, xlsx)
medi-105-e48953-s013.xlsx (34.9KB, xlsx)
medi-105-e48953-s014.xlsx (10.6KB, xlsx)
medi-105-e48953-s015.xlsx (16.8KB, xlsx)
medi-105-e48953-s018.xlsx (19.1KB, xlsx)
medi-105-e48953-s019.xlsx (14.2KB, xlsx)
medi-105-e48953-s020.xlsx (22.3KB, xlsx)
medi-105-e48953-s021.xlsx (15.3KB, xlsx)

Articles from Medicine are provided here courtesy of Wolters Kluwer Health

RESOURCES