Skip to main content
Springer logoLink to Springer
. 2026 Mar 21;28(9):4236–4252. doi: 10.1007/s12094-025-04158-8

Public Transcriptomic Data Mining for SCLC: From Candidate Ma rkers to Therapeutic Exploration

Hailin Liu 1,#, Fangyuan Qu 1,#, Guangyao Zhou 1,#, Yuechen Cui 1, Bo Yan 1, Lianmin Zhang 1, Chenguang Li 1, Zhenfa Zhang 1, Tingting Qin 1, Qiangzhe Zhang 2,3,✉
PMCID: PMC13499861  PMID: 41863696

Abstract

Background

Lung cancer, especially small-cell lung cancer (SCLC), is a widespread and deadly disease often detected at advanced stages, resulting in low five-year survival rates. This study aims to identify new genetic targets to enhance understanding of the genetic drivers of SCLC progression.

Methods

Data from 215 samples (82 normal, 133 tumor) across four datasets were retrieved from the GEO database. Using R software, we normalized and analyzed the data to assess correlations between differentially expressed genes (DEGs) and SCLC. Techniques included differential expression, expression quantitative trait loci (eQTL), and Mendelian randomization (MR) analyses. Functional and pathway analyses utilized Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG). Machine learning was applied to develop predictive models for disease diagnosis and progression.

Results

Analysis of 129 samples revealed 369 upregulated and 529 downregulated genes. Six genes with shared regions were significantly linked to SCLC. GO and KEGG analyses highlighted their roles in vital processes like organic hydroxy compound biosynthesis. CIBERSORT analysis emphasized immune cell variations in SCLC patients. Machine learning identified key genes, with survival analysis showing significant differences for COLEC12 and MUC1, validated by GSEA and qPCR.

Conclusion

COLEC12 and MUC1 are novel diagnostic markers and therapeutic targets for SCLC, offering potential for targeted treatments and future research.

Supplementary Information

The online version contains supplementary material available at 10.1007/s12094-025-04158-8.

Keywords: SCLC, Mendelian randomization analyses, Machine learning, Diagnostic markers

Introduction

Small cell lung carcinoma (SCLC) is a particularly aggressive and deadly form of lung cancer [1–3], making up around 15% of all lung cancer cases [4, 5].It is known for its highly invasive nature and unfavorable prognosis [6, 7]. The majority of SCLC patients would develop treatment resistance within a short period, resulting in a very low overall survival rate [8]. Immunotherapy has achieved a significant advancement in the management of small cell lung cancer [9, 12]. Patient survival has not significantly improved over the last few decades [13], and SCLC still remains outside the realm of precision medicine [10, 14, 15]. Furthermore, numerous challenges were faced during the development of novel pharmaceuticals for the treatment of SCLC. Therefore, it is particularly important to find effective therapeutic targets. This work aims to look into the molecular mechanisms and genetic elements underlying the development of SCLC, focusing on the potential of relevant exposure factors as therapeutic targets, studying their molecular mechanisms, and making up for the shortcomings of existing treatment methods for SCLC.

Mendelian randomization (MR) is a methodology used to establish the causal impact of risk factors on health outcomes by employing genetic variations as instrumental variables [16, 17]. It simulates randomized trials by using randomly assigned genetic variants that exist in nature, thereby eliminating problems associated with potential confounding factors and reverse causality in observational studies [18–20]. The primary aim of this study is to discover genes that are expressed differently (differentially expressed genes or DEGs [18]) in SCLC compared to normal samples. The objective of the study was to assess the correlation and causal connection between these genes and the mechanisms responsible for the development of SCLC through the use of expression quantitative trait loci (eQTL) and Mendelian randomization (MR) analysis [17]. The project will utilize gene enrichment analysis, Gene Ontology (GO), and KEGG to investigate the possible functional pathways and disease processes linked to these differentially expressed genes (DEGs) [21, 22]. Furthermore, the study seeks to investigate the intricate mechanisms behind SCLC, with a specific focus on the cellular and molecular elements that could provide novel targets for treatment. Our objective was to uncover potential biomarkers and assess the direct impacts of potential therapeutic targets in SCLC. This will aid in future studies on the mechanisms involved and the development of new drugs.

Machine learning is an important artificial intelligence technology that is often used to anticipate cancer diagnostic markers [23]. It also shows great potential for forecasting cancer prognosis [24–26]. However, the development of reliable cancer outcome prediction models for routine clinical use remains a formidable challenge. In this study, we identified differentially expressed genes in SCLC patients from the GEO databases. Subsequently, we performed MR analysis on these genes to identify 6 genes associated with SCLC. Finally, through comparative analysis of various machine learning methods, we have confirmed the potential therapeutic targets and diagnostic indicators for SCLC as COLEC12 and MUC1.

Materials and methods

Data availability and processing

SCLCs data were acquired in Gene Expression Omnibus (GEO) database (https://www.ncbi.nlm.nih.gov/geo/) [27] (accession number GSE149507, GSE40275, GSE30219, GSE60052). The GSE60052 [28] dataset includes clinical survival data that can be used to plot future survival curves.

Differentially expressed genes (DEGs) found

Subsequently, the datasets were combined, and the process of adjusting for batch effects and conducting comparative analysis was carried out on a total of 82 normal samples and 54 SCLC samples. The “limma” package was employed for conventional Bayesian data analysis to identify differentially expressed genes (DEGs). The significance criteria were established at P < 0.001 and log Fold Change (Log FC) > 1.5. The “pheatmap” program generated volcano plots and heatmaps for DEGs. PCA was conducted with the “prcomp” algorithm to eliminate batch effects, improve visualization, and assess significant genes that distinguish SCLC from healthy control samples.

Acquiring and filtering eQTLs

The study utilized eQTL data obtained from the GWAS Catalog website (https://gwas.mrcieu.ac.uk/). The R package “TwoSampleMR” [29] was utilized to identify instrumental variables that have a strong correlation with SNPs (p < 5e-08). Linkage disequilibrium parameters were set at r2 < 0.001 and clumping distance = 10,000 kb [19]. We apply a filtering process to the data, specifically selecting only those with an F-value greater than 10. Additionally, we eliminate any weak instrumental factors.

Acquiring for the outcome data

The outcome data were acquired from the genetic association database of the GWAS summary dataset (IEU) accessible at https://gwas.mrcieu.ac.uk/. [30] The GWAS summary statistics utilized in this study are freely available and can be obtained at no charge. The GWAS ID utilized was ieu-a-988 and GCST004746. Ieu-a-988 encompasses 2791 cases and 20,580 controls of European ancestry, with a total of 7438,318 SNPs. GCST004746 encompasses 2664 cases and 21,444 controls of European ancestry, with a total of 7,620,430 SNPs.

Mendelian randomization analysis

The study was conducted using the “TwoSampleMR” software program. Initially, we performed Mendelian randomization analysis on the exposure data using five distinct approaches: MR Egger, simple mode, weighted media, weighted mode methods, and inverse variance weighted (IVW) method [31]. The OR values of the results were computed, and genes exhibiting the same OR direction as the five techniques were extracted. An examination of heterogeneity and pleiotropy was conducted on the instrumental variable SNP. Genes with a P-value less than 0.05 were selected from the IVW. Genes were removed with pleiotropic p values less than 0.05. Ultimately, genes associated with the risk of SCLC were identified. Genes exhibiting nominal significance (P < 0.05) in the inverse-variance weighted (IVW) analysis were initially selected. Subsequently, variants demonstrating horizontal pleiotropy, as indicated intercept P-value of less than 0.05, were excluded. All remaining genes successfully passed heterogeneity testing, as assessed (P ≥ 0.05), as well as leave-one-out sensitivity analyses, thereby confirming the robustness of our causal estimates.Ultimately, genes associated with the risk of SCLC were identified.

Acquisition of intersection genes

A total of 75 genes with an OR > 1 and 83 genes with an OR < 1. Using differential analysis, we have discovered 369 genes that are upregulated and 529 genes that are downregulated, all of which are associated with SCLC. By taking the intersection of OR > 1 with upregulated genes, we identified 2 genes. Similarly, by intersecting OR < 1 with downregulated genes, we identified a total of 4 genes.

GO and KEGG enrichment analysis

GO and KEGG analyses for co-expression genes were implemented using R package “org.Hs.eg.db,” “enrichplot,” “clusterProfiler,” “circlize,” and “ggplot2” [27] to enrich biological process (BP), cellular component (CC), molecular function (MF), and signal pathways, with the study’s filtering criterion set at P value < 0.05.

Immune cell analysis

The CIBERSORT algorithm is a deconvolution method that uses standardized gene expression profiles to predict the proportional proportions of 22 subtypes of immune cells in tissue samples [32, 33].The gene expression matrices from the three datasets were transformed into 22 immune cell matrices, which correspond to the leukocyte signature matrix (LM22). The p-value criterion of < 0.05 for each sample indicates that the anticipated fraction of each invading immune cell subtype is statistically accurate and appropriate for further study.

Constructing machine learning models for SCLC

Fifteen machine learning algorithms were selected for this study for the binary classification variables. LASSO (Least Absolute Shrinkage and Selection Operator), RR (Ridge Regression), Step-wise, LR (Logistic Regression), ENR (Elastic Net Regression). These methods predominantly rely on linear relationships for modeling purposes and are particularly well-suited for handling data that is either linearly separable or approximately linear. SVM (support vector machine), LDA (linear discriminant analysis), QDA (quadratic discriminant analysis), NB (Naive Bayes), DT (Decision Tree), RF (Random Forest), GBM (gradient boosting machine), KNN (K-Nearest Neighbors). These methods are primarily utilized for classification tasks. XGB (eXtreme Gradient Boosting): an optimization algorithm based on the gradient boosting framework, which belongs to ensemble learning methods. NN (Neural Network): A mathematical model that mimics the structure and function of biological neural networks. A total of 120 algorithms were developed from the 15 machine learning algorithms. All model combinations were subsequently tested in training and testing cohorts. For the performance evaluation of each model, the AUC score across these cohorts was computed. The model with the highest average AUC within the training and testing cohorts was regarded as optimal. Feature genes were subsequently obtained via the best-performing model.

Survival analysis was conducted on the genes that intersected

We performed survival analysis on the genes that had significant p-values at the intersection, using data from GSE60052 for validation. The collection comprises 87 samples, with 7 being normal cases and 79 being tumor cases. Among these, 45 samples have survival statistics available. A survival study was performed on the genes of the 45 samples, revealing that six genes, specifically PSRC1, PSAT1, DHCR24, COLEC12, MUC1, and HP were linked to prognosis.

Gene enrichment

KEGG pathway enrichment analyses of survival-related interaction genes were performed with the “clusterProfiler” R package. GSEA was performed using the “enrichplot” R package.

RNA isolation and quantitative RT-PCR assay

Tianjin Medical University Cancer Institute & Hospital gathered seven paired SCLC and normal tissue samples from patients during a surgical surgery. The Ethics Committee of Tianjin Medical University Cancer Institute & Hospital approved the investigations, and participants provided signed informed permission. Total RNA was extracted from SCLC cells or tissues using the SPARKeasy Improved Tissue/Cell RNA Kit (SparkJade, AC0202). The manufacturer’s instructions were followed to synthesize complementary DNA (cDNA) using the RevertAid™ First Strand cDNA Synthesis Kit (Thermo Fisher Scientific). The SYBR Green PCR kit (Takara Bio, Otsu, Japan) was used to perform qRT-PCR using a Step One Real-Time PCR machine (Thermo Fisher Scientific). Gene expression levels were measured using the 2-△△CT technique. The Primers are: MUC1: F-TGCCGCCGAAAGAACTACG.

R- TGGGGTACTCGCTCATAGGAT

COLEC12: F-AATCCTTCGGTTACAAGCGGT.

R-ACTGTGATTGTTAGCAAGGCAC.

β-ACTIN:F-TGACGTGGACATCCGCAAAG.

R-CTGGAAGGTGGACAGCGAGG

Statistical analysis

Experimental data from triplicate independent trials are presented as mean ± SD. Parametric comparisons employed two-tailed Student's t-test (two groups) or one-way Analysis of Variance with Tukey's post-hoc correction (multiple groups). Statistical significance (p≤0.05) is indicated by asterisks, with "ns" denoting non-significance (p>0.05).

Results

DEGs identification

This study acquired four SCLCs microarray datasets from the GEO database to be used as the experimental group. Detailed information regarding the datasets included can be found in Table 1. Subsequently, the ensemble ID was transformed into gene symbol using Perl. The expression values of each gene in its corresponding dataset were rectified and merged utilizing R version 4.3.1. In addition, batch effects were eradicated via Principal Component Analysis (PCA). Figure 1A, B display images before and after batch correction, respectively. It is evident that, following batch correction, all samples in the datasets have achieved satisfactory uniformity.

Table 1.

Dataset information

Accession number Number of SCLC Number of normal lung Platform Array
GSE40275 15 43 GPL15974 Human Exon 1.0 ST Array
GSE30219 21 14 GPL570 Affymetrix Human Genome U133 Plus 2.0 Array
GSE149507 18 18 GPL23270 Affymetrix Human Genome U133 Plus 2.0 Array
GSE60052 79 7 GPL11154 Illumina HiSeq 2000 (Homo sapiens)

Fig. 1.

Fig. 1

The differentially expressed genes(DEGs) in three GEO datasets. (A) Before batch correction. (B) Afterbatch correction. (C) The DEGs are exhibited in volcano plots. (D) The DEGs are exhibited in a visualized heatmap

The distribution of upregulated or downregulated genes is exhibited in volcano plots and in a visualized heatmap (Fig. 1C, D). A total of 1956 differentially expressed genes (DEGs) were identified, consisting of 369 upregulated genes and 529 downregulated genes. The details can be found in the Supplementary Materials, Table S1. The cutoff criteria for the differentially expressed genes (DEGs) were |log2 fold change (FC)|> 1.5 and P-value < 0.001.

Mendelian randomization analysis

Mendelian randomization (MR) is a technique used to investigate the connections between potential risk factors and health outcomes. It involves employing genetic variants related to the specific exposure being studied, usually derived from genome-wide association studies (GWAS), as instrumental variables. Five methods for MR analysis were employed, including MR Egger, simple mode, weighted median, weighted mode methods, and inverse variance weighted (IVW) method. The OR values were obtained using these methods, and genes with consistent OR direction were retrieved across all five methods. An examination of heterogeneity and pleiotropy testing was conducted on the instrumental variable SNP. Genes with a P-value < 0.05 were selected for the IVW method, whereas genes with a P-value < 0.05 were eliminated for the pleiotropy. Ultimately, a total of 316 genes with an OR > 1 and 319 genes with an OR < 1 were identified for SCLC in ieu-a-988, as shown in Supplementary Table S2. Meanwhile, a total of 100 genes with an OR > 1 and 102 genes with an OR < 1 were identified for SCLC in GCST004746, as shown in Supplementary Table S3. Based on the results of Mendelian analysis and the intersection genes of differentially expressed genes, we obtained the intersection genes shown in Fig. 2A, B. Among them, 2 genes that were upregulated and OR > 1 were PSRC1 and PSAT1, while 4 genes that were downregulated and OR < 1 were DHCR24, COLEC12, MUC1, and HP. To further clarify the chromosomal distribution of the aforementioned genes, we visualized the co-expressed genes (Fig. 2C). Afterwards, the MR findings of these 6 genes were shown visually, and the OR and p values of each gene were shown using a forest plot, as shown in Fig. 3A. Two of the upregulated co-expressed genes showed a statistically significant positive causal relationship with SCLC, indicating that the upregulation of genes may increase the risk of developing SCLC. All 4 co-expressed genes that were downregulated had a substantial negative causal association with SCLC. Furthermore, genes that were downregulated act as a preventive factor against SCLC. Lower expression of these genes is associated with a decreased likelihood of getting SCLC. The results of the heterogeneity tests and pleiotropy tests for the co-expressed genes all indicated a p value > 0.05, suggesting no statistical significance. Therefore, there is no need to evaluate the impact of heterogeneity and pleiotropy on the results. The leave-one-out sensitivity analysis revealed that the effect sizes of the included independent variables (IVs) were similar to the overall effect size, indicating the analysis’s robustness.

Fig. 2.

Fig. 2

Mendelian randomization analysis and DEGs. A Venn diagram illustrating up-regulated genes and MR-identified OR > 1. B Venn diagram illustrating down-regulated genes and MR-identified OR < 1 C Circos plot of co-expressed genes

Fig. 3.

Fig. 3

Candidate hub genes. A MR forest plot for co-expressed genes. B KEGG enrichment analysis of candidate hub genes. C GO.circlize for candidate hub genes. D KEGG enrichment analysis of candidate hub genes

GO function and KEGG pathway enrichment analyses

To explore the function of the candidate hub genes in SCLC development, we carried out GO and KEGG pathway enrichment analyses. KEGG enrichment analysis indicated the candidate hub genes mainly affect Steroid biosynthesis signaling pathway, as shown in Fig. 3B. GO enrichment analysis revealed the candidate hub genes mainly affect organic hydroxy compound biosynthetic process and defense response to bacterium, as shown in Fig. 3C, D. Detailed data can be found in Supplementary Table S4.

Infiltration of immune cells

Since the tumor immune microenvironment plays an important role in tumorigenesis, we identified the abundance ratios (p value < 0.05) of 22 types of immune cells in normal and SCLC samples with the CIBERSORT algorithm. The results from three GEO datasets showed that macrophages M0 and mast cells resting were substantially higher than those in normal controls; M0 macrophages may contribute to tumor angiogenesis under specific conditions, supplying vital nutrients and oxygen for the proliferation and spread of lung cancer. Mast cells resting may promote the growth and metastasis of lung cancer by releasing pro-angiogenic factors and matrix metalloproteinases. High expression of both factors indicates a poor prognosis for patients with small cell lung cancer. B cells naïve, plasma cells, and T cells follicular helper in SCLC tissues are lower than those in normal controls (Fig. 4B). Naive B cells focus on recognizing antigens, presenting them, and forming immune memory, while plasma cells secrete antibodies to inhibit cancer growth and work with other immune cells to enhance anti-tumor responses. T follicular helper (Tfh) cells play a pivotal role in regulating the humoral immune response. A decline in these three parameters always implies a bad prognosis for patients with small cell lung cancer. Additionally, Fig. 4C shows that the co-expressed genes COLEC12 and MUC1 have negative correlations with T cells follicular helper and positive correlations with mast cells resting. However, COLEC12 has a positive correlation with T cells CD4 memory activated. MUC1 has a positive correlation with macrophages M0. T cells CD4 memory activated, M0 macrophages, and mast cells resting all play important roles in the immune system, working together to maintain immune defense and bodily homeostasis. Figure 4C shows the association between other co-expressed genes and immune cells. These patterns imply that SCLC tumors evade immune surveillance by suppressing anti-tumor immune cells (e.g., Naive B cells, Tfh cells) while enriching pro-tumorigenic subsets (e.g., M0 macrophages, mast cells). COLEC12 and MUC1 may serve as biomarkers of immune dysfunction, guiding immunotherapy strategies.

Fig. 4.

Fig. 4

Analysis of Immune Cell Infiltration in SCLC. A Stacked histogram of the proportions of immune cells between the SCLC group and the control group. B Box plot showing the comparison of 22 types of immune cells with GEO data. C Heatmap showing the correlation between 22 types of immune cells and hub genes

Verification of target genes other GEO databases

We validated the target genes for the aforementioned analysis using GSE60052 and the merged dataset (GSE149507, GSE40275, and GSE30219). PSAT1, DHCR24, COLEC12, and MUC1 exhibited a consistent pattern across both analyses, as illustrated in Fig. 5A,B.

Fig. 5.

Fig. 5

Validation group differential analysis. AThe Validation of hub genes with GSE149507,GSE30219 and GSE40275. B The Validation of hub genes with GSE60052

Machine learning-based integration

We use 15 machine learning methods to create prediction models that improve accuracy, prevent overfitting, and increase model stability. After the calculations, “the Support Vector MachineCross-Validation(SVM-CV)" yielded the best predictive performance among the 150 algorithms tested. It was constructed by applying the “the SVM-CV” algorithm to a SCLC consisting of 6 genes (PSRC1, PSAT1, DHCR24, COLEC12, MUC1, and HP) (Fig. 6A). The predicted probability fitting curve illustrates the alignment between the forecasted likelihoods and the actual results (Fig. 6B–E). The ROC values for diagnosing SCLC in four sets (GSE149507 cohort, GSE30219 cohort, GSE40275 cohort, and GSE60052 cohort) were 0.917, 1.000, 0.988, and 0.884, respectively (Fig. 7A). The model performed well in all testing sets, indicating that the “the SVM-CV” algorithm has a low risk of overfitting.

Fig. 6.

Fig. 6

Machine Learning-Based Integration. A The ROC values for hub genes using the 150 algorithms on the four sets are shown in a heatmap. The fitting curve of the predicted probability for B GSE149507. C GSE30219. D GSE40275. E GSE60052

Fig. 7.

Fig. 7

The ROC values and survival analysis for SCLC. A The ROC values for diagnosing SCLC in four sets. B survival analysis for hub genes

Perform survival analysis on the results

45 samples were gathered from the GEO database (GSE60052), representing survival data. Out of them, 79 samples had SCLC and 7 samples were normal. Initially, we merged the survival and gene expression data, and subsequently performed survival analysis on the target genes. We discovered that COLEC12 and MUC1 had a strong connection to survival (p < 0.05), as shown in Fig. 7B. Furthermore, an increased level of expression is associated with an extended duration of survival. All of these genes OR < 1, which indicates that they are protective genes. Thus, our MR findings align with our own, thereby providing further evidence to support the dependability of our results.

Gene enrichment and pathway analysis

GSEA enrichment analyses of the survival-related interaction genes (COLEC12 and MUC1) revealed significant enrichment in pathways related to autoimmune thyroid disease, complement and coagulation cascades, cytokine–cytokine receptor interaction, and hematopoietic cell lineage in the highly expressed group, as shown in Fig. 8A, B. Meanwhile, in the lowly expressed group, the enrichment in pathways related to cell cycle, DNA replication, RNA degradation, and spliceosome, as shown in Fig. 8C, D. These mechanisms are all linked to the stability, correctness, and regulatory expression of genetic information in living organisms.

Fig. 8.

Fig. 8

The role of COLEC12 and MUC1. A GSEA enrichment analysis s in the high COLEC12 expressing groups. B GSEA enrichment analysis s in the low COLEC12 expressing groups. C GSEA enrichment analysis s in the high MUC1 expressing groups. D GSEA enrichment analysis s in the low MUC1 expressing groups. E Differential expression of COLEC12 and MUC1 iBeas-2b and H446. F The Validation of COLEC12 in normal and tumour tissues. G The Validation of MUC1 in normal and tumour tissues

The experiment of survival-related interaction genes

In order to ascertain the dependability of the findings, we carried out further experiments. For this work, we cultivated normal lung epithelial cells known as BEAS-2B and a cell line of SCLC called H446. The expression of COLEC12 and MUC1 is low in H446 (Fig. 8E). Additionally, we obtained 7 pairs of samples of both normal and tumor tissues from patients diagnosed with SCLC. In those samples, we observed the same expression trends (Fig. 8F, G). As expected, the expression of COLEC12 and MUC1 is low in 7 pairs of SCLC tissue samples.

Discussion

Lung cancer is a common and deadly disease that affects people all over the world, but SCLC is more invasive than other types of lung cancer. Patients with SCLC exhibit sensitivity to treatment initially, but thereafter acquire resistance [34, 35], and frequently experience metastases in the early stages [36, 37]. Immunotherapy exhibits efficacy, but below anticipated levels, resulting in a poor five-year survival rate. Emerging targets are constantly being discovered, such as DLL, CD3, PARP [38, 39]. However, there are still no effective targets developed for SCLC currently. Despite a large number of clinical trials, results have been disappointing, and there are still no approved targeted drugs for SCLC [10, 15, 40, 41]. Novel technologies and methodologies are required to ascertain biomarkers for SCLC. Recent large-scale genetic investigations, such as GWAS [42, 43] have revealed several risk variants associated with SCLC. An early comprehension of the genetic variant information of small cells, timely prediction and intervention, and starving for surgical options are essential for SCLC patients. The timely identification of SCLC, as well as the exploration of novel diagnostic biomarkers and therapeutic agents, represents critical challenges that necessitate attention. In this study, we analyzed a large RNA dataset to assess genes as potential pathogenic factors and biomarkers for SCLC, followed by preliminary analysis and external validation. We found PSRC1, PSAT1, DHCR24, COLEC12, MUC1, and HP linked to SCLC. Following the validation of the risk genes COLEC12 and MUC1, it was observed that the two genes achieved a statistically significant P-value of less than 0.05 in the survival analysis.

Collectin subfamily member 12 (COLEC12), a transmembrane scavenger receptor C-type lectin [44]. Acts as a pattern recognition molecule and activates the complement system through alternative pathways [45]. COLEC12 was down-regulated in alveolar macrophages from tuberculosis patients [46], COLEC12 exhibits the capacity to eliminate damage-associated molecular patterns in addition to its innate immune defense capabilities against bacteria and fungi [47], which include a scavenging property against damage-associated molecular pattern [48]. The paucity of reports on COLEC12 research in lung cancer is notable. According to our thorough biological investigation, small cell lung cancer has a good prognosis and underexpressed COLEC12. The reliability of our study is further validated by this discovery, which is consistent with the findings of Mendelian randomization analysis.

MUC1, a key transmembrane mucin, is involved in several malignancies. It is significantly overexpressed in ovarian cancer [49], early-stage metastatic lung cancer [50], and colorectal cancer [51], but underexpressed in lung cancer [52], gastric cancer, and bladder cancer [53]. It is worth mentioning that research into MUC1 in small cell lung cancer has been limited till now. This study goes into this topic and discovers that MUC1 is likewise underexpressed in small cell lung cancer, which is closely associated with a better patient prognosis. This research reveals MUC1’s enormous potential as a diagnostic and prognostic marker for lung cancer, promising to give strong support for clinical decision-making and aid in the formulation of more accurate treatment strategies.

COLEC12 and MUC1 have emerged as promising diagnostic markers and therapeutic targets, offering extensive potential applications in the clinical management of SCLC. To validate these findings, we plan to conduct multi-center studies with large sample sizes and to standardize detection methods, including immunohistochemistry and quantitative PCR, to ensure the accuracy and reproducibility of test results. This research aims to incorporate the assessment of COLEC12 and MUC1 expression levels in tumor tissue samples obtained via surgery or biopsy into future diagnostic and therapeutic protocols, utilizing immunohistochemical staining or quantitative PCR techniques. Such assessments will enable risk stratification of SCLC patients and inform the development of personalized treatment plans. Based on gene expression levels, patients can be classified into high-risk and low-risk categories, with the former potentially necessitating more intensive treatment strategies. Furthermore, by analyzing tumor tissue collected before and after treatment, the expression status of COLEC12 and MUC1 may serve as predictive indicators of treatment response, thereby assisting clinicians in selecting the most effective therapeutic options.

The rapid progression of small cell lung cancer (SCLC) and its frequent diagnosis at advanced stages result in a low rate of surgical resection, leading to the collection of only seven tissue specimens over the three-year research period. This severe shortage of sample quantity not only limits the representativeness of the samples, hindering the ability to account for individual patient variability, but also results in insufficient statistical power. Furthermore, this scarcity of samples precludes the possibility of conducting in-depth subgroup analyses. Consequently, it imposes significant constraints on subsequent research endeavors, and it also obstructs the translation of research findings into clinical practice. Furthermore, Limited by the current laboratory resources, this study has not yet incorporated additional small cell lung cancer (SCLC) cell lines. Subsequent research will expand the cellular models to include, thereby enhancing the external validity of the experimental conclusions.

The results of immunological correlation analysis indicate a positive correlation between COLEC12 and MUC1expression levels and T cells CD4 memory activated, M0 macrophages, and mast cells resting. COLEC12 and MUC1 have negative correlations with T cells follicular helper. Survival analysis on the target genes reveals that COLEC12 and MUC1 had a strong connection to survival; elevated expression signifies a positive prognosis.

Conclusion

We performed a differential gene analysis on the Gene Expression Omnibus (GEO) dataset, followed by Mendelian randomization analysis utilizing single nucleotide polymorphisms (SNPs) as instrumental variables to assess the causal relationship between exposure factors and the outcomes of interest, ultimately identifying the co-expressed genes of interest. Following this, GO KEGG analysis, immune cell infiltration analysis, and survival analysis were conducted. We have effectively identified colec12 and muc1 as two critical key genes through a comprehensive and thorough bioinformatics analysis. These two genes are not only of significant importance at the bioinformatics level, but they also have the potential to serve as clinical diagnostic markers and accurately predict patient prognosis. It is anticipated that this discovery will offer new and potent support for clinical decision-making.

Supplementary Information

Below is the link to the electronic supplementary material.

Acknowledgements

We are very grateful for the data provided by databases such as TCGA and GEO. Thanks to the reviewers and editors for their sincere comments.

Author contributions

Hailin Liu, Fangyuan Qu, Guangyao Zhou were involved in the planning and design of the study; Yuechen Cui and Bo Yan conducted a comprehensive review of the relevant literature, completed the data analysis, and generated the visual representations of the final results. Lianmin Zhang, Chenguang Li, Zhenfa Zhang, Tingting Qin, and Qiangzhe Zhang contributed to the revision of the manuscript.

Funding

This study is supported in part by grants from National Natural Science Foundation of China [Grant No. 82373028], Joint Innovation Project of Hospital and Enterprise of Hebei Provincial Health Commission [Grant No. LH20250039], and the Tianjin Key Medical Discipline (Specialty) Construction Project [Grant No.TJYXZDXK-010A].

Data availability

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding authors.

Declarations

Conflict of interest

The authors declare that there is no conflict of interest.

Ethical approval and informed consent

The studies involving human participants were reviewed and approved by Tianjin Medical University Cancer Institute and Hospital Ethical Board, which gave the study its approval. The patients/participants provided their written informed consent to participate in this study.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Hailin Liu, Fangyuan Qu and Guangyao Zhou have contributed equally to this work and share first authorship.

Change history

4/18/2026

The original online version of this article was revised to correct Reference 49.

References

  • 1.Zimmerman S, Das A, Wang S, Julian R, Gandhi L, Wolf J. 2017-2018 Scientific advances in thoracic oncology: small cell lung cancer. J Thorac Oncol. 2019;14(5):768–83. 10.1016/j.jtho.2019.01.022. [DOI] [PubMed] [Google Scholar]
  • 2.Zhang C, Wang H. Accurate treatment of small cell lung cancer: current progress, new challenges and expectations. Biochimica et Biophysica Acta (BBA) Rev Cancer. 2022;1877(5):188798. 10.1016/j.bbcan.2022.188798. [DOI] [PubMed] [Google Scholar]
  • 3.Meder L, König K, Fassunke J, Ozretić L, Wolf J, Merkelbach-Bruse S, et al. Implementing amplicon-based next generation sequencing in the diagnosis of small cell lung carcinoma metastases. Exp Mol Pathol. 2015;99(3):682–6. 10.1016/j.yexmp.2015.11.002. [DOI] [PubMed] [Google Scholar]
  • 4.Wang J, Han F, Ma Y, Yang Y, Wu Y, Han Z, et al. Effect of segmental abutting esophagus-sparing technique to reduce severe esophagitis in limited-stage small-cell lung cancer patients treated with concurrent hypofractionated thoracic radiation and chemotherapy. Cancers. 2023. 10.3390/cancers15051487. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wasamoto S, Imai H, Tsuda T, Nagai Y, Minemura H, Yamada Y, et al. Pretreatment glasgow prognostic score predicts survival among patients administered first-line atezolizumab plus carboplatin and etoposide for small cell lung cancer. Front Oncol. 2022;12:1080729. 10.3389/fonc.2022.1080729. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Shao S, Li Z, Gao W, Yu G, Liu D, Pan F. ADAM-12 as a diagnostic marker for the proliferation, migration and invasion in patients with small cell lung cancer. PLoS ONE. 2014;9(1):e85936. 10.1371/journal.pone.0085936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Jahchan NS, Lim JS, Bola B, Morris K, Seitz G, Tran KQ, et al. Identification and targeting of long-term tumor-propagating cells in small cell lung cancer. Cell Rep. 2016;16(3):644–56. 10.1016/j.celrep.2016.06.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jin Y, Chen Y, Tang H, Hu X, Hubert SM, Li Q, et al. Activation of PI3K/AKT pathway is a potential mechanism of treatment resistance in small cell lung cancer. Clin Cancer Res. 2022;28(3):526–39. 10.1158/1078-0432.Ccr-21-1943. [DOI] [PubMed] [Google Scholar]
  • 9.Zhu L, Qin J. Predictive biomarkers for immunotherapy response in extensive-stage SCLC. J Cancer Res Clin Oncol. 2024;150(1):22. 10.1007/s00432-023-05544-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Yang S, Zhang Z, Wang Q. Emerging therapies for small cell lung cancer. J Hematol Oncol. 2019;12(1):47. 10.1186/s13045-019-0736-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Niu X, Chen L, Li Y, Hu Z, He F. Ferroptosis, necroptosis, and pyroptosis in the tumor microenvironment: perspectives for immunotherapy of SCLC. Semin Cancer Biol. 2022;86:273–85. 10.1016/j.semcancer.2022.03.009. [DOI] [PubMed] [Google Scholar]
  • 12.Meijer JJ, Leonetti A, Airò G, Tiseo M, Rolfo C, Giovannetti E, et al. Small cell lung cancer: novel treatments beyond immunotherapy. Semin Cancer Biol. 2022;86:376–85. 10.1016/j.semcancer.2022.05.004. [DOI] [PubMed] [Google Scholar]
  • 13.Rudin CM, Brambilla E, Faivre-Finn C, Sage J. Small-cell lung cancer. Nat Rev Dis Primers. 2021;7(1):3. 10.1038/s41572-020-00235-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Zheng Z, Liu J, Ma J, Kang R, Liu Z, Yu J. Advances in new targets for immunotherapy of small cell lung cancer. Thorac Cancer. 2024;15(1):3–14. 10.1111/1759-7714.15178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Teicher BA. Targets in small cell lung cancer. Biochem Pharmacol. 2014;87(2):211–9. 10.1016/j.bcp.2013.09.014. [DOI] [PubMed] [Google Scholar]
  • 16.Yang H, Liu D, Zhao C, Feng B, Lu W, Yang X, et al. Mendelian randomization integrating GWAS and eQTL data revealed genes pleiotropically associated with major depressive disorder. Transl Psychiatry. 2021;11(1):225. 10.1038/s41398-021-01348-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wu Y, Wang Z, Yang Y, Han C, Wang L, Kang K, et al. Exploration of potential novel drug targets and biomarkers for small cell lung cancer by plasma proteome screening. Front Pharmacol. 2023;14:1266782. 10.3389/fphar.2023.1266782. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Huang Z, Zheng H, Wang H, Ning H, Che A, Cai C. Identification of potential therapeutic targets for breast cancer using Mendelian randomization analysis and drug target prediction. Environ Toxicol. 2024. 10.1002/tox.24249. [DOI] [PubMed] [Google Scholar]
  • 19.Hu W, Xu Y. Transcriptomics in idiopathic pulmonary fibrosis unveiled: a new perspective from differentially expressed genes to therapeutic targets. Front Immunol. 2024;15:1375171. 10.3389/fimmu.2024.1375171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Davey Smith G, Hemani G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Hum Mol Genet. 2014;23:R89-98. 10.1093/hmg/ddu328. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Shen YY, Gu XK, Zhang RR, Qian TM, Li SY, Yi S. Biological characteristics of dynamic expression of nerve regeneration related growth factors in dorsal root ganglia after peripheral nerve injury. Neural Regen Res. 2020;15(8):1502–9. 10.4103/1673-5374.274343. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Liu H, Zong C, Sun J, Li H, Qin G, Wang X, et al. Bioinformatics analysis of lncRNAs in the occurrence and development of osteosarcoma. Transl Pediatr. 2022;11(7):1182–98. 10.21037/tp-22-253. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Zeng Y, Wang C, Ye Q, Liu G, Zhang L, Wan J, et al. Machine learning model of imipenem-resistant Klebsiella pneumoniae based on MALDI-TOF-MS platform: an observational study. Health Sci Rep. 2023;6(9):e1108. 10.1002/hsr2.1108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Zhu Y, Chen B, Zu Y. Identifying OGN as a biomarker covering multiple pathogenic pathways for diagnosing heart failure: from machine learning to mechanism interpretation. Biomolecules. 2024. 10.3390/biom14020179. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Dai L, Yuan W, Jiang R, Zhan Z, Zhang L, Xu X, et al. Machine learning-based integration identifies the ferroptosis hub genes in nonalcoholic steatohepatitis. Lipids Health Dis. 2024;23(1):23. 10.1186/s12944-023-01988-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Liu Z, Liu L, Weng S, Guo C, Dang Q, Xu H, et al. Machine learning-based integration develops an immune-derived lncRNA signature for improving outcomes in colorectal cancer. Nat Commun. 2022;13(1):816. 10.1038/s41467-022-28421-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Liu H, Yan B, Chen Y, Pang J, Li Y, Zhang Z, et al. Identification of potential prognostic biomarkers associated with monocyte infiltration in lung squamous cell carcinoma. Biomed Res Int. 2022;2022:6860510. 10.1155/2022/6860510. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Chao F, Zhang Y, Lv L, Wei Y, Dou X, Chang N, et al. Extracellular vesicles derived circSH3PXD2A inhibits chemoresistance of small cell lung cancer by miR-375-3p/YAP1. Int J Nanomedicine. 2023;18:2989–3006. 10.2147/IJN.S407116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Xin J, Gu D, Chen S, Ben S, Li H, Zhang Z, et al. SUMMER: a Mendelian randomization interactive server to systematically evaluate the causal effects of risk factors and circulating biomarkers on pan-cancer survival. Nucleic Acids Res. 2023;51(D1):D1160–7. 10.1093/nar/gkac677. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Wang Y, Gu X, Wang X, Zhu W, Su J. Exploring genetic associations between allergic diseases and indicators of COVID-19 using mendelian randomization. iScience. 2023;26(6):106936. 10.1016/j.isci.2023.106936. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Lei P, Xu W, Wang C, Lin G, Yu S, Guo Y. Mendelian randomization analysis reveals causal associations of polyunsaturated fatty acids with sepsis and mortality risk. Infect Dis Ther. 2023;12(7):1797–808. 10.1007/s40121-023-00831-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Chen B, Khodadoust MS, Liu CL, Newman AM, Alizadeh AA. Profiling tumor infiltrating immune cells with CIBERSORT. Methods Mol Biol. 2018;1711:243–59. 10.1007/978-1-4939-7493-1_12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Guan M, Jiao Y, Zhou L. Immune infiltration analysis with the CIBERSORT method in lung cancer. Dis Markers. 2022;2022:3186427. 10.1155/2022/3186427. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Furuta M, Kikuchi H, Shoji T, Takashima Y, Kikuchi E, Kikuchi J, et al. DLL3 regulates the migration and invasion of small cell lung cancer by modulating Snail. Cancer Sci. 2019;110(5):1599–608. 10.1111/cas.13997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Jin Y, Chen Y, Qin Z, Hu L, Guo C, Ji H. Understanding SCLC heterogeneity and plasticity in cancer metastasis and chemotherapy resistance. Acta Biochim Biophys Sin (Shanghai). 2023;55(6):948–55. 10.3724/abbs.2023080. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Uda S, Yamada T, Yoshimura A, Goto Y, Yoshimine K, Nakamura Y, et al. Clinical impact of amrubicin monotherapy in patients with relapsed small cell lung cancer: a multicenter retrospective study. Transl Lung Cancer Res. 2022;11(9):1847–57. 10.21037/tlcr-22-160. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Mohan S, Foy V, Ayub M, Leong HS, Schofield P, Sahoo S, et al. Profiling of circulating free DNA using targeted and genome-wide sequencing in patients with SCLC. J Thorac Oncol. 2020;15(2):216–30. 10.1016/j.jtho.2019.10.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Rudin CM, Reck M, Johnson ML, Blackhall F, Hann CL, Yang JC, et al. Emerging therapies targeting the delta-like ligand 3 (DLL3) in small cell lung cancer. J Hematol Oncol. 2023;16(1):66. 10.1186/s13045-023-01464-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Pillai RN, Owonikoko TK. Small cell lung cancer: therapies and targets. Semin Oncol. 2014;41(1):133–42. 10.1053/j.seminoncol.2013.12.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Yuan M, Zhao Y, Arkenau HT, Lao T, Chu L, Xu Q. Signal pathways and precision therapy of small-cell lung cancer. Signal Transduct Target Ther. 2022;7(1):187. 10.1038/s41392-022-01013-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Li L, Ng SR, Colon CI, Drapkin BJ, Hsu PP, Li Z, et al. Identification of DHODH as a therapeutic target in small cell lung cancer. Sci Transl Med. 2019. 10.1126/scitranslmed.aaw7852. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Smith-Byrne K, Cerani A, Guida F, Zhou S, Agudo A, Aleksandrova K, et al. Circulating isovalerylcarnitine and lung cancer risk: evidence from Mendelian randomization and prediagnostic blood measurements. Cancer Epidemiol Biomarkers Prev. 2022;31(10):1966–74. 10.1158/1055-9965.Epi-21-1033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Abifadel M, Boileau C. Genetic and molecular architecture of familial hypercholesterolemia. J Intern Med. 2023;293(2):144–65. 10.1111/joim.13577. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Sun X, Zhang Q, Shu P, Lin X, Gao X, Shen K. COLEC12 promotes tumor progression and is correlated with poor prognosis in gastric cancer. Technol Cancer Res Treat. 2023;22:15330338231218164. 10.1177/15330338231218163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Yu H, Si G, Si F. Mendelian randomization validates the immune landscape mediated by aggrephagy in esophageal squamous cell carcinoma patients from the perspectives of multi-omics. J Cancer. 2024;15(7):1940–53. 10.7150/jca.93376. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Lavalett L, Rodriguez H, Ortega H, Sadee W, Schlesinger LS, Barrera LF. Alveolar macrophages from tuberculosis patients display an altered inflammatory gene expression profile. Tuberculosis. 2017;107:156–67. 10.1016/j.tube.2017.08.012. [DOI] [PubMed] [Google Scholar]
  • 47.Ohtani K, Suzuki Y, Eda S, Kawai T, Kase T, Keshi H, et al. The membrane-type collectin CL-P1 is a scavenger receptor on vascular endothelial cells. J Biol Chem. 2001;276(47):44222–8. 10.1074/jbc.M103942200. [DOI] [PubMed] [Google Scholar]
  • 48.Zhang J, Li A, Yang CQ, Garred P, Ma YJ. Rapid and efficient purification of functional collectin-12 and its opsonic activity against fungal pathogens. J Immunol Res. 2019;2019:9164202. 10.1155/2019/9164202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Dong Y, Walsh MD, Cummings MC, Wright RG, Khoo SK, Parsons PG, et al. Expression of MUC1 and MUC2 mucins in epithelial ovarian tumours. J Pathol. 1997;183(3):311–7. 10.1002/(SICI)1096-9896(199711)183:3<311::AID-PATH917>3.0.CO;2-2 [DOI] [PubMed]
  • 50.Kaira K, Okumura T, Nakagawa K, Ohde Y, Takahashi T, Murakami H, et al. MUC1 expression in pulmonary metastatic tumors: a comparison of primary lung cancer. Pathol Oncol Res. 2012;18(2):439–47. 10.1007/s12253-011-9465-9. [DOI] [PubMed] [Google Scholar]
  • 51.Cox KE, Liu S, Lwin TM, Hoffman RM, Batra SK, Bouvet M. The mucin family of proteins: candidates as potential biomarkers for colon cancer. Cancers. 2023. 10.3390/cancers15051491. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Xu T, Li D, Wang H, Zheng T, Wang G, Xin Y. MUC1 downregulation inhibits non-small cell lung cancer progression in human cell lines. Exp Ther Med. 2017;14(5):4443–7. 10.3892/etm.2017.5062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Nielsen TO, Borre M, Nexo E, Sorensen BS. Co-expression of HER3 and MUC1 is associated with a favourable prognosis in patients with bladder cancer. BJU Int. 2015;115(1):163–5. 10.1111/bju.12658. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

SCLCs data were acquired in Gene Expression Omnibus (GEO) database (https://www.ncbi.nlm.nih.gov/geo/) [27] (accession number GSE149507, GSE40275, GSE30219, GSE60052). The GSE60052 [28] dataset includes clinical survival data that can be used to plot future survival curves.

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding authors.


Articles from Clinical & Translational Oncology are provided here courtesy of Springer

RESOURCES