Abstract
Background
Lung adenocarcinoma (LUAD) is the most common subtype of lung cancer and exhibits complex heterogeneity. Despite extensive research on driver gene mutations such as in EGFR and TP53, resistance mechanisms and co-mutations underscore the need to elucidate the molecular mechanisms of LUAD and to discover novel biomarkers for personalized treatment.
Methods
We integrated transcriptomic data from 1398 samples (1160 LUAD and 238 normal tissues) obtained from the GEO and TCGA databases. Analytical procedures included differential expression, functional enrichment (GO/KEGG/GSEA), and immune cell infiltration assessment. We applied 113 machine learning algorithms to identify core genes and develop a diagnostic model. The relationship between core gene expression and prognosis was assessed using TCGA data, and downregulation of these genes was experimentally verified by RT-PCR in A549 and Beas-2B cell lines.
Results
Our analysis revealed significant enrichment of the “cytoskeleton in muscle cells” pathway and activation of pyrimidine metabolism in LUAD. Immune profiling showed reduced proportions of monocytes and resting mast cells. We identified four core genes—FGD5, LRRC36, C8B, and MYOC—that were consistently downregulated in LUAD. A diagnostic model based on these genes demonstrated strong predictive performance, and low expression of each core gene was correlated with poor patient prognosis.Multivariate Cox regression confirmed their status as independent prognostic factors.
Conclusions
This study identifies FGD5, LRRC36, C8B, and MYOC as novel core genes and independent prognostic biomarkers in LUAD, and underscores the potential role of the “cytoskeleton in muscle cells” pathway. Our multi-faceted strategy provides new insights into LUAD pathogenesis and yields promising biomarker candidates for improving clinical diagnostics.
Supplementary Information
The online version contains supplementary material available at 10.1007/s12672-026-04505-3.
Keywords: Lung adenocarcinoma (LUAD), Diagnostic biomarkers, Machine learning, FGD5, LRRC36, C8B
Introduction
As early as 2020, a study published in The Lancet indicated that lung adenocarcinoma had become the most common histological subtype of lung cancer globally. In most countries, the incidence of adenocarcinoma in males has surpassed that of squamous cell carcinoma, while in females, the incidence of adenocarcinoma is consistently higher than that of squamous cell carcinoma across all nations [1]. The Global Burden of Disease Study in 2021 further emphasized that with the aging of the global population, the incidence and mortality of lung cancer continue to rise, establishing it as a leading cause of cancer-related deaths in both China and the United States [2–5].
The pathogenesis of lung adenocarcinoma is closely associated with complex biological mechanisms, including mutations in genes such as EGFR, KRAS, BRAF, and TP53, as well as rearrangements in ALK and ROS1 [6–8]. These genetic alterations play significant roles in tumorigenesis and progression [9–11]. Studies have shown that EGFR exhibits the highest mutation rate (61.36%), followed by TP53 (60.61%), BRAF (22.73%), KRAS (15.91%), and ROS1 (15.91%) [12]. Although numerous targeted drugs have been developed against key driver genes in lung adenocarcinoma, effective treatment remains unattainable for many patients due to the heterogeneity of genetic mutations, the emergence of resistance mechanisms, and the complexity of co-occurring mutations. Therefore, a deeper investigation into the genetic mechanisms of lung adenocarcinoma and the identification of novel biomarkers represent an urgent need for developing more effective personalized therapeutic strategies.
This study aims to comprehensively integrate transcriptomic data from lung adenocarcinoma. We utilized microarray data (representing first-generation sequencing technology) from the GEO database and high-throughput data (from next-generation sequencing) from the TCGA database to identify novel core genes and functional pathways. The expression changes of these core genes were subsequently validated using RT-PCR experiments. Furthermore, we constructed a diagnostic model for lung adenocarcinoma. The strengths of our study include the combined validation from two complementary databases, a large sample size (totaling 1398 samples, comprising 238 normal lung tissues and 1160 lung adenocarcinoma tissues), and the application of a comprehensive suite of machine learning algorithms (totaling 113 methods). We believe that these findings will provide new and profound insights into gene expression and the underlying molecular mechanisms of lung adenocarcinoma.
Materials and methods
Acquisition and analysis of lung adenocarcinoma data
The workflow of this study is illustrated in Fig. 1. Data utilized in this study were primarily sourced from the Gene Expression Omnibus (GEO, https://www.ncbi.nlm.nih.gov/geo/) and The Cancer Genome Atlas (TCGA, https://portal.gdc.cancer.gov/) databases. For the GEO database, our search strategy employed the keywords “Lung adenocarcinoma” and “normal”, with the organism filter set to *Homo sapiens* and the study type limited to “Expression profiling by array” or “Expression profiling by high throughput sequencing”. To ensure data validity, we included only datasets containing at least five samples each of normal lung tissue and lung adenocarcinoma tissue. Ultimately, we screened and downloaded eight lung adenocarcinoma-related datasets; their detailed information is summarized in Table 1. From the TCGA database, we selected Project TCGA-LUAD under the TCGA Program, finally acquiring transcriptome data for 53 normal control tissues and 481 lung adenocarcinoma tissues.
Fig. 1.
Workflow of the comprehensive transcriptomic analysis for lung adenocarcinoma
Table 1.
Sample information of the lung adenocarcinoma-related datasets
| Dataset | Platform ID | Normal lung tissues | Lung adenocarcinoma tissues | Species |
|---|---|---|---|---|
| GSE118370 | GPL570 | 6 | 6 | Human |
| GSE136043 | GPL13497 | 5 | 5 | Human |
| GSE140797 | GPL13497 | 7 | 7 | Human |
| GSE10072 | GPL96 | 49 | 58 | Human |
| GSE32863 | GPL6884 | 58 | 58 | Human |
| GSE31210 | GPL570 | 20 | 226 | Human |
| GSE30219 | GPL570 | 14 | 293 | Human |
| GSE7670 | GPL96 | 26 | 26 | Human |
For the subsequent machine learning analysis, the gene expression matrices from the eight GEO datasets were merged. To address potential batch effects arising from different experimental batches, we applied the ComBat algorithm from the sva R package to harmonize the combined dataset prior to model training This integration and batch-effect correction of multiple microarray datasets represents an established approach to strengthen the reliability of biomarker discovery. For model training and validation, the batch-corrected dataset GSE32863 was designated as the training set, while the remaining seven GEO datasets served as the initial validation set. Additionally, TCGA lung adenocarcinoma data were used as an independent cohort for secondary validation. The integration of diverse transcriptomic data from both microarray and RNA-seq platforms is a established approach to strengthen the reliability of biomarker discovery [13].
Differential gene expression analysis
We employed the “limma” R package for microarray data and the “DESeq2” package for high-throughput sequencing data to identify differentially expressed genes (DEGs) between normal lung tissues and lung adenocarcinoma tissues. DEGs were selected based on an adjusted p-value < 0.05 and an absolute log2 fold change (|log2FC|) ≥ 1.These stringent thresholds are widely adopted in transcriptomic studies to ensure the identification of robust and biologically relevant DEGs [14].The “ggplot2” and “pheatmap” packages were subsequently used to generate volcano plots and heatmaps, respectively, for visualizing significantly up- and down-regulated DEGs.
GO, KEGG enrichment and gene set enrichment analysis (GSEA)
Gene Ontology (GO) analysis, Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis, and Gene Set Enrichment Analysis (GSEA) were performed on all eight GEO datasets (including both training and validation sets) using the “clusterProfiler” R package. Terms or gene sets with an adjusted p-value < 0.05 were considered statistically significant.
Immune cell infiltration analysis
To assess the infiltration levels of various immune cell types in lung adenocarcinoma tissues across the eight GEO datasets, we utilized the “CIBERSORT” algorithm. Bar plots were generated to display the proportions of different immune cells in LUAD versus normal lung tissues, and violin plots were used to visually compare the differences in immune cell infiltration between LUAD patients and normal controls.
Screening of core genes using 113 machine learning algorithms and diagnostic model construction
Based on the DEGs identified from the training set (GSE32863), this study applied a combination of 113 machine learning algorithms to conduct an in-depth analysis across all datasets for screening the most critical core genes. This algorithmic combination included 12 base algorithms and their various ensembles. The 12 base algorithms were: LASSO regression, Support Vector Machine (SVM), Random Forest (RF), Generalized Linear Model with Boostin (glmBoost), Partial Least Squares Regression (plsRglm), Stepwise Generalized Linear Model (Stepglm), Elastic Net regression (Enet), Linear Discriminant Analysis (LDA), Gradient Boosting Machine (GBM), Ridge regression, Extreme Gradient Boosting (XGBoost), and Naive Bayes. The R packages employed included: openxlsx, seqinr, plyr, randomForestSRC, glmnet, plsRglm, gbm, caret, mboost, e1071, BART, MASS, snowfall, xgboost, ComplexHeatmap, RColorBrewer, and pROC. To enhance model performance, robustness, and reduce overfitting risk, we implemented: (1) hyperparameter tuning via grid search and random search, (2) an ensemble learning strategy that weighted and averaged prediction results from individual algorithms, and (3) K-fold cross-validation to evaluate the effectiveness of the algorithm combinations. The performance was assessed by calculating the Area Under the Receiver Operating Characteristic Curve (AUC) for each algorithm across all datasets, visualized via heatmaps, to identify the best-performing algorithm combination. The ability of the core genes selected by the optimal algorithm to distinguish lung adenocarcinoma from normal lung tissue was evaluated by plotting ROC curves and gene expression bar plots for these genes. To assess the potential clinical applicability of the core genes, we constructed a diagnostic model based on them. A nomogram was plotted using the “rms” package to visually demonstrate the importance of each core gene. A calibration curve was generated to evaluate the consistency between the model-predicted probabilities and the actual outcomes. Furthermore, Decision Curve Analysis (DCA) was performed to estimate the net clinical benefit of the model.
Secondary validation of core genes using TCGA database
Following the initial screening of core genes from the eight GEO datasets, we performed differential expression analysis on the TCGA dataset (53 normal controls vs. 481 LUAD tissues) to verify whether these core genes remained differentially expressed in an independent database. Additionally, Kaplan-Meier survival curves were plotted for the core genes to determine their impact on patient survival prognosis. To further evaluate the independent prognostic value of these genes while accounting for potential confounding factors, we performed both univariate and multivariate Cox proportional hazards regression analyses. Clinical variables including age, gender, tumor stage, molecular subtype, and smoking history were incorporated into the multivariate model to assess whether the core genes were independently associated with overall survival.
Experimental validation of core gene expression by RT-PCR
To validate the actual expression levels of the core genes in lung adenocarcinoma, RT-PCR was performed using the human bronchial epithelial cell line Beas-2B as the control group and the human lung adenocarcinoma cell line A549 as the experimental group. The human lung adenocarcinoma A549 cell line (Item No. BFN60800665) and the human normal lung epithelial BEAS-2B cell line (Item No. BFN6080086) were purchased from the Shanghai Cell Bank (Shanghai, China). Both cell lines were cultured in DMEM high-glucose medium supplemented with 10% fetal bovine serum (FBS) and 1% penicillin-streptomycin (containing 100 U/mL penicillin and 100 µg/mL streptomycin) at 37 °C in a 5% CO₂ atmosphere. Total RNA was extracted from cell samples using TRIzol reagent (Invitrogen) according to the manufacturer’s instructions. Subsequently, 1000 ng of total RNA was reverse transcribed into cDNA using the Evo M-MLV RT Kit (Accurate Biology, China). Quantitative real-time PCR (qRT-PCR) analysis was conducted using the SYBR Green Premix Pro Taq HS qPCR Kit (Accurate Biology, China) following the manufacturer’s standard protocol. Relative mRNA expression levels were calculated using the 2–ΔΔCt method, with GAPDH serving as the internal reference control. All samples were assayed in quintuplicate (n = 5). The primer sequences used in this study are listed in Table 2.
Table 2.
Quantitative reverse-transcription PCR primer sequences
| Genes | Forward (5’−3’) | Reverse (5’−3’) |
|---|---|---|
| GAPDH | GCAAATTCCATGGCACCGT | TCGCCCCACTTGATTTTGG |
| FGD5 | GACCTCCTCCTGCCTCACATC | GCCTTCCTCATCATCCTCCTCTG |
| C8B | CAGATGCGGAGTGTGGATGTTAC | GAGAGGGCTGGAGCAAGTAGG |
| LRRC36 | CAGACACAGGAAGTAGCAAGAAGG | GTGGACTGACGAGGCGAGTAG |
| MYOC | CAGTCAGTCGCCAATGCCTTC | GATACCTGTGCCTGTGTCATAAGC |
Results
Identification of differentially expressed genes (DEGs) in LUAD
Volcano plots illustrating the DEGs for each GEO microarray dataset are presented in Figs. 2a-h. We identified 468 DEGs in GSE10072, 816 in GSE118370, 1206 in GSE136043, 2088 in GSE140797, 957 in GSE32863, 1823 in GSE31210, 1724 in GSE30219, and 1222 in GSE7670. Corresponding heatmaps for each dataset are provided in Supplementary File 1.
Fig. 2.
Differentially Expressed Gene (DEG) analysis of eight datasets from the GEO database. A–H Volcano plots for the eight GEO datasets, respectively. Red dots represent genes significantly up-regulated in lung adenocarcinoma, green dots represent genes significantly down-regulated, and blue dots represent non-DEGs
Functional enrichment and GSEA reveal common pathway alterations in LUAD
GO enrichment analysis revealed remarkable consistency across all datasets (Fig. 3a-h, details in Table 3). Biological processes were significantly enriched for terms including extracellular matrix organization and cell-matrix adhesion, highlighting the potential importance of the tumor microenvironment. Enrichment in endothelial cell differentiation, vasculogenesis, and regulation of vascular permeability suggested active involvement of angiogenesis and invasive growth. Processes related to epithelial cell proliferation and wound healing indicated mechanisms underlying self-renewal and aggressive proliferation. For cellular components, the collagen-containing extracellular matrix was significantly enriched. Molecular functions were notably enriched in glycosaminoglycan binding and integrin binding, emphasizing the role of cell-ECM interactions in tumor progression.
Fig. 3.
Gene Ontology (GO) analysis of the training and validation sets. A–H Bar plots display the significantly enriched GO terms (Biological Process, Cellular Component, Molecular Function) for each dataset. Only the top 10 most significantly enriched terms in each category are shown
Table 3.
GO terms enriched concurrently across all datasets
| ONTOLOGY | ID | Description |
|---|---|---|
| BP | GO:0030198 | Extracellular matrix organization |
| BP | GO:0043062 | Extracellular structure organization |
| BP | GO:0045229 | External encapsulating structure organization |
| BP | GO:0072001 | Renal system development |
| BP | GO:0001822 | Kidney development |
| BP | GO:0003018 | Vascular process in circulatory system |
| BP | GO:0031589 | Cell-substrate adhesion |
| BP | GO:0050673 | Epithelial cell proliferation |
| BP | GO:1,903,034 | Regulation of response to wounding |
| BP | GO:0010810 | Regulation of cell-substrate adhesion |
| BP | GO:0032835 | Glomerulus development |
| BP | GO:0003206 | Cardiac chamber morphogenesis |
| BP | GO:0003007 | Heart morphogenesis |
| BP | GO:0001570 | Vasculogenesis |
| BP | GO:0061041 | Regulation of wound healing |
| BP | GO:0050678 | Regulation of epithelial cell proliferation |
| BP | GO:0043114 | Regulation of vascular permeability |
| BP | GO:0043116 | Negative regulation of vascular permeability |
| BP | GO:0072073 | Kidney epithelium development |
| BP | GO:0007160 | Cell-matrix adhesion |
| BP | GO:0010811 | Positive regulation of cell-substrate adhesion |
| BP | GO:0051145 | Smooth muscle cell differentiation |
| BP | GO:0050920 | Regulation of chemotaxis |
| BP | GO:0050921 | Positive regulation of chemotaxis |
| BP | GO:0003158 | Endothelium development |
| BP | GO:0001952 | Regulation of cell-matrix adhesion |
| BP | GO:0030324 | Lung development |
| BP | GO:0030323 | Respiratory tube development |
| BP | GO:0060840 | Artery development |
| BP | GO:0008037 | Cell recognition |
| BP | GO:0048754 | Branching morphogenesis of an epithelial tube |
| BP | GO:0061138 | Morphogenesis of a branching epithelium |
| BP | GO:0060541 | Respiratory system development |
| BP | GO:0001763 | Morphogenesis of a branching structure |
| BP | GO:0031032 | Actomyosin structure organization |
| BP | GO:0045446 | Endothelial cell differentiation |
| BP | GO:0001823 | Mesonephros development |
| CC | GO:0062023 | Collagen-containing extracellular matrix |
| MF | GO:0005539 | Glycosaminoglycan binding |
| MF | GO:0005178 | Integrin binding |
In KEGG pathway analysis, DEGs from all datasets were consistently and significantly enriched in the ‘Cytoskeleton in muscle cells’ pathway (Fig. 4a-h). Other commonly enriched pathways included Complement and coagulation cascades, Cell adhesion molecules, ECM-receptor interaction, and Leukocyte transendothelial migration.
Fig. 4.
Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis of the training and validation sets. A–H Bar plots display the significantly enriched KEGG pathways for each dataset. The top 10 most significantly enriched pathways are shown for each dataset (or all pathways if the total number is less than 10)
GSEA further demonstrated significant inhibition of the ‘Cytoskeleton in muscle cells’ pathway in three datasets (GSE140797, GSE30219, GSE31210; Supplementary1). Notably, all LUAD datasets indicated a significant activation of the ‘Pyrimidine Metabolism’ pathway (Fig. 5a-h). Collectively, the GO, KEGG, and GSEA analyses strongly suggest shared pathway mechanisms in LUAD.
Fig. 5.
Gene Set Enrichment Analysis (GSEA) of the training and validation sets. A–H GSEA results demonstrate that the ‘Pyrimidine metabolism’ pathway was significantly activated in lung adenocarcinoma tissues across all datasets
Immune cell infiltration in LUAD: decreased monocytes and resting mast cells
Immune infiltration analysis was performed to compare the immune landscape between LUAD and normal lung tissues. Bar plots in Supplementary File 1 display the relative percentages of 22 immune cell types, while violin plots (Fig. 6a-h) detail the differences. The analysis revealed a significant decrease in the proportions of Monocytes and resting Mast cells in LUAD compared to normal tissues across most datasets. Eosinophils and Neutrophils also showed decreasing trends in multiple datasets. In contrast, Plasma cells, T cells follicular helper, and M1 Macrophages were significantly increased in most datasets, indicating substantial alterations in the immune microenvironment of LUAD.
Fig. 6.
Comparison of immune cell infiltration between normal lung tissues and lung adenocarcinoma tissues. A–H Box plots illustrate the infiltration levels of 22 immune cell types analyzed by the CIESBORT algorithm in different groups. The proportions of monocytes and resting mast cells were significantly decreased in lung adenocarcinoma tissues
Screening of core genes and construction of a diagnostic model for LUAD
To identify robust core genes, we employed a combination of Random Forest (RF) and Naive Bayes algorithms for an integrative analysis of DEGs across all datasets. The RF + NaiveBayes model demonstrated superior performance, achieving a mean AUC of 100% [15] (Fig. 7a). This approach identified four core genes: FGD5, LRRC36, C8B, and MYOC. ROC curve analysis confirmed the discriminative power of these genes, with AUC values exceeding 0.74 across all datasets (Fig. 7b). Furthermore, their mRNA expression levels were significantly downregulated in LUAD tissues compared to normal controls (Fig. 8a-h).
Fig. 7.
Screening and validation of core genes based on machine learning algorithms. A Feature selection was performed using a combination of 113 machine learning algorithms. The model combining Random Forest and Naïve Bayes achieved the highest mean AUC value in both training and validation sets, demonstrating optimal performance. B This model identified four core genes (FGD5, LRRC36, C8B, MYOC). The Receiver Operating Characteristic (ROC) curves for these core genes in the training and validation sets are shown on the right, with all AUC values exceeding 0.74
Fig. 8.
Core gene expression and diagnostic model construction. A–H mRNA expression levels of the core genes FGD5, LRRC36, C8B, and MYOC in the training and validation sets. All four genes were significantly downregulated in lung adenocarcinoma tissues compared to normal tissues. I Schematic diagram of the diagnostic model for lung adenocarcinoma constructed based on the four core genes. J The calibration curve for the nomogram prediction of the diagnostic model, showing good consistency between predicted probability and actual probability. K Decision curve analysis indicates that the model provides significant net clinical benefit across all threshold probabilities
A diagnostic model based on these four genes was constructed and visualized using a nomogram, which indicated that FGD5, LRRC36, and MYOC contributed more substantially to the model than C8B (Fig. 8i). The calibration curve showed close alignment with the ideal line, indicating high predictive accuracy (Fig. 8j). Decision curve analysis (DCA) revealed that the diagnostic model provided a superior net clinical benefit across a wide range of probability thresholds compared to the ‘treat all’ or ‘treat none’ strategies (Fig. 8k).
Association between core gene expression and patient survival
Differential expression analysis of the TCGA-LUAD dataset (53 normal vs. 481 LUAD) confirmed significant downregulation of all four core genes in LUAD tissues (Fig. 9a). Kaplan-Meier survival analysis further revealed that low expression levels of these core genes were significantly associated with poorer overall survival in LUAD patients (all log-rank P < 0.05; Fig. 9b-e).
Fig. 9.
Validation using the TCGA database and experimental verification. A Volcano plot of DEGs from the TCGA-LUAD dataset. Red and green dots represent genes significantly up-regulated and down-regulated in lung adenocarcinoma, respectively. The four core genes (FGD5, LRRC36, C8B, MYOC) are annotated and were all down-regulated. B–E Kaplan–Meier survival analysis curves for the four core genes. Lower gene expression levels were significantly associated with poorer overall survival in patients. F Validation of core gene mRNA expression levels in lung adenocarcinoma cell lines via RT-PCR. All four core genes were significantly downregulated in cancer cell lines compared to normal lung epithelial cells
To ascertain whether the four core genes hold independent prognostic value beyond established clinical factors, we performed univariate and multivariate Cox proportional hazards regression analyses. The multivariate model was adjusted for key clinical variables, including age, gender, tumor stage, molecular subtype, and smoking history. The results are summarized in Table 4.
Table 4.
Univariate and multivariate Cox regression analyses of clinical variables and core genes for overall survival in the TCGA-LUAD cohort
| Variable | Univariate Cox regression | Multivariate Cox regression | ||||
|---|---|---|---|---|---|---|
| HR | 95% CI | P value | HR | 95% CI | P value | |
| Age | 1.014 | (0.998–1.030) | 0.080 | 1.018 | (1.002–1.034) | 0.084 |
| Gender | 1.148 | (0.849–1.553) | 0.369 | 0.994 | (0.725–1.362) | 0.969 |
| Stage | 1.633 | (1.417–1.881) | < 0.001 | 1.490 | (1.285–1.728) | < 0.001 |
| Molecular subtypes | 0.802 | (0.536–1.202) | 0.285 | 0.835 | (0.552–1.264) | 0.394 |
| Smoking | 1.135 | (0.615–1.420) | 0.751 | 1.187 | (0.582–1.444) | 0.707 |
| FGD5 | 0.833 | (0.699–0.993) | 0.042 | 0.850 | (0.675–0.995) | 0.045 |
| LRRC36 | 0.824 | (0.704–0.964) | 0.015 | 0.853 | (0.687–0.985) | 0.034 |
| C8B | 0.863 | (0.752–0.989) | 0.034 | 0.872 | (0.730–0.994) | 0.042 |
| MYOC | 0.732 | (0.411–0.968) | 0.035 | 0.771 | (0.366–1.012) | 0.049 |
In the univariate analysis, high expression of all four core genes (FGD5, LRRC36, C8B, and MYOC) was significantly associated with favorable overall survival (all P < 0.05). Among the clinical variables assessed, only advanced tumor stage emerged as a significant risk factor (P < 0.001), whereas age, gender, molecular subtype, and smoking history did not show significant associations.
Most importantly, in the subsequent multivariate Cox regression model adjusted for all the above covariates, the expression levels of FGD5, LRRC36, C8B, and MYOC each retained their status as significant independent predictors of better survival (all P < 0.05). Concurrently, tumor stage remained a strong independent prognostic factor. This confirms that the prognostic association of these four biomarkers is not confounded by the distribution of key clinical and molecular variables in the cohort.
Table 4. The multivariate Cox model was adjusted for age, gender, tumor stage, molecular subtype, and smoking history. The results demonstrate that high expression of the four core genes (FGD5, LRRC36, C8B, MYOC) are independent favorable prognostic factors.
Experimental validation of core gene expression
qRT-PCR validation was performed using the Beas-2B normal bronchial epithelial cell line and the A549 LUAD cell line, with five biological replicates. The results confirmed significant downregulation of FGD5, LRRC36, C8B, and MYOC at the mRNA level in A549 cells compared to Beas-2B controls (Fig. 9f), consistent with the trends observed in the public database analyses.
Discussion
By integrating eight microarray datasets from the GEO database and 534 samples from TCGA, we conducted a systematic analysis of the transcriptomic characteristics of lung adenocarcinoma (LUAD). Functional enrichment analysis revealed significant involvement of various biological processes, including extracellular matrix organization, cell-matrix adhesion, angiogenesis, and regulation of vascular permeability. Concurrently, cytoskeleton-related signaling pathways were suppressed, while the pyrimidine metabolism pathway was markedly activated. Tumor immune infiltration analysis further identified a distinctive immunological feature in the LUAD microenvironment: a significant reduction in the proportions of monocytes and resting mast cells. Through an extensive screening process employing 113 machine learning algorithms, we identified four novel core genes—FGD5, LRRC36, C8B, and MYOC—which were consistently downregulated in LUAD, a finding validated by qRT-PCR experiments. A diagnostic model constructed from these genes demonstrated robust discriminatory power in distinguishing LUAD from normal lung tissue, indicating promising diagnostic potential.
Compared to previous studies, we are the first to report the significant enrichment of the ‘Cytoskeleton in muscle cells’ pathway among DEGs in LUAD. A growing body of evidence in cancer biology links the ‘Cytoskeleton in muscle cells’ pathway to processes critical for tumor progression, including the organization of the extracellular matrix, cell-matrix adhesion, and the formation of new blood vessels [16].Prior research often focused on pathways such as Malaria, Complement and coagulation cascades, and Leukocyte transendothelial migration, paying less attention to the role of the cytoskeleton in tumor development [17]. This discrepancy may stem from previous studies not performing comprehensive enrichment analyses directly on DEGs, instead linking LUAD to mechanisms like pyroptosis or mitophagy, thereby limiting the scope of differential genes examined [18, 19]. Our GSEA results revealed significant activation of the ‘Pyrimidine metabolism’ pathway, aligning with literature describing the importance of pyrimidine metabolism in LUAD [20]. This finding suggests a key role for pyrimidine metabolism in tumor metabolic reprogramming, where its enhancement may not only promote tumor cell proliferation and survival but also influence responses to chemotherapy and immunotherapy. Thus, our study supports the consideration of pyrimidine metabolism-related signatures (PMRS) as potential prognostic and therapeutic biomarkers for LUAD, hinting at applications in personalized medicine.
Studies on the immune cell landscape in LUAD development have revealed that tumor progression is accompanied by an increase in B cells and CD4 + T cells, alongside a gradual decrease in NK cells and monocytes [21]; compared to normal tissues, the infiltration of resting mast cells is significantly reduced in NSCLC tissues, and their infiltration level is negatively correlated with patient risk—specifically, low-risk groups exhibit higher infiltration while high-risk groups show lower infiltration, with low infiltration indicating poor prognosis, and this marked reduction in infiltration of both monocytes and resting mast cells in lung adenocarcinoma is consistent with the bioinformatic analysis of this study [22]; notably, although our study found a significant increase in M1 macrophages in LUAD, the general literature often indicates that M1 macrophages are typically decreased in non-small cell lung cancer while M2 macrophages are increased [23]. This discrepancy may reflect the substantial heterogeneity in tumor-associated immune responses within LUAD, as recent studies have also reported context-dependent roles for immune cells in the tumor microenvironment [24].M1 macrophages are generally considered anti-tumor due to their strong phagocytic capacity and pro-inflammatory responses, whereas M2 macrophages are involved in immunosuppression, tumor angiogenesis, and stromal remodeling, promoting tumor development and metastasis [25]; furthermore, previous studies have found that peripheral blood neutrophils in lung adenocarcinoma patients can transmigrate across blood vessel walls into tumor tissue to form tumor-associated neutrophils (TANs) [26]; neutrophils play a significant role in promoting the progression of lung adenocarcinoma, and tumor cells enhance neutrophil infiltration by regulating the expression of TPM2 and ELANE, thereby accelerating tumor deterioration [27]; however, this contradicts the findings of our immune infiltration analysis; moreover, the infiltration status of tumor-associated neutrophils (TANs) in lung adenocarcinoma lacks a unified conclusion in the literature, and therefore, the overall trend of TANs in lung adenocarcinoma requires further research and discussion.
Studies have found that in the hypoxic microenvironment of LUAD, tumor cell-secreted exosomes are enriched with the lncRNA FGD5-AS1 [28]. Exosomal FGD5-AS1 can function as a competitive endogenous RNA (ceRNA) by directly binding and ‘sponging’ the tumor suppressor miR-1179, relieving its post-transcriptional repression of the downstream target gene CDH3. The interaction within the FGD5-AS1/miR-1179 axis ultimately leads to the upregulation of the oncoprotein P-cadherin (CDH3), significantly driving the proliferation, migration, and invasion of LUAD cells in vitro and in vivo [29]. FGD5-AS1 is highly expressed in NSCLC and can also promote the proliferation and metastasis of pancreatic cancer [30]. The core gene FGD5 identified in our study is itself an important protein-coding gene. Its encoded protein acts as a guanine nucleotide exchange factor, primarily activating CDC42, a key signaling molecule regulating the cytoskeleton, thereby participating in angiogenesis [31]. FGD5-AS1 is a long non-coding RNA transcribed from the antisense strand of the FGD5 locus. This “gene-antisense RNA” pairing suggests that FGD5-AS1 might directly influence FGD5 transcriptional activity through mechanisms like transcriptional interference or epigenetic modifications, thereby implicating FGD5 in its constructed oncogenic network.
Although the function of the complement component C8B in LUAD remains unclear, existing studies suggest its role is cancer-type specific. In hepatocellular carcinoma, C8B was identified as a key hub gene in HBV-related HCC, with its low expression significantly associated with poor patient prognosis, confirming its tumor-suppressive potential [32, 33]. However, in cholangiocarcinoma, C8B is significantly upregulated in specific subtypes [34]. As a key component of the complement membrane attack complex [35], C8B might exert an anti-tumor effect through mediating cell lysis, a notion supported by recent evidence highlighting its investigational potential in cancer immune surveillance and escape [36].We speculate that in LUAD, C8B might similarly play a role in immune surveillance, where its loss of expression could weaken the complement system’s ability to clear tumor cells. This paradoxical phenomenon makes C8B a new entry point for exploring immune escape mechanisms in LUAD.
A study on ALK-positive NSCLC found that pre-treatment detection of MYOC levels in patient plasma could help predict prognosis after ALK-TKI therapy, expanding the clinical application prospects of MYOC and corroborating its important role across different cancer types [37]. Substantial research shows that MYOC is consistently downregulated in various malignancies of different anatomical origins and molecular subtypes, suggesting it might be a pan-cancer tumor suppressor gene: it is listed as one of the most significantly downregulated genes in thymoma; it acts as a hub gene participating in metabolic inhibition via the AMPK/PPAR pathway in Luminal A breast cancer [38]; it is identified as a key secreted protein consistently downregulated with tumor progression in prostate cancer and regulated by specific miRNAs [39]; and in sarcoma, analysis based on TCGA data lists it as one of 11 hub survival genes, with its high expression significantly associated with better overall survival rates [40]. This evidence collectively outlines the biological portrait of MYOC as a broad-spectrum tumor suppressor gene. Its loss of expression in LUAD might disrupt protective functions in the tumor microenvironment, thereby influencing disease progression and treatment response, providing an important direction for subsequent mechanistic research and clinical translation.
Our study also identified a completely novel core gene, LRRC36, which has not been implicated in any previous research. This pure novelty positions LRRC36 as a ‘black box’ full of possibilities, and elucidating its function holds the potential for groundbreaking insights. Collectively, the significant association between low expression of our identified core genes (FGD5, LRRC36, C8B, and MYOC) and poorer overall survival provides clinically actionable prognostic implications. Notably, this association remained statistically significant even after rigorous adjustment for key clinical covariates including tumor stage and molecular subtypes in a multivariate Cox model, underscoring their independent prognostic value. These findings encourage the exploration of these genes as potential predictive biomarkers for risk stratification and outcome prediction in phenotypically diverse LUAD populations. This approach aligns with the growing emphasis on utilizing molecular signatures to guide personalized management strategies in lung adenocarcinoma [41].
In summary, this study identifies four novel core genes closely associated with LUAD and reveals the critical role of a previously underappreciated signaling pathway, ‘Cytoskeleton in muscle cells’. The integration of a multi-algorithm machine learning framework with subsequent validation of independent prognostic utility strengthens the reliability of our findings. However, a primary limitation of this study is that these findings remain at the stage of screening and association analysis. Their specific biological functions and molecular mechanisms urgently require elucidation through functional experiments. We therefore call for future research to focus on in-depth mechanistic exploration of these core genes. We anticipate that breakthroughs in this area will lay a solid foundation for elucidating the regulatory principles of LUAD occurrence and development at the genetic level.
Conclusion
This study systematically identified four novel core genes (FGD5, LRRC36, C8B, and MYOC) and the “Cytoskeleton in muscle cells” pathway in lung adenocarcinoma, highlighting their potential diagnostic and independent prognostic significance. While the current findings are derived from bioinformatic analyses and preliminary validation, we emphasize the necessity for future research to elucidate their precise biological roles through functional experiments, thereby advancing the development of precision diagnosis and treatment for LUAD.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
The authors acknowledge the support of the public resources.
Abbreviations
- AUC
Area under the curve
- BP
Biological process
- CC
Cell composition
- DEGs
Differentially expressed genes
- GEO
Gene Expression Omnibus
- GSEA
Gene set enrichment analysis
- GO
Gene ontology
- KEGG
Kyoto Encyclopedia of Genes and Genomes
- LUAD
Lung adenocarcinoma
- MF
Molecular function
- ROC
Receiver operating characteristic
Author contributions
Xumin Wu: Data curation; Methodology; Formal analysis; Writing - Original Draft.Xujie Wu: Methodology; Validation, Visualization.Qifan Chen: Conceptualization; Writing - Review & Editing.
Funding
Not applicable.
Data availability
The datasets generated and analyzed during the current study are available in the Gene Expression Omnibus (GEO) repository under accession numbers GSE118370, GSE136043, GSE140797, GSE10072, GSE32863, GSE31210, GSE30219, and GSE7670, as well as in The Cancer Genome Atlas (TCGA) Lung Adenocarcinoma (TCGA-LUAD) project. All data can be accessed through their respective official websites: GEO (https://www.ncbi.nlm.nih.gov/geo/) and the Genomic Data Commons Portal (https://portal.gdc.cancer.gov/).
Declarations
Ethics approval and consent to participate
This study utilized publicly available datasets and did not involve new experiments on human participants or animals performed by the authors. Therefore, ethical approval was not required for this secondary analysis. Given that the data analyzed are aggregated, de-identified, and publicly available for research purposes, and that all source studies obtained necessary ethical approvals and participant consent, this study is exempt from requiring additional consent in accordance with relevant national/institutional ethical guidelines..
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Xumin Wu and Xujie Wu are co-first authors.
References
- 1.Zhang Y, Vaccarella S, Morgan E, et al. Global variations in lung cancer incidence by histological subtype in 2020: a population-based study. Lancet Oncol. 2023;24(11):1206–18. [DOI] [PubMed] [Google Scholar]
- 2.Kuang Z, Wang J, Liu K, et al. Global, regional, and National burden of tracheal, bronchus, and lung cancer and its risk factors from 1990 to 2021: findings from the global burden of disease study 2021. EClinicalMedicine. 2024;75:102804. [DOI] [PMC free article] [PubMed]
- 3.Diao X, Guo C, Jin Y, et al. Cancer situation in china: an analysis based on the global epidemiological data released in 2024. Cancer Commun (Lond). 2025;45(2):178–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Luo G, Zhang Y, Rumgay H, et al. Estimated worldwide variation and trends in incidence of lung cancer by histological subtype in 2022 and over time: a population-based study. Lancet Respir Med. 2025;13(4):348–63. [DOI] [PubMed] [Google Scholar]
- 5.Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–63. [DOI] [PubMed] [Google Scholar]
- 6.Shinozaki T, Togasaki K, Hamamoto J, et al. Basal-shift transformation leads to EGFR therapy-resistance in human lung adenocarcinoma. Nat Commun. 2025;16(1):4369. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Kim H, Lee W, Kim Y, et al. Proteogenomic characterization identifies clinical subgroups in EGFR and ALK wild-type never-smoker lung adenocarcinoma. Exp Mol Med. 2024;56(9):2082–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Skoulidis F, Goldberg ME, Greenawalt DM, et al. STK11/LKB1 mutations and PD-1 inhibitor resistance in KRAS-Mutant lung adenocarcinoma. Cancer Discov. 2018;8(7):822–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Nieto P, Ambrogio C, Esteban-Burgos L, et al. A Braf kinase-inactive mutant induces lung adenocarcinoma. Nature. 2017;548(7666):239–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xu JY, Zhang C, Wang X, et al. Integrative proteomic characterization of human lung adenocarcinoma. Cell. 2020;182(1):245–61. [DOI] [PubMed] [Google Scholar]
- 11.Hofman V, Rouquette I, Long-Mira E, et al. Multicenter evaluation of a novel ROS1 immunohistochemistry assay (SP384) for detection of ROS1 rearrangements in a large cohort of lung adenocarcinoma patients. J Thorac Oncol. 2019;14(7):1204–12. [DOI] [PubMed] [Google Scholar]
- 12.Li J, Li X, Guo H, et al. Multi-gene panel sequencing reveals the relationship between driver gene mutation and clinical characteristics in lung adenocarcinoma. Discov Oncol. 2025;16(1):274. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Ma C, Zhang L. Comparison of small biopsy and cytology specimens: subtyping of pulmonary adenocarcinoma. Cytojournal. 2023;20:5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ramalingam PS, Priyadharshini A, Emerson IA, et al. Potential biomarkers uncovered by bioinformatics analysis in Sotorasib resistant-pancreatic ductal adenocarcinoma. Front Med (Lausanne). 2023;10:1107128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Li XY, Xiang J, Wu FX, et al. NetAUC: A network-based multi-biomarker identification method by AUC optimization. Methods. 2022;198:56–64. [DOI] [PubMed] [Google Scholar]
- 16.Hussain MS, Sharma S, Kumari A, et al. Role of long non-coding RNAs in neurofibromatosis and schwannomatosis: pathogenesis and therapeutic potential. Epigenomics. 2024;16(23–24):1453–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Wang S, Wang Q, Fan B, et al. Machine learning-based screening of the diagnostic genes and their relationship with immune-cell infiltration in patients with lung adenocarcinoma. J Thorac Dis. 2022;14(3):699–711. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Liu LP, Lu L, Zhao QQ, et al. Identification and validation of the Pyroptosis-Related molecular subtypes of lung adenocarcinoma by bioinformatics and machine learning. Front Cell Dev Biol. 2021;9:756340. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Wang B, Liu D, Shi D, et al. The role and machine learning analysis of mitochondrial autophagy-related gene expression in lung adenocarcinoma. Front Immunol. 2025;16:1509315. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Hu T, Shi R, Xu Y, et al. Multi-omics and single-cell analysis reveals machine learning-based pyrimidine metabolism-related signature in the prognosis of patients with lung adenocarcinoma. Int J Med Sci. 2025;22(6):1375–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Liu W, You W, Lan Z, et al. An immune cell map of human lung adenocarcinoma development reveals an anti-tumoral role of the Tfh-dependent tertiary lymphoid structure. Cell Rep Med. 2024;5(3):101448. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Yang Y, Qian W, Zhou J, et al. A mast cell-related prognostic model for non-small cell lung cancer. J Thorac Dis. 2023;15(4):1948–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Kang G, Song H, Bo L, et al. Nicotine promotes M2 macrophage polarization through alpha5-nAChR/SOX2/CSF-1 axis in lung adenocarcinoma. Cancer Immunol Immunother. 2024;74(1):11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Luo C, Yu Y, Zhu J, et al. Deubiquitinase PSMD7 facilitates pancreatic cancer progression through activating Nocth1 pathway via modifying SOX2 degradation. Cell Biosci. 2024;14(1):35. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Liu J, Geng X, Hou J, et al. New insights into M1/M2 macrophages: key modulators in cancer progression. Cancer Cell Int. 2021;21(1):389. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Qian L, Ji Z, Mei L, et al. IGF2BP2 promotes lung adenocarcinoma progression by regulating LOX1 and tumor-associated neutrophils. Immunol Res. 2024;73(1):16. [DOI] [PubMed] [Google Scholar]
- 27.Huang C, Qiu H, Xu C, et al. Downregulation of Tropomyosin 2 promotes the progression of lung adenocarcinoma by regulating neutrophil infiltration through neutrophil elastase. Cell Death Dis. 2025;16(1):264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Yang B, Tian Z, Luo Z, et al. Exosomal FGD5-AS1 promotes proliferation of lung cancer cells under hypoxia by inhibiting miR-1179 and activating P-cadherin. Hum Cell. 2025;38(5):145. [DOI] [PubMed] [Google Scholar]
- 29.He Z, Wang J, Zhu C, et al. Exosome-derived FGD5-AS1 promotes tumor-associated macrophage M2 polarization-mediated pancreatic cancer cell proliferation and metastasis. Cancer Lett. 2022;548:215751. [DOI] [PubMed] [Google Scholar]
- 30.Qin S, Liu Y, Zhang X, et al. LncRNA FGD5-AS1 is required for gastric cancer proliferation by inhibiting cell senescence and ROS production via stabilizing YBX1. J Exp Clin Cancer Res. 2024;43(1):188. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Azad AK, Farhan MA, Murray CR, et al. FGD5 regulates endothelial cell PI3 kinase-beta to promote neo-angiogenesis. FASEB J. 2022;36(1):e22080. [DOI] [PubMed] [Google Scholar]
- 32.Zhang Y, Chen X, Cao Y, et al. C8B in complement and coagulation cascades signaling pathway is a predictor for survival in HBV-Related hepatocellular carcinoma patients. Cancer Manag Res. 2021;13:3503–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Li Z, Xu J, Cui H, et al. Bioinformatics analysis of key biomarkers and potential molecular mechanisms in hepatocellular carcinoma induced by hepatitis B virus. Med (Baltim). 2020;99(20):e20302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Bernatz S, Schulze F, Bein J, et al. Small duct and large duct type intrahepatic cholangiocarcinoma reveal distinct patterns of immune signatures. J Cancer Res Clin Oncol. 2024;150(7):357. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Kaufmann T, Rittner C, Schneider PM. The human complement component C8B gene: structure and phylogenetic relationship. Hum Genet. 1993;92(1):69–75. [DOI] [PubMed] [Google Scholar]
- 36.Wu A, Yang H, Xiao T, et al. COPZ1 regulates ferroptosis through NCOA4-mediated ferritinophagy in lung adenocarcinoma. Biochim Biophys Acta Gen Subj. 2024;1868(11):130706. [DOI] [PubMed] [Google Scholar]
- 37.Yu L, Ke J, Du X, et al. Genetic characterization of thymoma. Sci Rep. 2019;9(1):2369. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Jia X, Lei H, Jiang X et al. Identification of crucial LncRNAs for luminal A breast cancer through RNA sequencing. Int J Endocrinol. 2022;2022:6577942. [DOI] [PMC free article] [PubMed]
- 39.Santos NJ, Camargo A, Carvalho HF et al. Prostate cancer secretome and membrane proteome from Pten conditional knockout mice identify potential biomarkers for disease progression. Int J Mol Sci. 2022;23(16):9224. [DOI] [PMC free article] [PubMed]
- 40.Dai D, Xie L, Shui Y, et al. Identification of tumor Microenvironment-Related prognostic genes in sarcoma. Front Genet. 2021;12:620705. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Liu X, Ren Y, Fei Z, et al. Relationship between qualitative and quantitative parameters of three-dimensional computed tomography, EGFR gene mutation, and ALK gene rearrangement in GGO-associated lung adenocarcinoma and their prognostic value. Pathol Oncol Res. 2025;31:1612081. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets generated and analyzed during the current study are available in the Gene Expression Omnibus (GEO) repository under accession numbers GSE118370, GSE136043, GSE140797, GSE10072, GSE32863, GSE31210, GSE30219, and GSE7670, as well as in The Cancer Genome Atlas (TCGA) Lung Adenocarcinoma (TCGA-LUAD) project. All data can be accessed through their respective official websites: GEO (https://www.ncbi.nlm.nih.gov/geo/) and the Genomic Data Commons Portal (https://portal.gdc.cancer.gov/).









