Abstract
Lung adenocarcinoma, the predominant pathogenic type of lung cancer, exhibits intricate biological behaviors and molecular pathways, which have consistently been a focal point and challenge in tumor research. The integration of multi-omics technologies and diverse analytical methodologies facilitates an in-depth study of lung adenocarcinoma’s characteristics at cellular, molecular, and clinical dimensions, providing a theoretical foundation and prospective targets for precise diagnosis and effective treatment. scRNA-seq and RNA-seq data were obtained from public databases. The epithelial cells of lung adenocarcinoma were analyzed by methods such as CNV level assessment, cell trajectory analysis, malignant cell identification, and cell communication analysis, and the marker genes of malignant cells were extracted. In the RNA-seq data, CoxBoost was used for prognostic modeling to mine the risk gene BAIAP2L2 for further analysis. The effect of BAIAP2L2 on lung adenocarcinoma was verified by cell experiments. There is obvious heterogeneity among the epithelial cells of lung adenocarcinoma, and significant differences exist between different cell subgroups. The malignant differentiation of lung adenocarcinoma cells has a clear trajectory, and epithelial cell subgroups may influence the differentiation direction through cell communication. The CoxBoost model established based on the marker genes of malignant cells has a relatively good predictive effect, and BAIAP2L2 can serve as a prognostic factor for lung adenocarcinoma. In vitro cellular experiments demonstrated that interfering with BAIAP2L2 can impede the proliferation, migration, and invasion of lung adenocarcinoma cells, along with the epithelial-mesenchymal transition of these cells. BAIAP2L2 is a prognostic factor for lung adenocarcinoma, and interfering with BAIAP2L2 can inhibit the growth, metastasis, and epithelial-mesenchymal transition of lung adenocarcinoma.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-025-33171-8.
Keywords: Lung adenocarcinoma, EMT, Multi-omics, Prognostic factor
Subject terms: Cancer, Molecular medicine, Genome informatics
Introduction
Lung adenocarcinoma (LUAD) represents the predominant subtype of non-small cell lung cancer (NSCLC) and constitutes a substantial fraction of cancer-related fatalities globally1,2.Despite continuous advancements in diagnostic techniques and treatment modalities, the prognosis of many lung adenocarcinoma patients remains poor, especially for those in the advanced stage3. A comprehensive understanding of the underlying biological mechanisms driving the progression of lung adenocarcinoma and the identification of reliable prognostic biomarkers are crucial for improving patient outcomes.
Lung adenocarcinoma exhibits significant heterogeneity at both molecular and cellular levels4. Tumor cells within a single lung adenocarcinoma can display significant genetic, epigenetic, and transcriptomic variations, resulting in diverse tumor phenotypes and treatment responses5. Conventional histopathological classification techniques, while essential for preliminary diagnosis, fail to comprehensively represent the intricacies of lung adenocarcinoma biology6.
In recent years, data from multiple omics levels, such as RNA transcriptomics and single-cell transcriptomics, have become powerful tools for dissecting the molecular landscape of lung adenocarcinoma7,8. These methods enable us to view the disease from a more comprehensive perspective, further revealing the complex biological characteristics of lung adenocarcinoma and providing additional information for disease research and treatment9.
Precisely identifying reliable prognostic biomarkers in lung adenocarcinoma is another key aspect of improving patient prognosis10. Prognostic markers can forecast disease progression, evaluate recurrence probability, and inform therapy choices. Conventional prognostic factors for lung adenocarcinoma include tumor stage, histological grade, and some clinical parameters11. Nonetheless, these factors have certain limitations in accurately predicting the prognosis of individual patients. The discovery of novel prognostic biomarkers with strong predictive capabilities and potential clinical applications will greatly enhance our ability to manage patients with lung cancer.
BAIAP2L2 (BAR/IMD Domain Containing Adaptor Protein 2 Like 2) has recently emerged as a promising candidate among the possible biomarkers for lung adenocarcinoma12. Proteins containing the Bin/amphiphysin/Rvs (BAR) domain play a critical role in shaping cellular membranes and are essential for various cellular processes, including endocytosis, cell motility, and morphogenesis13. The metastasis of lung adenocarcinoma is an important factor affecting the poor prognosis of patients14. We hope to clarify the influence of BAIAP2L2 on epithelial-mesenchymal transition through cell experiments and further explore the significance of this prognostic marker for lung adenocarcinoma. Although the hazard ratio (HR) of BAIAP2L2 was not the most prominent among all candidates, it was selected based on its consistent upregulation in LUAD across multiple RNA-seq datasets, its ability to stratify patient prognosis in both training and validation cohorts, and its biological role validated by in vitro experiments. BAIAP2L2 exhibited moderate but clinically meaningful predictive power and showed strong association with epithelial-mesenchymal transition (EMT), making it a biologically relevant and functionally important marker worthy of deeper investigation.
In this study, we aimed to identify different malignant subtypes and discover a novel prognostic biomarker, BAIAP2L2, by comprehensively exploring the multi-omics landscape of lung adenocarcinoma. Meanwhile, we verified its impact on the growth and metastasis of lung adenocarcinoma through cell experiments.
Methods
Data collection
The RNA-seq data used in this study were obtained from the TCGA database and the GEO database, including TCGA-LUAD, GSE3141, GSE8894, GSE13212, GSE19188, GSE26939, GSE29013, GSE29016, GSE30219, GSE31210, GSE37745, GSE50081, and GSE72094. The dataset GSE131907 from the GEO database was used for scRNA-seq analysis.
Data preprocessing
The combat function of the SVA package was used to remove batch effects from all RNA-seq datasets, and the data were grouped according to the source of the dataset. Data from the TCGA were used as the training set, and data from the GEO were used as the validation set. For the training set data, differential analysis was performed using the edgeR package, with the criteria for screening differentially expressed genes being logFC > 1 and FDR < 0.05.
Single-cell analysis
In this study, the criteria for cell selection were that the proportion of mitochondrial gene expression was less than 10%, the proportion of erythrocyte gene expression was less than 3%, and the number of expressed genes was between 200 and 8000. To ensure data standardization and comparability, the data were standardized using the “NFS” strategy, which encompasses three key steps: Normalizedata, Findvariablefeatures, and Scaledata. Meanwhile, the “cc.genes.updated.2019” dataset provided by the “Seurat” package was used to accurately normalize cell cycle genes to eliminate potential interference from different cell cycle stages on data analysis. In addition, the “DoubletFinder” package was used to effectively identify and remove potential double cells in the dataset, thereby ensuring data purity and reliability. The remaining cells were analyzed using the “FindVariableFeatures” function to detect 2000 highly variable genes. Subsequently, principal component analysis (PCA) was performed to reduce the dimensionality of the scRNA-seq data. The Harmony package was used to remove batch effects of samples from different patients within the dataset. The single-cell data were split into tumor cells and normal cells according to their source and clustered separately. Cell clustering was performed using the “FindClusters” function with a resolution parameter set to 0.81. Based on the annotation information provided by the dataset, the cells were annotated, and the cells were divided into various subgroups according to the results of dimensionality reduction and clustering.
The GSE131907 dataset included single-cell RNA-sequencing data from 11 patients diagnosed with lung adenocarcinoma. The samples represented a mixture of early- and late-stage tumors, enabling characterization across disease progression. According to dataset metadata, none of the patients received chemotherapy prior to sample collection, supporting unbiased evaluation of the tumor microenvironment and cell states.
Malignant cell identification
The inferCNV package uses variations in the expression levels of neighboring genes to infer CNV in cells, based on the hypothesis that copy number changes can cause significant changes in local gene expression. InferCNV does not use Bayesian methods but rather compares gene expression values with normal reference cells to identify abnormal expression patterns and infer copy number changes. It requires pre-specifying a set of normal cells as a reference, which are used to construct a baseline expression pattern and then compare with the expression pattern of tumor cells to infer CNV. The epithelial cell subgroup was extracted, and the epithelial cell subgroups derived from tumor cells and normal cells were included in the “inferCNV” package for calculation. To accurately locate malignant cells exhibiting clonal widespread chromosomal copy number variations (CNV), the CNV spectrum of the cells was inferred using the epithelial cell subgroup derived from normal cells as a reference. K-means clustering was used to perform secondary clustering on the inferCNV results, and the final number of clusters was selected according to the Elbow Method. The CNV score of each cluster was calculated and visualized.
The CopyKAT package was used for copy number variation (CNV) analysis to distinguish tumor cells from normal cells and parse subclonal structures. The basic logic of inferring DNA copy number events in RNA-seq data is to provide in-depth information through the expression levels of adjacent genes, thereby inferring the genomic copy number in that region. The copy number spectrum estimated by CopyKAT can achieve a high consistency of up to 80% with the actual DNA copy number obtained from whole-genome DNA sequencing. The principle of distinguishing tumor/normal cell states is that aneuploidy is common in 90% of cancers. Cells with widespread whole-genome copy number abnormalities (aneuploidy) are considered tumor cells, while stromal normal cells and immune cells usually have a 2 N diploid or near-diploid copy number spectrum.
Cell trajectory analysis
We used the Monocle2 package and the Slingshot package to infer the cell developmental trajectories of subgroups. The algorithm of the Monocle2 package is based on a simple and effective principle, namely clustering cells and then constructing a trajectory based on the center point of the average value of the cell population. At the same time, the distance between other cells and the assumed trajectory was calculated, and these cells were assigned to the cell population closest to them. The Slingshot package mainly consists of two steps, including identifying the global lineage structure using a minimum spanning tree (MST) based on clustering and fitting a principal curve to describe each lineage.
Cell communication
Cell communication analysis is a method for studying how different cell types communicate and regulate with each other through signaling molecules such as ligands and receptors. It is of great significance in revealing the mechanisms of cell-cell interactions and understanding how tissues and organs coordinate their functions. We used the CellChat package to perform cell communication analysis on each subgroup of epithelial cells and visualized the results.
Prognostic analysis
In this study, we used the CoxBoost model to calculate the risk scores of patients. The CoxBoost model was constructed using the CoxBoost package in R language. During the construction process, instead of using the gradient-based boosting method to estimate the Cox proportional hazards model, an offset-based boosting method was employed. Specifically, in each round of boosting iteration, the result obtained from the previous round of boosting was incorporated into the offset of the partial likelihood estimation and served as part of the penalty term. In this way, the model could update and adjust the covariates in each iteration.
It is worth noting that compared with the traditional gradient boosting method, the offset-based boosting method exhibits more flexible penalty structure characteristics. In the current implementation, this flexibility is mainly reflected in the treatment of covariates. Specifically, this method can impose an unpenalized constraint on forced covariates, enabling the coefficients of forced covariates to accumulate rapidly during the boosting steps. Meanwhile, appropriate penalties are still applied to other non-forced covariates to ensure the stability and accuracy of the model during the optimization process.
Functional analysis of BAIAP2L2 in lung adenocarcinoma
Based on the RNA-seq data of lung adenocarcinoma and adjacent normal tissues from the TCGA database, we used the limma package to perform differential analysis, thereby determining the differential expression of BAIAP2L2 in lung adenocarcinoma and normal lung tissues. The Kaplan-Meier survival curve method was used to group patients according to the expression level of BAIAP2L2, and the survival curves were drawn. The Log-rank test was used to compare the survival differences between the two groups. Then, the time-dependent ROC curve method was used to evaluate the prediction accuracy of the target gene at different time points, and the AUC value was calculated to quantify its prediction ability. The differences in Age, Gender, TNM stage, and stage between the high-expression group and the low-expression group of BAIAP2L2 were analyzed to determine the independent impact of BAIAP2L2 on TNM stage and stage. Univariate Cox analysis and multivariate Cox analysis were performed on clinical factors to independently analyze the prognosis, and to detect whether BAIAP2L2 could serve as an independent prognostic factor for lung adenocarcinoma, with the prognostic assessment power unaffected by different clinical stages. Immunohistochemical sections of BAIAP2L2 in lung cancer tissues and normal lung tissues were obtained from the HPA database (The Human Protein Atlas, HPA) to determine the expression location and relative expression level of BAIAP2L2. The pRRophetic package in R software was used to calculate the half-maximal inhibitory concentration (IC50) of BAIAP2L2 expression differences for 138 drugs, and the ggpubr package was used to draw drug sensitivity box plots15. The TCGAplot package was used to analyze the performance of BAIAP2L2 in pan-cancer data16.
Cell line selection and transfection
The lung adenocarcinoma cell lines A549 and PC9, and the normal lung cell line Beas-2B were selected. siRNA and liposomes (Lipofectamine 3000, Thermo Fisher Scientific) were used to transfect and knock down the BAIAP2L2 gene in the cell lines. Cells were seeded into six-well plates at a concentration of 2 × 10⁵ cells/well. After adding the liposome suspension, the cells were cultured statically for 24 h to complete the transfection. The siRNA sequences were as follows: sense (5’-3’): GGACAUGCAGUUCAUCAAATT, antisense (5’-3’): UUUGAUGAACUGCAUGUCCTT.
qRT-PCR
Total RNA was extracted from the cell lines using an RNA isolation kit (Axygen, Suzhou, China). Then, the total RNA was reverse transcribed into cDNA. Subsequently, qRT-PCR was performed using the SYBR Green kit, and the reaction conditions were set on a thermal cycler (Bio-Rad, Hercules, CA, USA) as follows: 600 s at 95 °C, followed by 10 s at 95 °C, 60 s at 65 °C, 1 s at 97 °C, and 30 s at 37 °C, for a total of 40 cycles. GAPDH was used as an internal reference, and the expression level of BAIAP2L2 was quantitatively analyzed using the 2 − ΔΔCt method. The PCR primer sequences for BAIAP2L2 were: Forward Primer: GCCGAGGTCTACTTCAGTGC, Reverse Primer: GCTGGGTGTCAGACATCTGC.
CCK8 assay
After the cell transfection, the transfected cells were seeded into a 96-well plate at a density of 3.5 × 10³ cells/well. Subsequently, the seeded cells were cultured and observed regularly. At 24 h, 48 h, 72 h, and 96 h of culture, the cell culture was terminated. After terminating the culture, an appropriate amount of Cell Counting Kit-8 reagent was added to each well, and the 96-well plate was incubated at room temperature for 2 h to ensure sufficient reaction between the reagent and the cells. After incubation, the absorbance of each well in the 96-well plate was measured using a microplate reader, with the measurement wavelength set at 450 nm. Based on the absorbance data obtained at different time points, the cell proliferation curve was fitted using data analysis software or relevant algorithms.
Colony formation assay
Cells were seeded into six-well plates at a density of 5 × 10³ cells/well. An appropriate amount of medium was added to the six-well plates to ensure that the cells could grow in a suitable environment. Then, the six-well plates were placed in an incubator at 37 °C with 5% CO₂ concentration, and the cells were cultured statically under these conditions for 10 days. The medium needed to be replaced every three days, that is, fresh medium was added to the six-well plates to replace the original medium. After 10 days of culture, the cells in the six-well plates were fixed with 4% paraformaldehyde solution. After fixation, the cells were stained with 0.1% crystal violet staining solution. After staining, the crystal violet was removed, and the number of colonies was photographed and recorded.
Cell scratch assay
Cells were routinely cultured in six-well plates. When the cells reached 80% confluence, a clear scratch was made in the cells using a 10 µl pipette tip. The detached cells in the scratch area were gently washed away with PBS, and the cells were cultured in basic medium. The initial image before scratching was observed and photographed under an inverted microscope as the baseline at 0 h. After 24 h, the image of the scratch area was photographed again to observe the cell migration and healing. The degree of cell healing was calculated using ImageJ image analysis software to obtain the cell migration percentage.
Transwell migration and invasion assays
To synchronize the cell growth state during the cell experiment, the cells were subjected to serum-free starvation culture with basic medium for 24 h. After this culture, the cells were digested with trypsin and collected. Subsequently, the original medium was discarded, and the cells were resuspended in basic medium for subsequent experimental operations.
For the migration assay, after installing the Transwell chamber in a 24-well plate, 2 × 10⁴ cells were seeded into the upper chamber of the chamber, and complete medium containing 10% fetal bovine serum was added to the lower chamber of the corresponding chamber to construct the migration assay environment. For the invasion assay, before inoculating the cells, Matrigel was added to the upper chamber of the Transwell chamber in advance. The subsequent operations were the same as those of the migration assay, that is, 2 × 10⁴ cells were seeded into the upper chamber containing Matrigel, and complete medium containing 10% fetal bovine serum was added to the lower chamber. After 24 h of culture, the Transwell chamber was removed from the culture system. To fix the cell morphology, 4% paraformaldehyde was added to the chamber for fixation. After fixation, the cells were stained with 0.1% crystal violet solution. After staining, the residual non-migrated or non-invading cells in the upper chamber of the chamber were removed to reduce the interference of irrelevant cells on the observation results. Finally, the treated Transwell chamber was observed under an inverted microscope and imaged to record the experimental results.
Hoechst 33,342 staining
Cells were seeded into six-well plates at a density of 2 × 10⁵ cells/well and cultured until the logarithmic growth phase, after which the culture was terminated. Dead cells were washed away with PBS, and 1 ml of Hoechst 33,342 dye was added to each well. The cells were incubated at 37 °C for 30 min to evenly stain the cell nuclei. The unbound dye was removed by washing three times, and then the cell nuclei morphology and number were observed and analyzed under a fluorescence microscope, and images were taken.
Immunofluorescence staining
Cells were seeded into 96-well plates at a density of 3.5 × 10³ cells/well and cultured for 24 h. The cells were fixed with 4% paraformaldehyde, and then the cell membrane was permeabilized with 0.1% Triton X-100 to ensure that the antibodies could enter the cells. The cells were blocked with 5% bovine serum albumin (BSA) in TBST solution for 1 h to block non-specific binding sites. The primary antibody was added, and the cells were incubated at 4 °C overnight. After the primary antibody was removed, the cells were incubated with FITC-labeled secondary antibody (Fluorescein (FITC)-conjugated Goat Anti-Rabbit IgG(H + L), 1:500, Proteintech, Cat No: SA00003-2) for 1 h. Unbound antibodies were washed away with TBST, and the cell nuclei were stained with DAPI. Finally, the images were observed and photographed under a fluorescence microscope.
Western blot analysis
Cells in six-well plates were lysed with RIPA lysis buffer containing protease and phosphatase inhibitors to obtain total protein. The protein concentration was measured using the BCA method, and the samples were adjusted to the same concentration. Equal amounts of protein samples were separated by SDS-PAGE gel electrophoresis and transferred to PVDF membranes. After transfer, the membranes were blocked with 5% BSA in TBST buffer at room temperature for 1 h. The primary antibody was added, and the membranes were incubated at 4 °C overnight. After the primary antibody was removed, the membranes were washed three times with TBST buffer for 5 min each time, and then the HRP-labeled secondary antibody was added, and the membranes were incubated at room temperature for 1 h. The membranes were developed with ECL, and the bands were recorded and the expression levels of the target proteins were analyzed using an image analysis system. The antibodies used in this part are as follows: BAIAP2L2 (BAIAP2L2 Polyclonal antibody, 1:1000, Proteintech, Cat No: 16738-1-AP), N-cadherin (N-cadherin Polyclonal antibody, 1:5000, Proteintech, Cat No: 22018-1-AP), E-cadherin (E-cadherin Polyclonal antibody, 1:5000, Proteintech, Cat No: 20874-1-AP), Vimentin (Vimentin Polyclonal antibody, 1:20000, Proteintech, Cat No: 10366-1-AP), GAPDH (GAPDH Monoclonal antibody, 1:20000, Proteintech, Cat No: 60004-1-Ig), Rabbit secondary antibody (HRP-conjugated Goat Anti-Rabbit IgG(H + L), 1:7500, Proteintech, Cat No: SA00001-2), Mouse secondary antibody (HRP-conjugated Goat Anti-Mouse IgG(H + L), 1:7500, Proteintech, Cat No: SA00001-1).
Uncropped, original blots with membrane edges visible are provided in Supplementary Figure S1 for all antibodies (E-cadherin, N-cadherin, Vimentin, BAIAP2L2, and GAPDH).
Statistical analysis
Bioinformatics analysis and related image generation were performed using R (version 4.1.0) software. Statistical analysis and image generation of experimental data were completed using GraphPad Prism (version 9.0) and ImageJ software. For comparisons between two groups of samples, the two-tailed Student’s t-test was used, and for comparisons among multiple groups of samples, one-way analysis of variance (ANOVA) was used. All results are presented as mean ± standard deviation (SD), and a p-value < 0.05 was considered statistically significant. The significance markers are as follows: *P < 0.05, **P < 0.01, ***P < 0.001, ****P < 0.0001.
Results
Identification of malignant subpopulations in lung adenocarcinoma
Cluster 6 was enriched in tumor samples and exhibited high aneuploidy, suggesting strong malignant potential. However, site-specific annotations such as brain metastasis were unavailable in the datasets used (TCGA/GEO), and N/M staging metadata were incomplete across samples. Therefore, brain metastasis–specific survival analysis could not be conducted. Instead, T stage analysis showed that BAIAP2L2 expression was significantly higher in T3 compared to T1–T2, highlighting its association with tumor progression.
After dimensionality reduction, clustering, and annotation of the scRNA-seq data of lung adenocarcinoma, 10 cell types were obtained. Nine cell types, namely “Oligodendrocytes”, “Endothelial cells”, “Fibroblasts cells”, “MAST cells”, “NK cells”, “Myeloid cells”, “T lymphocytes”, “Epithelial cells”, and “B lymphocytes”, were clearly clustered. Due to insufficient resolution in cell clustering and inadequate coverage precision of marker genes, some cells could not be precisely clustered. These cells were named “Undetermined” and excluded from further analysis (Fig. 1A). Copy number variation (CNV) heatmap visualization confirmed the malignant potential of epithelial cell clusters, distinguishing them from non-malignant populations (Fig. 1B).
Fig. 1.
Single cell landscape of lung adenocarcinoma. (A) Single cell UMAP diagram. (B) Cell type annotation. (C) Determine the number of K-means clustering based on the elbow diagram. (D) Epithelial cell subpopulation CNV score. (E) InferCNV dimensionality reduction clustering. (F) Percentage of CNV subpopulation derived cells. (G) Marker genes of CNV subgroups.
InferCNV analysis was performed on the epithelial cell subpopulations, and the results were subjected to K-means clustering. The number of clusters was determined according to the elbow method, resulting in 9 clusters (Fig. 1C). It was found that epithelial cells could be further divided into nine CNV subpopulations based on different CNV levels, and the distribution percentages of these subpopulations varied in tumors from different sources (Fig. 1D and F). Marker genes were extracted for each of the nine epithelial subpopulations, and significant differences were observed among these marker genes. This indicates substantial heterogeneity within lung adenocarcinoma epithelial cells, and different differentiation subtypes may lead to variations in the prognosis of lung adenocarcinoma (Fig. 1G).
Development trajectories and cell communication in lung adenocarcinoma
Cell trajectory analysis using the Monocle2 package revealed the differentiation trajectories of CNV subpopulations of lung adenocarcinoma epithelial cells (Fig. 2A). The pseudotime sequence was redefined to match the left-to-right progression of malignant transformation, with subpopulations 6 and 8, exhibiting higher CNV scores, located at the ends of two cell differentiation trajectories (Fig. 2B). Subpopulation 4, which contained normal cells, was positioned at the origin. This suggests that subpopulations 6 and 8 may represent more malignant differentiation endpoints, and the trajectory from subpopulation 4 to 6 or 8 symbolizes the process of tumor cells transforming into a malignant state.
Fig. 2.
(A) Pseudotime analysis of lung adenocarcinoma epithelial cells. (B) Developmental trajectory of CNV subclusters. (C) UMAP dimensionality reduction distribution of CNV subclusters. (D) Slingshot cell trajectory distribution. (E) Cell-cell communication in CNV subclusters. (F) CopyKAT analysis of cell malignancy.
To further clarify the complex differentiation relationships within the epithelial cell subpopulations, a secondary clustering of the epithelial cell subpopulations was performed. The UMAP dimensionality reduction plot showed relatively clear boundaries among the cell subpopulations, indicating that the K-means clustering based on CNV scores had a good classification effect (Fig. 2C).
To validate the trajectory analysis results of the Monocle2 package, the Slingshot package was used for cell trajectory analysis again. The results showed that subpopulations 6 and 8 were still at different differentiation endpoints, and cells in subpopulation 2 and part of subpopulation 4 were distributed in the middle of the cell trajectory, consistent with the previous analysis (Fig. 2D).
Although single-gene tracking along bifurcating pseudotime paths was not supported by the data format, BAIAP2L2 expression was found to be significantly enriched in subpopulations 6 and 8. This suggests an upregulation trend along the malignant differentiation trajectory, implicating BAIAP2L2 in the progression toward more malignant cell states.
Figure 2F shows the CopyKAT analysis of cell malignancy. Figure 2E shows cell-cell communication in CNV subclusters. Cell communication analysis of the CNV subpopulations of epithelial cells revealed extensive cell communication among these subpopulations, with significant differences between different subpopulations. This may indicate that tumor cells promote malignant transformation through cell communication. To further identify the most malignant cell subpopulation, the CopyKAT package was used to evaluate the malignancy of the CNV subpopulations of epithelial cells. The most malignant cells were selected and defined as “aneuploid” (Fig. 2F). Notably, most of the cells in the “aneuploid” subpopulation were enriched in CNV subpopulation 6, with a small portion enriched in other subpopulations. Marker genes of the “aneuploid” subpopulation were obtained for further analysis.
High expression of BAIAP2L2 indicates poor prognosis in lung adenocarcinoma
The marker genes of the “aneuploid” subpopulation were intersected with the differentially expressed genes calculated from the TCGA-LUAD dataset, resulting in 402 genes (Fig. 3A). Through univariate Cox regression analysis and LASSO penalized regression analysis, 13 genes were selected as prognostic risk genes, with the selection criteria being HR > 1 and P < 0.05 (Fig. 3B and C). The TCGA-LUAD dataset was used as the training set, while data from GEO datasets were used as the validation set. Across these independent datasets, patients with high BAIAP2L2 expression consistently showed poorer overall survival compared to low-expression groups. This cross-validation highlights the generalizability and reliability of BAIAP2L2 as a prognostic marker. The CoxBoost model was subsequently employed to score patients in both the training (TCGA) and validation (GEO) sets, demonstrating its efficacy in evaluating patient prognosis. Patients with high-risk scores had significantly worse prognoses than those with low-risk scores (Fig. 3D and E). BAIAP2L2, as a risk gene, had a relatively high HR value, but its functional mechanism in lung adenocarcinoma has not been widely studied. In the TCGA-LUAD cohort, patients were divided into high-expression and low-expression groups based on the median value of BAIAP2L2. Drug sensitivity analysis showed that patients with high BAIAP2L2 expression were more sensitive to gefitinib (Fig. 3F).
Fig. 3.
(A) Intersection genes Upset plot. (B) Univariate Cox regression forest plot. (C) LASSO dimensionality reduction plot. (D) Training set Kaplan-Meier curve. (E) Validation set Kaplan-Meier curve. (F) Gefitinib sensitivity.
Kaplan-Meier survival analysis revealed that patients with high BAIAP2L2 expression had significantly worse prognoses than those with low expression (Fig. 4A). Differential analysis showed that BAIAP2L2 was significantly overexpressed in tumor tissues (Fig. 4B). There were certain differences in BAIAP2L2 expression among patients in different T stages, with higher expression in T3 stage patients compared to T1 and T2 stage patients (Fig. 4C). The time-dependent ROC curve showed that BAIAP2L2 had a certain predictive power for the five-year survival rate of patients (Fig. 4D). To avoid the interference of clinical factors on the prognostic power of BAIAP2L2, independent prognostic analyses were performed. It was found that BAIAP2L2 had prognostic value in both the univariate Cox model and the multivariate Cox model, independent of other clinical factors (Fig. 4E and F).
Fig. 4.
(A) BAIAP2L2 single-gene Kaplan-Meier curve. (B) BAIAP2L2 differential analysis. (C) Correlation between BAIAP2L2 and T stage. (D) Time-dependent ROC curve. (E) Univariate COX regression for independent prognostic analysis. (F) Multivariate COX regression for independent prognostic analysis.
Expression and function of BAIAP2L2 in pan-cancer
To further explore the function of BAIAP2L2 in tumors, a pan-cancer analysis was conducted. Univariate Cox analysis showed that BAIAP2L2 had prognostic value in several cancers, including Adrenocortical carcinoma (ACC), Kidney Chromophobe (KICH), Acute Myeloid Leukemia (LAML), Lung adenocarcinoma (LUAD), Mesothelioma (MESO), Prostate adenocarcinoma (PRAD), Skin Cutaneous Melanoma (SKCM), and Uveal Melanoma (UVM) (Fig. 5A). Additionally, BAIAP2L2 showed significant differential expression in multiple tumors (Fig. 5B and C). Figure 5D displays the pan-cancer differential expression of BAIAP2L2. It does not represent immune cell infiltration data.
Fig. 5.
(A) Pan-cancer univariate COX analysis of BAIAP2L2. (B) Pan-cancer single-cell correlation of BAIAP2L2. (C) Pan-cancer immune checkpoint correlation of BAIAP2L2. (D) Pan-cancer differential analysis of BAIAP2L2.
Correlations between BAIAP2L2 and immune infiltration levels, immune checkpoint genes, and immune modulators are instead shown in Fig. 6A and D. These analyses revealed significant heterogeneity in the immune associations of BAIAP2L2 across tumor types, supporting its potential role in the tumor immune microenvironment. Although our analyses demonstrate these associations, mechanistic explanations of how BAIAP2L2 regulates immune modulation remain to be elucidated. Future studies may employ single-cell spatial transcriptomics or co-culture models to dissect these interactions.
Fig. 6.
(A) Pan-cancer immune stimulator correlation of BAIAP2L2. (B) Pan-cancer immune suppressor correlation of BAIAP2L2. (C) Pan-cancer chemokine correlation of BAIAP2L2. (D) Pan-cancer chemokine receptor correlation of BAIAP2L2.
Interference with BAIAP2L2 inhibits the growth and metastasis of lung adenocarcinoma
qRT-PCR analysis showed that the expression of BAIAP2L2 was significantly lower in the normal lung cell line BEAS-2B than in the lung adenocarcinoma cell lines A549 and PC9 (Fig. 7A). The siRNA could effectively interfere with the expression of BAIAP2L2, resulting in significant reductions in both mRNA and protein expression levels in the lung adenocarcinoma cell lines (Fig. 7B and F). In Fig. 7F, BAIAP2L2 protein expression in PC9 cells was significantly reduced after siRNA transfection (n = 3; p = 0.0034, two-tailed Student’s t-test). Immunohistochemical slices of BAIAP2L2 in lung cancer tissues and adjacent tissues from the HPA database showed significantly enhanced staining in lung cancer tissues, with staining mainly concentrated in the cytoplasm (Fig. 7G and F). Figure 7H highlights the low expression of BAIAP2L2 in adjacent normal tissues, further supporting its tumor-specific overexpression pattern. The CCK8 assay revealed that the growth rate of lung adenocarcinoma cells was significantly slowed down after interfering with BAIAP2L2 (Fig. 7I and J). Proliferation was measured at 24, 48, 72, and 96 h using triplicate replicates per group (n = 3), and ANOVA with post hoc Tukey’s test indicated statistically significant growth inhibition at 48, 72, and 96 h (p < 0.01).
Fig. 7.
(A) Relative expression level of BAIAP2L2. (B) Knockdown efficiency of BAIAP2L2 in A549 cells. (C) Knockdown efficiency of BAIAP2L2 in PC9 cells. (D) Verification of BAIAP2L2 expression by Western Blot. (Full-length uncropped blots are shown in Supplementary Fig. S1) (E) Statistical plot of BAIAP2L2 protein expression in interfering A549 cells. (F) Statistical plot of BAIAP2L2 protein expression in interfering PC9 cells. (G) Immunohistochemistry of BAIAP2L2 in tumor tissues. (H) Immunohistochemistry of BAIAP2L2 in adjacent normal tissues. (I) CCK8 assay of BAIAP2L2 in A549 cells. (J) CCK8 assay of BAIAP2L2 in PC9 cells.
The colony formation assay results indicated that the number of colonies formed was significantly decreased compared to the si-NC group, suggesting that the single-cell tumorigenic ability of lung adenocarcinoma cells was significantly suppressed (Fig. 8A and B). Transwell migration and invasion assays showed that the number of cells penetrating the membrane was significantly reduced after interfering with BAIAP2L2, and both migration and invasion abilities were markedly inhibited (Fig. 8C and G). The cell scratch assay showed that interfering with BAIAP2L2 could significantly inhibit the healing of cell scratches, further demonstrating the inhibition of lung adenocarcinoma cell migration (Fig. 8H and I). Hoechst 33,342 staining showed that interference with BAIP2L2 can significantly promote apoptosis of lung adenocarcinoma cells (Fig. 9A and B).
Fig. 8.
(A) Colony formation of A549 cells. (B) Colony formation of PC9 cells. (C) Transwell migration and invasion assay of A549 and PC9 cells. (D) Data analysis of A549 cell migration ability following BAIAP2L2 knockdown. (E) Data analysis of A549 cell invasion ability following BAIAP2L2 knockdown. (F) Data analysis of PC9 cell migration ability following BAIAP2L2 knockdown. (G) Data analysis of PC9 cell invasion ability following BAIAP2L2 knockdown (H) Wound healing assays in A549 cells. (I) Wound healing assays in PC9 cells.
Fig. 9.
(A) Hoechst 33,342 staining of A549 cells. (B) Hoechst 33,342 staining of A549 cells. (C) Western Blot of EMT markers after interfering with BAIAP2L2. (Full-length uncropped blots are shown in Supplementary Fig. S1) (D) Statistical plot of EMT markers in A549 cells after interfering with BAIAP2L2. (E) Statistical plot of EMT markers in PC9 cells after interfering with BAIAP2L2. (F) Immunofluorescence of EMT Markers in A549 Cells after interfering with BAIAP2L2. (G) Immunofluorescence of EMT Markers in PC9 Cells after interfering with BAIAP2L2.
Interference with BAIAP2L2 effectively inhibits epithelial-mesenchymal transition in lung adenocarcinoma
Immunofluorescence and Western blot analyses showed that interfering with BAIAP2L2 could significantly inhibit the expression of EMT markers in lung adenocarcinoma. After BAIAP2L2 knockdown, the expression levels of E-cadherin increased, whereas N-cadherin and Vimentin decreased compared to the si-NC group (Fig. 9C and E). Immunofluorescence analysis also showed that the fluorescence intensity of E-cadherin was enhanced, while Vimentin was reduced in lung adenocarcinoma cells (Fig. 9F and G). Consistently, Fig. 9G further confirms the reduction in Vimentin expression in PC9 cells, reinforcing the inhibitory effect of BAIAP2L2 silencing on EMT progression.
Although our data demonstrate that BAIAP2L2 influences EMT markers, the upstream molecular pathways by which BAIAP2L2 regulates EMT remain undefined. Identifying the intermediates or signaling axes, such as PI3K/AKT, Wnt/β-catenin, or TGF-β pathways, will be important for mechanistic understanding and therapeutic targeting in future investigations.
Discussion
Lung adenocarcinoma, a prevalent subtype of lung cancer, has consistently posed significant challenges in tumor research owing to its intricate biological behaviors and molecular causes17. Integrating multi-omics technologies and various analytical methods facilitates a thorough analysis of lung adenocarcinoma’s cellular, molecular, and clinical characteristics, yielding substantial theoretical evidence and potential targets for precise diagnosis and effective treatment18.
By analyzing the single-cell transcriptome data of lung adenocarcinoma tissues, we identified 10 different cell types, comprising 9 well-annotated kinds, including several immune and stromal cells. This provides a cellular basis for a deeper understanding of the tumor microenvironment of lung adenocarcinoma. Notably, we discovered 9 CNV subpopulations of lung adenocarcinoma epithelial cells, and the distribution percentages of these subpopulations varied significantly among different tumors, indicating substantial heterogeneity within lung adenocarcinoma epithelial cells. This heterogeneity may be an important cause of the differences in clinical outcomes among lung adenocarcinoma patients, as cells with different copy number states may have distinct proliferative, migratory, and invasive capabilities.
To further elucidate the evolutionary relationships among these different CNV subpopulations, we reconstructed the cell developmental trajectories. The results showed that subpopulations 6 and 8, with higher CNV levels, were located at the ends of the evolutionary trajectories, suggesting that they may represent highly malignant cell subpopulations. Subpopulation 4, serving as the starting point, indicates that normal cells may gradually transform into a highly malignant state through a series of molecular and cellular events. These findings are consistent with the traditional tumor evolution model, namely, that tumor cells gradually acquire malignant characteristics from normal cells. Simultaneously, this discovery provides potential targets for early tumor detection and intervention. For example, targeting cells in the middle of the evolutionary trajectory may prevent their further malignant transformation.
The expression of BAIAP2L2 in lung adenocarcinoma patients is closely related to clinical prognosis. Patients with high expression have a significantly worse prognosis, and this association is independent of other clinical factors. By comparing survival durations between high-risk and low-risk patients, we further substantiated the efficacy of BAIAP2L2 as an independent prognostic indicator. Furthermore, BAIAP2L2 is expressed at higher levels in patients with later T stages, indicating its potential involvement in critical processes during tumor growth.
In the pan-cancer analysis, we found that BAIAP2L2 exhibited prognostic value in multiple tumor types, indicating that it may play a similar role in the development and progression of various tumors. Meanwhile, the immune-related characteristics of BAIAP2L2 vary among different tumors, implying that it may have complex regulatory functions in the tumor immune microenvironment, not limited to lung adenocarcinoma. These findings expand the research scope of BAIAP2L2 and provide a theoretical basis for the development of multi-tumor targeted therapy strategies targeting BAIAP2L2.
Cellular experimental results showed that interfering with the expression of BAIAP2L2 can significantly inhibit multiple malignant biological behaviors of lung adenocarcinoma cells, including cell proliferation, migration, invasion, and colony formation ability, as well as the wound healing process. These results indicate that BAIAP2L2 plays a crucial role in maintaining the malignant phenotype of lung adenocarcinoma cells, likely through the regulation of multiple mechanisms such as cell stemness, proliferative signaling pathways, and intercellular communication.
Subsequent research demonstrated that BAIAP2L2 can markedly impede the epithelial-mesenchymal transition (EMT) in lung adenocarcinoma, a critical biological pathway via which tumor cells develop migratory and invasive properties. EMT endows tumor cells with migratory, invasive, and metastatic abilities, significantly affecting the progression, treatment resistance, and prognosis of lung adenocarcinoma19. During EMT, tumor cells acquire mesenchymal characteristics, enabling them to break through the basement membrane and invade blood vessels or the lymphatic system20. In lung adenocarcinoma, the downregulation of E-cadherin and upregulation of Vimentin are significantly associated with lymph node metastasis21. By regulating the expression of EMT-related markers such as E-cadherin, N-cadherin, and Vimentin, BAIAP2L2 may influence the epithelial-mesenchymal transition of tumor cells, thereby affecting their migratory and invasive abilities. This discovery not only reveals the mechanism by which BAIAP2L2 affects lung adenocarcinoma metastasis but also provides new insights into the development of therapeutic strategies targeting the EMT process.
As a valuable prognostic factor, BAIAP2L2 also plays an important role in other tumors. In hepatocellular carcinoma, BAIAP2L2 has been found to be overexpressed and promotes the migration and invasion of HCC cells22. BAIAP2L2 interacts with GABPB1 to inhibit its ubiquitin-mediated degradation and promote its nuclear translocation. Meanwhile, BAIAP2L2 regulates the level of reactive oxygen species (ROS) through GABPB1, thereby promoting the carcinogenic properties of HCC and reducing its sensitivity to lenvatinib23. Another study showed that BAIAP2L2 and JAK1 colocalize and interact within HCC cells, thereby enhancing the activation of the JAK1/STAT3 signaling pathway. The use of the JAK1 inhibitor Ruxolitinib can effectively reverse the cell proliferation, migration, invasion, and PD-L1 upregulation induced by BAIAP2L224. In gastric cancer, knocking down BAIAP2L2 has been found to inhibit the AKT/mTOR and Wnt3a/β-catenin signaling pathways25. Additionally, BAIAP2L2 has been associated with chemotherapy resistance, as it promotes the transfer of chemotherapy resistance from resistant cells to sensitive cells via extracellular vesicle proteins such as ANXA426. In prostate cancer, BAIAP2L2 is considered a poor prognostic factor, and the depletion of BAIAP2L2 can inactivate vascular endothelial growth factor (VEGF) in PCa cells and induce the apoptotic signaling pathway27,28. In osteosarcoma, the expression level of BAIAP2L2 has been found to be upregulated, and inhibiting BAIAP2L2 can lead to apoptosis of osteosarcoma cancer cells, inhibit their migration and invasion, and induce the inactivation of the Wnt/β-catenin pathway29.
The results of this study provide new strategies and targets for the clinical treatment of lung adenocarcinoma. Based on the expression level of BAIAP2L2, we can develop targeted drugs or diagnostic reagents for BAIAP2L2 to achieve precise diagnosis and personalized treatment of lung adenocarcinoma. For individuals with high BAIAP2L2 expression, prioritizing medicines that specifically target and inhibit BAIAP2L2 may enhance therapy outcomes. Moreover, considering the relationship between BAIAP2L2 and the tumor immune microenvironment, incorporating it as a factor in immunotherapy is also a direction for future research, which may provide new ideas for overcoming the problem of resistance to immune checkpoint inhibitors.
Although in vitro assays clearly support the oncogenic role of BAIAP2L2 in proliferation, migration, invasion, and EMT, animal models such as xenografts or orthotopic lung tumor models were not employed due to resource limitations. This represents a limitation of the current study. Future work incorporating in vivo systems will be essential to confirm the therapeutic potential of targeting BAIAP2L2 in lung adenocarcinoma.
Limitations
This study has some limitations. Public datasets, including TCGA and GEO, do not provide full N/M staging or organ-specific metastasis data, such as brain involvement. Consequently, we analyzed T stage only and could not test BAIAP2L2 against specific metastatic patterns or perform brain metastasis–based survival analysis. The data structure also did not support dynamic visualization along bifurcating pseudotime branches. In addition, vivo validation remains a valuable direction for future exploration. Future studies will leverage datasets with more detailed clinical and lineage information.
Conclusion
This study identified a novel biomarker, BAIAP2L2, for lung adenocarcinoma through the integration of multi-omics technologies. Knocking down the expression of BAIAP2L2 can inhibit the growth, migration, and invasion of lung adenocarcinoma cells, as well as the epithelial-mesenchymal transition of lung adenocarcinoma.
Supplementary Information
Below is the link to the electronic supplementary material.
Author contributions
All authors reviewed and approved the final version of the manuscript.
Data availability
The RNA-seq datasets analysed in this study are publicly available. The TCGA-LUAD dataset was obtained from The Cancer Genome Atlas (TCGA) database through the Genomic Data Commons (GDC) portal (https://portal.gdc.cancer.gov/projects/TCGA-LUAD; dbGaP Study Accession: phs000178; Project Name: Lung Adenocarcinoma). Additional datasets were obtained from the Gene Expression Omnibus (GEO) database, including GSE3141, GSE8894, GSE13212, GSE19188, GSE26939, GSE29013, GSE29016, GSE30219, GSE31210, GSE37745, GSE50081, and GSE72094. The single-cell RNA-seq dataset GSE131907 was also obtained from the GEO database (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE131907), with controlled access available at the European Genome-Phenome Archive (EGA) under accession number EGAD00001005054.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Bowei Jiang, Yuan Gan, Wanshuo Wei and Meichun Yang contributed equally to this work.
Contributor Information
Jianjun Wen, Email: 13877699698@163.com.
Zhongheng Wei, Email: Weizhongh1968@126.com.
Qijun Long, Email: longqijun248@163.com.
References
- 1.Bray, F. et al. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin.68 (6), 394–424 (2018). PMID 30207593. [DOI] [PubMed] [Google Scholar]
- 2.Vachhani, P. & Chen, H. Spotlight on pembrolizumab in non-small cell lung cancer: the evidence to date. Onco Targets Ther.9, 5855–5866. 10.2147/OTT.S97746 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Hou, H. et al. Distinctive targetable genotypes of younger patients with lung adenocarcinoma: a cBioPortal for cancer genomics data base analysis. Cancer Biol. Ther.21 (1), 26–33 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Altorki, N. K. et al. The lung microenvironment: an important regulator of tumour growth and metastasis. Nat. Rev. Cancer. 19 (1), 9–31. 10.1038/s41568-018-0081-9 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Nguyen, A., Yoshida, M., Goodarzi, H. & Tavazoie, S. F. Highly variable cancer subpopulations that exhibit enhanced transcriptome variability and metastatic fitness. Nat. Commun.7, 11246. 10.1038/ncomms11246 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Yu, K. H. et al. Predicting non-small cell lung cancer prognosis by fully automated microscopic pathology image features. Nat. Commun.7, 12474. 10.1038/ncomms12474 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Marguerat, S. & Bahler, J. RNA-seq: from technology to biology. Cell. Mol. Life Sci.67 (4), 569–579. 10.1007/s00018-009-0180-6 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Li, F. et al. In vivo epigenetic CRISPR screen identifies Asf1a as an immunotherapeutic target in Kras-Mutant lung adenocarcinoma. Cancer Discov.10 (2), 270–287. (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Hu, P., Zhang, W., Xin, H. & Deng, G. Single cell isolation and analysis. Front. Cell. Dev. Biol.4, 116. 10.3389/fcell.2016.00116 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hollern, D. P. et al. B cells and T follicular helper cells mediate response to checkpoint inhibitors in high mutation burden mouse models of breast cancer. Cell179 (5), 1191–1206e21. 10.1016/j.cell.2019.10.028 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Lim, S. B., Tan, S. J., Lim, W. T. & Lim, C. T. An extracellular matrix-related prognostic and predictive indicator for early-stage non-small cell lung cancer. Nat. Commun.8 (1), 1734. 10.1038/s41467-017-01430-6 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Xu, L. et al. BAI1–associated protein 2–like 2 is a potential biomarker in lung cancer. Oncol. Rep.41 (2), 1304–1312. 10.3892/or.2018.6883 (2019). [DOI] [PubMed] [Google Scholar]
- 13.Pykalainen, A. et al. Pinkbar is an epithelial-specific BAR domain protein that generates planar membrane structures. Nat. Struct. Mol. Biol.18 (8), 902–907. 10.1038/nsmb.2079 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Yu, J. et al. CCR7 promote lymph node metastasis via regulating VEGF-C/D-R3 pathway in lung adenocarcinoma. J. Cancer. 8 (11), 2060–2068. 10.7150/jca.19069 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Geeleher, P., Cox, N. & Huang, R. S. pRRophetic: an R package for prediction of clinical chemotherapeutic response from tumor gene expression levels. Plos One9 (9), e107468. 10.1371/journal.pone.0107468 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Liao, C. & Wang, X. TCGAplot: an R package for integrative pan-cancer analysis and visualization of TCGA multi-omics data. BMC Bioinform.24 (1), 483. 10.1186/s12859-023-05615-3 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Yin, N. et al. Protein kinase ciota and Wnt/beta-catenin signaling: Alternative pathways to Kras/Trp53-driven lung adenocarcinoma. Cancer Cell.36 (2), 156–167.e7 10.1016/j.ccell.2019.07.002 (2019). [DOI] [PMC free article] [PubMed]
- 18.Gillette, M. A. et al. Proteogenomic characterization reveals therapeutic vulnerabilities in lung adenocarcinoma. Cell182 (1), 200–225.e35. 10.1016/j.cell.2020.06.013 (2020). [DOI] [PMC free article] [PubMed]
- 19.Kim, Y. H. et al. Senescent tumor cells lead the collective invasion in thyroid cancer. Nat. Commun.8, 15208. 10.1038/ncomms15208 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wei, S. C. & Yang, J. Forcing through tumor metastasis: The interplay between tissue rigidity and epithelial-mesenchymal transition. Trends Cell Biol.26 (2), 111–120. 10.1016/j.tcb.2015.09.009 (2016). [DOI] [PMC free article] [PubMed]
- 21.Pacurari, M. et al. The microRNA-200 family targets multiple non-small cell lung cancer prognostic markers in H1299 cells and BEAS-2B cells. Int. J. Oncol.43 (2), 548–560. 10.3892/ijo.2013.1963 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Wei, H. et al. BAIAP2L2 is a novel prognostic biomarker related to migration and invasion of HCC and associated with cuprotosis. Sci. Rep.13 (1), 8692. 10.1038/s41598-023-35420-0 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Jia, W. et al. BAIAP2L2 promotes the malignancy of hepatocellular carcinoma via GABPB1-mediated reactive oxygen species imbalance. Cancer Gene Ther.31 (12), 1868–1883. 10.1038/s41417-024-00841-0 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Xie, Z. et al. BAIAP2L2 facilitates hepatocellular carcinoma progression and immune evasion of via targeting JAK1-mediated pathway and PD-L1 expression. Cancer Gene Ther.32 (4), 464–474. 10.1038/s41417-025-00890-z (2025). [DOI] [PubMed] [Google Scholar]
- 25.Liu, J., Shangguan, Y., Sun, J., Cong, W. & Xie, Y. BAIAP2L2 promotes the progression of gastric cancer via AKT/mTOR and Wnt3a/beta-catenin signaling pathways. Biomed. Pharmacother. 129, 110414. 10.1016/j.biopha.2020.110414 (2020). [DOI] [PubMed] [Google Scholar]
- 26.Liao, Y. et al. N6-methyladenosine RNA modified BAIAP2L2 facilitates extracellular vesicles-mediated chemoresistance transmission in gastric cancer. J. Transl Med.23 (1), 320. 10.1186/s12967-025-06340-6 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Song, Y., Zhuang, G., Li, J. & Zhang, M. BAIAP2L2 facilitates the malignancy of prostate cancer (PCa) via VEGF and apoptosis signaling pathways. Genes Genom.43 (4), 421–432. 10.1007/s13258-021-01061-8 (2021). [DOI] [PubMed] [Google Scholar]
- 28.Hu, W., Wang, G., Yarmus, L. B. & Wan, Y. Combined methylome and transcriptome analyses reveals potential therapeutic targets for EGFR wild type lung cancers with low PD-L1 expression. Cancers (Basel)12 (9). 10.3390/cancers12092496 (2020). [DOI] [PMC free article] [PubMed]
- 29.Guo, H. et al. BAIAP2L2 promotes the proliferation, migration and invasion of osteosarcoma associated with the Wnt/beta-catenin pathway. J. Bone Oncol.31, 100393. 10.1016/j.jbo.2021.100393 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The RNA-seq datasets analysed in this study are publicly available. The TCGA-LUAD dataset was obtained from The Cancer Genome Atlas (TCGA) database through the Genomic Data Commons (GDC) portal (https://portal.gdc.cancer.gov/projects/TCGA-LUAD; dbGaP Study Accession: phs000178; Project Name: Lung Adenocarcinoma). Additional datasets were obtained from the Gene Expression Omnibus (GEO) database, including GSE3141, GSE8894, GSE13212, GSE19188, GSE26939, GSE29013, GSE29016, GSE30219, GSE31210, GSE37745, GSE50081, and GSE72094. The single-cell RNA-seq dataset GSE131907 was also obtained from the GEO database (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi? acc=GSE131907), with controlled access available at the European Genome-Phenome Archive (EGA) under accession number EGAD00001005054.









