Skip to main content
Discover Oncology logoLink to Discover Oncology
. 2026 Jan 16;17:281. doi: 10.1007/s12672-026-04428-z

Integrating multiomics data using a correlation based graph attention network for subtype classification in lower grade glioma

Eman Mohammed Hamid 1, Murtada K Elbashir 2,✉, Nosiba Yousif Ahmed 1, Wafa Alameen Alsanousi 1,3, Abdulrahman Alyami 2,✉, Ayman Mohamed Mostafa 2, Mohanad Mohammed 5, Mohamed Elhafiz Musa 4, Mahmood A Mahmood 2
PMCID: PMC12891306  PMID: 41543639

Abstract

Accurate classification of cancer subtypes is crucial for personalised therapies and targeted interventions. In this study, we propose BioGAT-LGG, a deep learning framework that integrates multi-omics data, including mRNA, miRNA, and DNA methylation, using a correlation-based Graph Attention Network version 2 (GATv2) for biomarker discovery and Lower-Grade Glioma (LGG) subtype classification. Unlike existing methodologies that rely on external biological priors, such as protein-protein interaction networks or reference graphs, BioGAT-LGG constructs gene-driven correlation graphs, enabling the model to learn biologically meaningful molecular interactions. To improve feature interpretability and reduce dimensionality, LASSO regression is performed during model training. The model achieved 98.03% accuracy, with precision (98.12%), recall (97.74%), and F1-score (97.87%) in a stratified 10-fold cross-validation. Extensive analysis and enrichment of known cancer-related pathways, including PI3K-Akt signalling, Small Cell Lung Cancer, and Transcriptional Misregulation in Cancer, identified the biomarkers hsa-mir-3936, MTCO1P40, and CCND2, which were subsequently validated. These results indicate that BioGAT-LGG effectively captures biologically validated mechanisms and can enable clinically significant subtype classification and biomarker-guided decision-making. This framework thus lays a scalable foundation for multi-omics integration in oncology, which can be further adopted in other tumour types.

Keywords: Biomarker identification, GATv2, Multi-omics data, LGG, Correlation-based graph, Gene ontology, Cancer subtype classification, KEGG pathways

Introduction

LGG is a diverse group of primary brain tumours characterised by distinct molecular profiles, therapeutic responses, and clinical outcomes. To guide treatment decisions and improve patient outcomes, LGG subtypes must be accurately classified. However, the intricate relationships among the transcriptional, epigenetic, and genetic layers that contribute to tumour heterogeneity are beyond the reach of current methods [1].

Recent advances in deep learning and multi-omics profiling have created opportunities to address these issues. Multi-omics data provide a comprehensive representation of tumour biology, and deep learning models can capture high-dimensional, nonlinear relationships among molecular features [2]. However, most current frameworks rely on static biological networks, such as protein–protein interactions (PPIs), which may not capture interactions specific to LGG subtypes in a given context. By constructing correlation-based graphs over gene-level features derived from from patient data, subtype-specific modelling becomes possible without external databases, thereby reducing bias and increasing adaptability.

Tumour subtyping and biomarker discovery have advanced significantly through the integration of deep learning and cancer genomics. Previous studies have presented several approaches, each with distinct strengths and limitations, as outlined below, with an emphasis on the approaches, results, and drawbacks. DeepAutoGlioma, an autoencoder-based framework for LGG subtyping that integrates transcriptome and methylome data, was presented by Munquad et al. [3]. Cox survival analysis was used to guide feature selection. With an accuracy of 98.03% (± 0.06), the model demonstrated that deep latent multi-omics representations are effective for accurate subtype classification. However, limited interpretability and reliance on superficial fusion constrained biological insight into subtype-specific biology. Wu et al. [4] presented DeepMoIC, a supervised framework that used deep neural layers to model cross-modal dependencies. The model outperformed conventional baselines, achieving an accuracy of about 90%. However, its integration was constant across patients, limiting its ability to capture interactions specific to each sample. Another multi-omics diagnostic framework, MOGDx, was presented by Ryan et al. [5] and achieved an accuracy of 89.9%. Although it showed that heterogeneous integration was feasible, the model’s small cohort size and high-dimensional features led to overfitting, and it offered little mechanistic insight. Singh et al. [6] used the DIABLO framework for supervised multi-omics integration. It is based on sparse partial least squares. Compact biomarker signatures with better interpretability were identified. However, as a linear method, DIABLO underfits nonlinear dependencies and cannot model higher-order or context-specific interactions. Gao et al. [7] developed DeepCC, a deep learning framework that mapped gene expression data—and later multi-omics data—onto consensus cancer subtypes. It achieved reproducible and robust subtype assignments. Still, the architecture relied mainly on dense layers, limiting visibility into regulatory edges, such as miRNA–mRNA (Messenger RNA) or methylation–expression. Sun et al. [8] introduced SADLN (Self-Attentive Deep Learning Network), which employed attention to weight modalities during fusion. It improved representation quality and interpretability, but operated on vectorised inputs rather than explicit graphs, limiting its ability to capture relational structure and sample-level topology. DeepMO, a family of modality-specific encoders that project into a common latent space, was introduced by Lin et al. [9] with regularisation to improve stability. It performed well; however, due to the absence of explicit graph constraints, cross-omics dependencies remained implicit and harder to interpret. Early GNN research used GCN or GraphSAGE layers to propagate signals and built graphs from static resources (PPI, pathways). Although limited by incomplete reference networks that did not capture sample-specific rewiring, these approaches increased biological plausibility. To improve classification and survival prediction and provide feature-level interpretability, subsequent work employed GATs to learn edge-wise attention coefficients. However, the majority used static graphs, which might have missed dynamic and subtype-specific regulatory rewiring.

Recent techniques have enabled the inference of context-specific topology directly from data by substituting networks based on mutual information or correlation for static graphs. Despite their potential, these approaches struggle to filter out noisy edges, strike a balance between biological coverage and sparsity, and ensure robustness in high-dimensional multi-omics.

In conclusion, glioma subtyping has improved thanks to deep learning. Models such as DeepAutoGlioma, DeepMoIC, and MOGDx have demonstrated high accuracy, while DIABLO has improved interpretability, yet underfits nonlinear relationships. Although they did not explicitly model graph structure, attention-based models such as SADLN enhanced feature weighting. Although they mostly relied on static networks, GNNs and GATs introduced relational inductive biases. To address this limitation, adata-driven graph construction approach has been developed; however, issues with noise, stability, and interpretability remain. As a result, there is a glaring gap: current models either perform omics fusion without explicitly modelling cross-omics dependencies or rely on fixed graphs that cannot capture sample-specific regulation. Furthermore, rather than being incorporated into the model design, interpretability is often left to post-hoc analysis.

Here, we introduce BioGAT-LGG, a correlation-based graph attention framework that integrates multi-omics data for the classification of LGG subtypes. More significantly, the model captures biologically relevant inter-omics interactions related to tumorigenesis and cancer prognosis, while achieving strong predictive performance. BioGAT-LGG identifies new candidates beyond known subtype-specific biomarkers through attention-driven interpretability. Additionally, we used survival analyses to confirm the prognostic significance of these biomarkers and to examine their association with patient outcomes. Thus, in addition to precise classification, the framework provides clear insights into biomarker identification and potential prognostic evaluation in LGG.

Materials and methods

To classify LGG into three molecular subtypes—IDHmut-codel (n = 169), IDHmut-non-codel (n = 245), and IDH wildtype (IDHwt) (n = 95) rather than performing a binary LGG versus normal classification, we developed a deep learning framework (BioGAT-LGG). The framework, illustrated in Fig. 1, integrates RNA-seq, miRNA, and DNA methylation data within a correlation-based GATv2.

Fig. 1.

Fig. 1

Overview of the BioGAT-LGG model for LGG subtype classification using multi-omics data and GATv2-based learning

Deep learning models have demonstrated significant success in extracting hierarchical structure from complex biological data [10, 11]. Rather than developing networks from external datasets, we built correlation-based graphs from the omics datasets, enabling the model to learn molecular relationships specific to the biological context. LASSO regression was used to select features for modelling, reducing dimensionality while retaining informative biomarkers [12, 13]. The resulting graphs were then passed into the GATv2, which utilised attention-based mechanisms to model higher-order interactions. The biological relevance of the selected biomarkers was assessed using GO and KEGG enrichment analyses, as outlined by Kanehisa, Furumichi [14].

Omics data acquisition and preprocessing

LGG multi-omics data were obtained from TCGA-LGG using the TCGAbiolinks R package, following the methods in [15]. The data comprised RNA-seq, miRNA expression, and DNA methylation levels from normal and tumour samples. Preprocessing included normalisation, filtering, and annotation, performed using the GDCquery and GDCprepare functions to ensure data quality and consistency.

Gene expression counts were normalised by DESeq2 (v1.44.0) under a generalised linear model [16], and low-expressed genes were filtered out by count-based thresholds. For miRNA and DNA methylation, LIMMA (v3.60.6) was employed with voom transformation [17, 18]. For variance stabilisation and linear modelling. miRNA features with CPM > 1 were retained, reducing 1,881 features to 771, of which 588 were found to be significantly differentially expressed (adjusted p < 0.05). For DNA methylation, cross-reactive positions, sex-chromosome loci, and missing values in the probes were excluded, and there were 298,588 CpGs remaining. LIMMA analysis identified 2,838 significantly differentially methylated CpG positions (adjusted p < 0.05) [19].

These omics-specific filtering and preprocessing steps are tabulated in Table 1, showing the counts at each stage before data integration. Diagnostic plots of dispersion estimates, fold changes, and voom mean–variance trends for mRNA (a, b), miRNA (c), and DNA methylation (d) are shown in Fig. 2, providing visual confirmation of the preprocessing.

Table 1.

Feature filtering per omics data type (Pre-integration)

Step mRNA miRNA DNA methylation
Original features 60,660 1,881 330,974
Removing sex chromosomes 57,670 – 324,232
Differentially expressed analysis 19,193 771 256,584
LIMMA model (selected features) – 588 238,245
LASSO (global, pre-revision)* 3,395 161 90

This table summarises the preliminary filtering performed on each omics dataset before integration

Fig. 2.

Fig. 2

Diagnostic plots showing dispersion, fold change, and voom mean-variance trends for mRNA (a, b), miRNA (c), and DNA methylation (d) data

In addition to the initial filtering described above, a refined preprocessing and variance-based feature selection pipeline was applied to ensure data quality, remove redundancy, and maintain reproducibility across omics layers. The initial filtering steps are summarised in Table 1, and the final per-omics feature selection results are presented in Table 2.

Table 2.

Final feature filtering and selection per omics data type (Used in model Training)

Step mRNA miRNA DNA methylation
Filtered features (input to pipeline) 25,315 588 238,245
Low-variance filtering 10,500 310 7,800
Feature selection (per-omics) 6,328 180 1,073

This table shows the refined preprocessing pipeline adopted in the final BioGAT-LGG model. The resulting omics-specific feature sets were integrated into a unified multi-omics matrix containing 7,581 high-variance features across 509 samples

As per these protocols, the three omics layers were integrated at the feature level into a single multi-omics feature matrix and subjected to variance-threshold filtering, yielding 7,580 high-variance features. This reduction is shown in Table 2. The resulting filtered and integrated matrix served as input to per-fold LASSO feature selection (Sect. 2.2) and downstream correlation-graph construction.

To avoid ambiguity about preprocessing, we note that the variance-based filtering used here is an unsupervised, label-independent step. Because variance does not use subtype information, applying this filter before splitting does not provide the model with any class-specific signal. However, from a strict ML perspective, preprocessing the full dataset may introduce a minor theoretical dependency between the training and test sets. For full transparency, this design choice is now explicitly stated, and we emphasise that all supervised steps (LASSO, scaling, SMOTE, and graph construction) were performed strictly within each training fold to avoid data leakage. Subsequent refinement and variance-filtering steps were applied to obtain the final feature sets used for model training (see Table 2).

Feature selection with LASSO regression

Within each training fold, features were divided into omics-specific blocks (RNA-Seq, miRNA, and DNA methylation). LASSO logistic regression with cross-validation and L1 regularisation (SAGA solver) was applied only to the training data within each omics block to identify features with non-zero coefficients. The number of selected features varied across folds, reflecting inherent data-driven variability in feature relevance and enhancing the model’s generalisability.

To ensure methodological integrity and prevent data leakage, the feature matrix (X) and target labels (Y) were separated before any preprocessing or feature selection. For each fold of stratified cross-validation, LASSO feature selection, feature scaling, oversampling (SMOTE), and correlation graph construction were performed exclusively on the training set. The trained models and transformation parameters (e.g., scaling means and variances) were then applied to the test set without re-estimation. This procedure ensures that no knowledge of the test set influences model training or feature selection.

LASSO regression is a form of linear regression that employs L1 regularisation [20, 21], as defined in Eq. 1. It imposes a penalty on the absolute values of model coefficients, inducing sparsity by shrinking less significant coefficients towards zero. The objective function is expressed as:

graphic file with name d33e577.gif 1

whereInline graphicis the target variable (LGG subtype classification), Inline graphic represents the feature values for the sample Inline graphic, Inline graphic Are the regression coefficients and Inline graphic is the regularisation parameter controlling feature sparsity.

The LASSO logistic regression, optimised by the SAGA solver, was then trained individually on each training fold of a stratified 10-fold cross-validation. The same selected feature set was projected onto its corresponding test fold to ensure identical feature spaces without information leakage. The number of selected features varied across folds, as expected given data-driven variability, reaching up to 3,646 in some folds (Table 3). This fold-wise LASSO feature selection procedure improves model generalizability, reduces overfitting risk, and ensures that no test data were used during feature selection.

Table 3.

Final integrated feature reduction (Post-integration)

Step Number of Features
Integrated features after filtering 7,580
LASSO (per fold, max cap) ≤ 3,646
Sample size 509
Subtype distribution (IDHmut-non-codel / IDHmut-codel / IDHwt) 245 / 169 / 95

Feature scaling and normalisation

To ensure consistent integration across heterogeneous omics layers (mRNA, miRNA, DNA methylation), feature scaling was applied in a fold-wise manner using training data statistics to standardise value ranges. Standardisation centres and scales each feature to have zero mean and unit variance, thereby enhancing the performance of deep models such as GAT. The transformation is defined in Eq. 2. Each omics layer was scaled independently to account for biological variation. Unlike PCA-based approaches, this method first selected biologically meaningful features via differential expression analysis, preserving interpretability while improving the model’s numerical stability and robustness.

graphic file with name d33e653.gif 2

Where Inline graphic Is the original feature value, Inline graphic Is the mean of the feature, and Inline graphic Is the standard deviation of the feature.

For DNA methylation data, we used β-values directly as obtained from the TCGA legacy pipeline, without transforming to M-values. We opted not to apply Z-score normalisation to preserve the interpretability of methylation levels, as β-values represent the proportion of methylation at a given CpG site. This decision aligns with standard practices in integrative multi-omics analyses, which retain β-values for their direct biological meaning [22].

GATv2

The BioGAT-LGG model uses a two-layer GATv2 [23] to capture interactions among mRNA, miRNA, and DNA methylation data for LGG subtype prediction. Unlike GCN, GATv2 learns adaptive attention weights for neighbours, enabling it to focus on biologically significant relationships in high-dimensional multi-omics graphs [24] .

LASSO-based feature selection and standard scaling were applied, and SMOTE [25, 26] was used to balance the training data and address class imbalance. Synthetic samples were generated as described in Eq. 3. The model architecture comprises two GATv2 layers: one with eight attention heads and dropout (0.4), and the other with one head. A fully connected layer with ReLU activation completes the classification. Forward propagation is defined as in Eq. 4. Furthermore, the computation of αiⱼ, representing attention-based node interaction in multi-Omics, is given by Eq. 5.

graphic file with name d33e706.gif 3

where Inline graphicis a randomly selected sample from the minority class, ​Inline graphic is one of the k-nearest neighbours of Inline graphic ​, and L is a random number in the range (0,1) that determines the interpolation factor.

graphic file with name d33e724.gif 4

where ​Inline graphic represents the hidden representation of the node Inline graphic at layer Inline graphic, Inline graphic is the learnable weight matrix, and Inline graphicIs the attention coefficient, calculated as:

graphic file with name d33e751.gif 5

where a is the attention weight vector, this architecture effectively models multi-omics data while preserving essential biological relationships. W is a learned weight matrix applied to node features Inline graphicand Inline graphic, transforming them into a new representation.

In our graph formulation, nodes represent molecular attributes (genes, miRNAs, CpG sites) rather than patient samples. This enables the model to capture inter-feature biological relationships, which are essential for biomarker-based interpretation. Although hyperparameters (e.g., number of heads, dropout rate) were not tuned via grid search, the architecture was validated using stratified 10-fold cross-validation and stabilised with learning-rate scheduling and gradient clipping.

Compared with traditional approaches that rely on external biological priors, such as STRING PPI networks [27] the BioGAT-LGG model constructs a correlation-based graph directly from the combined multi-omics dataset. This approach captures gene–genes-dependent molecular dependencies without relying on curated interaction networks that may be incomplete, biased towards well-characterised proteins, or unrepresentative of brain or tumour-specific milieus [28]. While PPI-based graphs are biologically relevant, they are static and do not represent dynamic or condition-specific interactions [29]. In contrast, the correlation-based graphs capture true co-expression and co-methylation in the studied cohort, enabling the model to learn biologically relevant, context-specific interactions for LGG subtype determination. This makes the model more interpretable and favourable for biomarker discovery, further supported by the existing literature that prioritises data-driven networks in multi-omics research [30].

Model training and Cross-Validation

To evaluate the performance and generalisability of the BioGAT-LGG model, we employed 10-fold stratified cross-validation, ensuring balanced representation of LGG subtypes across all folds. Model training was optimised using the Adam optimiser with weight decay to mitigate overfitting. A learning-rate scheduler dynamically adjusted the step size based on validation loss, and gradient clipping was applied to stabilise training. The network parameters were updated using the GATv2 node update rule as defined in Eq. (6):

graphic file with name d33e788.gif 6

To optimise classification performance, the model was trained to minimise the cross-entropy loss function:

graphic file with name d33e794.gif 7

where N is the number of samples, C is the number of classes, Inline graphicis the actual output for the i-th sample and j-th class, and Inline graphic is the predicted output for the i-th sample and j-th class. To assess feature importance, we extracted the attention weights from the first GATv2 layer and computed their absolute mean values across cross-validation folds, as shown in Eq. (8):

graphic file with name d33e811.gif 8

Here, K represents the number of folds. This strategy enabled the identification of candidate biomarkers with strong discriminative power, thereby offering essential molecular insights for accurate LGG subtype classification.

Model stability

In addition to the default 10-fold setup, stratified 5-fold and 20-fold CV yielded consistent results (see Table 4), confirming robustness across fold numbers.

Table 4.

Performance stability across k-fold values (stratified CV, TCGA-LGG) for BioGAT-LGG; values are mean ± SD

k-fold Accuracy (mean ± SD) Precision (mean ± SD) Recall (mean ± SD) F1 (mean ± SD)
5 0.9725 ± 0.0114 0.9717 ± 0.0178 0.9720 ± 0.0132 0.9707 ± 0.0130
10 0.9803 (baseline) 0.9812 (baseline) 0.9774 (baseline) 0.9787 (baseline)
20 0.9883 ± 0.0216 0.9874 ± 0.0280 0.9882 ± 0.0228 0.9867 ± 0.0253

The 10-fold row reports the baseline

Hardware and training time

The BioGAT-LGG model was trained and evaluated on the Kaggle cloud platform, which provides a free computational environment with 13 GB RAM and an optional NVIDIA Tesla T4 GPU. Before training, the multi-omics datasets were downloaded and preprocessed on a local workstation with an Intel(R) Core(TM) i7-9700 CPU @ 3.00 GHz, 32 GB RAM, and an NVIDIA GeForce GTX 1080 GPU (8 GB). Each fold of the 10-fold stratified cross-validation took approximately 4 to 6 min to train, yielding a total training time of around 45 min. This setup ensured efficient computation while maintaining reproducibility. All code and processed datasets used in this study have been deposited on Zenodo for transparency and reproducibility (DOI: 10.5281/zenodo.15761178).

Results

The proposed BioGAT-LGG model uses a correlation-driven GATv2 to incorporate multi-omics features—mRNA, miRNA, and DNA methylation— for the subtype classification of LGG. The model’s design enables it to attend to biologically important correlations by dynamically adjusting attention weights across features. Synchronising correlation graphs with attention can convey a more informed structure and better optimise predictive power and interpretability. The model’s stability was evaluated using 10-fold stratified cross-validation, with balanced representation of all subtypes in both the training and validation sets. The subsequent sections present the model’s performance, validation metrics, and an interpretation of its learning dynamics.

Model performance and evaluation

BioGAT-LGG achieved strong performance in 10-fold cross-validation, with an accuracy of 98.03%, precision of 98.12%, recall of 97.74%, and F1-score of 97.87%, demonstrating its ability to represent complex molecular interactions and classify LGG subtypes with high confidence. The training and validation loss curves in Fig. 3(a) indicate stable convergence, while the adaptive learning rate strategy in Fig. 3(b) supported generalization by preventing early convergence. Furthermore, an additional held-out test set (34 IDHmut-codel, 49 IDHmut-non-codel, and 19 IDHwt) confirmed the model’s generalizability, as feature selection and graph construction were performed only on the training split.

Fig. 3.

Fig. 3

a Training vs. Validation Loss Curve – Demonstrates the model’s convergence, with training and validation losses stabilizing over epochs, indicating effective learning and minimal overfitting. b) Learning Rate Curve – Illustrates the adaptive learning rate adjustments, ensuring optimal training progression and preventing divergence

In addition to the stratified 10-fold cross-validation, we reserved 20% of the samples (34 IDHmut-codel, 49 IDHmut-non-codel, 19 IDHwt) as a separate held-out test set to provide an independent evaluation of model performance. Feature selection, scaling, and graph construction were performed using only the training split, and the held-out test set was used solely for final evaluation.

To further confirm the benefit of multi-omics integration, we systematically compared single-omics, pairwise, and full multi-omics models across multiple baselines (LASSO, PCA, Random Forest). As summarised in Table 5, the integrated multi-omics model consistently outperformed single-omics and pairwise approaches, with GATv2 + FeatureAttention achieving the highest overall accuracy (98.03%). These findings confirm the added predictive value of multi-omics integration.

Table 5.

Performance comparison of single-omics, pairwise, and multi-omics models (ACC, F1, AUC; mean ± SD)

Subset Model ACC (mean ± std) F1m (mean ± std) AUCm (mean ± std)
mRNA LASSO_LogReg 0.963 ± 0.021 0.958 ± 0.024 0.986 ± 0.012
mRNA PCA + LR 0.947 ± 0.025 0.942 ± 0.024 0.988 ± 0.010
mRNA RandomForest 0.961 ± 0.025 0.958 ± 0.028 0.993 ± 0.011
mRNA GATv2 + FeatureAttention 0.969 ± 0.022 0.946 ± 0.024 0.982 ± 0.013
miRNA LASSO_LogReg 0.949 ± 0.031 0.941 ± 0.030 0.988 ± 0.015
miRNA PCA + LR 0.939 ± 0.018 0.934 ± 0.020 0.987 ± 0.015
miRNA RandomForest 0.937 ± 0.033 0.934 ± 0.033 0.988 ± 0.012
miRNA GATv2 + FeatureAttention 0.951 ± 0.014 0.927 ± 0.015 0.979 ± 0.010
Methy LASSO_LogReg 0.970 ± 0.023 0.982 ± 0.022 0.996 ± 0.007
Methy PCA + LR 0.974 ± 0.018 0.986 ± 0.016 0.996 ± 0.006
Methy RandomForest 0.972 ± 0.020 0.983 ± 0.019 0.997 ± 0.005
Methy GATv2 + FeatureAttention 0.975 ± 0.018 0.955 ± 0.020 0.987 ± 0.008
mRNA + miRNA LASSO_LogReg 0.963 ± 0.021 0.958 ± 0.024 0.986 ± 0.012
mRNA + miRNA PCA + LR 0.945 ± 0.020 0.940 ± 0.021 0.988 ± 0.011
mRNA + miRNA RandomForest 0.965 ± 0.023 0.961 ± 0.027 0.992 ± 0.011
mRNA + miRNA GATv2 + FeatureAttention 0.975 ± 0.032 0.935 ± 0.040 0.984 ± 0.014
mRNA + Methy LASSO_LogReg 0.963 ± 0.024 0.958 ± 0.026 0.985 ± 0.011
mRNA + Methy PCA + LR 0.943 ± 0.018 0.937 ± 0.019 0.987 ± 0.010
mRNA + Methy RandomForest 0.967 ± 0.022 0.964 ± 0.025 0.993 ± 0.011
mRNA + Methy GATv2 + FeatureAttention 0.967 ± 0.011 0.964 ± 0.015 0.992 ± 0.006
miRNA + Methy LASSO_LogReg 0.970 ± 0.023 0.981 ± 0.023 0.996 ± 0.007
miRNA + Methy PCA + LR 0.970 ± 0.016 0.983 ± 0.014 0.995 ± 0.006
miRNA + Methy RandomForest 0.972 ± 0.020 0.983 ± 0.019 0.996 ± 0.005
miRNA + Methy GATv2 + FeatureAttention 0.977 ± 0.011 0.954 ± 0.013 0.989 ± 0.007
All LASSO_LogReg 0.965 ± 0.022 0.960 ± 0.023 0.990 ± 0.011
All PCA + LR 0.945 ± 0.017 0.940 ± 0.017 0.987 ± 0.011
All RandomForest 0.965 ± 0.023 0.961 ± 0.027 0.993 ± 0.009
All GATv2 + FeatureAttention 0.9803 ± 0.005 0.955 ± 0.008 0.989 ± 0.008

Evaluation metrics

Accurate LGG subtype classification is essential to advance precision oncology. BioGAT-LGG fuses mRNA, miRNA, and DNA meth—data within a correlation-based GATv2, enabling the model to capture complex molecular interactions. The model performed robustly in 10-fold stratified cross-validation: 98.03% accuracy, 98.12% precision, 97.74% recall, and 97.87% F1-score, with minimal variance across folds. The mean performance scores and their standard deviations are illustrated in Fig. 4.

Fig. 4.

Fig. 4

BioGAT-LGG model performance across evaluation metrics (mean ± standard deviation)

Receiver Operating Characteristic (ROC) analysis also indicated the model’s strength, with AUC values of 0.99 for IDH-mutant with 1p/19q co-deletion, 0.98 for IDH-mutant without co-deletion, and 1.00 for IDH wildtype, as shown in Fig. 5, demonstrating excellent discriminative power and minimal risk of misclassification.

Fig. 5.

Fig. 5

ROC curve for LGG subtype classification. The model shows excellent discriminative performance, with Area Under the Curve (AUC) scores of 0.998 for IDH-mutant with 1p/19q co-deletion, 0.995 for IDH-mutant without 1p/19q co-deletion, and 0.994 for IDH-wildtype, with a macro-average AUC of 0.996

The TCGA-LGG dataset comprised 509 samples across three subtypes (169 IDHmut-codel, 245 IDHmut-non-codel, and 95 IDHwt). Stratified cross-validation was used to preserve class ratios in each fold. For independent evaluation, an 80/20 stratified split was applied to maintain representative subtype distributions (TRAIN: 196/135/76; TEST: 49/34/19). Feature selection was performed exclusively on the training set using L1-penalised logistic regression with cross-validation, followed by scaling and graph construction, thereby avoiding information leakage from the test set. The classification results on the held-out test set are summarised in Table 6, showing near-perfect performance with only a few misclassifications.

Table 6.

Confusion matrix for LGG subtype classification using BioGAT-LGG

Actual / predicted IDH-mutant with 1p/19q co-deletion IDH-mutant without 1p/19q co-deletion IDH-wildtype Total
IDH-mutant with 1p/19q co-deletion 49 1 0 50
IDH-mutant without 1p/19q co-deletion 0 47 2 49
IDH-wildtype 0 0 49 49
Total 49 48 51 148

Biomarker discovery

Biomarkers play a role in the staging and treatment of LGG. IDH1/IDH2 mutations, either alone or in combination with 1p/19q co-deletion, are recognised as favourable prognostic markers, whereas a series of genetic and epigenetic alterations are implicated in glioma pathogenesis. Feature importance analysis identified significant biomarkers, including MTCO1P40, CCND2, PRAME, MT1G, PRG4, MTCO3P12, PTGS2, TRIM48, TIMP4, ADNP, C1QL2, RND3, RXRA, and PENK. These biomarkers are involved in crucial processes in LGG, such as gene expression regulation, miRNA regulation, and epigenetic modification. Statistical significance and supporting evidence for these biomarkers, along with the genes’ identifiers, names, and reference studies, are provided in Table 7. Specific LGG subtypes are significantly associated with these biomarkers, as indicated by lower p-values. The agreement of these results with previous cancer studies further confirms the potential of these biomarkers as therapeutic and diagnostic tools.

Table 7.

Key biomarkers identified for LGG subtype classification

Gene ID Gene name Evidence (REF)
ENSG00000181195 PENK PMID: 38,981,212 [31]
ENSG00000186350 RXRA PMID: 38,372,090 [32]
ENSG00000115963 RND3 PMID: 31,819,475 [33]
ENSG00000144119 C1QL2 PMID: 20,525,073 [34]
ENSG00000101126 ADNP PMID: 39,055,895 [35]
ENSG00000157150 TIMP4 PMID: 31,677,819 [36]
ENSG00000150244 TRIM48 PMID: 31,703,057 [37]
ENSG00000073756 PTGS2 PMID: 25,966,079 [38]
ENSG00000198744 MTCO3P12 -
ENSG00000116690 PRG4 PMID: 34,622,503 [39]
ENSG00000125144 MT1G PMID: 36,835,877 [40]
ENSG00000185686 PRAME PMID: 37,796,789 [41]
ENSG00000118971 CCND2 PMID: 27,923,660 [42]
ENSG00000262902 MTCO1P40 -

Survival analysis of candidate biomarkers

To evaluate the prognostic value of the identified biomarkers, we conducted survival analyses using Kaplan–Meier (KM) curves, log-rank tests, and Cox proportional hazards regression. Among the candidates tested, hsa-mir-4784 (log-rank p = 0.0026, FDR = 0.0212) and MTCO1P40 (ENSG00000262902.1) (log-rank p = 7.3 × 10⁻⁴, FDR = 0.0117) showed statistically significant differences in survival between high- and low-expression groups (Fig. 6). Additionally, CCND2 (ENSG00000118971.9) and MTCO3P12 (ENSG00000198744.5) showed borderline prognostic associations (p < 0.05 before FDR adjustment). Univariate Cox regression also established the predictive value of several markers (e.g., TRIM48, PRG4, CCND2, RND3, ADNP, PRAME), as indicated by consistently > 1 hazard ratios (Table 8). A multimarker signature (Z-score sum) also stratified patients into high- and low-risk groups with markedly different survival. These findings suggest that BioGAT-LGG not only identifies biologically relevant biomarkers but also prioritises clinically prognostic candidates for LGG.

Fig. 6.

Fig. 6

Kaplan–Meier survival curves for (a) MTCO1P40 (ENSG00000262902.1) and (b) hsa-mir-4784. Patients were stratified into high- and low-expression groups according to median expression levels. Both biomarkers showed significant differences in overall survival between groups (p = 7.3 × 10⁻⁴ and p = 0.0027, respectively, log-rank test)

Table 8.

Univariate Cox regression and log-rank test results for candidate biomarkers in LGG

Biomarker HR (95% CI) Cox p Log-rank p FDR
TRIM48 (ENSG00000150244.12) 1.0002 (1.00008–1.00023) 2.8e-05 0.404 0.539
PRG4 (ENSG00000116690.13) 1.0004 (1.0002–1.00065) 1.5e-04 0.280 0.491
CCND2 (ENSG00000118971.9) 1.00002 (1.00001–1.00003) 0.0015 0.0099 0.053
RND3 (ENSG00000115963.13) 1.0001 (1.00002–1.00021) 0.020 0.0488 0.156
ADNP (ENSG00000101126.18) 1.0001 (1.00001–1.00021) 0.033 0.224 0.448
PRAME (ENSG00000185686.18) 1.0001 (1.00001–1.00024) 0.037 0.760 0.779
MTCO3P12 (ENSG00000198744.5) 0.9983 (0.9964–1.0001) 0.065 0.0133 0.053
hsa-mir-4784 1.15 (0.99–1.33) 0.073 0.00265 0.021
MT1G (ENSG00000125144.14) 1.0002 (0.99997–1.00036) 0.090 0.201 0.448
MTCO1P40 (ENSG00000262902.1) 1.00004 (0.99996–1.00011) 0.299 7.3e-04 0.012
PENK (ENSG00000181195.11) 1.0001 (0.99988–1.00033) 0.372 0.654 0.747
PTGS2 (ENSG00000073756.12) 0.9999 (0.99963–1.00016) 0.432 0.779 0.779
C1QL2 (ENSG00000144119.4) 0.9999 (0.99967–1.00021) 0.657 0.0787 0.210
TIMP4 (ENSG00000157150.5) 1.0000 (0.99996–1.00003) 0.661 0.368 0.535
hsa-mir-1251 1.01 (0.90–1.14) 0.840 0.307 0.491
RXRA (ENSG00000186350.12) 1.0000 (0.99992–1.00009) 0.889 0.530 0.652

To further validate the efficiency of attention-based feature importance, we compared the top-ranked features identified by the attention mechanism with those selected by LASSO regression. As summarised in Table 9, several biomarkers were selected by both methods (Shared), while others were attention-specific or LASSO-specific. This confirms the reliability of our method, while attention-specific features reveal new candidate markers that may not be captured by conventional linear methods.

Table 9.

Comparison of features identified by LASSO regression and attention-based importance, highlighting shared and novel candidates

Feature att_weight abs_lasso Type Novelty Category
hsa-mir-4661 0.000413 0.0254 miRNA Shared Shared
hsa-mir-3615 0.000238 0.023223 miRNA Shared Shared
hsa-mir-6822 0.000238 0.031668 miRNA Shared Shared
cg21195256 0.000883 Attention-only Attention-only
cg08109136 0.000784 Attention-only Attention-only
cg04077695 0.000763 Attention-only Attention-only
cg14530834 0.00074 Attention-only Attention-only
cg23090824 0.000727 Attention-only Attention-only
cg13193548 0.0007 Attention-only Attention-only
cg18437792 0.000673 Attention-only Attention-only
cg24391122 0.000644 Attention-only Attention-only
cg13598010 0.000643 Attention-only Attention-only
cg07250222 0.000643 Attention-only Attention-only
cg15057442 0.080247 Methylation LASSO-only LASSO-only
cg01987516 0.068525 Methylation LASSO-only LASSO-only
hsa-mir-425 0.067563 miRNA LASSO-only LASSO-only
cg08535474 0.058317 Methylation LASSO-only LASSO-only
cg13912117 0.056736 Methylation LASSO-only LASSO-only
cg25293806 0.053911 Methylation LASSO-only LASSO-only
cg00359661 0.052115 Methylation LASSO-only LASSO-only
cg00699392 0.050529 Methylation LASSO-only LASSO-only
cg04546413 0.049291 Methylation LASSO-only LASSO-only
cg06610376 0.044823 Methylation LASSO-only LASSO-only

Cross-referencing top MiRNAs with curated interaction databases

To determine whether the top BioGAT-LGG miRNAs are part of known regulatory pathways, we examined curated databases of experimentally confirmed miRNA–mRNA interactions (miRTarBase, DIANA-TarBase) and CLIP-seq-based resources (ENCORI/starBase). For the specific miRNA–mRNA pairs highlighted in this study (top miRNAs × candidate mRNAs), we did not identify any curated experimental MTIs in our analysis (accessed Sep 17, 2025). Therefore, we report high-confidence predicted links from TargetScan and miRDB, clearly labelled with database scores/IDs (Table 10). For additional context, we also include validated targets for these same miRNAs in non-LGG cases (Table 11). Notably, two candidates—hsa-miR-4784 and MTCO1P40—showed FDR-significant survival associations (log-rank FDR = 0.0212 and 0.0117, respectively), while others (e.g., CCND2, MTCO3P12) showed nominal signals before FDR correction. These findings support the biological relevance of our biomarkers, even though there are currently no curated experimental MTIs for the exact pairs.

Table 10.

High-confidence predicted targets for top MiRNAs (for completeness; no overlap with study candidate mRNAs)

miRNA Target_Gene Evidence_Type Score/assay Database
hsa-miR-4784 ZNF652 Predicted miRDB score 95 miRDB
hsa-miR-4784 FOXJ2 Predicted miRDB score 94 miRDB
hsa-miR-4784 JPH4 Predicted miRDB score 94 miRDB
hsa-miR-4784 USB1 Predicted miRDB score 93 miRDB
hsa-miR-4784 WNT1 Predicted miRDB score 91 miRDB
hsa-miR-1251-5p TAOK1 Predicted miRDB score 96 miRDB
hsa-miR-1251-5p FMO1 Predicted miRDB score 94 miRDB
hsa-miR-1251-5p SETD3 Predicted miRDB score 91 miRDB
hsa-miR-1251-5p KPNA3 Predicted miRDB score 88 miRDB
hsa-miR-1251-5p MECP2 Predicted miRDB score 87 miRDB

Table 11.

Validated targets for the same MiRNAs in non-LGG contexts

miRNA Target mRNA Evidence type Assay PMID Cancer context
hsa-miR-4784 AHDC1 Experimental Luciferase / RIP 31,390,932 non-LGG
hsa-miR-4784 KDM5C Experimental Luciferase 33,099,720 non-LGG
hsa-miR-1251-5p TBCC Experimental Luciferase 31,278,033 Ovarian

Functional enrichment and pathway analysis

Enrichment analysis of LGG biomarkers revealed significant molecular processes underpinning tumour growth. The most significant enrichments in the Cellular Component (CC) category were the nuclear membrane (GO:0031965, p = 0.0087), the nuclear outer membrane (GO:0005640, p = 0.0111), and the Cul2-RING ubiquitin ligase complex (GO:0031462, p = 0.0132), as shown in Fig. 7a. In Biological Processes (BP), the analysis highlighted myeloid leukocyte differentiation (GO:0002573, p = 0.0011), PPAR signalling (GO:0035357, p = 0.0041), and VEGF production (GO:0010575, p = 0.0173), indicating their roles in immune escape, metabolic control, and neovascularisation (Fig. 7c).

Fig. 7.

Fig. 7

Functional enrichment analysis results. a Molecular function (MF), b Cellular component (CC), c Biological process (BP), and d KEGG pathway enrichment related to cancer and signalling pathways

The Molecular Function (MF) category (Fig. 7b) showed significant terms, including nuclear receptor binding (p = 0.0031) and opioid receptor binding (p = 0.0035), which are crucial for transcriptional regulation and tumour growth. KEGG pathway enrichment (Fig. 7d) was highly significant for Small Cell Lung Cancer (p = 0.0019), PI3K-Akt signalling (p = 0.0247), and transcriptional misregulation in Cancer (p = 0.0073), pathways that contribute to glioma growth and movement. The prominence of PI3K-Akt signalling, a key feature of glioma biology, further supports the validity of our method. Additionally, the detection of other cancer-related pathways, such as Small Cell Lung Cancer, is not surprising in multi-omics studies, as shared mechanisms and processes across cancers are often observed. These results demonstrate the reliability of known glioma-related pathways and our model’s ability to reveal broader cancer connections.

Comparison with previous studies

Recently, deep learning has become a major part of cancer computing. It has brought significant change and progress, enabling new ways to identify unique types of LGG. Table 12 presents an overview of the recent models developed between 2020 and 2024, detailing their data usage, modelling approaches, and the factors that contribute to the superiority of some models over others. Among these models, BioGAT-LGG achieves the highest score (98.03%) because it employs a novel data analysis approach that others have not yet used. It is superior because it can analyse data differences flexibly rather than relying on predefined criteria. The use of LASSO helps remove less informative features, making the model more accurate but more general. These features are also utilised in a GATv2 model, which helps illustrate the influence of other data features or nodes, ensuring that the model achieves optimal accuracy for the given situation.

Table 12.

Comparative summary of recent deep learning models for LGG subtype classification

Study Data type(s) Method Accuracy Year Cancer type Graph-based / omics integration
BioGAT-LGG (This Study) mRNA + miRNA + DNA meth. Correlation-based GATv2 98.03% 2025 LGG Yes / Joint
DeepAutoGlioma [3] mRNA + DNA meth. Autoencoder + CNN 98.03% 2023 LGG No / Joint
MOGONET [43] mRNA + miRNA + DNA meth. Multi-view GCN - 2021 Multiple Cancers Yes / Joint
i-Modern [44] mRNA + miRNA + DNA meth. + CNV + mutation Autoencoder + FC layers (Unsupervised) - 2022 Glioma No / Joint
DeepMoIC [45] mRNA + miRNA + DNA meth. Autoencoder + Deep GCN 73.2% (Grade II vs. III) 2024 LGG Yes / Joint
DeepGlioma [46] SRH histology images CNN + Transformer 91.5% 2023 LGG No / Single-Modality

Unlike models such as DeepAutoGlioma [28], which uses autoencoders with CNNs for mRNA- and methylation-based classes, or DeepGlioma [29], which uses a CNN-transformer pipeline for SRH histology images, BioGAT-LGG sits in the middle by combining cross-omic dependencies at both the feature and object levels.

Models such as MOGONET [30] and DeepMoIC [32] also used graph convolutional ideas but lacked the complex attention mechanism and feature-level fine-tuning seen in BioGAT-LGG. Crucially, BioGAT-LGG’s attention weights provide biological interpretability for potential biomarker discovery. In the future, that skill can be used to find the correct drug targets. The fact that we can tie graph-based learning to biologically underpinned sparseness and an omics blend shows an advance in how we build rule-based and translational models of glioma. Therefore, BioGAT-LGG is not just a new way to do rules but can also be used for value-based precision care, with an eye to both identifying early in the life of a subtype and identifying the key molecular targets for LGG.

Discussion

The integration of multi-omics data into machine learning models has revolutionised the classification and prognosis of LGG. The proposed BioGAT-LGG model, which employs a correlation-driven GATv2, demonstrates superior performance in subtype classification compared with earlier approaches. In the following discussion, we summarise the current advances of BioGAT-LGG in the context of shifts in methodology, performance, and biological interpretability relative to prior state-of-the-art models (i.e., MOGONET [30], DeepAutoGlioma [28], and DeepMoIC (Wu et al., 2024)). We also discuss the implications of biomarker identification and functional enrichment for improving LGG diagnostics.

BioGAT-LGG utilises dynamic attention mechanisms in GATv2 to model interactions among mRNA, miRNA, and DNA methylation features. Compared with classical concatenation-based methods (e.g., feature stacking) or ensemble models, BioGAT-LGG’s graph structure weights feature correlations adaptively, thereby enhancing biological interpretability. Compared with MOGONET, which learns omics-specific representations with Graph Convolutional Networks (GCNS), BioGAT-LGG retains the fine-grained attention mechanism of GATv2 [14]. MOGONET integrates cross-omics correlations using a View Correlation Discovery Network (VCDN). In contrast, BioGAT-LGG’s attention mechanism directly learns feature importance without access to the auxiliary network, as it is intended to be machine–learning–ready to reduce computational burden [8]. BioGAT-LGG uses LASSO regression for feature selection, which culls prognostically informative biomarkers (such as CCND2, PTGS2, and TIMP4). In this way, our approach differs from DeepAutoGlioma, which uses autoencoders for dimensionality reduction and imposes no explicit feature sparsity constraints [33]. Likewise, i-Modern employs unsupervised autoencoders (AEs) and does not include the supervised feature selection method used by BioGAT-LGG, which we believe can lead to less discriminative representations [8]. LASSO-GATv2 combines LASSO with GATv2, ensuring that only the most predictive features drive the model, thereby improving generalisability.

One distinctive feature of BioGAT-LGG is its ability to map attention weights onto biological molecular interactions to infer subtype-specific interaction processes. For instance, the model recognises MT1G and PRAME as representative biomarkers, reinforcing previous studies associating those genes with glioma progression [7]. By contrast, DeepGlioma (Hollon et al., 2023) achieves high accuracy (91.5% compared with molecular classification) as a molecular classifier built on convolutional neural networks using histology images. However, it lacks non-molecular interpretability. DeepMoIC also uses a deep GCN without explicit attention-based feature importance modelling, which hinders clinical translatability. According to comparative results, BioGAT-LGG achieved an average accuracy of 98.03% for LGG subtype classification, exceeding DeepMoIC (73.2%) and matching DeepAutoGlioma (98.03%). This distinction is also reflected in the ROC-AUC values for the model (0.98–1.00) between the IDH-mutant and IDH-wildtype subtypes. By contrast, MOGONET demonstrates competitive performance in multi-cancer classification but lacks specific accuracy numbers for LGG [9, 34]. The pairwise stability of BioGAT-LGG is more consistent across the 10-fold cross-validation than that of i-Modern, which has not been rigorously validated on independent datasets.

Heterogeneity across omics data (e.g., mRNA, miRNA, and methylation) is a fundamental challenge in LGG classification. This is also evident in the use of correlation-driven attention in BioGAT-LGG, whereas MOGONET applies similarity networks, which are insufficient for nonlinear dependencies [14]. DeepMoIC also struggles with grade II vs. III classification (73. 2% accuracy), potentially due to undermodelling of cross-omics interactions [34]. Notably, multi-omics integration visualisation is a new feature in BioGAT-LGG not available in other existing models, as shown in the hierarchical clustering heatmaps of several expression types. BioGAT-LGG uncovers high- confidence biomarkers (e.g., CCND2, PTGS2, TIMP4) consistent with established glioma pathways (e.g., PI3K- Akt signalling, VEGF production) [8]. These results support previous research [35], which found that CCND2 is associated with cell-cycle dysregulation in glioma [36]. Moreover, functional enrichment analysis of the model demonstrates involvement of nuclear receptor binding and transcriptional misregulation, indicating their roles in LGG progression—previous research, such as that by Babar et al. And Munquad et al., utilising two glioma cohorts, found that RND 3 expression enhanced survival prediction [28], confirming the importance of this BioGAT-LGG feature. By contrast, GA-based methods [37], driven, e.g., by coexpression networks, are less data-driven and more specific to biological interpretation. The LASSO- Gatv2 pipeline also ensures that biomarkers are selected in a data-driven yet biologically interpretable manner [17]. While BioGAT-LGG highlights miRNAs and mRNAs with robust prognostic signals, curated experimental MTIs linking the exact pairs prioritised here were not identified in current repositories. We thus report predicted links as hypotheses-generating evidence and explicitly encourage prospective validation (e.g., luciferase reporter assays or CLIP-seq) to confirm the proposed regulatory axes.

BioGAT-LGG represents a substantial advance in LGG classification, integrating multi-omics data, employing interpretable attention mechanisms, and identifying biomarkers. It outperforms earlier models, such as MOGONET and DeepAutoGlioma, in accuracy, stability, and biological relevance. Nevertheless, external validation in independent cohorts (e.g., CGGA and GEO datasets) is required to confirm clinical utility. We aim to align computational models closely with precision oncology and potentially integrate temporal omics to model LGG progression.

Limitations

Several limitations in our study should be noted. First, the sample size was restricted to the publicly available TCGA-LGG dataset, which may not fully capture the heterogeneity present in larger clinical populations. Second, no external validation on independent datasets, such as CGGA or local clinical cohorts, was performed, which would have provided a stronger demonstration of generalisability. Third, although the correlation-based graph captured feature–feature interactions across omics layers, it lacked additional biological priors, such as protein–protein interaction (PPI) networks or pathway-driven graphs, that could enhance the analysis. Fourth, the model’s interpretability was derived from attention weights, which are informative but do not necessarily reflect direct biological mechanisms and thus ought to be subjected to additional experimental validation. Finally, LASSO was applied for dimensionality reduction, and SMOTE was used to balance class imbalance to address high dimensionality and class imbalance; however, these preprocessing steps may introduce selection biases that can influence downstream results.

Conclusion

BioGAT-LGG presents a correlation-based GATv2 that provides deeper insights into LGG subtypes classification and identifying biomarkers from integrated mRNA, miRNA, and DNA methylation data. In the proposed model, predefined PPI networks are replaced with data-driven correlation graphs, and features are selected using LASSO to capture condition-specific molecular patterns. In 10-fold cross-validation trials, BioGAT-LGG achieved an overall accuracy of 98.03% (98.12% precision, 97.74% recall, 97.87% F1-score). Functional enrichment and KEGG analyses confirmed that the identified cancer biomarkers are biologically relevant to significant cancer pathways. Compared with existing deep learning models, BioGAT-LGG achieved improved performance, and attention-guided models also provided interpretability by learning across omics layers. Future directions would involve extending the omics types considered, applying the framework to other cancers, and exploring its potential clinical utility in precision oncology.

Acknowledgements

The authors would like to acknowledge all contributors who provided invaluable support in data collection, analysis, and manuscript preparation.

Institutional review board

Not applicable.

Author contributions

EMH and NYA designed the study, implemented the BioGAT-LGG model, performed data preprocessing, conducted statistical analyses and functional enrichment analysis, and drafted the manuscript. WAA contributed to the writing of the manuscript. MM developed the computational code and assisted with data processing. MKE supervised the study, provided guidance on methodological design, contributed to model implementation, and refined the computational analysis. AY, AMM, MEM, and MAM participated in the study design and contributed to manuscript drafting. All authors reviewed and approved the final manuscript.

Funding

This research was funded by the Deanship of Graduate Studies and Scientific Research at Jouf University under grant no. DGSSR-2024-02-02191.

Data availability

All relevant code and processed data supporting the findings of this study are publicly available on Zenodo to ensure transparency and reproducibility:- Code repository: https://doi.org/10.5281/zenodo.17842483- Dataset repository: [https://doi.org/10.5281/zenodo.15761467](https:/doi.org/10.5281/zenodo.15761467) These repositories include the full implementation of the BioGAT-LGG model, instructions for running the code, and the processed multi-omics data used in this study.

Declarations

Ethics approval and consent to participate

This study did not involve experiments on human participants or animals. All data were obtained from publicly available repositories (TCGA), and therefore ethical approval was not required. Not applicable. No human subjects were directly involved, and all datasets were fully anonymised and publicly accessible.

Consent for publication

Not applicable. No individual or identifiable human data are included, and therefore no consent for publication was required.

Informed consent

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Murtada K. Elbashir, Email: mkelfaki@ju.edu.sa

Abdulrahman Alyami, Email: am.yami@ju.edu.sa.

References

  • 1.Louis DN, et al. The 2021 WHO classification of tumors of the central nervous system: a summary. Neurooncology. 2021;23(8):1231–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Badkas A, De Landtsheer S, Sauter T. Construction and contextualization approaches for protein-protein interaction networks. Comput Struct Biotechnol J. 2022;20:3280–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Munquad S, Das AB. DeepAutoGlioma: a deep learning autoencoder-based multi-omics data integration and classification tools for glioma subtyping. BioData Min. 2023;16(1):32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Wu J, et al. DeepMoIC: multi-omics data integration via deep graph convolutional networks for cancer subtype classification. BMC Genomics. 2024;25(1):1209. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Ryan B, Marioni RE, Simpson TI. Multi-Omic graph diagnosis (MOGDx): a data integration tool to perform classification tasks for heterogeneous diseases. Bioinformatics. 2024;40(9):btae523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Singh A, et al. DIABLO: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinformatics. 2019;35(17):3055–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Gao F, et al. DeepCC: a novel deep learning-based framework for cancer molecular subtype classification. Oncogenesis. 2019;8(9):44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Sun Q, et al. SADLN: Self-attention based deep learning network of integrating multi-omics data for cancer subtype recognition. Front Genet. 2023;13:1032768. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Lin Y, et al. Classifying breast cancer subtypes using deep neural networks based on multi-omics data. Genes. 2020;11(8):888. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Ahmed NY, et al. An efficient deep learning approach for DNA-binding proteins classification from primary sequences. Int J Comput Intell Syst. 2024;17(1):88. [Google Scholar]
  • 11.Alsanousi WA, et al. A novel deep learning-assisted hybrid network for plasmodium falciparum parasite mitochondrial proteins classification. PLoS ONE. 2022;17(10):e0275195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Putro IH, Ahmad T. Feature Selection Using Pearson Correlation with Lasso Regression for Intrusion Detection System. in 2024 12th International Symposium on Digital Forensics and Security (ISDFS). 2024. IEEE.
  • 13.Wang S, et al. Diabetes risk analysis based on machine learning LASSO regression model. J Theory Pract Eng Sci. 2024;4(01):58–64. [Google Scholar]
  • 14.Kanehisa M, et al. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021;49(D1):D545–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Colaprico A, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44(8):e71–71. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Jiang G, et al. A comprehensive workflow for optimizing RNA-seq data analysis. BMC Genomics. 2024;25(1):631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Ritchie ME, et al. Limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43(7):e47–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Law CW, et al. Voom: precision weights unlock linear model analysis tools for RNA-seq read counts. Genome Biol. 2014;15:1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Babar S, et al. Informatics approaches to identify potential biomarkers for breast cancer therapeutic targets using microarray data. Journal of Computational Biophysics and Chemistry; 2025.
  • 20.Tibshirani R. Regression shrinkage and selection via the Lasso. J Royal Stat Soc Ser B: Stat Methodol. 1996;58(1):267–88. [Google Scholar]
  • 21.Qi L, et al. Multi-omics data fusion for cancer molecular subtyping using sparse canonical correlation analysis. Front Genet. 2021;12:607817. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Pidsley R, et al. A data-driven approach to preprocessing illumina 450K methylation array data. BMC Genomics. 2013;14:1–10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Brody S, Alon U, Yahav E. How attentive are graph attention networks? arXiv preprint arXiv:2105.14491, 2021.
  • 24.Baul S, et al. OmicsGAT: graph attention network for cancer subtype analyses. Int J Mol Sci. 2022;23(18):10220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Chawla NV, et al. SMOTE: synthetic minority over-sampling technique. J Artif Intell Res. 2002;16:321–57. [Google Scholar]
  • 26.Fernández A, et al. SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary. J Artif Intell Res. 2018;61:863–905. [Google Scholar]
  • 27.Szklarczyk D, et al. The STRING database in 2021: customizable protein–protein networks, and functional characterization of user-uploaded gene/measurement sets. Nucleic Acids Res. 2021;49(D1):D605–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Greene CS, et al. Understanding multicellular function and disease with human tissue-specific networks. Nat Genet. 2015;47(6):569–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Valdeolivas A, et al. Random walk with restart on multiplex and heterogeneous biological networks. Bioinformatics. 2019;35(3):497–505. [DOI] [PubMed] [Google Scholar]
  • 30.Zitnik M, Leskovec J. Predicting multicellular function through multi-layer tissue networks. Bioinformatics. 2017;33(14):i190–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Larriba E, et al. Identification of new targets for glioblastoma therapy based on a DNA expression microarray. Comput Biol Med. 2024;179:108833. [DOI] [PubMed] [Google Scholar]
  • 32.Zhang T, Wang G. Inhibition of RXRA-mediated PLA2G2A improves delirium in COPD mice by regulating Endoplasmic reticulum stress pathway and inhibiting cell apoptosis. Cell Mol Biol. 2024;70(1):226–32. [DOI] [PubMed] [Google Scholar]
  • 33.Wu L et al. Long non-coding RNA HOXA-AS2 enhances the malignant biological behaviors in glioma by epigenetically regulating RND3 expression. OncoTargets Therapy, 2019: pp. 9407–19. [DOI] [PMC free article] [PubMed]
  • 34.Iijima T, et al. Distinct expression of C1q-like family mRNAs in mouse brain and biochemical characterization of their encoded proteins. Eur J Neurosci. 2010;31(9):1606–15. [DOI] [PubMed] [Google Scholar]
  • 35.Vageli DP, et al. Hypoxia-inducible factor 1alpha and vascular endothelial growth factor in glioblastoma multiforme: a systematic review going beyond pathologic implications. Oncol Res. 2024;32(8):1239. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Solga R, et al. CRN2 binds to TIMP4 and MMP14 and promotes perivascular invasion of glioblastoma cells. Eur J Cell Biol. 2019;98(5–8):151046. [DOI] [PubMed] [Google Scholar]
  • 37.Xue L-p, et al. Overexpression of tripartite motif-containing 48 (TRIM48) inhibits growth of human glioblastoma cells by suppressing extracellular signal regulated kinase 1/2 (ERK1/2) pathway. Med Sci Monitor: Int Med J Experimental Clin Res. 2019;25:8422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Lin R, Yao C, Ren D. Association between genetic polymorphisms of PTGS2 and glioma in a Chinese population. Genet Mol Res. 2015;14(2):3142–8. [DOI] [PubMed] [Google Scholar]
  • 39.Gross I, et al. Systematic expression analysis of plasticity-related genes in mouse brain development brings PRG4 into play. Dev Dyn. 2022;251(4):714–28. [DOI] [PubMed] [Google Scholar]
  • 40.Chen W, et al. Prognostic prediction model for glioblastoma: A ferroptosis-related gene prediction model and independent external validation. J Clin Med. 2023;12(4):1341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Le M-K, et al. Molecular and clinicopathological implications of PRAME expression in adult glioma. PLoS ONE. 2023;18(10):e0290542. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Zhang H, et al. Highly expressed LncRNA CCND2-AS1 promotes glioma cell proliferation through Wnt/β-catenin signaling. Biochem Biophys Res Commun. 2017;482(4):1219–25. [DOI] [PubMed] [Google Scholar]
  • 43.Wang T, et al. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat Commun. 2021;12(1):3445. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Pan X, et al. i-Modern: integrated multi-omics network model identifies potential therapeutic targets in glioma by deep learning with interpretability. Comput Struct Biotechnol J. 2022;20:3511–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Wu J, et al. DeepMoIC: multi-omics data integration via deep graph convolutional networks for cancer subtype classification. BMC Genomics. 2024;25(1):1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Hollon T, et al. Artificial-intelligence-based molecular classification of diffuse gliomas using rapid, label-free optical imaging. Nat Med. 2023;29(4):828–32. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

All relevant code and processed data supporting the findings of this study are publicly available on Zenodo to ensure transparency and reproducibility:- Code repository: https://doi.org/10.5281/zenodo.17842483- Dataset repository: [https://doi.org/10.5281/zenodo.15761467](https:/doi.org/10.5281/zenodo.15761467) These repositories include the full implementation of the BioGAT-LGG model, instructions for running the code, and the processed multi-omics data used in this study.


Articles from Discover Oncology are provided here courtesy of Springer

RESOURCES