Abstract
Background
Prognostic prediction following gastric cancer surgery plays a pivotal role in postoperative management, helping to optimize therapeutic strategies and improve patient survival. Standard clinicopathological indicators, including tumor differentiation and lymph node metastasis, continue to serve as the basis for outcome evaluation; however, they do not adequately represent the host’s systemic inflammatory response and immunonutritional status, both of which significantly affect tumor progression and postoperative recovery. Systemic inflammatory markers, such as the Neutrophil-to-Lymphocyte Ratio (NLR) and Platelet-to-Lymphocyte Ratio (PLR), have emerged as reliable, noninvasive prognostic indicators. However, the complex and nonlinear interactions among inflammatory, clinical, and demographic variables pose a limitation for traditional statistical methods.
Methods
This study proposes a novel deep learning framework that integrates three major components: Gradient-Boosted Decision Tree, Tree-Driven Encoder (TDE), and one-dimensional Convolutional Neural Network (1D-CNN) for postoperative prognostic prediction in gastric cancer. The GBDT module captures intricate dependencies among clinical and inflammatory variables, the TDE transforms tree-based structures into unified binary embeddings, and the 1D-CNN component learns high-level feature representations from these embeddings to predict postoperative prognosis. The model’s performance was evaluated using cross-validation and compared with various traditional machine learning algorithms and advanced deep learning architectures for tabular data.
Results
Experimental findings demonstrate that the proposed hybrid framework consistently outperforms both traditional and general deep learning models in predicting postoperative prognosis. By combining tree-based feature structuring with deep representation learning, the model effectively captures nonlinear and hierarchical relationships among systemic inflammatory markers and clinicopathological features. This approach achieves high predictive accuracy, robustness, and generalization capability, particularly in identifying high-risk patients characterized by elevated inflammatory activity. Moreover, the model exhibited stable performance across multiple random seeds and data partitions, confirming its reproducibility and reliability under different experimental conditions.
Conclusions
This study presents a data-driven and interpretable deep learning framework for postoperative prognostic prediction in gastric cancer. By integrating the strengths of gradient-boosted tree modeling and deep neural representation learning, the proposed model provides a more comprehensive understanding of the interplay among inflammation, nutrition, and tumor biology, supporting personalized treatment planning and evidence-based clinical decision-making. Future research will focus on external validation using independent cohorts, real-time clinical application, and enhancing model explainability to facilitate clinical adoption.
Keywords: Gastric cancer prognosis, Postoperative outcomes, Inflammatory biomarkers, Deep learning, Hybrid models
Introduction
Gastric cancer remains a major global health challenge. According to the latest GLOBOCAN 2022 estimates [1], there were over 968,000 new cases of stomach cancer and nearly 660,000 deaths worldwide in 2022, ranking it as the fifth most common malignancy and the fifth leading cause of cancer-related mortality globally. Prognostic prediction following gastric cancer surgery represents a critical component of postoperative management, aiming to optimize patient survival and personalize therapeutic strategies. Although conventional clinicopathological indicators-such as the Tumor-Node-Metastasis (TNM) stage, tumor differentiation, and lymph node metastasis-remain fundamental for outcome assessment, they fail to fully capture the host’s systemic inflammatory response and immunonutritional status, both of which profoundly influence tumor progression and postoperative recovery. In this context, systemic inflammatory markers, particularly the Neutrophil-to-Lymphocyte Ratio (NLR) and Platelet-to-Lymphocyte Ratio (PLR), have emerged as reliable, noninvasive indicators reflecting the balance between pro-tumor inflammation and anti-tumor immunity. Elevated levels of these markers have been shown to correlate strongly with TNM stage, tumor invasion depth, and postoperative morbidity, thereby serving as surrogate indicators of systemic immune dysfunction and tumor aggressiveness. Yu et al. [2] demonstrated that indices such as NLR, PLR, and the C-reactive protein to albumin ratio are significantly associated with lymph node metastasis and tumor differentiation. Mechanistically, neutrophilia may promote tumor proliferation and angiogenesis through vascular endothelial growth factor secretion and matrix metalloproteinase activation, while lymphopenia indicates impaired cellular immunity, leading to diminished tumor surveillance and poor prognosis. Recent studies have further emphasized the value of multifactorial prognostic models that integrate inflammatory, nutritional, and tumor-related parameters. Liu et al. [3] reported that elderly patients with elevated preoperative NLR exhibit higher risks of severe postoperative complications and delayed recovery, highlighting the need for individualized preoperative risk stratification based on inflammatory indices. Moreover, systemic inflammation has been implicated not only in long-term prognosis but also in the development of postoperative complications. DSouza et al. [4] revealed that heightened inflammatory activity contributes to a postoperative complication rate as high as 34%, including infectious and hepatic dysfunctions, which adversely affect quality of life and prolong hospitalization. These findings suggest that inflammatory markers could serve as valuable predictive tools for identifying high-risk patients requiring enhanced perioperative monitoring.
Integrating inflammatory markers with clinicopathological and imaging data through deep learning-based prognostic frameworks offers a powerful, data-driven approach to postoperative outcome prediction. Deep learning models can effectively capture complex, nonlinear interactions between systemic inflammation, nutritional status, and tumor biology-providing robust risk stratification beyond conventional statistical models. The development of inflammatory marker-driven deep learning models therefore holds significant promise for refining personalized treatment planning, optimizing surgical decision-making, and ultimately improving survival outcomes for patients undergoing gastric cancer surgery.
The application of machine learning (ML) and deep learning (DL) in cancer prognosis prediction has gained substantial momentum in recent years. Traditional prognostic frameworks, such as the TNM staging system, have long been regarded as the clinical gold standard. However, these models are inherently limited by their reliance on anatomical parameters and their inability to account for the biological heterogeneity and dynamic tumor host interactions that drive disease progression [5]. Consequently, patients classified within the same TNM stage may exhibit markedly different outcomes, emphasizing the necessity for biologically informed and data-driven prognostic tools. Recent advances in machine learning and deep learning have begun to address these limitations by capturing complex nonlinear patterns across heterogeneous data modalities. For instance, Diao et al. leveraged a deep learning framework to extract interpretable pathological features from digitized histology images and correlate them with molecular phenotypes and survival outcomes [6]. Such approaches demonstrate the potential of artificial intelligence (AI) to integrate high-dimensional, multimodal datasets that traditional models often fail to utilize. ML algorithms, with their ability to model intricate dependencies among clinical, molecular, and imaging variables, consistently outperform linear statistical methods and conventional staging systems [7]. Kuwayama et al. [5] demonstrated that ML models using routine clinical variables significantly improved preoperative diagnostic accuracy and postoperative risk stratification, surpassing TNM-based predictions.
Similarly, Wang et al. [8] highlighted that integrating nutritional indicators and tumor-specific characteristics further enhanced predictive performance beyond conventional staging methods. Moreover, a particularly promising frontier involves the integration of systemic inflammatory markers within ML and DL frameworks. Inflammatory indices such as the NLR, PLR, and C-reactive protein-to-albumin ratio have been identified as strong prognostic biomarkers reflecting the host immune response to tumor progression [9]. Studies have demonstrated that combining these markers with ML-based predictive models enhances accuracy and provides a holistic assessment of the patient’s immune-nutritional status, which is crucial for personalized postoperative management. Despite these advancements, challenges persist concerning data quality, model interpretability, and cross-cohort generalization [7]. Deep learning models are powerful but often viewed as “black boxes”, which reduces clinicians’ confidence in using them for real-world decision support [10]. Hence, developing transparent, biologically grounded DL frameworks remains a critical research direction in computational oncology.
Recent advances in computational oncology have further expanded the role of artificial intelligence and mathematical modeling in tumor progression prediction. Several machine learning-based studies have demonstrated that integrating multimodal clinical and molecular data can substantially improve prognostic accuracy and individualized risk assessment in cancer patients. In parallel, mathematical and systems biology models have been increasingly applied to characterize tumor growth dynamics, treatment response, and disease evolution through quantitative simulation frameworks [11]. Emerging digital twin technologies additionally provide a promising paradigm for creating patient-specific virtual representations capable of continuously updating prognostic predictions and supporting adaptive therapeutic decision-making [12]. These developments highlight the growing convergence between deep learning, mechanistic modeling, and precision oncology, providing an important contextual foundation for the present inflammatory marker-driven prognostic framework.
The present study aims to develop a deep learning model driven by inflammatory markers to predict postoperative prognosis in gastric cancer patients. Conventional prognostic systems such as TNM classification overlook several biologically relevant features-including systemic inflammation and nutritional status-that substantially affect clinical outcomes [2]. By leveraging the representational power of DL, this study seeks to enhance the precision and reliability of prognostic predictions, offering clinicians a more comprehensive decision-support tool. This research also carries significant implications for precision oncology. Integrating inflammatory biomarkers such as NLR and CAR within a deep learning framework provides a biologically interpretable and clinically actionable model that complements existing classification systems [13]. As demonstrated by Huang et al., incorporating diverse clinical and molecular parameters through DL can identify optimal prognostic indicators that traditional statistical analyses might overlook [14]. Beyond prediction, such models pave the way for adaptive postoperative management, enabling continuous monitoring and individualized adjustment of therapeutic regimens based on patient-specific trajectories [15]. Ultimately, the proposed inflammatory marker-driven deep learning model aspires to bridge the gap between computational innovation and clinical applicability, thereby improving survival outcomes and postoperative quality of life in gastric cancer patients. The main contributions of this study are summarized as follows:
A hybrid deep learning framework that combines GBDT-based feature structuring with a Tree-Driven Encoder and a 1D CNN for gastric cancer prognosis.
Systematic integration of inflammatory biomarkers with clinicopathological features through tree-driven binary embeddings, enabling deep learning models to capture nonlinear interactions.
Literature review
Inflammatory markers and gastric cancer prognosis
The prognostic significance of systemic inflammatory markers has gained increasing attention in recent years, particularly in gastric cancer. Biomarkers such as the NLR and PLR have emerged as reliable indicators reflecting the systemic inflammatory response associated with tumor progression and host immunity. A comprehensive meta-analysis by Mellor et al. (2018) [16] demonstrated that elevated NLR levels were strongly correlated with poor overall survival and disease-free survival in gastric cancer patients, supporting its prognostic utility alongside C-reactive protein and albumin as inflammation-based indices of clinical relevance. Similarly, Lee et al. (2013) [17] highlighted that preoperative inflammatory parameters, including NLR and PLR, provide meaningful prognostic information independent of TNM staging. Their findings suggest that inflammation-based markers may serve as valuable complements to conventional pathological staging in predicting postoperative outcomes. Moreover, Han et al. (2023) [18] emphasized the role of the systemic immune-inflammation index as a composite marker capturing the complex interplay between immune and inflammatory pathways, further underscoring the prognostic implications of the inflammatory microenvironment in gastric carcinogenesis. Inflammatory markers such as C-reactive protein, fibrinogen, and apolipoprotein A1 have also been implicated in modulating tumor aggressiveness and survival outcomes. Elevated CRP and fibrinogen levels have been consistently associated with poor prognosis across multiple malignancies, including gastric cancer, while apolipoprotein A1 has been proposed to exert protective effects against tumor progression [19, 20]. Notably, Wu et al. (2019) [19] demonstrated that preoperative fibrinogen and pre-albumin levels serve as independent prognostic indicators for long-term survival, reinforcing the clinical relevance of inflammatory biomarkers in risk stratification.
Integration of machine learning in prognostic modeling
Despite the widespread use of the TNM classification system, its prognostic accuracy remains limited by its inability to account for systemic biological responses such as inflammation and immune activation. To address these limitations, recent research has incorporated machine learning and deep learning techniques to enhance predictive modeling in gastric cancer prognosis. Algorithms such as Support Vector Machines (SVM), Adaptive Boosting (AdaBoost), and Artificial Neural Networks (ANN) have been employed to analyze complex, high-dimensional datasets integrating inflammatory and clinical parameters [9]. These models outperform conventional approaches by capturing nonlinear interactions between biological and pathological features, thereby enabling individualized risk assessment. Furthermore, Shi et al. (2021) introduced nomograms incorporating systemic inflammatory markers into predictive frameworks, providing a more patient-specific estimation of survival probabilities. The incorporation of explainable AI approaches has further improved clinical interpretability. As shown by Shao et al. (2025) [21], explainable models allow clinicians to visualize the contribution of individual biomarkers to prognostic predictions, fostering greater trust and transparency in AI-assisted clinical decision-making. This advancement addresses one of the critical barriers to the adoption of AI in oncology - the need for interpretability alongside accuracy.
Clinical implications and research gaps
Emerging evidence increasingly supports the integration of inflammatory indices with machine learning-based frameworks to improve prognostic precision in gastric cancer. Studies by Qi and Zheng (2024) [22] and Xu et al. (2024) [23] demonstrated that incorporating systemic inflammation indicators significantly enhanced predictive accuracy in various malignancies, including lung and gastric cancers. In parallel, Radulescu et al. (2022) [24] reported that additional ratios such as the monocyte-to-lymphocyte ratio may provide incremental prognostic value by reflecting specific immune-inflammatory interactions. Importantly, inflammation not only influences tumor progression but also modulates therapeutic response. Tada et al. (2024) [25] observed that peripheral inflammatory markers could guide individualized treatment decisions, particularly in cases with heterogeneous response patterns to adjuvant therapy. Consequently, integrating such biological markers into AI-based predictive models holds substantial potential for refining postoperative management and optimizing treatment strategies. Despite these advancements, several research gaps remain. Most existing models rely on retrospective data with limited generalizability across populations, and the mechanistic link between systemic inflammation and molecular tumor behavior is not yet fully elucidated. Future studies should aim to develop multi-modal prognostic frameworks that integrate inflammatory markers, genomic data, and imaging features using interpretable deep learning architectures to achieve a more holistic and clinically actionable prediction of gastric cancer outcomes.
Dataset
Data description
The dataset used in this study was obtained from the EBI BioStudies (S-EPMC5373584). It comprises the clinical and pathological records of 1,330 patients diagnosed with gastric cancer who underwent gastrectomy. All patients were histologically confirmed as having Stage I-III gastric carcinoma, and tumor staging was performed according to the 7th edition of the American Joint Committee on Cancer (AJCC) TNM classification system. The patients included in this dataset were selected based on the following inclusion criteria: First, no history of neoadjuvant chemotherapy or radiotherapy before surgery. Second, availability of complete clinical, pathological, and follow-up data relevant to potential prognostic factors. Third, absence of recurrence, residual gastric cancer, or other synchronous malignancies. Finally, no acute infection or inflammatory disease within two weeks prior to surgery. After applying these criteria, only patients with complete hematological and outcome records were retained for analysis to ensure data consistency and reliability.
- Clinical and Laboratory Features: Preoperative blood samples were collected one week before surgery. The following variables were extracted for analysis:
- Demographic data: age and sex.
- Laboratory indices: complete preoperative hematological parameters.
- Tumor characteristics: tumor size, location, histological grade, and TNM stage.
- Outcome variable: overall survival, defined as death from any cause after surgery.
- Highly differentiated group: papillary and moderately differentiated gastric cancers.
- Poorly differentiated group: signet-ring cell, mucinous, and undifferentiated gastric cancers.
- Biomarker Definitions: The study included three inflammation-based biomarkers widely used in prognostic assessment of gastric cancer:
- Neutrophil-to-Lymphocyte Ratio: calculated as the absolute neutrophil count divided by the absolute lymphocyte count.
- Platelet-to-Lymphocyte Ratio: calculated as the absolute platelet count divided by the absolute lymphocyte count.
- Combined Platelet and Neutrophil-to-Lymphocyte Ratio: computed as follows: Patients with elevated platelet count
were assigned a score of 2. Those with one abnormal index were assigned 1, and those with both normal indices were assigned 0.
All patients were followed up every three months during the first two years after surgery and every six months thereafter. Survival status was recorded at each follow-up, and the primary endpoint of this study was overall mortality following gastrectomy.
Data preprocessing
The raw dataset, obtained from the BioStudies database (as described in section “Data description”), included clinical, pathological, and preoperative hematological information of patients who underwent gastrectomy for gastric cancer. A systematic preprocessing pipeline was implemented to ensure data quality and maintain reproducibility. Duplicate records were removed, and the integrity of patient identifiers was verified. Laboratory indices were checked for physiological plausibility, and values outside the normal range were inspected and corrected if necessary. Patients with missing key prognostic or survival information were excluded from further analysis. Categorical variables such as sex, tumor location, and histological grade were one-hot encoded. TNM stages were encoded as ordinal variables. Continuous variables were standardized using z-score normalization based on the mean and standard deviation computed from the training set. The final dataset contained 1,330 samples and 15 variables, including demographic, pathological, and hematological features. The target variable, Survival status, included 806 deceased patients (60.6%) and 524 alive patients (39.4%). The dataset was divided into training, validation, and test subsets in an 8:1:1 ratio with stratification based on survival status to maintain class balance across partitions as described in Table 1. This stratified splitting approach ensured consistent outcome distribution across all subsets, supporting reliable model training and evaluation.
Table 1.
Distribution of survival status across training, validation, and test sets
| Dataset | Total Samples | Alive (n, %) | Deceased (n, %) |
|---|---|---|---|
| Training | 1,064 | 419 (39.4%) | 645 (60.6%) |
| Validation | 133 | 53 (39.8%) | 80 (60.2%) |
| Test | 133 | 52 (39.1%) | 81 (60.9%) |
| Total | 1,330 | 524 (39.4%) | 806 (60.6%) |
Methodology
Overview
In this study, we propose a deep learning framework for predicting postoperative prognosis in patients with gastric cancer. The model is designed to effectively capture complex nonlinear relationships among demographic, clinical, and inflammatory biomarker features. As illustrated in Fig. 1, the proposed framework consists of three main components:
Gradient-Boosted Decision Tree (GBDT) module that models inter-feature dependencies
Tree-Driven Encoder (TDE) that transforms heterogeneous clinical data into uniform binary embeddings
1D Convolutional Neural Network (CNN1D) that is trained on these binary embeddings to predict clinical outcomes.
Fig. 1.

Overview of the proposed GBDT-TDE-CNN1D framework. Clinicopathological and inflammatory features are first fed into a Gradient-Boosted decision tree (GBDT) module to model nonlinear inter-feature dependencies. The tree-driven Encoder (TDE) then converts the resulting decision paths into uniform sparse binary embeddings, which are processed by a 1D convolutional neural network (CNN1D) with successive convolution and pooling layers to capture local feature interactions and produce the final postoperative survival prediction
Gradient-boosted decision tree for feature structuring
To effectively capture the nonlinear interactions among heterogeneous clinical variables, we draw inspiration from prior studies [26] on leveraging tabular data representations. We first employ the GBDT algorithm, a widely used ensemble method that constructs a strong predictive model by iteratively adding weak learners, typically shallow decision trees. This approach minimizes information loss and eliminates the need for complex preprocessing, making it particularly suitable for medical datasets that contain both numerical and categorical variables. Formally, a GBDT model can be represented as an ensemble of k decision trees:
![]() |
1 |
where each tree
contributes to the final prediction through an additive boosting process. Empirical studies have shown that the internal structure of these trees retains substantial information about the underlying data distribution, which can be extracted and utilized for subsequent representation learning.
Tree-driven encoder for knowledge extraction
After training the ensemble of GBDT trees, we use a TDE to distill structural knowledge from the decision trees into a unified binary feature space. The encoder transforms each input sample
into a consistent binary vector representation
by traversing all trees in the ensemble and encoding the activation of internal nodes. Let each decision tree be denoted as T = (V, E, μ), where V is the set of nodes,
is the set of edges, and
maps a sample to its corresponding child nodes.
For each internal node
, we define a Boolean function
as follows:
![]() |
2 |
By traversing all internal nodes of the tree in a breadth-first order, we obtain a binary feature vector:
![]() |
3 |
where |T| denotes the number of internal nodes in tree T. The final encoded representation of
across the entire ensemble of trees is then defined as:
![]() |
4 |
This transformation ensures that all clinical and inflammatory features are represented in a unified binary space, enabling the deep neural network to learn meaningful prognostic patterns from homogeneous inputs.
The GBDT module was implemented due to its computational efficiency and native support for tree structure extraction, which is essential for the Tree-Driven Encoder. The complete set of hyperparameters for both the GBDT module and the Tree-Driven Encoder is reported in Table 2.
Table 2.
Hyperparameter configuration of the GBDT module and Tree-Driven Encoder
| Component | Hyperparameter | Value |
|---|---|---|
| GBDT (LightGBM) | Number of trees (k) | 100 |
| Maximum tree depth | 6 | |
| Learning rate | 0.01 | |
| Feature fraction (colsample) | 0.8 | |
| Bagging fraction (subsample) | 0.8 | |
L1 regularization ( ) |
0.1 | |
L2 regularization ( ) |
1.0 | |
| Tree-Driven Encoder | Encoding scheme | Internal-node activation |
| Traversal order | Breadth-first | |
| Binary embedding dimension | Total internal nodes |
With k = 100 trees of maximum depth 6, each tree contains up to
internal nodes, yielding a maximum embedding dimension of
binary features. In practice, due to early stopping and the minimum-samples-per-leaf constraint, the actual embedding dimension was approximately 4,800. The GBDT was trained with early stopping based on validation log-loss (patience of 20 rounds). Although the resulting binary vectors exhibit high sparsity (approximately 90% zeros, as each sample activates only one path per tree), the dense representation was retained because the CNN1D convolutional filters inherently learn to identify informative activation patterns within local windows.
1D convolutional neural network for prognostic prediction
The binary feature vectors
generated by the Tree-Driven Encoder are subsequently fed into a one-dimensional convolutional neural network to learn high-level feature representations and predict postoperative prognosis. Unlike fully connected deep neural networks (DNNs), the CNN1D leverages convolutional operations to capture local dependencies among adjacent binary feature groups, reflecting latent correlations between encoded clinical and inflammatory features. The proposed CNN1D architecture consists of the following key components:
1D Convolutional Layers: Three consecutive convolutional layers are employed with kernel sizes of 7, 5, and 3, and with 128, 64, and 32 output channels, respectively. These layers extract local feature patterns within the binary embedding space.
Batch Normalization: A Batch Normalization layer follows each CNN1D layer to stabilize training and improve generalization.
Activation Function: The ReLU activation function is applied to introduce non-linearity and mitigate the vanishing gradient problem.
Global Average Pooling (GAP): A GAP layer compresses the temporal dimension (sequence length) to reduce the number of parameters while preserving global contextual information.
Dropout: A Dropout layer with a rate of p = 0.3 is included to prevent overfitting.
-
Fully Connected Layer: The final dense layer maps the aggregated features to the prognostic probability
as: 
5 where
denotes the predicted survival probability or risk score, and
represents the learnable parameters of the network.
This convolutional design enables the ours framework to effectively model structured relationships among encoded clinical features while maintaining computational efficiency and robustness against overfitting.
Experiment
Training environment and resources
All experiments were conducted on a workstation equipped with an Intel Xeon Gold 6226 R CPU (2.9 GHz), 32 GB RAM, and an NVIDIA RTX 3090 GPU (24 GB VRAM). The implementation was developed using Python 3.14, leveraging the PyTorch and scikit-learn libraries for deep learning and traditional machine learning components, respectively. For optimization, the Adam optimizer was used with an initial learning rate of
, batch size of 64, and early stopping after 25 epochs without improvement on the validation loss. The learning rate was reduced by a factor of 0.1 upon plateau detection. Dropout layers with a rate of 0.3 were applied to prevent overfitting.
Evaluation metrics
To comprehensively evaluate the predictive performance of the proposed model, five standard metrics were adopted: Accuracy (ACC), Precision, Recall, Area Under the Receiver Operating Characteristic Curve (AUC), and Area Under the Precision-Recall Curve (PRAUC). All metrics were calculated on the independent test set.
- ACC: measures the overall proportion of correctly classified samples.

6 - Precision: quantifies the proportion of correctly predicted positive cases among all predicted positives.

7 - Recall: measures the proportion of true positive cases that were correctly identified by the model.

8 - AUC: represents the probability that the model ranks a randomly chosen positive case higher than a randomly chosen negative case.

9 - PRAUC: evaluates the trade-off between precision and recall, particularly useful for imbalanced clinical datasets.
where TP, TN, FP, and FN denote the numbers of true positives, true negatives, false positives, and false negatives, respectively. TPR and FPR represent the true positive rate and false positive rate. All metrics were computed on the independent test set, and mean values across repeated experiments were reported for fair comparison.
10
Model evaluation
Baseline models
To establish a comparative benchmark for the proposed framework, we implemented a set of conventional machine learning models that are widely applied to clinical and biomedical datasets. These baseline models are designed to capture nonlinear dependencies and complex interactions among inflammatory and biochemical markers while maintaining interpretability.
Random Forest (RF) [27] - an ensemble learning algorithm that constructs a large number of independent decision trees using bootstrap sampling and random feature selection at each split. The final prediction is determined by majority voting among the trees. This mechanism reduces the risk of overfitting, enhances generalization, and captures feature interactions effectively. RF has been widely used in clinical prognosis due to its robustness to noise and capability to handle heterogeneous feature types such as categorical and continuous laboratory values.
Support Vector Machine (SVM) [28] - a margin-based classifier that seeks the optimal hyperplane separating data points of different prognostic classes. The radial basis function (RBF) kernel was employed to model nonlinear relationships among inflammatory markers. SVM is particularly effective in small- to medium-sized medical datasets, providing strong generalization even when feature dimensionality exceeds the number of samples.
k-Nearest Neighbors (kNN) [29] - a non-parametric algorithm that classifies a sample based on the majority label of its k nearest points in the feature space. kNN serves as an intuitive distance-based baseline, highlighting how feature space geometry influences prognostic prediction performance without the need for complex parameter tuning.
Decision Tree (DT) [30] - a single-tree model that recursively partitions the feature space into homogeneous subsets based on impurity minimization (Gini index). Although prone to overfitting, DT provides transparent decision paths that facilitate clinical interpretability, allowing visualization of threshold-based decisions on inflammatory indicators.
Extreme Gradient Boosting (XGB) [31] - an advanced boosting algorithm that sequentially builds an ensemble of weak decision trees to minimize prediction errors through gradient optimization. XGB incorporates regularization terms to prevent overfitting and supports parallel computation for scalability. It has become a standard choice in structured biomedical data analysis and is included here to represent a strong machine learning baseline.
Each model was trained using the same input features and cross-validation protocol. Hyperparameters such as tree depth, learning rate, and number of estimators were optimized using grid search on the validation set to ensure fair comparison. This collection of baselines provides a comprehensive view of how traditional algorithms perform in capturing the prognostic patterns driven by postoperative inflammatory markers.
The hyperparameter configurations for all traditional machine learning baselines are reported in Table 3. All models were implemented using scikit-learn and tuned via grid search with 5-fold cross-validation on the training set, selecting the configuration that maximized validation AUC.
Table 3.
Hyperparameter configurations of traditional machine learning baselines
| Model | Hyperparameter | Value |
|---|---|---|
| Decision Tree | Max depth | 8 |
| Min samples split | 10 | |
| Min samples leaf | 5 | |
| Criterion | Gini | |
| k-Nearest Neighbors | k (neighbors) | 11 |
| Distance metric | Euclidean | |
| Weights | Distance-weighted | |
| SVM | Kernel | RBF |
| C (regularization) | 10.0 | |
| γ | Scale (auto) | |
| Random Forest | Number of trees | 300 |
| Max depth | 12 | |
| Min samples split | 5 | |
| XGBoost | Number of trees | 200 |
| Max depth | 6 | |
| Learning rate | 0.05 | |
| Subsample | 0.8 |
Advanced tabular deep learning models
Beyond classical machine learning baselines, we benchmarked the proposed model against several state-of-the-art deep learning architectures that have recently demonstrated strong performance on structured and tabular datasets. These models integrate modern deep neural mechanisms such as attention, gating, and differentiable tree structures, allowing them to learn high-level feature representations that go beyond conventional feature engineering.
Deep Neural Network (DNN) - a multilayer perceptron with several fully connected hidden layers activated by ReLU. The DNN serves as the simplest deep learning baseline, directly learning mappings from normalized inflammatory and biochemical features to the outcome label.
DeepGBM [32] - a hybrid model that combines Gradient Boosted Decision Trees with neural networks. It first learns feature embeddings from GBDT-leaf indices and then feeds them into a deep neural network for representation refinement. This design enables DeepGBM to leverage both the interpretability of decision trees and the expressive power of deep learning, making it particularly suitable for heterogeneous clinical features.
TabNet [33] - a sequential attention-based architecture that dynamically selects relevant features for each decision step using sparse masks. Its interpretability stems from the ability to visualize which features contribute most to each prediction, aligning well with clinical reasoning processes. TabNet’s sparse attention mechanism improves generalization and robustness in datasets with high feature correlation, such as inflammatory marker panels.
Neural Oblivious Decision Ensembles (NODE) [34] - a differentiable ensemble of decision trees trained end-to-end using backpropagation. Unlike conventional trees, NODE allows smooth optimization of split parameters, thus integrating the hierarchical decision-making process of trees with the gradient learning capabilities of neural networks. This model captures complex feature hierarchies in tabular data and performs well under moderate dataset sizes typical in medical prognosis tasks.
TabTransformer [35] - employs a Transformer-based self-attention mechanism to learn contextual dependencies among both categorical and continuous variables. The model constructs embedding representations for categorical features and refines them via multiple Transformer encoder layers, enabling cross-feature interaction modeling. Its ability to capture relationships among inflammatory markers enhances predictive interpretability and discriminative performance in clinical settings.
All deep tabular models were reimplemented based on their original publications using publicly available repositories.Hyperparameter tuning, including embedding dimension, attention head size, and layer depth, was performed via grid search to maximize validation AUC. This comparative setup allows a comprehensive assessment of whether the proposed inflammation-driven deep learning model provides significant improvements beyond the current state-of-the-art tabular architectures.
The hyperparameter configurations for all deep tabular baselines are reported in Table 4. All models were reimplemented based on their original publications using publicly available repositories and tuned via manual tuning to maximize validation AUC.
Table 4.
Hyperparameter configurations of advanced deep tabular learning baselines
| Model | Hyperparameter | Value |
|---|---|---|
| DNN | Hidden layers | 3 (256 → 128 → 64) |
| Activation | ReLU | |
| Dropout rate | 0.3 | |
| Batch size | 64 | |
| Learning rate | ![]() |
|
| DeepGBM [32] | GBDT trees | 100 |
| GBDT max depth | 6 | |
| DNN hidden layers | 2 (128 → 64) | |
| Embedding dim | 64 | |
| Learning rate | ![]() |
|
| TabNet [33] | Decision steps | 5 |
| Attention embedding dim | 32 | |
| Relaxation factor | 1.5 | |
| Sparsity coefficient | ![]() |
|
| Learning rate | ![]() |
|
| NODE [34] | Number of layers | 4 |
| Trees per layer | 256 | |
| Tree depth | 6 | |
| Learning rate | ![]() |
|
| TabTransformer [35] | Transformer layers | 6 |
| Attention heads | 8 | |
| Embedding dim | 32 | |
| MLP hidden dim | 128 | |
| Learning rate | ![]() |
Results and discussion
Performance evaluation
The proposed inflammation marker-driven model achieved superior predictive performance for postoperative gastric cancer prognosis. Our model reached an AUC of 0.902, PRAUC of 0.873, ACC of 0.828, Precision of 0.822, and Recall of 0.741, surpassing both traditional machine learning and modern deep tabular baselines. Table 5 presents the performance comparison between our proposed model and conventional machine learning algorithms. Among the baselines, XGB achieved the best overall performance, followed by Random Forest (RF). However, both models were still outperformed by the proposed model, which demonstrated higher discriminative capability and better calibration on unseen samples.
Table 5.
Performance comparison of the proposed model with traditional machine learning baselines
| Model | ACC | Precision | Recall | AUC | PRAUC |
|---|---|---|---|---|---|
| Decision Tree | 0.711 | 0.693 | 0.625 | 0.774 | 0.742 |
| k-Nearest Neighbors | 0.728 | 0.715 | 0.652 | 0.781 | 0.752 |
| Support Vector Machine | 0.743 | 0.731 | 0.678 | 0.805 | 0.778 |
| Random Forest | 0.763 | 0.754 | 0.691 | 0.824 | 0.801 |
| Extreme Gradient Boosting | 0.778 | 0.766 | 0.702 | 0.837 | 0.814 |
| Ours | 0.828 | 0.822 | 0.741 | 0.902 | 0.873 |
Compared with the best-performing traditional model, our model achieved an improvement of 6.5% in AUC and 7.3% in PRAUC. This enhancement can be attributed to ability to capture nonlinear dependencies among inflammatory biomarkers through its hybrid feature encoding and deep representation layers. Traditional ensemble methods such as RF and XGB are powerful for structured data but are limited by their fixed tree-based partitioning, which may overlook complex cross-feature relationships between inflammatory indices like CRP, NLR, PLR, and albumin levels. In contrast, our model adapts to such dependencies automatically, leading to improved prognostic discrimination.
The results in Table 6 show that our model consistently achieved the highest accuracy and AUC among all advanced methods.
Table 6.
Performance comparison with advanced deep tabular learning models
| Model | ACC | Precision | Recall | AUC | PRAUC |
|---|---|---|---|---|---|
| Deep Neural Network (DNN) | 0.782 | 0.771 | 0.709 | 0.842 | 0.823 |
| DeepGBM [32] | 0.793 | 0.781 | 0.716 | 0.852 | 0.832 |
| TabNet [33] | 0.799 | 0.788 | 0.724 | 0.859 | 0.842 |
| NODE [34] | 0.808 | 0.799 | 0.731 | 0.865 | 0.851 |
| TabTransformer [35] | 0.817 | 0.808 | 0.735 | 0.873 | 0.860 |
| Ours | 0.828 | 0.822 | 0.741 | 0.902 | 0.873 |
Compared with TabTransformer, which leverages attention-based feature interactions, our model achieved a 3.3% gain in AUC and a 1.3% gain in PRAUC. This indicates that the integration of GBDT-based structure encoding (TreeDrivenEncoder) and deep neural representations allows for more expressive modeling of heterogeneous clinical biomarkers.Moreover, encoding process of our model transforms heterogeneous tabular features into homogeneous binary embeddings, making the learning process more stable and data-efficient.
To visualize the discriminative capability of our model, Fig. 2 illustrates the Receiver Operating Characteristic curves. As observed, our model consistently achieves the highest AUC in both settings, highlighting its strong generalization on unseen samples.
Fig. 2.

AUC comparison across models. (A) Classical machine-learning baselines vs. The proposed model. (B) Advanced deep tabular models vs. the proposed model
Model interpretability via SHAP analysis
To further interpret the contribution of individual inflammatory biomarkers to the model predictions, we employed SHapley Additive exPlanations (SHAP) to quantify feature importance at the global level. SHAP values provide a unified framework for explaining the output of complex models by attributing each feature’s contribution to the final prediction.
As illustrated in Fig. 3, the SHAP analysis reveals that systemic inflammatory markers such as NLR, CRP, and albumin level are among the most influential predictors of postoperative prognosis in gastric cancer patients. Higher CRP and NLR values were consistently associated with increased risk, while higher albumin levels exhibited a protective effect, contributing negatively to the predicted risk score.
Fig. 3.

SHAP summary plot showing global feature importance and direction of impact of inflammatory biomarkers on model predictions
These findings are consistent with established clinical evidence regarding the role of systemic inflammation and nutritional status in cancer progression. Importantly, the SHAP distribution also demonstrates heterogeneous patient-level effects, indicating that the same biomarker may contribute differently depending on the overall clinical context captured by the model.
Model stability and robustness
To assess the robustness and reproducibility of the proposed model, we conducted ten independent training runs using different random seeds. The results summarized in Table 7 demonstrate the stability of proposed model across multiple trials. The variation across seeds was minimal, indicating consistent convergence behavior and strong generalization ability.
Table 7.
Performance stability of the proposed model across 10 different random seeds
| Trial | ACC | Precision | Recall | AUC | PRAUC |
|---|---|---|---|---|---|
| 1 | 0.826 | 0.818 | 0.742 | 0.901 | 0.871 |
| 2 | 0.831 | 0.825 | 0.745 | 0.905 | 0.876 |
| 3 | 0.829 | 0.821 | 0.740 | 0.903 | 0.875 |
| 4 | 0.824 | 0.819 | 0.738 | 0.900 | 0.870 |
| 5 | 0.832 | 0.826 | 0.746 | 0.905 | 0.877 |
| 6 | 0.827 | 0.821 | 0.743 | 0.902 | 0.872 |
| 7 | 0.829 | 0.822 | 0.739 | 0.904 | 0.874 |
| 8 | 0.830 | 0.823 | 0.744 | 0.903 | 0.873 |
| 9 | 0.825 | 0.819 | 0.740 | 0.901 | 0.871 |
| 10 | 0.828 | 0.821 | 0.741 | 0.902 | 0.872 |
| Mean (Std) | 0.828 (0.003) | 0.822 (0.003) | 0.742 (0.002) | 0.903 (0.002) | 0.873 (0.002) |
| 95% CI | (0.825–0.831) | (0.819–0.825) | (0.739–0.745) | (0.901–0.905) | (0.871–0.875) |
These results confirm that the our model maintains high reproducibility, with the standard deviation of AUC and PRAUC both below 0.002. Such consistency is particularly important for clinical decision support systems, where predictive reliability and stability across training runs are critical. The robustness of proposed model can be attributed to its hybrid architecture combining tree-driven encoding and deep representation learning, which effectively mitigates overfitting while preserving discriminative power.
Prognostic implications
To better understand the prognostic mechanisms underlying the superior predictive performance of our model, we conducted an in-depth analysis of the learned features and their clinical implications. The inflammation-based biomarkers, including NLR, PLR, and COP-NLR, emerged as the most influential predictors in the prognostic estimation. Among these, NLR showed the strongest correlation with postoperative survival. The results revealed that patients with NLR > 3 had significantly lower survival rates, consistent with previous clinical findings that elevated NLR reflects an imbalance between tumor-promoting inflammation and antitumor immune response.
Similarly, PLR demonstrated substantial prognostic relevance. High PLR values often indicate a chronic inflammatory and pro-thrombotic state, which facilitates tumor angiogenesis and metastasis. Data analysis showed that patients in the highest PLR quartile exhibited a markedly higher postoperative mortality rate compared with those in the lowest quartile.
In addition to inflammatory biomarkers, TNM and histological grade remained fundamental prognostic determinants. However, when these pathological factors were modeled together with inflammation markers, our model uncovered multiple nonlinear relationships that traditional methods could not capture. For instance, some Stage II patients with elevated NLR or PLR levels exhibited a mortality risk comparable to Stage III patients, suggesting that systemic inflammation can increase tumor aggressiveness and biological toxicity independent of anatomical extent.
Overall, these findings indicate that postoperative prognosis in gastric cancer is not solely determined by tumor characteristics but is also profoundly influenced by the patient’s systemic inflammatory and nutritional status. The model’s ability to autonomously learn and exploit these complex relationships explains its superior performance over conventional machine learning approaches. From a clinical perspective, integrating inflammation-based biomarkers into postoperative monitoring systems could enable earlier identification of high-risk patients and support more personalized treatment and prognostic management strategies.
Limitations and future work
Although the proposed model demonstrates high accuracy and stability in predicting postoperative prognosis for gastric cancer patients, this study still has several limitations. First, the dataset was primarily collected from a limited number of medical institutions with a relatively small sample size, which may affect the model’s generalizability when applied to more diverse patient populations. Future studies should conduct external validations across multiple independent centers to assess the model’s transferability under different clinical conditions.
In addition, while the Tree-Driven Encoder effectively captures the structural information from decision trees to enhance feature representation, the interpretability of the deep learning components remains limited. Future research should focus on integrating explainable artificial intelligence techniques or attention-based visualization mechanisms to clarify the biological and clinical implications of the learned features, thereby improving the transparency and reliability of the model in medical applications.
Moreover, recent advances in deep learning have demonstrated the effectiveness of multimodal fusion and interpretable representation learning in complex application scenarios [36]. Integrating such techniques into clinical prognostic frameworks represents a promising direction for improving both prediction accuracy and model transparency in future studies.
Finally, the current model primarily relies on static preoperative and intraoperative variables without considering temporal dynamics or postoperative changes in inflammatory and nutritional indices. Incorporating longitudinal data and temporal deep learning architectures such as recurrent neural networks or Transformers could enhance predictive accuracy and better capture the real-world progression of disease outcomes.
Conclusion
This study introduces a novel deep learning framework that integrates three key components GBDT, TDE, and 1D CNN block for postoperative prognosis prediction in gastric cancer patients. By combining the structural representation capability of decision tree models with the feature learning power of deep neural networks, the proposed model effectively captures nonlinear and hierarchical relationships among clinical, pathological, and systemic inflammatory features. Experimental results demonstrate that our model outperforms traditional machine learning algorithms and several advanced tabular deep learning methods. Beyond improving prognostic prediction performance, the model also provides a more comprehensive understanding of the interactions among inflammation, nutrition, and tumor biology. This integration offers significant potential for supporting personalized postoperative treatment strategies and promoting data-driven clinical decision-making. In future work, expanding the dataset to include multiple independent centers, incorporating temporal information, and enhancing interpretability through explainable artificial intelligence techniques will be important steps toward real-world clinical application.
Author contributions
All authors designed the method. Q.Z. wrote the code and performed the experiments. All authors wrote and reviewed the manuscript.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
The dataset utilized in this study can be accessed at https://www.ebi.ac.uk/biostudies/studies?query=S-EPMC5373584.
Code availability
The code used in this study is available from the corresponding author on reasonable request.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229–63. 10.3322/caac.21834. [DOI] [PubMed] [Google Scholar]
- 2.Yu Q, Li K-Z, Fu Y-J, Tang Y, Liang X-Q, Liang Z-Q, et al. Clinical significance and prognostic value of C-reactive protein/albumin ratio in gastric cancer. Ann Surg Treat Res. 2021;100(6):338. 10.4174/astr.2021.100.6.338. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Liu D, Jiang Y, Gao M. Development and validation of multiparameter prognostic nomogram combining neutrophil-to-lymphocyte ratio, tumor burden, and nutritional status for predicting postoperative outcome in elderly gastric cancer patients. Front Oncol. 2025;15. 10.3389/fonc.2025.1620443. [DOI] [PMC free article] [PubMed]
- 4.D’Souza J, McCombie A, Roberts R. The influence of short-term postoperative outcomes on overall survival after gastric cancer surgery. ANZ J Surg. 2023;93(12):2875–84. 10.1111/ans.18613. [DOI] [PubMed] [Google Scholar]
- 5.Kuwayama N, Hoshino I, Mori Y, Yokota H, Iwatate Y, Uno T. Applying artificial intelligence using routine clinical data for preoperative diagnosis and prognosis evaluation of gastric cancer. Oncol Lett. 2023;26(5). 10.3892/ol.2023.14087. [DOI] [PMC free article] [PubMed]
- 6.Diao JA, Wang JK, Chui WF, Mountain V, Gullapally SC, Srinivasan R, et al. Human-interpretable image features derived from densely mapped cancer pathology slides predict diverse molecular phenotypes. Nat Commun. 2021;12(1). 10.1038/s41467-021-21896-9. [DOI] [PMC free article] [PubMed]
- 7.Kourou K, Exarchos TP, Exarchos KP, Karamouzis MV, Fotiadis DI. Machine learning applications in cancer prognosis and prediction. Comput Struct Biotechnol J. 2015;13:8–17. 10.1016/j.csbj.2014.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Wang D, Pan B, Huang J-C, Chen Q, Cui S-P, Lang R, et al. Development and validation of machine learning models for predicting prognosis and guiding individualized postoperative chemotherapy: a real-world study of distal cholangiocarcinoma. Front Oncol. 2023;13. 10.3389/fonc.2023.1106029. [DOI] [PMC free article] [PubMed]
- 9.Chen S, Zang Y, Xu B, Lu B, Ma R, Miao P, et al. An unsupervised deep learning-based model using multiomics data to predict prognosis of patients with stomach adenocarcinoma. Comput Math Method Med. 2022;2022:1–20. 10.1155/2022/5844846. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Wang Y-Y, Li L, Zhao Z-S, Wang Y-X, Ye Z-Y, Tao H-Q. L1 and epithelial cell adhesion molecules associated with gastric cancer progression and prognosis in examination of specimens from 601 patients. J Exp Clin Cancer Res. 2013;32(1). 10.1186/1756-9966-32-66. [DOI] [PMC free article] [PubMed]
- 11.Kirouac DC, Zmurchok C, Morris D. Making drugs from T cells: the quantitative pharmacology of engineered T cell therapeutics. NPJ Syst Biol Appl. 2024;10(1):31. 10.1038/s41540-024-00355-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Niarakis A, Laubenbacher R, An G, Ilan Y, Fisher J, Flobak Å, et al. Immune digital twins for complex human pathologies: applications, limitations, and challenges. NPJ Syst Biol Appl. 2024;10(1):141. 10.1038/s41540-024-00450-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Choi J-H, Suh Y-S, Choi Y, Han J, Kim TH, Park S-H, et al. Comprehensive analysis of the neutrophil-to-lymphocyte ratio for preoperative prognostic prediction nomogram in gastric cancer. World J Surg. 2018;42(8):2530–41. 10.1007/s00268-018-4510-4. [DOI] [PubMed] [Google Scholar]
- 14.Huang B, Tian S, Zhan N, Ma J, Huang Z, Zhang C, et al. Accurate diagnosis and prognosis prediction of gastric cancer using deep learning on digital pathological images: a retrospective multicentre study. EBioMedicine. 2021;73:103631. 10.1016/j.ebiom.2021.103631. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Tran KA, Kondrashova O, Bradley A, Williams ED, Pearson JV, Waddell N. Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med. 2021;13(1). 10.1186/s13073-021-00968-x. [DOI] [PMC free article] [PubMed]
- 16.Mellor KL, Powell AGMT, Lewis WG. Systematic review and meta-analysis of the prognostic significance of neutrophil-lymphocyte ratio (NLR) after r0 gastrectomy for cancer. J Gastrointest Canc. 2018;49(3):237–44. 10.1007/s12029-018-0127-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Lee DY, Hong SW, Chang YG, Lee WY, Lee B. Clinical significance of preoperative inflammatory parameters in gastric cancer patients. J Gastric Cancer. 2013;13(2):111. 10.5230/jgc.2013.13.2.111. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Han Y, Lv W, Guo J, Shang Y, Yang F, Zhang X, et al. Prognostic significance of inflammatory and nutritional index for serous ovary cancer. Preprint at Research Square:rs-3509733/v1. 2023.
- 19.Wu Z-J, Xu H, Wang R, Bu L-J, Ning J, Hao J-Q, et al. Cumulative score based on preoperative fibrinogen and pre-albumin could predict long-term survival for patients with resectable gastric cancer. J Cancer. 2019;10(25):6244–51. 10.7150/jca.35157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Dong X, Lou YB, Mu YC, Kang MX, Wu YL. Predictive factors for differentiating pancreatic cancer-associated diabetes mellitus from common type 2 diabetes mellitus for the early detection of pancreatic cancer. Digestion. 2018;98(4):209–16. 10.1159/000489169. [DOI] [PubMed] [Google Scholar]
- 21.Shao L-L, Li X, Wang L-F. Role of neutrophil-to-lymphocyte, platelet-to-lymphocyte, and monocyte-to-lymphocyte ratios in rectal cancer prognosis. World J Gastrointest Surg. 2025;17(6). 10.4240/wjgs.v17.i6.106813. [DOI] [PMC free article] [PubMed]
- 22.Zheng Q, Ge H. The role of mir96 in predicting CTC status and prognostic evaluation in gastric cancer patients. Cell Mol Biol (Noisy-le-Grand). 2024;70(2):264–74. 10.14715/cmb/2024.70.2.37. [DOI] [PubMed] [Google Scholar]
- 23.Xu W, Liu X, Yan C, Abdurahmane G, Lazibiek J, Zhang Y, et al. The prognostic value and model construction of inflammatory markers for patients with non-small cell lung cancer. Sci Rep. 2024;14(1). 10.1038/s41598-024-57814-4. [DOI] [PMC free article] [PubMed]
- 24.Radulescu PM, Davitoiu DV, Baleanu VD, Padureanu V, Ramboiu DS, Surlin MV, et al. Has COVID-19 modified the weight of known systemic inflammation indexes and the new ones (MCVL and IIC) in the assessment as predictive factors of complications and mortality in acute pancreatitis? Diagnostics. 2022;12(12):3118. 10.3390/diagnostics12123118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Tada H, Kawabata-Iwakawa R, Takahashi H, Chikamatsu K. Novel index based on inflammatory markers correlates with treatment efficacy of nivolumab for recurrent/metastatic head and neck cancer. Oncology. 2024;103(8):714–24. 10.1159/000542683. [DOI] [PubMed] [Google Scholar]
- 26.Borisov V, Broelemann K, Kasneci E, Kasneci G. DeepTLF: robust deep neural networks for heterogeneous tabular data. Int J Data Sci Anal. 2023;16(1):85–100. 10.1007/s41060-022-00350-z. [DOI] [Google Scholar]
- 27.Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. 10.1023/A:1010933404324. [DOI] [Google Scholar]
- 28.Hearst MA, Dumais ST, Osuna E, Platt J, Scholkopf B. Support vector machines. IEEE Intell Syst Their Appl. 1998;13(4):18–28. 10.1109/5254.708428. [DOI] [Google Scholar]
- 29.Cover T, Hart P. Nearest neighbor pattern classification. IEEE Trans Inf Theory. 1967;13(1):21–27. 10.1109/TIT.1967.1053964. [DOI] [Google Scholar]
- 30.Quinlan JR. Induction of decision trees. Mach Learn. 1986;1(1):81–106. 10.1023/A:1022643204877. [DOI] [Google Scholar]
- 31.Chen T, Guestrin C. Xgboost: a scalable tree Boosting System. Preprint at arXiv:1603.02754. 2016.
- 32.Ke G, Xu Z, Zhang J, Bian J, Liu T-Y. DeepGBM: a deep learning framework distilled by gbdt for online prediction tasks. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. New York: ACM; 2019, pp. 384–94.
- 33.Arik SO, Pfister T. TabNet: attentive interpretable tabular learning. Preprint at arXiv:1908.07442. 2019.
- 34.Popov S, Morozov S, Babenko A. Neural oblivious decision ensembles for deep learning on tabular data. Preprint at arXiv:1909.06312. 2019.
- 35.Huang X, Khetan A, Cvitkovic M, Karnin Z. TabTransformer: tabular data modeling using contextual embeddings. Preprint at arXiv:2012.06678. 2020.
- 36.Li G, Zhao B, Su X, Yang Y, Hu P, Zhou X, et al. Discovering consensus regions for interpretable identification of RNA N6-methyladenosine modification sites via graph contrastive clustering. IEEE J Biomed Health Inf. 2024;28(4):2362–72. 10.1109/JBHI.2024.3357979. [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The dataset utilized in this study can be accessed at https://www.ebi.ac.uk/biostudies/studies?query=S-EPMC5373584.
The code used in this study is available from the corresponding author on reasonable request.












