Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 26;16:23957. doi: 10.1038/s41598-026-54867-5

Explainable machine learning-driven identification of heart failure biomarkers: a multi-model feature selection approach with SHAP-based interpretability

Yuhe Zhao 1,#, Ruoyu Zhang 1,#, Kelan Zha 2,#, Yafei Li 1, Huan Li 1, Yong Wang 1, Shuren Dai 1,✉, Yu Zeng 1,✉
PMCID: PMC13434832  PMID: 42185599

Abstract

Heart failure (HF) remains a major clinical challenge due to its complex pathophysiology and the limitations of existing biomarkers. In this study, we developed a robust machine learning (ML) framework to identify novel transcriptomic signatures of HF. Three GEO RNA-seq datasets (GSE141910, GSE198945, GSE263297) were integrated and harmonized, followed by a “split-first” strategy for training (70%) and testing (30%). We employed a triphasic feature selection process—integrating LASSO, Random Forest (RF), and SVM-RFE—to identify candidate genes. A 10-model ensemble system was evaluated using Leave-One-Study-Out Cross-Validation (LOSO-CV) and interpreted via SHAP values. Findings were validated using an independent external cohort (GSE135055) and experimental RT-qPCR in a local clinical cohort. Three potential biomarkers—FNDC1, LPCAT3, and TIMP2—were prioritized. FNDC1 and TIMP2 were significantly upregulated, while LPCAT3 was suppressed in HF tissues (p < 0.001), patterns consistently confirmed by qPCR. The ML models demonstrated high diagnostic stability, with peak LOSO-CV AUCs reaching 0.973 and maintaining robustness in external validation (AUC up to 0.876). SHAP analysis identified FNDC1 as the most influential predictor. Functional enrichment linked these signatures to extracellular matrix remodeling and lipid metabolism. These findings suggest that FNDC1, LPCAT3, and TIMP2 may serve as potential biomarkers associated with the pathological mechanisms of HF.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-54867-5.

Keywords: Heart failure, Machine learning, SHAP, FNDC1, LPCAT3, TIMP2

Subject terms: Biomarkers, Cardiology, Computational biology and bioinformatics, Diseases

Introduction

Heart failure (HF) is a complex syndrome characterized by progressive deterioration of cardiac pumping function, with a global prevalence exceeding 64 million cases and a five-year mortality rate as high as 50%, posing a significant public health burden1–3. A hallmark of HF pathophysiology is cardiac fibrosis, marked by myocardial scarring and excessive extracellular matrix (ECM) deposition4. Fibrosis can be triggered by myocardial injury, such as infarction, chronic hypertension, or ischemia5. While initially compensatory, persistent fibrosis weakens the myocardium and ultimately drives HF progression. Mainstay therapies like β-blockers and angiotensin receptor inhibitors alleviate symptoms but fail to reverse cardiomyocyte loss or aberrant fibrosis-mediated irreversible remodeling6–8. Studies reveal that adult mammalian cardiomyocytes exhibit negligible regenerative capacity post-injury, leading to fibroblast-dominated repair responses9,10. This results in pathological ECM accumulation and scar formation, increasing ventricular stiffness and precipitating functional decompensation11. Although targeting ECM metabolism is a promising therapeutic strategy, the absence of dynamic fibrosis-specific biomarkers hinders precision interventions.

In recent years, machine learning has emerged as a transformative force in cardiovascular disease research. As emphasized in the 2024 Scientific Statement from the American Heart Association, artificial intelligence and machine learning are now central to redefining diagnostic and classification paradigms to improve heart disease outcomes12. Advanced algorithms such as support vector machines, RF, and deep neural networks can extract latent patterns from complex, high-dimensional data, offering robust support for disease diagnosis, risk stratification, and therapeutic decision-making. Yamasan BE et al. utilizing Support Vector Machine (SVM) and RF models demonstrated high predictive accuracy for hypertrophic cardiomyopathy using differentially expressed miRNAs, suggesting new paths for early diagnosis and personalized therapy13. Current studies indicate that these ML-driven approaches are particularly superior to traditional statistical methods in capturing the heterogeneous nature of HF survival and progression14. In diagnostic and risk prediction domains, Machine Learning models have significantly enhanced traditional biomarker applications. For instance, MIT’s CHAIS system achieves 87.5% accuracy in predicting HFrisk through single-lead ECG analysis of left atrial pressure, matching the efficacy of invasive right heart catheterization15. Similarly, convolutional neural networks can automatically identify myocardial fibrosis features in echocardiograms, quantifying ECM deposition for dynamic monitoring16,17.

While high-throughput omics technologies have identified numerous candidate genes, conventional analytical methods fail to distinguish driver pathological signals from secondary expression perturbations. Recent bibliometric and trend analyses highlight a paradigm shift in the field, moving from traditional clinical risk factors to the integration of high-throughput transcriptomic and proteomic data to decipher the molecular complexity of cardiomyopathies18. Moreover, they lack functional linkages between biomarkers and ECM remodeling or immune microenvironment dynamics. Thus, novel biomarker discovery systems integrating multi-dimensional data analytics are becoming pivotal for deciphering HF’s molecular mechanisms.

However, key challenges persist in high-dimensional genomic data analysis: small sample sizes (n < 100), batch effects, and class imbalance often compromise the stability of traditional biomarker screening methods in identifying pathologically interpretable targets. Addressing this requires innovative AI frameworks that balance algorithmic robustness with biological relevance.Ensemble learning methods such as XGBoost, RF have addressed this by prioritizing pathological driver genes through feature importance ranking in cardiac tissue transcriptomes. Recent Machine Learning approaches have identified core fibrosis biomarkers including: Potassium channel gene KCNQ119, Collagen metabolism gene COL1A120, Inflammatory regulator IL-6R21. Notably, the development of explainable artificial intelligence techniques, particularly SHAP (Shapley Additive Explanations) analysis, has significantly enhanced the interpretability of machine learning models in biomedical research22. However, it should be noted that while SHAP values can effectively quantify feature contributions to predictive models, their biological interpretation still requires integration with domain-specific analytical methods.

This study investigates key molecular mechanisms underlying HF progression by integrating machine learning and bioinformatics approaches. Leveraging bulk RNA sequencing data, we employed multiple machine learning methods to establish a multi-model consensus framework, thereby addressing biases inherent in single-omics analyses caused by high dimensionality, small sample sizes, and batch effects. Further refinement was achieved through SHAP interpretability modeling. Additional analyses exploring the relationship between HF-specific biomarkers and immune cell relative proportions will provide novel insights into HF pathogenesis. The study workflow is illustrated in Fig. 1.

Fig. 1.

Fig. 1

Workflow of the integrated bioinformatics and machine-learning pipeline for identifying diagnostic signatures in HF.

Materials and methods

Data information

The data used in this study were obtained from the Gene Expression Omnibus (GEO) database of the National Center for Biotechnology Information, incorporating four heart failure-related high-throughput transcriptomic datasets (GSE14191023, GSE19894524, and GSE263297,GSE13505525). Raw data were downloaded using the GEOquery R package (version 3.18.0)26, comprising human left ventricular myocardial tissue samples from both HF patients (n = 222) and non-heart failure controls (n = 235). The “non-heart failure controls” were specifically identified as non-failing donor hearts unsuitable for transplantation due to non-cardiac factors (e.g., age or size mismatch), serving as a healthy baseline for myocardial tissue analysis.

Data preprocessing included the following steps: First, probe annotation and gene symbol standardization were performed based on the corresponding GPL platform for each dataset, followed by background correction and quantile normalization using the limma package (version 3.56.2)27. Second, cross-study batch effects were corrected using the ComBat algorithm from the sva package (version 3.48.0) 28, To ensure the preservation of biological signals and prevent potential over-correction, biological group labels (HF vs. Control) were explicitly included as a covariate (model matrix) in the ComBat algorithm. This approach protects the variance associated with the biological “Condition” from being treated as a batch effect, thereby safeguarding the pathological signatures of interest. The correction efficacy was validated by principal component analysis (PCA). Finally, low-expression filtering was applied to the gene expression matrix (retaining genes with Counts Per Million ≥ 1 in > 50% of samples), followed by log2 transformation to generate a standardized merged expression profile.

Differential gene expression analysis

Differentially expressed genes (DEGs) were identified using the R package “limma” by comparing gene expression profiles between HF and control groups. To address the multiple testing problem inherent in high-throughput transcriptomic datasets—where testing thousands of genes simultaneously increases the risk of false-positive results—p-values were adjusted using the Benjamini-Hochberg (BH) method to control the False Discovery Rate (FDR). This procedure ranks all raw p-values (Pi) in ascending order and adjusts them such that the k-th smallest p-value is compared against (k/n) times α, where n is the total number of genes. The resulting adjusted p-value (FDR) represents the expected proportion of incorrectly rejected null hypotheses. Significance thresholds were set at an adjusted p-value (FDR) < 0.05 and a magnitude of change |log2 fold change (logFC)| > 1.

Feature gene selection

To prevent information leakage and ensure the robustness of the biomarkers, the integrated dataset was first randomly partitioned into a training set (70%) and an independent test set (30%) prior to any feature selection or modeling. Three machine learning algorithms were utilized for key feature gene selection: LASSO regression, RF29, and Support Vector Machine (SVM). LASSO regression30 was implemented via the R package “glmnet"31 (version 4.1-7), with 10-fold cross-validation to optimize the regularization parameter λ (lambda.min criterion), retaining genes with non-zero coefficients. RF analysis (R package “randomForest”, version 4.7–1.1)32 initially evaluated gene importance using the Gini index (preserving the top 20%). To assess and correct for potential Gini-related bias, we concurrently evaluated Permutation Importance (Mean Decrease in Accuracy). Most importantly, to strictly mitigate single-algorithm biases, we relied on a multi-algorithm consensus strategy. The SVM model (R package “e1071”, version 1.7–13) employed recursive feature elimination (SVM-RFE) to iteratively remove low-weight genes. Genes identified by all three methods were defined as key feature genes for subsequent analysis. Mann-Whitney U test (Wilcoxon rank-sum test) was applied to compare key gene expression between subgroups (HF/control).

Functional enrichment analysis

To elucidate the biological functions and pathway characteristics of differentially expressed genes (DEGs), this study integrated multi-level enrichment analysis methods. Gene Ontology (GO)33 and Kyoto Encyclopedia of Genes and Genomes (KEGG)34 enrichment analyses were performed using the R package “clusterProfiler” (version 4.8.1)35. GO analysis covered three categories: Biological Process (BP), Molecular Function (MF), and Cellular Component (CC), with significance thresholds set as an adjusted p-value (FDR) < 0.05 and an enrichment factor > 1.5. Gene Set Enrichment Analysis (GSEA)36 was conducted using the “fgsea” package (version 1.26.0), with genes ranked by log2 fold change (logFC). Gene Set Variation Analysis (GSVA)37 was implemented with the R package “GSVA” (version 1.48.3) to quantify sample-specific pathway activity using KEGG gene sets in an unsupervised manner. For both GSEA and GSVA, gene-set sizes were restricted to a minimum of 10 and a maximum of 500 genes to ensure statistical robustness. Multiple-testing adjustments were performed using the Benjamini-Hochberg (BH) method, and a FDR < 0.05 was considered statistically significant. Differential pathway activity between the HF and control groups was compared using the limma package (|logFC| > 0.5, FDR < 0.05). Results were visualized using the “ggplot2” and “enrichplot” packages.

Machine learning model performance evaluation and feature importance analysis

This study employed 10 machine learning algorithms to construct diagnostic models for HF, including Partial Least Squares (PLS), RF, Decision Tree (DTS), Support Vector Machine (SVM), Logistic Regression (Logistic), K-Nearest Neighbors (KNN), Extreme Gradient Boosting (XGBoost), Gradient Boosting Machine (GBM), Neural Network (NeuralNet), and Generalized Linear Model Boosting (glmBoost). Ten machine learning algorithms were utilized to ensure that the identified biomarkers remained robust across diverse computational architectures. This multi-model approach, covering linear, tree-based, and neural network methods, was implemented to eliminate algorithmic bias and identify the most effective diagnostic framework for high-dimensional cardiac transcriptomic data through rigorous performance benchmarking. All analyses were performed in R (version 4.3.1) using the “caret” package (version 6.0–94)38.

Due to the random 70/30 split of the integrated discovery cohort, samples from the same initial GEO dataset were inevitably present in both the internal training and testing sets. While separate ComBat corrections effectively mitigate major batch effects, theoretical risks of dataset-specific residual signals artificially inflating internal test performance remain. To strictly rule out this bias and rigorously assess cross-study generalizability, we implemented a Leave-One-Study-Out Cross-Validation (LOSO-CV) strategy. In this framework, the models were iteratively trained on two of the integrated datasets and validated on the omitted third dataset. This process was repeated for all datasets to obtain an unbiased estimate of internal robustness.

Additionally, for the discovery cohort, data were randomly split into a training set (70%) and a test set (30%) with a fixed seed (set.seed(12345)). All model training and hyperparameter optimization were performed using the training set with 5-fold repeated cross-validation. An external validation cohort (GSE135055) was reserved solely for final performance evaluation (AUC, sensitivity, and specificity) to provide an estimate of model generalizability. ROC curves were generated using the “pROC” package (version 1.18.4). The final optimized hyperparameters for all ten machine learning models, determined via grid search within the training set, are detailed in Supplementary Table S1.

For the optimal model, SHAP values were calculated using the “kernelshap” package (version 0.3.1) to interpret the contribution of key feature genes to predictions. SHAP summary plots and dependence plots were generated with the “shapviz” package (version 0.9.2).

To rigorously assess model robustness against overfitting, permutation tests were conducted. The training set labels were randomly shuffled 1,000 times to construct null distributions of the AUC for the top three models (GLMBOOST, PLS, SVM). An empirical p-value < 0.05, calculated as the proportion of permuted AUCs exceeding the actual AUC, indicated statistical significance.

Clinical sample collection and ethical approval

To validate the candidate biomarkers, endomyocardial biopsy (EMB) samples were obtained from a clinical cohort at the Seventh People’s Hospital of Chongqing, consisting of 10 HF patients and 10 healthy controls. The study was approved by the Institutional Ethics Committee of the Seventh People’s Hospital of Chongqing (No. 2025-741-08), and written informed consent was obtained from all participants or their legal representatives prior to the procedure. Inclusion criteria for the HF group included a clinical diagnosis of HF according to ESC guidelines, evidence of structural heart disease, and impaired cardiac function (LVEF < 40%). Exclusion criteria encompassed active inflammatory diseases, malignancies, severe renal or hepatic dysfunction, and other systemic conditions that might confound transcriptomic profiles. Control samples were obtained from organ/cadaveric donors with no history of cardiac dysfunction or significant systemic disease. Following collection, biopsy specimens were immediately stabilized and stored at -80 °C for subsequent RT-qPCR analysis.

Immune cell composition analysis

To systematically characterize the dynamics of the immune microenvironment in HF samples, the CIBERSORT algorithm39 was applied using the R package “CIBERSORT” (version 0.1.0) to deconvolute myocardial tissue transcriptomic data and quantify the relative abundance of immune cell subsets. Specifically, the deconvolution process was performed using the validated LM22 reference matrix, which contains 547 leukocyte gene signatures, as the basis for identifying specific immune cell subsets. A total of 500 permutation tests (p < 0.05) were conducted to validate immune cell relative proportions scores. The ESTIMATE algorithm40 was implemented with the R package “estimate” (version 1.1.0) to calculate immune scores (ImmuneScore) for each sample.

Spearman correlation analysis (using the “stats” package, version 4.3.1) was performed to assess associations between the expression levels of key feature genes and specific immune cell relative proportions (|ρ| > 0.3, FDR < 0.05). Wilcoxon rank-sum tests were used to compare differences between HF and control groups (p < 0.05). Boxplots illustrating immune cell subset distributions and correlation heatmaps were generated using “ggplot2” (version 3.4.4) and “corrplot” (version 0.92), respectively.

RNA extraction and RT-qPCR analysis

Total RNA was extracted from the cryopreserved myocardial biopsy specimens using TRIzol reagent (Invitrogen, USA) according to the manufacturer’s protocol. The concentration and purity of the isolated RNA were quantified using a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific, USA), ensuring that the A260/A280 ratio ranged between 1.8 and 2.0. Subsequently, 1 µg of total RNA was reverse-transcribed into complementary DNA using a PrimeScript RT Reagent Kit (Takara, Japan). Quantitative real-time PCR (RT-qPCR) was performed on a StepOnePlus Real-Time PCR System (Applied Biosystems, USA) using SYBR Green Master Mix (Takara, Japan). The cycling conditions consisted of an initial denaturation at 95 °C for 30 s, followed by 40 cycles of 95 °C for 5 s and 60 °C for 30 s. GAPDH was employed as the internal reference gene for normalization. The relative expression levels of the target genes (FNDC1, LPCAT3, and TIMP2) were calculated using the 2^-ΔΔCt method. All experiments were performed in triplicate. The specific primer sequences utilized in this study are provided in Supplementary Table S2.

Statistical analysis

All statistical analyses were conducted in R (version 4.3.1) using parametric and non-parametric methods. Normality of continuous variables (e.g., gene expression values, immune scores) was assessed by the Shapiro-Wilk test. Non-normally distributed data were compared between groups (HF vs. control) using the Wilcoxon rank-sum test, with a significance threshold of two-tailed p < 0.05. DEGs were identified using linear models (limma package) with Benjamini-Hochberg correction for multiple hypothesis testing (FDR < 0.05).

Results

Data processing

A total of four HF transcriptomic datasets (GSE141910, GSE198945, GSE263297, and GSE135055) were integrated, comprising 457 human left ventricular myocardial samples (222 HF patients and 235 non-HF controls). To harmonize the discovery datasets, batch effects were corrected using the ComBat algorithm. Principal Component Analysis (PCA) was employed to visualize the integration efficiency. In the training set (70% of the discovery data), samples initially exhibited distinct clustering based on their study of origin (Fig. 2A), which was effectively eliminated following correction, resulting in a homogenous distribution across batches (Fig. 2B). Similar results were observed in the testing set (30% of the discovery data), where the dominant batch-induced variance (Fig. 2C) was successfully harmonized (Fig. 2D). Variance decomposition analysis (Supplementary Fig. 1) further confirmed that the proportion of variance attributed to “Batch” was reduced to near-zero levels while preserving the biological signals. Differential expression analysis was subsequently performed to identify candidate signatures (Fig. 2E, F).

Fig. 2.

Fig. 2

Data integration, batch effect correction, and differential expression analysis. (A) PCA visualization of the training set before batch effect correction. (B) PCA visualization of the training set after ComBat-mediated batch effect removal. (C) PCA visualization of the testing set before batch effect correction. (D) PCA visualization of the testing set after batch effect correction. (E) Heatmap displaying the expression profiles of identified differentially expressed genes (DEGs) across integrated datasets. (F) Volcano plot summarizing transcriptomic changes, where red and blue points denote significantly upregulated and downregulated genes, respectively (|log_2FC| >1 and adj. P < 0.05).

Identification of differentially expressed genes (DEGs)

Based on the integrated normalized expression profiles, differential gene expression analysis (HF group vs. control group) identified 568 significant DEGs, including 319 upregulated genes and 249 downregulated genes (Fig. 2F). A heatmap (Fig. 2E) illustrated the clustering pattern of the top 50 DEGs (Table 1).

Table 1.

HF dataset information list.

Discovery cohort Validation cohort
Internal training and validation External validation
GSE141910 GSE198945 GSE263297 GSE135055
Platform GPL16791 GPL24676 GPL24676 GPL16791
Species Homo sapiens Homo sapiens Homo sapiens Homo sapiens
Tissue Left ventricle Left ventricle Left ventricle Left ventricle
HFs 167 20 14 21
controls 199 20 7 9

Functional enrichment analysis of DEGs

Functional enrichment analysis of DEGs was performed to elucidate their biological roles. Gene Ontology (GO) results (Fig. 3A, B) revealed significant enrichment in the following categories. Biological Processes (BP): Key processes included ECM organization, positive regulation of cell adhesion, and inflammatory responses (e.g., leukocyte cell-cell adhesion, chemotaxis), highlighting the central role of these genes in myocardial fibrosis and immune microenvironment remodeling. Cellular Components (CC): Significant enrichment was observed in collagen-containing extracellular matrix, collagen trimer, external side of plasma membrane, and endoplasmic reticulum lumen, linking these genes to ECM deposition and protein secretion pathways. Molecular Functions (MF): Enriched terms included extracellular matrix structural constituent, glycosaminoglycan binding, and Wnt-protein binding, suggesting involvement in ECM biomechanics and signal transduction during HF progression. The GO enrichment network (Fig. 3C, D) illustrated interconnectivity among biological processes. Immune-related processes (e.g., leukocyte adhesion, lymphocyte activation) and ECM organization pathways (e.g., external encapsulating structure organization) exhibited strong associations, indicating coordinated roles in extracellular environment remodeling.

Fig. 3.

Fig. 3

Comprehensive functional annotation and enrichment landscape of DEGs. (A) GO enrichment bar chart: displays the most significantly enriched Gene Ontology terms (biological process, cellular component, molecular function; q < 0.05), ranked by −log10(q-value). (B) GO category pie chart: clusters the significant GO terms from panel A into functional categories. (C) GO functional interaction network: nodes represent GO terms and edges denote functional relationships; node size scales with the number of enriched genes and color intensity reflects q-value. (D) Force-directed projection of the network in C, highlighting core functional modules and their interconnections. (E) KEGG pathway enrichment bubble plot: illustrates significantly enriched KEGG pathways. (F) KEGG pathway interaction network: nodes are pathways connected by shared genes; node size corresponds to gene count and color depth indicates q-value. (G) Force-directed projection of the network in F.

KEGG enrichment analysis (Fig. 3E) showed significant enrichment in pathways such as Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway. A KEGG network analysis (Fig. 3F, G) further visualized cross-talk between pathways, emphasizing interactions among cytokine signaling, ECM-receptor pathways, and inflammatory cascades.

Screening of HF-related key feature genes and validation of differential expression

Based on the identified DEGs, three distinct machine learning algorithms—LASSO regression (λ.min = 0.032), RF; top 15% of genes ranked by Gini index), and SVM-RFE—were utilized to screen for key feature genes (Fig. 4). LASSO regression identified 46 non-zero coefficient genes (Fig. 4A, B), SVM-RFE selected 24 high-discriminatory genes (Fig. 4C, D), and RF prioritized 43 genes (Fig. 4E, F) (Supplementary Table S3). Notably, supplementary evaluation using Permutation Importance (Mean Decrease in Accuracy) confirmed that the prioritized candidates were not artifacts of Gini index bias, as FNDC1, TIMP2, and LPCAT3 consistently maintained top-tier importance rankings when evaluated through feature permutation (Supplementary Fig. 6). The intersection of these feature sets yielded three core signatures: FNDC1, LPCAT3, and TIMP2 (Fig. 5A). The robustness of the feature selection process was further supported by stability analysis (Supplementary Fig. 4), which demonstrated consistent high selection frequencies for the identified biomarkers across multiple model iterations.

Fig. 4.

Fig. 4

LASSO. Feature selection using three machine learning methods for marker gene identification. (A) LASSO. Path coefficients across different values of the regularization parameter λ. (B) LASSO. Deviance plot showing the relationship between the log-likelihood and the number of features. (C) SVM-RFE. 10-fold cross-validation error as a function of the number of features. (D) 10-fold cross-validation accuracy, (E) Random Forest. Mean decrease in accuracy for each feature across 500 trees. (F) Random Forest. Importance plot of the top 20 features.

Fig. 5.

Fig. 5

Identification and multi-cohort validation of the three-gene signature. (A) Venn diagram illustrating the intersection of candidate feature genes identified by LASSO, RF, and SVM-RFE algorithms. (B) Box plots comparing the expression levels of FNDC1, LPCAT3, and TIMP2 in the training set. (C) Box plots comparing gene expression levels in the internal testing set. (D) Validation of the expression patterns in the independent external cohort (GSE135055). (E) Experimental validation of relative mRNA expression levels via RT-qPCR in the clinical cohort (n = 10 per group). (F) Volcano plot highlighting the fold change and statistical significance of the prioritized genes (FNDC1, LPCAT3, and TIMP2). (G) Circos plot depicting the chromosomal localization and genomic structural features of the three candidate genes. (*p < 0.05, **p < 0.01, ***p < 0.001).

Validation across integrated normalized expression profiles confirmed the significant upregulation of FNDC1 and TIMP2 (P < 0.001) and the downregulation of LPCAT3 (P < 0.001) in HF tissues across the training set (Fig. 5B) and internal testing set (Fig. 5C). In the independent external cohort (Fig. 5D), FNDC1 and LPCAT3 remained significantly altered (P < 0.05), whereas TIMP2 showed no statistically significant difference (P > 0.05). These patterns were further corroborated by RT-qPCR in the local clinical cohort (Fig. 5E), where all three genes achieved statistical significance. The global transcriptomic landscape visualized via a volcano plot further highlighted the distinct positioning of these three genes, exhibiting substantial fold changes and high statistical significance relative to the myocardial background (Fig. 5F). Furthermore, a Circos plot was constructed to map the chromosomal distribution of these candidates, revealing their specific genomic coordinates across the human genome (Fig. 5G).

Machine learning model performance evaluation and SHAP analysis of key feature genes

A 10-model ensemble system was implemented to evaluate the diagnostic stability of the three-gene signature. In the internal discovery cohort, all models exhibited robust performance under Leave-One-Study-Out Cross-Validation (LOSO-CV), with the glmBoost algorithm achieving the highest predictive accuracy (AUC = 0.973, 95% CI: 0.96–0.99), followed by PLS (AUC = 0.969) and SVM (AUC = 0.969) (Fig. 6A). To assess cross-dataset generalizability, the models were applied to an entirely independent external cohort (GSE135055). Peak performance was maintained by the KNN (AUC = 0.876) and PLS (AUC = 0.812) models (Fig. 6B). Permutation tests (1,000 iterations) statistically validated these findings. The actual AUCs for the top three models (GLMBOOST: 0.978; PLS: 0.978; SVM: 0.983) significantly outperformed the random null distributions (all empirical p < 0.01, Supplementary Fig. 5), confirming their robust predictive power is driven by biological signals rather than random chance.

Fig. 6.

Fig. 6

SHAP analysis for feature interpretation. (A) ROC curves comparing the performance of 10 machine learning models, including. (B) SHAP summary plot showing the distribution of SHAP values for FNDC1, LPCAT3, and TIMP2 across samples. (C) Bar plot representing the mean absolute SHAP values for FNDC1, LPCAT3, and TIMP2. (D) SHAP dependence plots for FNDC1, LPCAT3, and TIMP2.

Model interpretability was further elucidated through SHAP analysis. The SHAP summary plot prioritized FNDC1 as the most influential feature (mean |SHAP| = 0.293), followed by LPCAT3 (0.175) and TIMP2 (0.024) (Fig. 6C). High expression levels of FNDC1 (orange dots) were positively correlated with higher SHAP values, indicating a strong contribution to HF prediction, whereas LPCAT3 exhibited a negative correlation, consistent with its downregulated profile in HF tissues (Fig. 6C). SHAP dependence plots further visualized the individual marginal effects and interaction patterns of these three genes on model output (Fig. 6D). Additionally, a SHAP waterfall plot (Supplementary Fig. 2) illustrates the quantitative contribution of individual features (FNDC1 = 8.24 contributing − 0.335) to specific diagnostic predictions, confirming the model’s granular interpretability.

GSVA and GSEA enrichment analysis based on high- and low-expression groups of key feature genes

Gene Set Variation Analysis (GSVA) and Gene Set Enrichment Analysis (GSEA) were employed to investigate the biological pathways associated with the differential expression of core genes. Pathway scores for KEGG gene sets (restricted to sizes of 10–500 genes) were calculated via GSVA. Comparisons between high- and low-expression groups for each biomarker (defined by median split) were performed using the Benjamini-Hochberg (BH) method to adjust for multiple testing, with a significance threshold set at a FDR < 0.05.

As shown in Fig. 7, high expression of FNDC1 was significantly associated with the activation of extracellular matrix organization, focal adhesion, and Wnt signaling pathways (Fig. 7A). In contrast, the LPCAT3 high-expression group exhibited enrichment in metabolic signatures, including fatty acid metabolism and the PPAR signaling pathway (Fig. 7B). TIMP2 high expression was markedly linked to ECM-receptor interaction and TGF-β signaling (Fig. 7C). These findings were further corroborated by GSEA results, which demonstrated consistent enrichment patterns across the identified signatures (Supplementary Fig. 3). The observed enrichment patterns in high-expression groups provide additional insights into the potential mechanistic involvement of these biomarkers in HF-related pathological processes.

Fig. 7.

Fig. 7

GSVA analysis of FNDC1, LPCAT3, and TIMP2 across various biological pathways. (red for upregulated, blue for downregulated). (A) FNDC1: GSVA scores for FNDC1 across multiple KEGG pathways. (B) LPCAT3: GSVA scores for LPCAT3 across different KEGG pathways. (C) TIMP2: GSVA scores for TIMP2 across different KEGG pathways.

Immune relative proportions analysis based on integrated normalized expression profiles

The differences in immune cell relative proportions within the integrated normalized HF expression profile dataset were analyzed to assess immune cell relative proportions patterns in HF and their correlation with gene expression. Figure 8A displays the relative percentages of various immune cell types in control and HF groups, revealing significant alterations in the proportions of multiple immune cells in HF, particularly among macrophage and T-cell subsets. A heatmap of immune cell interactions was constructed using Spearman’s rank correlation coefficient (Fig. 8B), where the x- and y-axes represent 22 immune cell subtypes, and the color intensity of matrix blocks corresponds to correlation coefficient values. Box plots in Fig. 8C demonstrate differential relative proportions of specific immune cell types between control and HF groups, with nine immune cell types showing significant changes in HF. The correlation matrix in Fig. 8D indicates that TIMP2 expression positively correlates with M2 macrophage relative proportions (p < 0.01), while FNDC1 shows a negative correlation with CD8 + T cells (p < 0.05). It is important to note that these statistical correlations describe monotonic relationships between gene expression and cell fractions but do not demonstrate directionality or mechanistic causality.

Fig. 8.

Fig. 8

Immune relative proportions analysis. (A) Stacked bar chart for the relative proportions of various immune cell types between the HF group and the control group. (B) Heatmap depicting the correlation between different immune cell types, with colors ranging from blue (negative correlation) to red (positive correlation). (C) Box plot comparing the relative proportions of specific immune cell types between the HF group and the control group. (*p < 0.05, **p < 0.01, ***p < 0.001). (D) Correlation heatmap of the relationship between the expression of FNDC1, LPCAT3, and TIMP2 genes and various immune cell types. (*p < 0.05, **p < 0.01, ***p < 0.001)

Discussion

HF, a major global public health challenge, is pathologically characterized by a vicious cycle of progressive cardiac structural and functional decompensation. The core pathophysiological features of HF involve cardiomyocyte loss, pathological fibrosis, and chronic inflammation-induced cardiac remodeling41,42. This remodeling, which in many clinical contexts can be progressive, is driven by multidimensional interactions: cardiomyocyte apoptosis directly impairs contractile function, while fibrotic processes may increase ventricular stiffness through ECM deposition, collectively contributing to hemodynamic overload. Notably, evidence suggests that chronic inflammation in this context is not merely an epiphenomenon but can actively contribute to remodeling through the sustained release of pro-inflammatory cytokines (e.g., TNF-α, IL-6), which may activate fibroblast-to-myofibroblast transition and accelerate ECM remodeling43,44.

Although high-throughput omics technologies have unveiled the complexity of HF-associated molecular networks, distinguishing driver signals from compensatory responses in vast datasets remains a bottleneck for developing precision therapies. Emerging paradigms that integrate multi-omics data with machine learning algorithms now provide powerful tools to dissect the dynamic interplay among cardiomyocyte depletion, immune dysregulation, and fibrosis, offering novel insights into HF pathophysiology.

Machine Learning technologies have been successfully applied to early diagnostic biomarker discovery45, risk stratification (e.g., ECG feature extraction for sudden death risk prediction)46,47, and treatment response prediction48. This study integrated three machine learning algorithms—RF, Support Vector Machine, and XGBoost to identify FNDC1, LPCAT3, and TIMP2 as potential core feature genes for HF diagnosis. Notably, integrating SVM-based algorithms with rigorous statistical evaluations, such as the sigFeature framework, has been shown to significantly enhance both classification performance and the discovery of biologically relevant biomarkers49. Functional enrichment analysis of these genes revealed their potential involvement in multiple pathological mechanisms of HF.

LPCAT3 (Lysophosphatidylcholine Acyltransferase 3 ), a member of the lysophospholipid acyltransferase family, specifically regulates the abundance of distinct phosphatidylcholine species in various cells and tissues, playing a critical role in lipid metabolism and homeostasis50. LPCAT3 catalyzes the esterification of lysophospholipids and participates in ferroptosis regulation51–53. Studies indicate upregulated LPCAT3 expression in monocytes, B lymphocytes, CD8 + T cells, NK cells, CD4 + T cells, and lymph nodes54. HF, ferroptosis contributes to pathological progression55–57, and metabolic disorders severely impact cardiac function, suggest that LPCAT3 represents a potential candidate for further mechanistic investigation. LPCAT3 accounts for the majority of LPCAT activity in certain cell types58,59, and its reduced activity disrupts Golgi structure and limits plasma membrane extension and cell proliferation. LPCAT3 is a key modulator of free arachidonic acid levels, which are linked to atherosclerosis development60,61. Altered AA-PC and LPCAT3 levels in atherosclerotic tissues may drive changes in free AA. Additionally, LPCAT3 is highly expressed in vascular smooth muscle cells and plays a significant role in atherosclerosis progression62. These findings highlight LPCAT3 as a pivotal regulator of lipid metabolism and ferroptosis, warranting further investigation into its potential as a novel biomarker and therapeutic target in HF pathogenesis.

FNDC1 (Fibronectin Type III Domain Containing 1), a member of the fibronectin superfamily, plays critical roles in ECM interactions through its type III domain in various pathological processes63,64. This protein regulates angiogenesis by modulating the subcellular localization and kinase activity of the VEGFR2 receptor, a mechanism particularly prominent in hypoxic microenvironments. Animal studies have confirmed that myocardial ischemia-reperfusion injury induces FNDC1 upregulation65–67. Clinical genomic analyses demonstrate broad associations between FNDC1 genetic polymorphisms and human diseases: the rs3003174 variant correlates with coronary artery lesion risk in Kawasaki disease, and its mutation profile suggests involvement in hypertension and pediatric acute otitis media68,69. In cardiovascular research, recent evidence indicates that FNDC1 regulates myocardial cell hypoxia adaptation by remodeling G protein-coupled signaling pathways and acts as an ECM-targeted mediator influencing ECM remodeling67,70,71. Studies also confirm that hypoxia stress induces FNDC1 mRNA upregulation in rat aortic cells, aligning with our findings and suggesting its role as a molecular sensor of tissue microenvironmental changes70. Collectively, FNDC1 demonstrates potential as a high-priority candidate for future functional validation in HF pathogenesis through its integrated regulation of ECM remodeling, G protein signaling, and hypoxia microenvironment responses.

TIMP2 (TIMP Metallopeptidase Inhibitor 2), a natural inhibitor of matrix metalloproteinases (MMPs), regulates ECM remodeling. As a key modulator of MMPs, TIMP2 maintains ECM homeostasis through dual mechanisms: directly inhibiting the proteolytic activity of MMP-2/9 to prevent abnormal degradation of ECM components (e.g., collagen and elastin), and precisely controlling zymogen activation by forming the MT1-MMP/TIMP2/MMP-2 ternary complex, a “molecular switch” critical for tissue repair and pathological fibrosis. Reduced circulating TIMP-2 levels have been significantly associated with increased atrial fibrillation risk72. In atherosclerosis, unstable plaques exhibit higher MMP-2 expression compared to stable plaques72. TIMP2-mediated regulation of MMP-2 activity, which is selectively inhibited by TIMP1, modulates inflammatory responses41. Given that ECM remodeling, myocardial fibrosis, and inflammation are central pathophysiological processes in HF, TIMP2 likely plays a pivotal role in post-HF ECM remodeling and inflammatory cascades.

Recent studies have suggested that the immune regulatory network may play an important role in the pathological remodeling of the heart. As an immunologically active organ, the chronic inflammatory state and the continuous activation of immune cells in the heart are regarded as potential factors associated with the development of HF. In this study, exploratory immune relative proportions analysis indicated that the myocardial microenvironment may exhibit certain adaptive immune response characteristics, with T cell subsets such as CD4 memory resting cells, T follicular helper cells, and regulatory T cells (Tregs) being significantly enriched (FDR < 0.05), pointing toward a potential association between T lymphocyte-mediated responses and the disease process. Meanwhile, M2 - type macrophages were also significantly enriched in HR, possibly exacerbating myocardial fibrosis through multiple pathways. As key effector cells in anti - inflammatory/repair - related immune regulation, M2 - type macrophages maintain tissue homeostasis by secreting key factors such as IL − 10 and TGF - β. The loss of their function leads to dual pathological effects: on the one hand, the anti - inflammatory/ pro - inflammatory balance shifts towards the pro - inflammatory phenotype dominated by M1 - type macrophages, promoting the release of large amounts of pro - inflammatory factors such as TNF - α and IL − 6, and exacerbating myocardial inflammatory damage. Pro - inflammatory factors such as TNF - α, IFN - γ, and IL − 6 are secreted to form an inflammatory microenvironment, and the immune response is promoted to the adaptive stage. This cascading inflammatory response ultimately leads to a vicious cycle of myocardial fibrosis and ventricular remodeling. This exploratory evidence suggests that the chronicization of cardiac remodeling might involve immune activation, adaptive immune maintenance, and regulatory immune failure, which aligns with our preliminary results. However, we emphasize that these Spearman correlations are purely associative. These findings do not provide evidence of a causal relationship—for instance, whether upregulated FNDC1 actively recruits specific T-cells or if the altered immune landscape itself modulates gene expression remains to be elucidated. Furthermore, the lack of comprehensive clinical metadata in the public datasets limits our ability to control for all potential confounding factors. These results should therefore be viewed as biological clues that warrant further investigation using targeted co-culture experiments or longitudinal clinical cohorts.

Machine-learning-based biomarker screening systems face several technical bottlenecks in medical research. The core challenge arises from the interaction between algorithm characteristics and medical data features. First, the predictive nature of machine learning models makes them highly sensitive to data distribution. Limited data scales may amplify the impact of noise signals in the dataset, leading models to focus on sample-specific features rather than universal biological laws. This is particularly evident in non-parametric models like decision trees, which may generate overly complex classification rules (e.g., a single decision path tailored to a few abnormal samples), thereby weakening the model’s clinical extrapolation ability. To address this, this study implemented a dual-validation mechanism: (1) A K-fold cross-validation strategy (K = 10) was adopted. By performing multiple data splits and model reconstructions, this approach ensured the stability of feature importance evaluation and effectively reduced the dominant influence of outliers on the classifier. (2) A fixed random seed was set to control the reproducibility of the data splitting process, maintaining a balanced class distribution in both the training and validation sets. These measures supported model stabilityand facilitated robust evaluation across independent datasets. While internal validation techniques ensure algorithmic reproducibility, the ultimate generalizability of our findings was rigorously confirmed through the high diagnostic performance maintained in the independent external validation cohort (GSE135055). However, it should be noted that small sample sizes may still limit the biological interpretability of biomarker selection. Furthermore, while SHAP analysis provides insights into the algorithm’s decision-making process, it is crucial to recognize that it explains model behavior (statistical attribution) rather than biological causality. The feature importance identified reflects associative patterns within the dataset; thus, these genes should be regarded as high-probability candidates that require further experimental validation to confirm their mechanistic roles in HF pathology. Future studies need to further confirm the pathological relevance of key genes through multicenter cohort validation and in vitro functional experiments. The number and diversity of omics-based approaches are constantly expanding, and the convenience and affordability of biometric-based valuations have fueled the growing demand for integrating and meaningfully interpreting large and diverse patient datasets.

As multi-omics technologies and bioinformatics tools continue to evolve, the application of machine learning in this study provides additional insights into the molecular landscape of HF. However, it is important to contextualize the potential clinical utility of these findings. Since our analysis was primarily based on transcriptomic data from left ventricular myocardial tissues—frequently sourced from end-stage explanted hearts or post-mortem examinations—the identified genes (FNDC1, LPCAT3, and TIMP2) should be interpreted with caution. Furthermore a fundamental mismatch exists between the diagnostic goals of this study and the biological source material. Our primary markers were identified and validated in myocardial tissue (left ventricle), which is generally inaccessible for routine clinical screening or early diagnosis. They are perhaps better characterized as candidates for understanding disease mechanisms rather than as conventional diagnostic biomarkers for routine clinical screening in the general population. Their relevance may be more pronounced in specific clinical settings, such as the analysis of EMB samples, where they could potentially serve as molecular indicators to help evaluate the progression of fibrosis or immune-metabolic changes.Furthermore, these tissue-level signatures may offer a useful reference for future translational research. Subsequent studies are needed to investigate whether these intramyocardial changes correlate with biomarkers that can be measured through less invasive means, such as peripheral blood analysis. By integrating emerging techniques like single-cell sequencing or spatial omics, future research might further clarify the roles of these genes within the cardiac microenvironment. Ultimately, these findings contribute to a more detailed understanding of HFpathophysiology and may support the ongoing efforts to develop more targeted strategies for managing myocardial remodeling.

Conclusions

In this study, multi-model machine learning combined with SHAP analysis identified FNDC1, LPCAT3, and TIMP2 as candidate biomarkers for HF. Pathway analysis suggested their involvement in matrix remodeling and immune-metabolic imbalance, although these were only correlational findings; our stable and efficient prediction model (mean AUC = 0.965) warrants further functional and clinical validation.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (842.8KB, pdf)

Acknowledgements

We are very grateful to the GEO database for providing free data access.

Author contributions

YHZ, RZ, KZ and YZ contributed to study design. YHZ, RZ, KZ, YL, LH and YW extracted and analysed the data. All authors drafted the manuscript. YHZ, SD, and YZ contributed to a critical revision of the manuscript. All authors approved the submitted version.

Data availability

The bulk-seq datasets for heart failure are fully available in the GEO database (GSE141910: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE141910; GSE198945: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE198945; GSE263297: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE263297; GSE135055: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135055).

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yuhe Zhao, Ruoyu Zhang and Kelan Zha.

Contributor Information

Shuren Dai, Email: 422154191@qq.com.

Yu Zeng, Email: 402828201@qq.com.

References

  • 1.Trends in survival after a diagnosis of heart failure in the United Kingdom 2000–2017: population based cohort study. BMJ367, l5840. 10.1136/bmj.l5840 (2019). [DOI] [PMC free article] [PubMed]
  • 2.Roth, G. A. et al. Global Burden of Cardiovascular Diseases and Risk Factors, 1990–2019: Update From the GBD 2019 Study. J. Am. Coll. Cardiol.76, 2982–3021. 10.1016/j.jacc.2020.11.010 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Tong, Z., Xie, Y., Li, K., Yuan, R. & Zhang, L. The global burden and risk factors of cardiovascular diseases in adolescent and young adults, 1990–2019. BMC Public. Health. 24, 1017. 10.1186/s12889-024-18445-6 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Chakraborty, S., Dutta, A., Roy, A., Joshi, A. & Basak, T. The Theatrics of Collagens in the Myocardium: The Supreme Architect of the Fibrotic Heart. Am. J. Physiol. Cell. Physiol.10.1152/ajpcell.01043.2024 (2025). [DOI] [PubMed] [Google Scholar]
  • 5.Tikhomirov, R. et al. Exosomes: From potential culprits to new therapeutic promise in the setting of cardiac fibrosis. Cells9. 10.3390/cells9030592 (2020). [DOI] [PMC free article] [PubMed]
  • 6.Chimura, M. et al. Comprehensive Analysis of the Effects of Sacubitril/Valsartan According to Sex Among Patients With Heart Failure and Reduced Ejection Fraction in PARADIGM-HF. J. Am. Heart Assoc. (e038249). 10.1161/jaha.124.038249 (2025). [DOI] [PMC free article] [PubMed]
  • 7.Zhu, T. & Song, Y. Efficacy of Sacubitril/Valsartan Combined With Metoprolol on Cardiac Function, Cardiac Remodeling, and Endothelial Function in Patients With Coronary Heart Disease and Heart Failure. Br. J. Hosp. Med. (Lond). 86, 1–16. 10.12968/hmed.2025.0120 (2025). [DOI] [PubMed] [Google Scholar]
  • 8.Shin, J. & Johnson, J. A. Beta-blocker pharmacogenetics in heart failure. Heart Fail. Rev.15, 187–196. 10.1007/s10741-008-9094-x (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Senyo, S. E. et al. Mammalian heart renewal by pre-existing cardiomyocytes. Nature493, 433–436. 10.1038/nature11682 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Bergmann, O. et al. Evidence for cardiomyocyte renewal in humans. Science324, 98–102. 10.1126/science.1164680 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Sun, J. et al. VASN knockout induces myocardial fibrosis in mice by downregulating non-collagen fibers and promoting inflammation. Front. Pharmacol.15, 1500617. 10.3389/fphar.2024.1500617 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Armoundas, A. A. et al. Use of Artificial Intelligence in Improving Outcomes in Heart Disease: A Scientific Statement From the American Heart Association. Circulation149, e1028–e1050. 10.1161/cir.0000000000001201 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Yamasan, B. E. & Korkmaz, S. Unveiling miRNA biomarkers for hypertrophic cardiomyopathy through integrated bioinformatics and machine learning analysis. Acta Cardiol.80, 1149–1162. 10.1080/00015385.2025.2577015 (2025). [DOI] [PubMed] [Google Scholar]
  • 14.Kokori, E. et al. Machine learning in predicting heart failure survival: a review of current models and future prospects. Heart Fail. Rev.30, 431–442. 10.1007/s10741-024-10474-y (2025). [DOI] [PubMed] [Google Scholar]
  • 15.Yasmin, F. et al. Artificial intelligence in the diagnosis and detection of heart failure: the past, present, and future. Rev. Cardiovasc. Med.22, 1095–1113. 10.31083/j.rcm2204121 (2021). [DOI] [PubMed] [Google Scholar]
  • 16.Unterhuber, M. et al. Deep learning detects heart failure with preserved ejection fraction using a baseline electrocardiogram. Eur. Heart J. Digit. Health. 2, 699–703. 10.1093/ehjdh/ztab081 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wang, Z. et al. Using Deep Learning to Identify High-Risk Patients with Heart Failure with Reduced Ejection Fraction. J. Health Econ. Outcomes Res.8, 6–13. 10.36469/jheor.2021.25753 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Akram, M. J. et al. Machine learning based insights into cardiomyopathy and heart failure research: a bibliometric analysis from 2005 to 2024. Front. Med. (Lausanne). 12, 1602077. 10.3389/fmed.2025.1602077 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Draelos, R. L. et al. Gene-Specific Machine Learning Models for Variants of Uncertain Significance Found in Catecholaminergic Polymorphic Ventricular Tachycardia and Long QT Syndrome-Associated Genes. Circ. Arrhythm. Electrophysiol.15, e010326. 10.1161/circep.121.010326 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Zhang, P. et al. Identification of shared molecular mechanisms and diagnostic biomarkers between heart failure and idiopathic pulmonary fibrosis. Heliyon10, e30086. 10.1016/j.heliyon.2024.e30086 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wang, Z. et al. Construction of machine learning diagnostic models for cardiovascular pan-disease based on blood routine and biochemical detection data. Cardiovasc. Diabetol.23, 351. 10.1186/s12933-024-02439-0 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Lundberg, S. M. & Lee, S. I. In Proceedings of the 31st International Conference on Neural Information Processing Systems 4768–4777Curran Associates Inc., Long Beach, California, USA, (2017).
  • 23.Doan, K. V. et al. Cardiac NAD(+) depletion in mice promotes hypertrophic cardiomyopathy and arrhythmias prior to impaired bioenergetics. Nat. Cardiovasc. Res.3, 1236–1248. 10.1038/s44161-024-00542-9 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Madè, A. et al. circrna-mirna-mrna Deregulated Network in Ischemic Heart Failure Patients. Cells 12. 10.3390/cells12212578 (2023). [DOI] [PMC free article] [PubMed]
  • 25.Hua, X. et al. Multi-level transcriptome sequencing identifies COL1A1 as a candidate marker in human heart failure progression. BMC Med.18, 2. 10.1186/s12916-019-1469-4 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Davis, S. & Meltzer, P. S. GEOquery: a bridge between the Gene Expression Omnibus (GEO) and BioConductor. Bioinformatics23, 1846–1847. 10.1093/bioinformatics/btm254 (2007). [DOI] [PubMed] [Google Scholar]
  • 27.Barrett, T. et al. NCBI GEO: mining tens of millions of expression profiles–database and tools update. Nucleic Acids Res.35, D760–765. 10.1093/nar/gkl887 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Leek, J. T., Johnson, W. E., Parker, H. S., Jaffe, A. E. & Storey, J. D. The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics28, 882–883. 10.1093/bioinformatics/bts034 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wu, Y. et al. iLSGRN: inference of large-scale gene regulatory networks based on multi-model fusion. Bioinformatics3910.1093/bioinformatics/btad619 (2023). [DOI] [PMC free article] [PubMed]
  • 30.Cai, W. & van der Laan, M. Nonparametric bootstrap inference for the targeted highly adaptive least absolute shrinkage and selection operator (LASSO) estimator. Int. J. Biostat. 10.1515/ijb-2017-0070 (2020). [DOI] [PubMed] [Google Scholar]
  • 31.Engebretsen, S. & Bohlin, J. Statistical predictions with glmnet. Clin. Epigenetics. 11, 123. 10.1186/s13148-019-0730-1 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Liu, Y. & Zhao, H. Variable importance-weighted Random Forests. Quant. Biol.5, 338–351 (2017). [PMC free article] [PubMed] [Google Scholar]
  • 33.Zhao, Y. et al. A Literature Review of Gene Function Prediction by Modeling Gene Ontology. Front. Genet.11, 400. 10.3389/fgene.2020.00400 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kanehisa, M. & Goto, S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res.28, 27–30. 10.1093/nar/28.1.27 (2000). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Yu, G., Wang, L. G., Han, Y. & He, Q. Y. clusterProfiler: an R package for comparing biological themes among gene clusters. Omics16, 284–287. 10.1089/omi.2011.0118 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Subramanian, A. et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc. Natl. Acad. Sci. U S A. 102, 15545–15550. 10.1073/pnas.0506580102 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Hänzelmann, S., Castelo, R. & Guinney, J. GSVA: gene set variation analysis for microarray and RNA-seq data. BMC Bioinform.14, 7. 10.1186/1471-2105-14-7 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Kuhn, M. Building Predictive Models in R Using the caret Package. J. Stat. Softw.28, 1–26. 10.18637/jss.v028.i05 (2008).27774042 [Google Scholar]
  • 39.Newman, A. M. et al. Robust enumeration of cell subsets from tissue expression profiles. Nat. Methods. 12, 453–457. 10.1038/nmeth.3337 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Choi, H. & Na, K. J. Integrative analysis of imaging and transcriptomic data of the immune landscape associated with tumor metabolism in lung adenocarcinoma: Clinical and prognostic implications. Theranostics8, 1956–1965. 10.7150/thno.23767 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Kim, I. S., Yang, W. S. & Kim, C. H. Physiological Properties, Functions, and Trends in the Matrix Metalloproteinase Inhibitors in Inflammation-Mediated Human Diseases. Curr. Med. Chem.30, 2075–2112. 10.2174/0929867329666220823112731 (2023). [DOI] [PubMed] [Google Scholar]
  • 42.Kong, P., Christia, P. & Frangogiannis, N. G. The pathogenesis of cardiac fibrosis. Cell. Mol. Life Sci.71, 549–574. 10.1007/s00018-013-1349-6 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Packer, M. et al. Chronic Kidney Disease in Patients With Heart Failure With a Preserved Ejection Fraction: The Underlying Role of Visceral Adiposity. J. Am. Coll. Cardiol.86, 1900–1916. 10.1016/j.jacc.2025.08.086 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Van Linthout, S., Miteva, K. & Tschöpe, C. Crosstalk between fibroblasts and inflammatory cells. Cardiovasc. Res.102, 258–269. 10.1093/cvr/cvu062 (2014). [DOI] [PubMed] [Google Scholar]
  • 45.Akerman, A. P. et al. External validation of artificial intelligence for detection of heart failure with preserved ejection fraction. Nat. Commun.16, 2915. 10.1038/s41467-025-58283-7 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Moroz, H., Li, Y., Marelli, A. & hART Deep learning-informed lifespan heart failure risk trajectories. Int. J. Med. Inf.185, 105384. 10.1016/j.ijmedinf.2024.105384 (2024). [DOI] [PubMed] [Google Scholar]
  • 47.Huang, C. C. et al. Examining arterial pulsation to identify and risk-stratify heart failure subjects with deep neural network. Phys. Eng. Sci. Med.47, 477–489. 10.1007/s13246-023-01378-6 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Nirschl, J. J. et al. A deep-learning classifier identifies patients with clinical heart failure using whole-slide images of H&E tissue. PLoS One. 13, e0192726. 10.1371/journal.pone.0192726 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Das, P., Roychowdhury, A., Das, S., Roychoudhury, S. & Tripathy, S. sigFeature: Novel Significant Feature Selection Method for Classification of Gene Expression Data Using Support Vector Machine and t Statistic. Front. Genet.11, 247. 10.3389/fgene.2020.00247 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Shao, G. et al. Research progress in the role and mechanism of LPCAT3 in metabolic related diseases and cancer. J. Cancer. 13, 2430–2439. 10.7150/jca.71619 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Kagan, V. E. et al. Oxidized arachidonic and adrenic PEs navigate cells to ferroptosis. Nat. Chem. Biol.13, 81–90. 10.1038/nchembio.2238 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Doll, S. et al. ACSL4 dictates ferroptosis sensitivity by shaping cellular lipid composition. Nat. Chem. Biol.13, 91–98. 10.1038/nchembio.2239 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Reichert, C. O. et al. Ferroptosis mechanisms involved in neurodegenerative diseases. Int. J. Mol. Sci.2110.3390/ijms21228765 (2020). [DOI] [PMC free article] [PubMed]
  • 54.Viswanathan, V. S. et al. Dependency of a therapy-resistant state of cancer cells on a lipid peroxidase pathway. Nature547, 453–457. 10.1038/nature23007 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Li, Q. et al. Ferroptosis: The potential target in heart failure with preserved ejection fraction. Cells1110.3390/cells11182842 (2022). [DOI] [PMC free article] [PubMed]
  • 56.Kitakata, H. et al. Imeglimin prevents heart failure with preserved ejection fraction by recovering the impaired unfolded protein response in mice subjected to cardiometabolic stress. Biochem. Biophys. Res. Commun.572, 185–190. 10.1016/j.bbrc.2021.07.090 (2021). [DOI] [PubMed] [Google Scholar]
  • 57.Ma, S. et al. Canagliflozin mitigates ferroptosis and ameliorates heart failure in rats with preserved ejection fraction. Naunyn Schmiedebergs Arch. Pharmacol.395, 945–962. 10.1007/s00210-022-02243-1 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Zhao, Y. et al. Identification and characterization of a major liver lysophosphatidylcholine acyltransferase. J. Biol. Chem.283, 8258–8265. 10.1074/jbc.M710422200 (2008). [DOI] [PubMed] [Google Scholar]
  • 59.Hishikawa, D. et al. Discovery of a lysophospholipid acyltransferase family essential for membrane asymmetry and diversity. Proc. Natl. Acad. Sci. U S A. 105, 2830–2835. 10.1073/pnas.0712245105 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Pérez-Chacón, G., Astudillo, A. M., Balgoma, D., Balboa, M. A. & Balsinde, J. Control of free arachidonic acid levels by phospholipases A2 and lysophospholipid acyltransferases. Biochim. Biophys. Acta. 1791, 1103–1113. 10.1016/j.bbalip.2009.08.007 (2009). [DOI] [PubMed] [Google Scholar]
  • 61.Pérez-Chacón, G., Astudillo, A. M., Ruipérez, V., Balboa, M. A. & Balsinde, J. Signaling role for lysophosphatidylcholine acyltransferase 3 in receptor-regulated arachidonic acid reacylation reactions in human monocytes. J. Immunol.184, 1071–1078. 10.4049/jimmunol.0902257 (2010). [DOI] [PubMed] [Google Scholar]
  • 62.Tanaka, H. et al. Lysophosphatidylcholine Acyltransferase-3 Expression Is Associated with Atherosclerosis Progression. J. Vasc Res.54, 200–208. 10.1159/000473879 (2017). [DOI] [PubMed] [Google Scholar]
  • 63.Bentmann, A. et al. Circulating fibronectin affects bone matrix, whereas osteoblast fibronectin modulates osteoblast function. J. Bone Min. Res.25, 706–715. 10.1359/jbmr.091011 (2010). [DOI] [PubMed] [Google Scholar]
  • 64.Gao, M. et al. Structure and functional significance of mechanically unfolded fibronectin type III1 intermediates. Proc. Natl. Acad. Sci. U S A. 100, 14784–14789. 10.1073/pnas.2334390100 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Hayashi, H., Al Mamun, A., Sakima, M. & Sato, M. Activator of G-protein signaling 8 is involved in VEGF-mediated signal processing during angiogenesis. J. Cell. Sci.129, 1210–1222. 10.1242/jcs.181883 (2016). [DOI] [PubMed] [Google Scholar]
  • 66.Shibuya, M. Vascular Endothelial Growth Factor (VEGF) and Its Receptor (VEGFR) Signaling in Angiogenesis: A Crucial Target for Anti- and Pro-Angiogenic Therapies. Genes Cancer. 2, 1097–1105. 10.1177/1947601911423031 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Sato, M. et al. Identification of a receptor-independent activator of G protein signaling (AGS8) in ischemic heart and its interaction with Gbetagamma. Proc. Natl. Acad. Sci. U S A. 103, 797–802. 10.1073/pnas.0507467103 (2006). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.van Ingen, G. et al. Genome-wide association study for acute otitis media in children identifies FNDC1 as disease contributing gene. Nat. Commun.7, 12792. 10.1038/ncomms12792 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Lin, K. et al. FNDC1 Polymorphism (rs3003174 C > T) Increased the Incidence of Coronary Artery Aneurysm in Patients with Kawasaki Disease in a Southern Chinese Population. J. Inflamm. Res.14, 2633–2640. 10.2147/jir.S311956 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Yuan, C., Sato, M., Lanier, S. M. & Smrcka, A. V. Signaling by a non-dissociated complex of G protein βγ and α subunits stimulated by a receptor-independent activator of G protein signaling, AGS8. J. Biol. Chem.282, 19938–19947. 10.1074/jbc.M700396200 (2007). [DOI] [PubMed] [Google Scholar]
  • 71.Bouchareb, R. et al. Proteomic Architecture of Valvular Extracellular Matrix: FNDC1 and MXRA5 Are New Biomarkers of Aortic Stenosis. JACC Basic. Transl Sci.6, 25–39. 10.1016/j.jacbts.2020.11.008 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Scholtes, V. P. et al. Carotid atherosclerotic plaque matrix metalloproteinase-12-positive macrophage subpopulation predicts adverse outcome after endarterectomy. J. Am. Heart Assoc.1, e001040. 10.1161/jaha.112.001040 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (842.8KB, pdf)

Data Availability Statement

The bulk-seq datasets for heart failure are fully available in the GEO database (GSE141910: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE141910; GSE198945: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE198945; GSE263297: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE263297; GSE135055: https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135055).


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES