Skip to main content
Alzheimer's Research & Therapy logoLink to Alzheimer's Research & Therapy
. 2025 Jul 4;17:147. doi: 10.1186/s13195-025-01789-5

Machine-learning based strategy identifies a robust protein biomarker panel for Alzheimer’s disease in cerebrospinal fluid

Xiaosen Hou 1,#, Yunjie Qiu 1,#, Hui Li 1,2, Yan Yan 1,4, Dongxu Zhao 1, Simei Ji 2, Junjun Ni 1, Jun Zhang 1, Kefu Liu 3,, Hong Qing 1,2,, Zhenzhen Quan 1,
PMCID: PMC12232211  PMID: 40616179

Abstract

Background

The complex pathogenesis of Alzheimer’s disease (AD) has resulted in limited current biomarkers for its classification and diagnosis, necessitating further investigation into reliable universal biomarkers or combinations.

Methods

In this work, we collect multiple CSF proteomics datasets and build a universal diagnose model by SVM-RFECV method combined with equal sample size and standard normalization design. The model was training in 297_CSF and then test the effect in other datasets.

Results

Utilizing machine learning, we identify a 12-protein panel from cerebrospinal fluid proteomic datasets. The universal diagnosis model demonstrated strong diagnostic capability and high accuracy across ten different AD cohorts across different countries and different detection technologies. These proteins involved in various biological processes related to AD and shows a tight correlation with established AD pathogenic biomarkers, including amyloid-β, tau/p-tau, and the Montreal Cognitive Assessment score. The high accuracy in the model may due to multiple protein combination based on comprehensive pathogenesis and different AD progress. Furthermore, it effectively differentiates AD from mild cognitive impairment (MCI) and other neurodegenerative disorders, especially the frontotemporal dementia (FTD), which share similar pathogenesis as AD.

Conclusion

This study highlights a high accuracy, robustness and compatibility model of 12-protein panel whose detection is even based on label-free, TMT and DIA mass spectrometry or ELISA technologies, implicating its potential prospect in clinical application.

Supplementary Information

The online version contains supplementary material available at 10.1186/s13195-025-01789-5.

Keywords: Alzheimer’s disease, Machine learning, Cerebrospinal fluid, Biomarker

Introduction

Alzheimer’s disease (AD) is recognized as one of the most prevalent neurodegenerative disorders, primarily characterized by two classical pathological hallmarks: amyloid-β (Aβ) plaques and neurofibrillary tangles [1]. As global life expectancy continues to rise, the number of individuals with AD increases annually; however, a large percentage of elderly individuals worldwide exhibit accumulation of extensive amyloid pathology but maintain intact cognition [2, 3]. Cerebrospinal fluid (CSF) is regarded as the optimal source for identifying potential AD biomarkers, owing to its direct connection to the brain’s extracellular space and its ability to transport proteins associated with neurodegenerative conditions. It offers high diagnostic utility, clinical validation, and excellent analytical performance [4]. According to the pathogenesis of AD, the proteins that can reflect the total tau (t-tau), phosphorylated tau (p-tau), Aβ42 and Aβ42/40 ratio are initially identified as “core” CSF biomarkers for AD [5, 6]. Despite these advancements, current CSF biomarkers are limited in many respects due to the heterogeneous nature of AD’s pathogenesis and varied patient susceptibility. Challenges include inadequate assessments of AD progression, particularly in the preclinical stage, difficulties in distinguishing biological variations among AD patients, and shortcomings in facilitating therapeutic drug development [7]. Therefore, there is an urgent need to explore CSF datasets more comprehensively and identify more effective pathogenic factors for AD diagnosis and treatment in clinical practice as well as drug discovery.

While numerous studies have utilized proteomics to identify cerebrospinal fluid (CSF) biomarkers, the integration of machine learning (ML) techniques for disease diagnosis represents a recent advancement. Notably, the support vector machine (SVM) is a supervised ML method that excels in classification and subtyping tasks and has recently been applied in cancer genomics and cancer classification [811]. A recent study reported that a combination model developed through machine learning can be used to predict AD in a single-center clinical cohort based on classical AD biomarkers [12]. However, the application of ML to multiple-center CSF proteomic datasets for the identification of new protein biomarkers related to AD remains quite limited.

To explore robust and universal biomarkers for AD, we collected CSF proteomics datasets of AD from various regions, along with one CSF proteomic dataset that includes other neurodegenerative disorders and one cortex tissue proteomics dataset. Our aim was to develop a machine learning-based strategy for classifying AD by analyzing multiple CSF datasets and identifying promising combinations of protein biomarkers which can be compatible in different cohorts and technologies. We identified a 12-protein AD diagnosis biomarker panel which included multiple biological function and can comprehensively indicate AD heterogeneous progress in different individuals. Our diagnosis model showed high accuracy, robustness and compatibility in different cohorts crossing different countries and different detection technologies.

Methods

Data resources

The present study is a secondary analysis using the mass spectrometry-based proteomic datasets from recently published work [7, 1319]. The ethical approval protocols were offered by each paper. The ethical number and approved ethical committees were listed in Table 1 (Some ethical numbers are not reported in previous published work, so not offered in table). Thus, consent to participate from participants is deemed unnecessary for this study. The data were anonymized before its use. The inclusion criteria were the following: (1) published in English before November 2024; (2) must have human CSF samples and run proteomics analysis (3) CSF samples should have both control and AD; (4) provided downloadable quantification proteomics data. Any group less than 15 biological replication was excluded; (5) The quantification data is available at present. The gender ratio and age are relatively consistent between case and control in most datasets (Table 2) and the detailed information including age, gender, APOE genotype and race were provided in supplementary Tables 16. Take 297_CSF dataset for example, 297 stands for the numbers of samples, CSF stands for the origin of samples. The numbers of some datasets were not consistent with the origin data due to the following reasons: (1) We only kept reliable data according to the information in published work; (2) Data without labels or other disease were deleted; (3) In order to ensure the balance of samples, we retained samples through large data via repeatedly random sampling to cover all picked samples. The validation work using CSF samples (details in Table 3) by ELISA were gifted from Xuanwu hospital under the approval of medical ethics.

Table 1.

The proteomic datasets used in this study

Dataset Total Case Control AD Quantified Proteins ID Ethical approval Quantification methods Origin
297_CSF (Johnson et al., 2020) 297 147 150 532 syn20933797 Approved by the Institutional Review Board at Emory University TMT the ADRC and the Emory Cognitive Neurology Program, USA
114_CSF (Bader et al., 2020) 114 59 55 1448 PXD016278 Approved by the Gothenburg ethics committee, Otto-von-Guericke University Magdeburg ethics committee and University Hospital Schleswig-Holstein clinical care and ethics committee. DIA Magdeburg and Kiel, German & Sweden
36_CSF (Zhou et al., 2020) 38 20 18 41 syn21541022 Approved by the institutional review board PRM ADRC and Emory Cognitive Neurology Program, USA
39_CSF (Bader et al., 2020; Higginbotham et al., 2020) 39 19 20 2874 syn20821165 Approved by the Institutional Review Board at Emory University. TMT Emory ADRC, USA
120_CSF (Dayon et al., 2018) 120 78 42 697 syn20821165 Approved by the institutional ethics committee of the University Hospitals of Lausanne (no.171/2013) TMT University Hospitals of Lausanne, Switzerland
55_CSF (Tao et al., 2024) 55 21 34 3238 PXD039146 NaN TMT

Xuanwu Hospital affiliated with Capital Medical University,

the First and Second Affiliated Hospitals of Zhejiang University School of Medicine,

and the First Affiliated Hospital of Xiamen University

300_CSF (Dammer et al., 2024) 300 140 160 2334 syn52888250

Approved by the

Institutional Review Board at Emory University

TMT

the Emory Goizueta Alzheimer’s Disease Research Center

(ADRC) and Emory Healthy Brain Study (EHBS)

425_CSF (Marta et al., 2023) 425 190 235 665 syn52282088

Approved by the Institutional Ethical Review Boards

of each center, University of Pennsylvania

PEA

Amsterdam Dementia Cohort (ADC) and DEvELOP, the University of Pennsylvania

were included

Multi_disease CSF (Johnson et al., 2020) 78 18 AD:17 477 syn20821165 Approved by the Institutional Review Board at Emory University. TMT Emory ADRC, USA
ALS:19
FTD:11
PD:13
54_CSF_MCI (Johnson et al., 2020) 90 63 MCI:27 792 syn20821165 Approved by the Institutional Review Board at Emory University TMT Emory ADRC, USA
182_DLPFC (Johnson et al., 2020) 321 91 230 3335 syn20933797 Approved by the Institutional Review Board at Emory University LFQ Baltimore Longitudinal Study of Aging, Banner Sun Health Research Institute, the Mount Sinai School of Medicine Brain Bank, the Adult Changes in Thought Study, USA
296_CSF (Dammer et al., 2024) 296 139 157 4098 syn52888250 Approved by the Institutional Review Board at Emory University SomaScan

the Emory Goizueta Alzheimer’s Disease Research Center

(ADRC) and Emory Healthy Brain Study (EHBS)

425_CSF (del Campo et al., 2023) 425 190 235 665 syn52282088

Approved by the Institutional Ethical Review Boards

of each center, University of Pennsylvania

PEA

Amsterdam Dementia Cohort (ADC) and DEvELOP, the University of Pennsylvania

were included

Table 2.

Clinical information in proteomic datasets used in this study

Dataset Control (M/F) AD (M/F) Age of control (yr) Age of AD (yr) APOE RACE
297_CSF 41/106 70/80 65±8.13 68.2±8.3 in supplementary table in supplementary table
114_CSF 25/34 32/23 61.7±17.9 74±6.02 not available in supplementary table
36_CSF 12/8 13/5 69.9±7.6 72.2±9.5 not available in supplementary table
39_CSF 8/11 8/12 69.1±9.0 65.6±11.2 in supplementary table in supplementary table
120_CSF not available not available not available not available not available not available
55_CSF 12/9 16/18 63.6±8.68 62.4±10.2 not available not available
300_CSF 37/103 74/86 64.6±7.83 68.4±8.4 in supplementary table in supplementary table
Muti_disease CSF 11/7

AD: 10/7 ALS:11/8

FTD: 7/4 PD:12/1

68.6±8.9 61.8±10.8(all case) in supplementary table in supplementary table
54_CSF_MCI 26/37 11/16 63.3±6.9 72.3±7.04 not available in supplementary table
182_DLPFC 48/43 103/127 79.15±7.7 81±6.4 in supplementary table not available
296_CSF 37/102 72/85 64.5±7.85 68.3±8.4 in supplementary table in supplementary table
425_CSF 120/70 139/96 57.74±7.6 65.88±7.8 not available not available

Table 3.

CSF samples from Xuanwu hospital

Number Gender Age Diagnosis Aβ42(pg/ml) Aβ40(pg/ml) T-tau (pg/ml) P-tau181 (pg/ml) MMSE MoCA CDR Aβ PET
wu-561-A F 75 AD 686.19 10944.57 1199.5 >182.00 8 4 2 -
wu-479-A F 53 AD - - - - 2 - 3 Positive
wu-754-F F 58 AD 1541.9 8411.6 466.97 38.65 12 8 1 -
wu-736-A F 66 AD 586.71 6166.8 429.08 52.09 12 9 1 -
wu-647-A M 60 AD 565.96 5481.63 381.62 86.27 15 9 1 Positive
wu-603-A F 65 AD 1053.1 6852.51 700.97 110.94 20 16 1 Positive
1/wu-326-A F 58 MCI - - - - 22 17 0.5 -
wu-302-A M 65 MCI - - - - 18 14 0.5 -
wu-293-A F 63 MCI - - - - 22 16 0.5 -
wu-183-A M 45 MCI - - - - 23 19 0.5 -
wu-717-A M 57 MCI 540.26 5447.9 501.65 100.5 23 18 0.5 Positive
wu-676-A F 67 MCI 1252.1 8847.72 452.05 53.63 23 19 0.5 Positive
D1 M 56 Normal - - - - - - - -
D2 M 59 Normal
D4 M 45 Normal
D5 F 71 Normal
D6 F 48 Normal
D7 F 66 Normal

Data preprocessing

Proteins in all datasets with a missing value over 80% were deleted, and then the K-Nearest Neighbor method was employed to fill in the missing values for all datasets [20]. The datasets downloaded from SYNAPSE (https://www.synapse.org/#) are normalized data, apart from the 114_CSF dataset is original data with original abundance, which was logarithmized and normalized [21].

Identification of deps

DEPs were initially identified via the scipy.stats module (1.5.4) in Python (3.6.13) by two-side t.test. For analysis of every DEP, the Levene test was used to check the homogeneity of variance. Proteins with P > 0.05 in Levene test were analyzed by student’s t.test, proteins with P < 0.05 in Levene test were analyzed by Welch t.test. The DEPs were used with a loose filter standard: P < 0.05 for provided the preliminary candidate protein list [20].

The prioritization of biomarker combinations by SVM-RFECV

SVM is a supervised ML algorithm that can be used for classification or regression. It purposes to find out the optimal hyperplane (ωT x + b = 0), making the maximal margin between the nearest heterogeneous points (which is the support vectors) and the hyperplane [22]. In particular, ω is the normal vector, which is also the weight coefficient vector of the optimal hyperplane. The greater the weight, the more classification information the feature contains. SVM classification performance was previously used as an evaluation standard for gene selection, and then the SVM based on recursive feature elimination (SVM-RFE) was applied for feature ranking via weight coefficient parameter. The features with lower ranking were deleted, and then the SVM modeling for preserved features was reconstructed and feature weights were ranked via continuous iteration, until the required number of features were preserved [23]. Tandon R et al. have applied SVM-RFE strategy to classify AD from Control, which has verified the good classification ability of this method [24, 25].

In this study, based on SVM-RFE, we added cross-validation for data sets during every iteration process, which were named as SVM-RFECV, to calculate the model’s crossing score in different feature models, and then select feature numbers with the best cross-validation scores as our optimal feature subset.

We follow schematic diagram (Figs. 1A and 2A) to applied SVM-RFECV algorithm. Briefly, robust DEP list was selected by cohort 1(297_CSF) and cohort 2 (114_CSF) as candidate. The optimal protein combination was determined by selecting the maximal AUC in cohort 1 data. The model is constructed in cohort 1 by selected proteins which was consistent 12-protein panel. The 12-protein panel was evaluated the performance and generalization in cohort 1 and other 4 independent datasets by confusion matrix, accuracy and area under receiver operating curve (AUC). Confusion matrix was calculated predicted true or false percentage in control and disease group separately. The confusion matrix, accuracy and AUC were calculated by scikit-learn database (https://scikit-learn.org/stable/).

Fig. 1.

Fig. 1

Robust DEPs in cohort 1 and cohort 2 analyzed by GO and KEGG. A: Schematic representing the number of CSF proteins differentially expressed in control and AD. B: GO enrichment analysis of Robust DEPs shown in the term of biological processes (hypergeometric test; p < 0.001) and the number of counts (m > 4) GO terms were sorted by E-ratio. C: KEGG enrichment analysis of Robust DEPs (hypergeometric test; p < 0.001) and the number of counts (m > 3). KEGG terms were sorted by E-ratio. D: A robust DEPs Regulatory Network Associated with AD (0.4 confidence in STRING). The size of each red dot indicates the relative BC. Proteins with the largest BC are considered as ‘hub’ proteins in the Robust DEPs. Proteins highlighted in red are top 25% in the network

Fig. 2.

Fig. 2

Identification of potential biomarker combinations by machine-learning strategy. A: Workflow for machine learning approach, including selection step and validation step to prioritize highly potential biomarker combinations. B: The best performing panel based on the AUC curve using the RFE_SVM algorithm was selected in the Discovery cohort1. Y-axis, the area under the AUC curve; x-axis, proteins selected by the RFE_SVM algorithm. C: The accuracy from the 5-fold cross-validation for 12-protein panel using the Discovery cohort 1. Randomly selected control: accuracy = 62%, AUC = 0.67. Shuffling control: accuracy = 49%, AUC = 0.50. D: The confusion matrix of the 12-protein panel. E: The PCA plot using the 12-protein panel showed a good distinguish performance in cohort 1. F: AUC curve by the 12-protein panel showed a good performance in classification of control and AD patients on 4 Validation cohorts. G: The accuracy of control and AD were calculated by the 12-protein panel, all cohorts have > 78% accuracy in both groups. H: the 12-protein panel can specifically distinguish AD from other neurodegenerative disorders on Muti_disease_CSF cohort

Protein-protein interaction (PPI) network construction

The STRING database (version 11.5, https://string-db.org/) [26] was adopted for constructing PPI network with the medium confidence 0.4 or low confidence 0.15, and the Betweenness Centrality (BC) algorithm in Cytoscape software (3.9.1) [27] was used to construct the scoring PPI network [28].

Bioinformatics analysis

Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis were processed by clusterProfiler (4.8.1) [29] in R 4.2.0. Other analyzes are all run in Python (3.6.13). The Principal Component Analysis (PCA) analysis was implemented in Scikit-learn library (0.24.2) and Pearson’s correlation analysis was implemented by Scipy library (1.5.4).

AD associated protein lists which contain AD relevance score were obtained in GeneCards website (https://www.genecards.org/). Rank is examined based on the relevance scores. Permutation test was run 100 times via randomly picking 12 proteins in all human proteins and the mean of AD relevance scores, AD relevance ranks and STRING network degrees were calculated.

ELISA test

The concentrations of BASP1, SMOC1 and FN1 in CSF samples were measured using ELISA kits after appropriate dilution following the manufacturer’s protocol (BASP1: EH4068, FineTest; SMOC1: CSB-EL021842HU, CUSABIO; FN1: E-EL-H0179, Elabscience). The prediction of ELISA result was performed by logistic regression method (Fig. 5D).

Fig. 5.

Fig. 5

The good prediction performance cross multiple platforms. A: Workflow for using the 12-protein panel in multiple technologies prediction. B: AUC curve in SomaScan and PEA datasets. C: The prediction accuracy in SomaScan and PEA datasets. D: The prediction accuracy among three group of ELISA results. E-G: three core proteins of the 12-protein panel ELISA results: SMOC1 in E, BASP1 in F and FN1 in G

Statistical analysis

Statistical methods were described in each method section and data with P < 0.05 was considered statistically significant and reported in figures. Figure 2C were shown as the mean ± SE. Other bar plots were shown as the mean values. Figure 5E-G were plotted by using one-way ANOVA.

Results

Identification of the robust deps from AD CSF discovery cohort

Cohort 1 (control: n = 147; AD, n = 150) [13] was collected as a discovery cohort. To identify robust DEPs, cohort 2 (control: n = 59; AD, n = 55) [14] was also used in differential analysis. 150 DEPs were identified in cohort 1 and 631 DEPs identified in cohort 2. In particular, 81 DEPs were overlapped between cohort 1 and cohort 2. After removing opposite direction, we preserved 70 robust DEPs as the candidate protein list (Fig. 1A) and also showed their alteration in heatmap charts (Supplementary Fig. 1 and Supplementary Fig. 2). The robust DEPs were further analyzed using the GO and KEGG pathway enrichment analyses, revealing a significant enrichment in processes related to immune response, coagulation and fibrinolysis, inflammation, and notably, the HIF-1 signaling pathway and carbon metabolism (Fig. 1B, C). These results are consistent with previous findings of AD pathology [7, 13]. We also examined the PPI network of the robust DEPs, which revealed 13 hub proteins within the network (Fig. 1D). These findings indicate that these conserved proteins can be differentially expressed across multiple CSF datasets and may serve as potential biomarkers or biomarker clusters for AD.

ML-based detection of biomarker combinations for classifying AD cases

To accurately differentiate AD protein biomarkers from controls, we analyzed 70 robust DEPs using support vector machine with recursive feature elimination and cross-validation (SVM-RFECV) to identify potential protein biomarker combinations for classifying AD pathogenesis in the CSF proteomic data of cohort 1. In each iteration, the lowest-ranking protein was eliminated, and the AUC was calculated using 5-fold cross-validation, allowing us to determine the optimal protein combination based on the maximum AUC (Fig. 2A). Through this recursive feature elimination process, we identified a panel of 12 proteins, which included: 14-3-3 protein zeta/delta (YWHAZ), L-lactate dehydrogenase A chain (LDHA), pyruvate kinase PKM (PKM), chitinase-3-like protein 1 (CHI3L1), brain abundant membrane attached signal protein 1 (BASP1), SPARC-related modular calcium-binding protein 1 (SMOC1), gamma-enolase (ENO2), peroxiredoxin-1 (PRDX1), V-set and transmembrane domain-containing protein 2A (VSTM2A), prothrombin (F2), superoxide dismutase (SOD1), and fibronectin 1 (FN1), collectively referred to as the 12-protein panel (Fig. 2A, B). Equal sample size and standard normalization design was used to build a universal diagnosis model by SVM-RFECV method. This model achieved an accuracy of 92% (AUC = 0.96) in identifying AD from controls in cohort 1 using the 12-protein panel (Fig. 2B). For find out the best ML classification methods, we also tested Logistic Regression, Random Forest, Logistic Regression-RFE, Random Forest-RFE and SVM-RFE methods, and the SVM-RFE method produced the best performance in our case (Supplementary Table 7).

To validate the accuracy of the 12-protein panel, we conducted two control experiments. In the first, we created a diagnostic panel using 12 randomly selected proteins from a pool of 532 proteins detected in cohort 1, resulting in an accuracy of only 55% (AUC = 0.58, repeated 20 times). In the second experiment, proteins were selected by shuffling the true labels of the subjects within the training set, yielding an accuracy of 49% (AUC = 0.50). Notably, the AUC scores from the control experiments were significantly lower than those obtained with the 12-protein panel (Fig. 2C).

Further analysis using the confusion matrix (Fig. 2D) and PCA (Fig. 2E) demonstrated strong distinguishing performance, highlighting the reliability of the machine learning-based 12-protein panel in differentiating AD from controls. These results underscore the robustness of this ML-based 12-protein panel for differentiating AD from normal conditions.

To verify the generalization of the 12-protein panel via ML-based strategy, we further collected independent 6 datasets (36_CSF [16], 39_CSF [7], 114_CSF [14], 120_CSF [15], 55_CSF [18] and 300_CSF [19] datasets), 114_CSF data only offered DEPs list, but not involved in ML training process) for study. 3 of the 6 datasets (300_CSF, 36_CSF and 39_CSF) and cohort 1 are both from Emory ADRC project. Compared to the cohort 1, these datasets also exhibited similar high accuracy and AUC values (0.80–0.97) for identifying AD (Fig. 2F & G, Supplementary Fig. 3). Although 16 samples in 36_CSF, 33 samples in 39_CSF and 297 samples in 300_CSF are overlapped with cohort 1 patient ID (Supplementary Table 1), the AUC are not always showed at a very high level (0.94, 0.97 and 0.8). In particular, all 12 proteins in the panel were also observed in 114_CSF, 39_CSF and 300_CSF cases; 11 proteins in the panel found in 120_CSF and 55_CSF cases; 6 proteins found in 36_CSF with low missing rate (Supplementary Table 8, most protein have missing rate less than 1%, only PRDX1 have 47.47% missing rate in 297_CSF, VSTM2A have 24.56% in 114_CSF and FN1 have 20.51% in 39_CSF). It means that even though some proteins in the panel are missing, it still has good diagnostic performance. Moreover, to test the ML-based classification of AD from multiple diseases, we collected multi_disease CSF datasets [7] which have frontotemporal dementia (FTD), Parkinson’s disease (PD), amyotrophic lateral sclerosis (ALS) and AD. The accuracy for distinguishing Control vs. PD, Control vs. ALS and Control vs. FTD were remarkably reduced to 48%, 51% and 65%, respectively. But the accuracy for AD vs. PD, AD vs. ALS and AD vs. FTD were all larger than 0.8 (Fig. 2H), indicating that not only can the 12-protein panel discriminate control from AD, but also can recognize AD and other neurodegenerative disorders. Most importantly, it can also distinguish AD and FTD, though they have similar underlying pathogenesis that makes them hard to distinguish in other work [7]. Therefore, this test further validates the predicting accuracy of the ML model based on 12-protein panel.

The alterations of host proteins in CSF are linked with AD progress

To further validate the 12-protein panel as promising biomarker candidates for classifying AD, we examined the correlations between the expression levels of the 12 proteins and the concentrations of classical AD biomarkers (Aβ42 and tau) in CSF using Pearson’s correlation analysis. We utilized the first principal component (PC1) scores of these 12 proteins, which reflects major variance of these proteins. Our analysis revealed a highly significant correlation between the 12-protein panel and the classical AD biomarkers (297_CSF: positive correlation: Aβ42, R = 0.31, P = 7.5e-0.8; Aβ42/tau, R = 0.64, P = 6.6e-36; negative correlation: p-tau, R=-0.62, P = 7.5e-33; t-tau, R=-0.76, P = 2.4e-57; Fig. 3A). The meta-analysis was also applied to further verify the results (as seen in Supplementary Fig. 4, only 297_CSF, 114_CSF, 120_CSF and 425_CSF are used for meta-analysis due to the samples are independent among these datasets). Also, the 12-protein panel was positively correlated to the Montreal Cognitive Assessment Score (R = 0.41, P = 1.6e-13, no data for meta-analysis). These data suggest the 12-protein panel is moderately associated with the impairment of cognitive functions in AD and has potential to reflect AD progress (Single protein expression correlation to classical AD feature showed in Supplementary Fig. 5). In addition, we checked AD association with the 12 proteins by GeneCards. There were 10 proteins excluding SMOC1 and VSTM2A displaying relevance scores of more than 2 (Fig. 3B). We performed permutation test for 100 times in cohort 1 protein lists. The density map displayed the 12-protein panel has a relatively high score and rank (red stars) in the random distribution (P = 0.07, Fig. 3C; P = 0.03, Fig. 3D). Proteins SOD1, LDHA, ENO2 and FN1 exhibited largest degree correlation with reddest color and can be considered as ‘hub’ proteins in the 12-protein panel by PPI analysis (Fig. 3E, F). The PPI network connectivity also showed a high average degree in the protein panel (average degree = 5.8) compared with that of 100 times random PPI network (max average degree = 3.8) by permutation test (P < 0.01, Fig. 3G). These results indicate the 12-protein panel via SVM-RFECV algorithm is related to AD progression and share certain regulatory relationship.

Fig. 3.

Fig. 3

The PC1 feature of biomarker is highly correlated to AD classical biomarker and enriched in AD relevant gene lists. A: Scattered plot of the AD pathogenic molecule and PC1 represented 12-protein panel showed highly correlation. B: Relevance score and proportion of ranking from the 12-protein panel in GeneCard website. C: Density map of count distribution about the mean score in random 12-proteins from Cohort 1 protein list (repeated 100 times). D: Density map of count distribution about the mean rank in random 12-proteins from Cohort 1 protein list (repeated 100 times). E: A CSF 12-protein panel Regulatory Network Associated with AD (0.15 confidence in STRING). F: The color of each circle indicates the relative Degree Correlation. Proteins with the largest degree correlation are Redder and considered as ‘hub’ proteins in the 12-protein panel. G: Density map of count distribution about the mean degree in the random 12-protein panel from Cohort 1 protein list (repeated 100 times). The red star in C, D, E shows 12-proteins in the panel value

Validation of biomarkers from different tissues and different stages of AD cases

We further assessed whether the 12-protein panel could predict the mild cognitive impairment (MCI) stage of AD or AD itself through brain tissue analysis. In particular, we screened out 11 proteins in the panel detected from a 54_CSF_MCI [13] dataset to distinguish MCI from control, which showed a mean AUC value of 0.88 via 5-fold cross-validation (Fig. 4A). The PCA also demonstrated a clear classification of samples from MCI and control (Fig. 4B). The confusion matrix further showed a high accuracy (control: 0.79, MCI, 0.85) for distinguishing MCI to control by the 11-protein panel (Fig. 4C). Meanwhile, we also screened out 7 proteins in the panel from a 182_DLPFC [13] dataset to distinguish AD from control, which exhibited a mean AUC value of 0.7 via 5-fold cross-validation (Fig. 4D). The confusion matrix exhibited an accuracy of over 0.65 between predicted label and the true label (Fig. 4F). Though the PCA demonstrated unclear classification between AD and control groups, it might due to the samples from brain tissues, which might not be as significant as CSF samples (Fig. 4E). These data suggest not only the 12-protein panel screened by the SVM-RFECV algorithm can clearly classify AD from control, but also it is suitable for predicting CSF samples of MCI. The brain tissue result unveils that these proteins are also altered in brain tissue in AD compared to control, suggesting they are well correlated with AD pathogenesis.

Fig. 4.

Fig. 4

The 12-protein panel has good performance in post-mortem brain tissues and preclinical AD diagnosis. A: AUC curve by 12-protein panel showed a good performance in classification of control and AD patients on 54_CSF_MCI. B: The PCA plot using 12-protein panel showed a good distinguish performance in 54_CSF_MCI. C: The confusion matrix of the 12-protein panel. D: AUC curve by 12-protein panel showed a good performance in classification of control and AD patients on 182_DLPFC. E: The PCA plot using 12-protein panel showed a good distinguish performance in 182_DLPFC. F: The confusion matrix of the 12-protein panel

Core proteins from the panel show robust prediction performance cross technologies

In order to evaluate the generalization ability of the 12-protein-panel, we used another two datasets from two technologies to test the prediction performance (296_CSF dataset [19] by aptamer-based SomaScan and 425_CSF dataset [17] by antibody-based ligand proximity extension assay (PEA), which showed a AUC score of 0.83 and 0.68, respectively (Fig. 5B). Although the 296_CSF dataset is a subset of 300 CSF samples from the same patients analyzed by a different measurement platform as the 297_CSF and 300_CSF datasets, it displayed a high AUC score, indicating the 12-protein-panel model obtained based on different technologies have generalization ability. The relatively low AUC and accuracy of the 425_CSF dataset may be attributed to the limited number of 4 proteins that are utilized in its analysis. It implies that the 12-protein-panel not only can be used in MS-based proteomic data prediction, but also suitable for other technologies. Then, we further employed the ELISA test, a widely-used clinical method, to examine the prediction accuracy of the 12-protein panel. In combining the ML model important score with AD relevance score, SMOC1, BASP1 and FN1 were chosen for ELISA test. Each 6 CSF samples of clinically diagnosed AD, MCI and normal cognition were kindly gifted by Professor Wu Liyong from the Xuanwu hospital (Table 3). The cohort1 CSF proteomic data were used to train the model, and AD or MCI ELISA data were used to test the prediction ability separately (Fig. 5A). Excitingly, 11 out of 12 samples were predicted correctly in AD test (AUC, 0.92) and 10 out of 12 samples were well diagnosed in MCI test (AUC = 0.83) (Fig. 5D). The training and testing data are from different detection technologies, countries and races, but still showed robust prediction performance. It indicates the 12-protein-panel can cross over technologies to distinguish AD and have strong potential clinical application ability. The ELISA results also exhibited differences among three groups. Both SMOC1 and BASP1 showed significantly increased protein expressions in AD CSF samples in comparison to that in control, though their protein expressions were slightly increased in MCI CSF cases with no statistical significance (Fig. 5E-F). In contrary, FN1 exhibited much reduced protein expression level in AD CSF samples compared to that in control, as well as that in MCI vs. control (Fig. 5G). These data indicated that single protein is hard to well distinguish the disease status, but the protein-panel based on the ML algorithm can help to distinguish AD from MCI or control, showing its great potential in clinical application.

Discussion

Several studies have identified potential biochemical biomarkers in blood [30, 31] and CSF [12, 32], however, ML methods have seen limited application in the discovery of new biomarkers. In this study, we screened a 12-protein panel as a biomarker combination for identifying AD using a large-scale proteomic dataset and an ML-based SVM-RFECV algorithm. This algorithm utilizes support vector machines for recursive feature elimination, where support vectors serve as key factors for classification, allowing us to identify essential features while eliminating large but redundant ones. This approach is straightforward and also possesses considerable robustness.

Traditional proteomic analyses typically rely on univariate statistical methods, whereas ML-based analyses can evaluate all proteins simultaneously, making it possible to detect disease-related proteins, especially when considering interactions among them. Consequently, ML-based analysis offers higher sensitivity and greater specificity, even though proteins selected by ML may not always align with those identified through univariate methods. By applying the 12-protein panel across cohorts from different countries and regions, we validated its effectiveness as a universal tool with robust predictive capability. Though this result may due to the same batch of patients shared in some datasets like 297_CSF and 296_CSF (Table 1 and Supplementary Table 1), completely independent datasets (114_CSF, 120_CSF and 55_CSF datasets have different groups of patients compared with 297_CSF) still have high accuracy. The results implies that the 12-protein panel is a robust universal prediction tool. The proteins in this panel demonstrate a strong correlation with classical CSF biomarkers of AD, including Aβ42 and tau, which are often challenging to be detected using traditional mass spectrometry-based proteomic analyses due to their low abundance (approximately t-tau: 100–2000 pg/mL, p-tau: 20–200 pg/mL, Aβ: 200-1000pg/ml, Supplementary Fig. 6) in CSF [14]. The 12-protein panel displays a good detection rate in different CSF proteomic datasets, and shows a relatively high expression abundance (approximately 3-1000 ng/ml, Supplementary Fig. 5, estimated via DIA quantification intensity in 114_CSF datasets [14] and YWHAZ protein concentration in CSF [33]) and good stability in CSF, which can reflect the changes of classic biomarkers such as Aβ42 and tau to a certain extent. The 182_DLPFC brain dataset showed relatively low AUC. The possible reason may due to that only 6 proteins are detected in 182_DLPFC dataset and it has to be noted that, the CSF proteome have different compositions from the brain tissue (the CSF proteome is mainly consisted of secreted proteins from brain cells and permeable proteins from plasma, while brain tissue proteome is mostly composed of intracellular proteins). It highlights that using multiple protein combinations can reduce errors caused by abnormal detection of a single protein and improve the diagnostic performance of AD. Therefore, our study using ML-based analysis processes high accuracy, high AUC value and high stability in discovery of different cohorts, and demonstrates greater potential for AD biomarkers feature selection via multiple cohort validation.

In this panel, it is worth noting that 10 of 12 proteins from this panel have been previously reported for their correlations with AD [34]. Among these proteins, PKM and ENO2 are glycolysis-related proteins that have been identified from proteomic data of AD brain and CSF samples [14, 35, 36]. Recent work also reported changes in energy metabolism in AD, thus they are considered as diagnostic CSF biomarkers of AD in several studies [13, 37, 38]. Glycolysis and energy metabolism is also involved in pathogenesis of AD, evoking some scientists to considering AD as Type 3 Diabetes [39, 40]. Both PRDX1 and SOD1 are important antioxidant enzymes and are also identified previously as AD-associated proteins [41, 42]. In particular, PRDX1 is involved in glycolysis that is likely to be increased along with the induced glycolytic flux [13]. We observed that an increased level of SMOC1 in AD and it was negatively correlated with Aβ42 and Aβ42/tau and Montreal Cognitive Assessment score, but positively correlated with t-tau or p-tau. SMOC1 is a development-associated protein that has been identified by deep multilayer proteomics and its co-localization with plaques in human observed [4346]. CHI3L1 is functional in neural inflammation and tissue remodeling and has also previously identified as CSF AD biomarker [4750]. BASP1, LDHA and YWHAZ were also validated in a synaptic-related biomarker panel and F2 in a vascular-related biomarker panel [51] for diagnosing AD from over 500 CSF samples of multiple degenerative conditions across multiple replication analysis [7, 52, 53]. In particular, YWHAZ is a 14-3-3 family protein that involves in signal transduction by binding to phosphoserine-containing proteins. It was previously identified by bioinformatics in multiple AD models [5457]. Moreover, there are several unreported protein biomarkers, such as VSTM2A that plays a role in lipid storage and FN1 in blood-brain barrier penetrability, which might serve as novel newly biomarkers worthy of further study. These results indicate that the 12-protein panel is functionally involved in energy metabolism, inflammation, antioxidant and synapse related functions, reflecting different AD molecular subtype in pathogenesis of AD mechanism [5860]. This also explains why the biomarker panels have excellent and universal diagnosis capabilities for this highly heterogeneous pathophysiology of AD. Therefore, the 12-protein panel we screened by the SVM-RFECV algorithm can be considered as a robust and reliable biomarker combination for AD classification.

In addition to identifying AD and control, the 12-protein panel also exhibits a strong ability to distinguish AD from other neurodegenerative diseases, including ALS and PD, due to their distinct pathological differences from AD that demonstrated much lower accuracy. Most importantly, it can also well distinguish AD from FTD, though FTD and AD both belong to neurodegenerative diseases with common features in clinical manifestations, neuropathology, and genetic underpinnings, and have shown tighter proteomic relationships in many studies [7, 13, 61]. In addition, this method also displays high accuracy for identification of AD at MCI stage. These indicate the 12-protien panel produced by SVM-RFECV algorithm can be used as effective biomarker sets for distinguishing AD with specificity and reliability. More than that, SomaScan, PEA and ELISA datasets indicated that 12-protein-panel have great generalization ability. It can combine proteomics data and other technologies data in model prediction performance and have excellent clinical application feasibility.

Limitations

This study has several limitations. Since the proteomic data utilized were primarily drawn from published work, we lack comprehensive demographic details about the samples. For instance, information regarding chronic diseases is not available in all datasets, and some datasets do not provide data on sex, age, and ethnicity. This absence of information undoubtedly limits our ability to screen across multiple conditions. For some datasets that the same batch of patients were applied, the good performance could due to sample bias that might be a reason leading to the limitation of the results. In addition, In AD vs. other neurodegenerative disorders, only one dataset was used. So, the ability to discriminate AD and other neurodegenerative diseases needs to be further validated. Furthermore, this work mainly focused on establishing an AD diagnosis model, and we did not conduct further molecular mechanism of the screened proteins involved with AD through wet experiments. It will be essential to investigate these proteins using molecular experiments in brain tissues from both animal models and human samples in future studies. While the SVM-RFECV algorithm successfully identified a 12-protein panel as potential AD biomarkers, there is a need to further develop a kit based on this panel to analyze larger datasets, thereby enhancing the clinical application of this protein panel. In addition, although this study provides preliminary evidence supporting the predictive potential of the 12-protien panel as AD biomarkers, we acknowledge that the generalizability of our predictions remains to be validated based on longitudinal cohorts. The follow-up studies should be considered in the future work.

Conclusion

In summary, we gathered nearly all available CSF proteomic datasets for AD, consisting of over 1200 CSF samples from published research. We employed a ML-based approach to discover and validate an effective protein biomarker panel for AD, achieving an accuracy rate exceeding 90%. Not only can the 12-protein panel diagnosis model be suitable for mass spectrometry-based detection technologies including label-free, TMT and DIA, but also in ELISA technology. Combining with ML algorithm, the panel proteins which are involved different disease pathogenesis can increase diagnosis ability. In the future, longitudinal validation should be performed to assess the long-term robustness of the 12-protein panel predictions. For complex diseases, quantification in multiple proteins involved in different diseases’ pathogenesis using this ML-based algorithm may provide a better strategy in clinical molecular diagnosis.

Electronic supplementary material

Below is the link to the electronic supplementary material.

Supplementary Material 1 (2.9MB, docx)
Supplementary Material 2 (7.7MB, xlsx)

Acknowledgements

We thank Dr. Liyong Wu and Dr. Min Chu in Xuanwu Hospital for providing the CSF samples for validating the expressional changes of those proteins from the panel. We thank the Biological and Medical Engineering Core Facilities of Beijing Institute of Technology for supplying our experimental equipment.

Author contributions

XH, YQ, ZQ, KL and HQ conceived and designed the studies; ZQ and KL wrote the papers. XH developed the machine learning based method; YQ, KL, HL, YY and SJ performed the proteomic analysis; ZQ, JN, YQ and DZ analyzed the data. All authors contributed to the data analysis and presentation in the paper.

Funding

This work was supported by the Ministry of Science and Technology of China (STI2030-Major Projects 2022ZD0206800, Quan Zhenzhen) and the National Natural Science Foundation of China (Grant No. 82371446, Qing Hong; 82371441, Zhenzhen Quan), Beijing Municipal Natural Science Foundation (Grant No. IS23093, 7222113, Qing Hong), Beijing Nova Program (Grant No. 20220484083, 20230484436, Quan Zhenzhen), Sichuan Science and Technology Program (2024YFHZ0010, Quan Zhenzhen) and Innovation Team Project of Guangdong General Colleges and Universities (Natural Science, 2024KCXTD016).

Data availability

All data needed to evaluate the conclusions in the manuscript are present in the manuscript and /or supplementary materials. Code is uploaded to github (https://github.com/liukf10/ML_based_AD-CSF-12-protein-biomarker-panel).

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Xiaosen Hou and Yunjie Qiu contributed equally to this work.

Contributor Information

Kefu Liu, Email: liukefu@csu.edu.cn.

Hong Qing, Email: hqing@bit.edu.cn.

Zhenzhen Quan, Email: qzzbit2015@bit.edu.cn.

References

  • 1.Jack CR Jr., Bennett DA, Blennow K, Carrillo MC, Dunn B, Haeberlein SB, et al. NIA-AA research framework: toward a biological definition of alzheimer’s disease. Alzheimers Dement. 2018;14(4):535–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Fischer L, Adams JN, Molloy EN, Vockert N, Tremblay-Mercier J, Remz J, et al. Differential effects of aging, alzheimer’s pathology, and APOE4 on longitudinal functional connectivity and episodic memory in older adults. Alzheimers Res Ther. 2025;17(1):91. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Jansen WJ, Ossenkoppele R, Knol DL, Tijms BM, Scheltens P, Verhey FR, et al. Prevalence of cerebral amyloid pathology in persons without dementia: a meta-analysis. JAMA. 2015;313(19):1924–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Blennow K, Hampel H, Weiner M, Zetterberg H. Cerebrospinal fluid and plasma biomarkers in alzheimer disease. Nat Rev Neurol. 2010;6(3):131–44. [DOI] [PubMed] [Google Scholar]
  • 5.Hansson O, Batrla R, Brix B, Carrillo MC, Corradini V, Edelmayer RM, et al. The alzheimer’s association international guidelines for handling of cerebrospinal fluid for routine clinical measurements of amyloid beta and Tau. Alzheimer’s Dement J Alzheimer’s Assoc. 2021;17(9):1575–82. [DOI] [PubMed] [Google Scholar]
  • 6.Hansson O, Seibyl J, Stomrud E, Zetterberg H, Trojanowski JQ, Bittner T, et al. CSF biomarkers of alzheimer’s disease concord with amyloid-beta PET and predict clinical progression: A study of fully automated immunoassays in biofinder and ADNI cohorts. Alzheimer’s Dement J Alzheimer’s Assoc. 2018;14(11):1470–81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Higginbotham L, Ping L, Dammer EB, Duong DM, Zhou M, Gearing M et al. Integrated proteomics reveals brain-based cerebrospinal fluid biomarkers in asymptomatic and symptomatic Alzheimer’s disease. Sci Adv. 2020;6(43). [DOI] [PMC free article] [PubMed]
  • 8.Huang S, Cai N, Pacheco PP, Narrandes S, Wang Y, Xu W. Applications of support vector machine (SVM) learning in Cancer genomics. Cancer Genomics Proteom. 2018 Jan-Feb;15(1):41–51. [DOI] [PMC free article] [PubMed]
  • 9.Wang S, Cai Y. Identification of the functional alteration signatures across different cancer types with support vector machine and feature analysis. Biochim Biophys Acta Mol Basis Dis. 2018;1864(6 Pt B):2218–27. [DOI] [PubMed] [Google Scholar]
  • 10.Sammut SJ, Crispin-Ortuzar M, Chin SF, Provenzano E, Bardwell HA, Ma W, et al. Multi-omic machine learning predictor of breast cancer therapy response. Nature. 2022;601(7894):623–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Swanson K, Wu E, Zhang A, Alizadeh AA, Zou J. From patterns to patients: advances in clinical machine learning for cancer diagnosis, prognosis, and treatment. Cell. 2023;186(8):1772–91. [DOI] [PubMed] [Google Scholar]
  • 12.Gao F, Lv X, Dai L, Wang Q, Wang P, Cheng Z et al. A combination model of AD biomarkers revealed by machine learning precisely predicts Alzheimer’s dementia: China Aging and Neurodegenerative Initiative (CANDI) study. Alzheimers Dement. 2022 Jun 6. [DOI] [PubMed]
  • 13.Johnson ECB, Dammer EB, Duong DM, Ping L, Zhou M, Yin L, et al. Large-scale proteomic analysis of alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nat Med. 2020;26(5):769–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Bader JM, Geyer PE, Muller JB, Strauss MT, Koch M, Leypoldt F, et al. Proteome profiling in cerebrospinal fluid reveals novel biomarkers of alzheimer’s disease. Mol Syst Biol. 2020;16(6):e9356. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dayon L, Núñez Galindo A, Wojcik J, Cominetti O, Corthésy J, Oikonomidi A, et al. Alzheimer disease pathology and the cerebrospinal fluid proteome. Alzheimers Res Ther. 2018;10(1):66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Zhou M, Haque RU, Dammer EB, Duong DM, Ping L, Johnson ECB, et al. Targeted mass spectrometry to quantify brain-derived cerebrospinal fluid biomarkers in alzheimer’s disease. Clin Proteomics. 2020;17:19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Del Campo M, Vermunt L, Peeters CFW, Sieben A, Hok AHYS, Lleó A, et al. CSF proteome profiling reveals biomarkers to discriminate dementia with lewy bodies from alzheimer´s disease. Nat Commun. 2023;14(1):5635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Tao QQ, Cai X, Xue YY, Ge W, Yue L, Li XY, et al. Alzheimer’s disease early diagnostic and staging biomarkers revealed by large-scale cerebrospinal fluid and serum proteomic profiling. Innov (Cambridge (Mass)). 2024;5(1):100544. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Dammer EB, Shantaraman A, Ping L, Duong DM, Gerasimov ES, Ravindran SP, et al. Proteomic analysis of alzheimer’s disease cerebrospinal fluid reveals alterations associated with APOE ε4 and Atomoxetine treatment. Sci Transl Med. 2024;16(753):eadn3504. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Muraoka S, DeLeo AM, Sethi MK, Yukawa-Takamatsu K, Yang Z, Ko J, et al. Proteomic and biological profiling of extracellular vesicles from alzheimer’s disease human brain tissues. Alzheimer’s Dement. 2020;16(6):896–907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Shu T, Ning W, Wu D, Xu J, Han Q, Huang M, et al. Plasma proteomics identify biomarkers and pathogenesis of COVID-19. Immunity. 2020;53(5):1108–e225. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Cortes C, Vapnik V. Support-vector networks. Mach Learn. 1995;20(3):273–97. [Google Scholar]
  • 23.Guyon I, Weston J, Barnhill S, Vapnik V. Gene selection for Cancer classification using support vector machines. Mach Learn. 2002;46(1/3):389–422. [Google Scholar]
  • 24.Tandon R, Levey AI, Lah JJ, Seyfried NT, Mitchell CS. Machine learning selection of most predictive brain proteins suggests role of sugar metabolism in alzheimer’s disease. J Alzheimers Dis. 2023;92(2):411–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Tandon R, Zhao L, Watson CM, Elmor M, Heilman C, Sanders K et al. Predictors of cognitive decline in healthy middle-aged individuals with asymptomatic Alzheimer’s disease. Res Sq. 2023:rs.3.rs-2577025. 10.21203/rs.3.rs-2577025/v1
  • 26.Szklarczyk D, Gable AL, Nastou KC, Lyon D, Kirsch R, Pyysalo S, et al. The STRING database in 2021: customizable protein-protein networks, and functional characterization of user-uploaded gene/measurement sets. Nucleic Acids Res. 2021;49(D1):D605–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13(11):2498–504. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Muraoka S, Jedrychowski MP, Yanamandra K, Ikezu S, Gygi SP, Ikezu T. Proteomic Profiling of Extracellular Vesicles Derived from Cerebrospinal Fluid of Alzheimer’s Disease Patients: A Pilot Study. Cells. 2020;9(9). [DOI] [PMC free article] [PubMed]
  • 29.Wu T, Hu E, Xu S, Chen M, Guo P, Dai Z, et al. ClusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov (Camb). 2021;2(3):100141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Yu H, Liu Y, He B, He T, Chen C, He J, et al. Platelet biomarkers for a descending cognitive function: A proteomic approach. Aging Cell. 2021;20(5):e13358. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Walker KA, Chen J, Shi L, Yang Y, Fornage M, Zhou L, et al. Proteomics analysis of plasma from middle-aged adults identifies protein markers of dementia risk in later life. Sci Transl Med. 2023;15(705):eadf5681. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.van der Ende EL. Veld S, Hanskamp I, van der Lee S, Dijkstra JIR, Hok AHYS, CSF proteomics in autosomal dominant Alzheimer’s disease highlights parallels with sporadic disease. Brain. 2023 Jun 22. [DOI] [PMC free article] [PubMed]
  • 33.Sogorb-Esteve A, Nilsson J, Swift IJ, Heller C, Bocchetta M, Russell LL, et al. Differential impairment of cerebrospinal fluid synaptic biomarkers in the genetic forms of frontotemporal dementia. Alzheimers Res Ther. 2022;14(1):118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Li Y, Chen Z, Wang Q, Lv X, Cheng Z, Wu Y et al. Identification of hub proteins in cerebrospinal fluid as potential biomarkers of Alzheimer’s disease by integrated bioinformatics. J Neurol. 2023;270(3):1487–1500. 10.1007/s00415-022-11476-2 [DOI] [PubMed]
  • 35.de Geus MB, Leslie SN, Lam T, Wang W, Roux-Dalvai F, Droit A, et al. Mass spectrometry in cerebrospinal fluid uncovers association of Glycolysis biomarkers with alzheimer’s disease in a large clinical sample. Sci Rep. 2023;13(1):22406. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Del Campo M, Peeters CFW, Johnson ECB, Vermunt L, Hok AHYS, van Nee M, et al. CSF proteome profiling across the alzheimer’s disease spectrum reflects the multifactorial nature of the disease and identifies specific biomarker panels. Nat Aging. 2022;2(11):1040–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Sathe G, Na CH, Renuse S, Madugundu AK, Albert M, Moghekar A, et al. Quantitative proteomic profiling of cerebrospinal fluid to identify candidate biomarkers for alzheimer’s disease. Proteom Clin Appl. 2019;13(4):e1800105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Panyard DJ, McKetney J, Deming YK, Morrow AR, Ennis GE, Jonaitis EM, et al. Large-scale proteome and metabolome analysis of CSF implicates altered glucose and carbon metabolism and succinylcarnitine in alzheimer’s disease. Alzheimer’s Dement J Alzheimer’s Assoc. 2023;19(12):5447–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Jang M, Choi N, Kim HN. Hyperglycemic Neurovasculature-On-A-Chip to study the effect of SIRT1-Targeted therapy for the type 3 diabetes alzheimer’s disease. Adv Sci (Weinh). 2022;9(34):e2201882. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Cho SY, Kim EW, Park SJ, Phillips BU, Jeong J, Kim H, et al. Reconsidering repurposing: long-term Metformin treatment impairs cognition in alzheimer’s model mice. Translational Psychiatry. 2024;14(1):34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Shafiq K, Sanghai N, Guo Y, Kong J. Implication of post-translationally modified SOD1 in pathological aging. GeroScience. 2021;43(2):507–15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Pinz MP, de Oliveira RL, da Fonseca CAR, Voss GT, da Silva BP, Duarte LFB et al. A Purine Derivative Containing an Organoselenium Group Protects Against Memory Impairment, Sensitivity to Nociception, Oxidative Damage, and Neuroinflammation in a Mouse Model of Alzheimer’s Disease. Mol Neurobiol. 2022 Nov 24. [DOI] [PubMed]
  • 43.Bai B, Wang X, Li Y, Chen PC, Yu K, Dey KK, et al. Deep multilayer brain proteomics identifies molecular networks in alzheimer’s disease progression. Neuron. 2020;105(6):975–91. e7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Wang H, Dey KK, Chen PC, Li Y, Niu M, Cho JH, et al. Integrated analysis of ultra-deep proteomes in cortex, cerebrospinal fluid and serum reveals a mitochondrial signature in alzheimer’s disease. Mol Neurodegeneration. 2020;15(1):43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Guo Y, Chen SD, You J, Huang SY, Chen YL, Zhang Y, et al. Multiplex cerebrospinal fluid proteomics identifies biomarkers for diagnosis and prediction of alzheimer’s disease. Nat Hum Behav. 2024;8(10):2047–66. [DOI] [PubMed] [Google Scholar]
  • 46.Johnson ECB, Bian S, Haque RU, Carter EK, Watson CM, Gordon BA, et al. Cerebrospinal fluid proteomics define the natural history of autosomal dominant alzheimer’s disease. Nat Med. 2023;29(8):1979–88. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Connolly K, Lehoux M, O’Rourke R, Assetta B, Erdemir GA, Elias JA et al. Potential role of chitinase-3-like protein 1 (CHI3L1/YKL-40) in neurodegeneration and Alzheimer’s disease. Alzheimer’s & dementia: the journal of the Alzheimer’s Association. 2022 Mar 2. [DOI] [PMC free article] [PubMed]
  • 48.Lananna BV, McKee CA, King MW, Del-Aguila JL, Dimitry JM, Farias FHG et al. Chi3l1/YKL-40 is controlled by the astrocyte circadian clock and regulates neuroinflammation and Alzheimer’s disease pathogenesis. Sci Transl Med. 2020;12(574). [DOI] [PMC free article] [PubMed]
  • 49.Sanfilippo C, Castrogiovanni P, Imbesi R, Musumeci G, Vecchio M, Li Volti G, et al. Sex-dependent neuro-deconvolution analysis of alzheimer’s disease brain transcriptomes according to CHI3L1 expression levels. J Neuroimmunol. 2022;373:577977. [DOI] [PubMed] [Google Scholar]
  • 50.Mun DG, Budhraja R, Bhat FA, Zenka RM, Johnson KL, Moghekar A, et al. Four-dimensional proteomics analysis of human cerebrospinal fluid with trapped ion mobility spectrometry using PASEF. Proteomics. 2023;23(10):e2200507. [DOI] [PubMed] [Google Scholar]
  • 51.Ashton NJ, Nevado-Holgado AJ, Barber IS, Lynham S, Gupta V, Chatterjee P, et al. A plasma protein classifier for predicting amyloid burden for preclinical alzheimer’s disease. Sci Adv. 2019;5(2):eaau7220. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Chen Y, Luo Z, Sun Y, Li F, Han Z, Qi B, et al. Exercise improves choroid plexus epithelial cells metabolism to prevent glial cell-associated neurodegeneration. Front Pharmacol. 2022;13:1010785. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Bisht I, Ambasta RK, Kumar P. An integrated approach to unravel a putative crosstalk network in alzheimer’s disease and parkinson’s disease. Neuropeptides. 2020;83:102078. [DOI] [PubMed] [Google Scholar]
  • 54.Hoffman JL, Faccidomo S, Kim M, Taylor SM, Agoglia AE, May AM, et al. Alcohol drinking exacerbates neural and behavioral pathology in the 3xTg-AD mouse model of alzheimer’s disease. Int Rev Neurobiol. 2019;148:169–230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Park SA, Jung JM, Park JS, Lee JH, Park B, Kim HJ, et al. SWATH-MS analysis of cerebrospinal fluid to generate a robust battery of biomarkers for alzheimer’s disease. Sci Rep. 2020;10(1):7423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Quan X, Liang H, Chen Y, Qin Q, Wei Y, Liang Z. Related network and differential expression analyses identify nuclear genes and pathways in the Hippocampus of alzheimer disease. Med Sci Monitor: Int Med J Experimental Clin Res. 2020;26:e919311. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zhao Y, Xie YZ, Liu YS. Accelerated aging-related transcriptome alterations in neurovascular unit cells in the brain of alzheimer’s disease. Front Aging Neurosci. 2022;14:949074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Long J, Pan G, Ifeachor E, Belshaw R, Li X. Discovery of novel biomarkers for alzheimer’s disease from blood. Dis Markers. 2016;2016:4250480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Chen M, Zhu Y, Li H, Zhang Y, Han M. A quantitative proteomic approach explores the possible mechanisms by which the small molecule stemazole promotes the survival of human neural stem cells. Brain Sci. 2022;12(6). [DOI] [PMC free article] [PubMed]
  • 60.Neff RA, Wang M, Vatansever S, Guo L, Ming C, Wang Q et al. Molecular subtyping of Alzheimer’s disease using RNA sequencing data reveals novel mechanisms and targets. Sci Adv. 2021;7(2). [DOI] [PMC free article] [PubMed]
  • 61.Lalwani AK, Krishnan K, Bagabir SA, Alkhanani MF, Almalki AH, Haque S et al. Network Theoretical Approach to Explore Factors Affecting Signal Propagation and Stability in Dementia’s Protein-Protein Interaction Network. Biomolecules. 2022;12(3). [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (2.9MB, docx)
Supplementary Material 2 (7.7MB, xlsx)

Data Availability Statement

All data needed to evaluate the conclusions in the manuscript are present in the manuscript and /or supplementary materials. Code is uploaded to github (https://github.com/liukf10/ML_based_AD-CSF-12-protein-biomarker-panel).


Articles from Alzheimer's Research & Therapy are provided here courtesy of BMC

RESOURCES