Skip to main content
NPJ Parkinson's Disease logoLink to NPJ Parkinson's Disease
. 2026 Sep 4;12:213. doi: 10.1038/s41531-026-01497-3

Proteo-metabolomic integration identifies stage-specific candidate biomarkers for Parkinson’s disease

Aya Galal 1,2, Ahmed Moustafa 2,3, Mohamed Salama 1,4,5,✉
PMCID: PMC13545238  PMID: 42697887

Abstract

Parkinson’s disease (PD) is a progressive neurodegenerative disorder with a prolonged prodromal phase and complex motor symptoms. Despite improved clinical criteria, early diagnosis and longitudinal monitoring remain challenging. While cerebrospinal fluid (CSF) and plasma metabolites and proteins show biomarker potential, their utility in predictive models is insufficiently characterized. We employed a secondary computational approach to integrate proteometabolomic profiles from CSF and plasma samples of >1100 Parkinson’s Progression Markers Initiative (PPMI) participants. Using multi-omics machine learning, we identified biofluid-specific signatures and evaluated predictive performance. Twenty-one biomarker candidates were validated across three models (SVM, GLMNET, RF); SVM and GLMNET achieved the highest recall (83–86%) and AUCs of 0.84–0.89. Longitudinal mixed-effects modeling revealed eight candidates associated with progression across diagnostic stages. We identified a three-part molecular framework characterizing neurodegeneration: a diagnostic subpanel reflecting early microbiome dysregulation (secretory granins and metabolites) and synaptic breakdown; a second subpanel monitoring phenoconversion via neurogenesis precursors and extracellular matrix proteins; and a third subpanel tracking progression through chronic neuroinflammation and immune activation. This integrated multi-omics approach provides a robust framework for stage-specific PD monitoring and potential clinical deployment.

Subject terms: Biomarkers, Neurology, Neuroscience

Introduction

Parkinson’s Disease (PD) is a progressive neurodegenerative disorder characterized by declining motor function and a range of non-motor symptoms1. A hallmark feature of PD is its extended prodromal phase, during which subtle physiological and biochemical changes precede clinical diagnosis, by up to 10–20 years2,3. These early changes include mitochondrial dysfunction, neuroinflammation, synaptic dysregulation, and protein aggregation. By the time motor symptoms become clinically actionable, more than 50% of nigrostriatal dopaminergic neurons are typically lost, accompanied by advanced α-synuclein pathology within the substantia nigra4.

Cerebrospinal fluid (CSF) provides a direct window into the central nervous system pathology due to its close interaction with brain tissue and its enrichment in proteins, metabolites5, and other molecular indicators reflective of neurodegeneration6. Integrating multiple biofluids, such as CSF and plasma, offers a more comprehensive view of disease biology, as each captures complementary aspects of pathological processes.

In this context, the identification of robust biomarkers is particularly critical for PD, where reliable indicators are needed to distinguish disease stages, improving disease progression monitoring, supporting early diagnosis, and therapeutic development. Despite advances in neuroimaging, molecular profiling, and computational methods, the translation of these insights into clinically useful biomarkers remains limited7,8. Early diagnosis and longitudinal tracking are still constrained by the lack of sensitive, specific, and accessible markers, a challenge compounded by the marked heterogeneity of PD and its overlap with other neurodegenerative disorders2.

Traditional biomarker discovery has largely focused on single molecular targets, like dopaminergic metabolites in plasma or α-synuclein levels in CSF9,10. While informative, these approaches rarely achieve the sensitivity and specificity required for reliable disease stratification or progression tracking. Interpatient variability and disease-stage differences further limit their clinical utility, underscoring the need for strategies that capture the multi-dimensional nature of PD biology.

Emerging evidence supports the use of integrative, multi-omics strategies that combine data across biofluids and molecular layers6. By leveraging complementary information from genomics, transcriptomics, proteomics, and metabolomics11, these approaches enable the identification of stage-specific molecular signatures and improve robustness to biological variability6,12,13.

To address this need, we conducted a secondary computational analysis of proteomic and metabolomic data from CSF and plasma obtained through the Parkinson’s Progression Markers Initiative (PPMI). We evaluated whether multi-omics integration across biofluids improves biomarker discovery and classification performance, and identified molecular signatures associated with PD progression with potential clinical utility.

Results

PPMI data extraction and processing

We analyzed proteomic and metabolomic data from CSF and plasma samples collected from participants enrolled in the Parkinson’s Progression Markers Initiative (PPMI). The proteomic dataset included 661 individuals (177 controls, 101 prodromal, 383 PD) with 477 unique items measured (291 for CSF, and 186 for Plasma), and the metabolomic dataset included 1,136 individuals (232 controls, 300 prodromal, 604 PD), with 646 unique items measured (298 for CSF, and 348 for Plasma) (Table 1). All biomarkers were measured at the baseline clinical visit to ensure relevance for early-stage diagnosis and were reassessed at follow-up visits for up to 16 months for progression analysis.

Table 1.

PPMI-study participant counts and distribution

Omics Biofluid Control Prodromal PD Total patients
Proteomics
CSF 125 80 277 482
Plasma 52 21 106 179
Total patients 177 101 383
Metabolomics
CSF 112 117 300 529
Plasma 120 183 304 607
Total patients 232 300 604

Summary of Study Participants across Metabolomic and Proteomic datasets, each omics level is comprised of CSF and plasma biofluids and 3 cohorts (Control, Prodromal, and PD). N refers to unique items measured in each dataset.

For the following section, results from CSF and Plasma proteomic and metabolomic analysis will be presented. From here onwards, proteomics data will be presented and denoted as ProCSF or ProPlasma, while metabolomic data will be presented and denoted as MetCSF or MetPlasma, respectively.

Statistical modeling

Statistical pairwise comparison from limma modeling identified 124 significant metabolites in MetCSF (69 in Control vs. PD, 30 in Control vs. Prodromal, and 25 in PD vs. Prodromal). In MetPlasma 58, metabolites were found to be significant (10 in Control vs. PD, 12 in Control vs. Prodromal, and 36 in PD vs. Prodromal, (Detailed statistical summary of significant features i.e., t-test, p value, adjusted p-value, and B-statistics, etc., can be found in Data S1). Similarly, 41 significant proteins in ProCSF in Control vs. PD. With no protein passing the predetermined thresholds in plasma across all comparisons (Table 2). (Detailed statistical summary of significant features i.e., t-test, p value, adjusted p-value, and B-statistics, etc., can be found in Data S1 and Data S2).

Table 2.

Descriptive statistics of stratified univariate analysis findings conducted

Type Features analyzed Comparisons performed Number of significant comparisons
Ctrl vs. PD (%) Ctrl vs. Prodromal (%) PD vs. Prodromal (%)
Met CSF 249 747 69 (27.7) 30 (12.0) 25 (10.0)
MetPlasma 298 894 10 (3.3) 12 (4.0) 36 (12.0)
ProPlasma 186 558 0(0) 0(0) 0(0)
ProCSF 291 873 41 (14.0) 0(0) 0(0)

All metrics reported are counts with respective percentages (5%).

Type refers to the dataset type belonging to each analysis where ProCSF—Proteomic CSF, ProPlasma—Proteomic Plasma, MetCSF—Metabolomic CSF, MetPlasma—Metabolomic Plasma.

Significant features underwent subsequent functional pathway enrichment and annotation to identify the molecular categories represented among the top discriminative signals. Across comparisons in the CSF metabolomic dataset (Fig. 1A–C), functional characterization revealed enrichment in pathways involving amino acids and lipid metabolites. Among lipid classes, phospholipids and sphingolipids were most prominently represented, with features from these categories showing the strongest contribution to group differentiation across both biofluids. While in the control vs. PD comparison in plasma metabolomic dataset (Fig. 2A), neurotransmitters such as Dopa, dopamine -3-o-sulfate, dopamine-4-o-sulfate, and 3-methoxytyrosine and neuroprotective metabolites such as piperine, along with essential amino acids show altered expression. In contrast, control vs. prodromal (Fig. 2B) and PD vs. prodromal (Fig. 2C) comparisons show prominent phospholipids and sphingolipid class involvement in altered plasma metabolites. In the proteomic dataset (Fig. 3), enriched functions were predominantly associated with macromolecular binding, protein binding, as well as enzymatic processes related to serine-type endopeptidase activity and serine-type endopeptidase inhibitor activity. These functional categories accounted for the majority of significant protein-level differences and were consistently observed in CSF. Together, these analyses delineate the principal functional classes encompassed by the significant proteomic and metabolomic findings.

Fig. 1. Heatmaps of significant CSF metabolites.

Fig. 1

A Heatmap of statistically significant CSF metabolites in Control and PD comparison. Top 30 differentially expressed metabolites identified in CSF through Control vs. PD comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in PD compared to control highlighted in red the opposite in blue. B Heatmap of statistically significant CSF metabolites in Control and Prodromal comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in prodromal compared to control highlighted in red the opposite in blue. C Heatmap of statistically significant CSF metabolites in PD and Prodromal comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in PD compared to prodromal highlighted in red the opposite in blue.

Fig. 2. Heatmaps of significant Plasma metabolites.

Fig. 2

A Heatmap of statistically significant Plasma metabolites in Control and PD comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in PD compared to control highlighted in red the opposite in blue. B Heatmap of statistically significant Plasma metabolites in Control and Prodromal comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in prodromal compared to control highlighted in red the opposite in blue. C Heatmap of statistically significant Plasma metabolites in PD and Prodromal comparison. Cohort distinction is highlighted on the top row panel with Control highlighted in blue, Prodromal in yellow and PD in red. Expression values range increased expression in PD compared to prodromal highlighted in red the opposite in blue.

Fig. 3. Heatmap of statistically significant CSF proteins in control and PD comparison.

Fig. 3

Expression values range increased expression in PD compared to control highlighted in red the opposite in blue. Enriched functions were predominantly associated with macromolecular binding, protein binding, as well as enzymatic processes related to serine-type endopeptidase activity and serine-type endopeptidase inhibitor activity. These functional categories accounted for the majority of significant protein-level differences and were consistently observed in CSF.

Multivariate analysis

Single-omics analyses were performed for each algorithm on individual biofluids as well as on combined CSF and plasma datasets. This design yielded six model groupings: ProCSF, MetCSF, ProPlasma, MetPlasma, Pro_Combined, Met_Combined. Single omic performance varied substantially by both algorithm and biofluid (Table 3). GLMNET consistently outperformed other models, achieving the highest accuracies for proteomic datasets, particularly Pro_Combined (accuracy = 0.893, AUC = 0.892) and ProCSF (accuracy = 0.866, AUC = 0.886). Proteomic plasma (ProPlasma) also yielded strong performance (accuracy = 0.797, AUC = 0.824), driven by high sensitivity (0.891), while metabolomic datasets were less discriminative, with lower accuracy and AUCs (e.g., MetCSF accuracy = 0.612, AUC = 0.604). RF and SVM demonstrated moderate performance overall, with best results obtained on Pro_Combined (RF accuracy = 0.789, AUC = 0.747; SVM accuracy = 0.714, AUC = 0.695) and Met_Combined (SVM accuracy = 0.810, AUC = 0.773). XGBoost underperformed relative to other algorithms, with accuracies generally below 0.70, except for modest improvement in MetCSF (AUC = 0.679) and MetPlasma (AUC = 0.650). NN achieved competitive results for Pro_Combined (accuracy = 0.801, AUC = 0.764) and ProCSF (accuracy = 0.769, AUC = 0.682), though performance dropped sharply for plasma-only datasets. Overall, proteomic models particularly those leveraging combined CSF and plasma datasets consistently outperformed metabolomic models across algorithms, with CSF only models showing higher performance compared to plasma only models.

Table 3.

Performance evaluation metrics for single omics modeling: four machine learning algorithms were used on individual CSF, plasma biofluids as well as combined model for each omics

Model Accuracy % Sensitivity % Specificity % F1-score % AUC %
GLMNET
ProCSF 86.60 (0.56) 72.70 (0.77) 90.10 (1.22) 75.50 (1.23) 88.57 (1.13)
MetCSF 61.20 (0.99) 58.20 (0.54) 62.90 (0.38) 59.40 (1.55) 60.42 (1.15)
ProPlasma 79.7 (1.13) 89.10 (0.37) 84.30 (0.88) 68.20 (1.1) 82.40 (0.62)
MetPlasma 59.80 (1.17) 69.10 (0.65) 59.70 (1.03) 67.10 (0.35) 57.20 (0.63)
Pro_Combined 89.30 (0.63) 95.20 (0.74) 92.70 (1.04) 75.50 (1.35) 89.20 (0.47)
Met_Combined 78.40 (0.97) 92.70 (0.7) 72.70 (0.95) 66.70 (0.87) 73.50 (0.76)
Random Forest
ProCSF 68.20 (0.67) 62.90 (1.19) 62.50 (0.52) 65.00 (1.97) 60.60 (1.47)
MetCSF 67.30 (0.44) 67.50 (2.6) 52.50 (0.66) 61.50 (0.34) 53.50 (1.92)
ProPlasma 66.20 (0.71) 62.90 (0.8) 63.00 (1.79) 58.40 (0.85) 62.90 (0.63)
MetPlasma 69.70 (0.19) 86.00 (1.1) 84.30 (1.05) 67.60 (0.3) 61.60 (0.96)
Pro_Combined 78.90 (0.78) 79.30 (0.62) 77.10 (1.9) 66.70 (1.01) 74.70 (0.6)
Met_Combined 72.40 (0.98) 71.40 (0.52) 67.10 (0.52) 69.10 (2.09) 72.70 (1.11)
Support Vector Machine (SVM)
ProCSF 68.80 (0.97) 62.90 (0.95) 57.20 (2.07) 55.10 (1.26) 54.10 (0.84)
MetCSF 67.00 (0.77) 57.30 (1.84) 50.00 (1.14) 53.80 (1.01) 58.80 (0.8)
ProPlasma 70.00 (0.83) 71.40 (0.73) 73.30 (1.64) 62.80 (0.88) 66.70 (1.01)
MetPlasma 66.10 (0.79) 65.30 (0.8) 59.80 (0.2) 55.40 (1.79) 61.8(1)
Pro_Combined 71.40 (0.33) 79.80 (0.24) 77.10 (1.34) 75.00 (0.76) 60.50 (1.11)
Met_Combined 81.00 (0.74) 85.20 (1.51) 71.40 (0.47) 76.40 (0.79) 77.30 (0.83)
XGBoost
ProCSF 60.40 (0.87) 60.30 (1.51) 57.10 (0.47) 63.30 (0.67) 57.80 (0.72)
MetCSF 65.20 (0.83) 69.70 (0.87) 76.10 (1.67) 65.20 (0.88) 67.90 (0.83)
ProPlasma 51.40 (0.77) 47.30 (1.43) 45.00 (0.58) 46.70 (0.87) 47.30(2)
MetPlasma 66.90 (0.76) 61.80 (1.31) 71.40 (0.69) 73.90 (0.8) 65.00 (0.73)
Pro_Combined 68.50 (1.31) 71.10 (1.07) 60.00 (0.61) 61.50 (0.63) 67.10 (0.73)
Met_Combined 50.90 (0.81) 51.10 (0.74) 71.10 (0.8) 56.20 (0.33) 62.80 (0.63)
Neural Networks (NN)
ProCSF 76.90 (0.76) 76.40 (1.54) 76.20 (0.35) 75.50 (0.74) 68.20 (0.8)
MetCSF 60.00 (0.62) 67.00 (0.77) 57.40 (1.22) 59.40 (1.23) 47.30 (1.13)
ProPlasma 53.80 (0.79) 51.80 (0.54) 45.70 (0.74) 51.50 (1.55) 42.90 (1.17)
MetPlasma 59.30 (1.15) 58.20 (0.92) 57.10 (0.38) 66.70 (0.37) 56.20 (0.88)
Pro_Combined 80.10 (1.1) 86.30 (0.62) 80.00 (1.04) 70.50 (0.47) 76.42 (0.76)
Met_Combined 75.38 (0.99) 78.98 (0.78) 76.73 (0.63) 67.97 (0.38) 74.93 (0.54)

To assess individual and combined performance.

All metrics are presented as mean percentage with their respective standard deviations (SD).

GLMNET elastic net regularization, NN neural network, RF random forest, SVM support vector machine, ProCSF proteomic CSF, ProPlasma proteomic plasma, MetCSF metabolomic CSF, MetPlasma metabolomic plasma, Met_Combined combined metabolomic plasma and CSF dataset, Pro_Combined combined proteomic plasma and CSF dataset.

Integrative approach (multi fluid + multi omics)

Candidate feature sets incorporating biomarkers identified from stratified univariate analyses and top performing algorithm(s) from single-omics analysis were further reduced using LASSO. Feature selection was performed with five-fold cross-validation to optimize model sparsity while retaining predictive capacity. This approach reduced the initial panel of 1138 features to 318 non-zero coefficient features, comprising both proteomic and metabolomic measurements from CSF and plasma. The retained features represented an integrated multi-fluid, multi-omics signature that was subsequently used to train and evaluate ML classifiers.

To evaluate predictive performance across biological matrices in the combined dataset, we benchmarked four ML classifiers: GLMNET, NN, RF, SVM. ROC analysis demonstrated that all models achieved robust discriminative power, with AUC values ranging from 0.82 to 0.89 (Fig. 4). SVM and RF consistently outperformed the other classifiers, yielding higher sensitivities across a broad range of specificities.

Fig. 4. Comparative ROC curves for machine learning classifiers.

Fig. 4

Receiver operating characteristic (ROC) analysis comparing Elastic Net Regularization (GLMNET, cyan), Neural Network (NN, purple), Random Forest (RF, red), and Support Vector Machine (SVM, green). All models demonstrated strong discriminative power, with SVM achieving the highest overall performance (accuracy = 82%, sensitivity = 86%, specificity = 76%, F1 = 85%, AUC = 0.89), followed by RF (accuracy = 79%, AUC = 0.87), GLMNET (accuracy = 76%, AUC = 0.84), and NN (accuracy = 77%, AUC = 0.82).

Quantitative performance metrics further supported these trends (Table 4). SVM achieved the best overall performance (accuracy = 82%, sensitivity = 86%, specificity = 76%, F1-score = 85%, AUC = 0.89), followed closely by RF (accuracy = 79%, sensitivity = 82%, specificity = 75%, F1-score = 82%, AUC = 0.87). GLMNET and NN showed moderately lower accuracy (76% and 77%, respectively) and AUC (0.84 and 0.82, respectively), though both maintained balanced sensitivity and specificity. Together, these results indicate that ensemble- and margin-based classifiers provide the most reliable predictive performance for multi-omics integration across fluid types, while linear and neural network models yield competitive but slightly weaker results.

Table 4.

Performance evaluation metrics of machine learning classifiers applied on a multi-fluid, multi-omic dataset

Model Accuracy % Sensitivity % Specificity % F1-score % AUC %
GLMNET 76.00 (1.04) 83.00 (1.17) 66.83 (0.88) 80.26 (0.76) 84.23 (0.54)
NN 77.38 (1.55) 81.00 (0.92) 71.37 (0.37) 80.24 (0.76) 82.70 (0.23)
RF 79.13 (0.71) 82.72 (0.92) 75.00 (1.2) 82.06 (0.32) 87.17 (1.07)
SVM 82.42 (0.41) 86.38 (1.2) 76.03 (0.58) 85.32 (0.24) 89.00 (0.78)

Summary of predictive performance for Elastic Net Regularization (GLMNET), Neural Network (NN), Random Forest (RF), and Support Vector Machine (SVM). Reported metrics include overall accuracy, sensitivity, specificity, F1-score, and area under the ROC curve (AUC). SVM achieved the highest accuracy (82%) and AUC (89%), followed by RF (79%, 87%), GLMNET (76%, 84%), and NN (77%, 82%), highlighting the superior discriminative performance of SVM and RF across evaluation measures.

All metrics are presented as mean percentages with their respective standard deviations (SD)

GLMNET elastic net regularization, NN neural network, RF random forest, SVM support vector machine.

A total of 21 cross-fluid, multi-omics disease-specific features were identified and validated across at least two classification models, and their relative feature importance was extracted (Table 5, and Data S3). These features included established markers of PD as well as proteins and metabolites linked to key processes implicated in PD pathophysiology, such as neuroinflammation, immune response, and neurotransmitter regulation. The panel encompassed both metabolite and protein biomarkers, including cadaverine, PE(P-16:0/22:4), dopamine 3-O-sulfate, and 3-methoxytyrosine, alongside neurosecretory and synaptic proteins (Neurosecretory protein—VGF [O15240], secretogranin [P05060], cadherin-13 [P55290], Limbic system-associated membrane protein (LSAMP) [Q13449]), and enzymatic regulators (peptidyl-glycine alpha-amidating monooxygenase (PAM) [P19021], CD59 glycoprotein [P13987]). Additional features included neuroendocrine protein 7B2 [P05408], lumican [P51884], UDP-glycoprotein glucosyltransferase 1 [Q9NYU2], cytoplasmic malate dehydrogenase [P40925], galectin-3 binding protein [Q08380], and immunoglobulin heavy constant gamma 4 [P01861]. Circulating plasma proteins such as kininogen-1 [P01042], chitinase-3-like protein 1 [P36222], phosphatidylethanolamine-binding protein 4 [Q96S96], haptoglobin [P00738], and corticosteroid-binding globulin [P08185] were also retained. Figure 5: Cross-fluid, multi-omics candidate biomarker panel retained across at least two ML models ranked by average importance.

Table 5.

List of 21 cross-fluid, multi-omic candidate biomarkers identified form integrative approaches

Name Other Identifier Biofluid Omic Associated Process/Function Previously reported in NDD
3-Methyltyrosine Metyrosine Plasma Metabolite Tyrosine hydroxylase enzyme inhibitor Alters dopamine levels in PD10
Dopamine-3-osulfate - Plasma Metabolite Deactivated, sulfated form of dopamine Progression PD marker39
PE(P-16:0/22:4) Phosphatidylethanolamine Plasma Metabolite Cell membrane and signaling Lipid class unexplored role in PD40
Cadaverine - Plasma Metabolite Polyamine metabolite of the microbiome Endogenous H4R agonist in the brain Increased in PD6,28,41,42
O15240 Neurosecretory protein—VGF CSF Proteomics Neurosecretory protein Decreased in PD6,43, NDD44, AD45
P05060 Secretogranin-1 (CHGB) CSF Proteomics Neuroendocrine secretory granule protein Interaction with LARRK246
P55290 Cadherin-13 CSF Proteomics Calcium-dependent cell adhesion proteins Protective role in interneuron development47
Q13449 Limbic system-associated membrane protein –(LSAMP) CSF Proteomics Selective neuronal growth and axon target mediator Increased in PD6,31
P19021 Peptidyl-glycine alpha-amidating monooxygenase –(PAM) CSF Proteomics Catalyze amidation of C-terminus of proteins Unexplored role in neuronal integration37
P13987 CD59 glycoprotein CSF Proteomics Inhibitor of the complement MAC action Potential role in NDD48
P05408 Neuroendocrine protein 7B2 CSF Proteomics Molecular chaperone for PCSK2/PC2 in secretory pathway Decreased in PD6,30
P51884 Lumican CSF Proteomics Proteoglycan Inflammatory cell response in NDD49, PD50
Q9NY02 UDP-glycoprotein glucosyltransferase 1 CSF Proteomics Glycosyltransferase Misfolding of neuroproteins in NDD51
P40925 Cytoplasmic malate dehydrogenase (MDH1) CSF Proteomics Role in malate-aspartate shuttle Potential role in metabolic dysfunction in PD52
Q08380 Galactin-3 binding protein Plasma Proteomics Integrin-mediated cell adhesion Pathophysiology of PD53, potential biomarker of idopathic PD54
P01861 Immunoglobulin heavy constant gamma 4 CSF Proteomics Immunoglobulin Increased in PD6
P01042 Kininogen-1 CSF Proteomics Inhibitors of thiol proteases Early cognitive impairment in PD55
P36222 Chitinase-3-like protein 1 CSF Proteomics Tissue remodeling Decreased in PD6,38
P00738 Haptoglobin (HPT) CSF Proteomics Hepatic recycling of heme Decreased in PD56
P08185 Corticosteroid-binding globulin CSF Proteomics Transport of glucocorticoids and progestins Inflammation in PD57
Q96S96 Phosphatidylethanolamine-binding protein 4 CSF Proteomics AKT phosphorylation, PI3K-AKT signaling pathway Altered levels in AD58,59

NDD neurodegenerative disorders, PD Parkinson disease, AD Alzheimer disease, MAC membrane attack complex, H4R Histamine 4 Receptor.

Fig. 5. Validated cross-fluid, multi-omics candidate biomarker panel.

Fig. 5

Validated features retained across at least two ML models ranked by average importance. Figure showing the 21 features retained across at least two classifiers, ranked by average importance score. Features validated in three models are indicated in dark blue, while those retained in two models are shown in light blue. Identified features span proteomic and metabolomic markers from cerebrospinal fluid and plasma, including known PD-related metabolites and proteins.

Longitudinal progression analysis

To assess biomarker dynamics across disease progression and estimate fixed and random effects we applied linear mixed-effects models to quantify longitudinal trajectories in both CSF and plasma (Figure S1). Among the 21 candidate biomarkers identified from prior analyses, only 5 (24%) met our inclusion criteria in both CSF and plasma biofluids (defined as data availability from at least 5 patients with a minimum of 3 visits). An additional 12 biomarkers (57%) met criteria for CSF-only analyses, while 4 (19%) for plasma-only analyses. Detailed statistical information regarding sample size per biomarker can be found in Data S4. Individual trajectory plots of candidate biomarkers are shown in Figure S2. Longitudinal trajectories of candidate biomarkers over 16 months, showing both individual patient-level variability and cohort-level population trends. Spaghetti plots represent individual trajectories, while overlaid regression lines reflect mixed-effects model estimates, highlighting differences in progression patterns across diagnostic groups from baseline to 16 months.

Many CSF biomarkers show distinctly positive slopes in prodromal compared to control and PD, shows negative or small slopes in PD and control. Including 015240, P05060, P05408, P40925 (Fig. 6). This suggests a more pronounced biomarker progression within CSF in preliminary stages of disease development compared to PD or controls. Others such as P00738 show stepwise decline across disease progression (Fig. 6). Conversely in plasma, several candidates show controls with the strongest positive slopes, intermediate PD, and negative or flat prodromal. Including, 3-Methoxytyrosine exhibiting a strong PD slope and a weak prodromal signal. Suggesting a potential inverse progression pattern in plasma.

Fig. 6. Disease progression rates of candidate biomarkers.

Fig. 6

Showing mean slope (log10 units per month), 21 candidate biomarkers in both biofluids (CSF and plasma) and cohorts (Control (green), Prodromal (Yellow), PD (Red)). Each bar shows the disease progression rate as well as the direction of change.

Progression rate analysis revealed several subsets of biomarkers with stage-specific utility (Fig. 6) metabolites exhibit a distinct pattern between cohorts with lower prodromal slopes when compared to PD cohort. Cadaverine displayed the most striking difference. In the prodromal cohort, it showed the highest positive slope (5.95 × 10⁻⁵ log₁₀ units per month), indicating a sharp increase. Conversely, the PD cohort showed a negative slope (−1.00 × 10⁻⁴ log₁₀ units per month), suggesting a decrease or a different rate of change in established disease. In contrast 3-methoxytyrosine and dopamine-3-o-sulfate shows the opposite trend with a lower slope in the prodromal cohort (1.74 × 10⁻⁴, 5.60 × 10⁻⁴ log₁₀ units per month respectively) compared to the significantly higher slope in the PD cohort (8.93 × 10⁻³, 4.15 × 10⁻³ log₁₀ units per month respectively). PE(P-16:0/22:4) similarly demonstrated a lower slope in the prodromal cohort (1.80 × 10⁻³ log₁₀ units per month) compared to the higher slope in the PD cohort (3.22 × 10⁻³ log₁₀ units per month).

In proteins, P00738 (haptoglobin) displays the most striking difference, with control showing the strongest positive slope (1.14 × 10⁻² log₁₀ units per month), while PD demonstrated a decline in slope (7.17 × 10⁻³) and the prodromal showed an even lower trajectory (3.11 × 10⁻³ log₁₀ units per month). A similar pattern was observed for Q08380 (Galactin-3 binding protein), which increased in controls (1.03 × 10⁻² log₁₀ units per month) but declined in PD (−2.50 × 10⁻³ log₁₀ units per month), suggesting a reversal of trend with disease progression. In contrast, P01042 (Kininogen-1) and plasma P08185 (Corticosteroid-binding globulin) displayed negative slopes in PD (−3.69 × 10⁻³ and −5.40 × 10⁻³ log₁₀ units per month) and drops in prodromal plasma (−1.10 × 10⁻² and −1.75 × 10⁻³ log₁₀ units per month), whereas both remained positive in controls (4.41 × 10⁻³ and −5.40 × 10⁻³ log₁₀ units per month, respectively). Lumican (P51884) also shifted from positive in controls (1.83 × 10⁻³ log₁₀ units per month) to strongly negative in prodromal plasma (−7.88 × 10⁻³ log₁₀ units per month), with intermediate slope in PD (3.33 × 10⁻³ log₁₀ units per month).

In CSF, the clearest stage distinction was observed for P40925 (Cytoplasmic malate dehydrogenase) and Q13449 (LSAMP), both showing the highest positive slopes in the prodromal cohort (6.11 × 10⁻³ and 6.04 × 10⁻³ log₁₀ units per month, respectively), followed by attenuated slopes in PD (1.19 × 10⁻³ and 4.65 × 10⁻⁴ log₁₀ units per month). Conversely to its plasma levels, P00738 (haptoglobin) in CSF was relatively stable in controls (3.61 × 10⁻³) but shifted to a negative slope in prodromal patients (−5.91 × 10⁻⁴ log₁₀ units per month), before partially recovering in PD (2.32 × 10⁻³ log₁₀ units per month). These trajectories highlight proteins such as P00738, Q08380, and P02749 as key candidates reflecting differential progression dynamics across biofluids and disease stages.

In PD CSF, most proteins exhibited negative slopes compared to prodromal. P36222 declined (4.33 × 10⁻³ log₁₀ units per month) in prodromal to (3.27 × 10⁻³ log₁₀ units per month) in PD, while P19021 decreased further (1.97 × 10⁻³ in prodromal, 1.33 × 10⁻³ log₁₀ units per month in PD). Proteins such as P05060 (−1.16 × 10⁻³ log₁₀ units per month), P08185 (−1.30 × 10⁻³ log₁₀ units per month), P55290 (5.04 × 10⁻⁴ log₁₀ units per month), Q13449 (4.65 × 10⁻⁴ log₁₀ units per month), and Q9NYU2 (3.26 × 10⁻⁴ log₁₀ units per month) displayed reduced slopes, with some transitioning from positive in prodromal to negative in PD. Similarly, P05408 shifted from strong positive in prodromal (5.40 × 10⁻³ log₁₀ units per month) to nearly flat in PD (1.08 × 10⁻⁴ log₁₀ units per month). Taken together these patterns reveal the large magnitude shifts with controls exhibiting consistent higher positive slopes, while prodromal and PD cohorts are classified by the inverse.

Trajectory plots and disease progression rates were used to investigate longitudinal changes across disease stages. To quantify disease-related differences in biomarker progression, individual delta progression slopes representing the difference in the rate of change (log₁₀ units per month) between the study cohorts (prodromal and PD) and the control group. Several candidate biomarkers show a negative delta progression rate in both prodromal and PD cohorts indicating rapid decrease when compared to control (Fig. 7). Particularly pronounced in P01042 (in both CSF and plasma) and P51884, were prodromal cohort exhibits a more rapid decline than PD. This suggests pronounced decrease in levels in early stages of disease. Conversely P05060 demonstrates positive slope in both cohorts with slightly more pronounced increase in PD.

Fig. 7. Delta progression slopes for candidate biomarkers.

Fig. 7

For each biomarker and biofluid, patient-specific slopes were grouped by cohort. Cohort-level mean slopes were contrasted against control groups, bar plots showing differences in log10-scaled progression rates per month relative to control. PD-Control (red), Prodromal-Control (yellow).

From the initial 21 candidates, we prioritized features showing the largest absolute slopes and those with functions implicated in neurological or brain-related processes. This filtering enabled the identification of stage- and fluid-specific patterns, where several proteins displayed strong divergence between prodromal and PD cohorts, suggesting utility as early indicators of disease trajectory. Figure 8- shows the disease progression rates for the subset showing stage specific utility and potential clinical implications. The 21 candidate biomarkers were subclassified into 3 subpanels (early diagnosis markers, prodromal to PD conversion, and PD progression) based on a combination of the direction and magnitude of longitudinal slopes, together with their differential behavior across clinical groups. Owing to the differences in scale and variability between analytes, classification was based on relative effect sizes and consistency of slope direction, and statistical significance of longitudinal trends within each biomarker.

Fig. 8. Disease progression rates of candidate subpanels.

Fig. 8

This figure depicts the progression rates of the eight candidate biomarkers stage-specific utility (those divided across the three subpanels) and potential clinical implications. Progression rates show changes per month across disease cohorts - Control (navy), Prodromal (purple), PD (yellow).

Among those showing noticeable slopes during prodromal stages and minimal during PD includes cadaverine (CSF) positive slope in prodromal (+5.5 × 10e-5/month), with negative slope in PD (−0.0001/month). P40925 (CSF)—shows a positive prodromal and PD slope (0.00067/month, 0.001846/month), P05060 (CSF)—prodromal shows a rise (+0.0032/month) while PD remains flat (−0.0012/month). P05408 (CSF)—strong prodromal increase (+0.0054/month), minimal PD change (+0.0001/month). P40925 (CSF)—prodromal increase (+0.0061/month), small PD slope (+0.0012/month). Q13449 (CSF)—prodromal increase (+0.0060/month), minimal PD slope (+0.0005/month). Suggesting potential early detection markers, useful in identifying patients in the prodromal stage.

Others are candidates to be conversion (Prodromal-PD) markers, including O15240 (CSF) where PD patients began at lower levels and declined (-0.0033/month), whereas prodromal patients showed a compensatory increase (+0.0035/month). P51884 (Plasma)—PD shows mild increase (+0.0033/month), prodromal shows decline (−0.0079/month) (Fig. 8). Progression markers: P01042 (Plasma) and Q08380 (Plasma) showed the largest negative slopes in both PD and prodromal groups, with greater declines than controls—potentially reflecting neurodegeneration, protein loss, or altered clearance. Notably, prodromal groups exhibited even steeper declines, supporting their utility as progression biomarkers for early disease.

Candidate biomarkers were filtered down to eight potential stage specific candidates based on their slope, potential clinical utility, and biological relevance these can be divided into subpanels with the following potential clinical utility. Subpanel I early diagnostic markers: - Cadaverine, P05060 (Secretogranin), P05408 (Neuroendocrine protein 7B2), Q13449 (Limbic system-associated membrane protein). Subpanel II—Prodromal to PD conversion: O15240 (Neurosecretory protein VGF) P51884 (Lumican). Subpanel III—Sensitive Progression markers: P01042 (Kininogen-1), Q08380 (Galectin 3 binding protein).

In addition to the eight candidates in our subpanels, a further six from our initial candidate biomarker panel show future potential as disease monitoring biomarkers, due to parallel changes in slope between prodromal and PD, however they would require further investigation and additional data to increase the sample size in longitudinal tracking to support the claims. (Figure S3). These include P19021 (PAM), P55290 (cadherin-13), and P36222 (Chitinase-3-like protein 1), all 3 CSF proteins showing positive slope for both prodromal and PD, P08185 (Corticosteroid-binding globulin) shows positive PD/prodromal slope in plasma and negative slope in both in CSF. Q9NYU2 (UDP-glycoprotein glucosyltransferase 1) and Q96S96 (Phosphatidylethanolamine-binding protein 4) in CSF both show negative slopes across both PD and Prodromal.

Discussion

While most PD studies rely on single-modality clinical features (e.g., gait14, sleep15, and motor data15,16), we present an integrative ML framework combining metabolomic and proteomic data across multiple biofluids. This multi-omics approach complements existing models based on neuroimaging or clinical-demographic data to improve predictive and monitoring accuracy across the PD life cycle.

Our univariate analysis and subsequent functional enrichment identified key molecular axes of distinction across cohorts. Within the metabolomic dataset, robust enrichment in amino acid metabolism and lipid biochemistry emerged (Fig. 1A–C), with phospholipids and sphingolipids serving as dominant contributors. Given their roles in membrane dynamics, cell signaling, and neuroinflammation (Fig. 2A–B), their consistent discriminative power across both biofluids suggests they may act as cross-compartment biomarkers bridging metabolic and proteomic dysregulation. Within the proteomic dataset, enrichment highlighted macromolecular binding and enzymatic activities, including serine-type endopeptidase activity and its regulatory mechanisms (Fig. 3), which may drive proteome-level cohort separation.

Differential expression analysis revealed a potential molecular transition from neuromodulation to neuroinflammation during disease progression. In the CSF, this signature is characterized by proteins involved in neuromodulation, proteostatic regulation, and innate immunity. Notably, the neuron-specific glycoprotein NELL2, vital for neuronal proliferation and synaptic function via the MAPK pathway17,18, exhibited significantly altered expression in PD participants.

Additionally, endocrine, and homeostatic regulatory proteins showed altered expression: GRP78, an endoplasmic reticulum chaperone, suggests a potential response to proteotoxic stress19, while alterations in peptidyl-glycine alpha-amidating monooxygenase (PAM) and chromogranin A align with the neurotransmitter dysregulation characteristics of PD20,21.

These findings indicate that centra nervous proteins involved in synaptic architecture and neuroinflammation display higher discriminatory power in proximal CSF than in plasma, were systemic dilution, metabolic turnover, or peripheral compensation may attenuate these signs. Conversely, systemic proteins like serotransferrin, and APOC3 contributed more prominently to peripheral cohort separation. APOC3 may influence pathology indirectly through lipid metabolism22, while serotransferrin mediates cellular iron delivery23,24 and neuroimmune interactions25,26. Rather than compartment-specific pathology, these peripheral markers likely reflect individual metabolic variations, disease susceptibility, or overall patient heterogeneity.

Our machine learning multivariate analysis identified a 21-candidate biomarker cross-fluid panel (Table 5), filtered down to eight candidates stratified into three exploratory subpanels based on longitudinal trajectories and biological plausibility (Fig. 8).

Subpanel I (Early diagnosis) candidates characterized by altered trajectories during prodromal stages and minimal during clinical PD. This includes one CSF metabolite (cadaverine) and three CSF proteins (Neuroendocrine protein 7B2, Cytoplasmic malate dehydrogenase, and LSAMP). These markers are associated with early homeostatic dysregulation and synaptic loss. Elevated polyamines (e.g., cadaverine, and putrescine) are frequently reported in PD CSF and plasma27,28; under oxidative stress, polyamines can promote toxic alpha-synuclein aggregation6,29,30. Speculatively, the non-linear trajectories of CSF cadaverine may reflect an exploratory link between systemic metabolic or barrier alterations and central pathology. Furthermore, the secretory chaperone 7B2 co-localizes with protein aggregated in the PD brain, while the adhesion molecule LSAMP indicates early regional vulnerability in limbic structures31,32.

Subpanel II (Prodromal to PD conversion) features the neuropeptide precursor VGF and the extracellular matrix proteoglycan lumican, which show an initial increase during prodromal phase followed by low levels and subsequent decline over time. These features capture extracellular matrix (ECM), structural integrity, and neuroplasticity. VGF is essential for neurogenesis, while lumican regulates tissue remodeling and inflammatory signaling6,33–35. Combined, this exploratory panel may reflect a biological threshold where ongoing neuroinflammation begins to outpace the brain’s compensatory neuroplastic mechanisms, marking a potential conversion zone to clinical PD.

Subpanel III (PD progression) reflects neuroinflammatory intensity and vascular-tissue interactions. Composed of kininogen-1 and galectin-3-binding protein, exhibiting the largest negative slopes in both PD and prodromal groups compared to controls. These markers participate in broad inflammatory networks, correlating with neuroinflammatory intensity and cognitive decline over time36, thereby offering candidate metrics for tracking progressive neurodegenerative burden. Together, these markers provide complementary mechanistic insights that reflect neuroinflammatory activation and disease progression.

An additional six proteins demonstrated parallel trajectories between prodromal and PD stages, suggesting potential utility for long-term disease monitoring. However, these findings require validation in larger longitudinal cohorts to substantiate their clinical relevance (Fig. S3). These include PAM, cadherin-13, Chitinase-3-like protein 1, UDP-glycoprotein glucosyl transferase 1 and phosphatidylethanolamine-binding protein 4. These markers are broadly linked to extracellular homeostasis37 and proteostatic quality control6,38, though expanding longitudinal datasets will be required to define their precise temporal specificity.

Leveraging an integrative strategy enabled the identification of candidate biomarkers that could be used to monitor the entire life cycle of PD disease journey. However, additional studies are necessary to evaluate the utility of this approach in distinguishing prodromal cases from healthy controls. Prodromal PD is inherently heterogenous, with some individuals exhibiting molecular profiles that more closely resemble controls and others showing patterns more similar to established PD. Given this variability, large-scale cohorts are essential to accurately resolve a distinct signature associated with prodromal phase.

A limitation in our study is the absence of validation in a fully independent external cohort. Model performance was assessed using cross-validation in combination with a held-out test subset, while this approach provides an additional layer of internal validation and enables validation on unseen samples, it remains derived from the same underlying dataset and therefore does not fully reflect variability across independent populations. Consequently, the reported performance metrics may still be optimistic and further work on validating these findings in independent cohorts and across analytical platforms to assess robustness, reproducibility and generalizability is required to determine the translational potential of the identified molecular signatures for early diagnosis and disease monitoring.

In our longitudinal disease progression analysis, we investigated disease progression over a period of 16 months, while we acknowledged that this period is relatively short for a chronic progressive disorder such as PD, this time frame captures early disease dynamics, which may reflect initial pathological changes and compensatory biological mechanisms. Longer follow-up will be essential in validating our observations. Additionally, we employed strict inclusion criteria for the candidate biomarkers (≥5 patients with ≥3 visits), only 5 of the 21 candidate biomarkers had sufficient data in both CSF and plasma. Additional subsets of biomarkers met the inclusion criteria in only one biofluid (CSF or plasma), reflecting variability in data availability across cohorts. We prioritized biomarkers with measurements in both biofluids to enable cross-compartment comparisons. We acknowledge that this limited sample size may reduce statistical power and may increase variability in slope estimates, thereby affecting the robustness and generalizability of the findings. Additionally, we utilized a mix of statistical descriptive frameworks to classify candidate biomarkers into subpanels as the aim was to provide a biologically interpretable and hypothesis-generating approach for organizing candidate biomarkers according to their observed longitudinal behavior across cohorts, rather than to establish definitive diagnostic categories. These results should therefore be interpreted as exploratory and hypothesis generation. Further studies with larger cohorts and longer follow-up will be necessary to confirm these longitudinal trends suggested here.

This study illustrates that integrating multi-omics modalities across multiple biofluids improves predictive performance for some ML frameworks. Our data-driven approach maps candidate biomarker trajectories across the PD life cycle. Ultimately, these findings underscore the value of a multi-analyte panel that strategically aligns biomarker selection with the biological compartment most reflective of central disease pathology.

Methods

Data extraction and preparation

PPMI—ethical information

This study is a secondary computational analysis of existing data obtained from the PPMI database. No new experimental data was generated. This secondary computational analysis was conducted in compliance with PPMI data use agreements, internal institutional review board approval was granted by the American University in Cairo institutional review board (AUC-IRB) (Ethics Approval # 2021-2022-058 and 2021-2022-203) pertains solely to the use and analysis of existing data.

The original PPMI study was conducted in accordance with the Declaration of Helsinki and the Good Clinical Practice guidelines after approval of the local ethics committees of each of the participating sites. Written informed consent was obtained from all the participants in the PPMI study, including the genetic research part. The study was conducted under the ethical standards of the Helsinki Declaration of 1975.

PPMI—data collection

Data was obtained from the PPMI. Details on the specified experimental study methodology can be found at https://www.ppmi-info.org. Additionally, data was acquired from the Laboratory of NeuroImaging (LONI, www.loni.usc.edu) and are available for download by qualified investigators by making an access request.

PPMI—study participants

For proteomic and metabolomic datasets, we extracted data from 2 biofluids (CSF and Plasma) and 3 Cohorts (Control, Prodromal, and PD). Each cohort is defined as follows: Controls are participants who at baseline present no neurologic disorder(s) (current or clinically significant symptoms), no first-degree relative with PD, and normal dopamine transporter (DAT) imaging. PD cohort comprises participants with early, untreated sporadic PD or PD with pathogenic genetic variant(s). The Prodromal cohort contains patients who are at risk of PD based on clinical features, genetic variants, or other biomarkers.

Data preprocessing

The computational and statistical processing pipelines of proteomic and metabolomic datasets from both CSF and plasma were analyzed in parallel. All computational analyses were performed in R (version 4.5.1) using the most recent CRAN releases of relevant packages. Raw intensity values were log2-transformed (with a pseudo count of 1) to stabilize variance and approximate normality. Within each dataset (CSF-proteome, plasma-proteome, CSF-metabolome, plasma-metabolome), entries with >20% missing data were excluded, and duplicate measurements were aggregated using mean values creating entity level matrices for each proteomics and metabolomic, respectively. Clinical grouping variables were defined using diagnostic cohort labels (Control, PD, and Prodromal) which were then encoded as an unordered factor with control defined as the reference group in downstream contrast analyses. Samples were aligned with corresponding clinical metadata based on participant identifiers prior to statistical modeling to ensure correct sample-label matching. Samples with missing or inconsistent cohort annotations were excluded from the analysis.

Statistical modeling

Univariate analyses were conducted separately within each omics–biofluid dataset to detect disease-associated changes across diagnostic contrasts and account for differences in sample availability and avoid bias introduced by incomplete cross-biofluid matching. Differential abundance analysis was performed using a model-based method limma (Linear Models for Microarray Data) with explicit group factors (biofluid and cohort). Prior to modeling, features with excessive missingness or insufficient observation coverage were excluded to reduce variance estimation. Missing values were implicitly handled within the linear modeling framework under the assumption that data were missing at random, consistent with standard limma workflows for high-throughput molecular data. Differential abundance testing was conducted using limma empirical Bayes framework, where for each bio-fluid specific dataset, a design matrix was constructed using cohort as a categorical factor with three levels: Control, Prodromal, and PD. A no-intercept parameterization (~0+ group) was used to estimate group-specific means directly. Pairwise comparisons were defined using linear contrasts (Control vs. PD, Control vs Prodromal, PD vs Prodromal). Linear modeling was performed using lmFit, followed by contrast fitting and empirical Bayes moderation using eBayes. Multiple testing correction was performed using Benjamini-Hochberg false discovery rate (FDR) procedure. Features were considered significantly differentially abundant at FDR < 0.05, with an additional biological effect size threshold of log2FoldChange ≥ 1 applied for downstream interpretation. Batch effects were not explicitly modeled due to lack of technical covariates, where PCA-based quality control did not indicate dominant non-biological clustering. Therefore, batch correction was not applied, however, potential confounding by technical effects were mitigated by stratification by biofluid type, ensuring that systematic differences between CSF and plasma were not modeled jointly. While a combined interaction model including biofluid by cohort interactions was initially considered, however, due to incomplete overlap in feature coverage and missing observations across biofluids, a stratified modeling strategy was adopted, enabling robust within biofluid inference while preserving statistical power. For each contrast and biofluid, the number of significantly altered features passing FDR and effect size threshold was quantified. These counts were used to summarize the overall differential signal burden across cohort and biofluids.

To visualize patterns of differential abundance across samples, heatmaps were generated for each biofluid and diagnostic contrast using top ranked significant features identified from the differential analysis. Features were ranked by adjusted p-value and plotted using hierarchical clustering applied to features and samples, enabling unsupervised grouping based on similarity in expression profiles. Sample level metadata were incorporated through column annotations, specifically diagnostic cohort, allowing visual assessment of group specific clustering patterns.

Multivariate testing

While univariate analyses provide valuable insights into individual associations, they are limited in resolving interdependencies of biological data. Many molecular features co-vary, multivariate testing evaluates variables jointly, enabling the detection of cross-fluid patterns that otherwise might not be detected. In the present analysis, this framework is essential in distinguishing independent contributions from correlated features and constructing cross-fluid, cross-omics signatures with greater predictive and biological relevance. The detailed description workflow of the multivariate integrative ML approach can be found in Fig. 9.

Fig. 9. Integrated workflow for multi-omics biomarker discovery and validation.

Fig. 9

Data was downloaded, organized, and preprocessed through cleaning, label encoding, and normalization. Downstream analyses were performed using both single-omics and multi-omics approaches. Feature importance was first assessed within individual biofluids and modalities, followed by multi-fluid integration. Candidate biomarkers were identified through a combination of statistical testing and systematic feature selection using LASSO regression, with exclusion and reintegration of non-contributing features. Selected features were incorporated into stratified machine learning pipelines for predictive modeling. Candidate biomarkers were further validated using linear mixed-effects models to evaluate stage-specific progression patterns, enabling the identification of robust, disease stage–specific signatures across biofluids and omics layers.

Single-omics modeling

To assess the independent predictive value of each biomarker, single-omics classification models were trained separately for each dataset. This design yielded six model groupings: ProCSF, MetCSF, ProPlasma, MetPlasma, Pro_Combined, Met_Combined. For each omics, biofluids were analyzed separately and in combination. We benchmarked five ML algorithms, elastic net regularization (GLMNET), random forest (RF), support vector machine (SVM), Neural Networks (NN), and extreme gradient boosting (XGBoost) were implemented. Owing to the class imbalance in the datasets (2:1:1 – PD: Control: Prodromal), random down-sampling using SMOTE was implemented. Models were trained following an 80%-20% stratified data split and evaluated using stratified five-fold cross-validation to maintain diagnostic group balance. Performance was quantified using receiver operating characteristic (ROC) curves, area under the curve (AUC), precision-recall curves, and accuracy. Feature inputs were restricted to those of feature importance scores derived from RF, SVM, and XGBoost were compared with regression coefficients to evaluate consistency across models. Features from the highest performing models were subsequently prioritized for further integrative approach.

Integrative approach (multi-fluid + multi-omics)

Integrative analysis employed similar statistical parameters as the single-omic approach. First features were separated from sample identifiers and clinical metadata, retaining only the proteomic and metabolomic features needed. Proteomic and metabolomic features from both CSF and plasma were concatenated into combined feature matrices, and class labels were defined. To prevent data leakage, a stratified training and test split was performed (80–20%). Along with a held-out test set for downstream model validation, ensuring proportional representation of all cohorts in both sets. Candidate feature sets were reduced using the least absolute shrinkage and selection operator (LASSO) exclusively on the training set, retaining biomarkers with non-zero coefficients following five-fold cross-validation and stratified data split. Performance metrics reported are based on a held-out test subset.

Multi-omics classification was performed using generalized linear models with GLMNET, RF, SVM, and NN. All models were validated with stratified five-fold cross-validation, and predictive performance was compared across classifiers. Hyperparameters were tuned within each model class. For computational efficiency, additional lightweight models (RF and penalized logistic regression) were trained using reduced cross-validation folds and fewer trees. To improve interpretability, selected features from the best-performing models were cross-referenced with univariate significance and biofluid-specific consistency. Model performance was evaluated on held-out test sets. Predictions were generated for both probability estimates and class labels. Accuracy was computed directly, and per-class area under the ROC curve values were averaged to obtain macro-ROC scores. Confusion matrices were generated for the best-performing models, and performance was further characterized by using detailed per-class metrics, including AUC, sensitivity, specificity, and F1 scores (Table 6).

Table 6.

Definition of evaluation metrics used in machine learning algorithms

Performance metric Definitions
Sensitivity [%] TP/(TP + FN) × 100
Accuracy [%] (TP + TN)/(TP + FP + FN + TN) × 100
F1-score [%] (2 x TP)/(2 × TP + FP + FN) × 100
Specificity [%] TN/(TN + FP) × 100
Area under the ROC curve (AUC) [%] ∫01 TPR (fpr) d(FPR) × 100

TP true positive, TN true negative, TPR true positive rate, FP false positive, FN false negative, FPR false positive rate.

Feature importance was extracted from all classifiers using appropriate importance scores. Top features validated across at least 2 models were extracted for potential inclusion in the candidate biomarker panel. Importance scores were visualized and annotated for downstream progression analysis.

Longitudinal disease progression analysis

To assess longitudinal biomarker dynamics, we constructed a progression dataset from clinical and molecular profiling data encompassing both CSF and plasma from both proteomic and metabolomic data. Clinical visits were first enumerated to determine temporal coverage across cohorts, and visit distributions were assessed both at the cohort level and per patient. A standardized visit order (baseline through follow-up visits) was defined and mapped to approximate months from baseline, enabling consistent temporal alignment of proteomic and metabolomic measurements. Only visits with clear temporal order were retained. Patients with at least three longitudinal visits were included in progression analyses to ensure sufficient temporal resolution.

The resulting dataset was structured by patient, cohort, biofluid, and biomarker, with visit timepoints expressed as months from baseline. Biomarker values were log-transformed to normalize distributions and reduce skewness. To model progression, linear mixed-effects models were fit using patient-level random effects to capture within-patient correlation across visits. The primary model included fixed effects for cohort, time (months from baseline), and their interaction, with random slopes and intercepts specified per patient. This framework allowed testing for differential temporal trajectories across diagnostic groups while accounting for repeated measures within individuals. Statistical inference regarding differences in longitudinal trajectories was based on the mixed-effects model estimates and the Cohort × Time interaction term, where early diagnostic candidates showed significant prodromal vs control difference (FDR < 0.05) and no significant progression effect, the conversion candidates exhibited significant PD vs prodromal differences (FDR < 0.05), and the progression candidates showed significant Cohort x Time interaction in the mixed model (FDR < 0.05). Relative effect sizes (difference > 0.02 between respective groups) were used at specific disease stages to reflect magnitude of difference. Patient-level slope estimates were summarized descriptively within cohorts to facilitate visualization of progression heterogeneity.

Subsequent progression modeling focused on a curated candidate biomarker panel of features (proteins/metabolites), identified from a previous integrative modeling approach, and selected for predictive classification capabilities and relevance to disease biology. For each biomarker–biofluid combination, models were fit separately, and key parameters (including fixed effect estimates, patient counts, and variance components) were extracted. Only biomarkers with sufficient sample size were advanced to detailed modeling. A comprehensive summary plot was produced, displaying candidate biomarker-specific progression rates across cohorts, with standard errors and interquartile ranges to highlight variability.

The candidate biomarkers were subclassified into 3 subpanels (early diagnosis markers, prodromal to PD conversion, and PD progression) based on a combination of the direction and magnitude of longitudinal slopes, together with their differential behavior across clinical groups. Owing to the differences in scale and variability between analytes, classification was based on relative effect sizes (difference > 0.02 between respective groups) and consistency of slope direction, and statistical significance of longitudinal trends within each biomarker as described previously. Specifically, those in the early diagnosis subpanel exhibited consistent baseline group differences with relatively stable or non-progressive longitudinal trends. Those in the conversion subpanel were defined by directional changes in longitudinal slopes. Sensitive progression candidates showed consistent and measurable longitudinal change (non-zero slope) within PD patients, reflecting disease progression.

Delta progression analysis

To quantify disease-related differences in biomarker progression over time points, individual patient-level slopes (previously derived from linear regressions of log-transformed values against time) were summarized across diagnostic groups. For each biomarker and biofluid, cohort grouped patient-specific slopes, and summary statistics were computed (i.e., mean slope, standard error, and patient counts). Biomarkers represented by fewer than 20 patients per group were excluded to ensure sufficient statistical support.

For visualization and interpretation, cohort-level mean slopes were contrasted against the control group. Delta slopes were calculated as the difference between PD and control means, and between prodromal and control means, for each biomarker and biofluid. These values were assembled into a delta progression dataset and expressed as differences in log10-scaled progression rates per month relative to controls.

To illustrate these results, bar plots were generated showing cohort-specific deltas for each candidate biomarker, with separate panels for CSF and plasma (when available). Candidate biomarkers were then filtered down and categorized into stage-specific classes, including early diagnostic indicators, conversion-associated features (Prodromal to PD conversion), and progression-sensitive markers, based on their slope patterns and relevance to CNS-related functions across the prodromal, PD, and control groups in both CSF and plasma.

Supplementary information

Supplementary Information (236.2KB, pdf)
Supplementary Data (113KB, xls)

Acknowledgements

A.G. was supported by the Mohamed Bin Abdulkarim Allehedan Ph.D. fellowship. M.S. was supported by the Bartlett Fund for Critical Challenges. Agreement Number: 2—Cycle 3. The funding sources had no role in study design, review, data curation and analysis, interpretation, decision to publish, or preparation of the manuscript.

Author contributions

A.M. and M.S. planned and supervised the project. A.G. was responsible for the design, execution, writing of the manuscript draft and visualizations. A.G., A.M., and M.S. reviewed and edited the final manuscript.

Data availability

The data used in the present study are available from the PPMI (Parkinson’s Progression Markers Initiative) database is publicly available, and researchers can apply for access. To request data, please refer to the PPMI Data User Guide at: https://www.ppmi-info.org/sites/default/files/docs/PPMI%20Data%20User%20Guide.pdf. This guide provides detailed instructions on how to apply for access and outlines the associated usage policies.

Code availability

All computational analyses were performed using R software (v4.5.1) and Python (v3.13). Code is available upon request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s41531-026-01497-3.

References

  • 1.Hayes, M. T. Parkinson’s disease and Parkinsonism. Am. J. Med.132, 802–807 10.1016/J.AMJMED.2019.03.001 (2019). [DOI] [PubMed] [Google Scholar]
  • 2.Postuma, R. B. & Berg, D. Advances in markers of prodromal Parkinson disease. Nat. Rev. Neurol.12, 622–634 10.1038/NRNEUROL.2016.152 (2016). [DOI] [PubMed] [Google Scholar]
  • 3.Boura, I. et al. Prodromal Parkinson’s disease: a snapshot of the landscape. Neurol. Clin.43, 209–228 10.1016/J.NCL.2024.12.004 (2025). [DOI] [PubMed] [Google Scholar]
  • 4.Surmeier, D. J., Obeso, J. A. & Halliday, G. M. Selective neuronal vulnerability in Parkinson disease. Nat. Rev. Neurosci.18, 101–113 10.1038/NRN.2016.178 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Faizan, M. et al. Cerebrospinal fluid protein biomarkers in Parkinson’s disease. Clin. Chim. Acta. 556 10.1016/j.cca.2024.117848 (2024). [DOI] [PubMed]
  • 6.Kim, S. G. et al. Integrative metabolome and proteome analysis of cerebrospinal fluid in Parkinson’s disease. Int. J. Mol. Sci.25, 11406 10.3390/IJMS252111406/S1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Caudle, W. M., Bammler, T. K., Lin, Y., Pan, S. & Zhang, J. Using ‘omics’ to define pathogenesis and biomarkers of Parkinson’s disease. Expert Rev. Neurother.10, 925 10.1586/ERN.10.54 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Bao, Y., Wang, L., Yu, F., Yang, J. & Huang, D. Parkinson’s disease gene biomarkers screened by the LASSO and SVM algorithms. Brain Sci. 13 10.3390/BRAINSCI13020175 (2023). [DOI] [PMC free article] [PubMed]
  • 9.Krawczuk, D., Groblewska, M., Mroczko, J., Winkel, I. & Mroczko, B. The role of α-synuclein in etiology of neurodegenerative diseases. Int. J. Mol. Sci. 25 10.3390/IJMS25179197 (2024). [DOI] [PMC free article] [PubMed]
  • 10.Sackner-Bernstein, J. Rethinking Parkinson’s disease: could dopamine reduction therapy have clinical utility? J. Neurol.271, 5687 10.1007/S00415-024-12526-7 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Schilder, B. M., Navarro, E. & Raj, T. Multi-omic insights into Parkinson’s disease: from genetic associations to functional mechanisms. Neurobiol. Dis.163, 105580 10.1016/J.NBD.2021.105580 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Sharma, P. & Dhamija, R. K. The quest for Parkinson’s disease biomarkers: traditional and emerging multi-omics approaches. Mol. Biol. Rep. 52 10.1007/S11033-025-10929-X (2025). [DOI] [PubMed]
  • 13.Carrillo, F. et al. Multiomics approach discloses lipids and metabolites profiles associated to Parkinson’s disease stages and applied therapies. Neurobiol. Dis.202 106698 10.1016/J.NBD.2024.106698 (2024). [DOI] [PubMed] [Google Scholar]
  • 14.Di Biase, L. et al. Gait analysis in Parkinson’s disease: an overview of the most accurate markers for diagnosis and symptoms monitoring. Sensors20, 3529 10.3390/S20123529 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Sotirakis, C. et al. Identification of motor progression in Parkinson’s disease using wearable sensors and machine learning. NPJ Parkinsons Dis.9, 142 10.1038/S41531-023-00581-2 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Arnab, S. B. et al. Analysis of different modality of data to diagnose Parkinson’s disease using machine learning and deep learning approaches: a review. Expert Syst.42, e13790 10.1111/EXSY.13790 (2025). [DOI] [Google Scholar]
  • 17.Aihara, K. et al. A neuron-specific EGF family protein, NELL2, promotes survival of neurons through mitogen-activated protein kinases. Mol. Brain Res.116, 86–93 10.1016/S0169-328X(03)00256-0 (2003). [DOI] [PubMed] [Google Scholar]
  • 18.Phung, D. M. et al. Meta-analysis of differentially expressed genes in the substantia nigra in Parkinson’s disease supports phenotype-specific transcriptome changes. Front. Neurosci.14, 596105 10.3389/FNINS.2020.596105/FULL (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Mnich, K. et al. Endoplasmic reticulum stress-regulated chaperones as a serum biomarker panel for Parkinson’s disease. Mol. Neurobiol.60, 1476 10.1007/S12035-022-03139-0 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Kaiserova, M. et al. Cerebrospinal fluid levels of chromogranin A in Parkinson’s Disease and multiple system atrophy. Brain Sci.11, 141 10.3390/BRAINSCI11020141 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Liu, Y. et al. Chromogranin A promotes the pathological conversion of α-synuclein at the synapse in Parkinson’s disease. Cell Rep.44, 116562 10.1016/J.CELREP.2025.116562 (2025). [DOI] [PubMed] [Google Scholar]
  • 22.Packard, C. J., Pirillo, A., Tsimikas, S., Ference, B. A. & Catapano, A. L. Exploring apolipoprotein C-III: pathophysiological and pharmacological relevance. Cardiovasc Res119, 2843 10.1093/CVR/CVAD177 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Campos-Escamilla, C. The role of transferrins and iron-related proteins in brain iron transport: applications to neurological diseases. Adv. Protein Chem. Struct. Biol.123, 133–162 10.1016/bs.apcsb.2020.09.002 (2021). [DOI] [PubMed] [Google Scholar]
  • 24.Ayton, S., Lei, P., McLean, C., Bush, A. I. & Finkelstein, D. I. Transferrin protects against Parkinsonian neurotoxicity and is deficient in Parkinson’s substantia nigra. Signal Transduct. Target. Ther. 1 10.1038/SIGTRANS.2016.15 (2016). [DOI] [PMC free article] [PubMed]
  • 25.Tripathi, C. B., Gangania, M., Kushwaha, S. & Agarwal, R. Evidence-Based discriminant analysis: a new insight into iron profile for the diagnosis of Parkinson’s disease. Ann. Indian Acad. Neurol.24, 234–238 10.4103/AIAN.AIAN_419_20 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Morrison, J. I., Metzendorf, N. G., Liu, J. & Hultqvist, G. Serotransferrin enhances transferrin receptor-mediated brain uptake of antibodies. Drug Deliv. Transl. Res.15, 3321 10.1007/S13346-025-01811-1 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Yan, Z. et al. Alterations of gut microbiota and metabolome with Parkinson’s disease. Microb. Pathog. 160 10.1016/J.MICPATH.2021.105187 (2021). [DOI] [PubMed]
  • 28.Vascellari, S. et al. Gut microbiota and metabolome alterations associated with Parkinson’s disease. mSystems5 10.1128/MSYSTEMS.00561-20 (2020). [DOI] [PMC free article] [PubMed]
  • 29.Helwig, M. et al. The neuroendocrine protein 7B2 suppresses the aggregation of neurodegenerative disease-related proteins. J. Biol. Chem.288, 1114 10.1074/JBC.M112.417071 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Caldi Gomes, L. et al. MicroRNAs from extracellular vesicles as a signature for Parkinson’s disease. Clin. Transl. Med.11, e357 10.1002/CTM2.357 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Heywood, W. E. et al. Identification of novel CSF biomarkers for neurodegeneration and their validation by a high-throughput multiplexed targeted proteomic assay. Mol. Neurodegener.10, 64 10.1186/S13024-015-0059-Y (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Guo, Y. et al. Defining specific cell states of MPTP-induced Parkinson’s disease by single-nucleus RNA sequencing. Int. J. Mol. Sci.23, 10774 10.3390/IJMS231810774 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Bartl, M. et al. Lysosomal and synaptic dysfunction markers in longitudinal cerebrospinal fluid of de novo Parkinson’s disease. NPJ Parkinsons Dis.10 1–13 10.1038/S41531-024-00714-1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Rotunno, M. S. et al. Cerebrospinal fluid proteomics implicates the granin family in Parkinson’s disease. Sci. Rep.10, 1–11 10.1038/S41598-020-59414-4 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Karayel, O. et al. Proteome profiling of cerebrospinal fluid reveals biomarker candidates for Parkinson’s disease. Cell Rep. Med.3, 100661 10.1016/j.xcrm.2022.100661 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Costa, J. et al. Investigating LGALS3BP/90 K glycoprotein in the cerebrospinal fluid of patients with neurological diseases. Sci. Rep.10, 1–9 10.1038/S41598-020-62592-W (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Powers, K. G., Ma, X. M., Eipper, B. A. & Mains, R. E. Cell-type specific knockout of peptidylglycine α-amidating monooxygenase reveals specific behavioral roles in excitatory forebrain neurons and cardiomyocytes. Genes Brain Behav.20, e12699 10.1111/GBB.12699 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Qu, Y. et al. A systematic review and meta-analysis of inflammatory biomarkers in Parkinson’s disease. NPJ Parkinsons Dis.9, 18 10.1038/S41531-023-00449-5 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Zhang, Y., Yan, Y., Kong, X., Zhang, H. & Su, S. Potential cerebrospinal fluid metabolomic biomarkers and early prediction model for Parkinson’s disease. Front. Aging Neurosci.17 1582362 10.3389/FNAGI.2025.1582362/BIBTEX (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Zhang, J., Zhang, X., Wang, L. & Yang, C. High performance liquid chromatography-mass spectrometry (LC-MS) based quantitative lipidomics study of ganglioside-NANA-3 plasma to establish its association with Parkinson’s disease patients. Med. Sci. Monit.23, 5345 10.12659/MSM.904399 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Munoz-Pinto, M. F., Empadinhas, N. & Cardoso, S. M. The neuromicrobiology of Parkinson’s disease: A unifying theory. Ageing Res. Rev.70, 101396 10.1016/J.ARR.2021.101396 (2021). [DOI] [PubMed] [Google Scholar]
  • 42.Peng, C., Trojanowski, J. Q. & Lee, V. M. Y. Protein transmission in neurodegenerative disease. Nat. Rev. Neurol.16, 199 10.1038/S41582-020-0333-7 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Filippini, F. et al. Secretion of VGF relies on the interplay between LRRK2 and post-Golgi v-SNAREs. Cell Rep.42 112221 10.1016/J.CELREP.2023.112221 (2023). [DOI] [PubMed] [Google Scholar]
  • 44.Quinn, J. P., Kandigian, S. E., Trombetta, B. A., Arnold, S. E. & Carlyle, B. C. VGF as a biomarker and therapeutic target in neurodegenerative and psychiatric diseases. Brain Commun.3, fcab261 10.1093/BRAINCOMMS/FCAB261 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Barba, L. et al. CSF neurosecretory proteins VGF and neuroserpin in patients with Alzheimer’s and Lewy body diseases. J. Neurol. Sci.462, 123059 10.1016/J.JNS.2024.123059 (2024). [DOI] [PubMed] [Google Scholar]
  • 46.Verma, A. et al. In silico comparative analysis of LRRK2 interactomes from brain, kidney and lung. Brain Res. 1765 10.1016/J.BRAINRES.2021.147503 (2021). [DOI] [PMC free article] [PubMed]
  • 47.Killen, A. C. et al. Protective role of Cadherin 13 in interneuron development. Brain Struct. Funct.222, 3567 10.1007/S00429-017-1418-Y (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Hwang, H. et al. Glycoproteomics in neurodegenerative diseases. Mass Spectrom. Rev.29, 79 10.1002/MAS.20221 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Nikitovic, D., Papoutsidakis, A., Karamanos, N. K. & Tzanakakis, G. N. Lumican affects tumor cell functions, tumor-ECM interactions, angiogenesis and inflammatory response. Matrix Biol.35, 206–214 10.1016/J.MATBIO.2013.09.003 (2014). [DOI] [PubMed] [Google Scholar]
  • 50.Abdi, I. Y. et al. Cross-sectional proteomic expression in Parkinson’s disease-related proteins in drug-naïve patients vs healthy controls with longitudinal clinical follow-up. Neurobiol. Dis.177, 105997 10.1016/J.NBD.2023.105997 (2023). [DOI] [PubMed] [Google Scholar]
  • 51.Vicente, J. B. et al. Glycosyltransferase 8 domain-containing protein 1 (GLT8D1) is a UDP-dependent galactosyltransferase. Sci. Rep.13, 21684 10.1038/S41598-023-48605-4 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Anandhan, A. et al. Metabolic disorder dysfunction in Parkinson’s disease: bioenergetics, redox homeostasis and central carbon metabolism. Brain Res. Bull.133, 12 10.1016/J.BRAINRESBULL.2017.03.009 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.García-Revilla, J. et al. Galectin-3 shapes toxic alpha-synuclein strains in Parkinson’s disease. Acta Neuropathol.146, 51–75 10.1007/S00401-023-02585-X (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Cengiz, T., Türkboyları, S., Gençler, O. S. & Anlar, Ö. The roles of galectin-3 and galectin-4 in the idiopatic Parkinson disease and its progression. Clin. Neurol. Neurosurg. 184 10.1016/j.clineuro.2019.105373 (2019). [DOI] [PubMed]
  • 55.Markaki, I. et al. Cerebrospinal fluid levels of Kininogen-1 Indicate early cognitive impairment in Parkinson’s disease. Mov. Disord.35, 2101–2106 10.1002/MDS.28192 (2020). [DOI] [PubMed] [Google Scholar]
  • 56.Guo, J. et al. Proteomic analysis of the cerebrospinal fluid of Parkinson’s disease patients. Cell Res19, 1401–1403 10.1038/CR.2009.131 (2009). [DOI] [PubMed] [Google Scholar]
  • 57.Herrero, M. T., Estrada, C., Maatouk, L. & Vyas, S. Inflammation in Parkinson’s disease: role of glucocorticoids. Front. Neuroanat.9, 32 10.3389/FNANA.2015.00032 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Wojdała, A. L. et al. Phosphatidylethanolamine Binding Protein 1 (PEBP1) in Alzheimer’s Disease: ELISA Development and Clinical Validation. J. Alzheimers Dis.88, 1459–1468 10.3233/JAD-220323 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.George, A. J. et al. Decreased phosphatidylethanolamine binding protein expression correlates with Aβ accumulation in the Tg2576 mouse model of Alzheimer’s disease. Neurobiol. Aging27, 614–623 10.1016/J.NEUROBIOLAGING.2005.03.014 (2006). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (236.2KB, pdf)
Supplementary Data (113KB, xls)

Data Availability Statement

The data used in the present study are available from the PPMI (Parkinson’s Progression Markers Initiative) database is publicly available, and researchers can apply for access. To request data, please refer to the PPMI Data User Guide at: https://www.ppmi-info.org/sites/default/files/docs/PPMI%20Data%20User%20Guide.pdf. This guide provides detailed instructions on how to apply for access and outlines the associated usage policies.

All computational analyses were performed using R software (v4.5.1) and Python (v3.13). Code is available upon request.


Articles from NPJ Parkinson's Disease are provided here courtesy of Nature Publishing Group

RESOURCES