Skip to main content
Brain Pathology logoLink to Brain Pathology
. 2026 Mar 3;36(5):e70089. doi: 10.1111/bpa.70089

Proteomic profile of CSF obtained at the time of diagnosis determines amyotrophic lateral sclerosis progression and survival: CXCL7 levels in disease prognosis and survival

Sergio Roca‐Pereira 1,2, Yaiza López‐Sampere 1,2, Pol Mengod‐Soler 1,2, Mario Peña‐Fonteboa 1, Carla Marco 1,3, Alejandro Caravaca‐Puchades 1,2,3, Raúl Domínguez 1,2,3, Juan Francisco Vázquez‐Costa 4,5, Enrique Santamaría 6, Joaquín Fernández‐Irigoyen 6, María J Colomina 7, Isabel León Moreno 1,8, Antonio Martínez Yélamos 1,8,9, Sergio Martínez Yélamos 1,8,9, Gabriel Santpere 10, Elia Obis 11, Manuel Portero‐Otín 11, Gerard Piñol‐Ripoll 12,13,14, Mónica Povedano 1,2,3, Pol Andrés‐Benito 2,12,
PMCID: PMC13429306  PMID: 41776751

Abstract

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease primarily affecting motor neurons. Neurofilament light chain (NfL) is the most established prognostic biomarker; however, its diagnostic resolution is limited, particularly within intermediate concentration ranges, and it does not capture the molecular heterogeneity of ALS. This study aimed to identify complementary cerebrospinal fluid (CSF) biomarkers and pathway‐specific signatures through a non‐targeted multiomic approach. We performed SWATH‐MS‐based proteomics and LC–MS/MS lipidomics on CSF from ALS patients stratified by survival (ALS‐SS and ALS‐LS) and healthy controls. Weighted protein co‐expression network analysis (WPCNA) was applied to identify biologically coherent protein modules associated with disease phenotype and progression. Top biomarker candidates were further evaluated using immunoassays in an independent cohort. Post‐mortem ALS spinal cord tissues were analyzed to explore the pathophysiological relevance of identified proteins. CSF proteomic profiles robustly distinguished ALS patients from controls and stratified patient subgroups by survival, revealing a molecular signature characterized by inflammation, downregulation of detoxification mechanisms, and synaptic dysregulation in aggressive disease forms. In contrast, lipidomic profiles showed limited discriminatory power. WPCNA identified modular proteomic signatures capturing ALS heterogeneity, and machine learning models based on these profiles yielded optimal biomarker panels for diagnosis and prognosis. CXCL7 emerged as a promising complementary biomarker, and shed light in disease physiopathology. Immunoassay validation supported the diagnostic and prognostic potential of CXCL7 and its association with survival time. Histopathological analysis further confirmed CXCL7 localization in anterior horn motor neurons, despite no detectable changes in whole spinal cord lysates at late disease stages. Comprehensive CSF proteomic profiling, combined with network‐based analysis, enhances our understanding of ALS molecular heterogeneity and provides a framework for precision biomarker discovery. CXCL7 complements NfL as a diagnostic and prognostic biomarker, supporting improved patient stratification and advancing the development of personalized therapeutic strategies in ALS.

Keywords: ALS, biomarkers, CSF, CXCL7, lipidomics, prognosis, proteomics


Untargeted multiomic profiling of cerebrospinal fluid reveals that proteomic, but not lipidomic, signatures robustly distinguish ALS patients from controls and stratify individuals by survival, highlighting marked molecular differences between short survival and long survival disease. Network‐based and machine learning analyses identify CXCL7 as a novel biomarker with reduced CSF levels in ALS, associated with disease progression (fast or slow) and survival. Post‐mortem validation localizes CXCL7 to anterior horn motor neurons, linking its CSF alterations to ALS pathophysiology and supporting its role as a complementary diagnostic and prognostic biomarker alongside NfL.

graphic file with name BPA-36-e70089-g008.jpg

1. INTRODUCTION

Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease that affects upper motor neurons (UMN) of the primary motor cortex and lower motor neurons (LMN) of the lower brainstem and anterior horn of the spinal cord [1, 2]. This causes rapidly progressive muscle weakness and atrophy that ultimately lead to death within 5 years in most of the affected individuals [3]. ALS is a clinically heterogeneous neurodegenerative disease, presenting with considerable variability in onset, progression, symptoms, and prognosis [4]. The disease can begin in different regions—most commonly the limbs (limb‐onset) or the bulbar muscles (bulbar‐onset), with the latter often progressing more rapidly. ALS also varies in the degree of upper and lower motor neuron involvement, giving rise to subtypes such as primary lateral sclerosis (PLS) (UMN‐predominant), progressive muscular atrophy (PMA) (LMN‐predominant), and flail limb variants, each with distinct prognoses. Cognitive and behavioral changes occur in up to 50% of patients, with a subset developing frontotemporal dementia (ALS‐FTD), particularly in association with certain genetic mutations like C9ORF72. While most cases are sporadic, about 5%–10% are familial, involving mutations such as SOD1, TARDBP, and FUS, which influence both clinical presentation and progression [5, 6]. Age at onset, sex, genetic background, and site of symptom emergence all contribute to the wide clinical spectrum observed in ALS, complicating diagnosis and management but also highlighting the need for personalized approaches to care and research [4].

In addition, although different mechanisms have been postulated as main contributors to disease etiopathology such as neuronal and oligodendrocyte dysfunctions, mitochondrial damage, defective axonal transport, oxidative stress, or neuroinflammation, among others [7, 8, 9], it remains unknown whether these alterations are a consequence of neuronal injury or whether they actively contribute to the development and progression of the disease. Moreover, it is not known whether these changes are specific to particular phenotypic outcomes or shared across different clinical forms.

As a result of this variability, the diagnosis of ALS relies on the detection of clinical or neurophysiological signs of upper and/or lower motor neuron involvement, combined with additional tests (such as biochemical analyses, EMG, MRI, and genetic testing) to rule out other conditions. On average, it takes 10–16 months from symptom onset to confirm an ALS diagnosis [10]. Biomarkers may help at this point, reflecting biological pathways specifically altered in a particular disease or being reactive to disease progression. They can be used to facilitate earlier diagnosis, predict disease prognosis, identify novel therapeutically targetable pathways, and anticipate pharmacological responses [11].

After years of research, the incorporation to the clinical practice of neurofilament light chain (NfL) levels in cerebrospinal fluid (CSF) or serum as a diagnostic and prognostic biomarker in neurological diseases [12, 13, 14, 15, 16, 17], such as ALS, has been proposed. NfLs are neuron‐specific cytoskeletal proteins, their levels increasing in biological fluids proportionally to the degree of axonal damage. However, despite its numerous advancements in the clinical arena, NfL does have certain limitations as a biomarker. As a diagnostic biomarker, NfL has a notable limitation in its ability to distinguish between neurological disorders [18, 19], including ALS, from other mimicking diseases, such as Guillain‐Barré Syndrome, myelopathies, or multiple sclerosis (MS) with motor symptoms, among others [20, 21, 22]. This is because all these conditions exhibit elevated levels of NfL compared to healthy individuals, due to the common characteristic of axonal degeneration, resulting in overlapping among them [12, 23], in addition to other co‐morbidities that can influence NfL levels [24, 25]. Thus, NfL has a limited diagnostic value when weighted against clinical expertise [26]. In addition, although NfL has emerged as a strong prognostic marker, in the mid‐range of values its prognostic value is limited [27]. Axonal degeneration is not a specific alteration for disease or disease phenotypes.

To overcome these limitations and identify physiopathological signatures associated with different phenotypical outcomes in ALS, the present study aims to discover new specific altered pathways and biomarkers using a multiomics approach combining non‐targeted proteomics and lipidomics analyses of CSF from sporadic ALS patients with extreme survival times, as well as from control subjects. By applying a multiomics strategy, untargeted proteomics can reveal dysregulated proteins involved in pathways, such as neuroinflammation or synaptic components, while untargeted lipidomics can uncover alterations in lipid metabolism and membrane integrity—both of which are increasingly recognized as key contributors to ALS pathogenesis [28, 29]. This approach led to identifying several specific altered pathways components and promising candidates using machine learning models. The potential candidates that exhibited biological and statistical significance were validated with more precise molecular biology techniques, such as ELISA tests, and assessed as potential biomarkers. Finally, proteins showing significant differences in ELISA results were investigated in human post‐mortem tissue using western blot, qPCR, and immunohistochemistry to gain insights into their role in the disease's pathophysiology.

2. MATERIALS AND METHODS

2.1. Study design, participants, and clinical data collection

We designed an exploratory retrospective, longitudinal, observational, case and control study including patients diagnosed with ALS, who were prospectively followed up in the ALS unit of a tertiary hospital with available CSF samples at the time of initial patient evaluation and diagnosis, when clinical signs start. The STROBE guidelines for observational studies were used to conduct this study.

Sporadic ALS patients were diagnosed according to updated El Escorial criteria [30]. Patients were evaluated clinically and categorized according to disease survival as short survival (SS) (survived less than 3 years) and long survival (LS) (still alive after 5 years). The ALS Functional Rating Scale Revised (ALS‐FRS‐R, version May 2015) was used to calculate the disease progression rate (DPR) (the change in ALS‐FRS‐R/month) at the time of diagnosis. Fast progressors (FP) were defined as patients with a DPR ≥1.5 points/month. In contrast, slow progressors (SP) were defined as patients with a DPR ≤1 point/month. CSF from healthy control cases was obtained from patients at the need for knee surgical procedures under spinal anesthesia without neurological or infection conditions in the Traumatology Unit of the Bellvitge Hospital. They were selected by sex and age matching the patient groups to avoid possible bias.

The exclusion criteria were the presence of neurological comorbidities (peripheral or central nervous system). In addition, only ALS patients who did not need respiratory support or gastrostomy during the first 5 years of the disease were included. Only patients without a family history of ALS or mutations in C9ORF72, TARDBP, ATXN2, FUS, UNC13A, and SOD1 were included. In both patients and control groups, cases with samples not meeting the technical criteria (presence of particles or hemolyzed samples) were also excluded. The variables collected were sex, date of birth, date of onset of ALS, and CSF sampling; and ALS‐FRS‐R at the diagnose time.

Two different cohorts were used for the study: (a) discovery cohort, and (b) validation cohort. In the discovery cohort, a total of 15 healthy control patients and 28 ALS patients were included (Table 1). Among the ALS group, 15 patients were classified as SS and FP as demonstrated by the survival time (<3 years) and the DPR (≥1.5 points); while 13 patients were LS (>5 years) and SP as demonstrated by the DPR (≤1/month). The study groups were restricted to spinal‐onset ALS cases (including 2 spinal PMA cases in ALS‐LS group). This choice was primarily driven by sample availability, as sufficient numbers of well‐characterized CSF samples from non‐spinal phenotypes were not available to support robust subgroup analyses. This constraint, however, also offered an important analytical advantage. ALS is characterized by substantial phenotypic heterogeneity, and inclusion of bulbar or other non‐spinal forms may introduce additional clinical and biological variability, potentially obscuring molecular signals related to disease progression. By focusing on spinal‐onset cases, we were able to assemble a more homogeneous cohort, thereby reducing background noise and increasing sensitivity for the detection of progression‐associated molecular differences within the extreme phenotype framework adopted in this study.

TABLE 1.

List of cases used for CSF studies. Including: Group, number of cases (n), age (mean ± standard deviation), sex (female/male), disease subtypes according to prognosis and number of associated cases.

Group n Age Sex (f/m) Disease progression subtype
Discovery cohort
HC 15 65 ± 11.5 (8/7)
sALS 28 63 ± 12.3 (13/15) SS (15/28)
LS (13/28)
Discovery and validation cohort
HC 28 70 ± 8.57 (16/12)
sALS 54 64 ± 12.9 (24/30) SP (27/54)
FP (27/54)
IP 26 50 ± 20.1 (8/18)
Myelopathy 14 51 ± 17.6 (9/5)

Abbreviations: ALS‐SS, amyotrophic lateral sclerosis short‐survival; ALS‐LS, amyotrophic lateral sclerosis long survival; ALS‐FP, amyotrophic lateral sclerosis fast progressors; ALS‐SP, amyotrophic lateral sclerosis slow progressors; HC, healthy control; IP, inflammatory polyneuropathy; sALS, sporadic amyotrophic lateral sclerosis.

Due to the difficulty in obtaining more samples from long‐survival patients, as they represent only 20% of cases and require 5 years of disease progression, we validated our proteomic cohort based on progression rates, adding fast‐progressor (survived less than 3 years) and slow‐progressor patients, who share the same progression pattern as the survival groups. The cohort used for the validation of proteomics in CSF included a total of 82 samples; 54 were sporadic ALS cases (27 slow progression (SP) (22 spinal and 5 bulbar) and 27 fast progression (FP) (19 spinal, 7 bulbar and 1 respiratory)) and 28 were healthy controls. The samples used for the proteomic analysis (discovery cohort) were also part of this validation cohort.

In addition, CSF samples from mimic disorders from the Neurology Service of the Bellvitge Hospital were used to test validated biomarkers for differential diagnosis applications; these included 14 samples from patients with myelopathies and 26 samples from patients with inflammatory polyneuropathy (IP). A summary of cases is exposed in Table 1.

2.2. CSF sample collection

CSF was prospectively collected from patients undergoing lumbar puncture, at the first visit, as part of routine clinical assessment due to clinical suspicion of motor neuron disease at the functional unit of ALS of the Neurology Service of the Bellvitge University Hospital and Hospital la Fe de Valencia. In all patients, 1.5 ± 0.5 mL of CSF was collected in polypropylene tubes. CSF was centrifuged at 2000 rpm for 1 min at room temperature in the range of 30 min and 180 min after obtention. The supernatant was collected and aliquoted in volumes of 220 μL and stored at −80°C until use.

2.3. CSF proteomics analysis

CSF samples from control Raw files were processed with MaxQuan v2.0.1 [31] using the integrated Andromeda Search engine [32]. All data were searched against a target/decoy version of the Human Uniprot Reference Proteome with isoforms (Proteome ID: UP000005640_9606; March 2021). First search peptide tolerance was set to 20 ppm, main search peptide tolerance was set to 4.5 ppm. Fragment mass tolerance was set to 20 ppm. Trypsin was specified as enzyme, cleaving after all lysine and arginine residues, and allowing up to two missed cleavages. Carbamidomethylation of cysteine was specified as fixed modification and peptide N‐terminal acetylation, oxidation of methionine, deamidation of asparagine, glutamine and pyro‐glutamate formation from glutamine and glutamate, and phosphorylation of serine, threonine and tyrosine were considered variable modifications, with a total of 2 variable modifications per peptide. “Maximum peptide mass” were set to 7500 Da, the “modified peptide minimum score” and “unmodified peptide minimum score” were set to 25 and everything else was set to the default values, including the false discovery rate (FDR) limit of 1% on both the peptide and protein level nd ALS cases were mixed with lysis buffer containing 7 M urea, 2 M thiourea, and 50 mM DTT (v/v; 1:1). Protein extraction, in‐solution digestion, peptide purification, and reconstitution prior to mass spectrometric analysis were performed as previously described [33]. The Perseus software (version 1.6.14.0) [34] was used for statistical analysis and data visualization of proteomic data.

Two‐sample t‐test based on permutation‐based FDR statistics was applied (250 permutations); proteins with a p‐value lower than 0.05, and absolute fold‐change of <0.77 (down‐regulation) or >1.3 (up‐regulation) in linear scale were considered significantly differentially expressed and proposed for further evaluation. Functional analysis of significant deregulated proteins was performed using ShinyGO (v0.80) software [35]. ShinyGO's ranking system for gene enrichment analysis prioritizes results based on statistical significance (p‐value and FDR) and the magnitude of enrichment (fold enrichment). Here, the “Select by FDR, sort by fold enrichment” method was applied, which first filters gene sets based on FDR significance and then ranks the filtered results by fold enrichment.

2.3.1. Weighted protein Coexpression network analysis

Weighted Protein Coexpression Network Analysis (WPCNA) was done in R using the WGCNA library [36]. We first constructed a protein coexpression network based on pair‐wise correlation of protein expression using all samples at the same time. Therefore, all subsequent analysis based on the coexpression network constructed with all samples used a soft‐power threshold of 12. We identified modules of proteins based on their topological overlap dissimilarity with their connection strengths in the weighted network [37, 38]. Using the dynamic tree‐pruning algorithm, 16 modules were obtained with highly correlated eigengenes (Pearson correlation >/= 0.9); proteins not assigned to any module were labeled in gray. Module eigengenes were correlated to clinical groups. The p‐values were obtained from a general univariate lineal model. GO enrichment analysis was performed using GOstats [39]. Differentially expressed proteins with corrected p‐values <0.05 for each contrast were used for functional analysis. All proteins belonging to a particular protein coexpression module were also used for independent functional analyses. P‐values for categories were adjusted considering FDR using BH with the p.adjust function in R.

2.4. CSF lipidomic analysis

The lipidomic analysis was based on a previously validated methodology [40]. Briefly, we performed.

2.4.1. Preparation of lipid standards

Lipid standards consisting of isotopically labeled lipids were used for external standardization (i.e., lipid family assignment) and internal standardization (i.e., for adjustment of potential inter‐and intra‐assay variances). Stock solutions were prepared by dissolving standards in methyl tert‐butyl ether (MTBE) at a concentration of 1 mg/mL, and working solutions were diluted to 2.5 μg/mL in MTBE. These internal standards of each lipid class are added before lipid extraction (see below).

2.4.2. Lipid extraction

Five microliters of Mili‐Q water and 20 μL of methanol were added to 10 μL of the CSF sample and vortex‐mixed for 2 min to precipitate proteins. For lipid extraction, 250 μL of MTBE containing internal standards was added, and tubes containing the samples were immersed in a water bath (ATU Ultrasonidos, Valencia, Spain) with an ultrasound frequency and power of 40 kHz and 100 W at 10°C for 30 min. Then, 75 μL of Mili Q water was added to the mixture, and the organic phase was separated by centrifugation at 3000× rpm at 10°C for 10 min. 170 μL of the upper phase containing lipid extracts was collected and stored in vials at −20°C. A pool of 30 μL obtained from all samples was used as quality control (QC sample) in multivariate analyses.

2.4.3. LC–MS method

Lipid extracts were analyzed by LC/MS using a liquid chromatograph Agilent UPLC 1290 coupled to a mass spectrometer Agilent Q‐TOF MS/MS 6520 (Agilent Technologies, Barcelona, Spain). The analysis was based on a published method [34]. The order for the injection of samples was randomized, and QC samples were distributed every seven samples to control instrumental drift. Data were collected with MassHunter Data Acquisition software (Agilent Technologies, Barcelona, Spain). Relative SD was calculated from the intensity of these internal standards across the QC following [41], and median values of this relative SD allowed us to estimate the interassay and intraassay coefficient of variance, which was below 9%.

2.4.4. Data processing and statistical analyses

Before data processing, we checked each lipid standard's relative abundances in all the samples for the quality assessment using specific software. All samples shown passed this quality assessment, supporting the robustness of the analyses. According to the specific metabolomics software, the intensity of features (measured as arbitrary units) is normalized using one representative internal standard (MassProfiler Professional, Agilent). The sample with the highest internal standard abundance is used as a reference, and the other samples are scaled, multiplying by the relative internal standard abundance of each sample compared to the reference. All investigation was performed blinded. After all the analyses were completed, identities of clinical characteristics were attributed to the corresponding samples after grouping the data. Molecular features (i.e., groups of ions with similar chromatographic behavior and compatible with the same molecular composition, accounting adduct formation) were extracted with MassHunter Qualitative Analysis (Agilent Technologies, Barcelona, Spain), as previously detailed [42]. Molecular features in the same retention time [i.e., within a 0.1% of total run time (±0.25 min)] and in the same mass window [i.e., within 30 ppm (±2 mDa)] were considered the same to account for instrumental drift. We chose the 30 ppm tolerance based on actual conditions obtained by authentic standards. Only common features (found in at least 50% of the samples of the same condition compared) were considered to minimize individual bias. Peak intensities were relativized by internal standard peak intensity. In order to choose the most reliable identity of the lipid when there was no MS/MS confirmation and there was more than one option, we followed the following criteria according to previous experience and the composition of the mobile phases: for positive ionization, M + H, M + NH4, M + H − H2O, and M + Na were employed; for negative ionization, M − H, and M + CH3OO were used. Mass to charge ratios (m/z) shown are those majoritarian in the ion profile of given species, and they could be derived from molecular ion (i.e., +H+ in positive ionization or –H+ in negative ionization) or ionization of adducts (e.g., +NH4+ in positive ionization or + COOH− in negative ionization).

Text files (.txt) containing relative intensities of every metabolite in every sample were formatted for later statistical analysis with specific R‐based scripts. Non‐parametric tests (Mann–Whitney and Kruskal–Wallis tests), Spearman correlations, and multivariate analyses were performed with these scripts, and ROC curves were obtained using Metaboanalyst software [43]. In univariate analyses, Benjamini–Hochberg correction was adopted to correct for type 1 error, being shown as FDR corrected p‐value. Based on the semiquantitative nature of non‐targeted lipidomics, fold‐change differences are rounded to <0.1 or >10 in the figures to avoid the introduction of overestimation of differences.

2.4.5. Annotation and pathway analysis

Molecules showing statistically significant expression (with p‐value <0.05 in differential analyses) were annotated by comparing their exact mass, retention time, and isotopic distribution with specific databases [44, 45] to obtain potential identities. Identities were confirmed by comparing resulting MS/MS spectra against representative class standards using the LipidMatch workflow [46]. The MS/MS spectrum allowed the annotation of lipid families based on the presence of specific fragments. Annotation of lipid species has been performed based on exact mass and retention time, which allows the acyl chain's total sum report. In all cases, results and discussion comprise both positively and negatively ionized lipids.

Pathway analyses were performed using the LIPEA platform [47] and the Homo sapiens as lipidome background, and by employing the enrichment analysis option in the Consensus PathDB‐human platform [48] against the Kyoto Encyclopedia of Genes and Genomes (KEGG) database [49], with a minimum overlap of two lipids with those of reference, and a p‐value cutoff of 0.01.

2.5. Omic data machine learning processing

In the context of proteomic biomarker discovery, PLS‐DA and Random Forest were selected due to their complementary strengths in handling complex, high‐dimensional datasets typical of mass spectrometry‐based studies.

2.5.1. Partial least‐squares discriminant analysis (PLS‐DA)

PLS‐DA is particularly well suited for dimensionality reduction and classification in datasets where the number of variables (proteins) far exceeds the number of samples. It projects high‐dimensional data into a lower‐dimensional space by maximizing the covariance between protein expression levels and the predefined class labels (e.g., disease vs. control). This not only facilitates classification but also highlights proteins contributing most significantly to the observed group separation, aiding in biomarker selection [50], as well as for classification in a supervised manner [51, 52, 53]. Furthermore, PLS‐DA is robust to multicollinearity, a common feature in proteomic data due to co‐regulated protein expression. Variable importance parameters (VIP) scores estimate the importance of each variable, such as a metabolite feature, in the PLS‐DA model. These scores indicate how much each molecule contributes to the classification between different groups. A higher VIP score signifies that the variable plays a more influential role in explaining the separation among groups. Here, we applied a VIP score threshold greater than 1.5 to increase stringency and boost confidence in selecting relevant features, and we ranked the top 25 compounds based on their contribution to the model.

PLS‐DA was applied, as the proteomic dataset featured a moderate number of quantified proteins with relatively low collinearity and acceptable sample‐to‐variable ratio. This allowed us to retain the full set of variables and explore global patterns associated with ALS diagnosis and progression while ensuring model stability. In contrast, for the lipidomics data, we used sparse PLS‐DA (sPLS‐DA), a penalized version of PLS‐DA, due to the higher dimensionality and strong inter‐variable correlations inherent in lipidomic profiles. sPLS‐DA enabled simultaneous classification and feature selection, enhancing model interpretability and reducing the risk of overfitting. Both models were implemented using cross‐validation procedures to assess classification performance and identify the most discriminative molecular features for each omic layer.

2.5.2. Random forest regression model

We used the random forest regression (RFR) model to predict the diagnosis and progression of ALS patients using proteomic and lipidomic CSF profiling at the beginning of the disease. We implemented RFR as a supervised machine learning algorithm that constructs an ensemble of decision trees, using bootstrap aggregation and random feature selection to enhance classification accuracy and robustness [54, 55]. The RF algorithm was implemented by the Metaboanalyst software [43]. Data preprocessing was conducted within the software, including transformation and autoscaling to stabilize variance and normalize the intensity distributions. Missing values were imputed using the platform's default approach, which replaces missing entries with a small value calculated as half the minimum of the non‐zero entries for each feature. Data filtering was applied to remove proteins with low interquartile range variability, reducing noise and dimensionality. The RF model was constructed using 500 decision trees with default parameters. For classification, the software applies internal cross‐validation by using out‐of‐bag (OOB) error estimation, which provides an unbiased estimate of model accuracy. Classification performance was assessed by OOB error rate, accuracy, sensitivity, and specificity metrics.

In the RF model, feature importance was assessed using the mean decrease accuracy (MDA) method. This approach estimates the contribution of each feature to the predictive performance of the model. Specifically, after training the RF classifier, the values of each feature are randomly permuted across samples, one feature at a time, while keeping the others unchanged. The model's accuracy is then re‐evaluated on the OOB data. A substantial decrease in accuracy following permutation indicates that the feature was important for the model's predictive capability. Features were ranked according to the magnitude of this decrease, with higher MDA values indicating greater importance in distinguishing between classes. In this case, the top 15 proteins were ranked by their importance scores and were interpreted as possible candidate biomarkers. A misclassification table was generated to show the proportion of correctly classified observations in the data set.

2.6. CSF biomarkers candidate analysis

C‐X‐C motif chemokine ligand 7 (CXCL7) was quantified using the Human CXCL7 Quantikine ELISA Kit (NBP1‐89921) from R&D Systems according to the manufacturer's instructions (R&D Systems, Inc., Minneapolis, MN, USA). In addition, NfL was quantified as a standardized biomarker for neuroaxonal damage in several neurological disorders. NfL levels were detected in CSF using the ELISA Neurofilament Light Kit Assay from UmanDiagnostics according to the manufacturer's instructions (UmanDiangostics (Quanterix Company), Umeå, Sweeden). All samples were analyzed at the second thaw and freeze cycle. Test performers were blinded to clinical information.

2.7. Human tissue samples: Spinal cord and frontal cortex area 8

Post‐mortem samples were obtained from the Institute of Neuropathology HUB‐ICO‐IDIBELL Biobank (C.0008091) following the guidelines of Spanish legislation on this matter and the approval of the CEIC of the Bellvitge University Hospital (Ref. PR010/22). The post‐mortem interval between death and tissue processing was between 2 and 17 h and not belong to CSF patient's cohort. Tissue processing was carried out as detailed elsewhere [56]. All cases met the neuropathological criteria for classical ALS. Detailed clinical information on DPR (fast vs. slow) was not available for all post‐mortem cases. Patients with any associated pathology were excluded. Age‐matched cases, who had not suffered from neurologic or psychiatric diseases and did not have abnormalities in the neuropathological examination, were also assessed as controls. A summary of ALS and control cases is shown in Tables 2, 3, 4.

TABLE 2.

Summary of the 39 cases analyzed by q‐PCR corresponding to the spinal cord of 17 controls (age 61.8 ± 9.71 years) and 22 sALS cases (age 65.3 ± 10.13 years).

Case Age Gender Diagnosis PM delay RIN
1 70 M ALS 03 h 00 min 7.00
2 77 M ALS 04 h 30 min 7.50
3 83 F ALS 15 h 15 min 7.20
4 56 F ALS 03 h 45 min 8.10
5 56 M ALS 10 h 50 min 6.60
6 76 M ALS 12 h 40 min 7.00
7 69 M ALS 02 h 00 min 7.00
8 63 F ALS 13 h 50 min 6.50
9 N/A M ALS N/A 8.70
10 65 F ALS 04 h 10 min 7.70
11 50 M ALS 10 h 10 min 5.30
12 71 M ALS 03 h 25 min 8.10
13 54 M ALS 04 h 50 min 8.80
14 64 M ALS 16 h 30 min 6.70
15 75 F ALS 04 h 05 min 8.50
16 76 F ALS 13 h 00 min 8.10
17 57 F ALS 10 h 00 min 7.00
18 79 F ALS 02 h 10 min 8.10
19 57 F ALS 04 h 00 min 6.20
20 46 M ALS 07 h 00 min 7.00
21 69 F ALS 17 h 00 min 6.40
22 59 M ALS 03 h 15 min 6.80
23 66 M Control 14 h 00 min 5.00
24 46 M Control 15 h 00 min 5.70
25 66 M Control 05 h 00 min 5.40
26 77 F Control 08 h 30 min 5.10
27 64 F Control 05 h 00 min 7.00
28 60 F Control 09 h 40 min 5.80
29 52 M Control 03 h 00 min 5.00
30 67 M Control 07 h 00 min 5.50
31 47 M Control 04 h 55 min 5.60
32 64 F Control 11 h 20 min 6.20
33 56 M Control 07 h 10 min 6.10
34 71 F Control 08 h 30 min 5.90
35 55 M Control 09 h 45 min 5.30
36 75 M Control 07 h 30 min 6.60
37 51 F Control 04 h 00 min 6.30
38 59 M Control 12 h 05 min 6.40
39 75 F Control 10 h 30 min 5.20

Abbreviations: ALS, amyotrophic lateral sclerosis; F, female; M, male; N/A, not available; PM, post‐mortem delay (hours, minutes); RIN, RNA integrity number; SC, anterior horn of the spinal cord lumbar level.

TABLE 3.

Summary of the 32 cases analyzed by q‐PCR corresponding to the frontal cortex area 8 of 17 control cases (age 64.8 ± 8.9 years) and 15 sALS cases (age 61 ± 8.63 years).

Case Sex Age Diagnosis PMD RIN
1 M 66 Control 18 h 00 min 6.4
2 M 61 Control 03 h 40 min 7.0
3 M 62 Control 05 h 45 min 5.0
4 M 74 Control 06 h 40 min 7.2
5 M 65 Control 05 h 15 min 6.8
6 F 64 Control 02 h 15 min 5.0
7 M 63 Control 08 h 05 min 7.1
8 F 79 Control 03 h 35 min 6.8
9 F 67 Control 05 h 20 min 6.2
10 M 70 Control 03 h 45 min 7.2
11 M 52 Control 04 h 40 min 7.2
12 F 52 Control 05 h 45 min 5.1
13 F 82 Control 07 h 35 min 5.2
14 F 74 Control 02 h 45 min 5.7
15 M 55 Control 05 h 40 min 7.7
16 M 59 Control 07 h 05 min 7.8
17 M 56 Control 03 h 50 min 7.6
18 M 70 ALS 03 h 00 min 7.0
19 F 56 ALS 03 h 45 min 7.7
20 M 59 ALS 03 h 15 min 7.7
21 F 63 ALS 13 h 50 min 8.2
22 F 59 ALS 14 h 15 min 6.7
23 M 54 ALS 04 h 50 min 7.8
24 M 76 ALS 12 h 40 min 7.4
25 M 64 ALS 16 h 30 min 7.3
26 F 57 ALS 04 h 00 min 8.6
27 F 75 ALS 04 h 05 min 6.8
28 F 57 ALS 10 h 00 min 7.1
29 M 50 ALS 10 h 10 min 5.9
30 F 59 ALS 02 h 30 min 7.5
31 M 46 ALS 07 h 00 min 8.0
32 F 69 ALS 17 h 00 min 6.3

Abbreviations: ALS, amyotrophic lateral sclerosis; F, female; M, male; N/A, not available; PM, post‐mortem delay (hours, minutes); RIN, RNA integrity number; SC, Anterior horn of the spinal cord lumbar level.

TABLE 4.

Summary of the 14 cases analyzed by Western blotting corresponding to the spinal cord (n = 7 cases/group) and frontal cortex area 8 (n = 6 cases/group) of control cases (age 57.3 ± 13.01 years) and sALS cases (age 64.1 ± 11.16 years).

Case Sex Age Diagnosis PMD Region
1 M 55 Control 09 h 45 min SC
2 M 43 Control 05 h 55 min SC + FC
3 F 77 Control 08 h 30 min SC
4 M 66 Control 14 h 00 min SC
5 F 60 Control 09 h 40 min SC
6 M 82 Control 04 h 00 min SC
7 F 62 Control 05 h 45 min SC
8 F 45 Control 14 h 40 min FC
9 M 49 Control 07 h 35 min FC
10 M 40 Control 18 h 30 min FC
11 M 52 Control 04 h 40 min FC
12 M 56 Control 03 h 45 min FC
13 M 70 ALS 03 h 00 min SC + FC
14 F 83 ALS 15 h 15 min SC
15 M 59 ALS 03 h 15 min FC
16 F 63 ALS 13 h 50 min SC + FC
17 M 54 ALS 04 h 50 min FC
18 M 64 ALS 16 h 30 min SC + FC
19 M 46 ALS 07 h 00 min SC
20 F 76 ALS 13 h 00 min FC
21 M 55 ALS N/A SC
22 M 71 ALS 03 h 25 min SC

Abbreviations: ALS, amyotrophic lateral sclerosis; F, female; M, male; N/A, Not available; PM, post‐mortem delay (hours, minutes); RIN, RNA integrity number; SC, anterior horn of the spinal cord lumbar level.

2.8. RNA extraction and RT‐qPCR validation

Frozen samples of the anterior horn of the lumbar spinal cord [sALS (n = 22, mean age ± SD: 62.4 ± 10.8) and controls (n = 17, mean age ± SD: 64.9 ± 10.6 years)] and the frontal cortex area 8 [sALS (n = 15, mean age ± SD: 60.9 ± 8.63) and controls (n = 17, mean age ± SD: 64.8 ± 8.9 years)] were obtained for RNA extraction using RNeasy Mini Kit following the instructions of the supplier (Qiagen® GmbH, Hilden, Germany). RNA integrity, concentration, and 28S/18S ratios were determined with the Agilent Bioanalyzer (Agilent Technologies Inc., Santa Clara, CA, USA) (Tables 2 and 3). Complementary DNA (cDNA) preparation used the High‐capacity cDNA Reverse Transcription kit (Applied Biosystems, Foster City, CA, USA) following the protocol provided by the supplier. Parallel reactions for each RNA sample were run in the absence of MultiScribe Reverse Transcriptase to assess the lack of contamination of genomic DNA. TaqMan RT‐qPCR assays were performed in duplicate for each gene on cDNA samples in 384‐well optical plates using an ABI Prism 7900 Sequence Detection system (Applied Biosystems, Life Technologies, Waltham, MA, USA). For each 5 μL TaqMan reaction, 2.25 μL cDNA was mixed with 0.25 μL 20× TaqMan Gene Expression Assays and 2.5 μL of 2× TaqMan Universal PCR Master Mix (Applied Biosystems).

Analyzed genes included the following TaqMan probes: Chemokine (C‐X‐C motif) ligand 7 (CXCL7, Hs00234077_m1), C‐X‐C Motif Chemokine Receptor 1 (CXCR1, Hs01921207_s1), C‐X‐C Motif Chemokine Receptor 2 (CXCR2, Hs01891184_s1), allograft inflammatory factor (AIF1, Hs00741549_g1) coding for IBA1 protein, and glial fibrillary acidic protein (GFAP, Hs00909233_m1). The value of one house‐keeping gene, hypoxanthine‐guanine phosphoribosyltransferase 1 (HPRT1, Hs02800695_m1), was used as the internal control for normalization for spinal cord, and β‐glucuronidase (GUSβ, Hs00939627_m1) was used as the internal control for normalization for the frontal cortex area 8, as optimized previously [57]. The parameters of the reactions were 50°C for 2 min, 95°C for 10 min, and 40 cycles of 95°C for 15 s and 60°C for 1 min. Finally, the capture of all TaqMan PCR data was with the Sequence Detection Software (SDS version 2.2.2, Applied Biosystems). The double‐delta cycle threshold (ΔΔCT) method was used to analyze the data.

2.9. Gel electrophoresis and immunoblotting

Frozen samples of the anterior horn of the lumbar spinal cord (n = 7 sALS, n = 7 controls) and frontal cortex area 8 (n = 6 sALS, n = 6 controls) were homogenized using radioimmuno‐precipitation assay buffer (Table 4). The homogenates were centrifuged for 20 min at 20,000 rpm. Protein concentration was determined with the BCA method (Thermo Scientific). Equal amounts of protein (12 μg) for each sample were loaded and separated by electrophoresis on 15% SDS‐polyacrylamide gel electrophoresis (SDS‐PAGE) gels and then transferred onto nitrocellulose membranes (Amersham, Freiburg, GE). Non‐specific binding was blocked by incubation in 5% milk in PBS containing 0.2% Tween for 1 h at room temperature. After washing, the membranes were incubated overnight at 4°C with antibodies against CXCL7 (NBP1‐89921, dil. 1/250, R&D Systems, Bio‐Techne, MN, USA). Protein loading was monitored using antibodies against β‐actin (42 kDa, 1:40,000, Sigma, Barcelona, Spain). Membranes were incubated for 1 h with appropriate HRP‐conjugated secondary antibodies (1:2000, Dako); the immunoreaction was revealed with a chemiluminescence reagent (ECL, Amersham). Densitometric quantification was carried out with the ImageLab v4.5.2 software (BioRad), using β‐actin for normalization.

2.10. Immunohistochemistry

Single‐label immunohistochemistry was performed on de‐waxed sections of spinal cord and the frontal cortex area 8 of ALS (n = 4) and control cases (n = 4) that were 4‐μm‐thick. The sections were subjected to boiling in citrate buffer (pH = 6.0) for 20 min to enhance antigenicity. Subsequently, endogenous peroxidases were blocked by treating the sections with Dako Real peroxidase‐blocking solution for 15 min at room temperature. Next, Dako Real diluent solution (Dako) was applied to block non‐specific binding, permeabilize the tissue, and dilute the primary antibodies CXCL7 (NBP1‐89921, dil. 1/200, R&D Systems, Bio‐Techne, MN, USA). The sections were incubated with the primary antibodies overnight at 4°C. Following incubation, the slides were washed and then incubated with a biotinylated secondary antibody, “Ready to use” (Abcam), for 30 min at room temperature. This was followed by incubation with Streptavidin Peroxidase, also “Ready to use” (Abcam), for an additional 30 min at room temperature. The peroxidase reaction was visualized using diaminobenzidine and H2O2, resulting in a brown precipitate indicating immunoreactivity. The sections were counterstained with hematoxylin. To ensure the specificity of the immunostaining, a control was performed by omitting the primary antibody, which resulted in no signal upon incubation with only the secondary antibodies.

2.11. Statistical analysis

The normality of data distribution was assessed using the D'Agostino and Pearson test for each analysis. Categorical variables were described using frequencies. For normally distributed data, comparisons between two groups were performed using the unpaired Student's t‐test, while one‐way analysis of variance followed by Tukey's post hoc test was used for comparing CSF protein levels across multiple groups. For non‐normally distributed data, the Mann–Whitney U test was used for two‐group comparisons, and the Kruskal–Wallis test for comparisons among more than two groups. For the analysis of CSF CXCL7 levels in ALS and ALS‐mimic disorders, age was included as a covariate. Because CXCL7 concentrations showed a skewed distribution and heteroscedasticity, values were log‐transformed and group differences were assessed using an ANCOVA on the log‐transformed data.

Receiver operating characteristic (ROC) curve analysis was conducted to evaluate the classification performance of individual molecules. Features were primarily ranked based on their area under the ROC curve (AUC), which quantifies overall diagnostic accuracy. A higher AUC indicates better discriminative ability between groups. In cases where two or more features yielded identical AUC values, secondary ranking criteria were applied: features were further ranked based on sensitivity and specificity at the optimal cutoff (typically determined by Youden's index), the shape of the ROC curve across thresholds, and, if needed, the original input or internal processing order as a final tiebreaker.

Group differences in CSF CXCL7 and NfL concentrations were evaluated using a multivariate general linear model (GLM) to explore their potential association with ALS progression. To assess the predictive value of these biomarkers, multiple linear regression models were used to quantify the independent and combined associations of CXCL7 and NfL with ALS disease progression, with ALSFRS‐R scores included as the dependent variable. Effect sizes were reported using partial eta squared (η 2), with thresholds for small (η 2 = 0.01), medium (η 2 = 0.06), and large effects (η 2 ≥ 0.14). Multiple linear regression analyses were used to assess the predictive value of CSF CXCL7 and NfL levels on ALS disease progression (ALS‐FRS‐R slope). Three models were tested: [1] CXCL7‐only, [2] NfL‐only, and [3] a combined model including both biomarkers. Each model included biomarker levels as independent variables and ALS‐FRS‐R slope as the dependent variable. Model fit was evaluated using R 2, adjusted R 2, and F‐statistics. Regression coefficients (β) and their statistical significance were reported for each predictor. All assumptions for linear regression—including linearity, independence of errors, homoscedasticity, and absence of multicollinearity—were verified through residual plots and variance inflation factor diagnostics.

Associations between CXCL7 concentrations and clinical progression, operationalized as the ALS Functional Rating Scale‐Revised (ALS‐FRS‐R) slope and disease duration, were assessed using correlation analyses. Pearson's correlation was used for normally distributed variables, while Spearman's rank correlation was applied for non‐parametric data. Correlation coefficients (r) were interpreted using conventional thresholds: small (|r| = 0.1), moderate (|r| = 0.3), and strong (|r|≥ 0.5). To evaluate potential multicollinearity between biomarkers, we calculated bivariate correlations between CSF levels of CXCL7 and NfL.

All statistical analyses and graphical representations were performed using GraphPad Prism version 9.5 (La Jolla, CA, USA). Outliers were identified using GraphPad QuickCalcs (p <0.05). Data are presented as mean ± SEM, and statistical significance versus the control group was defined as *p <0.05, **p<0.01, and ***p <0.001.

3. RESULTS

3.1. Patient cohort

The discovery cohort was composed of CSF samples obtained from a total of 43 individuals, including 15 neurologically healthy control subjects and 28 patients diagnosed with sporadic ALS. Within the ALS group, patients were further stratified based on disease progression into two subgroups: 15 individuals with a short survival phenotype and 13 with a long survival phenotype, as determined by clinical follow‐up.

For validation of candidate protein biomarkers identified in the discovery phase, other cases were added to the discovery cohort, according to disease progression at the sample obtention time, comprising 82 individuals. This validation cohort included 54 patients with sporadic ALS, subdivided into 27 cases with a SP phenotype and 27 cases with a FP phenotype, alongside 28 neurologically healthy control subjects. Quantification of target proteins in the validation cohort was carried out using commercially available ELISA kits according to the manufacturer's protocols.

Importantly, no statistically significant differences in demographic variables, including age or sex distribution, were observed between the clinical groups in either the discovery or validation cohorts. In the discovery cohort, no significant differences were observed in sex (p = 0.44) or age (p = 0.31) among healthy controls and ALS subtypes. Additionally, ALS subtype cases did not show significant differences in age of onset (p = 0.89) or sex distribution (p = 0.81). In the validation cohort, no significant differences in age (p = 0.23) or sex (p = 0.31) were observed between the groups. Similarly, among ALS subtypes, no significant differences were found in age of onset (p = 0.41) or sex distribution (p = 0.81). No significant differences in sex (p = 0.15) distribution were observed between groups when comparing ALS cases with ALS‐mimic disorders cases. In contrast, age showed a significant difference among groups (p = 0.001), being significantly increased in sALS when compared to mimic‐ALS disorders groups.

A detailed summary of the demographic and clinical characteristics for all study participants is provided in Table 1.

3.2. Comprehensive CSF proteomic profiling reveals distinct molecular signatures between ALS and controls and across ALS progression subtypes

In the discovery cohort, a total of 1059 proteins were quantified across all CSF samples using SWATH‐MS. To explore global variation in protein expression profiles and detect potential grouping patterns among samples, both multivariate and univariate statistical analyses were performed. PCA was employed to reduce data dimensionality and uncover underlying structure in the dataset by capturing the major sources of variance across all samples. This analysis revealed a clear separation between ALS and control samples, with the first two principal components explaining 31.9% (PC1) and 16.6% (PC2) of the total variance, respectively (Figure 1A). Univariate analysis identified 386 differentially expressed proteins, with 199 proteins up‐regulated and 187 proteins down‐regulated. These results are visualized in a volcano plot highlighting proteins with significant fold changes (<0.77 or >1.3) and unadjusted p‐values (p <0.05) (Figure 1B). A heatmap further illustrates the expression patterns of all quantified proteins across samples (Figure 1C). Exact fold‐change values and statistical details are provided in Supplementary Data S1A.

FIGURE 1.

FIGURE 1

Proteomic profile results in the CSF of ALS and control cases. (A) PCA analysis for the identified proteins in control and ALS, utilizing only two principal components explaining 31.9% (PC1) and 16.6% (PC2) of the total variability. (B) Volcano plot displaying differentially expressed proteins in CSF between ALS patients and controls. (C) Heat‐map representing the 386 differentially expressed proteins and their relative expression levels across control and ALS subgroups. Specific patterns are observed within the ALS groups. (D), (E) Functional enrichment analysis of the proteins differentially up‐ and down‐regulated against GO biological process terms when comparing ALS against controls. (F) PCA analysis for the identified proteins in ALS subtypes classified as ALS‐LS and ALS‐SS, utilizing only two principal components explaining 44.9% (PC1) and 10.5% (PC2) of the total variability. (G) Volcano plot displaying differentially expressed proteins in CSF between ALS patient subtypes based on survival times. (H) Heat‐map with proteomic data representing the 386 differentially expressed proteins and their relative expression levels across ALS subgroups. (I), (J) Functional enrichment analysis of the proteins differentially up‐ and down‐regulated against GO biological process terms when comparing ALS‐SS against ALS‐LS. ALS‐SS, ALS‐short survival; ALS‐LS, ALS‐long survival, and control (healthy individuals). Volcano's plot legend: The x‐axis represents the log2 fold change, while the y‐axis indicates the –log10 adjusted p‐value. Proteins with both statistical significance (adjusted p‐value <0.05) and biologically relevant fold changes (|log2FC| >1.33) are highlighted. Upregulated and downregulated proteins are indicated in red and green, respectively. Heatmaps legend (C and F): Increased levels are indicated in the red spectrum, whereas decreased levels are noted in the green spectrum. Functional analysis legend (D and G): Protein grouping FDR is indicated as its log10 in a colored legend.

The functional significance of differential proteomes was assessed by functional enrichment analysis against gene ontology (GO) biological processes terms. When comparing the control cases to the ALS group, the up‐regulated proteins suggest a global enhancement of functional clusters in ALS cases linked to (i) immune cell system function, (ii) exocytosis mechanisms and (iii) cell adhesion/export mechanisms. Specifically, we found significant altered proteins linked to functional GO terms, including in “homophilic cell adhesion via plasma membrane adhesion molecules,” “cell‐cell adhesion via plasma‐membrane adhesion molecules,” “neutrophil degranulation,” “neutrophil activation involved in immune response,” “neutrophil mediated immunity,” “neutrophil activation,” “myeloid cell activation involved in immune response,” “granulocyte activation,” “regulated exocytosis,” “leukocyte degranulation,” “myeloid leukocyte mediated immunity,” “myeloid leukocyte activation,” “exocytosis,” “cell–cell adhesion,” “cell adhesion,” “biological adhesion,” “secretion by cell,” “export from cell,” and “secretion” (Figure 1D). When evaluated down‐regulated proteins in the ALS group, GO biological process terms identified proteins linked to (i) immune cell system function—similarly to upregulated processes, (ii) exocytosis mechanisms and (iii) oxidative stress mechanisms. The down‐regulated biological processes identified included “gas transport,” “hydrogen peroxide catabolic process,” “cellular detoxification,” “cellular response to toxic substance,” “receptor‐mediated endocytosis,” “regulation of body fluid levels,” “endocytosis,” “response to wounding,” “regulated exocytosis,” “response to lipid,” “secretion,” “negative regulation of multicellular organismal process,” “secretion by cell,” “response to oxygen‐containing compound,” “export from cell,” “biological adhesion,” “cellular response to oxygen‐containing compound,” “cell adhesion,” and “homeostatic process” (Figure 1E).

Building upon the results of the proteomic PCA, we further examined potential molecular differences between the ALS subgroups defined by survival duration. Specifically, we evaluated whether the global CSF proteomic profiles could distinguish patients with ALS‐SS from those with ALS‐LS. PCA conducted exclusively within the ALS cohort revealed a clear separation between the two phenotypes, with the first and second principal components accounting for 44.9% (PC1) and 10.5% (PC2) of the total variance, respectively (Figure 1F). This finding suggests that distinct proteomic signatures are associated with disease progression dynamics within ALS.

In line with this observation, differential expression analysis identified a total of 386 proteins showing statistically significant differences between the ALS subgroups, being 177 proteins up‐regulated and 209 proteins down‐regulated. These were determined using the same cut‐off criteria applied in the initial ALS versus control comparison. The results are visualized in a volcano plot (Figure 1G), which highlights proteins with phenotype‐specific regulation. Furthermore, the expression patterns of these proteins demonstrated clear clustering behaviors that correlated with the clinical phenotype, underscoring their potential relevance as progression‐related biomarkers. A heatmap further illustrates the expression patterns of all quantified proteins across samples (Figure 1H). Detailed fold‐change values and statistical parameters for all quantified proteins are provided in Supplementary Data Set S1B.

Functional enrichment analysis of proteins up‐regulated in ALS‐SS compared to ALS‐LS revealed a strong immune and secretory signature. Key enriched biological processes included “neutrophil degranulation,” “neutrophil activation involved in immune response,” and “leukocyte degranulation,” reflecting the involvement of innate immune mechanisms. Several processes associated with myeloid lineage activity, such as “myeloid leukocyte mediated immunity” and “leukocyte activation involved in immune response”, were also significantly enriched. In addition, pathways linked to vesicular transport and release, including “regulated exocytosis,” “exocytosis,” “secretion by cell,”, “export from cell,”, and “secretion,”, were up‐regulated. Processes related to cellular motility and adhesion, such as “regulation of cellular component movement,”, “regulation of locomotion,” “locomotion,” “cell motility,” “cell adhesion,” and “biological adhesion,” were also overrepresented in ALS‐SS. Collectively, these results suggest a relative activation of immune effector and secretory functions in the short‐survival group (Figure 1I).

Conversely, proteins down‐regulated in ALS‐SS relative to ALS‐LS were enriched for processes associated with neuronal structure and differentiation. The top enriched terms included “interleukin‐7–mediated signaling pathway” and “DNA replication‐dependent nucleosome assembly,” which suggest modulation of signaling and transcriptional activity. Neuronal morphogenesis was strongly represented, with terms such as “axonogenesis,” “axon development,” “cell morphogenesis involved in neuron differentiation,” “plasma membrane bounded cell projection morphogenesis,” “cell projection morphogenesis,” and “neuron projection morphogenesis.” Broader morphogenetic categories were also highlighted, including “cell part morphogenesis,” “cell morphogenesis involved in differentiation,” and “cellular component morphogenesis.” Processes supporting neuronal development, such as “neuron projection development,” “cell morphogenesis,” “neuron differentiation,” and “generation of neurons,” were also enriched. Finally, terms linked to structural connectivity, including “cell adhesion,” “biological adhesion,” and “locomotion,” reflect cytoskeletal and adhesion dynamics associated with neuronal remodeling. Together, these findings indicate that the short‐survival ALS group exhibits a proteomic shift toward neuronal differentiation and axonal morphogenesis, contrasting with the immune activation observed in up‐regulated proteins (Figure 1J).

3.3. Weighted protein Co‐expression network analysis identifies phenotype‐associated protein modules in ALS

We performed a WPCNA to identify clusters of co‐expressed proteins that reflect underlying biological processes and pathways. This network‐based approach enables the detection of functionally coherent modules and their association with sample traits, providing insights beyond individual protein changes. A weighted protein coexpression network was constructed using protein‐level values of 1059 proteins exhibiting variable expression across samples. 16 uncorrelated (r <0.9) gene modules were identified and labeled by colors and numbers (M1–M15). The proteins not assigned to any particular coexpression module were assigned to M0 or “gray.” Hierarchical clustering of both proteins and samples revealed a clear modular structure, with proteins forming coherent clusters that correspond to the identified modules (Figure 2A). The co‐expression heatmap provides an overview of the global protein expression patterns captured in the WPCNA (Figure 2B). In this, samples are grouped according to their overall proteomic similarity, demonstrating that the network‐based modules reflect biologically meaningful patterns of protein co‐expression. This visual representation confirms that the modular organization of the proteome captures major biological trends and sample‐specific signatures.

FIGURE 2.

FIGURE 2

Weighted protein coexpression network analysis (WPCNA) of CSF proteome of ALS and control cases. (A) Weighted protein coexpression network analysis of the CSF proteome using 1059 protein expression values of 43 assessed samples identifies 15 protein modules. Modules are labeled by color and number (M1 to M15). Proteins not assigned to any particular coexpression modules were labeled M0 or gray. Dendrogram obtained by hierarchical clustering of proteins based on their topological overlap is shown at the top. (B) Module eigengene correlation heatmap generated by WGCNA displays the relationships among all identified modules and samples. Each row represents a module eigengene and each column represents a sample, with color intensity reflecting the strength and direction of the Pearson correlation. Blue indicates negative correlations, whereas red indicates positive correlations. (C) Wilcoxon test analysis revealed distinct significant changes in module eigengene expression among groups in eight different modules. (D) Identification of gene enriched GO terms in each significant module was performed for modules characterization.

We obtained the most representative pattern of protein expression across all samples for each one of these modules by calculating the eigengene (i.e., the first principal component). The eigengene of each module was then correlated with the phenotype obtaining an eigengene significance for each module. Wilcoxon test analysis revealed significant differences in module eigengene expression among the groups (Figure 2C). To characterize these modules, gene‐enriched GO terms were identified for each significant module (Figure 2D).

Proteins included in the M1‐Black module (196 proteins) showed a significant upregulation in both ALS phenotypes compared to the control group (p<0.0001) (Figure 2C). The “black module” is mainly enriched in glycoprotein‐related processes. The top enriched terms include “glycoprotein metabolic process,” “negative regulation of tumor necrosis factor production,” and “glycoprotein biosynthetic process.” Additional pathways such as “protein O‐linked glycosylation,” “macromolecule glycosylation,” and “O‐glycan processing,” highlight its association with post‐translational modification and immune‐related signaling. The adjusted p‐values range approximately from 2 × 10−6 to 2 × 10−4, indicating strong enrichment (Figure 2D).

Proteins included in the M2‐Greenyellow module (134 proteins) were significantly upregulated in the ALS‐LS phenotype (p <0.0001) and significantly downregulated in the ALS‐SS phenotype (p <0.0001) compared with the control group, with a significant difference also observed between ALS‐LS and ALS‐SS (p <0.0001) (Figure 2C). The “greenyellow module” shows a strong enrichment in structural and developmental processes. Key terms include “intermediate filament organization,” “intermediate filament cytoskeleton organization,” and “keratinocyte differentiation.” Pathways related to tissue formation and adhesion such as “skin development,” “homophilic cell adhesion via plasma membrane adhesion molecules,” and “keratinization” are also prominent. P‐values range from 0.01 to 0.02, suggesting moderate significance (Figure 2D).

Proteins included in the M3‐Pink module (109 proteins) were significantly downregulated in the ALS‐SS phenotype (p <0.0001) compared with both the control and ALS‐LS groups, while the ALS‐LS phenotype showed a significant upregulation relative to the control group (p <0.01) (Figure 2C). The “pink module” is dominated by neuronal and synaptic pathways. Top GO terms include “axon guidance,” “neuron projection guidance,” and “cell‐cell adhesion via plasma membrane adhesion molecules.” Additional enriched processes such as “postsynaptic membrane organization,” “artery development,” and “postsynapse organization” indicate a role in neurodevelopment and vascular interactions. P‐values lie between 0.01 and 0.02, suggesting a moderate significance (Figure 2D).

Proteins included in the M4‐Green module (83 proteins) were significantly downregulated in the ALS‐SS phenotype (p <0.001) compared with the ALS‐LS group, while no significant differences were observed in the other comparisons (Figure 2C). The “green module” demonstrates robust enrichment in neuronal development and growth regulation. The leading pathways are “axonogenesis,” “axon guidance,” and “neuron projection guidance.” Regulatory processes such as “regulation of neuron projection development,” “regulation of axonogenesis,” and “synapse maturation” further emphasize its role in neurogenesis and synaptic plasticity. P‐values range from 1 × 10−7 to 5 × 10−7, reflecting highly significant enrichment (Figure 2D).

Proteins included in the M5‐Magenta module (71 proteins) were significantly downregulated in the ALS‐SS phenotype (p <0.01) compared with both the control and ALS‐LS groups (Figure 2C). The “magenta module” is enriched in synaptic organization and adhesion processes. Top enriched GO terms include “synaptic membrane adhesion,” “positive regulation of synapse assembly,” and “cell‐cell adhesion via plasma‐membrane adhesion molecules.” Supporting terms such as “synapse assembly,” “regulation of synapse organization,” and “regulation of synapse structure or activity” underscore its role in synaptic connectivity. P‐values range from 0.01 to 0.04, suggesting a moderate significance (Figure 2D).

Proteins included in the M10‐Purple module (43 proteins) were significantly downregulated in the ALS‐LS phenotype (p <0.0001) compared with both the control and ALS‐LS groups, while the ALS‐SS phenotype showed a significant upregulation relative to the control group (p <0.01) (Figure 2C). The “purple module” is associated with immune cell migration and adhesion. Its main pathways are “homotypic cell‐cell adhesion,” “myeloid leukocyte migration,” and “neutrophil chemotaxis.” Complementary processes include “cell chemotaxis,” “leukocyte chemotaxis,” and “desmosome organization,” indicating a functional link to immune response and cell–cell interactions. P‐values range from 0.005 to 0.02, suggesting moderate to strong significance (Figure 2D).

Proteins included in the M11‐Yellow module (42 proteins) were significantly upregulated in the ALS‐SS phenotype (p <0.0001) compared with both the control and ALS‐LS groups, while no significant differences were observed in the other comparisons (Figure 2C). The “yellow module” presents a simpler enrichment profile focused on neurodevelopment, with top terms “axon guidance” and “neuron projection guidance,” and p‐values around 3 × 10−5, suggesting a strong significance (Figure 2D).

Finally, proteins included in the M14‐Brown module (11 proteins) were significantly downregulated in both ALS‐LS and ALS‐SS phenotypes (p <0.0001) compared with the control group. A significant difference was also observed between the ALS phenotypes, with a greater decrease in the ALS‐SS group compared with ALS‐LS (p <0.05) (Figure 2C). The “brown module” is clearly enriched in oxidative and detoxification processes. Leading terms include “hydrogen peroxide catabolic process,”“gas transport,” and “oxygen transport,” followed by “cellular detoxification,” “cellular response to toxic substance,” and “cellular oxidant detoxification.” This suggests a prominent role in redox regulation and response to oxidative stress. P‐values span from 10−4 to 4 × 10−4, highlighting a very robust enrichment (Figure 2D).

3.4. Lipidomic profiling reveals subtle alterations in CSF lipids with limited discriminative power between ALS and control groups

Simultaneously, we conducted an MS‐based lipidomic analysis to explore novel CSF biomarkers associated with ALS, aiming to determine whether lipidomic alterations could provide new insights into ALS pathophysiology, prognosis, and patient management. Lipidomic profiling was performed on the same CSF samples used for proteomics, although some samples were unavailable due to their complete consumption during proteomic analysis. As a result, the lipidomic analysis included 11 sporadic ALS‐LS cases, 13 sporadic ALS‐SS cases, and 13 control samples.

The CSF lipidome analysis revealed the presence of 1119 distinct m/z entities across all study groups. Both multivariate and univariate statistical approaches were applied to assess potential lipidomic differences among ALS subtypes and control samples. Initial unsupervised analysis using PCA on lipid profiles demonstrated no clear separation between ALS patients and controls. The PCA plot showed that the first principal component (PC1) explained 100% of the total variance, while the second component (PC2) contributed 0.00% (Figure 3A), indicating that the lipidomic data are essentially one‐dimensional with minimal group separation and limited discriminatory power. Moreover, the overlapping distribution of samples in the PCA score plot suggests no meaningful separation between ALS subtypes and controls, indicating that their CSF lipidomic profiles are highly similar. Any underlying differences appear too subtle to detect through unsupervised analysis or may be masked by noise or technical variation.

FIGURE 3.

FIGURE 3

Lipidomic profile results in the CSF of ALS and control cases. (A) PCA analysis for the identified lipids in control and ALS, utilizing only two principal components explaining 100.0% (PC1) and 0.0% (PC2) of the total variability. It indicates the lack of differentiation between groups. (B) Heat‐map representing the 15 differentially expressed lipids and their relative expression levels across control and ALS subgroups. (C) List of differentially expressed lipids identified and annotated when comparing the control group against the ALS group. (D) PCA analysis for the identified proteins in ALS subtypes classified as ALS‐LS and ALS‐SS, utilizing only two principal components explaining 27.4% (PC1) and 11.7% (PC2) of the total variability. (E) Heat‐map with lipidomic data representing the three differentially expressed lipids and their relative expression levels across ALS subgroups. Increased levels are indicated in the red spectrum, whereas decreased levels are noted in the blue spectrum. (F) List of differentially expressed lipids identified and annotated when comparing ALS subtypes. (G) The main lipid components of CSF capable of distinguishing between ALS patients and controls as determined through the ROC model. The results just showed that four candidates were able to differentiate between groups with an AUC value higher than 0.70. ALS‐SS, ALS‐short survival; ALS‐LS, ALS‐long survival, control (healthy individuals); CSF, cerebrospinal fluid. Heatmap legend: Increased levels are indicated in the red spectrum, whereas decreased levels are noted in the blue spectrum.

However, despite the lack of group separation in the PCA, univariate statistical analysis comparing the entire ALS group to controls identified 15 lipids with nominal p‐values <0.05. Among these, one lipid species remained statistically significant after FDR correction (FDR <0.05), while the remaining 14 exhibited unadjusted p‐values below 0.05. A hierarchical clustering heatmap of the 15 most significant lipid species (Figure 3B) confirmed the limited discriminatory power of lipidomic signatures in differentiating ALS cases from controls.

At the functional level, we offer a potential annotation of 7 out of the 15 deregulated lipids in ALS subtypes and controls, while the remaining 8 were not identified. The annotated lipids belonged to the following categories: carboxylic acids (n = 2), fatty acids (n = 3), organooxygen compounds (n = 1), and glycosphingolipids (n = 1) (Figure 3C). The clusters of differentially down‐ and up‐regulated lipids in ALS subtypes and control CSF cases are limited. Consequently, no functional analysis was possible for the comparison between ALS and control cases, as well as between subtypes.

To further investigate potential group‐specific lipidomic signatures, the ALS cohort was stratified into two clinical subgroups: ALS‐LS and ALS‐SS. Despite the poor classification observed in the overall comparison, we assessed whether lipidomic profiles could distinguish between these ALS subtypes. PCA revealed limited separation, with PC1 and PC2 explaining 27.4% and 11.7% of the total variance, respectively (Figure 3D). The close clustering and substantial overlap of samples in the PCA score plot indicate a high degree of similarity in lipidomic profiles between ALS‐LS and ALS‐SS patients, suggesting a lack of clear subgroup‐specific patterns. In line with these findings, univariate analysis identified only three lipids with nominal p‐values <0.05. A hierarchical clustering heatmap of these three most significant lipid species (Figure 3E) further confirmed the absence of a whole substantial distinct lipidomic signature capable of differentiating ALS subtypes. At the functional level, when comparing ALS subtypes, we identified only 2 out of the 3 differential compounds, specifically 1 lipid and lipid‐like molecule, and 1 organooxygen compound (Figure 3F).

Using lipidomic data obtained from MS/MS analysis, we also evaluated the potential of lipids to discriminate between control and ALS cases by applying a ROC curve model to the differentially expressed lipid species. Among the 15 lipids that were significantly altered between the control and ALS groups, four exhibited an AUC value greater than 0.70. These lipids demonstrated modest discriminative ability, suggesting potential as diagnostic markers; however, further validations in independent cohorts and applying other technical approaches are necessary to confirm their clinical utility (Figure 3G).

3.5. Machine learning models uncover robust and distinct proteomic signatures differentiating ALS subtypes and controls

To address the complexity of mass spectrometry‐based datasets, we selected machine learning models such as PLS‐DA and Random Forest for their complementary strengths in handling high‐dimensional data. Through these tools, we aimed to identify the features most capable of differentiating between ALS patients and controls, as well as distinguishing among ALS subtypes.

Using a PLS‐DA model to classify samples based on ALS subtypes and control groups, we observed a clear separation of the samples into three distinct and well‐defined clusters (Figure 4A). Evaluation of the model's performance based on 300 permutations yielded a statistically significant result (p‐value = 0.03), indicating that the model was robust and unlikely to be the result of random chance. The validity of the model was further supported by a high Q 2 value of 0.972 for four components, along with strong performance metrics (Accuracy = 1.0; R 2 = 0.995), suggesting that the model is not overfitted and has strong predictive capability. Within this model, 15 ALS‐SS and 14 ALS‐LS subjects were accurately classified as ALS (100%), alongside controls, reinforcing the discriminative power of the protein expression profiles. The proteins contributing most significantly to class separation included SPAT2, ZN449, CXCL7, ZP2, MYLB6, VIM, FLG, MPO, WDR87, BPGM, IGKV3‐7, APOB, SERPINB1, CA2, PIANP, ZNF518A, GALNT18, BLVRB, ARSA, CRABP1, MB, H1‐4, HSP90B1, ABCA1, and PODXL2 (Figure 4B). This specific distribution of samples supports the notion that ALS subgroups—particularly those with short or LS times—exhibit distinct and homogeneous proteomic signatures, clearly differentiating them from control individuals.

FIGURE 4.

FIGURE 4

Multivariate classification of ALS phenotypes using PLS‐DA and random forest analysis of CSF proteins. (A) PLS‐DA analysis for the identified proteins in all studied groups, selecting two components which explain 27.6% (Component 1) and 15.1% (Component 2). (B) Classification of proteins with the greatest influence on class separation depending on the VIP score. (C) Multivariate Random Forest capable of identifying which proteins belong to each phenotype with an OOB error equal to 0.00. (D) Panel of the top proteins that discriminate between ALS subtypes and controls with 100% accuracy in the RF model. ALS‐LS, ALS‐long Survival; ALS‐SS, ALS‐short survival, control (healthy individuals); OOB, out‐of‐bag; PLS‐DA, partial least squares‐discriminant analysis; VIP, variable importance in projection.

The functional enrichment analysis of proteins ranked by the PLS‐DA machine learning model revealed several significantly overrepresented biological processes based on GO terms. Notable enriched biological processes include “gas transport” and “oxygen transport,” which showed the highest fold enrichment and statistical significance. Several processes related to lipid handling—such as “lipoprotein particle remodeling,” “cholesterol storage,” and “macrophage differentiation”—highlight key roles in lipid metabolism and immune system regulation. Additionally, responses to external stimuli, including “lipopolysaccharide” and “molecules of bacterial origin,” suggest an involvement of these proteins in immune response mechanisms. The complete list of enriched GO biological process terms is: “gas transport,” “oxygen transport,” “low‐density lipoprotein particle remodeling,” “reg. of cholesterol storage,” “cholesterol storage,” “reg. of macrophage derived foam cell differentiation,” “plasma lipoprotein particle organization,” “protein‐lipid complex remodeling,” “protein‐lipid complex subunit organization,” “protein‐containing complex remodeling,”, “reg. of plasma lipoprotein particle levels,” “response to lipopolysaccharide,” “response to molecule of bacterial origin,” and “response to oxygen‐containing compound.”

We next investigated whether the observed protein expression profiles could be used to construct a RF model capable of accurately classifying individuals into ALS subtypes or control groups. The resulting multivariate RF model successfully predicted all samples into their correct classes, achieving 100% classification accuracy (Figure 4C). Feature selection identified a panel of 15 proteins that contributed most to the model's discriminative power: FAM174A, DSC1, SERPINB12, CREG1, CRABP1, IGKV3‐7, PCDHGC5, KRT86, ARSA, C1QTNF1, TYRO3, MB, CACNA2D3, CLSTN2, and ELANE (Figure 4D). This highly accurate classification further supports the presence of distinct proteomic signatures among ALS subtypes and controls, and highlights the potential of this protein panel for use in diagnostic or stratification models in ALS. To further characterize the biological functions associated with the proteins ranked by the RF machine learning model, a functional enrichment analysis was conducted using GO biological process terms. The results revealed a specific enrichment in processes related to cell adhesion. Notably, “cell–cell adhesion” and “homophilic cell adhesion via plasma membrane adhesion molecules” were significantly overrepresented, suggesting that these proteins may play a role in mediating intercellular connectivity and maintaining tissue architecture. These findings may point to functional implications in structural organization and cell communication.

Notably, four of the top‐ranked proteins overlapped between the PLS‐DA and Random Forest models—IGKV3‐7, ARSA, CRABP1, and MB—indicating shared discriminative potential across both approaches. However, the differences observed between the sets of ranked proteins identified by the PLS‐DA and RF models are expected and can be attributed to the distinct mathematical frameworks and feature selection strategies employed by each algorithm. PLS‐DA is a linear projection method that identifies proteins contributing most strongly to the separation between groups by maximizing the covariance between the predictor variables and class membership. In contrast, RF is a non‐linear, ensemble‐based classification algorithm that ranks features based on their contribution to reducing classification error across multiple decision trees, often capturing complex interactions and non‐linear patterns in the data.

As a result, PLS‐DA tends to highlight proteins that exhibit consistent, group‐correlated variation, whereas RF may prioritize a smaller set of highly discriminative proteins that offer the most predictive power, even if some of them are not strongly linearly associated with the outcome. Additionally, the RF model is less influenced by collinearity, often selecting a single representative from a group of correlated features, while PLS‐DA may rank multiple correlated proteins highly due to their shared contribution to group separation.

These methodological differences justify the partial divergence in protein rankings and demonstrate the value of applying complementary machine learning approaches to gain a broader and more nuanced understanding of the biological signals within the dataset.

3.6. Machine learning analysis of lipidomic data not reveals discriminatory power for ALS classification

Similarly, we applied the same machine learning approaches to the lipidomics dataset. A discriminant sparse PLS‐DA (sPLS‐DA) model was constructed using ALS and control samples, based on the small number of informative biomarkers. However, in this case, the samples clustered into three poorly separated and non‐distinct classes, and model evaluation based on 300 permutations indicated that the model was not statistically significant (p‐value = 0.7; Figure 5A). Among the differential lipids, those contributing most strongly to class separation were mainly positively ionized species, as illustrated in Figure 5B.

FIGURE 5.

FIGURE 5

Multivariate classification of ALS phenotypes using PLS‐DA and random forest analysis of CSF lipids. (A) sPLS‐DA analysis for the identified proteins in all studied groups selecting two components, which explain 8.1% (Component 1) and 5.8% (Component 2) of the total variability. (B) The sPLS‐DA VIP scores classify the variables depending on the contribution of each one to the differentiation between classes; variables with high VIP values are considered more important for the classification task. (C) The random forest model trying to identify which lipids belong to the different classes. (D) Panel of lipids that discriminate better between ALS subtypes and controls. ALS‐LS, ALS‐long survival; ALS‐SS, ALS‐short survival, control (healthy individuals); OOB, out‐of‐bag; PLS‐DA, partial least squares‐discriminant analysis; VIP, variable importance in projection.

We further investigated whether the lipid signature could be used to construct a multivariate Random Forest model for classifying individuals into ALS subtypes or control groups. However, classification performance was poor, with most samples misclassified (Figure 5C) and an overall error rate of 70.3%, indicating limited discriminatory power of the selected lipids. Despite this, the feature selection process identified a subset of 15 lipid predictors that showed the greatest potential for distinguishing between the groups (Figure 5D).

3.7. Validation of CXCL7 as a candidate CSF biomarker reveals moderate diagnostic potential compared to NfL in ALS

Based on its statistical significance in the exploratory proteomic analysis, putative biological relevance, multiple machine learning models results, compatibility with targeted quantification techniques and its role in inflammatory processes, CXCL7 was selected for further evaluation as a candidate biomarker. Its CSF levels were measured in a validation cohort to assess its diagnostic and/or prognostic potential. In parallel, CSF levels of NfL, a well‐established biomarker in neurodegenerative diseases, were also quantified to serve as a reference standard for comparative analysis.

Significantly lower CXCL7 protein levels were detected in sporadic ALS (127.20 ± 10.49 pg/mL) compared with controls (256.73 ± 46.73 pg/mL) (U = 386, p = 0.0065), the effect size is moderate (rank‐biserial correlation rₑ≈0.38), indicating a moderate to large biological impact but also highlighting considerable variance between individuals and not just statistical noise (Figure 6A). To calculate the clinical accuracy of CXCL7 in discriminating between ALS and the control group, the AUC value resulted in 0.6907 ± 0.067, 95% CI: 0.5599–0.8215 (p = 0.007) (Figure 6C). Considering the optimal cut‐off at 311.4 pg/mL, defined by the Youden index (0.307), an overall sensitivity of 100% and specificity of 30.77% were predicted. In the same samples, NfL levels were also measured. Significantly higher NfL protein levels were detected in sporadic ALS (5605.32 ± 609.07 pg/mL) compared with controls (483.36 ± 43.21 pg/mL) (U = 6, p = 0.000) (Figure 6B). To calculate the clinical accuracy of NfL in discriminating between sporadic ALS and the control group, we estimated the AUC value of 0.9952 ± 0.00457, 95% CI: 0.9863–1.000 (p = 0.000) (Figure 6C). Considering the optimal cut‐off at 730.4 pg/mL, defined by the Youden index (0.933), an overall sensitivity of 100% and specificity of 93.33% were predicted.

FIGURE 6.

FIGURE 6

Biomarker's validation using ELISA techniques. Quantification of CXCL7 (pg/mL) (A) and NfL (pg/mL) (B) protein levels in the CSF in ALS and control cases using ELISA. (C) ROC curves of CXCL7 and NfL quantification in the differential diagnosis of ALS compared with controls. (D) CSF CXCL7 concentrations across sALS, inflammatory polyneuropathy (IP), and myelopathies. Although raw CXCL7 levels were highest in the IP group, no significant differences were observed between groups after log‐transformation and adjustment for age. Quantification of CXCL7 (pg/mL) (E) and NfL (pg/mL) (F) protein levels in the CSF in ALS cases of stratified patients by progression using ELISA. (G) Receiver operating characteristic (ROC) curves of CXCL7 and NfL quantification in the differential prognosis between ALS‐SP and ALS‐FP. (H) Correlation between CSF CXCL7 levels and disease duration. Each dot represents an individual patient. A moderate positive association was observed (Spearman's ρ = 0.37, 95% CI 0.08–0.60; p = 0.012), indicating higher CXCL7 levels with longer disease duration. Mean ± SEM is represented in the graphs. Statistical significance was set at *p ≤ 0.05, **p ≤ 0.01, ***p ≤ 0.001 and ****p ≤ 0.0001. AUC values correspond to the area under ROC curves. ALS‐SP, ALS‐slow progressors; ALS‐FP, ALS‐fast progressors; ROC, receiver operating characteristic.

Finally, to assess the disease specificity of CXCL7 CSF concentrations in sALS, we analyzed CSF samples from patients with clinically relevant mimic disorders, including inflammatory polyneuropathy (IP; n = 26) and myelopathies (n = 14). For ALS‐mimic disease comparisons, CXCL7 concentrations were log‐transformed prior to analysis to improve normality. Descriptively, raw CSF CXCL7 concentrations were highest in the IP group (297.38 ± 47.85 pg/mL), followed by the sALS group (115.34 ± 31.84 pg/mL) and the myelopathies group (103.62 ± 65.59 pg/mL) (Figure 6D). An ANCOVA was performed with log‐transformed CXCL7 as the dependent variable, diagnostic group as the fixed factor, and age as a covariate. The overall model did not reach statistical significance (F (3, 76) = 2.51, p = 0.065), explaining 9.0% of the variance in log‐transformed CXCL7 levels (adjusted R 2 = 0.054). Age did not show a significant effect on CXCL7 concentrations (F (1, 76) = 3.20, p = 0.078), and no significant main effect of diagnostic group was observed after adjusting for age (F (2, 76) = 1.92, p = 0.153).

3.8. CXCL7 and NfL demonstrate prognostic value in distinguishing fast and slow progression phenotypes in sporadic ALS

CXCL7, together with NfL levels, were also evaluated as a prognostic tool in the different sporadic ALS phenotypes included within the same cohort used for diagnostic purposes, that included ALS‐SP and ALS‐FP cases. CXCL7 CSF levels showed significant differences between groups when comparing ALS‐SP (152.4 ± 17.47 pg/mL) with ALS‐FP (102.0 ± 10.51 pg/mL) using t‐test (t = 2.469, p = 0.017) (Figure 6E); demonstrating clinical accuracy in discriminating between the two extreme progression groups in sporadic ALS. This yields a Cohen's d of 0.72, indicating a medium to large effect size, suggesting that the observed difference in CXCL7 levels between ALS subgroups is not only statistically significant but also potentially meaningful in a biological or clinical context. To further assess this, a ROC curve test was conducted, resulting in an AUC value of 0.67 ± 0.08055, with a 95% CI of 0.5097–0.8254 (p = 0.046) (Figure 6G). Considering the optimal cut‐off at 146.3 pg/mL, defined by the Youden index (0.375), an overall sensitivity of 54.17% and specificity of 83.4% were predicted.

Finally, when sporadic ALS patients were stratified by survival phenotypes, the Mann–Whitney test reported significant differences among NfL levels between sporadic ALS phenotypes (U = 58.5, p = 0.0001), indicating high levels of NfL in the CSF of patients with ALS‐FP (8320.8 ± 804.8 pg/mL) when compared to those with ALS‐SP (3152.6 ± 514.53 pg/mL) (Figure 6F). Importantly, the rank‐biserial correlation (rₑ) was calculated to be 0.735, reflecting a very large effect size and suggesting a strong separation between the ALS‐FP and ALS‐SP distributions. This result not only confirms statistical significance but also supports the biological relevance of the observed differences between these patient subgroups. To calculate the clinical accuracy presented by NfL in discriminating between the two extreme progression groups in sporadic ALS, a ROC curve test was applied, obtaining an AUC value of 0.8650 ± 0.058, with a 95% CI of 0.7528–0.9819 (p = 0.0000) (Figure 6G). Considering the optimal cut‐off point at 6484 pg/mL, defined by the Youden index (0.667), a global sensitivity of 90.48% and specificity of 76.19% were predicted.

To directly compare the prognostic performance of CXCL7 and NfL in distinguishing fast versus slow ALS progressors, a DeLong test was performed to assess the statistical difference between their respective ROC curves. The test revealed that the AUC for NfL (0.8650) was significantly higher than that of CXCL7 (0.6700), with a DeLong p‐value of 0.017. This result confirms that NfL has superior prognostic accuracy compared to CXCL7 in differentiating ALS progression phenotypes, further supporting its utility as a more robust biomarker in this clinical context. Nonetheless, the moderate AUC of CXCL7, along with its significant group differences and effect size, suggests it may still offer complementary prognostic value when used alongside NfL.

3.9. CXCL7 and NfL independently associate with ALS progression, with NfL showing stronger predictive power in multivariate models

To investigate the potential role of CSF biomarkers in ALS progression, we first conducted a multivariate GLM to assess group differences in CXCL7 and NfL levels. The analysis revealed that both CXCL7 and NfL showed statistically significant differences between the groups. However, NfL demonstrated a stronger effect size (η 2 = 0.415, p < 0.001) compared to CXCL7 (η 2 = 0.179, p = 0.006), indicating that NfL possesses a greater discriminative ability. Specifically, group membership explained 41.5% of the variance in NfL levels, while CXCL7 levels accounted for 17.9% of the variance. Despite this difference in effect size, CXCL7 also contributed significantly to group discrimination, suggesting its potential value as a complementary biomarker.

Next, associations between biomarker levels and clinical progression were examined using the ALS‐FRS‐R slope. Pearson correlation analysis showed a negative association between ALS‐FRS‐R slope and CSF CXCL7 levels (r = −0.309, p = 0.05), indicating that lower CXCL7 levels were associated with faster disease progression. In contrast, Spearman correlation analysis revealed a positive association between ALS‐FRS‐R slope and NfL levels (r = 0.497, p = 0.002), consistent with higher NfL levels being linked to more rapid clinical decline. To assess potential collinearity between biomarkers, we also examined the association between NfL and CXCL7 levels. However, Spearman correlation analysis did not find significant correlation (r = −0.221, p = 0.171), supporting their evaluation as independent predictors.

To further determine the predictive value of these biomarkers, we performed a multiple linear regression analysis. This statistical approach quantifies the individual and combined contributions of CXCL7 and NfL to ALS disease progression using the ALS‐FRS‐R score. Three models were assessed: one including CXCL7 alone, one with NfL alone, and a combined model incorporating both. The CXCL7‐only model was statistically significant [F (1, 40) = 3.914, p = 0.05], explaining 9% of the variance in ALS progression (R 2 = 0.09). Lower CXCL7 levels were significantly associated with faster progression (β = −0.299, p = 0.05). The NfL‐only model showed stronger results [F (1, 35) = 13.06, p = 0.001], accounting for 27% of the variance (R 2 = 0.27), with higher NfL levels significantly predicting faster progression (β = 0.521, p = 0.001). In the combined model [F (2, 32) = 9.06, p < 0.001], both biomarkers together explained 36% of the variance (R 2 = 0.36). In this model, NfL remained a significant predictor (β = 0.475, p = 0.003), whereas the association between CXCL7 and progression did not reach significance (β = −0.256, p = 0.09).

3.10. Association of CSF CXCL7 levels with disease survival time, but not with site of onset

To further characterize the relationship between CXCL7 levels and clinical features of ALS, we performed additional analyses aimed at assessing the influence of cohort composition, survival‐based stratification, disease duration, and site of disease onset. CXCL7 levels were therefore compared across survival subgroups from the discovery and validation cohorts, and their associations with disease duration and clinical phenotype were examined.

CXCL7 levels were compared across cohorts‐based subgroups from the discovery and validation cohorts to evaluate the potential impact of cohort overlap. No significant difference in CXCL7 levels was observed between the ALS‐SS group from the discovery cohort and the ALS‐FP group from the validation cohort (Mann–Whitney U = 66.5; p = 0.85). Mean CXCL7 levels were comparable between the two groups, measuring 100.76 ± 14.54 pg/mL in ALS‐SS patients and 102.86 ± 15.19 pg/mL in ALS‐FP patients. In contrast, a significant difference in CXCL7 levels was observed between long‐survival subgroups. CXCL7 levels were higher in ALS‐LS patients compared with ALS‐SP patients (Mann–Whitney U = 35; p = 0.035). Mean CXCL7 levels were 190.2 ± 22.68 pg/mL in ALS‐LS and 120.42 ± 23.12 pg/mL in ALS‐SP patients. In consequence, CSF CXCL7 levels were evaluated in relation to disease duration and site of disease onset. A moderate positive correlation was observed between CXCL7 levels and disease duration (Spearman's ρ = 0.37, 95% CI 0.08–0.60; p = 0.012; n = 46), with higher CXCL7 levels associated with longer disease duration (Figure 6H).

Finally, CXCL7 levels were also compared between spinal‐onset and bulbar‐onset ALS patients. The single case with respiratory onset was excluded from subgroup analyses due to insufficient sample size. No significant difference in CXCL7 levels was detected between spinal‐onset and bulbar‐onset patients (Mann–Whitney U = 181.5; p = 0.69). Mean CXCL7 levels were 126.6 ± 12.6 pg/mL in spinal‐onset patients and 134.4 ± 22.5 pg/mL in bulbar‐onset patients.

3.11. Region‐specific analysis of post‐mortem tissue reveals differential CXCL7 mRNA expression without significant protein changes in ALS

We evaluated, at different molecular levels, CXCL7, CXCL7‐receptors, and other inflammatory protein markers in a region‐dependent manner in a separate set of samples of post‐mortem tissue of ALS and control cases, at the end of disease, focusing on the anterior horn of the spinal cord and the frontal cortex area 8 due to their implications in the ALS‐FTD disease spectrum.

In the anterior horn of the spinal cord, mRNA levels of CXCL7 protein were not significantly altered in sporadic ALS (p = 0.72) when compared with controls in the anterior horn of the spinal cord. However, notably, variability intragroup was reported in control and ALS cases. mRNA expression levels of CXCL7 main receptors, CXCR1 and CXCR2, were undetermined and unaltered (p = 0.11), respectively. Astrocytic and microglial genes were evaluated, including GFAP and AIF1 transcripts as inflammatory markers. Increased mRNA levels of AIF1 (p = 0.000) but not of GFAP (p = 0.18) were noted in sporadic ALS cases when compared with controls (Figure 7A).

FIGURE 7.

FIGURE 7

Neuropathological study of CXCL7 in post‐mortem tissue of sporadic ALS patients. Gene expression of CXCL7 and other genes encoding inflammatory markers in the anterior horn of the spinal cord (A) and the FC area 8 (B) of ALS and control cases. CXCL7 levels were unaltered in the anterior horn of the spinal cord when compared to controls but showed increased levels in ALS cases in the FC. Transcripts encoding CXCL7 receptors, CXCR1, and CXCR2, did not show significant differences between groups in any of the studied regions. AIF1 mRNA expression is significantly increased in the anterior horn of the lumbar spinal cord in ALS when compared with controls. GFAP mRNA expression is not modified in ALS when compared with controls in any region. Protein levels in the anterior horn of the spinal cord (C) and the FC area 8 (D) of ALS and control cases. Western blot analysis of CXCL7 in the anterior horn of the spinal cord and FC of ALS cases not showed significant differences when compared with controls in any region. β‐Actin was used for normalization. Graphical representation expresses fold changes in ALS relative to control cases. Mean ± SEM is represented in the graphs. Differences between ALS patients and controls were assessed with a t‐test and significance level was set at **p <0.01 and ***p <0.001. (E) Immunohistochemistry for CXCL7 in the anterior horn of the spinal cord in control and ALS cases (a–d) and in the FC area 8 (e–f). CXCL7 immunoreactivity is observed in the gray matter (GM) of the spinal cord in ALS (a), with high reactivity in motor neurons in both ALS cases (a–c) and controls (d). CXCL7 positivity is present in a few neurons in the gray matter of the frontal cortex of control (e) and sALS (f) cases. Paraffin sections are lightly counterstained with hematoxylin, and bar scale is indicated in the images. Au, arbitrary units; FC, frontal cortex; sALS, sporadic ALS.

In the frontal cortex area 8, t‐test statistical analysis revealed that mRNA levels of CXCL7 protein (p = 0.008) were significantly up‐regulated in sporadic ALS when compared with controls in the frontal cortex area 8. Intragroup variability was also reported in control and ALS cases. mRNA expression levels of CXCL7 receptors, CXCR1 and CXCR2, were unaltered (p = 0.81 and p = 0.63, respectively). Astrocytic and microglial genes were also evaluated as inflammatory markers, GFAP and AIF1 transcripts, but no significant changes were reported (p = 0.96 and p = 0.93, respectively) (Figure 7B).

Additionally, relative protein levels were determined in post‐mortem tissue samples of the anterior horn of the spinal cord (n = 7) and the frontal cortex area 8 (n = 6) from controls and sporadic ALS cases. Gel electrophoresis and Western blotting of the anterior horn of the lumbar spinal cord showed no significant altered levels in CXCL7 protein expression between control and sporadic ALS cases in the spinal cord (p = 0.60) and frontal cortex area 8 (p = 0.85) (Figure 7C,D). Intragroup variability was also observed in control and ALS cases at the protein level in both regions.

Immunohistochemistry revealed that CXCL7 is localized in the cytoplasm of motor neurons in both control and ALS cases, staining the grey matter region (Figure 7E(a)). CXCL7 immunoreactivity is present in the motor neurons of the anterior horn of the spinal cord and the Clarke's column nuclei, as well as in the neuronal prolongations of the neuropil (Figure 7(b–d)). In the frontal cortex area, CXCL7 positivity is present in a few neurons in the gray matter of control and sporadic ALS cases (Figure 7E(e,f, respectively)). Importantly, these IHC findings map CXCL7's spatial distribution and should not be interpreted as quantifying protein levels or indicating differences in abundance between groups.

4. DISCUSSION

In this study, we employed an integrative, non‐targeted proteomic and lipidomic approach to uncover novel CSF biomarkers for ALS, with a particular focus on addressing the limitations of current biomarkers and providing valuable information about the molecular changes associated with the disease and disease progression outcomes. A key strength of our study lies in the strategic focus on extreme ALS phenotypes—patients with markedly rapid versus slow disease progression. This design enhanced the contrast between clinical subgroups, thereby increasing the signal‐to‐noise ratio and facilitating the identification of prognostically relevant biomarkers. By analyzing individuals at opposite ends of the survival spectrum, we were able to uncover distinct molecular signatures that may otherwise be obscured in more heterogeneous or intermediate cases.

Moreover, the design of the study is aimed to demonstrate that CSF obtained at the time of ALS diagnosis carries prognostic information capable of predicting disease progression and survival. The capacity of a single CSF sample collected at diagnosis to inform both disease classification and survival prognosis has significant implications for clinical management and trial stratification. These possible findings reinforce the importance of early‐stage molecular profiling for enhancing diagnostic accuracy and personalizing therapeutic strategies in ALS.

4.1. CSF proteome signatures distinguish ALS diagnosis and prognostic subgroups based on survival time

SWATH‐MS enables the simultaneous identification and quantification of peptides within a sample in a single step, eliminating the necessity for multiple scans. As a result, SWATH exhibits enhanced throughput, accuracy, and reduced error rates compared to other mass spectrometry methods [58]. Our proteomic analysis of CSF identified 1059 proteins and revealed the existence of 386 proteins that were differentially expressed when compared to ALS and controls. Subsequently, ShinyGO database analysis was applied to the 386 significantly deregulated proteins identified by SWATH‐MS to explore their role in biological processes and determine what pathways the proteins mapped to. The analysis revealed up‐regulated proteins included in functional clusters related to processes of inflammation and immune cells activation and functioning, in addition to down‐regulated proteins linked to detoxification mechanisms components in ALS cases. Our proteomic dataset identified a significant increase in the main immunoglobulin classes and both κ and λ light chains, reflective of the pro‐inflammatory nature of the disease. These data are according to altered inflammatory responses and mechanisms highlighted by the presence of activated microglia and astrocytes, and infiltrating lymphocytes at sites of motoneuron injury, that also contributes to this pathogenic process observed in ALS post‐mortem samples [8, 9, 59, 60, 61, 62, 63, 64, 65] and ALS animal models [66, 67, 68]. In addition, our results are able to successfully replicate some of these prior findings, including the identification of C‐X‐C Motif Chemokine Ligand 12 (CXCL12) [69, 70], Glycoprotein nonmetastatic melanoma protein B (GPNMB) [71], YKL‐40 [72] or transferrin protein [73], among others. However, simultaneously, they also reveal several novel molecular candidates and their corresponding pathways as significant contributors to ALS pathophysiology.

More specifically, our WPCNA identified distinct protein co‐expression modules associated with ALS phenotypes, revealing a clear molecular separation between ALS‐SS, ALS‐LS, and control samples. The progressive changes observed in several modules indicate that ALS progression involves coordinated alterations in multiple biological systems. Modules showing stepwise upregulation in ALS‐SS and ALS‐LS, such as M1‐Black, M3‐Pink, M4‐Green, and M5‐Magenta, suggest the activation of pathways linked to neuro‐synaptic remodeling and post‐translational protein modifications. Conversely, the downregulation of M2‐Greenyellow and M14‐Brown modules highlights early loss of structural and homeostatic functions, including cytoskeletal organization and oxidative stress defense, particularly pronounced in ALS‐SS. The bidirectional patterns seen in modules such as M2‐Greenyellow, M3‐Pink, and M10‐Purple emphasize that ALS‐SS and ALS‐LS phenotypes exhibit distinct proteomic responses, suggesting differential involvement of neurodegenerative and immune processes depending on disease subtype. Together, these findings indicate that ALS pathophysiology is driven by a dynamic interplay of neurodegenerative, immune, and metabolic pathways, with ALS‐LS showing the most extensive proteomic remodeling, whereas ALS‐SS reflects an early breakdown of structural and protective cellular functions. This modular proteomic signature not only provides mechanistic insight into ALS heterogeneity but may also help prioritize protein networks for future biomarker discovery and therapeutic targeting.

The data obtained in our study enhance the current understanding of protein alterations occurring at the onset of the disease and postulate that the protein composition in CSF is capable of fully distinguishing between healthy subjects and ALS patients at the beginning of the disease. Previous studies have also unveiled significant alterations in the immune response components and other altered pathways such as synaptic components or extracellular matrix components in the CSF samples when compared with both conditions; however, the data generated in these studies were less able to differentiate between healthy controls and ALS patients [69, 70, 74, 75, 76, 77, 78, 79, 80].

Additionally, our study was designed to establish novel prognostic biomarkers that would be used potentially in the baseline assessment of a patient with a confirmed diagnosis to predict the clinical outcome based on patients' survival time and disease progression. Current approximations for prognosis determination are restricted to the utilization of a patient classification system based on the site of symptom onset (spinal, respiratory, and bulbar), in addition to other biochemical (such as CSF‐total protein levels) and physiological parameters (such as BMI and muscle denervation), which poorly predict differences in patient pathology, survival, treatment responsiveness, and symptom progression [19, 81, 82, 83]. Only the very recent inclusion of CSF detection of NfL levels has enhanced the management of prognosis in ALS patients, albeit with certain limitations [12] that must be complemented by novel biomarkers. As a result, there is a critical demand for sensitive and specific markers in clinical practice.

To address this gap, a goal of this research was also to identify a molecular signature capable of differentiating individuals with long and short ALS survival times and explore the potential diagnostic value of these markers in ALS. Upon stratifying the CSF proteomic data by ALS survival phenotype, we also discovered alterations in CSF proteomics that involve immune pathways like those previously observed in the comparison between ALS and control cases. Enriched GO terms were partially overlapped when compared comparisons between control group and ALS patients and between ALS phenotypes, mainly in those related to the immune system response that were increased in ALS context, with a more significant impact in the ALS‐SS phenotype. Studied in depth, this further distinguishes between ALS phenotypes, with ALS‐SS cases exhibiting a specific up‐regulated pattern of immune response. ShinyGO database analysis was applied and determined functional clusters related to neutrophil and leukocyte activation, exocytosis mechanisms, and cell adhesion processes; and down‐regulation of protein clusters related to cell projection morphogenesis, cell adhesion, and chromatin structure and nuclear components.

This type of experimental design, based on precision medicine, has been less applied in the context of proteomic approaches in CSF of ALS patients. Only a few studies delve deeper into the investigation of ALS in a personalized medicine‐dependent manner, such as phenotype‐dependent manner [84], genotype‐dependent manner [85] or progression‐dependent manner [86]; but show lesser altered proteins and predictive capabilities. The latter work is more aligned with our experimental design [86], stratifying patients based on their prognosis. However, the discovery and validation cohorts' size were smaller (n = 5 SP and n = 6 fast progressors), and the number of identified altered molecules was also limited (n = 59). Pathways associated with inflammatory responses were also significantly upregulated in the ALS‐FP in this work. Meanwhile, pathways related to synaptogenesis and glycolysis/gluconeogenesis were upregulated in ALS‐SP cases, suggesting the presence of distinct molecular signatures contributing to the progression of the disease [86].

Our results exhibit increased abundance of many proteins of blood origin or related to the peripheral immune system in CSF of ALS, being increased in fast progressor patients. Here, we hypothesize that this fact may be related to brain–blood‐barrier (BBB) and blood–spinal cord barrier (BSCB) disruption that occurred in ALS [8, 87, 88, 89]. A recent study by our group agrees with this hypothesis, demonstrating compromised vascular integrity and increased pouring of peripheral blood components, total protein levels, and albumin levels, to the CNS; being able to act as biomarkers and define prognosis and survival time in ALS patients [83]. In addition, increased levels of inflammatory mediators in CSF and blood of ALS patients have also been associated with disease and/or its progression [69, 90, 91].

Thus, our results introduce unreported molecular data that utilizes proteomic profile characterization to differentiate between different ALS‐survival times, and in consequence, in prognostic outcomes in ALS using 386 differentially expressed proteins. To explore the potential of this data, we developed two machine learning models capable of distinguishing individuals with short and long ALS survival time patients, encompassing both ALS and control patients simultaneously.

The application of machine learning not only enabled effective stratification of ALS subtypes but also facilitated the identification of a minimal and optimized biomarker panel with strong predictive power. By integrating high‐dimensional proteomic data with robust classification algorithms, we were able to reduce complexity while retaining biological relevance and clinical applicability. Importantly, the models demonstrated consistent performance across cross‐validation folds, underscoring their reliability and potential for generalization. These results support the feasibility of leveraging data‐driven models to inform early clinical decision‐making and suggest that integrating machine learning with proteomic profiling could pave the way for more accurate, personalized prognostic tools in ALS. Future work should focus on external validation in independent cohorts and longitudinal datasets to further assess the robustness and temporal applicability of these predictive models.

The machine learning analysis revealed a diverse set of proteins that significantly contributed to distinguishing ALS patients from controls and stratifying survival outcomes, reflecting the multifaceted nature of ALS pathophysiology. Among these, chemokines and inflammatory mediators such as CXCL7, SERPINB1, ELANE, and C1QTNF1 highlight the role of immune activation and neuroinflammation. Structural and cytoskeletal proteins including VIM, PODXL2, DSC1, and PCDHGC5 point to alterations in cellular integrity and adhesion mechanisms. Additionally, several enzymes and metabolic regulators like MPO, ARSA, and MB emphasize dysregulated oxidative stress and metabolic pathways. The panel also comprises immune‐related factors such as IGKV3‐7 and key neurobiological molecules including CRABP1 and TYRO3, which may be involved in neuronal function and survival. Together, these proteins form a comprehensive biomarker panel that captures the complexity of ALS and offers promising candidates for further validation in diagnostic and prognostic applications.

4.2. CSF lipidomic profiles fail to differentiate between ALS patients and healthy controls, or among prognostic subgroups defined by survival time

Conventional lipid concentrations have been examined in readily accessible plasma samples from patients; being some of them proposed as promising biomarkers for the diseases with discouraging results in the ALS context [29, 91]. Nevertheless, numerous lipid species and intermediates go undetected by conventional analytical tools. The advancement of quantitative lipidomics has undoubtedly contributed to a clearer understanding of their characterization and functions. In consequence, more comprehensive lipidomic studies have enabled the identification of a more intricate profile in samples from patient blood, CSF, and even neuronal tissue [92, 93, 94, 95, 96, 97, 98, 99].

In our lipidomic study of the CSF from patients with ALS, we aimed to identify lipid profiles that could aid in diagnosing the disease, predicting its progression, and exploring the underlying pathological mechanisms. The analysis of the lipidome in the same cohort of CSF samples studied in a proteomic approach revealed the identification of 1119 distinct lipid species across the studied groups, but no clear differentiation was observed in CSF lipid profiles among ALS and control samples, nor between ALS subtypes. The number of lipids detected in CSF in the present study is in the same order of magnitude as independent, comprehensive LC/MS CSF analyses [99]. However, only a subset of 15 lipids exhibited notable differences between ALS and control cases in CSF, and just 3 lipids were differentially expressed between ALS‐SS and ALS‐LS. Identified lipids and metabolites have been limited, without any apparent disturbance in pathway functionality, indicating specific and isolated alterations in components of the lipid landscape.

Previously, a few lipidomic studies have been conducted in the CSF of ALS patients compared to controls or ALS subtypes [94]. The reduced modifications observed in lipid composition align with findings from similar studies conducted by our group in the past, where we only identified 17 altered lipids using 23 ALS and 10 control cases [94]. In our earlier research, a noteworthy observation was that plasma lipidomes exhibit a more significant quantitative impact compared to CSF lipidomes. This discovery was unexpected, considering that ALS is primarily characterized as a neurodegenerative disease. The observation was not attributed to a lower lipid variety in CSF; the specificity of CSF changes and plasma variations is further supported by the almost total absence of common differentially annotated lipids between the two compartments. The lipid content (10–13 μg per mL) of human CSF is ca 500 times lower than that of plasma [99]. Instead, these results suggest an overarching process in ALS pathogenesis, specifically in the development of a systemic hypermetabolic state [94]. Similarly, other studies revealed altered lipid components in CSF (n = 48), but it is important to note that control patients from this last study were gender‐ and age‐matched controls from other neurological disorders (Sol et al., 2021) [94].

Our lipidomic data, indicate the limited potential of using CSF lipidomic composition as valuable predictive tools for early clinical prognostic information in the context of ALS. With a medium discriminating capacity, similar to the observed in ROC analysis between ALS and control cases using lipidomic profiling, previous researchers successfully developed significant models to distinguish between ALS patients and controls, as well as to predict the variation in the ALS‐FRS‐R score using the biosigner algorithm [100]. In this study, the most effective discriminant lipids distinguishing ALS patients from controls, albeit with only a 65% accuracy in the test set, were PC(36:4p), PC(36:4e), and SM(d43:2). These lipids exhibited elevated levels in ALS cases, as indicated by the biosigner analysis. Additionally, SM(d34:0), identified as another discriminant lipid by at least two analytical methods, showed higher levels in ALS cases. Conversely, TG(16:1/18:1/18:2) displayed lower levels in ALS cases, further emphasizing its significance as a discriminant lipid [100]. The lipids demonstrating the most robust prognostic capacity for ALS, achieving an accuracy of 71% in the test set, were identified through the biosigner analysis. Specifically, SM(d43:2), TG(16:0/16:0/18:1), and TG(18:0/16:0/18:1) emerged as key indicators, with higher levels of sphingomyelins (SM) and lower levels of triglycerides (TG) associated with a more favorable disease progression. The models predicting the evolution of ALS also revealed the involvement of 7 PCs, importantly, PC(32:1p), which is positively associated with ALS and the deleterious progression of the disease [100].

Our results are also in parallel to those observed in post‐mortem tissue of ALS and FTD patients, where alterations in lipid composition are less abundant [29] than changes reported using other‐omic approaches [9, 101, 102]. For example, previous lipidome analyses conducted on the spinal cord and motor cortex of SOD1‐G93A animals reported only a few altered cholesterol esters (CE) linked to polyunsaturated fatty acids in the spinal cord. However, no significant changes related to ALS were found in the motor cortex [103]. Another example is found in a previous study carried out by our group, focused on the frontal cortex area 8 lipidome at different clinical points of ALS‐FTD disease spectrum, showed quite specific alterations in lipid composition and gene expression of some components of peroxisomes and related lipid pathways [93].

With all of this, we must continue to state that these changes seem to reflect an altered metabolic environment in the CNS in the context of ALS, but probably due to a general process ongoing with ALS pathogenesis and, particularly, with the development of a hypermetabolic state. These alterations, less reflected in CSF, would be appropriate to meet the high energy demands imposed by the increased metabolic rate present in the disease; defined as an unbalance in energy homeostasis [104], with calorie uptake often lowered [105], and energy expenditure increased in a proportion of patients [106].

So, the understanding of the mechanisms underlying this grade of dyslipidemia is still insufficient, and none of them still could be translated into effective tools in clinical practice, but it leaves open the possibility for the discovery of novel derived patients' management tools and contribute to unveil, partially, the physiopathological mechanisms implied in ALS.

4.3. Proteomic data validation highlights CXCL7 levels in CSF as a complementary diagnostic and progression tool in ALS

Upon analyzing the ‐omic data and the discriminative capacity of various molecular profiles, we focused further investigation on proteomic alterations. However, a key limitation of our study lies in the validation phase, where it was not feasible to experimentally confirm the entire panel of top‐ranked proteins identified by the machine learning models. Although our multivariate analysis highlighted several promising candidates, we prioritized CXCL7 for validation due to a combination of technical feasibility—such as the availability of reliable immunoassays—and its biological relevance, particularly its known involvement in inflammatory processes. Consequently, we were unable to assess the full diagnostic and prognostic utility of the broader panel in an independent cohort. While this selective validation approach was pragmatic, it may limit the immediate clinical translatability of the complete model and warrants caution when interpreting the generalizability of our findings. Future studies should aim to validate the full biomarker panel in larger, multicenter cohorts using standardized assay platforms to confirm their predictive value and to explore potential synergistic effects among multiple markers.

Thus, based on proteomic data, statistical significance, machine learning models, and acknowledging the biological inflammatory role of CXCL7, we delved deeper into this chemokine and his potential use as a diagnostic and prognostic biomarker in the CSF of control subjects and individuals with ALS. Chemokines, a family of small molecular weight proteins renowned for their role as leukocyte chemotactic agents, serve as key regulators of the inflammatory response [107, 108]. CXCL7 is a chemokine released by platelet activation in other contexts, but also is expressed in neutrophils, megakaryocytes, natural killer cells, and lymphocytes. The precursor protein of CXCL7, known as the basic proplatelet protein (PPBP), undergoes cleavage to produce two distinct entities: the connective tissue activating peptide III (CTAP‐III) and the β‐thromboglobulin antigen (βTG‐Ag) [109, 110]. CXCL7 exhibits various structural forms including monomers, dimers, and tetramers, which can be interconverted in different states [111]. Monomers are dominant at lower concentrations, while tetramers prevail at higher concentrations. Dimers of CXCL7 are associated with glycosaminoglycans [112]. Additionally, CXCL7 has the capacity to form heterodimers with other chemokines, and the activity of CXCL7 is modulated by glycosaminoglycans [109]. CXCL7 plays an important role in inflammation as chemotactic agent and neutrophil activator, with pro‐inflammatory and pro‐angiogenic effects regulating the activation and function of different immune cells [110, 113, 114, 115]. CXCL7 binds to receptors C‐X‐C chemokine receptor 1 (CXCR1) and/or C‐X‐C chemokine receptor 2 (CXCR2), which are expressed by several inmune cell, including neutrophils and macrophages, as well as non‐inmune cells. This interaction plays a relevant role in inflammation, through signaling pathways such as PI3K/AKT/mTOR, NF‐κB, or MAPK. The consequential activation of these pathways contributes to mediating the role of CXCL7 in inflammation. Notably, abnormal production of the CXCL7 chemokine has been identified in many inflammatory diseases, such as acute lung injury/acute respiratory distress syndrome, autoimmune diseases, graft‐versus‐host disease, and viral infections [116].

In the context of ALS, validation of CXCL7 proteomic findings through ELISA quantification revealed significantly reduced levels in ALS patients compared to controls, supporting its potential as a diagnostic biomarker consistent with the proteomic data. Additionally, CXCL7 distinguished between ALS progression groups and correlated with the ALS‐FRS‐R slope score, underscoring its promise as a prognostic marker. Concurrently, NfL levels were elevated in ALS patients, aligning with numerous prior studies [117]. Despite these findings, NfL outperformed CXCL7 in diagnostic and prognostic assessments, as evidenced by superior ROC curve and AUC values. NfL consistently exhibits greater statistical power and effect size across evaluations, reaffirming its role as a robust biomarker of axonal injury. Nevertheless, CXCL7 remains a biologically relevant and potentially complementary biomarker, particularly in light of the complexity and heterogeneity of ALS pathophysiology, as it may reflect alterations in disease‐related pathways distinct from neuroaxonal damage. CXCL7 provides additional mechanistic insights into non‐neuronal contributors such as immune dysregulation, vascular pathology, and BBB disruption—processes not captured by NfL but potentially critical for patient subtyping, therapeutic targeting, and understanding disease heterogeneity. In consequence, integrating CXCL7 with NfL could enhance both diagnostic accuracy and prognostic precision, particularly when employed in multi‐marker panels. This combined approach may be especially valuable in early‐stage ALS, or atypical presentations.

Moreover, although derived from a small cohort with a specific clinical phenotype, our findings suggest that elevated CSF CXCL7 levels may be associated with longer survival times and prolonged disease duration before the need for ventilatory support in sALS patients. Nevertheless, the limited sample size and phenotypic restriction warrant cautious interpretation, as ALS heterogeneity may limit the generalizability of these results; for example, not all SP exhibit a long lifespan. Further validation in larger, longitudinal, and phenotypically diverse cohorts will be essential to clarify the temporal dynamics, biological significance, and clinical utility of CSF CXCL7 in ALS.

Regarding differential diagnostic capacity, NfL is a general marker of axonal damage, making it highly sensitive but not disease‐specific, as it is elevated across a broad range of neurodegenerative and inflammatory conditions, including multiple sclerosis, traumatic brain injury, and peripheral neuropathies. In this context, our analysis of CXCL7 in ALS mimic disorders did not reveal disease‐specific alterations when compared to myelopathies and IP, mirroring the low accuracy and specificity for differential diagnose that we reported for NfL in the same cohort of mimic disorders [70].

Altered CXCL7 levels are according to the existent literature, where aberrant expression of CXCL7 has been reported in inflammatory diseases and tumors and may serve as a biomarker for disease diagnosis and prognosis [116]. For instance, in hematological diseases, a decline in plasma CXCL7 levels serves as a diagnostic marker for advanced stages of myelodysplastic syndrome refractory anemia with excess blasts (RAEB) and RAEB in transformation (RAEB‐t) according to the FAB classification, or RAEB‐1 and RAEB‐2 [118]. Reduced CXCL7 levels also allow differentiation between pediatric acute lymphoblastic leukemia patients and healthy controls, as well as pediatric acute myeloid leukemia patients [119]. Furthermore, plasma CXCL7 levels exhibit a significant decrease in patients with pancreatic cancer, manifesting at all stages, including the early stages (stages I and II) [120]. Additionally, CXCL7 may play a crucial role in the early diagnosis of gynecological tumors. CTAPIII/CXCL7 is expressed in normal cervical epithelial cells; however, its expression gradually diminishes with the onset and progression of cervical cancer [121]. In the context of neurology, CSF levels of CXCL7 were significantly elevated in Neuromyelitis Optica (NMO) when compared to MS patients and healthy control groups, but not correlated with the patient clinical severity [122, 123].

4.4. Physiopathological role of CXCL7 in post‐mortem human brain tissue of ALS

Despite these modulations in their levels in different pathological contexts, unveiling the exact role of CXCL7 in the pathogenesis of inflammatory diseases remains complex and challenging to elucidate [124, 125]. In consequence, to study its physiopathological implications in ALS, we decided to investigate the levels of mRNA and protein, as well as their localization in post‐mortem human tissue, during the final stages of the disease. It is important to consider that this time‐point represents an advanced or end phase of ALS, which may not fully reflect earlier molecular alterations captured in the CSF at diagnosis.

In sporadic ALS post‐mortem tissue, CXCL7 mRNA levels were found to be increased in the frontal cortex (area 8), but not in the anterior horn of the lumbar spinal cord. No changes in CXCL7 protein levels were observed in either region. Altered CXCL7 transcript levels in the frontal cortex may be associated with early events in the disease continuum and with propagation to secondary brain regions. These changes may not be directly associated with clinical manifestations but may precede them, as our group and other researchers previously observed in ALS in the same region [9, 126], suggesting that CXCL7 dysregulation could represent an early feature of ALS pathology. Additionally, the expression levels of CXCL7 receptors remained unchanged in both regions, with CXCR1 mRNA being undetectable in the spinal cord. However, to better understand this molecular landscape, further studies are necessary.

Notably, this molecular profile did not align with the expression patterns of other inflammation‐related mRNAs and proteins associated, as we have already shown previously, where mRNA levels not correlate with the translated proteins that are increased in the disease [72]. It may be explained due to GFAP protein accumulation may result from post‐translational modifications or impaired degradation rather than increased transcription. In contrast, elevated AIF1 mRNA levels directly reflect microglial activation.

Importantly, for the first time, intense CXCL7 protein staining was observed within the human spinal cord, primarily in MNs, as well as in neuronal processes within the neuropil and in some neurons located in the deeper layers of the frontal cortex (area 8). Despite the presence of CXCL7 immunoreactivity in the remaining motor neurons of sALS cases, the overall protein levels in spinal cord homogenates remained within the range observed in control samples.

These results mark the initial documentation of CXCL7 production by MNs, revealing novel potential functions for this protein within the central nervous system. To the best of our knowledge, this is the first time CXCL7 has been documented in CSF and neural tissue. The observed reduction of CXCL7 in the CSF of sporadic ALS patients may be associated with neuronal damage and the rapid progression of the disease. The observed low levels of soluble CXCL7 in CSF of sporadic ALS patients could be indicative of MN dysfunction. These facts led us to hypothesize that MNs serve as producers of soluble CXCL7, triggering CXCL7 signaling, which may play a role in attenuating CNS inflammation and/or preventing neuromuscular denervation. Nevertheless, additional molecular studies are required to validate this hypothesis. In addition, its role has been also related to motor impairment but related to ossification processes, where it has been identified that CXCL7‐null mice presented with a phenotype of ossification of the posterior longitudinal ligament (OPLL), showing motor impairment, heterotopic ossification in the posterior ligament tissue, and osteoporosis in vertebrate tissue [127].

The observed phenomenon aligns with the case of another chemokine, CXCL13, synthesized by MNs, which undergoes alterations in the context of ALS. This chemokine plays a crucial role in recruitment of B and T lymphocyte and mediates neuroinflammation by interacting with CXCR5 receptor. In ALS SOD1‐G93A (mSOD1) mice, especially in those with fast disease progression, MNs exhibit high expression of CXCL13 during the disease course. In addition, CSF levels of CXCL13 are reduced in ALS patients [128]. An additional chemokine synthesized by MNs and other glial cells is CXCL12, a molecule previously described by our research group [69, 70]. Specifically, we postulated that it is more closely linked to the neuroinflammatory response than neuronal damage. This proposition states that CXCL12‐immunoreactive glial cells are observed in sporadic ALS but absent in control subjects. Nevertheless, CXCL12 introduces a more intricate situation characterized by different events. The distribution of CXCR4 in axonal spheroids and degenerating motor neurons, related with its existence in a subset of oligodendroglial cells and/or precursors in the pyramidal tracts, along with the presence of CXCR7 in motor neurons and reactive astrocytes in the pyramidal tracts. This scenario suggests neuron/axon preservation, oligodendroglial activation, and astrocyte signaling in sporadic ALS.

Thus, all these works reinforce the important and distinct roles of chemokines present in MNs and in the context of ALS, shedding light on the intricate interplay between chemokine signaling, neuroinflammatory processes, and the pathophysiology of ALS. These findings not only deepen our understanding of the molecular mechanisms underlying ALS but also emphasize the need for further investigation into the specific contributions of individual chemokines in modulating neuroinflammation and promoting or mitigating neurodegeneration. Such insights hold promise for the development of targeted therapeutic strategies that could potentially intervene in the progression of ALS by manipulating the chemokine pathways associated with motor neuron health and survival.

4.5. Clinical cohort and study limitations

While the discovery cohort included a relatively limited number of individuals (n = 43), this sample size is consistent with previous studies involving CSF‐based proteomic and lipidomic profiling, especially in the context of rare neurodegenerative diseases such as ALS. The collection of high‐quality CSF samples is inherently challenging due to the invasive nature of lumbar puncture and the requirement for rigorous clinical characterization. The cohort comprised neurologically healthy controls (n = 15) and patients with sporadic ALS (n = 28), further stratified into SS (ALS‐SS, n = 15) and LS (ALS‐LS, n = 13) subgroups. Notably, the LS group was defined based on a minimum clinical follow‐up of 5 years post‐diagnosis—an outcome that represents a relatively rare phenotype within the ALS population, where median survival typically ranges from 2 to 5 years.

Although the sample size is modest, the inclusion of this well‐defined and clinically rare subgroup enhances the biological relevance and translational value of the study. Furthermore, the integration of multi‐omics data—combining both proteomic and lipidomic analyses—with machine learning models such as PLS‐DA and RF enabled the identification of molecular features associated with disease progression and survival. Given the rarity of ALS and the difficulty of assembling well‐characterized, longitudinally monitored CSF cohorts, this dataset represents a valuable resource. Nevertheless, the findings should be considered exploratory and will require validation in larger, independent cohorts to confirm their robustness and clinical applicability.

5. CONCLUSIONS

The application of multi‐omic approaches is crucial for advancing our understanding of the pathophysiology of complex disorders. In this study, we integrated CSF proteomic and lipidomic analyses to explore molecular heterogeneity in ALS. Comprehensive CSF proteomic profiling successfully distinguished ALS patients from healthy controls and differentiated patient subgroups according to survival, revealing a pronounced inflammatory signature, downregulation of detoxification mechanisms, and dysregulation of synaptic organization, components, and structural establishment in more aggressive disease forms. Our findings demonstrate that modular proteomic signatures effectively capture ALS heterogeneity, providing a framework for precision biomarker discovery and the development of targeted therapeutic strategies. In contrast, lipidomic profiles lacked discriminatory power across these groups.

Machine learning models based on CSF proteomic data identified CXCL7 as a promising candidate biomarker for ALS diagnosis and prognosis, and associating its CSF levels with disease duration and survival. Validation studies confirmed that NfL remains the most robust and consistent predictor of ALS progression, yet CXCL7 shows potential as a complementary biomarker, offering additional insights into ALS biology and supporting refined clinical stratification. Future studies in larger, independent cohorts are warranted to elucidate the mechanistic role of CXCL7 and confirm its clinical relevance in broader settings.

AUTHOR CONTRIBUTIONS

SPR, YLS, PMS, and MPF contributed to experimental analysis, data representation, and interpretation. JFVC, MJC, CM, ACP, RD, ILM, AMY, and SMY contributed to clinical data and sample acquisition, data interpretation, and supervision. ES, JFI, GS, EO, MPO, and GPR contributed to omic data processing, data acquisition, resources, and critical revision of the manuscript. MP and PAB conceived and designed the study, obtained funding, supervised the project, interpreted the data, and drafted and revised the manuscript. All authors reviewed and approved the final version of the manuscript.

FUNDING INFORMATION

PAB and MP were supported by the “Fundació Catalana d'Esclerosi Lateral Amiotròfica (ELA) Miquel Valls,” and more specifically to the fundraising campaigns: “Elles contra l'Esclerosi Lateral Amiotròfica (ELA),” “Swim for Esclerosi Lateral Amiotròfica (ELA)” and “Mà amiga.” MPO was supported by grants from the Instituto de Salud Carlos III (PI20/000155, PI23/00176), and PAB was supported by grants from the same governmental institution (PI24/01516, CD23/00097). The project that gave rise to these results received the support of a fellowship to PMS from Fundación Ramón Areces. We want to particularly acknowledge patients and Biobank HUB‐ICO‐IDIBELL (PT20/00171) integrated in the ISCIII Biobanks and Biomodels Platform and Xarxa Banc de Tumors de Catalunya (XBTC) for their collaboration. We thank CERCA Programme/Generalitat de Catalunya for institutional support.

CONFLICT OF INTEREST STATEMENT

The authors declare that they have no competing interests related to this publication.

ETHICS STATEMENT

This study was conducted in accordance with the principles of the Declaration of Helsinki. Ethics approval was obtained from the Clinical Research Ethics Committee of Bellvitge University Hospital, reference number PR010/22.

CONSENT TO PARTICIPATE

Written informed consent was obtained from all participants (or their legal guardians, where applicable) prior to their inclusion in the study.

Supporting information

Data S1: Supporting Information

BPA-36-e70089-s001.xlsx (500.8KB, xlsx)

DATA AVAILABILITY STATEMENT

The data that support the findings of this study are available from the corresponding author upon reasonable request.

REFERENCES

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data S1: Supporting Information

BPA-36-e70089-s001.xlsx (500.8KB, xlsx)

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Brain Pathology are provided here courtesy of Wiley

RESOURCES