Skip to main content
GigaScience logoLink to GigaScience
. 2026 May 21;15:giag059. doi: 10.1093/gigascience/giag059

Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer’s and Parkinson’s disease

Anwesha Bhattacharya 1,2,3, Edward A Fon 4, Alain Dagher 5,6, Yasser Iturria-Medina 7,8,9, Jo Anne Stratton 10, Chloe Savignac 11,12,13, Jack Stanley 14,15,16, Liam Hodgson 17,18,19, Badr Ait Hammou 20,21,22, David A Bennett 23, Danilo Bzdok 24,25,26,27,✉
PMCID: PMC13289754  PMID: 42166149

Abstract

Aa. Background

Historically, Alzheimer’s disease (AD) and Parkinson’s disease (PD) have been investigated as 2 distinct disorders of the brain. However, a few similarities in neuropathology and clinical symptoms have been documented over the years. Traditional single-gene centric studies, such as differential gene expression analyses, have struggled to unravel the molecular basis for the observed pathological links between AD and PD.

Ab. Results

We tailor a latent factor framework to analyze synchronous gene co-expression changes in AD or PD at sub-cell-type resolution. Utilizing large, single-nucleus transcriptomics datasets in AD (70,634 nuclei) and PD (340,902 nuclei) from postmortem human brains, we systematically extract and juxtapose disease-critical molecular signatures in the brain. Our transcriptomic analysis reveals shared molecular programs between AD and PD that localize to specific glial and neuronal cell types. In neurons, convergent gene groups in AD and PD relate to cytoskeletal dynamics and mitochondrial stress mechanisms. In microglia, overlapping gene modules implicate T cell activation mechanisms and synapse pruning pathways. In parallel, AD- and PD-associated gene groups in astrocytes are involved in heavy metal processing; oligodendrocytes highlight convergent dysregulation in myelin synthesis. Additionally, our analysis reveals apolipoprotein E gene (an AD risk gene), and the synuclein alpha gene (a PD risk gene) to have disease predictive roles in both AD- and PD-associated gene modules.

Ac. Conclusion

Our multi-module sub-cell-type approach offers novel insights into the molecular basis of shared neuropathology in AD and PD.

Keywords: latent gene programs, latent gene programs, cytoskeleton dynamics, heavy metal processing, synapse pruning, cross-disorder analysis

Introduction

Alzheimer’s disease (AD) and Parkinson’s disease (PD) are 2 of the most prevalent disorders in today’s aging societies [1, 2]. There has been intensive research with the grand aim of altering and ultimately halting the course of these diseases. Despite educated forecasts predicting significant advances by this decade [3], AD and PD remain challenging to unravel. This difficulty is compounded by a historically entrenched dichotomy that has limited transfer of research insights from one disease to the other. AD and PD are considered distinct entities due to differences in primary brain regions affected [4–7], age of onset, clinical progression, and treatment response. PD is notably responsive to therapeutics that do not affect cognition [8], and AD is without any “hard-currency” therapeutic to date [9].

This dichotomy between AD and PD continues to be reinforced by genomics and polygenic risk studies, which show minimal to no overlap of genes between AD and PD [10, 11]. Indeed, aggregating prior findings, a Neuron review recently concluded, “There is intriguingly little overlap between the risk genes for AD and PD, providing genetic evidence for different disease onset and progression mechanisms ” [12].

By contrast, autopsy examinations show that over half of PD patients have aggregates of tau [13] and around 30% of PD patients develop cognitive impairment, with many going on to dementia [14]. Conversely, AD pathology and Lewy bodies co-occur more frequently than by chance, and Lewy bodies are associated with cognitive decline [7, 15, 16]. Further, neurofibrillary tangles in the substantia nigra can give rise to symptoms of Parkinsonism [17]. Together, these observations raise the possibility of shared disease mechanisms underlying the pathology of AD and PD [18, 19].

More broadly, recent systems-level analysis further suggests that AD and PD share overlapping biological pathways [18]. Proposed mechanisms include mitochondrial dysfunction, neuroinflammation, and dysregulated protein homeostasis [20]. Involvement of these processes has also been suggested in several other neurodegenerative disorders, reinforcing the view that neurodegenerative mechanisms may arise from partially overlapping manifestations of shared molecular networks rather than completely independent disease mechanisms [19, 21].

Despite this evidence, most genomics studies directly questioning the genetic and molecular basis of AD and PD overlap have not identified clear underlying mechanisms. Methodologically, these studies have largely relied on univariate approaches that focus on contribution of individual genes to diease [22–25]. In reality, however, gene expression occurs within tightly regulated environments where gene products interact in highly combinatorial ways [26–29]. Thus, the pathogenesis of neurodegenerative diseases is likely driven by dysregulation of gene networks rather than by isolated gene anomalies [30–32].

Moreover, the effects of dysregulated genes are not identical across different cell types. Large-scale single-nucleus transcriptomic studies in AD [33, 34], and PD [4, 35], have reported disease-associated transcriptional changes that are highly cell type-specific. For example, in AD, APOE expression is increased in microglia and decreased in astrocytes and oligodendrocyte precursor cells (OPCs) [36, 37]. Given such complexity at the cellular level, it is crucial to account for cell type heterogeneity when comparing AD- and PD-related changes. This is evidenced by the minimal AD-PD overlap derived from analysis of bulk RNA sequencing data [38]. Recent advances in single-nucleus RNA sequencing (snRNA-seq) have enabled the resolution of transcriptional signatures with high cell-type specificity. However, in contrast to recent proteomics- and CSF-based AD-PD comparative analyses [39], no single transcriptomics study has employed multivariate approaches at the single-cell resolution to investigate AD and PD simultaneously at scale.

In this present investigation, we systematically revisited the problem of identifying candidate molecular mechanisms that overlap between AD and PD through the lens of transcriptomics. Comprehensive snRNA-seq datasets from AD (70,634 nuclei from 48 postmortem brains) [36] and PD (340,902 nuclei from 15 postmortem brains) [40] allowed us to leverage advanced machine learning techniques for our analysis [41]. Enabled by a supervised multivariate model [42], we extracted biologically interpretable and diease relevant gene modules at sub-cell type granularity from the AD and PD transcriptomes (16,936 protein-coding transcripts). By linking our gene modules to curated biological pathways, we identified shared candidate mechanisms of neurodegeneration. In addition, we validated our main findings and conclusions using an independent AD-PD snRNA-seq dataset pair. Overall, our supervised pattern learning-based comparative approach provided a statistically rigorous and unified framework for the automatic discovery and comparison of disease-associated molecular signatures across disorders.

Results

Rationale

Traditional approaches in single-cell transcriptomics, such as DGE, typically focus on individual genes in isolation. This offers a fragmented view of disease-related transcriptional changes. Recent studies employing univariate frameworks to identify disease-associated genes have reported modest overlap in AD- and PD-relevant genes, primarily limited to glial populations [43]. However, we hypothesized that such methods may fail to capture co-regulated molecular alterations in AD and PD. Further, these molecular changes likley have a cell type-specific component, which are missed by bulk sequencing methods. In addition, we reasoned that disease progression within each cell type was not driven by a single molecular axis. Rather, multiple distinct transcriptional programs were altered. Therefore, a biologically coherent AD and PD comparison would entail examining multiple gene programs within distinct cell types to capture the full complexity of their molecular convergence.

To address these dimensions, here we interrogated AD- and PD-transcriptional changes using gene modules representing coordinated gene activity across the transcriptome. We extracted disease-relevant gene modules using a supervised latent factor modeling framework, previously validated in AD [42]. This method combined latent structure discovery strength while simultaneously being aware of valuable contextual information—the disease state. Using these discovered gene modules, we a rigorous cross-disorder comparative framework to study AD- and PD-associated molecular changes.

Notably, by examining the entire recorded transcriptome, our data-driven pipeline enabled an unbiased assessment of genes, without prior assumptions about AD- or PD-association. In addition, our approach could assign a gene as being relevant to multiple gene modules within a given cell type—a feature absent from many previous gene network analysis techniques [31]. Overall, our multivariate analysis is an unbiased investigator of the transcriptomic landscape of AD and PD related molecular mechanisms.

Disease predictive, cell type-specific gene modules identified using latent factor modeling

We explored the possibility of molecular overlap in AD and PD brains through transcriptional alterations captured in computationally derived gene modules (groups of co-expressed genes). The main analyses were conducted on 2 snRNA-seq datasets—Religious Orders Study and Memory and Aging Project (ROSMAP)-AD [36] (70,634 nuclei across 8 major cell types; the “Methods” section) and Kamath-PD [40] (340,902 nuclei across 11 cell types). Our analytical framework employed partial least squares discriminant analysis (PLS-DA) to gain an overview of 16,936 protein-coding genes (see the “Methods” section).

In either AD or PD, we fitted cell-type-level PLScell models across nuclei from all donors in the dataset (Fig. 1; the “Methods” section). This yielded gene modules as latent projections of gene expression matrices that assisted in distinguishing cells of patients from controls. Comparative assessments of these thus-derived gene modules highlighted shared molecular mechanisms between AD and PD that were more stable than expected by chance (Fig. 2A).

Figure 1.

Workflow figure detailing the steps for disease-relevant latent gene module identification from single nucleus RNA sequencing datasets, followed by comparative analysis between Alzheimers and Parkinsons disease. Gene modules are identified independently for each cell type within the AD or PD datasets.

Workflow diagram to test for AD-PD overlap: bottom-up approach. Overview of workflow. (A) Single nucleus RNA sequencing (snRNA-seq) datasets for AD and PD were downloaded from public databases. Filtering, quality control, and cell type annotations were performed by source authors. (B) Preprocessing step. Performed independently for each cell type, this step followed recommended guidelines for data transformation, removed lowly expressed genes, and corrected for disease versus control class imbalance. (C) Gene module identification step. PLS-DA was performed per cell type to extract weighted gene lists, referred to as gene modules, that were disease predictive. (D) Comparative analysis. We aggregated the results from the 2 analysis arms and assessed overlap using parallel methods—(i) direct correlation of cross-disease gene module pairs and (ii) gene set enrichment analysis to identify overlapping biological processes, molecular functions, and cellular components. We discover significant molecular similarities between AD and PD across cell types. PLS, partial least squares; AD, Alzheimer’s disease; PD, Parkinson’s disease; GSEA, gene set enrichment analysis; GO, gene ontology.

Figure 2.

Top left, graph quantifying the overlap between Alzheimer's and Parkinson's disease. Oligodendrocyte gene modules show the most significant overlaps between AD and PD. Microglia and astrocytes also show strong correlations between cross-disease gene modules. Graphs on the right and bottom left depict examples of genes and gene modules extracted from the independent AD and PD datasets.

Convergence of transcriptomic signatures in Alzheimer’s and Parkinson’s disease across cell types when zoomed in on disease-predictive gene modules. We probed 2 snRNA-seq datasets (ROSMAP-AD, Kamath-PD) to explore the transcriptomic overlap between AD and PD. By training 15 PLS models (one for each cell type in AD (6) or PD (9)), we extracted latent representations of gene expression (gene modules) that maximized the separation between disease and control nuclei. (A) The colored squares represent Kendall’s tau-b (τb), quantifying the degree of association between AD and PD gene modules. Darker colors represent stronger genetic associations, indicating similar ranking trends of robust genes in the module pair. Highlighted boxes with borders denote statistically significant τb exceeding 2.5/97.5% CIs based on label-shuffled permutation tests (n = 1,000). Square size is proportional to FDR-corrected statistical significance of overlap. The wide range of correlation strengths suggests that AD and PD share significant molecular similarity spanning the cell type landscape. (B) An illustration of the robustness assessment of genes in a gene module. The 30 genes with the highest empirical predictive weights from the first PLS component in ROSMAP-AD microglia are shown. Violins represent the loading distribution of genes derived from a BS resampling scheme (n = 500). Robust genes have non-zero predictive weights (2.5/97.5% CIs). Genes that do not meet this criterion are grayed out. APOE, a gene strongly associated with AD, was highlighted as a major disease predictor in this component. (C) The colored portions of the bars represent the number of genes with robust disease-predictive weights for each gene module. The gray regions denote the total number of analyzed genes for the cell type. Robust genes were assessed based on a 500-iteration BS resampling scheme. Ast, astrocyte; Ex, excitatory neuron; In, inhibitory neuron; Mic, microglia; Oli, oligodendrocyte; OPC, oligodendrocyte precursor cell; PLS, partial least squares.

As a first modelling step, we independently analyzed the transcriptomes of AD and PD to identify disease- and cell type-specific modeling hyperparameters. Specifically, the optimal number of gene modules that best distinguished diseased from healthy cells (in unseen cells) was selected using a 10-fold nested cross-validation (CV) scheme (see the “Methods” section). In ROSMAP-AD, 2 gene modules per cell type were determined to be optimal for all 6 cell categories. In Kamath-PD, the optimal number of gene modules was determined to be 2 for each major cell type, except endothelial cells (3 modules) and CALB1 dopaminergic neurons (3 modules).

The fitted cell type specific PLScell models (6 AD and 9 PD models; optimal hyperparameters) exhibited robust above-chance out-of-sample accuracy in differentiating disease samples from control (Fig. S1A). Unbiased classification accuracy was estimated based on a patient-partitioned CV scheme (see the “Methods” section). Specifically, in the AD versus control group contrast, the predictive power measured by AUROC ranged highest for microglia (AUROC: 0.66 ± 0.06 std across partitions) to lowest for OPCs (0.56 ± 0.08 std across partitions). For the PD versus control group contrast, the highest AUROC was for endothelial cells (0.89 ± 0.13 std across partitions), and the lowest was for excitatory neurons (0.69 ± 0.35 std across partitions). These observations strongly supported the role of gene modules in directly informing the disease phenotype across all examined cell types and conditions.

Our derived gene modules (PLScell components) thus featured a combination of several genes whose co-expression signature was associated with a disease state (AD versus control or PD versus control). We assessed the statistical significance of each module by comparing its empirical disease prediction performance to a null distribution of performance metrics derived by a label-shuffling permutation procedure (the “Methods” section). Only significant gene modules were considered for subsequent analyses (empirical module ρ > 97.5th percentile of permutation derived ρ distribution; Fig. S1B). In particular, 2 cell types in ROSMAP-AD, pericytes and ependymal cells, did not pass our significance test and were removed from further analysis. Two cell types in Kamath-PD, macrophages and ependymal cells, were similarly removed.

We identified stable genes within each PLScell gene module (12 AD modules, 20 PD modules). Specifically, using a bootstrap (BS) resampling technique (see the “Methods” section), we assessed which gene effects were statistically robust (zero not included in the 2.5/97.5% confidence interval (CI) of the BS distribution of each gene), and thus, reliably affected prediction outcomes. As an illustration, we visualized the PLSMic loadings for the first gene module from ROSMAP-AD microglial cells (Fig. 2B). Overall, each gene module yielded a variable set of robust genes distributed across the transcriptome (Fig. 2C).

We further investigated the characteristics of the derived modules using a clustering algorithm. For a cell type, we assigned each observed nucleus to exactly one of its modules based on the component harboring the maximum PLScell score. In doing so, we were able to visualize the distribution of the gene modules assigned to the nuclei in a 2-dimensional embedding space (Fig. S2; PHATE embedding space; the “Methods” section). No clear clustering among the components was observed suggesting that our gene modules did not necessarily correspond to cellular subtypes. Instead, they likely corresponded more closely with different functional programs within a given major cell type. That is, any given cell belonging to a type could exhibit several of our distinct gene programs to various continuous degrees.

In a stringent external validation analysis using untouched datasets, we derived independent AD and PD gene modules in 2 additional snRNA-seq datasets—Seattle-AD and Smajić-PD (detailed cohort and sample description in the “Methods” section). We repeated all main analyses, from scratch, and derived sub-cell-level disease predictive gene modules from these datasets (Fig. S3B; the “Methods” section). In Seattle-AD, 15 significant gene modules emerged across 10 examined cell types. Independently, in Smajić-PD, 25 significant gene modules emerged across 7 cell types. Significance of modules was determined, as before, using a label-shuffle permutation test. These independently derived gene modules were subsequently used to corroborate our findings from the primary AD-PD overlap analysis (see next section).

We assessed whether the latent gene modules were influenced by potential demographic variables. Concretely, we examined the contribution of age and sex to variation in module scores, quantified by the proportion of variance explained using a linear regression model (the “Methods” section). Across gene modules, the contribution of sex toward explaining the variance in module scores was low within each dataset (mean ± standard deviation across modules): ROSMAP-AD, R2 = 0.01 ± 0.01; Kamath-PD, R2 = 0.07 ± 0.09; Seattle-AD, R2 = 0.017 ± 0.018; Smajić-PD, R2 = 0.12 ± 0.10. Similarly, age explained little variance in module scores: ROSMAP-AD, R2 = 0.003 ± 0.006; Kamath-PD, R2 = 0.04 ± 0.08; Seattle-AD, R2 = 0.018 ± 0.020; Smajić-PD, R2 = 0.04 ± 0.06. Overall, across all gene modules from the 4 datasets, age and sex explained only a small fraction of the variance in module scores relative to the variance captured by diagnosis, the primary variable of interest (Table S10).

Importantly, in all analyses so far, the gene expression samples from AD and PD were not merged. Instead, these were conducted in parallel with independent supervision targets (AD-control or PD-control). Thus, we identified robust AD- and PD-predictive gene modules in a cell type-specific manner. These thus derived modules, specific to AD or PD, enabled subsequent comparative analyses aimed at examining the shared molecular alterations between AD and PD.

Shared molecular signatures uncovered between AD and PD, across major brain cell types

We subsequently moved to our comparative analysis between AD and PD. To quantify the coupling between AD- and PD-derived gene modules, we used Kendall’s tau-b (τb) metric ( the “Methods” section). This calculated the degree of similarity between 2 vectors of PLScell predictive weights of the overlapping robust genes (not the gene expression measurements) from an AD-PD gene module pair. By comparing the empirical correlation strength with a permutation-derived null distribution (obtained by correlating permutation modules derived from a label-shuffled dataset), we identified the module pairs with significant associations (empirical τb more extreme than 2.5/97.5% CI of the null distribution). Notably, we reported only those associations that survived this exhaustive permutation-based validation.

As the most important results of our investigation so far, we noted strong and robust associations across several AD-PD module pairs (robust to 1,000 iterations label shuffle permutation test; τb false discovery rate (FDR) q-value < 0.05; Fig. 2A; Table S1; the “Methods” section). Across all pairwise combinations, the most significant correlation (smallest q-value) emerged between the first oligodendrocyte module (represented as Oli 1) from ROSMAP-AD and the second oligodendrocyte module (Oli 2) from Kamath-PD (represented as Oli 1_Oli 2; τb, abs = 0.44, number of shared genes = 1,351, q = 5.4e-128). In addition to Oli 1_Oli 2, other top significant pairs included Oli 1_Ex 1, Oli 1_Ex 2, In 2_Ex 1 (exhibited τb, abs > 0.3, and shared over 500 robust genes).

In contrast, the strongest associations (ranked by absolute τb) were observed between glial modules—astrocytes, microglia, and OPCs. Specifically, the highest association strength was observed between Mic 2_Mic 1 (τb, abs= 0.67; q = 0.009), followed by Ast 2_Ast 2 (τb, abs = 0.62; q = 4.2e-5). These results suggest that in glial cells, disease associated molecular programs are broadly shared between AD and PD.

On the flipside, inter-neuron module pairs (Ex and In pairs) between AD and PD showed the least overlap, across all module pairs. However, dopaminergic neurons in PD—CALB1, and SOX6—showed significant similarities with AD-critical gene programs in neurons (Ex 1_Calb1 1, τb, abs = 0.35, q = 3.1e-20; In 1_Calb1 1, τb, abs = 0.32, q = 1e-5), along with AD-associated oligodendrocyte and astrocyte modules (Oli 1_Calb1 3, τb, abs = 0.19, q = 8.21e-5). Thus, in contrast to the major recorded PD neuronal populations (Ex and In), disease-critical PD-neurons (highlighted in original study [40]) had significant molecular similarities with several AD-associated molecular programs.

As a sanity check, we assessed how the cross-cell type module associations compared within a single disease. Specifically, we looked at the correlation effect sizes (statistically significant under a label shuffled permutation test, q < 0.05; the “Methods” section) across gene modules in a ROSMAP-AD versus ROSMAP-AD comparison, and a Kamath-PD versus Kamath-PD comparison (Fig. S1D). The presence of robust association signatures between gene modules from different cell types within the same disease supports the cross-cell type module-level overlaps observed between AD and PD.

We validated our primary findings by examining molecular similarity between Seattle-AD- and Smajić-PD-derived gene modules. Our analysis revealed significant module-level overlaps between AD and PD, scattered across different cell types (Fig. S3A; Table S2). Across all modules, inter- and intra-oligodendrocyte module pairs from AD and PD took center stage, aligning with our primary analysis. The strongest association was observed between oligodendrocytes from Seattle-AD and Smajić-PD (Oli 1_Oli 1, τb, abs= 0.51, q = 5.5e-26), followed closely by Ast 1_Oli 1 (τb, abs= 0.26, q = 9.2e-16). Strong significant associations were also observed between different combinations of neuron and glial cell-derived modules (Ast 1_L4_it 1, τb, abs= 0.45, q = 1.6e-9; Oli 2_In 3, τb, abs= 0.86; q = 0.05). Further, inhibitory neuron-derived modules showed sparse similarities between AD and PD, similar to our observations within the primary datasets. Overall, this external validation of shared transcriptomic signatures between AD and PD indicated that the observed overlaps were unlikely to be driven by dataset-specific factors such as transcriptomic platform, cohort composition, or brain region selection.

We further assessed the generalizability of our comparison model by randomly partitioning the ROSMAP-AD and Kamath-PD datasets into 2 non-overlapping subset pairs (split-half test; see the “Methods” section). We then repeated the full analysis pipeline on each subset and compared the resulting pairwise associations between the subset-derived gene modules. Our analysis revealed strong associations across different realizations of the partitioned, but otherwise identical, analyses (Pearson’s rho = 0.92 ± 0.02 standard deviation; Fig. S1C). This result provided additional quantitative support for the associations from the full datasets, indicating that the observed relationships were robust to sample size.

We assessed the specificity of our findings from the AD-PD comparative analysis with respect to other systemic diseases. Specifically, we performed a comparison of AD- or PD-derived modules with chronic obstructive pulmonary disease (COPD) of the lung (neurological versus non-neurological). COPD was selected as a negative control condition given the expectedly different cellular composition and tissue context relative to the brain. Following our established analytical pipeline, we computed pairwise similarities between AD- and PD-associated gene programs with COPD-associated gene programs across major recorded lung cell populations (AD-Lung, PD-Lung; derived following Fig. 1 pipeline; Tables S5–S8).

Overall, we observed limited overlap between brain and lung disease modules. The modest similarities that were detected primarily involved immune-related cell populations and endothelial cells from the lung and the brain. Specifically, neuronal populations from all 4 AD and PD datasets showed little overlap with the lung gene modules (maximum Mye 1_Ex 1 (Lung_Kamath-PD), τb, abs = 0.49, FDR q = 1.43e-181). In contrast to the neuronal cells from the brain, lung myeloid (Mye) and lymphoid (Lymp) cells showing significant overlap with brain glial cells (maximum Lymp 1_Mic 1 (Lung_ROSMAP-AD), τb, abs = 0.74, FDR q = 6.4e-4). Taken together, these results suggest that neuronal gene programs identified in AD and PD are largely brain-specific, while a subset of immune-related signatures are shared across tissues and disorders.

Collectively, these findings revealed significant molecular overlap between AD and PD at a sub-cell type resolution. The degree and specificity of these overlaps varied between cell type-specific gene module pairs, with oligodendrocyte and neuron-based gene module combinations in AD and PD signaling the strongest similarities. Our external validation experiments replicated these core findings in independent datasets, further substantiating our conclusions regarding the shared genetic architecture between these neurodegenerative diseases.

Cell type-specific gene modules reveal GWAS derived genes as key predictors of disease

We contextualized our gene modules post hoc to understand their relationship with known risk genes from genomic studies. Drawing from the most recent GWAS that reported AD [44] or PD [45], we investigated 164 genes (Table S3; see the “Methods” section), locating them within our gene modules. Most genes implicated by major GWAS risk loci (here refered to as GWAS genes) showed robust disease-predictive loadings in at least one gene module (Fig. 3). The top gene module harboring the most AD GWAS genes was in AD Ex 1, with 25.6% of AD GWAS genes present, followed by AD Ast 1 with 24.3% GWAS genes present. In PD, the top modules with most PD GWAS genes was Oli 1 with 58.9% PD GWAS genes followed by Ex 2 with 42%. Moreover, we found clear cell type localization of these genes within our modular framework. For example, APOE, a broadly accepted AD risk gene, showed robust predictive loadings in Ast 1 and Mic 1 ROSMAP-AD modules. Similarly, LRRK2, one of the major PD risk genes, was implicated in distinct Kamath-PD gene modules. The strongest effect was observed in PD Mic 2.

Figure 3.

Graphs localizing genes implicated by major genome wide association study risk loci in AD or PD within cell type-specific gene modules indentified in this study.

Cell type-specific gene modules reveal GWAS-associated genes as key disease predictors. We mapped the contribution of 164 candidate GWAS-nominated genes, compiled from the largest AD and PD GWAS studies, within our gene modules. Color saturation represents the strength of the predictive weight of a gene in a module as determined by the PLS model. Robust genes are highlighted with black stars, defined as those with BS-derived 2.5/97.5% CIs that do not include zero. (A–C) Genes are grouped based on the nominating disease. Group (A) contains 5 common genes nominated independently in both AD and PD. (B) shows the genes implicated in AD GWAS. (C) shows genes implicated in PD GWAS. (D) The bars on the right summarize the connection of a module with GWAS-mapped genes. The lengths indicate the percentage of GWAS-mapped genes with robust loadings in the module. Cell type-specific module-level localization of risk genes was observed. APOE, a major AD risk gene, showed strong predictive loading in ROSMAP-AD microglial and astrocyte modules. For PD, SNCA had strong predictive signals in Kamath-PD oligodendrocyte, microglia, excitatory neurons, and CALB1 dopaminergic neuron modules. Several genes that were implicated as being AD-relevant in GWAS had robust predictive loadings in Kamath-PD gene modules and vice versa. For example, APOE, APP, and other AD genes mapped to risk loci had robust predictive weights in Kamath-PD modules. Likewise, SNCA, MAPT, and other PD risk genes had robust predictive weights in ROSMAP-AD modules. Ast, astrocyte; End, endothelial; Per, pericyte; Ex, excitatory neuron; In, inhibitory neuron; Mic, microglia; Oli, oligodendrocyte; OPC, oligodendrocyte precursor cell; GWAS, genome-wide association study.

Next, we analyzed the effects of GWAS genes from one disease within gene modules linked to the other neurodegenerative disease. In other words, we investigated whether genes mapped to AD GWAS risk loci were highlighted in any PD modules and vice versa. Among the AD GWAS genes, top genes, including APOE, APP, BIN1, and CLU (based on P-value from GWAS [44]) had robust disease predictive loadings in gene modules associated with PD (Kamath-PD modules). Concretely, APOE had strong predictive weight in PD Mic 2 (gene loading = −0.06, max absolute loading for any gene in this module was 0.07, absolute rank of this gene = 23), APP in PD Mic 1 (loading = 0.04, maxabs = 0.07, rank = 262), BIN1 was found in PD Calb1 2 (loading = −0.02, maxabs = 0.04, rank = 1,626), and CLU had the strongest robust weight in PD Oli 2 (loading = 0.03, maxabs = 0.08, rank = 405). It is important to note that these genes were not the highest-ranked in the respective PD modules. In other words, they were not the primary disease-indicative genes (rank = 1). Instead, these genes likely played supporting roles that become apparent only within the context of the gene groups.

Conversely, we made similar observations for previously reported genes implicated in PD GWAS studies [46] within our (ROSMAP-)AD modules. The gene enconding synuclein alpha (SNCA) had strong predictive loading in AD Ex 1 (−0.03, maxabs = 0.07, rank = 397), microtubule-associated protein tau (MAPT) in AD In 1 (−0.02, maxabs = 0.07, rank = 979), and TMEM175 in AD Ex 1 (−0.02, maxabs = 0.07, rank = 1,869). Notably, while the genes implicated in AD GWAS had very clear cell type localizations, the GWAS genes associated with PD tended to be distributed across modules derived from multiple cell types, with particularly high effect sizes in neuronal modules. This observation is consistent with a previous genomic enrichment study showing that PD risk loci are not confined to specific cell types. Instead, they are associated with broad cellular processes observable across multiple cell types [47]. Collectively, our observations suggested that the bona fide GWAS genes not only tracked the disease they were implicated in, but also proved relevant in gene modules associated with the other disease.

In our external validation analysis, GWAS-mapped genes exhibited consistently strong disease-predictive loadings in gene modules derived from the Seattle-AD and Smajić-PD datasets (Fig. S4). Notably, all key observations from our primary dataset analysis were replicated. First, the genes tracked disease-corresponding modules; microglia and astrocyte Seattle-AD modules faithfully tracked APOE (interestingly, we observed robust predictive loading for APOE in Seattle-AD OPC modules, consistent with a previous study [36]). Similarly, PD GWAS genes, such as SNCA and LRRK2, were tracked by Smajić-PD gene modules. Second, mirroring our primary dataset pair, several GWAS-mapped genes were enriched in modules associated with the opposite disease (for instance, APOE tracked Smajić-PD modules and SNCA tracked Seattle-AD modules), further reinforcing the presence of cross-disease molecular convergence.

In summary, here we contextualized the suspected genes from AD- or PD-GWAS risk loci within the scope of transcriptome derived gene modules. This led us to identify significant effects of GWAS genes beyond the disease in which they were initially reported. For example, we observed significant impact of APOE not only in AD (microglia and astrocyte-specific modules), but also in PD-derived modules (microglia, astrocytes, and CALB1 DA neurons). These cross-disease results were replicated in independent validation datasets. Crucially, these observations were only possible through our approach, which evaluated the joint contribution of statistically meaningful gene sets to disease status. Thus, for example, although APOE may not show a strong univariate effect in PD, it might play a significant role when considered within the context of co-expressed genes in PD gene modules.

Overlapping biological functions between AD and PD gene modules

We contextualized the biological relevance of the observed AD-PD overlap using comprehensive gene set enrichment analyses (GSEAs). GSEA was performed independently for each gene module (12 ROSMAP-AD and 20 Kamath-PD). Notably, in our analysis, a single gene could contribute significantly to disease via multiple modules within the same cell type. This enabled us to capture likely gene effects on complementary, co-regulated pathways.

By screening the widely relied upon gene ontology (GO) databases corresponding to 3 complementary domains—biological processes (BPs), molecular functions (MFs), and cellular components (CCs)—we identified AD or PD relevant GO terms. Across all gene modules from ROSMAP-AD, 284 BP, 56 MF, and 102 CC terms were identified. In Kamath-PD, we obtained 715 BP, 152 MF, and 209 CC terms. In total, a universe of 27,993 GO BP, 11,271 GO MF, and 4,039 CC terms was analyzed.

Among all terms across gene modules, 158 BP, 33 MF, and 80 CC terms were shared between AD and PD (Table S4). The relatively small subset of overlapping terms highlighted the non-random and specific nature of our gene modules (Fig. 4A). The greatest number of GO term overlaps emerged in neuron–neuron AD-PD module pairs, along with pairs involving oligodendrocytes and OPCs (>160 terms per module pair, BP, MF, and CC combined; Fig. S5A, note Fig. 4B horizontal axis). Additionally, both AD and PD microglial gene modules shared, on average, 20 terms with other gene modules. In contrast, module pairs involving astrocytes featured lower overlaps, with the maximum number of shared terms being 13 between AD Ex 1 and PD Ast 1.

Figure 4.

Graphs showing gene set enrichment analysis results of Alzheimer's and Parkinson's disease relevant gene modules. Gene ontology terms indentified highlight cell type specific overlapping processes between AD and PD. Biological processess related to altered cytoskeleton dynamics, impaired mitochondrial function, and apoptotic signaling are enriched across gene modules from neurons in both AD and PD. Immune response, synapse maintenance, and lipid transport-related terms are enriched in one or more glial cell modules in both AD and PD.

GO terms mapped to gene modules are shared between AD and PD. GSEA results for the derived gene modules. For each gene module derived from ROSMAP-AD or Kamath-PD, we mapped the ranked genes (based on predictive weights) to terms in the GO database. (A) Overlapping terms in GO biological process (BP), cellular component (CC), and molecular function (MF) databases. The Venn diagrams depict the number of unique and shared terms across all gene modules, grouped by AD or PD. (B) Top 30 most frequently shared GO terms in AD-PD gene-module combinations. Solid black dots indicate that a term (vertical axis) is enriched in the corresponding gene-module pair (horizontal axis). Bar plots on the horizontal axes show counts of the total number of shared terms for the gene-module pair. Bar plots on the vertical axis show the total number of cross-disease gene-module pairs in which a term is present in (top 10 terms). (C) Graph visualization of top shared GO biological themes across AD and PD. Nodes are GO terms, and colors represent disease label. Node size indicates the GO-term gene-set size. Group names summarize the main themes from the terms in the group. Top AD-PD shared genes (robust PLS loadings) are annotated within each group. This overview highlights key BPs involved in both AD and PD. (D) Shared GO BP terms between AD and PD that are unique to either neurons or glia. The inner circle denotes the cell type group from AD, while the outer circle denotes the PD group. Darker shades represent terms enriched exclusively in neuronal modules (excitatory and inhibitory neurons, CALB1, SOX6), and lighter shades represent terms enriched exclusively in glial cell modules (microglia, astrocyte, oligodendrocyte, OPC, endothelial cells). BPs related to altered cytoskeleton dynamics, impaired mitochondrial function, and apoptotic signaling are enriched across gene modules from neurons in both AD and PD. Immune response, synapse maintenance, and lipid transport-related terms are enriched in one or more glial cell modules in both AD and PD.

By sorting all BPs based on their frequency of shared occurrence across AD-PD module pairs, we identified the commonly shared pathways. These were related to protein translation, cellular respiration, and mitochondrial energy synthesis (Fig. 4C). To further summarize the terms systematically, we devised a visualization procedure to obtain a synoptic summary of the overarching BPs (see the “Methods” section). Key shared biological themes emerged between AD and PD (Fig. 4D): protein synthesis and misfolding (top shared genes included RPLP1, RPL7A, RPL14, EIF1B, RPS8, SRP14), immune response (TMSB4X, CD59, RPL13A, GSTP1, FAU, APLP2, CALM3, RRAGA, GAPDH, TUBB2A/4B, RPS19, RAC1, DDT, CALM1, CFL1, CLU), lysosome acidification (ATP6V/6A, LAMTOR1, SNX3), glucose metabolism (ENO1, ENO2, MDH1, MDH2, PGAM1, PGK1, TPI1), mitochondrial dysfunction (COX6C/4I1/7A2/7C/5B, NDUFB9/A1/A4/B2/B11/S5, UQCRQ/10/B, CYCS), and myelination (SOD1, CD9, PTN, CXCR4, KLK6, CLU, PMP22, GPM6B, PLLP).

To investigate cell type localized BPs, we designed a probe to filter out general cellular injury-related pathways. We first categorized the gene modules into 2 groups—neuronal modules and glial modules. Given the fundamental anatomical and functional differences between neurons and glia, these groups can be expected to exhibit distinct responses to disease. The neuronal group included gene modules from excitatory neurons, inhibitory neurons, CALB1, and SOX6 cell types. The glial group comprised astrocytes, microglia, oligodendrocytes, and OPCs. We then removed all GO BP terms that were enriched in both neuronal and glial modules. This resulted in a refined set of GO terms that were exclusive to either neurons or glia (Fig. 4E; cf. Fig. S5B for gene module level grouping). By comparing these terms between AD and PD, we identified neuron and glia specific shared mechanisms of overlap.

Neurons shared the greatest number of exclusive (unique to cell types belonging to this category) GO BP terms between AD and PD. We identified several shared terms associated with microtubule depolymerization and cytoskeleton dynamics between Kamath-PD and ROSMAP-AD neuron modules (PD Calb1 1, Ex 1, Ex 2 and AD In 1, In 2, and Ex 1). Shared genes associated with these terms included MAPT, FKBP4, GBA2, MAP1A, MAP1B, MAP1S, MAP2, MAPRE3, STMN1, STMN2, STMN3, and STMN4. We also observed terms related to mitochondrial release of cytochrome c regulation (PINK1, PRELID1, CLU, BNIP3, DNM1L, GHITM, GPX1, MFF, MLLT11, MOAP1). Specifically, Ex 1 and Ex 2 in Kamath-PD, and In 1 and In 2 in ROSMAP-AD highlighted these terms. Terms related to iron homeostasis were noted in Ex 1 from PD and In 1 and Ex 1 from AD (SOD1, several ATP genes, CCDC115, FTH1, FTL, ISCU, NDFIP1, and SLC22A17).

We observed that glia-exclusive terms had several themes centered around the immune and complement systems along with synapse pruning, lipid transport, metal ion homeostasis, and thyroid hormone (TH) balance. T cell activation pathways were enriched in PD Mic 2, End 2, and AD Mic 2. These modules shared genes including B2M, HLA-A, HLA-B, HLA-C, HLA-DPA1, HLA-DRA, HLA-DRB1, HLA-DRB5, and HLA-E. Lipid and phospholipid transport showed up exclusively in microglial modules in AD Mic 1 and PD Mic 2 (APOE, TSPO, and PRELID1). Response to TH was recorded in PD Mic 2 and AD Mic 1 (CTSB, CTSH).

We also noted a few terms exclusive to opposite categories in AD and PD. For example, “response to copper ion” was identified exclusively in AD glia modules and PD neuron modules. However, functionally related terms, like cellular response to copper ion, copper ion binding, and copper ion homeostasis, were enriched in AD Ast 1, Opc 1, Ex 1, and Ex 2 and in PD Mic 1, End 1, Ex 1, Ex 2, and In 1. A closer inspection of all terms belonging to cross-category modules, AD neuron-PD glia or AD glia-PD neuron suggested that these differences largely reflect the granularity and naming conventions of GO annotations. In contrast, for AD neuron-PD neuron and AD glia-PD glia specific terms were functionally distinct, reflecting meaningful biological differences rather than annotation-related effects. For instance, a manual search for “cytoskeleton” or “microtubule” highlighted only neuronal modules in both AD and PD. These neuron and glia specific overlaps were further supported by replication in independent datasets (Seattle-AD and Smajić-PD).

For validation of these results, we turned to our external dataset pair. Analogous to our first analysis pair, we applied the GSEA pipeline to the gene modules derived from the Seattle-AD and Smajić-PD datasets (Fig. S6A). We confirmed that AD and PD neurons, oligodendrocytes, and OPC modules had the largest occurrence of shared terms (Fig. S6C; Table S9). Moreover, these results aligned with the broad biological themes highlighted from the primary analysis (Fig. S6B)—the most frequent AD and PD shared terms were related to mitochondrial energy metabolism (CD44, HSPA1A, CLU), myelination (SOD1, TENM4, PLP1), and glucose metabolism (APOD, RORA, HSPA5).

Thus, our GSEA successfully identified a variety of matching biological, cellular, and molecular processes across AD and PD. Our observations highlighted key themes that localized to certain cell types in AD and PD. Neurons in AD and PD demonstrated enrichment of terms related to cytoskeleton structural integrity, mitochondrial transport, and mitochondrial energy synthesis. Alternatively, glia-derived gene modules highlighted terms related to several regulatory mechanisms, including synapse pruning, lipid transport, immune, and inflammatory systems.

Latent gene modules reveal higher cross-disease transcriptomic convergence than traditional differential gene expression

We compared our gene module-based comparative framework to the widely adopted univariate method in RNA-seq—differential gene expression analysis. This method quantifies differences in gene expression between 2 groups by comparing expression profiles in diseased versus neurotypical cell states. We conducted a parallel AD-PD overlap analysis based on differentially expressed genes (DGEs; MAST; the “Methods” section) derived independently in the ROSMAP-AD (adDGEs) and Kamath-PD (pdDGEs) datasets. Within each dataset, DGEs were computed separately for each cell type (analogous to our main analysis; the “Methods” section). Statistically significant DEGs (FDR-corrected P-value < 0.05; the “Methods” section) were subject to further comparison between AD and PD cell type pairs.

We examined the overlap between the adDEGs and pdDEGs across AD and PD cell types (Kendall’s tau-b associations; Fig. 5A; the “Methods” section). Across 54 pairwise comparisons (6 AD and 9 PD cell types), the highest significant association observed was 0.4, occurring between microglia adDEGs and pdDEGs (FDR q-value < 0.05; P-values corrected for multiple comparisons). Notably, this highest association among all possible cell type pairings was significantly lower than the maximum similarity observed from our gene module analysis (cf. Fig. 2A).

Figure 5.

Graphs quantifying overlap between differentially expressed genes in Alzheimer's and Parkinson's disease. This widely used univariate method identifies fewer cross-disease overlap when compared to the gene module-level overlap identified in this study.

Differential expression analysis revealed modest similarity between AD- and PD-associated transcriptomic signatures. We benchmarked our latent factor model-derived AD-PD overlap with differential gene expression-derived AD-PD overlap. (A) Pairwise associations between AD and PD DGEs are shown (Kendall’s tau-b). For each cell type pair, statistical significance of association was assessed using a permutation test. Colored squares indicate significant associations (FDR < 0.05). Increasing intensity denotes greater similarity (+1) or anti-similarity (-1) between the log-fold changes of significant DEGs from an AD-PD cell type pair. Square size is proportional to −log10(FDR). Compared to our PLS gene-module-based overlap analysis (Fig. 2A), significantly smaller correlations emerge from this univariate approach. The maximum correlation observed was 0.4 between AD and PD microglia. (B) Dots represent shared GO terms between AD and PD from gene set enrichment of DGEs. Significant terms in AD or PD from GSEA were assessed (FDR q < 0.1). In contrast to 213 shared terms across pairwise AD-PD gene modules (Fig. 4A), only 7 shared terms emerged between AD and PD from our DGE analysis. Ast, astrocyte; Ex, excitatory neuron; In, inhibitory neuron; Mic, microglia; Oli, oligodendrocyte; OPC, oligodendrocyte precursor cell; End, endothelial.

We then systematically tested the overall difference in mean correlations between AD and PD associations based on pairwise gene module τb from PLS (12 × 20) versus pairwise cell type τb from DGE (6 × 9). Using Welch’s t-test, which accounts for unequal sample sizes, we observed a significant difference in correlation strengths between the 2 methods. Across pairwise comparisons, PLScell τb values were higher than DGE-derived τb values (Welch’s t = 5.51; P-value < 0.001). This finding suggested that associations between AD and PD similarity were systematically stronger when using our gene module approach compared to the classical DGE method.

Next, we performed a GSEA of the DGEs using GO databases (GO BP, MF, and CC). Independently for each disease, we created our ranked gene list based on the fold change significance level (conditioned on cell type; the “Methods” section) and used GSEA to identify significantly enriched terms (FDR q < 0.1). We observed 7 common terms, in total, between AD and PD across all 3 GO databases (Fig. 2B). Overall, these terms represented only a small subset of the broader set of shared AD-PD terms identified through our gene-module-based analyses (213 GO terms). Notably, the shared terms emerging from DGE analysis (e.g., cytoplasmic translation) appeared among the most frequently recurring terms across gene modules from the PLS analysis (cf. Fig. 4C), suggesting that DGE may primarily capture the strongest disease overlaps from the gene expression matrices.

We reassessed AD-PD overlap with DGEs from Seattle-AD and Smajić-PD dataset pair (Fig. S7). Across all pairwise combinations of relevant AD-PD cell types, the maximum observed correlation was between DEGs and between endothelial cells (−0.5; false discovery rate (FDR) q-value < 0.05) and microglia (0.3; FDR q-value < 0.05). In addition, a GSEA comparative analysis between AD and PD revealed extremely low number of shared terms (2 terms; Fig. S7B). Taken together, these results suggested that the univariate DEGs identified limited molecular overlap between AD and PD.

In summary, these findings demonstrated that DGE captured only modest overlaps between the molecular signatures of AD and PD brains. In contrast, our supervised latent factor modeling approach not only recapitulated the strongest DGE-derived AD-PD molecular overlaps (cross-disorder microglia), but also revealed a substantially broader set of shared, disease-relevant signatures. This comparative analysis highlighted the added value of identifying cross-disease associations through a multivariate, transcriptome-wide modeling of gene modules, rather than relying solely on individual gene-level differences.

GWAS-seeded co-expression networks also indicate AD-PD transcriptomic overlap

In an alternative set of analyses, we pursued the same research question, the extent of AD-PD molecular overlap, using a complementary quantitative workflow. Devising a top-down framework (Fig. 6A) seeded with 164 genes mapped from GWAS risk loci (GWAS in AD or PD), we constructed disease-specific differential gene co-expression networks (DGCNs) using the ROSMAP-AD and Kamath-PD datasets. The notion of gene co-expression networks (GCN) rests on the assumption that genes with co-varying expression profiles across cell transcriptomes often share functional or regulatory relationships [48–50].

Figure 6.

Workflow diagram and graphs depicting Alzheimer's and Parkinson's disease overlap using a top-down approach of identifying differential gene co-expression networks. Oligodendrocytes from AD and PD have similar disease relevant co-expression networks seeded with GWAS risk loci implicated genes.

DGCN identifies hints of overlap between AD and PD. To corroborate our findings in a technical replicate, we contrasted DGCNs between AD and PD. We constructed DGCNs based on genes whose expression patterns were co-expressed with seed genes (164 GWAS-nominated genes from AD or PD), but whose synchrony was altered in disease compared with control conditions. (A) Overview of workflow. (i) Single-nucleus RNA seq datasets were downloaded from open-source repositories. (ii) Local processing was performed independently for each cell type, following recommended guidelines. (iii) GCN analysis. For each gene from a set of GWAS-implicated genes from AD and PD, a co-expression network was created for the disease and control groups independently. The results were subtracted to generate the DGCNs. (iv) The disease-specific DGCNs were compared between AD and PD to identify cross-disease cell type associations. (B) Hierarchical clustering of cell types based on the similarity of the co-expression patterns captured by DGCN. The dendrogram captures the average Euclidean distance of DGCNs across all cell type pairs between AD and PD. (C) Kendall’s tau-b was used to calculate the pairwise associations between DGCNs. Cell type pair loadings for the top 3 principal components are shown. The pairs are presented in descending order, based on the loading magnitude. Dots represent the empirical PCA loading. Error bars represent the 20/80% CI from a 1,000-iteration BS analysis, which sampled rows from a DGCN with replacement. Pairs that did not include zero in their BS interval are shown. The first and second principal components are dominated by cell type pairings involving mainly excitatory and inhibitory neurons in AD and PD. The second component emphasizes PD dopaminergic neurons. The third component focuses on glial and vascular cell types from PD across cell types from AD.

We computed DGCNs independently for the ROSMAP-AD and Kamath-PD datasets and separately for each cell type within a dataset (the “Methods” section). By quantifying the alignment between our 2 analysis arms, we measured the correspondence between the relevant genes from DGCN (top-down) and the PLS-derived gene modules (bottom-up) (see the “Methods” section). Specifically, we looked at the number of robust genes that were common in each gene module—DGCN pair (Fig. S8A). We observed strong PLScell module—DGCN alignment within the same cell type pairs. For example, co-expression signatures from astrocytes shared, on average, 20% of genes (significant ρ at FDR < 0.01) with PLScell ROSMAP-AD Ast 2. Similarly, excitatory neurons shared, on average, 58.4% of genes (significant ρ at FDR < 0.01) with PLScell ROSMAP-AD Ex 1 and 42.2% of genes with PLScell ROSMAP-AD Ex 1.

We quantified the overlap between AD and PD at the DGCN level using a hierarchical clustering analysis (the “Methods” section). This grouped cell types from AD and PD based on their gene co-variation patterns centered on GWAS genes (Fig. 6B). Importantly, as in our primary analysis arm, we did not combine the gene expression matrices; instead, these DGCNs were derived separately for AD and PD. As an overarching observation, this secondary analytical approach recapitulated the main observation from the gene-module-level findings. Specifically, AD-oligodendrocytes and PD-oligodendrocytes were grouped together, suggesting similar co-expression profile alterations in AD and PD.

We then analyzed the extent of AD and PD overlap in yet another way. Using Kendall’s tau-b, we computed the similarity between the same seed-derived DGCNs across AD-PD cell type pairs (see the “Methods” section). Embeddings derived from this similarity matrix (164 GWAS genes × 54 AD-PD cell type pairs; PCA; the “Methods” section) revealed systematic AD-PD cell pairings that shared similar co-expression profiles (Fig. 6C). The first principal component was dominated by robust AD-PD neuronal cell types (25.6% explained variance of co-expression patterns; subjected to BS robustness check; see the “Methods” section). The second component grouped DA neurons in PD (CALB1 and SOX6) with excitatory and inhibitory neurons from AD (14.4% of total variance). The third component was dominated by combinations of microglia in AD and oligodendrocytes and OPCs in PD (9.5% of total variance).

Taken together, these results from our DGCN analysis recapitulated the gene module-based observations, confirming the strongest AD-PD overlap between oligodendrocytes and specific neurons. Thus, this parallel analysis reinforced the robustness of our earlier results through a technical replication.

Discussion

In this study, we examined the molecular ties between AD and PD using single-cell transcriptomics data. By analyzing the entire protein-coding transcriptome, our multivariate approach uncovered AD- and PD-deviant genes forming co-expressed modules at sub-cell type granularity. Our mappings of these gene modules to disease-relevant biological programs illuminated complex, cell type-specific mechanisms that might lead to shared disease phenotypes in AD and PD. On a broader level, we provide single-cell genomics scientists with a tool to compare any pair of diseases from a global transcriptome perspective.

Our primary analytical protocol was enabled by access to the ROSMAP-AD dataset with 70,634 nuclei from 8 major cell types, and the Kamath-PD dataset with 340,902 recorded nuclei from 11 major cell types. In both datasets, we extracted multiple gene modules within each cell type. These gene modules signified distinct modes of disease-related transcriptional changes at the sub-cell type level. Our grading of the alignment of gene importance between pairwise gene modules from AD and PD demonstrated sizable overlaps. In the spotlight were modules derived from AD and PD oligodendrocytes, astrocytes, and microglia. Significant, albeit more limited, overlap also emerged between AD and PD neuron-derived gene programs.

Similarly high degree of mirrored transcriptional changes between AD and PD were also highlighted in a secondary analysis arm which compared GCNs between the 2 diseases. Additionally, we were able to replicate these findings in an external AD-PD snRNA-seq dataset pair (Seattle-AD, Smajić-PD). Overall, our multivariate transcriptomic analyses offered rigorous and systematic evidence of shared molecular signatures between AD and PD.

Our investigations of GWAS gene hits situated several of these genes within our transcriptome-derived gene modules. For example, APOE, a major AD risk gene, emerged in AD modules and SNCA, a PD risk gene, emerged in PD modules [36, 37, 54].

Central to our investigation, several AD-relevant GWAS genes surfaced in PD-associated gene modules and vice versa, highlighting the interconnectedness between the 2 disease categories. For example, APOE had robust PD associations through several PD gene modules, including those pulled from microglia, astrocytes, oligodendrocytes, and neurons. Indeed, prior clinical studies have highlighted APOE to be predictive of cognitive decline in PD patients [51, 52]. Additionally, the APOE genotype has been shown to exacerbate synuclein pathology in transgenic mouse models [53]. As another example, we considered SNCA. In the PD brain, misfolded SNCA protein is a primary neuropathological marker [55]. Here, along with several PD neuron and glial modules, SNCA emerged as a strong contributor in AD excitatory neuron modules. Previously, APP transgenic mice with SNCA knockout have demonstrated a significant reduction in amyloid burden, hinting at connections between this PD gene and AD pathology [56, 57]. Taken together, this study situated GWAS AD or PD genes within the fuller context of disease-specific gene modules, further expanding the implication of these genes to potential neurodegeneration [12].

Drawing biological insights from pre-curated gene ontologies, we identified overlapping disease mechanisms between AD and PD. Our gene modules were related to shared alterations in cellular energy metabolism and stress response, inflammation, lipid signaling, protein folding, and protein degradation cascades. Some of these overarching disruptions in BPs have been discussed in previous reviews summarizing the collective understanding from decades of neurodegeneration research [57].

For example, a shared feature of neurodegeneration is the presence of abnormal protein aggregates [58, 59]. Consistent with this, our enrichment analysis revealed that protein misfolding (ER stress) and its associated BPs were widespread across all major cell types, diseases, and datasets. We also detected extensive alterations in molecular pathways related to protein degradation (ubiquitin protein ligase binding and clathrin-mediated endocytosis), indicating a potential breakdown in protein disposal systems across multiple diseased cell types. These results align with findings from animal models where malfunctions in ubiquitin-dependent protein clearance are linked to general neurodegeneration [60–62]. Overall, using a clean bottom-up computational workflow, this study confirms several prior findings, while going beyond them to carefully localize disease-relevant molecular processes to cell-type-specific gene modules.

Impaired molecular mechanisms shared between AD and PD in neurons

Neurons are particularly sensitive to proteasomal turnover due to their longevity and delicate synaptic regulatory requirements [62]. Zooming in on neuronal gene modules, we identified neuron exclusive mechanisms that were shared between AD and PD. Microtubule-associated processes localized almost uniquely to neuronal modules in both AD and PD (genes included APP, NEFL, TUBA1B, GAPDH, TUBB2A, CALM3, and MAPT). Several lines of prior evidence, including microscopy and genetic studies, have cited defects in cytoskeleton dynamics as major contributors to neuronal death [63]. Further, alterations in post-translational modifications of microtubule acetylation levels have been reported in in vitro and in vivo studies of both AD and PD [64, 65].

In addition, our neuronal modules were also enriched for terms related to alterations in mitochondrial functions, including axonal transport of mitochondria, and protein localization to mitochondria (all 4 datasets). This alludes to a vicious cycle between dysregulated microtubules and impaired mitochondrial transport [66]. Previous research has shown that alterations to these dynamics lead to the overproduction of mitochondrial reactive oxidative species (ROS) [67, 68]. ROS, in turn, exacerbates the levels of free tubulin [69], which in turn has been shown to interact with proteins like α-synuclein, promoting the formation of oligomeric aggregates in the form of Lewy bodies [70], or tau tangles [68].

The involvement of the MAPT gene in disease-relevant neuron modules from both AD and PD was also noteworthy. This gene encodes the protein tau and is responsible for stabilizing axon microtubules. Disruptions and alterations in this gene have been previously associated with multiple tauopathies and PD. Tau, in its aggregated form, demonstrates prion-like behavior, passing from neuron to neuron across synapses, a mechanism increasingly recognized in both AD and PD [71–75]. In our analysis, the emergence of MAPT across both disorders highlights a shared therapeutic target, motivating cross-disease strategies aimed at limiting pathological protein spread and neuronal death.

Apoptosis and cytochrome c regulation pathways

The emergence of terms related to the regulation of cytochrome c release and the apoptotic signaling pathway in several AD and PD neuronal gene modules was notable. In an intact human cell, cytochrome c is present in the mitochondrial intramembranous space. However, oligomeric forms of amyloid-β, α-synuclein, and tau have been shown to increase mitochondrial membrane permeability, causing leakage of cytochrome c into the cell cytosol [76]. This loss disrupts cytochrome c oxidase function and compromises ATP production. Once in the cytosol, cytochrome c initiates mitochondria-mediated apoptosis [77], a form of neuronal death long suspected as an early event in the pathophysiological cascade leading up to AD [78] and PD [79, 80]. Specific to PD, in a self-reinforcing manner, impaired cytochrome c release from mitochondria is thought to escalate α-synuclein oligomerization via radical formation [81]. Whether similar mechanisms exacerbate plaque and tangle aggregation in AD remains an important avenue for further investigation.

Oligodendrocyte and oligodendrocyte precursor cell modules in AD and PD

In oligodendrocytes and OPCs, we located gene modules with high similarity between AD-PD. Three key observations stand out related to these modules.

First, we observed that a sizeable portion of GWAS genes localized to cross-disease oligodendrocyte modules. Notably, PD oligodendrocyte modules contained more GWAS genes than modules from other cell types. Previous transcriptomic studies have pointed out that several disease risk loci are associated with oligodendrocytes in both AD [82] and PD [83]. Our study confirms and expands on this observation.

Second, in our enrichment analysis, several BPs revolving around myelination and regulation of axonogenesis were specific to these cell type modules and were shared between AD and PD. Prior correlative macroscopic brain-imaging studies have linked deteriorating myelin health to AD progression [84]. Further, a recent invasive experiment posited a causal link between aging myelin and AD [85]. In a few AD mouse models and a human PD model, transcriptomic analysis identified changes in oligodendrocyte transcription specifically related to impaired myelination [35, 86, 87].

Finally, in addition to strong intra-cell-type cross-disorder module associations, the AD oligodendrocyte modules also exhibited a high degree of association with AD and PD excitatory neuron gene modules. Such similarity in transcriptional modifications between oligodendrocyte and excitatory neurons was previously reported in the ROSMAP-AD transcriptomic study, in terms of shared DEGs [36]. The extension of this inter-cell-type overlap to cross-disorder settings might suggest pervasive crosstalk between excitatory neurons and oligodendrocytes in neurodegeneration.

Shared role of heavy metals is highlighted between AD and PD

In our study, gene modules from astrocytes showed exclusive enrichment for response to zinc ions in both AD and PD. Several metallothionein (MT) genes were common between these AD and PD modules (MT1E, MT1G, MT2A). Astrocytes, which remove excess heavy metals from the brain parenchyma [90], are susceptible to A1-type reactive astrogliosis in the presence of excess zinc and promote synaptic degeneration in neurons [91]. In general, human and animal research has shown that dysregulated homeostasis of multiple heavy metal ions can lead to an increased risk of onset and progression of neurodegenerative diseases [88, 89].

Another metal ion, copper (Cu), was notably enriched across our AD and PD neuronal and glial modules. This is consistent with a previously implicated metal dyshomeostasis axis linked to ROS and protein misfolding [92–94]. Biophysical and biochemical experiments have underscored alterations in Cu ion levels that trigger misfolding of α-synuclein [95] and amyloid-β [96].

However, here, a key distinction emerged between the neuronal and glial copper-handling gene modules—glial modules included genes related to buffering/detoxification responses (via MT1E, MT2A, MT3, and APP), whereas excitatory neuron modules uniquely enriched copper-binding genes tied to oxidative stress and proteostasis (via critical antioxidative enzymes [SOD1, PARK], SNCA, and copper chaperone protein [ATOX1]). Prior in silico analyses of microarray data on brain tissue had reported a similar grouping of copper-handling genes (metallothionein group and the enzyme binding group) [97]. Therapeutically, these findings support cell type-targeted interventions; for example, enhancing glial copper-buffering capacity (e.g., boosting metallothionein pathways) to stabilize extracellular redox balance, while simultaneously protecting neurons with copper-modulating and antioxidant strategies (e.g., targeting SOD1- or ATOX1-lined pathways) to reduce ROS-driven proteotoxicity. Thus, our cell-type-specific module-based approach provides important clues for precise drug treatment design.

In parallel, a third metal ion, iron ion, homeostasis terms were found to be enriched in neuron gene modules from both AD and PD. Across all datasets, excitatory and inhibitory neuron modules were involved. In the brain, iron plays a key role in myelin synthesis, neurotransmitter production, and overall metabolism [98]. However, elevated levels of redox-active iron, often originating from degenerating mitochondria, accumulate in several neurodegenerative diseases as evidenced by biochemical, and transgenic animal studies [99–102].

Together, these findings position disruptions of heavy metal pathways as shared therapeutic targets in AD and PD.

Shared microglia-specific molecular changes in AD and PD

Functionally, microglia are the primary immune cells of the CNS [103], and neuroinflammation and immune system dysfunction are believed to be key components of neurodegeneration [104]. Consistently, microglial gene modules from our study showed enrichment for several immune-related processes. For example, T cell activation terms were associated with microglial modules from both AD and PD (all 4 datasets). Microglia, upon activation by neuronal stress, are thought to release pro-inflammatory cytokines and upregulate MHC class I and II molecules [105]. Further, the inflammatory cytokines can induce the expression of adhesion molecules on brain endothelial cells, compromising the integrity of the blood–brain barrier (BBB). This BBB breakdown accelerates peripheral immune cell entry. Thus, in a chicken-egg scenario, microglia and endothelial cells drive a chain reaction of T cell activation, oxidative stress, and neuroinflammation [106–108]. This domino effect may have been captured in one of our PD endothelial modules whose pathways were associated with T cell activation terms as well as association with BPs like leukocyte adhesion to vascular cells, blood vessel morphogenesis, and diameter maintenance, pointing to the dysregulation of the BBB. This cascade of immune response events might exacerbate ROS production and neuronal damage [106, 107].

Further, we found a robust response to TH-related terms in microglial gene modules from both AD and PD. Within the brain, TH imbalances are an important contributing factor to increased ROS [109]. Epidemiological evidence has linked multiple thyroid-related autoimmune diseases to increased prevalence of both AD [110, 111] and PD [112]. However, the exact contribution of TH to either AD or PD pathophysiology has not been fully established. Recently, an AD mouse model has linked brain hypothyroidism with reduced microglial reactions to inflammatory stimuli and aberrant amyloid-β [113]. In PD, an α-synuclein PD mouse model found connections between TH and glucocerebrosidase activity in microglia [114]. Our findings, which highlight TH involvement in both AD and PD microglial gene modules, warrant further investigation of the thyroid axis.

We also identified microglial gene modules linked to synapse pruning in AD and PD. These modules implicated key genes, including TREM2 and several C1q genes from the complement system, an integral component for synaptic refinement. Synapse loss is frequently observed as an early event of neurodegeneration in animal models [115–117]. Variants of the TREM2 protein and aberrant activation of C1q genes have been shown to cause irregular synapse pruning in studies on AD mouse models [117–119]. In a PD mouse model, suppressing TREM2 gene products in microglial cells was shown to accelerate the loss of DA neurons [120, 121]. Yet, the totality of glia-synapse interactions, especially the role of the complement system in PD, is under-investigated [122, 123]. Once again, our bottom-up approach identified a common transcriptomic signature in AD and PD in the form of genes related to synapse loss—a potential area for further investigation.

In addition, lipid transport regulation terms were enriched in both AD and PD microglial gene modules (lipid, phospholipid, cholesterol, and sterol transfer terms) and involved key genes including APOE, TSPO, TMEM30A, several ATP-binding cassette subfamilies (ABC) genes (ABCG1, ABCA5), and NPC2. Excess lipid has been reported to accumulate as droplets in human iPSC-derived microglia, reducing their phagocytic capabilities and increasing the secretion of pro-inflammatory cytokines [124, 125]. These microglia mediated impairments ultimately exacerbate ROS burden and neurotoxic buildups leading to neurodegeneration.

In summary, our proof-of-principle investigation of the transcriptomic terrain intersecting AD and PD identified and characterized several shared molecular basis of neurodegeneration. Specifically, we were able to (i) quantify the degree of transcriptomic overlap between AD- and PD-related changes and (ii) localized these overlapping molecular changes to distinct cell types. In addition, the cell type-specific genes identified within convergent AD-PD–associated modules may represent potential therapeutic targets warranting further investigation.

Several limitations should be considered when interpreting our findings. While the present study focuses on the computational identification of disease-relevant gene programs and a subsequent comparison between AD and PD, future experimental work using cellular systems or animal models targeting the highlighted modules and shared pathways will be necessary to further investigate AD-PD convergence in vivo. For example, recent animal research on gut–brain axis disruptions in AD identified common therapeutic intervention strategies that might be applicable across neurodegenerative disorders [32, 126]. In addition, although the datasets analyzed in this study are among the largest currently available, larger and balanced cohorts will likely enable more robust estimation of gene modules, providing more granular insights into sex- and cell type-specific transcriptional basis of AD-PD overlap.

Future work can expand our study to include a greater diversity of neurodegenerative, neurodevelopmental, and psychiatric diseases. Moreover, applying a similar computational framework to different, readily accessible source transcriptomes, like blood or cerebrospinal fluid, can identify critical biomarkers for neurodegeneration.

Materials and methods

Single genomics data resources

Primary datasets

ROSMAP Alzheimer’s dataset [36]: The snRNA-seq dataset was derived from postmortem brain tissue from the prefrontal cortex (BA10) of individuals participating in the ROSMAP [127]. The dataset was collected from 48 subjects who were carefully matched in terms of age and sex (24 males and 24 females). Of these, 24 individuals were diagnosed with AD, and 24 were control subjects. The mean age was 85 years. The recorded cell types included excitatory neurons (n = 34,976), oligodendrocytes (n = 18,235), inhibitory neurons (n = 9,196), astrocytes (n = 3392), OPC (n = 2,627), microglia (n = 1,920), pericytes (n = 167), and endothelial cells (n = 121). The dataset encompassed transcript counts for 17,926 protein-coding genes, aligned with the human reference transcriptome hg38 (GRCh38.p5).

PD, Kamath dataset [40], GEO accession number GSE178265: This dataset included snRNA transcriptomes from postmortem human midbrain and dorsal striatum (caudate nucleus and substantia nigra) tissue. Tissue samples were derived from an age- and sex-matched cohort of 21 subjects in total (control, PD, and Lewy body disease). Samples derived from patients with Lewy body disease (4) were excluded from our analysis. To match the mean age of the ROSMAP-AD cohort, 3 patients <50 years of age were excluded. The final dataset consisted of 6 PD patients, while the remaining 7 were controls. Among this group, 6 were males, and 7 were females. The mean age of included individuals was 83 years. Major cell types in this dataset included oligodendrocyte (n = 134,940), excitatory neuron (n = 40,956), inhibitory neuron (n = 31,545), SOX6 (n = 25,482), astrocytes (n = 24,475), microglia (n = 24,038), CALB1 (n = 15,871), endothelial (n = 12,609), OPC (n = 9,603), macrophage (n = 955), and ependyma (n = 167). The dataset provided transcript counts for 33,692 genes, which were aligned to the hg19 genome.

External validation datasets

Seattle Alzheimer’s dataset [37]: The authors of this resource sourced brain specimens from the Adult Changes in Thought Study and the University of Washington’s Alzheimer’s Disease Research Center. Brain tissue samples were drawn from the middle temporal gyrus. The study included participants from age groups ranging from <65 years to more than 90 years. To maintain consistency with our other datasets, we considered only participants over the 65–77-year age bracket. This gave us 76 individuals, with 47 females and 29 males. Out of these, 36 subjects had recorded dementia, and 40 were controls. The mean age of the individuals was 88 years. The preprocessed datasets for the cell types labeled “L4 IT,” “L5 IT,” “Vip,” “Pvalb,” “Sst,” “Sncg,” “Oligodendrocyte,” “Microglia,” “Astrocyte,” and “OPC” were used. This resulted in the following nuclei being available for our analysis: sncg (n = 22,168), sst (n = 58,265), pvalb (n = 90,804), astro (n = 70,009), endo (n = 2,069), opc (n = 32,493), micro (n = 40,000), oligo (n = 111,194), l4_it (n = 168,860), l5_it (n = 128,090), and vip (n = 104,514). In total 36,517 genes were recorded and mapped to hg38 (GRCh38-2020-A) human reference genome.

PD, Smajić dataset [35], GEO accession number GSE157783: The creators of this resource worked with postmortem midbrain tissue sections that were linked with clinical and neuropathological data from the Parkinson’s UK Brain Bank and the Newcastle Brain Tissue Resource. The dataset included age and sex-matched nuclei samples from 6 controls and 5 idiopathic PD patients, all of whom exhibited severe neuronal loss in the substantia nigra and had no family history of the disease. The mean age of the subjects was 80 years. The major cell types were oligodendrocytes (n = 21,268), astrocytes (n = 4,708), microglia (n = 3,903), excitatory neurons (n = 3,037), OPC (n = 2,754), endothelial cells (n = 1,723), inhibitory neurons (n = 1,548), and pericytes (n = 1,229). Four cell types had fewer than 500 nuclei and were excluded from our analysis—ependymal cells (n = 536), GABA (n = 535), CADPS2+ neurons (n = 120), and DaNs (n = 74). The total number of genes was 24,005, and these were mapped to hg38 reference genome.

Control dataset

Lung disease, Kaminski dataset [128], GEO accession number GSE136831: The authors of this COPD dataset sourced nuclei from human distal lung parenchyma specimens. In total, 18 COPD patients and 28 control donor lungs were sampled. This resulted in the following nuclei being available for our analysis: myeloid (n = 174,146), multiplet (n = 3,765), lymphoid (n = 34,626), stromal (n = 6,607), endothelial (n = 2,069), epithelial (n =18,030). In total, 32,922 genes were recorded and mapped to hg38 (GRCh38) human reference.

Preprocessing pipeline at source

We relied on the preprocessed datasets from the authors responsible for the data collection (cf. above). This maximized reproducibility and compatibility with other studies working with these resources. The transcriptomic datasets were processed in the source studies using standardized snRNA-seq processing pipelines. These included quality control for cell inclusion, including doublet detection, the removal of low-quality and outlier cells, the removal of lowly expressed genes, and sample-level batch correction procedures.

Cell type classification—which cell belongs to which cell population—was taken from the original studies as a basis for our investigations. Exact details can be found in the Methods section of the independent research.

Primary local preprocessing

Different cell types perform widely diverse functions, each arising from the functional recruitment of distinct gene groups. In our pursuit to discover biologically meaningful gene groups, we implemented our quantitative analysis pipeline for a given disease on a cell type-by-cell type basis. This agenda enabled us to extract coherent latent components (latent factor/loading vector, hereby referred to as gene module) specific to each cell type.

To ensure an even sample size in the disease and control group, we randomly sub-sampled transcriptomes of a given cell type to have comparable counts of nuclei from both disease and control instances. On this sub-sampled data, we removed genes that were captured in fewer than 1 out of 1,000 cells to reduce our model’s degrees of freedom (scanpy.pp.filter_genes), which is in line with previous research [129, 130]. To reduce technical variation from sequencing depth, we normalized the data by dividing the raw UMI count by the total number of detected UMIs in each cell (scanpy.pp.normalize_total(data, target_sum=1e4)). To account for the heteroskedasticity originating from differences in highly expressed versus lowly expressed genes, we log-transformed the normalized gene expression data (scanpy.pp.log1p(data)). Taken together, these dataset transformations have been shown to work well as a preparatory step for downstream dimensionality reduction [131].

In an additional data cleaning step, postmortem interval (PMI) was included as a covariate to account for potential confounding effects on gene expression. PMI has been shown to be a potential source of confounding elsewhere [132, 133]. Specifically, here we regressed out the variation in gene expression attributable to differences in PMI. The resulting adjusted, cleaned, and standardized transcriptomic profiles for each examined cell were used for subsequent steps in our modeling pipeline.

We restricted the feature space to protein-coding genes in each dataset. This step reduced the ambient gene space dimensionality, which is beneficial for any high-dimensional analysis [134, 135]. Specifically, for each dataset, we considered the genes that overlapped with the set of 17,019 protein-coding genes in ROSMAP-AD. This filtering resulted in 16,936 genes in the Kamath-PD dataset, 12,681 genes in Seattle-AD, and 15,137 genes in Smajić-PD. As a result, the analyzed datasets were embedded in ambient feature space of comparable dimensionality, with a similar set and order of genes, with broadly consistent biological properties being evaluated across analysis.

Crucially, the transcriptomic datasets from different studies were at no point merged or jointly integrated for our analyses. Instead, each dataset was analyzed independently to derive disease-associated gene modules. Within each dataset, model fitting and hyperparameter optimization were performed separately. Thus, by avoiding cross-dataset integration, our analysis circumvented the need for cross-dataset batch-effect corrections. In summary, comparisons between AD and PD were performed at the level of model-derived gene loadings, and not experimentally recorded gene expression.

Identifying gene modules: supervised latent factor modeling

At the heart of this study, we sought to identify synchronous gene expression changes that occur in the brain in association with disease state, and compared them between AD and PD. To achieve this, we employed a latent factor approach to identify gene programs in a multivariate framework, rather than examining individual genes independently. Unsupervised latent factor models have bee previously used in snRNA-seq analyses to identify hidden gene programs [136–138]. In contrast, our implementation uses a supervised approach that identifies structured patterns in the high-dimensional gene space while simultaneously modeling the relationship between gene expression and the associated disease status.

This work builds on a prior study in which the multivariate method, PLS-DA [139], was employed to derive AD-predictive gene modules across major brain cell types [42]. In contrast to the original study, where the analyzed dataset had favorable observations-to-features ratio, several datasets analyzed here comprised substantially fewer samples within individual cell types. In these cases, PLS-DA is susceptible to overfitting, particularly given the inherent noise and sparsity of snRNA-seq data [140]. To address this limitation, we applied principal component analysis (PCA) denoising on the input gene expression matrix, prior to model fitting. Under a low-rank assumption on the input feature space (gene expression), dimensionality reduction with PCA yields an optimal low-rank approximation that is robust to noise and sparsity [27, 141]. In practice, such transformations are routinely applied in single-cell and single-nucleus analysis, where the ambient gene space is well approximated by a low-rank structure [142].

Specifically, for each cell type within individual datasets, we applied PCA to the input gene expression matrix. These components were then used as input variables for PLS-DA to obtain a projection maximizing disease versus control class separation. To guard against spurious derived structures [143], we performed strict model validations (see below), including label-shuffled permutation tests to assess the statistical significance of derived components (Figs S1B and S3B).

Formally, let Inline graphic be the input gene expression matrix, and Inline graphic, where M is the number of genes, and Inline graphic the number of observations (nuclei) for a cell type c. Y represents the disease label (+1 for disease and −1 for control) for each nucleus. Let Inline graphic be the number of chosen PCA components. Then, the input datasets to the PLS-DA models can be denoted as Inline graphic.

Concretely, PLS-DA can be viewed to consist of 2 key equations:

graphic file with name TM0008.gif
graphic file with name TM0009.gif

where T and U are Inline graphic score matrices of kc extracted components, P is a Inline graphic loading matrix (effect size) of X, and Q is a Inline graphic loading vector of Y, respectively. E and F are the residual matrices of X and Y, respectively. The decomposition of X and Y is set to the solution of the optimization objective:

graphic file with name TM0025.gif

where Inline graphic is the captured covariance Inline graphic, w and h are weight vectors that are extracted using the NIPALS algorithm [144].

To transition from the PCA embedding space (Inline graphic) back to the gene space (M), we projected the PLS loadings back into the original gene expression space (approximated using sklearn PCA.inverse_transform). This ensured that the domain interpretability of the PLS estimates was preserved in the biological ambient space—high absolute gene loadings signaled strong contributions (positive to the target disease, negative to the control group), while near-zero loadings indicated minimal impact. To assess the statistical robustness of derived gene-level loadings, we focused on genes that were consistently selected across model refits based on BS resamples of the dataset (see below). This procedure allowed us to identify genes that were stable across different realizations of the data, further reducing the likelihood that the reported results were driven by sampling noise or unstable features.

Model selection, training, and performance assessment

After preprocessing and cleaning the transcriptomic data resources, we carried out the training of our supervised learning models. First, the number of PCA components was chosen using Inline graphic, where Inline graphic was the number of recorded nuclei for the given cell type c in the dataset. Next, the optimal number of latent components for each cell type-specific PLS model was determined using a rigorous 10-fold CV scheme. Selecting this optimal number of latent components was crucial—choosing too few components implied losing out on important information and too many could lead to overfitting. To do this hyperparameter selection, the set of cells was randomly split into 10 equal-sized data point subsets. We ensured that the disease-to-control ratio of cells for each subset reflected that of the full dataset. Screening a range of component choices (1–8) in each iteration, 9 out of these 10 data subsets were combined and used for training a PLS model, while the out-of-bag subset was used to assess the component number choice. The model’s performance was evaluated based on the area under the receiver operating characteristic curve (AUROC) in disease discrimination. This was performed for all combinations of training and validation subsets (scikit-learn model_selection.GridSearchCV function with PLSRegression as the “estimator,” “scoring” set to “roc_auc”, and “n_components” parameter set to 1–8). The number of components yielding the maximum mean AUROC over the CV subsets was noted as optimal for the given cell type.

For each cell type, we then fitted optimal PLS models on the full set of transcriptome observations (PLS model specified using sklearn.cross_decomposition module PLSRegression). To audit the performance of the individual PLS models, we employed AUROC of disease classification as our evaluation metric. Given that cell samples from a patient can exhibit significant autocorrelation, it was crucial to account for this when evaluating the model. Traditional test-train splitting methods often involve blindly partitioning the dataset. This can result in overly optimistic test performance and makes it challenging to detect overfitting during the testing phase [145]. In light of this, we employed a variation of CV combined with bootstrapped Latin partitions [146] that ensured patient-level stratification.

Concretely, in each iteration, a random sample of subjects (not cells) was drawn with replacement. This formed the basis of the training dataset’s nuclei source. Based on the subset of patients, a random sample of cells was drawn, with replacement, while ensuring that the disease-to-control ratio of nuclei was reflective of the empirical dataset. The percentage of cell samples from any given subject was also preserved in each iteration. This analytical protocol ensured that transcription signatures from the same patient were not present in the train and test sets at the same time. Individual PLS models were fitted on the re-sampled train dataset. This fitted model was then evaluated based on AUROC scores on the test set nuclei, i.e., the transcripts from subjects that were not included in the training step. We performed 1,000 iterations of this BS-based model disease classification performance evaluation. This allowed for a principled assessment of the disease discrimination strength of the PLS solutions based on the cell type-specific gene transcription signatures.

Statistical tests of PLS-derived gene modules and gene loadings

Our cell type-specific hyperparameter selection allowed us to independently derive the number of PLS-DA components (gene modules) that maximized disease-versus-control identification within each cell type. The statistical significance of an overall gene module was assessed in a principled, non-parametric permutation procedure. In 1,000 permutation iterations, the transcriptome signatures were held constant, while the disease labels (outcome of model) were shuffled randomly. The resulting surrogate datasets preserved the statistical structure of gene expression profiles while selectively destroying the association of the transcription profiles (model input) with diagnosis (model output). This approach generated a null distribution with minimal modeling assumptions [147, 148].

The empirical covariance (test statistic) between the gene expression and disease signature captured by each module (Inline graphic defined above; t=model.x_scores_; u=model.y_scores_) was compared with the resulting permutation distribution. This distribution reflected the null hypothesis of random association between gene transcription and the disease designation, against which the actual model instance was tested. We deemed significant a module’s input-output covariance in the latent space if fewer than 5% of the null models yielded a better covariance strength than the original covariance from the actual model instance (Fig. S1B). In case a module failed to pass this label-shuffling permutation test, it was dropped from further analysis, along with all underlying gene modules for that cell type. Thus, in a data-driven approach, we were able to determine which gene modules in a cell type at hand carried enough information that allowed us to discern a biological signal from noise.

To identify the subset of the examined genes that robustly contributed to disease detection in each gene module from a cell type, we implemented a 500-iteration BS scheme. The BS resampling was done by selecting nuclei, with replacement, from the cell type observations before applying dimensionality reduction. This approach simulated random nuclei sample draws that could have been derived from the broader cell population. Dimensionality reduction (PCA) and PLS model estimation were performed on the resampled BS dataset in an identical fashion (cf. above).

An inherent ambiguity of the class of latent factor models (i.e., aspects of model non-identifiability), including PCA and PLS, is the reflection invariance of derived latent vectors. To remedy this source of indeterminacy, we computed the cosine similarity (γ; range −1 to +1), between a BS loading vector and its corresponding empirical loading vector. In the scenario where γ was <0, indicating a flipped (“mirrored”) loading vector, we multiplied the loading vector of the BS model elementwise by −1 to align it with the original loading vector. This method has been employed by previous authors to address the issue of reflection [42].

The resulting distribution of loadings for a gene in a module was compared to its counterpart in the original model estimate. We disregarded any genes whose model coefficient (PLS loading corresponding to this gene) included zero in its 5/95% BS-CI for that module. That is, the gene effect was removed by setting the loading value to zero. The resulting “robust” gene modules are denoted as Inline graphic (dimension Inline graphic; cf. PLS definition section above) in future references.

To evaluate the potential influence of demographic confounding factors, specifically age and sex, we modeled the component scores for each latent gene module as a function of diagnosis, age, and sex. For each gene module, component scores were regressed on diagnosis (primary variable of interest) together with age and sex (potential confounders) using ordinary least squares regression from the Python statsmodel package (v0.14.4). Following model fitting, an analysis of variance was performed to quantify the proportion of variance in component scores (R2) that was explained by each predictor (anova_lm, typ=2; statsmodel v0.14.4). In total, 72 gene modules were analyzed (12 ROSMAP-AD, 20 Kamath-PD, 25 Seattle-AD, and 15 Smajić-PD). This step provided a unique estimate of the contribution by each predictor while accounting for other variables in the models.

PHATE visualization

To gain a synopsis of the uniqueness of different gene modules from a single cell type, we created a concise, low-dimensional representation of high-dimensional cellular transcriptomes. For this purpose, we applied dimensionality reduction using PHATE [149]. Compared to prevalent visualization techniques like tSNE or UMAP, PHATE is well suited to preserve both local and global structures in a dataset and can capture non-linear relationships in the transcriptomic information. We estimated a separate PHATE model for each cell type and projected their transcriptomes to independent low-rank spaces (using the scanpy external.tl.phate function with default parameters except n_pca = 500). We colored the cells in the PHATE embedding space based on their PLS score from each gene module (PLSRegression x_scores_).

Quantifying degree of associations between cross-disease latent gene modules

To quantify the similarity between AD and PD at the level of gene expression patterns, we computed the association between the AD and PD gene modules. Importantly, a similarity metric (correlation) was computed across the derived model’s predictive rules for a disease. That is, we did not pit the raw gene expression measurements against each other.

Formally, for a disease d and cell type c, the ith robust gene module (cf. above) was denoted as Inline graphic, where p* is a vector of dimension Inline graphic (number of genes in the input feature space). A robust gene module could contain tens to thousands of genes with non-zero loadings (out of ∼17,000 genes), while the remaining loadings were zero (cf. above). To minimize the impact of tied zero loadings in the correlation metric calculation between 2 modules, we considered only those genes that had non-zero weights in both modules (AND conjunction). This approach worked well to identify groups of genes with similar disease contributions between 2 modules while ignoring genes that had robust effect sizes in one disease but not the other.

Concretely, we employed Kendall’s tau-b ranked correlation metric (Inline graphic) to evaluate pairwise correlations between AD and PD gene modules. Kendall’s tau-b correlation effectively handled tied ranks and provided a more accurate measure of ordinal association between gene modules. This was unlike Pearson’s r, which assumes monotonicity and is unstable, or Spearman’s r, which is biased and difficult to interpret [150]. For all pairwise gene modules from AD (12 modules across all cell types) and PD (20 modules across all cell types), Kendall’s tau-b rank correlation coefficient, Inline graphic, was calculated using the scipy function stats.kendalltau(Inline graphic), where Inline graphic was the ith gene module for the cell type Inline graphic present in the AD dataset, Inline graphic was the jth gene module from the cell type Inline graphic present in the PD dataset.

To independently assess the statistical significance of each coupled module association, we employed a non-parametric permutation procedure with the null hypothesis of random association between gene modules from different diseases. For each gene module pair, we reutilized the label shuffling derived module weights (cf. above) to calculate a null distribution from 1,000 permutation iterations. We only interpreted a module-pair’s correlation coefficient that emerged as statistically significant against a 5/95% CI threshold.

In parallel, we reported the FDR corrected q-values for the computed pairwise Inline graphic, with multiple testing controlled using the Benjamini–Hochberg procedure [151] across all pairwise comparisons.

A split-half test was also conducted to estimate the statistical sensitivity of sample size to the coupled associations of cross-disease gene modules. Across 500 iterations, we randomly bisected the empirical AD and PD datasets (before cleaning and standardizing) into 2 pairs of AD-PD subsets. We ensured to preserve the original proportion of cell type nuclei and the disease-control ratio for each cell type. Since the bisection reduced the effective observation size of each dataset to half, the downstream PLS fit enabled a robustness check of sample size for derived gene modules. For each AD-PD subset pair, we ran our workflow steps A-D in parallel (Fig. 1), resulting in 2 analogous sets of gene modules (a couple of 12 AD and 20 PD modules). We then performed Kendall’s tau-b correlation across these modules, giving us 2 correlation matrices of dimension 12 × 20. We unraveled these matrices and calculated Pearson correlated (Inline graphic) of the absolute Inline graphic values. We used absolute values since we were interested only in association strength, not direction. This analysis allowed us to compare the Inline graphic correlation levels between different gene module pairs derived from smaller subsets of the AD and PD datasets.

Differential gene expression

Differential gene expression is ubiquitously used in snRNA-seq analysis to identify genes that show statistically significant differences in expression levels between 2 conditions or groups [152]. It is a univariate method, meaning it examines each gene individually, thus losing key information hidden in gene co-expression patterns. We employed this traditional method to serve as an acid test for the proposed approach in this study.

Differential gene expression was performed using MAST (v1.36.0), implemented in R and accessed from Python using rpy2 [153]. Within each dataset, log-normalized single-nucleus expression data were analyzed separately for each cell type. For each gene, a generalized linear hurdle model was fitted with diagnosis (disease versus control) as the primary variable of interest. Cellular detection rates were included to account for differences in gene detection across cells. Age, PMI, and sex were included as covariates to account for potential confounding effects.

Significance was assessed using likelihood ratio tests as implemented in MAST. Multiple testing correction was performed across all assessed genes using the Benjamini–Hochberg method [151]. Effect sizes were quantified using the estimated log2 fold change between disease and control groups. Genes meeting the statistical significance threshold after correction were designated as DGEs. The final DEGs were referred to as adDEGs for AD and pdDEGs for PD.

To estimate the pairwise association between cell-type-specific adDEGs and pdDEGs, we computed Kendall’s tau-b using the log-fold change values of genes in the AD-PD AND conjunction set (see above). This resulted in a similarity matrix of dimension 6 × 9, corresponding to the number of AD (6) and PD (9) cell types. Statistical significance for each association was evaluated using a permutation test with 1,000 iterations, in which the fold change values of overlapping genes were randomly shuffled. FDR correction was performed across all comparisons (54) using the Benjamini–Hochberg procedure.

Quantifying difference in association strengths between PLS- and DGE-derived conclusions: Welch’s t-test

To formally quantify the cross-disease association information extracted by our latent factor approach versus DGE, we used a statistical test to compare the central tendencies of the respective correlation measures. Welch’s t-test was a natural choice of method here as the number of correlated combinations being compared were different (240 Inline graphic gene module combinations from PLS and 54 Inline graphic cell type combinations from DGE), and the variances were not assumed to be equal [154, 155]. Welch’s t-test can be formally computed as

graphic file with name TM0060.gif

In our case, Inline graphic was the unraveled PLS correlation vector (1 × 240; cf. Fig. 2A), and Inline graphic was the unraveled DGE correlation vector (1 × 54; cf. Fig. 5). Inline graphic and Inline graphic were the variances of these vectors, Inline graphic was the total number of cross-disease gene module pairs from the latent factor analysis, and Inline graphic was the number of cross-disease cell type pairs considered in the DGE analysis.

Identifying biological signaling pathways from gene modules: GO enrichment analysis

To query the biological meaning of our gene modules, we performed a GSEA [28]. Here, we used the GSEApy Python package [156], which itself uses Enrichr [157]. GSEApy is designed to extract statistically over-represented gene sets (example pathways) from a ranked gene list encompassing the whole genome. We used the GO BP, MF, and CC [158, 159] databases as gene sets of interest. We focused on GO, as collectively, they cover the largest fraction of the genome. Concretely, for each gene module, we fed the PLS gene loadings for the entire transcriptome recorded in our datasets to the enrichment tool (gseapy.Prerank tool with parameters rnk = gene loadings, min_size = 15, max_size = 1,500, and permutation = 1,000 for significance testing). We reported the pathways that had an FDR threshold of at most 0.05. This step was repeated identically and independently for all gene modules from each cell type, across all datasets.

To verify that the gene set enrichment results were not an artifact of noise in the PLS modeling but rather had actual biological relevance, we turned to our permutation test-derived gene modules. In 1,000 permutation iterations, we destroyed the relation between gene expression and disease label. Thus, the extracted gene modules captured noise. We fed the gene loadings from these modules into our GSEA pipeline to verify the specificity of the empirically enriched terms.

Gene network visualization

GO terms are organized in the form of a hierarchical tree [160] which can be downloaded here (OBO 1.4). This hierarchical organization often results in hundreds of hits from an enrichment analysis. One of the techniques widely used to crunch down this dense information is via network visualization [161, 162]. This technique can help identify broad groupings of terms based on a chosen parameter of interest, for example, shared genes between different terms.

For illustration purposes, we utilized Cytoscape [163] to create a structured network of disease-relevant GO-BPs identified by an enrichment analysis (cf. above). Each node in the network represented a GO BP hit. The edges were formed based on predefined relationships between nodes conditioned on shared genes and whether they were part of the same regulatory network. The resulting network layout (yFiles.organic layout) automatically clustered the enriched terms into biologically meaningful groups, allowing us to identify major functional themes that were shared between AD and PD.

Differential GCN

As a complementary analytical pipeline, we sought to explore the transcription profile of the RNA-seq datasets in a top-down GCN analysis approach. Specifically, we started with candidate genes mapped to AD or PD GWAS risk loci. Using these genes as seeds, we created networks of correlated genes; that is, we identified groups of genes whose differential expression change between disease and control closely matched a seed gene. Seeded DGCNs have been previously used to identify regulatory changes in gene expressions across various conditions [48, 50]. By taking a contrastive approach between disease and control (differential), the effects of housekeeping genes could be limited, and the residual patterns of covariation could be attributed to the effects of a disease.

To identify our set of seed genes, we utilized the most recent GWAS studies that reported AD- or PD-associated risk genes significant at the whole genome level. We identified 108 GWAS hits associated with AD [44] and 129 GWAS hits associated with PD [45]. Out of these 237 genes, 5 were common (CTSB, WNT3, BCKDK, HLA-DQA1, and HLA-DRB1) between AD and PD, giving us 232 unique genes. Of these, a total of 164 genes (out of the 232 genes) were present across all considered transcriptomic datasets.

Next, we split each individual dataset into disease and control groups based on the diagnosis labels provided. We then subdivided each of these groups based on cell types. For each cell type Inline graphic, the co-expression vector for a single GWAS gene i with another gene j recorded in the snRNA-seq dataset was calculated as Inline graphic, where Inline graphic is the Kendall’s tau-b correlation metric, Inline graphic is the read count vector for Inline graphic observations (nuclei), and Inline graphic is the read count vector of Inline graphic observations (nuclei).

Evaluating Inline graphic across all recorded genes gave us Inline graphic, where M was the number of genes common to both AD and PD datasets (M = 16,936). Thus, each element of the matrix Inline graphic was a numerical value between −1 and 1, capturing the degree of correlation of gene j with GWAS gene i. Stacking the vectors for all GWAS genes gave us gene co-expression matrices Inline graphic and Inline graphic for the control and disease groups, respectively, where N was the number of GWAS genes (N = 164). From this, we formally computed the differential gene co-expression matrix for a single cell type Inline graphic as follows,

graphic file with name TM0086.gif

Next, we systematically explored the mutual relationships between the DGCNs across different cell types, without discriminating them based on disease. To this end, we employed a hierarchical clustering analysis. Our goal was to probe for clusters of cell types that featured similar genome-wide co-deviation of gene transcription.

Concretely, we unraveled Inline graphic into a vector Inline graphic, where Inline graphic and combined them across 6 AD and 9 PD cell types to get Inline graphic, where Inline graphic. We computed a linkage matrix based on the Euclidean distance between 2 unraveled differential co-expression vectors for each cell type (scipy.cluster.hierarchy.linkage, parameters method = “average,” metric = “euclidean”). The linkage algorithm hierarchically clustered the 15 cell types, across AD and PD, with the cluster groups indicating cell types with the closest co-expression patterns. We visualized these clusters as a dendrogram in Python (scipy.cluster.hierarchy.dendrogram).

We refined our clustering-based qualitative approach to rigorously quantify the association between cross-disease DGCNs. Toward this goal, for the ith GWAS gene, we computed Kendall’s tau-b correlation metric (Inline graphic) between Inline graphic and Inline graphic giving us differential co-expression correlation matrix Inline graphic. Each of these matrices encoded a similarity (−1 to +1; 0 being no association) between AD and PD co-expression networks for one gene.

To congregate this cross-association information encoded by 164 GWAS risk genes, we vertically stacked the unraveled matrix Inline graphic (unraveled to Inline graphic) into a matrix Inline graphic. Thus, P, in essence, captured multiple modes of information condensed into one matrix: (i) differential gene co-expression between disease and control, (ii) quantified similarity of the expression changes between AD and PD stratified at the level of cell types. Note that the GWAS genes acted as seeds not only for the disease in which they were identified but also for the other disease. We finally distilled P using PCA. This uncovered linear combinations of cross-disease cell type pairs that were most related in terms of their alterations to transcription in response to disease.

The number of significant latent factors that captured biologically meaningful information was determined using a principled permutation testing framework. In 100 permutation iterations, we randomly shuffled the unraveled correlation vector Inline graphic, individually for each i, thus breaking the inherent meaningful patterns of covariation across cell type pairs. Across these 100 iterations, we fitted individual PCA models and computed the explained variances of the derived components. After comparing the permutation variances with our empirical component variances, we retained 4 latent factors as statistically significant based on the 5/95% CI (Fig. S8B). These 4 embeddings were by construction uncorrelated and rank-ordered, with the first component capturing the highest amount of variance in P.

We conducted a BS analysis on the extracted latent embeddings to formally assess the robustness of the cell type pairs that emerged as being closely associated with each other. Across 1,000 BS iterations, we sampled different rows (encapsulating all pairwise cell type co-deviations for a GWAS gene) with replacement to simulate a random seed gene collection that could have been sampled from the empirical population. We fitted individual PCA models to each of the thus-derived samples.

To handle the inherent order invariance (changed sequence, especially for later components with small explained variance) and reflection invariance (sign flipping of derived singular vectors) of PCA components [164], we applied the Jonker–Volgenant algorithm for component matching and Pearson’s correlation (Inline graphic) for sign matching. The Jonker–Volgenant algorithm is a widely used technique [165] that can identify a one-to-one mapping of latent embeddings derived from 2 separate BS iterations. The similarity between a pair of components from 2 runs was scored using the cosine similarity (cf. above). Subsequently, we solved this optimization problem to maximize the similarity between component orderings from 2 runs across all pairwise combinations of the first 10 empirical and BS-derived PCA components (scipy.optimize.linear_sum_assignment, maximize = True). To align directionality, Inline graphic was computed between the empirical PCA component loadings and the BS-component loadings. For cases where Inline graphic was <1, the latent vector loadings were multiplied by −1.

Thus, in a complementary data-driven approach to our latent factor modeling, we identified potential combinations of cell types that had the closest associations of gene expression changes between disease and control states in AD and PD.

Additional files

Supplementary Figure S1: Gene modules derived by supervised latent factor modeling perform above chance in out-of-sample disease classification. (A) Unbiased classification performance of PLS models fitted on individual cell types in Seattle-AD and Smajić-PD snRNA-seq datasets. Box plots represent clustered bootstrap performance results (n = 1,000). Each dot is one bootstrap iteration, where a random sample of donors was chosen with replacement from the bag of all donors. Disease and control donors were chosen in equal proportion. For the subset of donors, nuclei were chosen with replacement. The model performance was evaluated on the held-out donor nuclei. (B) Disease predictive power of PLS components. (Left) Kamath-PD, (right) ROSMAP-AD. Component-disease alignment was quantified using Pearson’s ρ between component’s disease prediction (x-score) and true disease representation (y-score). The black dot represents the model’s empirical ρ. Violins depict null distributions of ρ generated from label permutation of the empirical dataset (n = 1,000). Dashed lines represent 2.5/97.5% CI. (C) Kendall’s tau-b (τb) correlation between ROSMAP-AD ROSMAP-AD gene modules (top) and Kamath-PD Kamath-PD gene modules (bottom). Each pairwise τb is statistically significant, exceeding the 2.5/97.5% confidence interval (CI) based on a 1,000-iteration permutation test. Darker colors indicate stronger association strengths, while larger square sizes indicate a greater number of shared genes. Strong correlation (absolute magnitude) is observed for the same cell type-derived gene modules. Significant cross cell type within the gene modules from the same disease correlations are also observed. (D) Stability and consistency of gene module correlations across AD and PD were tested in a split-half analysis. The initial datasets were bisected into 2 pairs of AD-PD sub-datasets. The histogram illustrates the distribution of Pearson’s rho obtained by comparing the overlap of cross-disease gene module pairs from split pairs (n = 500). This analysis underscored the robustness of our analytical pipeline, revealing a stable mean correlation coefficient of 0.92 ± 0.02. The mean is represented by a red dashed line.

Supplementary Figure S2: Gene modules do not necessarily represent cell subtypes. PHATE visualization of transcriptomes for all nuclei per each cell type. Each nucleus (dot) is colored based on the component with the highest disease-predictive score. No distinct separation of cells was observed, highlighting that the gene modules captured different modes of transcription changes that were pervasive across cells of a given type. AD- represents PHATE plots of ROSMAP-AD cell types. PD- represent PHATE plots of Kamath-PD cell types.

Supplementary Figure S3: AD-PD overlap replicated in an independent snRNA-seq dataset pair. Independent disease predicting PLS models were fitted for each Seattle-AD and Smajić-PD cell type. (A) Kendall’s tau-b (τb) correlation between Seattle-AD and Smajić-PD gene modules. Darker colors indicate stronger association strengths, while square size indicates statistical significance (FDR corrected P-values). Black boxes represent statistically significant τb (exceeding 2.5/97.5% confidence interval) derived from permutation models fitted on the original datasets (n = 1,000). (B) Disease predictive power of PLS components. (Red) Smajić-PD cell type, (blue) Seattle-AD cell type. Component-disease alignment was quantified using Pearson’s ρ between component’s disease prediction (x-score) and true disease representation (y-score). The black dot represents the model’s empirical ρ. Violins depict null distributions of ρ generated from label permutation of the empirical dataset (n = 1,000). Dashed lines represent 2.5/97.5% CI.

Supplementary Figure S4: Situating AD and PD GWAS risk genes within AD and PD gene modules from an independent snRNA-seq dataset pair. (A–C) Mapping 164 GWAS genes (from both AD and PD GWAS) within Seattle-AD and Smajić-PD gene modules. Color intensity reflects the loading magnitude of each gene within the module. Only robust genes passing the 2.5/97.5% CI in a bootstrap permutation test are displayed. (D) Percentage of GWAS genes with robust disease-predictive weights across all gene modules (blue = AD gene modules, red = PD gene modules). Cellular localization of genes was observed; for example, APOE showed strong predictive loading in astrocytes, microglia, and oligodendrocyte precursor cell modules in AD.

Supplementary Figure S5: Gene Ontology terms shared between AD PD across cell types. (A) Number of shared GO terms between ROSMAP-AD and Kamath-PD gene module pairs is shown. Brighter colors represent a higher number of shared terms. White grids represent zero overlapping terms. Neuron gene modules in both AD and PD had the highest number of shared terms, closely followed by oligodendrocyte-related module combinations from PD. (B) Shared GO BP terms between ROSMAP-AD and Kamath-PD that are unique to cell type groups are shown. The inner circle denotes the cell type from AD, while the outer circle denotes the PD cell type. Each wedge represents a GO BP term shared between one AD and one PD gene module. Colors represent distinct cell types, and the numerical values indicate the PLS component in which the GO term was identified for that cell type. All GO terms shown are exclusive to the corresponding AD-PD module pair (that is, not occurring in any other module pair combination). This highlighted specific, non-redundant biological processes that converged across AD and PD in a cell type dependent manner.

Supplementary Figure S6: GSEA of gene modules identifies shared terms between Seattle-AD and Smajić-PD, an external validation snRNA-seq dataset pair. Shared GO terms across pairwise gene modules from AD and PD. Gene ontology (GO) biological process, molecular function, and cellular component terms are combined. Brighter color indicates a higher number of shared terms, while the white grid represents no overlap. (A) Total number of overlapping terms across all gene modules in AD (blue) or PD (red). The intersection region shows the number of shared terms. (B) Graph visualization of selected biological processes from the 2023 Gene Ontology database, focusing on both AD and PD. The largest subnetworks from the full GO hierarchical tree are shown. Node colors correspond to the disease label, and node size reflects the gene-set size. Core mechanisms such as immune processes, glucose metabolism, and ATP synthesis match the results from the ROSMAP-AD and Kamath-PD analysis arm. The top shared genes (robust genes per module are considered) across all AD-PD gene modules enriching for the representative terms are annotated. (C) Ranked GO BP terms based on their frequency of occurrence across AD-PD module pairs. Mitochondrial energy synthesis terms rank highest among all GO terms.

Supplementary Figure S7: Comparison between Seattle-AD- and Smajić-PD-associated differentially expressed genes. (A) Pairwise associations between Seattle-AD and Smajić-PD differentially expressed genes (DEGs) are shown (Kendall’s tau-b). For each cell type pair, statistical significance of association was assessed using a permutation test. Colored squares indicate significant associations (FDR < 0.05). Darker green (brown) denotes greater similarity (anti) between the log-fold change of significant DEGs from an AD-PD cell type pair. Square size is proportional to −log10(FDR). (B) Shared GO terms from gene set enrichment analysis of DEGs between Seattle-AD and Smajić-PD. Significant terms in AD or PD from GSEA analysis were assessed (FDR q < 0.1).

Supplementary Figure S8: Implicated genes are paralleled between the 2 analysis arms- DGCN and PLS. (A) The degree of similarity between PLS modules (Fig. 2A) and differential gene co-expression network (DGCN) is highlighted here. The fraction of robust genes shared between a PLS gene module and a GWAS-seeded DGCN, relative to the total number of genes in the GWAS-seeded DGCN for a given cell type, was used to quantify the degree of similarity. Left: ROSMAP-AD; Right: Kamath-PD. Each large grid represents a cell type combination (e.g., the top left grid shows the overlap of genes between excitatory neuron modules and excitatory neuron-derived co-expression networks). Lines within a larger grid correspond to a unique GWAS-seeded DGCN. Brighter colors indicate a higher degree of similarity, that is, a greater number of shared genes. Strong signatures along the diagonal suggested that similar gene cliques were derived for the same cell types using 2 orthogonal approaches, validating both analysis arms. (B) Distribution of explained variances for the first 12 principal components based on permutation analysis. Each vertical bar represents the mean explained variance for each PCA component across 100 permutations of the data. The error bars represent the 5/95% CI, indicating the range of variability under permutation. The empirical explained variance is marked with a red star for each component.

Supplementary_table_revision.xlsx.

Supplementary Material

giag059_Supplemental_Files
giag059_Authors_Response_To_Reviewer_Comments_original_submission
giag059_GIGA-D-25-00403_Original_Submission
giag059_GIGA-D-25-00403_Revision_1
giag059_Reviewer_1_Report_original_submission

Reviewer 1 -- 11/9/2025

giag059_Reviewer_2_Report_original_submission

Reviewer 2 -- 12/29/2025

giag059_Reviewer_2_Report_revision_1

Reviewer 2 -- 4/28/2026

Contributor Information

Anwesha Bhattacharya, Department of Biological and Biomedical Engineering, McGill University, 3775, rue University, Montréal, QC H3A 2B4, Canada; Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada.

Edward A Fon, Department of Neurology and Neurosurgery, MNI, McGill University, 3801 Rue University, Ville-Marie, Montréal, QC H3A 2B4, Canada.

Alain Dagher, Department of Psychology, MNI, McGill University, 2001 McGill College Ave, Montreal H3A 1G1, Canada; McConnell Brain Imaging Centre (BIC), MNI, 3801 University Street, Montreal H3A 2B4, Canada.

Yasser Iturria-Medina, Department of Neurology and Neurosurgery, MNI, McGill University, 3801 Rue University, Ville-Marie, Montréal, QC H3A 2B4, Canada; McConnell Brain Imaging Centre (BIC), MNI, 3801 University Street, Montreal H3A 2B4, Canada; Ludmer Centre for Neuroinformatics and Mental Health, 2001, McGill College Ave, suite 1310, Montreal H3A 1G1, Canada.

Jo Anne Stratton, Department of Neurology and Neurosurgery, MNI, McGill University, 3801 Rue University, Ville-Marie, Montréal, QC H3A 2B4, Canada.

Chloe Savignac, Department of Biological and Biomedical Engineering, McGill University, 3775, rue University, Montréal, QC H3A 2B4, Canada; Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada.

Jack Stanley, Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada; Quantitative Life Sciences, McGill University, 550 Sherbrooke West, Montreal H3A 1E3, Canada.

Liam Hodgson, Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada; School of Computer Science, McGill University, 3480 University Street, Montreal H3A 0E9, Canada.

Badr Ait Hammou, Department of Biological and Biomedical Engineering, McGill University, 3775, rue University, Montréal, QC H3A 2B4, Canada; Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada.

David A Bennett, Rush Alzheimer’s Disease Center, Rush University Medical Center, 1750 West Harrison Street, Suite 1000, Chicago, IL 60612, USA.

Danilo Bzdok, Department of Biological and Biomedical Engineering, McGill University, 3775, rue University, Montréal, QC H3A 2B4, Canada; Mila - Quebec Artificial Intelligence Institute, 6666 Saint-Urbain Street, Montréal H2S 3H1, Canada; The Neuro, Montreal Neurological Institute (MNI), McGill University, 3801 Rue University, Ville-Marie, Montreal, QC H3A 2B4, Canada; School of Computer Science, McGill University, 3480 University Street, Montreal H3A 0E9, Canada.

Author contributions

A.B. and D.B. conceptualized the project, planned the experiments, and analyzed the results. All authors helped write the manuscript and analyze the results. D.B. led data analysis.

Funding

National Institute on Aging, ROSMAP, P30AG10161; National Institute on Aging, ROSMAP, P30AG72975; National Institute on Aging, ROSMAP, R01AG15819; National Institute on Aging, ROSMAP, R01AG17917; National Institute on Aging, ROSMAP, U01AG46152; National Institute on Aging, ROSMAP, U01AG61356; Brain Canada Foundation, Danilo Bzdok; Canada Brain Research Fund, Health Canada, Danilo Bzdok; National Institutes of Health, NIH R01 AG068563A, Danilo Bzdok; National Institutes of Health, NIH R01 DA053301-01A1, Danilo Bzdok; National Institute of Health, NIH R01 MH129858-01A1, Danilo Bzdok; Canadian Institute of Health Research, CIHR 438531, Danilo Bzdok; Canadian Institute of Health Research, CIHR 470425, Danilo Bzdok; Healthy Brains Healthy Lives initiative, Canada First Research Excellence fund, Danilo Bzdok; IVADO R3AI initiative, Canada First Research Excellence fund, Danilo Bzdok; CIFAR Artificial Intelligence Chairs program Canada Institute for Advanced Research, Danilo Bzdok.

Data availability

The snRNA-seq PFC data originated from Mathys et al. [36] are available through Synapse under the doi 10.7303/syn184851755. The data are available under controlled use conditions set by human privacy regulations. The snRNA-seq MTG data originating from Gabitto et al. [37] are available through SEA-AD consortium’s web portal at SEA-AD.org. The snRNA-seq substantia nigra data originating from Kamath et al. [40] are available at the Gene Expression Omnibus (GEO) with accession number GSE178265. The snRNA-seq midbrain data originating from Smajić et al. [35] are available for download from the Gene Expression Omnibus (GEO) with accession number GSE157783. The snRNA-seq COPD dataset originating from Adams et al. [128] are available for download from the Gene Expression Omnibus (GEO) with accession number GSE136831. All custom analysis code is available on GitHub at https://github.com/dblabs-mcgill-mila/AD-PD-overlap-study. DOME-ML (Data, Optimization, Model, and Evaluation in Machine Learning) annotations are available via the DOME registry (accession v09w7a6a5r) [166]. All additional supplementary materials are available in the GigaScience repository, GigaDB [167].

Competing interests

D.B. is an equity holder at MindState Design Labs, USA. The authors declare no other competing interests.

References

  • 1. Bloem  B R, Okun  M S, Klein  C. Parkinson’s disease. The Lancet. 2021;397:2284–303. 10.1016/S0140-6736(21)00218-X. [DOI] [PubMed] [Google Scholar]
  • 2. Gustavsson  A, Norton  N, Fast  T  et al.  Global estimates on the number of persons across the Alzheimer’s disease continuum. Alzheimers Dement. 2023;19:658–70. 10.1002/alz.12694. [DOI] [PubMed] [Google Scholar]
  • 3. Melnikova  I. Therapies for Alzheimer’s disease. Nat Rev Drug Discovery. 2007;6:341–42. 10.1038/nrd2314. [DOI] [PubMed] [Google Scholar]
  • 4. Kamath  T, Macosko  E Z. Insights into neurodegeneration in Parkinson’s disease from single-cell and spatial genomics. Mov Disord. 2023;38:518–25. 10.1002/mds.29374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Albers  M W, Gilmore  G C, Kaye  J  et al.  At the interface of sensory and motor dysfunctions and Alzheimer’s disease. Alzheimer’s & Dementia. 2015;11:70–98. 10.1016/j.jalz.2014.04.514. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Oldham  M C, Konopka  G, Iwamoto  K  et al.  Functional organization of the transcriptome in human brain. Nat Neurosci. 2008;11:1271–82. 10.1038/nn.2207. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Twohig  D, Nielsen  H M. α-Synuclein in the pathophysiology of Alzheimer’s disease. Mol Neurodegener. 2019;14:23. 10.1186/s13024-019-0320-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Aarsland  D, Batzu  L, Halliday  G M  et al.  Parkinson disease-associated cognitive impairment. Nat Rev Dis Primers. 2021;7:1. 10.1038/s41572-021-00280-3. [DOI] [PubMed] [Google Scholar]
  • 9. Cummings  J, Lee  G, Ritter  A  et al.  Alzheimer’s disease drug development pipeline: 2020. Alzheimers Dement. 2020;6:e12050. 10.1002/trc2.12050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. The Brainstorm Consortium, Anttila  V, Bulik-Sullivan  B  et al.  Analysis of shared heritability in common disorders of the brain. Science. 2018;360:eaap8757. 10.1126/science.aap8757. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Wightman  D P, Savage  J E, Tissink  E  et al.  The genetic overlap between Alzheimer’s disease, amyotrophic lateral sclerosis, lewy body dementia, and Parkinson’s disease. Neurobiol Aging. 2023;127:99–112. 10.1016/j.neurobiolaging.2023.03.004. [DOI] [PubMed] [Google Scholar]
  • 12. Balusu  S, Praschberger  R, Lauwers  E. Neurodegeneration cell per cell. Neuron. 2023;111:767–86. 10.1016/j.neuron.2023.01.016. [DOI] [PubMed] [Google Scholar]
  • 13. Zhang  X, Gao  F, Wang  D  et al.  Tau pathology in Parkinson’s disease. Front Neurol. 2018;9:809. 10.3389/fneur.2018.00809. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Aarsland  D, Kurz  M W. The epidemiology of dementia associated with Parkinson disease. J Neurol Sci. 2010;289:18–22. 10.1016/j.jns.2009.08.034. [DOI] [PubMed] [Google Scholar]
  • 15. Schneider  J A, Arvanitakis  Z, Leurgans  S E  et al.  The neuropathology of probable Alzheimer’s disease and mild cognitive impairment. Ann Neurol. 2009;66:200–08. 10.1002/ana.21706. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Schneider  J A, Arvanitakis  Z, Yu  L  et al.  Cognitive impairment, decline and fluctuations in older community-dwelling subjects with lewy bodies. Brain. 2012;135:3005–14. 10.1093/brain/aws234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Schneider  J A, Li  J L, Li  Y  et al.  Substantia nigra tangles are related to gait impairment in older persons. Ann Neurol. 2006;59:166–73. 10.1002/ana.20723. [DOI] [PubMed] [Google Scholar]
  • 18. Armstrong  R A, Lantos  P L, Cairns  N J. Overlap between neurodegenerative disorders. Neuropathology. 2005;25:111–24. 10.1111/j.1440-1789.2005.00605.x. [DOI] [PubMed] [Google Scholar]
  • 19. Perl  D P, Warren  C O, Calne  D. Alzheimer’s disease and parkinson’s disease: distinct entities or extremes of a spectrum of neurodegeneration?. Ann Neurol. 1998;44:S19–31. 10.1002/ana.410440705. [DOI] [PubMed] [Google Scholar]
  • 20. Wu  Y, Sun  R, Ren  S  et al.  Neuronal reshaping of the tumor microenvironment in tumorigenesis and metastasis: bench to clinic. Medicine Advances. 2025;3:364–71. 10.1002/med4.70044. [DOI] [Google Scholar]
  • 21. Lei  H Y, Pi  G L, He  T  et al.  Targeting vulnerable microcircuits in the ventral hippocampus of male transgenic mice to rescue Alzheimer-like social memory loss. Military Med Res. 2024;11:16. 10.1186/s40779-024-00512-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Desikan  R S, Schork  A J, Wang  Y  et al.  Genetic overlap between Alzheimer’s disease and Parkinson’s disease at the MAPT locus. Mol Psychiatry. 2015;20:1588–95. 10.1038/mp.2015.6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Sadeghi  I, Gispert  J D, Palumbo  E  et al.  Brain transcriptomic profiling reveals common alterations across neurodegenerative and psychiatric disorders. Comput Struct Biotechnol J. 2022;20:4549–61. 10.1016/j.csbj.2022.08.037. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Wingo  T S, Liu  Y, Gerasimov  E S  et al.  Shared mechanisms across the major psychiatric and neurodegenerative diseases. Nat Commun. 2022;13:1. 10.1038/s41467-022-31873-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Le Bars  S, Glaab  E. Single-cell cortical transcriptomics reveals common and distinct changes in cell–cell communication in Alzheimer’s and Parkinson’s disease. Mol Neurobiol. 2025;62:2655–73. 10.1007/s12035-024-04419-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Barabási  A L, Gulbahce  N, Loscalzo  J. Network medicine: a network-based approach to human disease. Nat Rev Genet. 2011;12:56–68. 10.1038/nrg2918. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Crow  M, Gillis  J. Co-expression in single-cell analysis: saving Grace or Original Sin?. Trends Genet. 2018;34:823–31. 10.1016/j.tig.2018.07.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Subramanian  A, Tamayo  P, Mootha  V K  et al.  Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci. 2005;102:15545–50. 10.1073/pnas.0506580102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Zhang  B, Horvath  S. A general framework for weighted gene co-expression network analysis. Stat Appl Genet Mol Biol. 2005;4:17. 10.2202/1544-6115.1128. [DOI] [PubMed] [Google Scholar]
  • 30. Gerstein  M B, Kundaje  A, Hariharan  M  et al.  Architecture of the human regulatory network derived from ENCODE data. Nature. 2012;489:91–100. 10.1038/nature11245. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Saelens  W, Cannoodt  R, Saeys  Y. A comprehensive evaluation of module detection methods for gene expression data. Nat Commun. 2018;9:1090. 10.1038/s41467-018-03424-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Zhao  Y, Jia  M, Ding  C  et al.  Time-restricted feeding mitigates Alzheimer’s disease-associated cognitive impairments via a B. pseudolongum-propionic acid-FFAR3 axis. IMeta. 2025;4:e70006. 10.1002/imt2.70006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Xiong  X, James  B T, Boix  C A  et al.  Epigenomic dissection of Alzheimer’s disease pinpoints causal variants and reveals epigenome erosion. Cell. 2023;186:4422–37.e21. 10.1016/j.cell.2023.08.040. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Mathys  H, Peng  Z, Boix  C A  et al.  Single-cell atlas reveals correlates of high cognitive function, dementia, and resilience to Alzheimer’s disease pathology. Cell. 2023;186:4365–4385.e27. 10.1016/j.cell.2023.08.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Smajić  S, Prada-Medina  C A, Landoulsi  Z  et al.  Single-cell sequencing of human midbrain reveals glial activation and a Parkinson-specific neuronal state. Brain. 2022;145:964–78. 10.1093/brain/awab446. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Mathys  H, Davila-Velderrain  J, Peng  Z  et al.  Single-cell transcriptomic analysis of Alzheimer’s disease. Nature. 2019;570:332–37. 10.1038/s41586-019-1195-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Gabitto  M I, Travaglini  K J, Rachleff  V M  et al.  Integrated multimodal cell atlas of Alzheimer’s disease. Nat Neurosci. 2024; 27:2366–83. 10.1038/s41593-024-01774-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Pak  V, Adewale  Q, Bzdok  D  et al.  Distinctive whole-brain cell types predict tissue damage patterns in thirteen neurodegenerative conditions. eLife. 2024;12:RP89368. 10.7554/eLife.89368. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Ali  M, Timsina  J, Xu  Y  et al.  Large-scale CSF and plasma proteomics reveal immune, synaptic, and extracellular matrix disruptions across neurodegenerative diseases. Neuron. 2026. 10.1016/j.neuron.2026.02.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Kamath  T, Abdulraouf  A, Burris  S J  et al.  Single-cell genomic profiling of human dopamine neurons identifies a population that selectively degenerates in Parkinson’s disease. Nat Neurosci. 2022;25:588–95. 10.1038/s41593-022-01061-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Bzdok  D, Ioannidis  JPA. Exploration, inference, and prediction in neuroscience and biomedicine. Trends Neurosci. 2019;42:251–62. 10.1016/j.tins.2019.02.001. [DOI] [PubMed] [Google Scholar]
  • 42. Hodgson  L, Li  Y, Iturria-Medina  Y  et al.  Supervised latent factor modeling isolates cell-type-specific transcriptomic modules that underlie Alzheimer’s disease progression. Commun Biol. 2024;7:1–19. 10.1038/s42003-024-06273-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Zhu  B, Park  J M, Coffey  S  et al.  Single-cell transcriptomic and proteomic analysis of Parkinson’s disease brains. Sci Transl Med. 2024;16:eabo1997. 10.1101/2022.02.14.480397. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Bellenguez  C, Küçükali  F, Jansen  I E  et al.  New insights into the genetic etiology of Alzheimer’s disease and related dementias. Nat Genet. 2022;54:4. 10.1038/s41588-022-01024-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Nalls  M A, Blauwendraat  C, Vallerga  C L  et al.  Identification of novel risk loci, causal insights, and heritable risk for Parkinson’s disease: a meta-genome wide association study. Lancet Neurol. 2019;18:1091–1102. 10.1016/S1474-4422(19)30320-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Billingsley  K J, Bandres-Ciga  S, Saez-Atienzar  S  et al.  Genetic risk factors in Parkinson’s disease. Cell Tissue Res. 2018;373:9–20. 10.1007/s00441-018-2817-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Reynolds  R H, Botía  J, Nalls  M A  et al.  Moving beyond neurons: the role of cell type-specific gene regulation in Parkinson’s disease heritability. NPJ Parkinsons Dis. 2019;5:1–14. 10.1038/s41531-019-0076-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Langfelder  P, Horvath  S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinf. 2008;9:559. 10.1186/1471-2105-9-559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Roy  S, Lagree  S, Hou  Z  et al.  Integrated module and gene-specific regulatory inference implicates upstream signaling networks. PLoS Comput Biol. 2013;9:e1003252. 10.1371/journal.pcbi.1003252. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Watson  M. CoXpress: differential co-expression in gene expression data. BMC Bioinf. 2006;7:509. 10.1186/1471-2105-7-509. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Szwedo  A A, Dalen  I, Pedersen  K F  et al.  GBA and APOE impact cognitive decline in Parkinson’s disease: a 10-year population-based study. Mov Disord. 2022;37:1016–27. 10.1002/mds.28932. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52. Zenuni  H, Bovenzi  R, Bissacco  J  et al.  Clinical and neurochemical correlates of the APOE genotype in early-stage Parkinson’s disease. Neurobiol Aging. 2023;131:24–28. 10.1016/j.neurobiolaging.2023.07.011. [DOI] [PubMed] [Google Scholar]
  • 53. Davis  A A, Inman  C E, Wargel  Z M  et al.  APOE genotype regulates pathology and disease progression in synucleinopathy. Sci Transl Med. 2020;12:eaay3069. 10.1126/scitranslmed.aay3069. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Krüger  R, Kuhn  W, Müller  T  et al.  AlaSOPro mutation in the gene encoding α-synuclein in Parkinson’s disease. Nat Genet. 1998;18:106–08. 10.1038/ng0298-106. [DOI] [PubMed] [Google Scholar]
  • 55. Devine  M J, Gwinn  K, Singleton  A  et al.  Parkinson’s disease and α-synuclein expression. Mov Disord. 2011;26:2160–68. 10.1002/mds.23948. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Khan  S S, LaCroix  M, Boyle  G  et al.  Bidirectional modulation of Alzheimer phenotype by alpha-synuclein in mice and primary neurons. Acta Neuropathol. 2018;136:589–605. 10.1007/s00401-018-1886-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Larson  M E, Sherman  M A, Greimel  S  et al.  Soluble α-synuclein is a novel modulator of Alzheimer’s disease pathophysiology. J Neurosci. 2012;32:10253–66. 10.1523/JNEUROSCI.0581-12.2012. [DOI] [PMC free article] [PubMed] [Google Scholar] [Retracted]
  • 58. Gan  L, Cookson  M R, Petrucelli  L  et al.  Converging pathways in neurodegeneration, from genetics to mechanisms. Nat Neurosci. 2018;21:10. 10.1038/s41593-018-0237-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59. Gao  F B, Richter  J D, Cleveland  D W. Rethinking unconventional translation in neurodegeneration. Cell. 2017;171:994–1000. 10.1016/j.cell.2017.10.042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60. Bence  N F, Sampat  R M, Kopito  R R. Impairment of the ubiquitin-proteasome system by protein aggregation. Science. 2001;292:1552–55. 10.1126/science.292.5521.1552. [DOI] [PubMed] [Google Scholar]
  • 61. Qadir  A, Kumar  A, Nagpal  R  et al.  Understanding the ubiquitin proteasome system: history and revolution. In: Nandave  M, Jain  P, eds. PROTAC-Mediated Protein Degradation: A Paradigm Shift in Cancer Therapeutics. Singapore: Springer; 2024:1–20. 10.1007/978-981-97-5077-1_1. [DOI] [Google Scholar]
  • 62. Tai  H C, Schuman  E M. Ubiquitin, the proteasome and protein degradation in neuronal function and dysfunction. Nat Rev Neurosci. 2008;9:826–38. 10.1038/nrn2499. [DOI] [PubMed] [Google Scholar]
  • 63. Millecamps  S, Julien  J P. Axonal transport deficits and neurodegenerative diseases. Nat Rev Neurosci. 2013;14:161–76. 10.1038/nrn3380. [DOI] [PubMed] [Google Scholar]
  • 64. Andreu-Carbó  M, Egoldt  C, Velluz  M C  et al.  Microtubule damage shapes the acetylation gradient. Nat Commun. 2024;15:2029. 10.1038/s41467-024-46379-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Naren  P, Samim  K S, Tryphena  K P  et al.  Microtubule acetylation dyshomeostasis in Parkinson’s disease. Transl Neurodegener. 2023;12:20. 10.1186/s40035-023-00354-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66. Guedes-Dias  P, Holzbaur  ELF. Axonal transport: driving synaptic function. Science. 2019;366:eaaw9997. 10.1126/science.aaw9997. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Lin  M T, Beal  M F. Mitochondrial dysfunction and oxidative stress in neurodegenerative diseases. Nature. 2006;443:787–95. 10.1038/nature05292. [DOI] [PubMed] [Google Scholar]
  • 68. Quntanilla  R A, Tapia-Monsalves  C. The role of mitochondrial impairment in Alzheimer’s disease neurodegeneration: the tau connection. Curr Neuropharmacol. 2020;18:1076–91. 10.2174/1570159×18666200525020259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Pellegrini  L, Wetzel  A, Grannó  S  et al.  Back to the tubule: microtubule dynamics in Parkinson’s disease. Cell Mol Life Sci. 2017;74:409–34. 10.1007/s00018-016-2351-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70. Esteves  A R, Arduino  D M, Swerdlow  R H  et al.  Microtubule depolymerization potentiates alpha-synuclein oligomerization. Front Aging Neurosci. 2010;1:5. 10.3389/neuro.24.005.2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71. Jaunmuktane  Z, Brandner  S. Invited review: the role of prion-like mechanisms in neurodegenerative diseases. Neuropathol Appl Neurobiol. 2020;46:522–45. 10.1111/nan.12592. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Walker  L C, Jucker  M. The prion principle and Alzheimer’s disease. Science. 2024;385:1278–79. 10.1126/science.adq5252. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Diamond  M I. Travels with tau prions. Cytoskeleton. 2024;81:83–88. 10.1002/cm.21806. [DOI] [PubMed] [Google Scholar]
  • 74. Kaufman  S K, Sanders  D W, Thomas  T L  et al.  Tau prion strains dictate patterns of cell pathology, progression rate, and regional vulnerability in vivo. Neuron. 2016;92:796–812. 10.1016/j.neuron.2016.09.055. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Rauch  J N, Olson  S H, Gestwicki  J E. Interactions between microtubule-associated protein tau (MAPT) and small molecules. Cold Spring Harb Perspect Med. 2017;7:a024034. 10.1101/cshperspect.a024034. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76. Camilleri  A, Zarb  C, Caruana  M  et al.  Mitochondrial membrane permeabilisation by amyloid aggregates and protection by polyphenols. Biochim Biophys Acta. 2013;1828:2532–43. 10.1016/j.bbamem.2013.06.026. [DOI] [PubMed] [Google Scholar]
  • 77. Cui  J, Zhao  S, Li  Y  et al.  Regulated cell death: discovery, features and implications for neurodegenerative diseases. Cell Commun Signal. 2021;19:120. 10.1186/s12964-021-00799-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78. Datta  S R, Dudek  H, Tao  X  et al.  Akt phosphorylation of BAD couples survival signals to the cell-intrinsic death machinery. Cell. 1997;91:231–41. 10.1016/s0092-8674(00)80405-5. [DOI] [PubMed] [Google Scholar]
  • 79. Erekat  N S. Apoptosis and its role in Parkinson’s disease. In: Stoker  T B, Greenland  J C, eds. Parkinson’s Disease: Pathogenesis and Clinical Aspects. Brisbane: Codon Publications; 2018. Accessed August 1, 2024. http://www.ncbi.nlm.nih.gov/books/NBK536724/. [Google Scholar]
  • 80. Mochizuki  H, Goto  K, Mori  H  et al.  Histochemical detection of apoptosis in Parkinson’s disease. J Neurol Sci. 1996;137:120–23. 10.1016/0022-510x(95)00336-z. [DOI] [PubMed] [Google Scholar]
  • 81. Kumar  A, Ganini  D, Mason  R P. Role of cytochrome c in α-synuclein radical formation: implications of α-synuclein in neuronal death in Maneb- and paraquat-induced model of Parkinson’s disease. Mol Neurodegen. 2016;11:70. 10.1186/s13024-016-0135-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82. McKenzie  A T, Moyon  S, Wang  M  et al.  Multiscale network modeling of oligodendrocytes reveals molecular components of myelin dysregulation in Alzheimer’s disease. Mol Neurodegener. 2017;12:82. 10.1186/s13024-017-0219-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83. Agarwal  D, Sandor  C, Volpato  V  et al.  A single-cell atlas of the human substantia nigra reveals cell-specific pathways associated with neurological disorders. Nat Commun. 2020;11:4183. 10.1038/s41467-020-17876-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84. Bartzokis  G. Age-related myelin breakdown: a developmental model of cognitive decline and Alzheimer’s disease. Neurobiol Aging. 2004;25:5–18. author reply 49-62. 10.1016/j.neurobiolaging.2003.03.001. [DOI] [PubMed] [Google Scholar]
  • 85. Depp  C, Sun  T, Sasmita  A O  et al.  Myelin dysfunction drives amyloid-β deposition in models of Alzheimer’s disease. Nature. 2023;618:349–57. 10.1038/s41586-023-06120-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86. Kenigsbuch  M, Bost  P, Halevi  S  et al.  A shared disease-associated oligodendrocyte signature among multiple CNS pathologies. Nat Neurosci. 2022;25:876–86. 10.1038/s41593-022-01104-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87. Zhou  Y, Song  W M, Andhey  P S  et al.  Human and mouse single-nucleus transcriptomics reveal TREM2-dependent and TREM2-independent cellular responses in Alzheimer’s disease. Nat Med. 2020;26:131–42. 10.1038/s41591-019-0695-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88. Balali-Mood  M, Naseri  K, Tahergorabi  Z  et al.  Toxic mechanisms of five heavy metals: mercury, lead, chromium, cadmium, and arsenic. Front Pharmacol. 2021;12:643972. 10.3389/fphar.2021.643972. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89. Haidar  Z, Fatema  K, Shoily  S S  et al.  Disease-associated metabolic pathways affected by heavy metals and metalloid. Toxicol Rep. 2023;10:554–70. 10.1016/j.toxrep.2023.04.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90. Li  B, Xia  M, Zorec  R  et al.  Astrocytes in heavy metal neurotoxicity and neurodegeneration. Brain Res. 2021;1752:147234. 10.1016/j.brainres.2020.147234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91. Huiliang  Z, Mengzhe  Y, Xiaochuan  W  et al.  Zinc induces reactive astrogliosis through ERK-dependent activation of Stat3 and promotes synaptic degeneration. J Neurochem. 2021;159:1016–27. 10.1111/jnc.15531. [DOI] [PubMed] [Google Scholar]
  • 92. Gamez  P, Caballero  A B. Copper in Alzheimer’s disease: implications in amyloid aggregation and neurotoxicity. AIP Adv. 2015;5:092503. 10.1063/1.4921314. [DOI] [Google Scholar]
  • 93. Pal  A, Rani  I, Pawar  A  et al.  Microglia and astrocytes in Alzheimer’s disease in the context of the aberrant copper homeostasis hypothesis. Biomolecules. 2021;11:1598. 10.3390/biom11111598. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94. Zhou  Q, Zhang  Y, Lu  L  et al.  Copper induces microglia-mediated neuroinflammation through ROS/NF-κb pathway and mitophagy disorder. Food Chem Toxicol. 2022;168:113369. 10.1016/j.fct.2022.113369. [DOI] [PubMed] [Google Scholar]
  • 95. Uversky  V N, Li  J, Fink  A L. Metal-triggered structural transformations, aggregation, and fibrillation of human alpha-synuclein. A possible molecular NK between Parkinson’s disease and heavy metal exposure. J Biol Chem. 2001;276:44284–96. 10.1074/jbc.M105343200. [DOI] [PubMed] [Google Scholar]
  • 96. Sarell  C J, Wilkinson  S R, Viles  J H. Substoichiometric levels of Cu2+ ions accelerate the kinetics of fiber formation and promote cell toxicity of amyloid-β from Alzheimer disease *. J Biol Chem. 2010;285:41533–40. 10.1074/jbc.M110.171355. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97. Myhre  O, Utkilen  H, Duale  N  et al.  Metal dyshomeostasis and inflammation in Alzheimer’s and Parkinson’s diseases: possible impact of environmental exposures. Oxid Med Cell Longev. 2013;2013:726954. 10.1155/2013/726954. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98. Acosta-Cabronero  J, Betts  M J, Cardenas-Blanco  A  et al.  In vivo MRI mapping of brain iron deposition across the adult lifespan. J Neurosci. 2016;36:364–74. 10.1523/JNEUROSCI.1907-15.2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99. Belaidi  A A, Bush  A I. Iron neurochemistry in Alzheimer’s disease and Parkinson’s disease: targets for therapeutics. J Neurochem. 2016;139:179–97. 10.1111/jnc.13425. [DOI] [PubMed] [Google Scholar]
  • 100. Bjørklund  G, Hofer  T, Nurchi  V M  et al.  Iron and other metals in the pathogenesis of Parkinson’s disease: toxic effects and possible detoxification. J Inorg Biochem. 2019;199:110717. 10.1016/j.jinorgbio.2019.110717. [DOI] [PubMed] [Google Scholar]
  • 101. Thomas  GEC, Leyland  L A, Schrag  A E  et al.  Brain iron deposition is linked with cognitive severity in Parkinson’s disease. J Neurol Neurosurg Psychiatry. 2020;91:418–25. 10.1136/jnnp-2019-322042. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102. Ward  R J, Zucca  F A, Duyn  J H  et al.  The role of iron in brain ageing and neurodegenerative disorders. Lancet Neurol. 2014;13:1045–60. 10.1016/S1474-4422(14)70117-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103. Lenz  K M, Nelson  L H. Microglia and beyond: innate immune cells As regulators of brain development and behavioral function. Front Immunol. 2018;9:698. 10.3389/fimmu.2018.00698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104. Tansey  M G, Wallings  R L, Houser  M C  et al.  Inflammation and immune dysfunction in Parkinson disease. Nat Rev Immunol. 2022;22:657–73. 10.1038/s41577-022-00684-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105. Khoury  J E, Luster  A D. Mechanisms of microglia accumulation in Alzheimer’s disease: therapeutic implications. Trends Pharmacol Sci. 2008;29:626–32. 10.1016/j.tips.2008.08.004. [DOI] [PubMed] [Google Scholar]
  • 106. Chen  X, Firulyova  M, Manis  M  et al.  Microglia-mediated T cell infiltration drives neurodegeneration in tauopathy. Nature. 2023;615:668–77. 10.1038/s41586-023-05788-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107. González  H, Pacheco  R. T-cell-mediated regulation of neuroinflammation involved in neurodegenerative diseases. J Neuroinflam. 2014;11:201. 10.1186/s12974-014-0201-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108. Xu  Y, Li  Y, Wang  C  et al.  The reciprocal interactions between microglia and T cells in Parkinson’s disease: a double-edged sword. J Neuroinflam. 2023;20:33. 10.1186/s12974-023-02723-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109. Villanueva  I, Alva-Sánchez  C, Pacheco-Rosado  J. The role of thyroid hormones as inductors of oxidative stress and neurodegeneration. Oxid Med Cell Longev. 2013;2013:218145. 10.1155/2013/218145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110. Ewins  D L, Rossor  M N, Butler  J  et al.  Association between autoimmune thyroid disease and familial Alzheimers disease. Clin Endocrinol. 1991;35:93–96. 10.1111/j.1365-2265.1991.tb03502.x. [DOI] [PubMed] [Google Scholar]
  • 111. Kalmijn  S, Mehta  K M, Pols  H A  et al.  Subclinical hyperthyroidism and the risk of dementia. The Rotterdam study. Clin Endocrinol (Oxf). 2000;53:733–37. 10.1046/j.1365-2265.2000.01146.x. [DOI] [PubMed] [Google Scholar]
  • 112. Mohammadi  S, Dolatshahi  M, Rahmani  F. Shedding light on thyroid hormone disorders and Parkinson disease pathology: mechanisms and risk factors. J Endocrinol Invest. 2021;44:1–13. 10.1007/s40618-020-01314-5. [DOI] [PubMed] [Google Scholar]
  • 113. Kim  D K, Choi  H, Lee  W  et al.  Brain hypothyroidism silences the immune response of microglia in Alzheimer’s disease animal model. Sci Adv. 2024;10:eadi1863. 10.1126/sciadv.adi1863. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114. Seo  B A, Kim  D, Hwang  H  et al.  TRIP12 ubiquitination of glucocerebrosidase contributes to neurodegeneration in Parkinson’s disease. Neuron. 2021;109:3758–3774.e11. 10.1016/j.neuron.2021.09.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115. Hong  S, Beja-Glasser  V F, Nfonoyim  B M  et al.  Complement and microglia mediate early synapse loss in Alzheimer mouse models. Science. 2016;352:712–16. 10.1126/science.aad8373. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116. Song  P, Peng  W, Sauve  V  et al.  Parkinson’s disease-linked parkin mutation disrupts recycling of synaptic vesicles in human dopaminergic neurons. Neuron. 2023;111:3775–3788.e7. 10.1016/j.neuron.2023.08.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117. Tzioras  M, Daniels  MJD, Davies  C  et al.  Human astrocytes and microglia show augmented ingestion of synapses in Alzheimer’s disease via MFG-E8. CR Med. 2023;4:101175. 10.1016/j.xcrm.2023.101175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 118. Das  M, Mao  W, Voskobiynyk  Y  et al.  Alzheimer risk-increasing TREM2 variant causes aberrant cortical synapse density and promotes network hyperexcitability in mouse models. Neurobiol Dis. 2023;186:106263. 10.1016/j.nbd.2023.106263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119. Filipello  F, Morini  R, Corradini  I  et al.  The microglial innate immune receptor TREM2 is required for synapse elimination and normal brain connectivity. Immunity. 2018;48:979–991.e8. 10.1016/j.immuni.2018.04.016. [DOI] [PubMed] [Google Scholar]
  • 120. Guo  Y, Wei  X, Yan  H  et al.  TREM2 deficiency aggravates α-synuclein-induced neurodegeneration and neuroinflammation in Parkinson’s disease models. FASEB J. 2019;33:12164–74. 10.1096/fj.201900992R. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121. Shafi  S, Singh  A, Ibrahim  A M  et al.  Role of triggering receptor expressed on myeloid cells 2 (TREM2) in neurodegenerative dementias. Eur J Neurosci. 2021;53:3294–10. 10.1111/ejn.15215. [DOI] [PubMed] [Google Scholar]
  • 122. Alexander  J J, Anderson  A J, Barnum  S R  et al.  The complement cascade: Yin–Yang in neuroinflammation—neuro-protection and -degeneration. J Neurochem. 2008;107:1169–87. 10.1111/j.1471-4159.2008.05668.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123. Carbutt  S, Duff  J, Yarnall  A  et al.  Variation in complement protein C1q is not a major contributor to cognitive impairment in Parkinson’s disease. Neurosci Lett. 2015;594:66–69. 10.1016/j.neulet.2015.03.048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124. Hallett  P J, Engelender  S, Isacson  O. Lipid and immune abnormalities causing age-dependent neurodegeneration and Parkinson’s disease. J Neuroinflam. 2019;16:153. 10.1186/s12974-019-1532-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125. Sienski  G, Narayan  P, Bonner  J M  et al.  APOE4 disrupts intracellular lipid homeostasis in human iPSC-derived glia. Sci Transl Med. 2021;13:eaaz4564. 10.1126/scitranslmed.aaz4564. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126. Hou  Yf, Shan  C, Zhuang  Sy  et al.  Gut microbiota-derived propionate mediates the neuroprotective effect of osteocalcin in a mouse model of Parkinson’s disease. Microbiome. 2021;9:34. 10.1186/s40168-020-00988-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127. Bennett  D A, Buchman  A S, Boyle  P A  et al.  Religious orders study and rush memory and aging project. J Alzheimers Dis. 2018;64:S161–89. 10.3233/JAD-179939. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128. Adams  T S, Schupp  J C, Poli  S  et al.  Single-cell RNA-seq reveals ectopic and aberrant lung-resident cell populations in idiopathic pulmonary fibrosis. Sci Adv. 2020;6:eaba1983. 10.1126/sciadv.aba1983. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129. Krishnaswami  S R, Grindberg  R V, Novotny  M  et al.  Using single nuclei for RNA-seq to capture the transcriptome of postmortem neurons. Nat Protoc. 2016;11:499–524. 10.1038/nprot.2016.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130. Ma  P, Liu  X, Xu  Z  et al.  Joint profiling of gene expression and chromatin accessibility during amphioxus development at single-cell resolution. Cell Rep. 2022;39:110979. 10.1016/j.celrep.2022.110979. [DOI] [PubMed] [Google Scholar]
  • 131. Ahlmann-Eltze  C, Huber  W. Comparison of transformations for single-cell RNA-seq data. Nat Methods. 2023;20:665–72. 10.1038/s41592-023-01814-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132. Zhu  Y, Wang  L, Yin  Y  et al.  Systematic analysis of gene expression patterns associated with postmortem interval in human tissues. Sci Rep. 2017;7:1. 10.1038/s41598-017-05882-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133. Ferreira  P G, Muñoz-Aguirre  M, Reverter  F  et al.  The effects of death and post-mortem cold ischemia on human tissue transcriptomes. Nat Commun. 2018;9:1. 10.1038/s41467-017-02772-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134. Hughes  G. On the mean accuracy of statistical pattern recognizers. IEEE Trans Inf Theory. 1968;14:55–63. 10.1109/TIT.1968.1054102. [DOI] [Google Scholar]
  • 135. Jia  W, Sun  M, Lian  J  et al.  Feature dimensionality reduction: a review. Complex Intell Syst. 2022;8:2663–93. 10.1007/s40747-021-00637-x. [DOI] [Google Scholar]
  • 136. Thompson  J R, Nelson  E D, Tippani  M  et al.  An integrated single-nucleus and spatial transcriptomics atlas reveals the molecular landscape of the human hippocampus. Nat Neurosci. 2025; 28:1990–2004. 10.1038/s41593-025-02022-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137. Kotliar  D, Veres  A, Nagy  M A  et al.  Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-seq. eLife. 2019;8:e43803. 10.7554/eLife.43803. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138. Mathys  H, Boix  C A, Akay  L A  et al.  Single-cell multiregion dissection of Alzheimer’s disease. Nature. 2024;632:858–68. 10.1038/s41586-024-07606-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139. Barker  M, Rayens  W. Partial least squares for discrimination. J Chemom. 2003;17:166–73. 10.1002/cem.785. [DOI] [Google Scholar]
  • 140. Chun  H, Keleş  S. Sparse partial least squares regression for simultaneous dimension reduction and variable selection. J R Stat Soc Series B Stat Methodol. 2010;72:3–25. 10.1111/j.1467-9868.2009.00723.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141. Agarwal  A, Shah  D, Shen  D  et al.  On robustness of principal component regression. J Am Stat Assoc. 2021;116:1731–1745. 10.48550/arXiv.1902.10920. [DOI] [Google Scholar]
  • 142. Luecken  M D, Theis  F J. Current best practices in single-cell RNA-seq analysis: a tutorial. Mol Syst Biol. 2019;15:e8746. 10.15252/msb.20188746. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143. Lee  L C, Liong  C Y, Jemain  A A. Partial least squares-discriminant analysis (PLS-DA) for classification of high-dimensional (HD) data: a review of contemporary practice strategies and knowledge gaps. Analyst. 2018;143:3526–39. 10.1039/C8AN00599K. [DOI] [PubMed] [Google Scholar]
  • 144. Wold  HOA. Nonlinear Iterative Partial Least Squares (NIPALS) Modelling: Some Current Developments. 1973. https://api.semanticscholar.org/CorpusID:118962244. Accessed May 12, 2026.
  • 145. Rodríguez-Pérez  R, Fernández  L, Marco  S. Overoptimism in cross-validation when using partial least squares-discriminant analysis for omics data: a systematic study. Anal Bioanal Chem. 2018;410:5981–92. 10.1007/s00216-018-1217-1. [DOI] [PubMed] [Google Scholar]
  • 146. de Boves Harrington  P. Statistical validation of classification and calibration models using bootstrapped Latin partitions. TrAC Trends Anal Chem. 2006;25:1112–24. 10.1016/j.trac.2006.10.010. [DOI] [Google Scholar]
  • 147. Miller  K L, Alfaro-Almagro  F, Bangerter  N K  et al.  Multimodal population brain imaging in the UK Biobank prospective epidemiological study. Nat Neurosci. 2016;19:1523–36. 10.1038/nn.4393. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148. Spreng  R N, Dimas  E, Mwilambwe-Tshilobo  L  et al.  The default network of the human brain is associated with perceived social isolation. Nat Commun. 2020;11:6393. 10.1038/s41467-020-20039-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149. Moon  K R, van Dijk  D, Wang  Z  et al.  Visualizing structure and transitions in high-dimensional biological data. Nat Biotechnol. 2019;37:1482–92. 10.1038/s41587-019-0336-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 150. Arndt  S, Turvey  C, Andreasen  N C. Correlating and predicting psychiatric symptom ratings: Spearman’s r versus Kendall’s tau correlation. J Psychiatr Res. 1999;33:97–104. 10.1016/s0022-3956(98)90046-2. [DOI] [PubMed] [Google Scholar]
  • 151. Benjamini  Y, Hochberg  Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol. 1995;57:289–300. 10.1111/j.2517-6161.1995.tb02031.x. [DOI] [Google Scholar]
  • 152. Stark  R, Grzelak  M, Hadfield  J. RNA sequencing: the teenage years. Nat Rev Genet. 2019;20:631–56. 10.1038/s41576-019-0150-2. [DOI] [PubMed] [Google Scholar]
  • 153. Finak  G, McDavid  A, Yajima  M  et al.  MAST: a flexible statistical framework for assessing transcriptional changes and characterizing heterogeneity in single-cell RNA sequencing data. Genome Biol. 2015;16:278. 10.1186/s13059-015-0844-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154. Ruxton  G D. The unequal variance t-test is an underused alternative to Student’s t-test and the Mann–Whitney U test. Behav Ecol. 2006;17:688–90. 10.1093/beheco/ark016. [DOI] [Google Scholar]
  • 155. Welch  B L. The generalization of “student’s” problem when several different population varlances are involved. Biometrika. 1947;34:28–35. 10.1093/biomet/34.1-2.28. [DOI] [PubMed] [Google Scholar]
  • 156. Fang  Z, Liu  X, Peltz  G. GSEApy: a comprehensive package for performing gene set enrichment analysis in Python. Bioinformatics. 2023;39:btac757. 10.1093/bioinformatics/btac757. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 157. Chen  E Y, Tan  C M, Kou  Y  et al.  Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinf. 2013;14:128. 10.1186/1471-2105-14-128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 158. Ashburner  M, Ball  C A, Blake  J A  et al.  Gene ontology: tool for the unification of biology. Nat Genet. 2000;25:25–29. 10.1038/75556. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 159. The Gene Ontology Consortium, Aleksander  S A, Balhoff  J  et al.  The gene ontology knowledgebase in 2023. Genetics. 2023;224:iyad031. 10.1093/genetics/iyad031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 160. Carbon  S, Ireland  A, Mungall  C J  et al.  AmiGO: online access to ontology and annotation data. Bioinformatics. 2009;25:288–89. 10.1093/bioinformatics/btn615. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 161. Merico  D, Isserlin  R, Stueker  O  et al.  Enrichment map: a network-based method for gene-set Enrichment visualization and interpretation. PLoS One. 2010;5:e13984. 10.1371/journal.pone.0013984. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 162. Reimand  J, Isserlin  R, Voisin  V  et al.  Pathway enrichment analysis and visualization of omics data using g:profiler, GSEA, Cytoscape and EnrichmentMap. Nat Protoc. 2019;14:2. 10.1038/s41596-018-0103-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 163. Shannon  P, Markiel  A, Ozier  O  et al.  Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13:2498–2504. 10.1101/gr.1239303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 164. Saltoun  K, Adolphs  R, Paul  L K  et al.  Dissociable brain structural asymmetry patterns reveal unique phenome-wide profiles. Nat Hum Behav. 2023;7:251–68. 10.1038/s41562-022-01461-0. [DOI] [PubMed] [Google Scholar]
  • 165. Crouse  D F. On implementing 2D rectangular assignment algorithms. IEEE Trans Aerosp Electron Syst. 2016;52:1679–96. 10.1109/TAES.2016.140952. [DOI] [Google Scholar]
  • 166. Bhattacharya  A, Fon  E A, Dagher  A  et al.  Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer’s and Parkinson’s disease. [DOME-ML Annotations]. DOME-ML Registry. 2025. GigaScience, ahead of print. https://doi.org/GIGA-D-25-00403. [DOI] [PMC free article] [PubMed]
  • 167. Bhattacharya  A, Fon  E A, Dagher  A  et al.  et al. Supporting data for "Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer's and Parkinson's disease". GigaScience Database. 10.5524/102822. [DOI] [PMC free article] [PubMed]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Citations

  1. Bhattacharya  A, Fon  E A, Dagher  A  et al.  Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer’s and Parkinson’s disease. [DOME-ML Annotations]. DOME-ML Registry. 2025. GigaScience, ahead of print. https://doi.org/GIGA-D-25-00403. [DOI] [PMC free article] [PubMed]
  2. Bhattacharya  A, Fon  E A, Dagher  A  et al.  et al. Supporting data for "Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer's and Parkinson's disease". GigaScience Database. 10.5524/102822. [DOI] [PMC free article] [PubMed]

Supplementary Materials

giag059_Supplemental_Files
giag059_Authors_Response_To_Reviewer_Comments_original_submission
giag059_GIGA-D-25-00403_Original_Submission
giag059_GIGA-D-25-00403_Revision_1
giag059_Reviewer_1_Report_original_submission

Reviewer 1 -- 11/9/2025

giag059_Reviewer_2_Report_original_submission

Reviewer 2 -- 12/29/2025

giag059_Reviewer_2_Report_revision_1

Reviewer 2 -- 4/28/2026

Data Availability Statement

The snRNA-seq PFC data originated from Mathys et al. [36] are available through Synapse under the doi 10.7303/syn184851755. The data are available under controlled use conditions set by human privacy regulations. The snRNA-seq MTG data originating from Gabitto et al. [37] are available through SEA-AD consortium’s web portal at SEA-AD.org. The snRNA-seq substantia nigra data originating from Kamath et al. [40] are available at the Gene Expression Omnibus (GEO) with accession number GSE178265. The snRNA-seq midbrain data originating from Smajić et al. [35] are available for download from the Gene Expression Omnibus (GEO) with accession number GSE157783. The snRNA-seq COPD dataset originating from Adams et al. [128] are available for download from the Gene Expression Omnibus (GEO) with accession number GSE136831. All custom analysis code is available on GitHub at https://github.com/dblabs-mcgill-mila/AD-PD-overlap-study. DOME-ML (Data, Optimization, Model, and Evaluation in Machine Learning) annotations are available via the DOME registry (accession v09w7a6a5r) [166]. All additional supplementary materials are available in the GigaScience repository, GigaDB [167].


Articles from GigaScience are provided here courtesy of Oxford University Press

RESOURCES