Abstract
Alzheimer’s disease is a complex neurodegenerative disorder characterized by progressive cognitive decline and neuroinflammation. Although its molecular hallmarks are well documented, cell-type-specific mechanisms driving gene dysregulation remain elusive. While single-cell RNA sequencing resolves cellular states, most studies focus on individual genes rather than coordinated programs. Moreover, a gap persists between interpretable network-based models and artificial intelligence foundation models, which capture complex interactions but lack mechanistic transparency. Whether these approaches converge or provide complementary views remains unclear. We present an integrated study combining SCANet, for reconstructing co-expression and gene regulatory networks, with scGPT foundation model. Applied to over 1.3 million cells across 18 cell types, this approach revealed that Alzheimer-associated transcriptional changes concentrate within coherent co-expression modules, for extracellular matrix organization, immune signaling, and neuronal communication. Genes prioritized by scGPT were largely embedded within SCANet modules, indicating convergence at the gene level; however, higher-order architecture agreement was limited and cell-type-specific. Thus, scGPT highlights influential genes, whereas SCANet resolves their modular organization, providing complementary information. By integrating both methods, we recovered known Alzheimer-related pathways and identified novel regulatory candidates, including the CEBPB-CENPQ axis in vulnerable SST-GABA interneurons. This demonstrates that combining network biology with foundation models enables gene prioritization and mechanistic interpretation.
Graphical Abstract
Graphical Abstract.

Introduction
Alzheimer’s disease (AD) is a hallmarked complex neurodegenerative disorder caused by amyloid-β accumulation, tau phosphorylation, synaptic dysfunction, and chronic neuroinflammation [1], and its genetic landscape in AD implicates immune response, lipid metabolism, and complement pathways [2]. Its prevalence continues to rise due to global aging [3], making mechanistic understanding increasingly urgent [4]. However, the mechanisms by which AD disrupts gene regulation within specific cell types remain unclear. Patients vary across age of onset, progression rate, symptom profile, and response to treatments [5], suggesting that these clinical manifestations likely mirror the underlying molecular heterogeneity. Stratified genetic analyses in late-onset AD show that genetic risk and heritability differ between younger onset (60–79) and older onset (≥ 80) groups [6], indicating heterogeneous genetic architecture across age ranges. Moreover, the human brain is highly heterogeneous, and the patterns of gene co-expression and regulation likely differ by cell type. As disease progression alters cellular composition and relative abundance within the brain tissue, resolving AD pathogenesis requires molecular analyses at single-cell resolution to accurately capture these dynamic, cell-type-specific changes.
Traditional bulk transcriptomics approaches average signals across large numbers of cells, thereby obscuring the heterogeneity that is a defining feature of the AD brain. Different neuronal and glial populations follow distinct molecular trajectories during disease progression, highlighting the need for single-cell resolution. In this context, single-cell RNA sequencing (scRNA-seq) has transformed the field by enabling characterization of transcriptional states at the individual cell levels. In Alzheimer’s disease research, this transformative technology has enabled the discovery of disease-associated microglial activation states [7], subtype-specific astrocyte reactivity, and transcriptional shifts in neuronal populations [8]. For instance, large-scale scRNA-seq atlases now exist that integrate dozens of AD and control samples, linking cell subpopulations to neuropathology and genetic risk variants [9], and uncovering previously hidden heterogeneity in AD brains, enabling inference of cellular trajectories and cross-cell-type interactions in disease contexts [8–10]. However, single-cell datasets are sparse, suffer from batch effects, and have limited read depth [11]. Furthermore, interpreting differential gene expression is insufficient; network-level inference is required to elucidate how gene modules or regulatory circuits are disrupted in disease, offering a systems-level perspective on disease biology.
Network medicine provides a framework for addressing this challenge by modeling genes as interconnected modules and regulatory circuits. In the field of AD, gene co-expression network (GCNs) analysis has identified modules correlated with clinical severity or pathology, revealing modules enriched in immune response, synaptic signaling, mitochondrial function, and glial activation [12]. Gene regulatory network (GRN) approaches extend this framework by inferring transcription factors (TFs) that are likely to drive the activity of the genes within these modules. GRNs often integrate prior knowledge (TF binding, motif information) and temporal or pseudotime data [13], and they have been used to predict regulatory networks in multiple brain cell types in AD, uncovering modules enriched in AD risk genes and candidate drug targets [14]. Tools such as SCANet [15] enable the reconstruction of both co-expression modules and GRNs directly from scRNA-seq data, thereby enhancing hub gene identification and therapeutic target prediction.
Concurrently, the emergence of foundational models designed for single-cell analysis has introduced a powerful, complementary perspective. Models such as scGPT [16], trained on millions of cells, learn latent representations of genes and cells that encode nonlinear dependencies and subtle patterns of covariation beyond pairwise correlations. When applied to disease contexts, scGPT can quantify how the influence of a gene on its neighbors changes between conditions by analyzing differences in attention scores, suggesting that gene–gene interactions are most informative of variation across conditions. However, powerful foundation models often act as black boxes, offering limited direct interpretability in terms of modules or mechanistic circuits. Additionally, recent evaluations (e.g. in zero-shot contexts) have cautioned about reliability and sensitivity to dataset bias [17].
These considerations lead to the central question guiding this study, namely, how Alzheimer’s disease alters the cell-type-specific structure of gene expression and regulatory networks. However, despite recent advances, a clear methodological gap remains. Most studies operate either in the network medicine regime, inferring co-expression or regulatory modules, or in the foundation model regime, learning latent influence embeddings. Few have attempted to bridge the two: by asking whether the genes deemed most influential by artificial intelligence (AI) models lie within the structured modules discovered by network inference. Attempts such as “GenoHoption” [18] and foundation model-graph neuronal network integrations [19] remain limited in scope. Integrative frameworks combining molecular networks and advanced learning models are essential for bridging predictive signals with mechanistic insights [20]. Without such integration, one loses either interpretability or sensitivity to nonlinear patterns. In the Alzheimer’s context, this gap is especially critical, as we do not yet know whether the molecular drivers of disease are concentrated in rewired modules or dispersed across modules through subtle influence. Addressing this question is essential for linking predictive signals to mechanistic understanding, and for identifying cell-type-specific regulatory vulnerabilities that may underlie disease heterogeneity.
To bridge this gap, we introduce a two-layer analytical pipeline that integrates SCANet and scGPT to characterize AD-associated transcriptional dysregulation across major brain cell types. SCANet is used to reconstruct cell-type-specific co-expression modules and regulatory networks in both healthy and diseased conditions, while scGPT quantifies disease-driven changes in gene influence using a transformer model trained on large-scale single-cell datasets. An overview of the proposed analytical comparative workflow is presented in Fig. 1.
Figure 1.

Overview of the integrative SCANet-scGPT analytical pipeline: (A) Preprocessed SEA-AD scRNA-seq data are filtered, and subset by cell type and condition (Reference and High Alzheimer’s disease). Highly variable genes (HVGs) are selected and representative cells (RCs) are sampled to ensure scalability and robustness. (B) SCANet workflow for each cell type and condition, including the reconstruction of gene co-expression networks (GCNs), identification of co-expression modules, and inference of gene regulatory networks (GRNs) within each module. (C) scGPT-based analysis, where a pretrained transformer model is used to compute differential attention scores between Reference and High conditions, yielding cell-type-specific gene influence networks and directed TF–target interactions. (D) Cross-analysis between methods. Node- and edge-level overlaps between SCANet- and scGPT-derived networks were evaluated, and functional concordance was assessed through comparison of enriched Gene set enrichment analysis (GSEA) terms across methods. (E) Benchmark validation. Genes identified by SCANet and scGPT were cross-referenced with external Alzheimer’s disease gene sets from Open Targets and CTD. Overlap analyses and network-based validation highlighted both previously validated Alzheimer’s genes and novel candidate genes supported by literature evidence.
Materials and methods
Ethics statement
This study involved the secondary analysis of publicly available, de-identified human single-nucleus RNA-sequencing data from the Seattle Alzheimer’s Disease Brain Cell Atlas (SEA-AD). No participants were recruited, no human samples were collected, and no interventions were performed as part of the present study. Therefore, additional ethical approval and informed consent were not required for this secondary analysis. The original SEA-AD study was conducted in accordance with applicable ethical standards. Informed consent for research brain donation was obtained under protocols approved by the University of Washington and Kaiser Permanente Washington Health Research Institute Institutional Review Boards, as reported by Gabitto et al. [21]. The present study was conducted in accordance with the principles of the Declaration of Helsinki insofar as applicable to the secondary analysis of previously collected, de-identified human data.
Dataset and cohort definition
The dataset used in this study was obtained from the SEA-AD [21] (Seattle Alzheimer’s Disease Brain Cell Atlas) consortium, funded by the National Institute on Ageing (NIA, NIH), which aims to construct the most detailed multimodal cellular atlas of the human brain affected by Alzheimer’s disease, with an emphasis on the early stages of the disease and normal aging.
This analysis focused on single-nucleus RNA sequencing (snRNA-seq) data obtained from the middle temporal gyrus (MTG), a brain region critical for semantic memory and language comprehension that is particularly vulnerable in the early AD, and the most extensively sampled area in SEA-AD, with >2 million nuclei profiled using the BICCN (BRAIN Initiative Cell Census Network) cellular classification reference.
We worked with the preprocessed and curated file “Whole Taxonomy - MTG: Seattle Alzheimer’s Disease Atlas (SEA-AD)” containing 1 378 211 cells and 36 412 genes, extracted from the CELLxGENE [22] platform. This file was selected for its high-quality, standardized cell annotation and preprocessing, which includes normalization, logarithmic transformation, and initial dimensionality reduction. The analyzed dataset came from 84 adult donors covering the entire neuropathological spectrum of AD, supplemented by a reference group of five young neurotypical donors. The SEA-AD cohort was specifically designed to capture a continuous neuropathological spectrum rather than a simple case-control comparison. In the original study, donor-level quality control demonstrated highly consistent cellular representation across the cohort, with significant cell abundance analyses restricted to cell populations detected in at least 75% of donors. Moreover, the original analyses accounted for relevant biological and technical covariates, including age, sex, technology, and APOE4 status, and key findings were independently replicated in an additional brain region and across 10 public datasets comprising 707 donors [21]. Transcriptomic data were generated using the 10x Genomics Chromium platform (v3.1). The dataset includes detailed neuropathological annotations (Braak stage, CERAD score, and Thal phase), cognitive metrics, and demographic variables.
Five analytical groups were defined according to the Overall Alzheimer’s Disease Neuropathological Change (ADNC) score, which classifies the degree of Alzheimer’s pathology according to NIA-AA guidelines. The Reference group (n = 137 303 cells) corresponded to canonical healthy controls from the SEA-AD study and was used as a comparative baseline. The Not AD group (n = 158 238 cells) included individuals with variable Braak stages (0–V) but with a CERAD score of zero and Thal phase 0 [23], indicating an absence of β-amyloid deposition and therefore not meeting the NIA-AA criteria for Alzheimer’s diagnosis. The remaining three groups represented patients with confirmed Alzheimer’s disease, stratified by severity into Low (n = 181 162 cells), Intermediate (n = 358 256 cells), and High (n = 543 252 cells) ADNC categories.
Analyses were conducted across the 18 cell types defined in the CELLxGENE annotation, encompassing glutamatergic neurons (L2IT, NP, CT, L6B, L5EP), GABAergic subtypes (PVALB, VIP, SST, CHAN, CGE, SNC, LMP), and major glial classes such as astrocytes (AST), microglia (MGC), oligodendrocytes (OLG), and oligodendrocyte precursor cells (OPC), and vascular cells (VLC).
Quality control and feature selection
The initial CELLxGENE dataset had already undergone quality control, normalization, log transformation, and dimensionality reduction. Additional filtering steps were applied to ensure data integrity. The initial dataset contained 36 412 genes, and genes detected in fewer than 1% of all cells (14 379 genes) were excluded, resulting in a filtered dataset of 22 033 genes. Cells with mitochondrial transcript counts exceeding 5% of total counts were also removed.
Highly variable genes (HVGs) were identified using the highly_variable_genes function implemented in Scanpy, which ranks genes according to their normalized dispersion relative to mean expression, thereby selecting genes that capture biologically meaningful cell-to-cell variability while accounting for the dependence of variance on expression level. Different HVG set sizes were evaluated, and the top 6000 HVGs were retained as the final feature set. Cells were further filtered to include only those annotated with ADNC conditions “Reference” and “High”. The preprocessing workflow is summarized in Supplementary Fig. 1A. Each of the 18 cell types was analyzed separately under these two conditions (Supplementary Fig. 1B), yielding a working dataset of 681 275 cells and 6000 genes.
Cell-type-specific gene co-expression and regulatory network inference using SCANet
Gene co-expression and gene regulatory networks were reconstructed using SCANet [15], a previously published framework for scRNA-seq network analysis. Full methodological details are provided in the original SCANet publication [15]. SCANet parameters were optimized to ensure scalability across cell-type-specific networks by evaluating a representative subset of neuronal and glial populations under both disease conditions while systematically varying the number of representative cells (RCs), obtained by aggregating neighboring cells with similar transcriptional profiles.
Briefly, SCANet reconstructs gene co-expression networks using WGCNA to identify co-expressed gene modules. Module eigengenes were correlated independently with the Reference and High AD conditions to identify disease-associated co-expression patterns. Disease-relevant modules were defined as those exhibiting an absolute difference greater than 0.5 between the module-condition correlation coefficients (|RefCorr − HighCorr| > 0.5). This threshold was selected to prioritize modules showing marked shifts in coordinated gene expression between conditions, while excluding modules displaying only modest changes. Gene regulatory networks are subsequently inferred by integrating co-expression modules with transcription factor information, and predicted regulatory interactions are filtered using motif enrichment analysis based on the RcisTarget database.
Foundation model-based inference of gene interaction and regulatory networks with scGPT
Gene interaction networks were inferred using scGPT, a transformer-based foundational model pretrained on large-scale scRNA-seq data. We used the optimized scGPT Brain model, trained on 13.2 million brain cells. Detailed methodological descriptions are provided in the original publication [16].
scGPT was applied independently to each cell type using highly variable preselected genes. Specifically, each of the 18 cell-type-specific datasets was processed separately, and 6000 HVGs were used as input for the model. Disease-associated gene–gene influence networks were derived using the “Difference” configuration, which quantifies changes in attention patterns between AD and reference conditions. For each cell type, the top 100 most influential genes were selected, and pairwise attention-difference scores (10 000 gene pairs) were used to construct weighted, undirected influence networks, collapsing reciprocal edges by retaining the maximum score to generate a nonredundant, weighted co-influence network per cell type.
From these networks, directed gene regulatory networks (scGPT-GRNs) were derived by retaining edges where the source gene matched a curated TF list [24, 25], yielding attention-derived TF → target gene relationships prioritized by the foundation model under Alzheimer’s perturbation.
Functional enrichment of network-derived gene sets
GSEA was performed to functionally characterize the gene sets derived from both significant co-expression analyses in a unified network. SCANet, enrichment was conducted for each co-expression module identified per cell type using Enrichr. Modules were classified as significant when they contained >three enriched terms with an adjusted P-value < .05 across the evaluated gene set collections (Gene Ontology Biological Process, Cellular Component, Molecular Function, and KEGG), indicating consistent functional enrichment. For scGPT, enrichment was applied to the 100 most influential genes for each cell type. To improve biological relevance and statistical sensitivity, enrichment analyses were performed using the 22 033 genes detected in the scRNA-seq dataset, a strategy that reduced P-values and increased the detection of significantly enriched terms by restricting the universe of comparison to genes that were actually expressed in the data.
Integration of SCANet and scGPT frameworks
To uncover multiscale patterns of transcriptional reorganization in AD, SCANet-derived co-expression and regulatory modules were integrated with scGPT-derived gene influence networks. For each cell type, the top 100 genes ranked by scGPT attention difference were intersected with SCANet-defined modules under both “Reference” and “High” conditions. Overlap was evaluated at the gene, network, and functional levels.
Gene-level overlap was defined as the number of shared genes between the scGPT-prioritized set and each SCANet module. Network-level overlap was computed as the proportion of shared gene–gene connections. Functional overlap was quantified using the shared GO and KEGG terms obtained via GSEA. Additionally, GRNs inferred by SCANet were compared with scGPT-derived GRNs to evaluate concordance at the regulatory layer and to identify modules containing genes strongly influenced by disease according to scGPT. SCANet GRNs were generated for each significant module containing TFs within the corresponding GCN, producing one or two GRNs per module depending on the condition availability, with each condition analyzed separately. In contrast, scGPT generates a single GRN per cell type by extracting GCNs weighted by attention scores and assigning directionality using TF annotations. Comparative analysis was performed for each cell type. Finally, to assess whether the observed overlap between scGPT-prioritized sets and SCANet modules exceeded random expectation, one-sided Fisher’s exact tests were conducted against a background of all expressed genes, followed by Benjamini–Hochberg FDR correction across cell types (FDR < 0.05). For multilayer comparisons (network and functional levels), standardized Z scores were calculated relative to size-matched null distributions.
Benchmarking against curated Alzheimer’s disease gene resources
Results were benchmarked against two curated AD repositories: OpenTargets [26] (release 2024.09), genes with globalScore ≥ 0.2 and at least one genetic or experimental evidence, n = 1 116), and the Comparative Toxicogenomics Database [27] (CTD, 2025 release, 110 curated disease–gene associations annotated to “Alzheimer’s Disease”). Jaccard indices were computed to quantify overlap, enabling identification of genes absent from current databases and highlighting putative novel regulators. Enriched GO and KEGG terms from both approaches were contrasted with OpenTargets and CTD-annotated AD terms to evaluate functional concordance.
Gene-level intersections were computed across SCANet, scGPT, OpenTargets, and CTD, and cross-referenced with GCN-, GRN-, and GSEA-level results. Finally, to validate identified regulatory hubs, external cellular vulnerability profiles from Gabitto et al. [21] were used as an independent benchmark, focusing on SST-GABA interneurons. Vulnerability scores from the SEA-AD dataset were compared with the direction of regulatory influence inferred from scGPT-based attention analysis.
To further evaluate representative regulatory interactions identified by scGPT, we quantified the joint activity of the CEBPB-CENPQ and EMX2-IL6R gene pairs in SST-GABA cells using UCell [28]. UCell scores were computed from the original expression matrix and compared across pathology conditions (Reference versus High) and APOE4 carrier status (Reference versus APOE4 N versus APOE4 Y) using two-sided Mann–Whitney U tests.
Results
To determine whether interpretable network reconstruction and transformer-based foundation models converge on shared disease programs or provide distinct perspectives, we systematically integrated SCANet-derived co-expression and regulatory networks with scGPT-derived influence networks across 17 cell types retained after filtering from the initial 18 SEA-AD annotations from 1.3 million cells. Our central finding is that the two analytical approaches strongly converge at the gene level yet diverge at the level of network organization and regulatory wiring, thereby providing complementary and nonredundant views of AD-associated transcriptional alterations.
SCANet reconstructs modular, cell-type-specific regulatory programs altered in Alzheimer’s disease
Cell type-specific reconstruction of co-expression networks using SCANet displayed a clear modular organization across all the analyzed cell populations. Across the full dataset, 715 modules were identified, with a median of 42 modules per cell type (Fig. 2A). The number of detected modules varied considerably across cell types, ranging from 23 to 72, and was not proportional to the number of profiled cells, suggesting that module complexity reflects cell-type-specific transcriptional organization rather than sample size alone. Among these, 132 modules met the predefined disease-relevance criterion, indicating substantial differences in module-level expression correlations between Reference and High (AD) conditions. The number of detected modules did not correlate with cell abundance, indicating that SCANet captures intrinsic transcriptional heterogeneity rather than the sampling depth (Supplementary Table 1).
Figure 2.

SCANet-derived co-expression modules, shared transcription factors, and enriched Alzheimer’s-related pathways across brain cell types. (A) Number of SCANet-derived co-expression modules reconstructed in each of the 17 analyzed brain cell types under Reference and High (AD) conditions. (B) Chord diagram showing the number of transcription factors shared across the SCANet-inferred gene regulatory networks (GRNs) of different cell types. Each chord connects cell types that share one or more TF regulators, with thicker connections representing a greater number of shared TFs. (C) Dotplot showing the top biologically relevant terms enriched in SCANet-derived modules across multiple cell types in Alzheimer’s disease. Terms were filtered to include only those containing at least one gene associated with Alzheimer's (from OpenTargets and CTD) and ranked by recurrence across cell types and statistical significance (−log10 adjusted P-value). Dot size corresponds to the number of cell types in which the term is enriched, while dot color indicates the term’s statistical significance.
Building on modular co-expression networks, GRNs were inferred separately for the Reference and High conditions to enable direct comparison of regulatory organization. After stringent reliability filtering, each GRN retained between 1 and 12 transcription factors, representing high-confidence regulatory hubs. Interneurons and astrocytes shared several TFs associated with stress and inflammatory responses, including ATF3, JUN/JUND, STAT3, and NFE2L2, suggesting activation of partially conserved regulatory programs. In contrast, excitatory neural populations displayed more distinct TF repertories, indicating that similar disease-associated biological processes may be regulated through cell-type-specific regulatory mechanisms. These findings support the central premise of our study, that AD disrupts shared molecular pathways while engaging distinct regulatory architectures across different cell types (Fig. 2B). Moreover, Alzheimer-associated transcriptional changes were not diffusely distributed across the network but instead concentrated within discrete, coherent modules. The predominant enriched pathways involved extracellular matrix organization, inflammatory signaling, neuronal communication, proteostasis, and mitochondrial function, consistent with core molecular processes underlying AD pathogenesis (Fig. 2C; Supplementary Table 2).
scGPT captures influence-based gene dependencies under Alzheimer’s perturbation
To contrast correlation-based modular reconstruction with transformer-based modeling, we derived gene influence networks from the scGPT attention weights. Differential attention scores between the Reference and High conditions were aggregated for each cell type to construct networks in which edges reflected condition-associated changes in learned gene dependencies.
These influence networks were highly connected and compact across all cell types. Community detection identified between three and seven modules per cell type with modest modularity values. Although most gene–gene interactions showed low-magnitude attention values, each cell type displayed a tail of higher weights corresponding to the most influential gene interactions under AD perturbation (Fig. 3A). In contrast to SCANet, which emphasizes modular separation, scGPT-derived networks display dense interconnectivity, consistent with the distributed dependency structures learned by the transformer model.
Figure 3.

scGPT attention patterns, transcription factor influence, and Alzheimer’s-related pathway enrichment across brain cell types: (A) Distribution of scGPT attention scores across all gene–gene pairs per cell type. (B) TF influence across cell types from scGPT-directed GRNs. Rows represent TFs shared between up to two cell types, connected by gray lines. Dots indicate TFs in specific cell types, with size and color reflecting the number of target genes (global scale). (C) Dotplot showing the top biologically relevant terms enriched in scGPT-most influenced genes across multiple cell types in Alzheimer’s disease. Terms were filtered to include only those containing at least one gene associated with Alzheimer’s (from OpenTargets and CTD) and ranked by recurrence across cell types and statistical significance (−log10 adjusted P-value). Dot size corresponds to the number of cell types in which the term is enriched, while dot color indicates the term’s statistical significance.
Restricting the influenced edges to transcription factor-target relationships generated directed regulatory networks for each cell type. Between 2 and 12 TFs were identified per network, with each regulating network for each cell type. Shared TFs across cell types were rare and, when present, they regulated largely distinct target sets (Fig. 3B). Thus, scGPT emphasizes highly cell-type-specific influence hierarchies, with minimal reuse of regulatory rewiring across populations. Functional enrichment of scGPT-prioritized genes predominantly highlighted inflammatory and cytokine-mediated signaling, cell adhesion, antigen presentation, and intracellular signaling pathways, supporting the involvement of immune and regulatory programs associated with Alzheimer’s disease (Fig. 3C; Supplementary Table 3). To illustrate how scGPT highlights genes with broad but cell-type-dependent influence, we examined MTMR2, the only gene among the top recurrent set that showed strong attention-based connectivity primarily in microglia and selected GABAergic neurons, suggesting highly cell-type-specific regulatory relationships (Supplementary Result 1).
Quantitative integration of SCANet and scGPT reveals complementary network perspectives
To determine whether transformer-derived influence patterns converge with modular co-expression programs—and to gain deeper insight into the cellular regulatory architecture—we systematically integrated outputs from SCANet and scGPT across cell types and quantified their agreement at multiple levels.
Agreement between SCANet and scGPT was evaluated at the gene, co-expression, and regulatory network levels. Because SCANet modules contained between 31 and 2047 genes (median of 106.5 genes), whereas scGPT prioritized the top 100 genes per cell type, direct overlap comparisons would be biased by the substantial difference in gene size. The complete overlap statistics for each significant SCANet module are provided in Supplementary Table 4.
Therefore, similarity was assessed using the proportion of scGPT-prioritized genes recovered by SCANet (containment), together with network-based metrics. Specifically, rather than assessing the raw overlap between gene sets, we evaluated edge overlap within GCNs and GRNs to determine which specific gene–gene co-expression and regulatory interactions were consistently preserved across both approaches. To further determine whether the genes prioritized by scGPT simply reflected conventional differential expression, we compared the top 100 scGPT-influenced genes with significant DEGs identified using a standard Wilcoxon rank-sum test (adjusted P-value < .05, |log2FC| > 0.25). Between 41% and 69% of scGPT-prioritized genes (86% in VLC) were not identified as significant DEGs, indicating that a substantial fraction of the transformer-prioritized genes cannot be explained by differential expression alone. Nevertheless, the observed overlaps between scGPT-prioritized genes and DEGs were significantly greater than expected by chance in every cell type (Fisher’s exact test, all FDR < 0.05; Supplementary Table 5), supporting that scGPT captures biologically relevant signals while providing complementary regulatory information beyond expression changes alone.
At a gene level, scGPT-prioritized genes were consistently embedded within SCANet-derived modules, with recall values ranging from 0.99 to 1.00, indicating that nearly all genes prioritized by scGPT were recovered across the set of significant SCANet modules reconstructed for the corresponding cell type. In contrast, Jaccard similarity remained uniformly low (≈0.016), reflecting the expected effect of the substantial difference in gene set size rather than a lack of biological concordance. To exclude a potential sample-size bias, we evaluated whether gene-level agreement was associated with the number of cells available for each cell type. Neither recall nor Jaccard similarity showed a significant association with cell number (Spearman’s ρ = −0.20, P-value = 0.432 for both metrics), indicating that the recovery of scGPT-prioritized genes by SCANet is independent of sample size.
However, these descriptive metrics alone do not account for the overlap expected under random sampling. To determine whether the observed containment simply reflected the larger size of SCANet modules, we evaluated its statistical significance using one-sided Fisher’s exact tests against a background of all expressed genes (Supplementary Table 6). The overlap between scGPT-prioritized genes and SCANet modules was significantly greater than expected by chance in 16 of the 17 analyzed cell types after FDR correction (FDR < 0.05), with odds ratios ranging from 3.34 to 7.29. Together, these analyses indicate that scGPT-prioritized genes are consistently embedded within broader SCANet-derived co-expression programs, and that this enrichment cannot be explained solely by differences in module size or cell-type abundance.
On the other hand, the proportion of shared GCN nodes and edges displayed pronounced cell-type specificity (Fig. 4A). Unlike gene-level agreement, node and edge overlap increased with the number of cells available for each cell type (node overlap: Spearman’s ρ = 0.79, P-value = .0002; edge overlap: Spearman’s ρ = 0.92, P-value < .0001), reflecting the greater structural complexity of SCANet networks reconstructed from larger cellular populations rather than increased biological concordance between both methods. Node overlap ranged from 0.04 in VLC to 0.55 in L2IT-GLU, while edge overlap remained consistently low, between 0 and 0.019. This pattern indicates that although scGPT-influenced genes are frequently present within SCANet modules, their specific pairwise co-expression relationships are only partially conserved. Thus, scGPT captures gene sets aligned with SCANet modules but emphasizes distinct interaction patterns. To confirm these structural convergences were not driven by chance, permutation testing (1000 randomizations) revealed that specific cell types, such as L2IT-GLU, exhibit significantly higher edge and node overlap than expected by random sampling (Edge overlap Z score = 2.4; Node overlap Z score = 2.0).
Figure 4.

Comparing SCANet and scGPT networks: (A) Agreement at the co-expression network (GCN) level, quantified as the proportion of scGPT-derived nodes and edges recovered within SCANet networks for each cell type. (B) Agreement at the gene regulatory network (GRN) level, quantified as the number of shared transcription factors and TF–target regulatory interactions between SCANet and scGPT for each cell type. (C) UpSet plot showing the overlap of transcription factors (TFs) identified by SCANet and scGPT across all analyzed cell types. (D) Dotplot showing the top biologically relevant terms enriched in SCANet modules and scGPT-most influenced genes across multiple cell types in Alzheimer’s disease. Terms were filtered to include only those containing at least one gene associated with Alzheimer’s (from OpenTargets and CTD) and ranked by recurrence across cell types and statistical significance (−log10 adjusted P-value). Dot size corresponds to the number of cell types in which the term is enriched, while dot color indicates the term’s statistical significance. (E) Venn diagram showing the overlap between Alzheimer-related genes from four different sources: SCANet (blue), scGPT (green), OpenTargets (yellow), and CTD (purple). (F) Integrative regulatory case study focusing on shared TFs (CEBPB and EMX2) and their supported target genes. Purple edges represent TF → TG relationships detected by both SCANet and scGPT methods, and the edge width represents the amount of cell types in which that relationship is present.
Regulatory network comparison revealed a selective convergence (Fig. 4B). Sixty transcription factors were shared between strategies, with the strongest overlap observed in SST-GABA and VIP-GABA interneurons (Fig. 4C), a finding statistically supported by permutation analysis demonstrating highly significant regulatory conservation in SST-GABA (TF-TG edge overlap Z score = 3.3; TF overlap Z score = 3.0). Although the fraction of shared regulatory edges per GRN was small (0.4% of all edges), several TFs, including KLF6, CEBPB, MEIS2, SOX8, EMX2, and KLF15, appeared recurrently in both GRN sets. These regulators represent a reproducible core of transcriptional control detected by both tools, whereas the low edge-level overlap indicates regulatory wiring is inferred differently by each method.
To assess whether scGPT-prioritized genes were randomly distributed across SCANet modular space, we quantified their distribution across SCANet modules for each cell type (Supplementary Fig. 3), revealing a nonuniform pattern with preferential concentration in specific modules. To provide an integrated view of agreement across molecular layers across cell types, we summarized gene- (recall), co-expression-, and regulatory-level metrics in a unified representation (Supplementary Fig. 4), showing cell-type-dependent concordance between SCANet and scGPT signals.
At the functional annotation level, restricting analyses of Alzheimer’s-associated genes supported by external benchmarks further demonstrated convergence in core disease programs. The pathways most consistently identified by both approaches included extracellular matrix organization and cell adhesion, together with neuroactive ligand-receptor interactions, calcium signaling, cytokine signaling and immune-related pathways, indicating that both SCANet and scGPT prioritize biologically coherent mechanisms despite capturing distinct regulatory layers (Fig. 4D). In total, nine modules across five cell types (L2IT-GLU, SST-GABA, CGE-INT, AST, and LMP-GABA) exhibited overlap in both the co-expression and regulatory structure (Table 1). SST-GABA interneurons displayed the strongest regulatory consistency across conditions, whereas excitatory neurons and astrocytes showed a more limited but stable overlap.
Table 1.
Cross-reference between SCANet and scGPT results
| CellType | SCANet module | GCN Nodes overlap | GCN Edges overlap | GSEA overlap | Condition | GRN TF overlap | GRN edges overlap |
|---|---|---|---|---|---|---|---|
| SST-GABA | M22 | 18 | 15 | 3 | Reference | 3 | 12 |
| SST-GABA | M22 | 18 | 15 | 3 | High | 2 | 3 |
| L2IT-GLU | M8 | 7 | 10 | 3 | High | 1 | 1 |
| SST-GABA | M2 | 10 | 1 | 2 | High | 1 | 2 |
| CGE-INT | M32 | 9 | 0 | 1 | High | 1 | 1 |
| L2IT-GLU | M16 | 21 | 50 | 7 | Reference | 1 | 1 |
| AST | M46 | 24 | 27 | 2 | Reference | 1 | 1 |
| LMP-GABA | M10 | 12 | 11 | 0 | Reference | 1 | 1 |
| L2IT-GLU | M8 | 7 | 10 | 3 | Reference | 1 | 1 |
Modules of cell types which were detected by both methods and also had at least one GRN for one condition with SCANet are represented.
Benchmarking against Alzheimer’s reference sources
This benchmarking strategy serves two primary purposes. First, it confirmed that the reconstructed co-expression and regulatory networks successfully recovered known Alzheimer’s biology, thereby ensuring the external validity of our computational results. Second, it enabled the differentiation between validated, well-characterized pathways and potentially novel disease modules uncovered by our integrative approach.
External validation against Open Targets and CTD
To evaluate the external validity, genes identified by SCANet and scGPT were compared with Alzheimer’s-associated gene sets curated in OpenTargets and CTD. Jaccard indices indicated modest but consistent overlap, with higher concordance observed for SCANet-derived gene sets than for scGPT top-ranked genes, reflecting differences in selection size and prioritization strategy (Supplementary Result 2; Supplementary Fig. 5).
The intersection of SCANet, scGPT, and OpenTargets revealed a core set of 96 genes supported by both computational approaches and a stringent benchmark but absent from CTD (Fig. 4E). Thirty-one of these genes were present in shared GCNs, indicating their participation in regulatory modules reconstructed independently by both models. Several TFs highlighted by the integrated GRNs, including CEBPB, KLF6, MEIS2, and SOX8, were absent from benchmark annotations, consistent with the limited representation of upstream regulators in these resources, which have not yet been curated.
Within SST-GABA interneurons, scGPT identified CEBPB and EMX2 as convergent regulatory hubs that link inflammatory and developmental programs to disease-related stress states (Fig. 4F). Although direct regulatory interactions remain experimentally unconfirmed, their co-occurrence within disease-associated modules suggests a novel axis of interneuron vulnerability. The inflammatory driver CEBPB (CCAAT/enhancer-binding protein beta), which has an experimentally supported role in Alzheimer’s disease [29–32], connects to the centromeric protein CENPQ, leading us to hypothesize that CEBPB activation in vulnerable SST neurons triggers a stress-coupled transcriptional state, intersecting neuroinflammation with DNA-damage or cell-cycle re-entry programs, rather than canonical neurodegeneration pathways. Similarly, scGPT identifies a novel link between the developmental TF EMX2 (Empty Spiracles Homeobox 2), yet with limited and mostly indirect evidence connecting it to AD [33, 34], and the inflammatory receptor IL6R (Interleukin-6 receptor), a major driver of neuroinflammation that has been linked to cognitive decline, amyloid pathology, and glial activation in AD [35–37]. This suggests a state-dependent interaction in which neuronal identity programs modulate sensitivity to inflammatory signaling. A more detailed discussion of this hypothesis can be found in Supplementary Result 3.
To further evaluate these representative regulatory interactions using the original expression matrix, we quantified the joint activity of the CEBPB-CENPQ and EMX2-IL6R pairs in SST-GABA interneurons using UCell [28]. Both gene pairs showed statistically significant differences in UCell scores between the Reference and High groups (Mann-Whitney test, P-value = 1.39. × 10⁻⁴ and P-value = 8.57 × 10⁻¹⁸, respectively), although the observed shifts in mean UCell scores were modest (Supplementary Table 7), consistent with altered activity of these gene pairs across disease conditions. Importantly, stratification by APOE4 status showed that CEBPB-CENPQ UCell scores were significantly lower in both APOE4-negative and APOE4-positive AD cases than in Reference controls (Mann–Whitney test, P-value < .01 for both comparisons), whereas no significant difference was observed between the two AD groups (P-value = 0.786) (Supplementary Table 8), suggesting that the observed reduction in CEBPB-CENPQ activity is present in AD irrespective of APOE4 carrier status.
Validation of regulatory hubs using SEA-AD cellular vulnerability profiles
To provide independent biological support for our findings, we cross-referenced the identified hubs with the SEA-AD cellular vulnerability dataset, provided in Supplementary Table 7 of Gabitto et al. [21] as an external benchmark for validation. This comparison focused on SST-GABA interneurons, a population identified by both our pipeline and the recent literature as highly susceptible to AD pathology.
Table 2 shows a striking alignment with the SEA-AD “affected” supertypes: the master regulator CEBPB was significantly enriched in these vulnerable cells (0.86), while its predicted target CENPQ was markedly depleted (−0.30). This inverse relationship, captured by our scGPT-based attention analysis, suggests a regulatory breakdown specifically within the neurons that are most prone to loss. This complementary validation confirms that our computational pipeline identifies regulatory drivers that are directly linked to cellular failure in the AD brain.
Table 2.
Validation of integrative regulatory hubs in SST-GABA interneurons using SEA-AD cellular vulnerability profiles
| Gene symbol | SCANet/scGPT role | SEA-AD enrichment (SST affected) | Biological implication |
|---|---|---|---|
| CEBPB | Master regulatory hub | 0.869 | Pro-inflammatory/stress response activation |
| CENPQ | scGPT influenced target | −0.307 | Potential repression/loss of centromere integrity |
| MTMR2 | Multilineage network hub | 1.229 | Membrane trafficking and endolysosomal stress |
| IL6R | High attention difference | 1.438 | SST-specific cytokine signaling vulnerability |
SEA-AD enrichment values correspond to the dimensionless gene enrichment scores reported in Supplementary Table 7 of Gabitto et al. [21] for affected SST-GABA neuronal subtypes.
Discussion
In this study, we introduce a comparative strategy that combines network medicine with foundation models to investigate cell-type-specific transcriptional dysregulation in Alzheimer’s disease. We systematically compared SCANet-derived modular networks and scGPT-derived influence networks to determine whether these complementary modeling paradigms converge on shared disease programs or capture distinct layers of transcriptional reorganization. Our results revealed that AD-associated perturbations are consistently localized to specific gene sets across cell types, while their organization into co-expression and regulatory architectures differs in a manner that reflects multiple levels of disease-related control.
Our results consistently show that AD-associated perturbations are concentrated within specific gene sets across cell types, while preserving much of the underlying network organization. This convergence between SCANet-derived co-expression modules and scGPT-derived regulatory signals suggests that both approaches capture complementary aspects of the same disease-associated biological programs. However, these genes have different structure. SCANet resolves them into discrete co-expression modules associated with specific biological pathways, whereas scGPT highlights compact influence networks that reflect learned dependency patterns between genes. This convergence at the component level but divergence at the organization level suggests that AD-related transcriptional changes are governed by both coordinated pathway-level programs and distributed regulatory dependencies that are not reducible to simple co-expression relationships.
The localization of scGPT-prioritized genes within specific SCANet modules further reinforced this complementarity. When influential genes are concentrated within a given module, this supports the interpretation of that module as a coherent disease-associated program. In contrast, more dispersed influence patterns may reflect indirect, secondary, or indirect perturbations across multiple biological processes. Integration therefore enhances the interpretability of scGPT outputs while prioritizating within SCANet-derived modular structures, enabling disease-relevant modules to be distinguished from background transcriptional variation.
At the regulatory level, the overlap between SCANet- and scGPT-derived GRNs is limited in terms of shared edges but selective in terms of shared hubs. This pattern is biologically informative, as it suggests regulatory rewiring, in which core transcriptional controllers are preserved across models, whereas their downstream targets vary across cell types and disease states. This configuration is consistent with regulatory systems in which master regulators coordinate broad responses through context-dependent target sets. Transcription factors including CEBPB, KLF6, MEIS2, SOX8, and EMX2, recurred as shared hubs, indicating a conserved regulatory layer underlying AD-associated programs. Notably, several of these regulators are underrepresented in curated AD databases, which tend to catalogue target genes rather than upstream drivers, highlighting the added value of integrative network approaches for identifying regulatory mechanisms.
SST-GABA interneurons have emerged as the cell type with the strongest regulatory convergence between methodologies, consistent with accumulating evidence that inhibitory interneurons are selectively vulnerable in AD [38]. Within this population, convergent prioritization of CEBPB and its predicted target CENPQ, together with independent validation using SEA-AD vulnerability profiles, point to a regulatory imbalance in neurons prone to degeneration. Rather than implying a direct mechanistic pathway, these observations suggest that inflammation-associated TFs and chromatin-related genes may be coordinately altered in vulnerable interneurons in AD pathology. Similarly, the co-occurrence of EMX2 and IL6R within convergent disease-associated modules indicates that developmental transcriptional regulators may modulate inflammatory sensitivity in a cell-type-specific context. These findings expand the regulatory landscape of AD beyond canonical microglial and excitatory neuron paradigms, and emphasize the contribution of interneuron-specific programs.
Benchmarking against OpenTargets and CTD confirms that our strategy robustly recovers established AD-associated genes while also identifying convergent regulators absent from curated sources. This highlights the added value of integrative network modeling in uncovering upstream drivers that are not yet represented in disease databases.
This study has several limitations. Both SCANet and scGPT rely on inferred networks rather than direct experimental perturbations and predicted regulatory interactions. Therefore, the remaining hypotheses require validation. In addition, attention-derived influence scores reflect learned dependency patterns and do not directly encode causality. The present analysis captures static disease contrasts and does not address temporal progression or dynamic state transitions.
Future work may extend this analytical approach by integrating complementary modalities, such as chromatin accessibility (scATAC-seq) or spatial transcriptomics, if suitable databases become available, to provide orthogonal support for regulatory interactions. Additionally, patient-specific network approaches and donor-level validation strategies (e.g. pseudobulk or leave-one-donor-out analyses) will help assess whether these regulatory programs generalize across donors or reflect subgroup-specific heterogeneity. Application to longitudinal or stage-resolved AD datasets, as these datasets become available, could further clarify how network perturbations emerge during disease progression. The convergent modules and regulators identified here provide a focused starting point for experimental validation, particularly in vulnerable neuronal populations such as interneurons.
This study demonstrates that jointly applying network medicine with foundation models provides complementary and non-redundant insights into the molecular architecture of Alzheimer’s disease. While both approaches converge on core disease-relevant genes, they reveal distinct layers of transcriptional organization, spanning modular co-expression structures and influence-based regulatory dependencies.
By bridging interpretability and predictive power, this pipeline enables robust gene prioritization and mechanistic understanding of cell-type-specific dysregulation. Its application to large-scale single-cell data highlights conserved extracellular matrix, immune, and neuronal signaling programs while identifying novel regulatory candidates, underscoring its potential as a scalable strategy for studying complex neurodegenerative disorders and guiding future experimental investigation.
Supplementary Material
Acknowledgments
Author contributions Andrea Álvarez-Pérez (Conceptualization [equal], Data curation [equal], Formal analysis [equal], Investigation [equal], Methodology [equal], Validation [equal], Visualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Alejandro Rodríguez-González (Conceptualization [equal], Funding acquisition [equal], Supervision [equal]), Jan Baumbach (Conceptualization [equal], Funding acquisition [equal], Resources [equal], Supervision [equal]), Lucía Prieto Santamaría (Conceptualization [equal], Supervision [equal], Writing—review & editing [equal]), Mhaned Oubounyt (Conceptualization [equal], Methodology [equal], Resources [equal], Software [equal], Validation [equal], Writing—review & editing [equal]). Andrea Álvarez-Pérez: Conceptualization, Methodology, Formal Analysis, Investigation, Data curation, Visualization, Validation, Writing—original draft, Writing—review & editing. Alejandro Rodríguez-González: Conceptualization, Supervision, Funding Acquisition. Jan Baumbach: Conceptualization, Resources, Supervision, Funding Acquisition. Lucía Prieto-Santamaría: Conceptualization, Supervision, Writing—review & editing. Mhaned Oubounyt: Conceptualization, Methodology, Resources, Software, Validation, Writing—Review & Editing.
Contributor Information
Andrea Álvarez-Pérez, Centro de Tecnología Biomédica, Universidad Politécnica de Madrid, Pozuelo de Alarcón, Madrid 28233, Spain; ETS de Ingenieros Informáticos, Universidad Politécnica de Madrid, Boadilla del Monte, Madrid 28660, Spain.
Alejandro Rodríguez-González, Centro de Tecnología Biomédica, Universidad Politécnica de Madrid, Pozuelo de Alarcón, Madrid 28233, Spain; ETS de Ingenieros Informáticos, Universidad Politécnica de Madrid, Boadilla del Monte, Madrid 28660, Spain.
Jan Baumbach, Institute for Computational Systems Biomedicine, University of Hamburg, Hamburg 22761, Germany; Computational BioMedicine lab, University of Southern Denmark, Campusvej 55, 5000 Odense C, Denmark.
Lucía Prieto-Santamaría, Centro de Tecnología Biomédica, Universidad Politécnica de Madrid, Pozuelo de Alarcón, Madrid 28233, Spain; ETS de Ingenieros Informáticos, Universidad Politécnica de Madrid, Boadilla del Monte, Madrid 28660, Spain.
Mhaned Oubounyt, Institute for Computational Systems Biomedicine, University of Hamburg, Hamburg 22761, Germany.
Supplementary data
Supplementary data is available at NAR Genomics & Bioinformatics online.
Conflict of interest
The authors declare that no competing interests exist.
Funding
This work was supported by the Spanish Ministry of Science, Innovation and Universities/State Research Agency (MICIU/AEI/10.13039/501100011033) [Data-driven drug repositioning applying graph neural networks (3DR-GNN) grant number PID2021-122659OB-I00]; European Regional Development Fund (“ERDF A way of making Europe”); German Federal Ministry of Research, Technology and Space (BMFTR) [NetMap project grant number 031L0309B]; and Universidad Politécnica de Madrid and Banco Santander [predoctoral “Programa Propio” grant to A.A.-P.].
Data availability
The data and computer code produced in this study are available in the following databases:
Gabitto MI, Travaglini KJ, Rachleff VM, Kaplan ES, Long B, Ariza J, et al. Integrated multimodal cell atlas of Alzheimer’s disease. Nat Neurosci 2024;27:2366–83 (https://doi.org/10.1038/s41593-024-01774-5).
CELLxGENE dataset for SEA-AD project: https://cellxgene.cziscience.com/collections/1ca90a2d-2943-483d-b678-b809bf464c30.
The code supporting the results and the datasets generated during and/or analysed during the current study are available in the: https://medal.ctb.upm.es/internal/gitlab/disnet/network-medicine/network-medicine-and-scgpt-in-alzheimer-single-cell. Access to peer review can be provided upon request. The repository will be made public upon formal publication.
Supplementary Data are available at NAR Genomics and Bioinformatics Online.
References
- 1. Chandra S, Sisodia SS, Vassar RJ. The gut microbiome in Alzheimer’s disease: what we know and what remains to be explored. Mol Neurodegeneration. 2023;18:9. 10.1186/s13024-023-00595-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Carmona S, Hardy J, Guerreiro R. The genetic landscape of Alzheimer disease. In: Geschwind D.H., Paulson H.L., Klein C. (eds.), Handbook of Clinical Neurology. Vol. 148 (Neurogenetics, Part II). Amsterdam: Elsevier, 2018, pp. 395–408. 10.1016/B978-0-444-64076-5.00026-0 [DOI] [PubMed] [Google Scholar]
- 3. Meijer E, Casanova M, Kim H et al. Economic costs of dementia in 11 countries in Europe: estimates from nationally representative cohorts of a panel study. The Lancet Regional Health - Europe. 2022;20:100445. 10.1016/j.lanepe.2022.100445 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Alzheimer's Disease International, Evans-Lacko S, Aguzzoli E et al. World Alzheimer Report 2024: Global changes in attitudes to dementia. 2024; London, UK: Alzheimer's Disease International. https://www.alzint.org/resource/world-alzheimer-report-2024/ (30-09-2026). [Google Scholar]
- 5. Guo T, Zhang D, Zeng Y et al. Molecular and cellular mechanisms underlying the pathogenesis of Alzheimer’s disease. Mol Neurodegeneration. 2020;15:40. 10.1186/s13024-020-00391-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Lo M-T, Kauppi K, Fan C-C et al. Identification of Genetic Heterogeneity of Alzheimer’s Disease across Age. Neurobiol Aging. 2019;84:243.e1–e9. 10.1016/j.neurobiolaging.2019.02.022 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. He Y, Lu W, Zhou X et al. Unraveling Alzheimer’s disease: insights from single-cell sequencing and spatial transcriptomic. Front Neurol. 2024;15:1515981. 10.3389/fneur.2024.1515981 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Ma Y, Xia Y, Karako K et al. Decoding Alzheimer’s disease: single-Cell sequencing uncovers Brain Cell Heterogeneity and Pathogenesis. Mol Neurobiol. 2025;62:14459–73. 10.1007/s12035-025-04997-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Wang C, Acosta D, McNutt M et al. A single-cell and spatial RNA-seq database for Alzheimer’s disease (ssREAD). Nat Commun. 2024;15:4710. 10.1038/s41467-024-49133-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Jin W, Pei J, Roy JR et al. Comprehensive review on single-cell RNA sequencing: a new frontier in Alzheimer’s disease research. Ageing Res Rev. 2024;100:102454. 10.1016/j.arr.2024.102454 [DOI] [PubMed] [Google Scholar]
- 11. Heumos L, Schaar AC, Lance C et al. Best practices for single-cell analysis across modalities. Nat Rev Genet. 2023;24:550–72. 10.1038/s41576-023-00586-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Chen Y, Li Z, Ge X et al. Identification of novel hub genes for Alzheimer’s disease associated with the hippocampus using WGCNA and differential gene analysis. Front Neurosci. 2024;18:1359631. 10.3389/fnins.2024.1359631 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Kim H, Choi H, Lee D et al. A review on gene regulatory network reconstruction algorithms based on single cell RNA sequencing. Genes Genom. 2023;46:1–11. 10.1007/s13258-023-01473-8 [DOI] [PubMed] [Google Scholar]
- 14. Gupta C, Xu J, Jin T et al. Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer’s disease. PLoS Comput Biol. 2022;18:e1010287. 10.1371/journal.pcbi.1010287 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Oubounyt M, Adlung L, Patroni F et al. Inference of differential key regulatory networks and mechanistic drug repurposing candidates from scRNA-seq data with SCANet. Bioinformatics. 2023;39:btad644. 10.1093/bioinformatics/btad644 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Cui H, Wang C, Maan H et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat Methods. 2024;21:1470–80. 10.1038/s41592-024-02201-0 [DOI] [PubMed] [Google Scholar]
- 17. Kedzierska KZ, Crawford L, Amini AP et al. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol. 2025;26:101. 10.1186/s13059-025-03574-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Cheng J, Li J, Yang K et al. GenoHoption: bridging gene network graphs and single-cell foundation models. In: 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). Lisbon, Portugal: IEEE. 2024. pp. 1453–1456. 10.1109/BIBM62325.2024.10822153 [DOI] [Google Scholar]
- 19. Rossner T, Li Z, Balke J et al. Integrating single-cell foundation models with graph neural networks for drug response prediction. arXiv, arXiv:2504.14361, 13 May 2025, preprint: not peer reviewed.
- 20. Álvarez-Pérez A, Prieto-Santamaría L, Casas AI et al. Navigating the computational landscape for drug repurposing. Annu Rev Pharmacol Toxicol. 2026;66:149–70. 10.1146/annurev-pharmtox-121924-042636 [DOI] [PubMed] [Google Scholar]
- 21. Gabitto MI, Travaglini KJ, Rachleff VM et al. Integrated multimodal cell atlas of Alzheimer’s disease. Nat Neurosci. 2024;27:2366–83. 10.1038/s41593-024-01774-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Megill C, Martin B, Weaver C et al. cellxgene: a performant, scalable exploration platform for high dimensional sparse matrices. bioRxiv, 6 April 2021, preprint: not peer reviewed. 10.1101/2021.04.05.438318. [DOI]
- 23. Justo AFO, Paes VR, Leite REP et al. Evaluating Amyloid Pathology and Cognitive Outcomes in AD: insights from CERAD and Thal Staging. Alzheimer's & Dementia. 2025;21:e100810. 10.1002/alz70856_100810 [DOI] [Google Scholar]
- 24. Aibar S, González-Blas CB, Aerts S, RcisTarget: identify transcription factor binding motifs enriched on a list of genes or genomic regions.Bioconductor R package. 10.18129/B9.bioc.RcisTarget [DOI]
- 25. Aibar S, González-Blas CB, Moerman T et al. SCENIC: single-cell regulatory network inference and clustering. Nat Methods. 2017;14:1083–6. 10.1038/nmeth.4463 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Buniello A, Suveges D, Cruz-Castillo C et al. Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Res. 2025;53:D1467–75. 10.1093/nar/gkae1128 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Davis AP, Wiegers TC, Johnson RJ et al. Comparative Toxicogenomics Database (CTD): update 2023. Nucleic Acids Res. 2023;51:D1257–62. 10.1093/nar/gkac833 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Andreatta M, Carmona SJ. UCell: robust and scalable single-cell gene signature scoring. Comput Struct Biotechnol J. 2021;19:3796–8. 10.1016/j.csbj.2021.06.043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Yao Q, Long C, Yi P et al. C/ebp β : a transcription factor associated with the irreversible progression of Alzheimer’s disease. CNS Neurosci Ther. 2024;30:e14721. 10.1111/cns.14721 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Ndoja A, Reja R, Lee S-H et al. Ubiquitin ligase COP1 suppresses neuroinflammation by degrading c/EBPβ in microglia. Cell. 2020;182:1156–1169.e12. 10.1016/j.cell.2020.07.011 [DOI] [PubMed] [Google Scholar]
- 31. Wang Z-H, Gong K, Liu X et al. C/EBPβ regulates delta-secretase expression and mediates pathogenesis in mouse models of Alzheimer’s disease. Nat Commun. 2018;9:1784. 10.1038/s41467-018-04120-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32. Yao Y, Kang SS, Xia Y et al. A delta-secretase-truncated APP fragment activates CEBPB, mediating Alzheimer’s disease pathologies. Brain. 2021;144:1833–52. 10.1093/brain/awab062 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Moreno JA, Dudchenko O, Feigin CY et al. Emx2 underlies the development and evolution of marsupial gliding membranes. Nature. 2024;629:127–35. 10.1038/s41586-024-07305-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Gangemi RMR, Daga A, Marubbi D et al. Emx2 in adult neural precursor cells. Mech Dev. 2001;109:323–9. 10.1016/S0925-4773(01)00546-9 [DOI] [PubMed] [Google Scholar]
- 35. Quillen D, Hughes TM, Craft S et al. Levels of soluble interleukin 6 receptor and Asp358Ala are associated with cognitive performance and Alzheimer disease biomarkers. Neurol Neuroimmunol Neuroinflamm. 2023;10:e200095. 10.1212/NXI.0000000000200095 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. The Alzheimer’s Disease Neuroimaging Initiative, The CHARGE Consortium, EPIGEN, IMAGEN, SYS, Hibar DP, Stein JL, Renteria ME et al. Common genetic variants influence human subcortical brain structures. Nature. 2015;520:224–9. 10.1038/nature14101 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Lyra E Silva NM, Gonçalves RA, Pascoal TA et al. Pro-inflammatory interleukin-6 signaling links cognitive impairments and peripheral metabolic alterations in Alzheimer’s disease. Transl Psychiatry. 2021;11:251. 10.1038/s41398-021-01349-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Leng K, Li E, Eser R et al. Molecular characterization of selectively vulnerable neurons in Alzheimer’s disease. Nat Neurosci. 2021;24:276–87. 10.1038/s41593-020-00764-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The data and computer code produced in this study are available in the following databases:
Gabitto MI, Travaglini KJ, Rachleff VM, Kaplan ES, Long B, Ariza J, et al. Integrated multimodal cell atlas of Alzheimer’s disease. Nat Neurosci 2024;27:2366–83 (https://doi.org/10.1038/s41593-024-01774-5).
CELLxGENE dataset for SEA-AD project: https://cellxgene.cziscience.com/collections/1ca90a2d-2943-483d-b678-b809bf464c30.
The code supporting the results and the datasets generated during and/or analysed during the current study are available in the: https://medal.ctb.upm.es/internal/gitlab/disnet/network-medicine/network-medicine-and-scgpt-in-alzheimer-single-cell. Access to peer review can be provided upon request. The repository will be made public upon formal publication.
Supplementary Data are available at NAR Genomics and Bioinformatics Online.
