Skip to main content
Stem Cell Reports logoLink to Stem Cell Reports
. 2025 Aug 21;20(9):102606. doi: 10.1016/j.stemcr.2025.102606

Computationally resolved neuroprogenitor cell biomarkers associate with human disorders

Gerarda Cappuccio 1,2,9, William T Choi 1,2,3,9, Fatih Semerci 1,2,3,9, Jill A Rosenfeld 4,5, Toni Claire Tacorda 2,6, Guantong Qi 3, Anthony W Zoghbi 7, Yi Zhong 1,2, Hu Chen 1,2, Pengfei Liu 4,5, Zhandong Liu 1,2,, Mirjana Maletić-Savatić 1,2,8,10,∗∗
PMCID: PMC12447325  PMID: 40845852

Summary

Adult hippocampal neurogenesis, the process of generating new neurons, relies on a rare population of neural stem and progenitor cells (NPCs) within the dentate gyrus complex microenvironment. Discovering the specific genes that define these cells is vital yet challenging due to overlapping expression patterns, limiting detection of rare cell populations using traditional approaches. By employing the computational digital sorting algorithm (DSA) that deconvolves complex gene expression data based on pattern recognition, we identified 129 genes enriched in murine NPCs. We validated these genes against published single-cell RNA sequencing (scRNA-seq) data and discovered that 25 human orthologs were known to cause Mendelian neurological conditions. In addition, leveraging a variety of computational tools and clinical and population databases, we identified 15 genes bearing novel damaging variants linked to neurological phenotypes, suggesting their potential role in contributing to human phenotypes. These discoveries illuminate NPC molecular underpinnings and underscore their relevance to both brain development and disease.

Keywords: dentate gyrus, neural stem cell, neuroprogenitors, neurogenesis, digital sorting, gene discovery, Mendelian diseases, undiagnosed diseases

Highlights

  • Digital sorting identified 129 genes enriched in neuroprogenitors

  • Fifteen genes were linked to novel neurological conditions


This study of Maletic-Savatic and colleagues revealed novel neural stem/progenitor genes potentially implicated in brain function and disease.

Introduction

The brain’s cellular diversity, coupled with its complex functions, poses significant challenges for identifying unique molecular signatures of specific cell types, especially rare populations such as neural stem and progenitor cells (collectively named NPCs here) (Guo et al., 2022). NPCs play a pivotal role in adult neurogenesis (Barde et al., 2025; Wu et al., 2025), yet their identification is hindered by overlapping gene expression patterns with other cell types. For example, radial neural stem cells (type I) are identified by SOX2 expression in the nucleus and GFAP in the radial process. Yet, both markers are also expressed in astrocytes, requiring expertise to disambiguate these cells using immunohistochemical staining. Although single-cell RNA sequencing (scRNA-seq) (Wang et al., 2025) has revolutionized our ability to profile transcriptomes at the single-cell level (Voineagu et al., 2011), its sensitivity is often insufficient to capture rare cell populations fully, particularly with overlapping markers, leaving gaps in our understanding of NPC-specific biology (Kolodziejczyk et al., 2015).

Recently, computational approaches, such as the semi-supervised digital sorting algorithm (DSA), have been developed to dissect heterogeneous transcriptome data (Zhong et al., 2013a; Zhong and Liu, 2011). Unlike conventional methods that require extensive a priori knowledge, DSA leverages statistical learning to identify unique gene expression patterns within complex datasets. Compared to supervised deconvolution algorithms, DSA does not require the entire transcriptome profiles for individual cell types as the input but only the knowledge of some cell-specific genes. We speculated that the versatility of this approach would make it suitable for discerning the different gene expression profiles of rare cell types within a given tissue such as the dentate gyrus (DG). As the site of adult neurogenesis, dentate gyrus complexity is augmented by the presence of a sparse population of NPCs and their progeny, compared to other regions of the brain (Chang and Hen, 2024; Gage, 2025; Li et al., 2024; Lucassen et al., 2020; Tosun et al., 2019). Thus, the challenge to disambiguate such tissue complexity in the context of different cell types and to identify genes specific to a very small population of cells is substantial.

Using DSA, we deconstructed the transcriptome of the murine dentate gyrus and identified 129 genes putatively enriched in NPCs. The accuracy of these findings was corroborated through multiple validation methods, including scRNA-seq and in situ hybridization data. Furthermore, we mapped these genes to their human orthologs and explored their potential involvement in neurological conditions. Among the identified genes, 25 were associated with neurological Mendelian disorders, while 15 others were linked to neurological phenotypes without prior disease associations. This dual-pronged approach not only sheds light on the molecular architecture of NPCs but also provides a valuable resource for studying the links between neural stem cell biology and human disorders.

Results

DSA identifies candidate genes enriched in NPCs

Gene expression deconvolution is a challenging problem in transcriptome analysis and many methods have been developed to address it (Abbas et al., 2009; Gong et al., 2011; Liebner et al., 2014; Newman et al., 2015; Qiao et al., 2012; Shen-Orr and Gaujoux, 2013; Zhong et al., 2013a; Zuckerman et al., 2013). However, all these methods require a significant amount of prior knowledge of the underlying cell types, except for the DSA (Zhong et al., 2013a; Zhong and Liu, 2011). Indeed, in a recent comparative analysis (Newman et al., 2015), DSA has been demonstrated to be a reliable approach for gene expression deconvolution. To identify genes specific to cell types present in the dentate gyrus, particularly those that are rare and with limited prior knowledge (Tosun et al., 2019), we deconvoluted gene expression profiles obtained from the dentate gyri of mice aged between 4 and 12 weeks (Figures 1A–1C). The DSA deconvolution requires two inputs: (1) the data matrix of the mixture, and (2) the marker/cell-group assignments to aid the deconvolution (Figure 1A). In our datasets, the cell types are present in different proportions depending on the age of the mouse. Namely, while most postmitotic cell numbers remain relatively constant during the age period investigated (4- to 12-week-old mice), the number of NPCs declines (Beccari et al., 2017; Kuhn et al., 1996; 2005; Semerci et al., 2017). Namely, the NPC number drops from 49,571 ± 3,948 at 3 weeks of age to 29,937 ± 2,132 at 8 weeks of age and further down to 11,097 ± 1,033 at 20 weeks of age, after which it plateaus (Encinas et al., 2011). Given this dramatic decline in NPC numbers during young adulthood, we selected five time points for our study, namely 4, 6, 8, 10, and 12 weeks of age (n= 10 mice per age group). For the marker/cell-group assignments, we focused on three different cell populations in the dentate gyrus: (1) NPCs (Battiste et al., 2007; Fukuda et al., 2003; Kim et al., 2007; Mignone et al., 2004; Nacher et al., 2005), (2) immature neurons (INs) (Duan et al., 2007; Overstreet et al., 2004; Rogers et al., 2013), and (3) a group of mature cell types we termed OMEGA (Oligodendrocytes, Microglia, Endothelial cells, Granule cells, and Astrocytes) (Albelda et al., 1990; Hachem et al., 2005; Imai et al., 1996; Levine et al., 1993; Miettinen et al., 1992). Each cell type could be identified by at least one genetic marker: NPCs—Nes, Pax6, and Ascl1; INs—Sox3, Pomc, and Disc1; and OMEGA—Cspg4, Aif1, Pecam1, Calb1, and S100β (Semerci et al., 2021; Semerci and Maletic-Savatic, 2016). The specificity of these seed genes for a given cell type has been well established. In addition, we purposefully did not use genes that are shared between different groups: for example, GFAP is a well-known marker of both neural stem cells (NSCs) and astrocytes. In our study, NSCs were part of the NPC group, while astrocytes were part of the OMEGA group. Thus, we did not choose GFAP as a seed gene as this would confound data interpretation.

Figure 1.

Figure 1

DSA deconvolutes transcriptome data from the mouse dentate gyrus into cell-specific gene profiles

(A) The DSA algorithm is organized into two categories: input and output. The input data to DSA are the expression profiles of dissected dentate gyri from wild-type C56BL/6J mice of various ages (4–12 weeks old, n = 10 mice per time point, two replicas). The input also includes user-defined cell groups and their respective markers (seed genes). The proportion of neural stem and progenitor cells (NPCs) and immature neurons (INs = neuroblasts [NBs] and immature neurons [INs]) decreases over time, while the proportion of OMEGA, a group of mature postmitotic cell types (oligodendrocytes, microglia, endothelial cells, granule cells, and astrocytes) does not change. DSA is based on the theory of the mixing process described as a linear model: O = S × W. O is an n × p matrix where n is the number of genes and p is the number of samples. S is an n × k matrix where n is the number of genes and k is the number of cell groups. W is a k × p matrix where k is the number of cell groups and p is the number of samples. DSA deconvolves a data matrix of gene expression profiles from the observed tissue (O) into a matrix representing gene profiles of individual cell types or sources (S) and a matrix representing the cellular proportions or weights (W). The DSA deconvolution produces two outputs: estimated gene profiles for each cell group and the estimated cellular proportions.

(B) Based on the proportions of the cell groups, DSA calculates the gene expression signature of each cell group. DSA-derived expression profiles of 25,556 genes are plotted as points in a 3D scattered plot.

(C) 2D scattered plots correspond to the gene expressions shown in the 3D plot. DSA-derived expression is transformed in log2 space. The color gradient of each point represents the relative enrichment of the corresponding gene in the three groups of cell types. Green, blue, and red colors correspond to NPCs, INs, and OMEGA cell groups, respectively. Bright red points represent the genes enriched in NPCs based on the DSA deconvolution.

(D) Heatmap shows the stability of DSA results. In the probe stabilization analysis, algorithmic resampling of the 80% of datasets identified 145 genes enriched in the NPCs, 159 genes enriched in INs, and 151 genes enriched in the OMEGA group. t test with Bonferroni correction (p < 0.001) sets the threshold for stabilization. The stabilization of the genes for all iterations is represented as a row Z score heatmap.

(E) The histograms represent the correlation of the stable genes with the DSA weights for each cell group and the expression levels from the tissue samples. Genes enriched in NPCs (green), INs (blue), and OMEGAs (red) highly correlate with the DSA-derived weights from the cell-group markers.

From the inputs, we calculated cellular proportions and relative gene expression levels. Overall, we assessed the relative expression of 25,556 genes in three cell groups (Figures 1B and 1C). Each gene was given a value for the expression level in each of the three cell groups. The quadratic programming calculation used for the deconvolution scaled the relative expression for each gene between 2 and 14 (low and high levels of expression, respectively). The specificity for a given gene was determined by the relative expression calculated by DSA. For example, the gene for Nes, a marker of NPCs, had relative expression values of 10.65, 2.50, and 3.47 for the NPC, IN, and OMEGA cell groups, respectively, and was thus assigned as specific for NPCs. Similar specificity was obtained for all other markers assigned to the respective cell types: Sox3, Pomc, and Disc1 had the highest expression values in INs, and Cspg4, Aif1, Pecam1, Calb1, and S100β had the highest expression values in the OMEGA group. These findings indicated that DSA can accurately identify the three groups of cell types we set out to investigate.

However, every gene in the dataset has a degree of relative enrichment among the three cell groups. Thus, it is very important to better determine which genes are specific for a given cell group. Toward this, we performed stabilization analysis, which aims to find genes that give stable and reproducible results when DSA is applied to randomly sampled transcriptome datasets (Figure 1D). The stabilization analysis consists of applying the DSA algorithm to only 80% of the dataset randomly. Each random sampling will provide a certain set of results. We randomly sampled the dataset 50 times and ran the DSA algorithm, thus providing 50 different sets of results of relative gene expressions for each of the three groups. Genes that have similar results across the 50 re-samplings are considered more stable than those with variable results. The variability was determined by pairwise t test where we compared the results of all 50 re-samples from one cell group to another cell group. To limit family-wise error rates, we adjusted the p values by Bonferroni correction (p < 0.001) and discarded probes above a certain noise threshold, determined by the static signal from the dataset. Cell groups that are “stable” have better p value than those with variable results. Out of 25,556 genes (including 11 seed genes), the algorithm identified those associated with the change in expression of the seed genes and assigned 145 genes as enriched in NPCs, 159 genes as enriched in INs, and 151 genes as enriched in the OMEGA group (Figure 1D).

To further ascertain the specificity of our results, we performed a Pearson’s correlation on the genes enriched in each cell group. The Pearson’s correlation was done on the whole set of time-series data (n = 10 datasets). The genes enriched in NPCs were expected to correlate with the assigned NPC markers (Nes, Pax6, and Ascl1) and not IN (Sox3, Pomc, and Disc1) or OMEGA (Cspg4, Aif1, Pecam1, Calb1, and S100β) markers. Indeed, the correlation analysis demonstrated that the distribution of genes identified by DSA as enriched in NPCs had 90% correlation with markers assigned to the NPC group (Figure 1E). Conversely, the distribution of genes identified by DSA as enriched in IN and OMEGA groups had minimal or negative correlation with the NPC markers (Figure 1E). Similar distributions were noted when gene candidates were interrogated with respect to correlations with IN or OMEGA genes. These data strongly suggest that the computational algorithm is specific in identifying genes enriched in each cell group.

Validation of NPC markers identified by DSA

We then turned to biological validation of the gene discoveries, focusing on the NPCs. The 145 identified genes represent 129 unique mouse genes. As NPCs are localized in the subgranular zone (SGZ) of the dentate gyrus, genes putatively enriched in NPCs should thus be found in the SGZ. To determine the spatial expression of the computationally discovered NPC genes, we utilized the Allen Brain Atlas mouse in situ hybridization database, which contains mRNA expression patterns for approximately 20,000 genes. We based our result on the intensity of the in situ probes present in the SGZ compared to the intensity in the granular zone of the dentate gyrus (Figure S1). In Figure 2A, we highlighted 11 genes that showed signals exclusively in the SGZ, suggesting correlation with the DSA-identified NPC genes. Furthermore, we performed quantitative PCR analysis of the selected genes in NPCs obtained from the dentate gyri of 2-month-old Nes-GFP mice (Mignone et al., 2004) isolated via florescence-assisted cell sorting (FACS; n = 3 samples of FACS runs with >95% cell viability; each sample is generated from n = 3–4 mice). Nes-GFP (+) cells served as the NPC population, while Nes-GFP (−) cells represented other cell types in the dentate gyrus. DSA-derived NPC genes showed higher gene expression in Nes-GFP (+) cells than in Nes-GFP (−) cells (Figure 2B), confirming the discovery of new genes selectively expressed in NPCs.

Figure 2.

Figure 2

DSA identifies 129 mouse genes putatively enriched in NPCs

(A) Allen Brain Atlas mouse in situ hybridization datasets were used for spatial analysis of the transcript expression. Hybridization signals in the SGZ support enrichment of the corresponding genes in the NPC group. Histologic representation of the dentate gyrus is shown in the top left box. M, molecular layer; G, granule cell layer; H, hilus; and SGZ, subgranular zone (arrow).

(B) Relative expression of genes from (A) confirms enrichment in NPCs. Fluorescence-assisted cell sorting (FACS) isolated cells from the dentate gyrus of 2-month-old Nes-GFP mice (Nes-GFP [+] contains NPCs and Nes-GFP [−] contains other cell types in the dentate gyrus). Relative expression of each gene is normalized to the expression of Ppia. Bars represent mean ± SEM of samples from three independent FACS-derived cell populations, together consisting of 16 dentate gyri. Student’s t test of p < 0.01 for all genes except for Sox4. Each reaction qPCR reaction was performed in triplicates.

(C) NPC genes identified by DSA relate to neurological phenotypes, as determined by the mammalian phenotype expression network. A subset of NPC genes identified by DSA share a similar phenotype based on Genemania analysis. Left: each node represents a gene, and the edges connect the two nodes that share a phenotype. 17/23 nodes (black) indicate genes that have an enriched neurological phenotype while 6/23 nodes (gray) do not. Right: neurological phenotype enrichment terms based on MemPhe are calculated by −log10 of the p value. Red line indicates the threshold of enrichment (<1 indicates no enrichment).

Furthermore, we interrogated Mouse Genome Informatics (MGI), an online database that centralizes information on the genetics, genomics, and biological phenotypes of the laboratory mouse (Mus musculus), particularly focusing on phenotypes associated with genetic mutations and their implications for biology and disease. MGI indicated that 23 of the DSA-derived genes were involved in known mouse pathology (Figure 2C), and specifically, 17 genes were associated with various neurological abnormalities affecting brain size, neuronal development, and glial development. The other 6 genes were associated with cardiac and bone development abnormalities.

We next sought to determine whether the NPC candidate genes from the DSA deconvolution were related to each other. Using Genemania (Warde-Farley et al., 2010), we found that 125 of the 129 mouse NPC genes tend to have coordinated expression and can be grouped into four different modules (Figure 3A). Based on gene ontology-biological processes terms, modules 1 and 4 were enriched in genes related to cell fate while modules 2 and 3 were related to cell cycle and cell division (Figure 3B). Interestingly, module 1 contained genes involved in myelin formation (Ugt8a, Cldn11, Oligo2, and Gjc3), implying that NPCs may differentiate into oligodendrocytes, while module 4 was enriched for genes involved in both neuronal and glial differentiation (Ascl1, Bmp4, Eomes, Dpysl5, Neurod4, L1cam, Olig2, Dcx, and Cd24a). Not surprisingly, modules 2 and 3 genes enriched in the NPC group—Ncaph, Kif11, Cks2, Tpx2, Bub1, Kif18a, Nusap1, Racgap1, and Ccna2—are involved in mitosis and cell cycle. Interestingly, these genes may reflect symmetric and/or asymmetric division of type-1 NSCs (Bonaguidi et al., 2011; Encinas et al., 2011). Thus, the genes discovered through digital sorting not only confirm existing data but also reveal potentially new cell cycle and cell fate determinants that could open new avenues for research in neurogenesis.

Figure 3.

Figure 3

NPC genes identified by DSA relate to cell cycle and cell fate functions

(A) Coordinated gene expression network modules of DSA-derived NPC genes, listed according to their respective modules.

(B) Coordinated gene expression network shows that 125/129 NPC genes identified by DSA share a similar expression pattern base, delineated by four modules. Detailed enrichment terms for each module are calculated by –log10 of the p value. Red line indicates the threshold of enrichment (<1 indicates no enrichment).

Finally, we conducted an extensive comparison of a published mouse dentate gyrus scRNA-seq dataset (Hochgerner et al., 2018) and the DSA biomarker genes categorized into the four modules (Figure 4A). Our analysis unveiled significant enrichment of DSA genes (modules 2, 3, and 4) among intermediate progenitor clusters (intermediate progenitor cells [IPCs] and perinatal neuronal IPCs), which are progenitors directly derived from NSCs. This indicates that our predicted NPC markers show concordance with the public dentate gyrus single-cell dataset (Figure 4B).

Figure 4.

Figure 4

Progenitor clusters from mouse dentate gyrus scRNA-seq are enriched in NPC genes identified by DSA

(A) UMAP enrichment scores. scRNA-seq data from GSE104323 are visualized by UMAP dimension reduction with cell annotations labeled (left). Enrichment scores of the 4 different modules are visualized using the UMAP embeddings (right).

(B) Boxplot for module enrichment score in different cell types identified in scRNA-seq performed in mice dentate gyrus (nIPCs, neuronal intermediate progenitor cells; perin nIPC, perinatal nIPC; VLMC, vascular and leptomeningeal cell; PVM, perivascular macrophage; MOLs, mature oligodendrocytes; OL, oligodendrocyte; NFOL, newly formed oligodendrocyte; OPC, oligodendrocyte precursor cell) showed overlap with genes listed in modules 2, 3, and 4. Boxplots show median value, the interquartile range (25–75th percentile), and whiskers to the most extreme points within the 1.5x interquartale range.

Collectively, these analyses indicate that the NPC-related genes identified by DSA are closely associated with the NPC biological features. Overall, our study introduces a novel set of genes potentially specific to NPCs, warranting further investigation into their biological functions and potential as targets for manipulating NPC biology.

Computationally identified NPC genes are associated with human neurological disorders

We next sought to determine whether the 129 mouse genes associated with NPCs are relevant to human health. 120 of the Mus musculus genes have human orthologs. All but four genes (116/120) are expressed in the human brain, suggesting that the putative NPC genes are evolutionarily conserved between mice and humans (p < 0.0001), as previously reported for rodents and non-human primates (Miller et al., 2013).

We examined the probability of pathogenicity associated to monoallelic and/or biallelic models for all the 120 genes. Computationally, genes highly likely associated to a monoallelic condition were TLE3, TET1, CCNA2, and ASCL1, while those potentially implicated to a biallelic model of disease were KIF18A, TDG, SLC17A6, and ARHGAP11A (Table S1).

To determine whether human orthologs have known associations to human disease, we first analyzed the Online Mendelian Inheritance in Man (OMIM) database (McKusick, 2007). 32 out of 120 genes were annotated in OMIM and, while only 7 genes were linked to non-neurological phenotypes, 25 genes have been reported to cause monogenic neurological disorders (Wang et al., 2017) (Figures 5A and 5B; Table S2), significantly different from 50:50 distribution (the chi-square 5.8887, p < 0.05). This leaves 88 genes that have not yet been associated with any documented human phenotype. To scrutinize these genes, we used population datasets such as the Database of Genomic Variants and Decipher (Table S3), which collect controls and affected individuals bearing copy number variation (CNV) and/or single-nucleotide variants (SNVs). We also used genomics metrics to assess the functional impact and tolerance of genetic variants within the selected genes such as loss-of-function intolerance (pLI) (measures how intolerant a gene is to loss-of-function mutations), missense Z score (evaluates the observed versus expected missense mutations in a gene), and Loss-of-function Observed/Expected Upper bound Fraction (LOEUF) score (quantifies the burden of missense mutations in a gene, indicating depletion or enrichment) (Table S3). Out of 88 genes, 19 genes were selected due to their potential involvement in novel Mendelian conditions, based on (1) the absence or very low frequency of benign CNVs, (2) the presence of CNVs encompassing specific gene in patients with reported neurological phenotype, and (3) the high probability of the pLI and/or missense and LOEUF score from GnomAD (v.2.1.1) (Gudmundsson et al., 2022). The prioritized 19 genes shared constraint metrics, supporting greater intolerance to the loss-of-function and missense variants (Figure 5A; Table S4).

Figure 5.

Figure 5

DSA-derived NPC genes may lead to human neurological disorders

(A) Filtering pipeline applied to the computationally identified genes. 25 human orthologs (out of 120) are known to cause neurological Mendelian phenotypes, while 88 genes are not listed in OMIM.

(B) Out of 120 identified genes, 25 (20.83%) are linked to monogenic neurological disorders, 7 (5.83%) to non-neurological conditions, and 88 (73.33%) have not been associated with any known human phenotype. The distribution of neurological vs. non-neurological associations is significantly different from a 50:50 ratio (p < 0.05).

(C) 15 candidate genes identified by DSA to be selectively enriched in NPCs showed variants contributing to enriched neurological phenotypes in patients. See Tables S6A and S6B for additional data.

(D) Dot plot shows the expression profiles of the 15 human orthologs in the subgranular zone (SGZ) of the human dentate gyrus, based on spatial transcriptomics data (Ramnauth et al., 2025).

SNVs and CNVs encompassing selected NPCs genes may contribute to neurological disorders

We subsequently investigated whether any of the 19 selected NPCs genes, not associated with Mendelian diseases, bear variants potentially contributing to human neurological disorders. This inquiry leveraged large clinical datasets incorporating exome sequencing (ES) and genome sequencing (GS) technologies that have revolutionized the diagnostic approach to undiagnosed rare and ultra-rare Mendelian diseases. These technologies have facilitated a vast number of diagnoses, changes in medical management, new treatments, and the discovery of novel disease genes. ES has led to an exponential increase in gene discovery and our understanding of how disease develops and manifests phenotypically. It primarily detects SNVs and small insertions/deletions. In contrast, microarray-based comparative genomic hybridization (aCGH) is effective in identifying CNVs, including intragenic exonic deletions and duplications (Wojcik et al., 2023).

To identify undiagnosed individuals with highly prevalent neurological phenotypes who underwent exome/genome and/or aCGH study, and who carried rare, potentially damaging variants (SNVs or CNVs), we interrogated large clinical databases (more than 40,000 patients) containing ES and GS data as well as aCGH data in unsolved patients at Baylor Genetics Laboratories and through the online tool Decipher (Figure 5A) (Tables S3, S6A, and S6B). For SNVs, we conducted high-quality variant interpretation using a comprehensive approach. This included analysis of population and clinical databases (gnomAD, ClinVar), bioinformatic predictions of functional impact, examination of family segregation patterns, and a detailed study of associated phenotypes. The initial filtering step begins by excluding variants commonly observed in population databases, as these are unlikely to cause rare diseases. Specifically, we focused on variants absent or at an extremely low frequency in controls in the gnomAD database (allele frequency less than <4.0 E−05) (Choi et al., 2024). Variants were deemed potentially damaging only when multiple bioinformatic prediction tools consistently indicated their harmful effects. In particular, when (1) Combined Annotation-Dependent Depletion (CADD) scored 20 or greater (the larger the score, the more likely the SNP has a damaging effect), (2) the Sorting Intolerant From Tolerant (SIFT) score was low (the lower the score, the more deleterious), and (3) Polymorphism Phenotyping (PolyPhen) had high values (values closer to 1.0 are predicted to be deleterious). We also considered parental inheritance (if known) and selected variants in allele states consistent with the potential inheritance pattern and neurological phenotypes.

ES/aCGH eligible patients are affected by multiple congenital anomalies and/or neurocognitive disabilities. Patients presenting with primary neurological conditions, including neurodevelopmental delay, intellectual disability, epilepsy, neuromuscular disease, and cerebellar ataxia, constitute a significant portion of the clinical indications for ES, averaging around 63% across different studies (Abul-Husn et al., 2023; Arteche-López et al., 2022; Kagan et al., 2023; Slavotinek et al., 2023) (in these studies, the percentage of neurological phenotypes tested with ES was 89%, 75.6%, 56%, and 30.74%, respectively). We specifically selected genes known to contain damaging SNVs that are globally associated with a higher prevalence of neurological phenotypes, i.e., surpassing the frequency observed among individuals undergoing clinical ES/GS (63%) (Table S6A). We selected CNVs encompassing DSA-selected genes (smaller than 10 Mb) in undiagnosed patients with reported neurological phenotypes, in line with family segregation and absent from the control CNV database (Table S6B). Finally, we detected 152 SNVs and 26 CNVs in 15 genes plausibly and potentially associated to novel neurological Mendelian conditions (Figures 5A and 5C; Tables S6A and S6B). We closely examined the literature on these 15 candidate genes, all of which have been linked to neural cell functions (Table S7).

In addition, we examined whether the human orthologs of the 15 candidate genes mapped to the human SGZ neurogenic niche. While analyzing human neurogenesis in postmortem tissues presents its own set of challenges, we successfully identified all candidate genes but SFRP1 (Figure 5D) in two spatial transcriptomics datasets that specifically focused on the SGZ (Ramnauth et al., 2025; Thompson et al., 2024). Notably, six of these genes—TET1, ASCL1, SEMA6A, ZRANB1, CDH6, and SRSF6—were also significantly enriched in the radial glia cluster within the developing human brain (Figure S2) (Eze et al., 2021). Although the SGZ includes neuroblasts and INs in addition to NPCs—making precise attribution to NPCs challenging—these data underscore the notion that these genes are not mere incidental findings but rather likely play conserved roles during normal neurodevelopment and in the pathogenesis of neurodevelopmental disorders. Furthermore, these data provide a robust foundation for deeper investigation into the mechanistic contributions of these genes in neural stem/progenitor cells in both embryonic and adult neurogenesis, advancing our understanding of how they influence brain development and disease. This work thus highlights that computational analysis of tissue complexity can uncover insights relevant not only to understanding cell biology in mouse models but also to elucidating human pathology.

Discussion

In this study, we set out to discover genes specific for cell types inhabiting the dentate gyrus, focusing on the rare but heterogeneous population of NPCs. These cells are difficult to discern because of their low numbers and even more cumbersome to sort as pure populations because of the lack of specific markers (Semerci et al., 2021; Semerci and Maletic-Savatic, 2016; Tosun et al., 2019). To accomplish our goal, we resorted to a computational approach, digital sorting, which has shown robust performance for cell cultures (Voineagu et al., 2011; Wang et al., 2025). The key contributions of this work include (1) expanding the repertoire of NPC-specific genes that may serve as their biomarkers, (2) bridging biology and pathology, as we link novel putative NPC genes to human neurological conditions, (3) demonstrating the power of DSA in resolving complex tissue transcriptome, particularly for rare cell populations, and (4) providing a foundation for further exploration into the roles of hippocampal NPCs in brain development and disease.

Digital sorting approach for identifying biomarkers of rare cell populations

The discovery of NPC-specific markers has been impeded by the lack of tools capable of resolving rare cell populations within heterogeneous tissues. Traditional models, such as transgenic models (Semerci et al., 2021; Semerci and Maletic-Savatic, 2016) and scRNA-seq, have provided many insights but remain limited. For example, widely used markers like Nestin-GFP (Mignone et al., 2004) and GFAP-GFP (Brenner et al., 1994; Suh et al., 2007) exhibit expression in multiple cell types, complicating efforts to specifically isolate NPCs. Similarly, while scRNA-seq offers unparalleled resolution, it is hampered by technical variability, high cost, and difficulties in analyzing rare cell types. Moreover, careful validation of scRNA-seq findings is essential to ensure that identified biomarkers truly reflect biological differences and are not artifacts of the experimental process. Consequently, there is a continuous effort to identify genes that exhibit greater specificity to NPCs than existing markers (Manganas et al., 2021; Semerci et al., 2017). DSA addresses these limitations by employing a probabilistic approach to deconvolute transcriptomic data, providing an unbiased and scalable method to identify cell-specific molecular signatures. By focusing on genes that consistently demonstrate differential expression across individual cells, DSA facilitates discovering biologically significant NPC-specific biomarkers that may be critical for understanding human pathology.

Insights into NPC-related disorders

Given the importance of NPCs for brain development not only in utero (Abraham et al., 2013; Abraham et al., 2013) but also throughout life—as they generate new neurons in the center for learning and memory—we reasoned that NPC-specific genes, if mutated, could play a critical role in human neurological disorders. Indeed, NPC dysfunction has been implicated in a range of neurological and mental health conditions, including decreased learning and memory, Alzheimer’s disease, depression (Boldrini et al., 2012; Gandy et al., 2017; Miller and Hen, 2015), addiction (Noonan et al., 2010), and schizophrenia (Christian et al., 2010). Here, to examine whether there are any neurophenotypes that could be associated with the computationally identified NPC genes, we conducted thorough analyses of databases such as ES/GS and aCGH and identified 25 genes (78% of the DSA-identified OMIM genes) associated with Mendelian neuro-conditions. Moreover, we found potentially pathogenic variants in 15 NPC genes that are not currently recognized as causing disease by existing disease-related databases such as OMIM (McKusick, 2007). Importantly, these genes also mapped to human radial glia in the embryonic brain and the SGZ in the postnatal brain. Collectively, these multi-stage and multi-modal analyses demonstrate convergent evidence that these candidate genes are expressed in human neurogenic populations, substantiating their biological and pathological relevance.

Study limitations

An inherent limitation of using DSA for identifying cell-specific markers lies in the selection of seed markers. While the seed genes for this study were chosen for their high specificity and characteristic expression in the cell types of interest, they remain susceptible to biases introduced by biological variability, environmental influences, and technical artifacts inherent in sequencing data. These factors may lead to skewing of the initial input and subsequently affect the deconvolution process. To address this issue, diversifying marker selection by assembling an expanded, comprehensive list of genes derived from multiple independent resources and datasets can help capture biological variability and reduce potential biases. Furthermore, iterative refinement of the deconvolution process, guided by reconstruction loss, offers an additional safeguard against bias. By dynamically recalibrating markers—introducing supplementary ones or removing those contributing to higher loss—this approach ensures a more robust and accurate deconvolution outcome. Finally, cross-validating the DSA-derived signatures against independent spatial and scRNA-seq datasets is highly recommended to ensure reproducibility and biological relevance.

Utility and future directions

Although our objective in this study was to discover NPC-specific profiles using expression data of the adult mouse dentate gyrus, DSA could be extended to other high-dimensional omics datasets from a wide variety of tissues, including human postmortem tissue. A number of guiding principles should be considered when embarking on such a study. First, the selection of seed markers is very important, as it provides the foundation for the DSA-based computation of variables of interest (genes, transcripts, proteins, etc.) that are associated with those markers. Second, computational validation and experimental corroboration are essential to minimize false positives and ensure biological relevance. Third, screening for markers (i.e., genes) that identify broad cell groups rather than individual cells casts a wider net that permits identification of a large repertoire of associations with those cell groups. Such a strategy should lead to discoveries of new cell subtypes defined by specific genes. In addition, it may lead to discoveries of genes enriched in a specific cell group, which was not possible when individual cells are studied in isolation. Finally, the ability of DSA to highlight cell-specific molecular signatures also holds promise for refining the phenotypic characterization of human genetic disorders. By associating specific genes with rare cell types, DSA provides insights that are not accessible through single-cell approaches alone.

Altogether, DSA is a powerful computational tool that can disambiguate cell type-specific molecular signatures when the sources of the data are heterogeneous tissues and the cells of interest represent a rare population within the tissue. By identifying novel NPC-specific genes and uncovering their connections to human pathology, we have provided a valuable resource for future research into neurogenesis, rare cell populations, and their roles in health and disease. Our findings illustrate how computational innovation can bridge the gap between molecular discovery and clinical insights, offering new pathways to understanding and potentially treating neurological conditions.

Methods

DSA analysis

All DSA analysis was done in RStudio (v.0.98.501) from previously published source code (Zhong et al., 2013a; Zhong and Liu, 2011). Robust Multi-array Average (RMA) gene-level data were filtered to remove all probes with annotations of “NA,” leaving 25,556 probes for analysis. Probe stabilization analysis, scatterplots, and heatmaps were also performed and generated in RStudio using R packages gplots, matrixStats, RColorBrewer, rgl, and scatterplot3d. The computational complexity is O(k3+k2p) for estimating cell type frequency matrix W, and O(n3k3) for estimating cell type-specific expression profile matrix S, where k is the number of cell types, p is the number of samples, and n is the number of genes.

Comparison on scRNA-seq of mouse dentate gyrus vs. NPC genes identified by DSA

scRNA-seq data of the mouse dentate gyrus were downloaded from the NCBI Gene Expression Omnibus (GEO) repository with the accession number database: GSE104323 (Hochgerner et al., 2018). We performed data normalization, scaling, and log transformation using Seurat. Top 20 principal components were used to construct uniform manifold approximation and projection (UMAP) embeddings. The enrichment scores of the 4 modules were calculated using the AddModuleScore function in Seurat (Hao et al., 2024).

Exome/genome and aCGH database query

All studies were approved by the Baylor College of Medicine Institutional Review Board and performed in accordance with the institutional and federal guidelines (protocols H-36612, H-41191). We cross-referenced the selected gene names with the clinical database of more than 40,000 patients undergoing ES, GS, or aCGH. The ES pipeline is based on methodologies outlined in the prior publication (Liu et al., 2019; Yang et al., 2014). The whole-genome microarrays were custom-designed by BCM MGL and manufactured by Agilent Technologies as described previously (Boone et al., 2010; Wiszniewska et al., 2014).

Visualization of predicted marker genes in spatial transcriptomic datasets

Two spatial transcriptomic datasets were employed to visualize the expression patterns of our predicted marker genes (Ramnauth et al., 2025). The first dataset, generated by Ramnauth et al. (2025) was retrieved from Zenodo (Zenodo Data: https://zenodo.org/records/10126688). There, the libraries were prepared according to the Visium Spatial Gene Expression User Guide (CG000239 Rev C, 10× Genomics) and sequenced using NovaSeq 6000 System (Illumina) with a depth of a minimum of 60,000 reads per Visium spot (Ramnauth et al., 2025). The second dataset, generated by Thompson et al. (2024), was accessed via the R package humanHippocampus2024 (version 0.99.8). There, slides were processed using Visium Spatial Gene Expression (protocol CG000160, revision B, 10× Genomics) to include dentate gyrus. Libraries were loaded at 300 pM and sequenced on a NovaSeq System (Illumina) and the number of genes × the number of spots was 31,483 × 150,917 (Thompson et al., 2024).

Visualization of predicted marker genes in scRNA-seq from the developing brain

We obtained a scRNA-seq dataset (Eze et al., 2021) from the UCSC Cell Browser (https://cells-test.gi.ucsc.edu/?ds=early-brain) together with its published cell type annotations, which we used without alteration. All 45,156 cells were included in our analysis. The raw expression matrix was normalized by log transformation. Predicted marker-gene expression across annotated cell types was visualized using Seurat’s Dotplot function (v.5.0.0). To illustrate developmental progression, cell types were ordered manually along the x axis according to their maturation stage in the trajectory plot.

Statistical analysis

Statistical analysis was performed using Statistical Package for the Social Sciences (SPSS). Experiments involving 2 groups were compared using unpaired Student’s t test. Experiments involving more than 2 groups were compared by multiple-way analysis of variance, followed by post hoc analysis with least significant difference. Significance was defined as p < 0.05.

Resource availability

Lead contact

The lead contact for data/resource availability is Mirjana Maletic-Savatic (maletics@bcm.edu).

Materials availability

This study did not generate new unique reagents.

Data and code availability

The microarray source data are available at GEO repository. The accession number for this dataset reported in this paper is: GSE303501. The code and the DSA framework are accessible at the GitHub link: https://github.com/geri86Ale/DSA-framework. Our DSA source code is already available in previously published papers (Zhong et al., 2013a; Zhong and Liu, 2011), while the source code for scRNA-seq vs. DSA genes comparison is available at GitHub with the link: https://github.com/zhandong/DSA.

Acknowledgments

We thank the members of Maletic-Savatic lab for comments and critical reading of the paper. Special thanks to Drs. Juan Jose Deudero, Juan Manuel Encinas, Amanda Sierra, Pavel Stankiewicz, Sau Wai Cheung, and Penelope Bonnen for contributions to this project at its initial stages. This work was supported in part by grants from the National Institute on Aging (1R01AG076942 to M.M.-S.), the Eunice Kennedy Shriver National Institute of Child Health & Human Development of the National Institutes of Health (P50HD103555) for use of the Microscopy Core facility, and the Genomic and RNA Profiling Core at Baylor College of Medicine with the expert assistance of Lisa D. White, PhD. The content of this paper is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. In addition, support was provided by the Autism Speaks (G.C.), Simons Foundation Shenoy Undergraduate Research Fellowship in Neuroscience (SURFiN) (T.C.T.) and Cynthia and Antony Petrello Endowment (M.M.-S.). W.T.C. was supported by the National Library of Medicine Training Program in Biomedical Informatics (T15LM007093), Developmental Biology Training Program (T32HD055200), and Baylor College of Medicine Medical Scientist Training Program.

Author contributions

G.C. designed and performed genetic analysis, analyzed and interpreted the human genetic data, and wrote parts of the manuscript. W.T.C. designed and performed the bioinformatics experiments, collected, analyzed and interpreted the data, and wrote parts of the manuscript. F.S. designed and performed the mouse experiments and collected, analyzed and interpreted the data. J.A.R. collected, analyzed, and interpreted genetic data. T.C.T. participated in genetic analyses and wrote part of the manuscript. H.C. and G.Q. performed analysis of the scRNA-seq dataset. A.W.Z. analyzed human genetics data. Z.L. and M.M.-S. designed and supervised all the experiments, analyzed and interpreted the data, provided financial support, and wrote the manuscript. All authors agreed with the final version of the manuscript.

Declaration of interests

The Department of Molecular and Human Genetics at Baylor College of Medicine receives revenue from clinical genetic testing completed at Baylor Genetics Laboratories.

Published: August 21, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.stemcr.2025.102606.

Contributor Information

Zhandong Liu, Email: zhandonl@bcm.edu.

Mirjana Maletić-Savatić, Email: maletics@bcm.edu.

Supplemental information

Document S1. Figures S1, S2, and Table S7
mmc1.pdf (37.4MB, pdf)
Data S1. Tables S1–S6
mmc2.xlsx (118.3KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (53.2MB, pdf)

References

  1. Abbas A.R., Wolslegel K., Seshasayee D., Modrusan Z., Clark H.F. Deconvolution of blood microarray data identifies cellular activation patterns in systemic lupus erythematosus. PLoS One. 2009;4 doi: 10.1371/journal.pone.0006098. [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Abraham A.B., Bronstein R., Chen E.I., Koller A., Ronfani L., Maletic-Savatic M., Tsirka S.E. Members of the high mobility group B protein family are dynamically expressed in embryonic neural stem cells. Proteome Sci. 2013;11:18. doi: 10.1186/1477-5956-11-18. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Abraham A.B., Bronstein R., Reddy A.S., Maletic-Savatic M., Aguirre A., Tsirka S.E. Aberrant neural stem cell proliferation and increased adult neurogenesis in mice lacking chromatin protein HMGB2. PLoS One. 2013;8 doi: 10.1371/journal.pone.0084838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Abul-Husn N.S., Marathe P.N., Kelly N.R., Bonini K.E., Sebastin M., Odgis J.A., Abhyankar A., Brown K., Di Biase M., Gallagher K.M., et al. Molecular diagnostic yield of genome sequencing versus targeted gene panel testing in racially and ethnically diverse pediatric patients. medRxiv. 2023 doi: 10.1101/2023.03.18.23286992. Preprint at. [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Albelda S.M., Oliver P.D., Romer L.H., Buck C.A. EndoCAM: a novel endothelial cell-cell adhesion molecule. J. Cell Biol. 1990;110:1227–1237. doi: 10.1083/jcb.110.4.1227. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Arteche-López A., Ávila-Fernández A., Riveiro Álvarez R., Almoguera B., Bustamante Aragonés A., Martin-Merida I., López Martínez M.A., Giménez Pardo A., Vélez-Monsalve C., Gallego Merlo J., et al. Five years’ experience of the clinical exome sequencing in a Spanish single center. Sci. Rep. 2022;12 doi: 10.1038/s41598-022-23786-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Barde W., Renner J., Emery B., Khanzada S., Hu X., Garthe A., Rünker A.E., Amin H., Kempermann G. Beyond nature, nurture, and chance: Individual agency shapes divergent learning biographies and brain connectome. Sci. Adv. 2025;11 doi: 10.1126/sciadv.ads7297. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Battiste J., Helms A.W., Kim E.J., Savage T.K., Lagace D.C., Mandyam C.D., Eisch A.J., Miyoshi G., Johnson J.E. Ascl1 defines sequentially generated lineage-restricted neuronal and oligodendrocyte precursor cells in the spinal cord. Development (Cambridge, England) 2007;134:285–293. doi: 10.1242/dev.02727. [DOI] [PubMed] [Google Scholar]
  9. Beccari S., Valero J., Maletic-Savatic M., Sierra A. A simulation model of neuroprogenitor proliferation dynamics predicts age-related loss of hippocampal neurogenesis but not astrogenesis. Sci. Rep. 2017;7 doi: 10.1038/s41598-017-16466-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Boldrini M., Hen R., Underwood M.D., Rosoklija G.B., Dwork A.J., Mann J.J., Arango V. Hippocampal angiogenesis and progenitor cell proliferation are increased with antidepressant use in major depression. Biol. Psychiatry. 2012;72:562–571. doi: 10.1016/j.biopsych.2012.04.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Bonaguidi M.A., Wheeler M.A., Shapiro J.S., Stadel R.P., Sun G.J., Ming G.l., Song H. In vivo clonal analysis reveals self-renewing and multipotent adult neural stem cell characteristics. Cell. 2011;145:1142–1155. doi: 10.1016/j.cell.2011.05.024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Boone P.M., Bacino C.A., Shaw C.A., Eng P.A., Hixson P.M., Pursley A.N., Kang S.-H.L., Yang Y., Wiszniewska J., Nowakowska B.A., et al. Detection of clinically relevant exonic copy-number changes by array CGH. Hum. Mutat. 2010;31:1326–1342. doi: 10.1002/humu.21360. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Brenner M., Kisseberth W.C., Su Y., Besnard F., Messing A. GFAP promoter directs astrocyte-specific expression in transgenic mice. J. Neurosci. 1994;14:1030–1037. doi: 10.1523/JNEUROSCI.14-03-01030.1994. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Chang W.-L., Hen R. Adult Neurogenesis, Context Encoding, and Pattern Separation: A Pathway for Treating Overgeneralization. Adv. Neurobiol. 2024;38:163–193. doi: 10.1007/978-3-031-62983-9_10. [DOI] [PubMed] [Google Scholar]
  15. Choi W.-J., Kim S.-H., Lee S.R., Oh S.-H., Kim S.W., Shin H.Y., Park H.J. Global carrier frequency and predicted genetic prevalence of patients with pathogenic sequence variants in autosomal recessive genetic neuromuscular diseases. Sci. Rep. 2024;14:3806. doi: 10.1038/s41598-024-54413-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Christian K., Song H., Ming G.-L. Adult neurogenesis as a cellular model to study schizophrenia. Cell Cycle. 2010;9:636–637. doi: 10.4161/cc.9.4.10932. [DOI] [PubMed] [Google Scholar]
  17. Duan X., Chang J.H., Ge S., Faulkner R.L., Kim J.Y., Kitabatake Y., Liu X.b., Yang C.-H., Jordan J.D., Ma D.K., et al. Disrupted-In-Schizophrenia 1 regulates integration of newly generated neurons in the adult brain. Cell. 2007;130:1146–1158. doi: 10.1016/j.cell.2007.07.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Encinas J.M., Michurina T.V., Peunova N., Park J.-H., Tordo J., Peterson D.A., Fishell G., Koulakov A., Enikolopov G. Division-coupled astrocytic differentiation and age-related depletion of neural stem cells in the adult hippocampus. Cell Stem Cell. 2011;8:566–579. doi: 10.1016/j.stem.2011.03.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Eze U.C., Bhaduri A., Haeussler M., Nowakowski T.J., Kriegstein A.R. Single-cell atlas of early human brain development highlights heterogeneity of human neuroepithelial cells and early radial glia. Nat. Neurosci. 2021;24:584–594. doi: 10.1038/s41593-020-00794-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Fukuda S., Kato F., Tozuka Y., Yamaguchi M., Miyamoto Y., Hisatsune T. Two distinct subpopulations of nestin-positive cells in adult mouse dentate gyrus. J. Neurosci. 2003;23:9357–9366. doi: 10.1523/JNEUROSCI.23-28-09357.2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Gage F.H. Adult Neurogenesis in the Human Dentate Gyrus. Hippocampus. 2025;35 doi: 10.1002/hipo.23655. [DOI] [PubMed] [Google Scholar]
  22. Gandy K., Kim S., Sharp C., Dindo L., Maletic-Savatic M., Calarge C. Pattern Separation: A Potential Marker of Impaired Hippocampal Adult Neurogenesis in Major Depressive Disorder. Front. Neurosci. 2017;11:571. doi: 10.3389/fnins.2017.00571. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Gong T., Hartmann N., Kohane I.S., Brinkmann V., Staedtler F., Letzkus M., Bongiovanni S., Szustakowski J.D. Optimal deconvolution of transcriptional profiling data using quadratic programming with application to complex clinical blood samples. PLoS One. 2011;6 doi: 10.1371/journal.pone.0027156. [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Gudmundsson S., Singer-Berk M., Watts N.A., Phu W., Goodrich J.K., Solomonson M., Genome Aggregation Database Consortium. Rehm H.L., MacArthur D.G., O’Donnell-Luria A. Variant interpretation using population databases: Lessons from gnomAD. Hum. Mutat. 2022;43:1012–1030. doi: 10.1002/humu.24309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Guo N., McDermott K.D., Shih Y.-T., Zanga H., Ghosh D., Herber C., Meara W.R., Coleman J., Zagouras A., Wong L.P., et al. Transcriptional regulation of neural stem cell expansion in the adult hippocampus. eLife. 2022;11 doi: 10.7554/eLife.72195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Hachem S., Aguirre A., Vives V., Marks A., Gallo V., Legraverend C. Spatial and temporal expression of S100B in cells of oligodendrocyte lineage. Glia. 2005;51:81–97. doi: 10.1002/glia.20184. [DOI] [PubMed] [Google Scholar]
  27. Hao Y., Stuart T., Kowalski M.H., Choudhary S., Hoffman P., Hartman A., Srivastava A., Molla G., Madad S., Fernandez-Granda C., Satija R. Dictionary learning for integrative, multimodal and scalable single-cell analysis. Nat. Biotechnol. 2024;42:293–304. doi: 10.1038/s41587-023-01767-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Hochgerner H., Zeisel A., Lönnerberg P., Linnarsson S. Conserved properties of dentate gyrus neurogenesis across postnatal development revealed by single-cell RNA sequencing. Nat. Neurosci. 2018;21:290–299. doi: 10.1038/s41593-017-0056-2. [DOI] [PubMed] [Google Scholar]
  29. Imai Y., Ibata I., Ito D., Ohsawa K., Kohsaka S. A novel gene iba1 in the major histocompatibility complex class III region encoding an EF hand protein expressed in a monocytic lineage. Biochem. Biophys. Res. Commun. 1996;224:855–862. doi: 10.1006/bbrc.1996.1112. [DOI] [PubMed] [Google Scholar]
  30. Kagan M., Semo-Oz R., Ben Moshe Y., Atias-Varon D., Tirosh I., Stern-Zimmer M., Eliyahu A., Raas-Rothschild A., Bivas M., Shlomovitz O., et al. Clinical impact of exome sequencing in the setting of a general pediatric ward for hospitalized children with suspected genetic disorders. Front. Genet. 2022;13 doi: 10.3389/fgene.2022.1018062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Kim E.J., Leung C.T., Reed R.R., Johnson J.E. In vivo analysis of Ascl1 defined progenitors reveals distinct developmental dynamics during adult neurogenesis and gliogenesis. J. Neurosci. 2007;27:12764–12774. doi: 10.1523/JNEUROSCI.3178-07.2007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kolodziejczyk A.A., Kim J.K., Svensson V., Marioni J.C., Teichmann S.A. The technology and biology of single-cell RNA sequencing. Mol. Cell. 2015;58:610–620. doi: 10.1016/j.molcel.2015.04.005. [DOI] [PubMed] [Google Scholar]
  33. Kuhn H.G., Biebl M., Wilhelm D., Li M., Friedlander R.M., Winkler J. Increased generation of granule cells in adult Bcl-2-overexpressing mice: a role for cell death during continued hippocampal neurogenesis. Eur. J. Neurosci. 2005;22:1907–1915. doi: 10.1111/j.1460-9568.2005.04377.x. [DOI] [PubMed] [Google Scholar]
  34. Kuhn H.G., Dickinson-Anson H., Gage F.H. Neurogenesis in the dentate gyrus of the adult rat: age-related decrease of neuronal progenitor proliferation. J. Neurosci. 1996;16:2027–2033. doi: 10.1523/JNEUROSCI.16-06-02027.1996. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Levine J.M., Stincone F., Lee Y.S. Development and differentiation of glial precursor cells in the rat cerebellum. Glia. 1993;7:307–321. doi: 10.1002/glia.440070406. [DOI] [PubMed] [Google Scholar]
  36. Li Y., Tang C., Song Y. Protocol to establish a demyelinated animal model to study hippocampal neurogenesis and cognitive function in adult rodents. STAR Protoc. 2024;5 doi: 10.1016/j.xpro.2024.103242. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Liebner D.A., Huang K., Parvin J.D. MMAD: microarray microdissection with analysis of differences is a computational tool for deconvoluting cell type-specific contributions from tissue samples. Bioinformatics (Oxford, England) 2014;30:682–689. doi: 10.1093/bioinformatics/btt566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Liu P., Meng L., Normand E.A., Xia F., Song X., Ghazi A., Rosenfeld J., Magoulas P.L., Braxton A., Ward P., et al. Reanalysis of Clinical Exome Sequencing Data. N. Engl. J. Med. 2019;380:2478–2480. doi: 10.1056/NEJMc1812033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Lucassen P.J., Fitzsimons C.P., Salta E., Maletic-Savatic M. Adult neurogenesis, human after all (again): Classic, optimized, and future approaches. Behav. Brain Res. 2020;381 doi: 10.1016/j.bbr.2019.112458. [DOI] [PubMed] [Google Scholar]
  40. Manganas L.N., Durá I., Osenberg S., Semerci F., Tosun M., Mishra R., Parkitny L., Encinas J.M., Maletic-Savatic M. BASP1 labels neural stem cells in the neurogenic niches of mammalian brain. Sci. Rep. 2021;11:5546. doi: 10.1038/s41598-021-85129-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. McKusick V.A. Mendelian Inheritance in Man and its online version, OMIM. Am. J. Hum. Genet. 2007;80:588–604. doi: 10.1086/514346. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Miettinen R., Gulyás A.I., Baimbridge K.G., Jacobowitz D.M., Freund T.F. Calretinin is present in non-pyramidal cells of the rat hippocampus--II. Co-existence with other calcium binding proteins and GABA. Neuroscience. 1992;48:29–43. doi: 10.1016/0306-4522(92)90335-y. [DOI] [PubMed] [Google Scholar]
  43. Mignone J.L., Kukekov V., Chiang A.-S., Steindler D., Enikolopov G. Neural stem and progenitor cells in nestin-GFP transgenic mice. J. Comp. Neurol. 2004;469:311–324. doi: 10.1002/cne.10964. [DOI] [PubMed] [Google Scholar]
  44. Miller B.R., Hen R. The current state of the neurogenic theory of depression and anxiety. Curr. Opin. Neurobiol. 2015;30:51–58. doi: 10.1016/j.conb.2014.08.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Miller J.A., Nathanson J., Franjic D., Shim S., Dalley R.A., Shapouri S., Smith K.A., Sunkin S.M., Bernard A., Bennett J.L., et al. Conserved molecular signatures of neurogenesis in the hippocampal subgranular zone of rodents and primates. Development (Cambridge, England) 2013;140:4633–4644. doi: 10.1242/dev.097212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Nacher J., Varea E., Blasco-Ibañez J.M., Castillo-Gomez E., Crespo C., Martinez-Guijarro F.J., McEwen B.S. Expression of the transcription factor Pax 6 in the adult rat dentate gyrus. J. Neurosci. Res. 2005;81:753–761. doi: 10.1002/jnr.20596. [DOI] [PubMed] [Google Scholar]
  47. Newman A.M., Liu C.L., Green M.R., Gentles A.J., Feng W., Xu Y., Hoang C.D., Diehn M., Alizadeh A.A. Robust enumeration of cell subsets from tissue expression profiles. Nat. Methods. 2015;12:453–457. doi: 10.1038/nmeth.3337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Noonan M.A., Bulin S.E., Fuller D.C., Eisch A.J. Reduction of adult hippocampal neurogenesis confers vulnerability in an animal model of cocaine addiction. J. Neurosci. 2010;30:304–315. doi: 10.1523/JNEUROSCI.4256-09.2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Overstreet L.S., Hentges S.T., Bumaschny V.F., de Souza F.S.J., Smart J.L., Santangelo A.M., Low M.J., Westbrook G.L., Rubinstein M. A transgenic marker for newly born granule cells in dentate gyrus. J. Neurosci. 2004;24:3251–3259. doi: 10.1523/JNEUROSCI.5173-03.2004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Qiao W., Quon G., Csaszar E., Yu M., Morris Q., Zandstra P.W. PERT: a method for expression deconvolution of human blood samples from varied microenvironmental and developmental conditions. PLoS Comput. Biol. 2012;8 doi: 10.1371/journal.pcbi.1002838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Ramnauth A.D., Tippani M., Divecha H.R., Papariello A.R., Miller R.A., Nelson E.D., Thompson J.R., Pattie E.A., Kleinman J.E., Maynard K.R., et al. Spatiotemporal analysis of gene expression in the human dentate gyrus reveals age-associated changes in cellular maturation and neuroinflammation. Cell Rep. 2025;44 doi: 10.1016/j.celrep.2025.115300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Rogers N., Cheah P.-S., Szarek E., Banerjee K., Schwartz J., Thomas P. Expression of the murine transcription factor SOX3 during embryonic and adult neurogenesis. Gene Expr. Patterns. 2013;13:240–248. doi: 10.1016/j.gep.2013.04.004. [DOI] [PubMed] [Google Scholar]
  53. Semerci F., Choi W.T.-S., Bajic A., Thakkar A., Encinas J.M., Depreux F., Segil N., Groves A.K., Maletic-Savatic M. Lunatic fringe-mediated Notch signaling regulates adult hippocampal neural stem cell maintenance. eLife. 2017;6 doi: 10.7554/eLife.24660. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Semerci F., Maletic-Savatic M. Transgenic mouse models for studying adult neurogenesis. Front. Biol. 2016;11:151–167. doi: 10.1007/s11515-016-1405-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. Semerci F., Parkitny L., Maletic-Savatic M. Mouse Models for Studying Hippocampal Adult Neural Stem Cell Biology. Methods Mol. Biol. 2021;2224:61–74. doi: 10.1007/978-1-0716-1008-4_4. [DOI] [PubMed] [Google Scholar]
  56. Shen-Orr S.S., Gaujoux R. Computational deconvolution: extracting cell type-specific information from heterogeneous samples. Curr. Opin. Immunol. 2013;25:571–578. doi: 10.1016/j.coi.2013.09.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  57. Slavotinek A., Rego S., Sahin-Hodoglugil N., Kvale M., Lianoglou B., Yip T., Hoban H., Outram S., Anguiano B., Chen F., et al. Diagnostic yield of pediatric and prenatal exome sequencing in a diverse population. NPJ Genom. Med. 2023;8:10. doi: 10.1038/s41525-023-00353-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. Suh H., Consiglio A., Ray J., Sawai T., D’Amour K.A., Gage F.H. In vivo fate analysis reveals the multipotent and self-renewal capacities of Sox2+ neural stem cells in the adult hippocampus. Cell Stem Cell. 2007;1:515–528. doi: 10.1016/j.stem.2007.09.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Thompson J.R., Nelson E.D., Tippani M., Ramnauth A.D., Divecha H.R., Miller R.A., Eagles N.J., Pattie E.A., Kwon S.H., Bach S.V., et al. An integrated single-nucleus and spatial transcriptomics atlas reveals the molecular landscape of the human hippocampus. bioRxiv. 2024 doi: 10.1101/2024.04.26.590643. Preprint at. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Tosun M., Semerci F., Maletic-Savatic M. Heterogeneity of Stem Cells in the Hippocampus. Adv. Exp. Med. Biol. 2019;1169:31–53. doi: 10.1007/978-3-030-24108-7_2. [DOI] [PubMed] [Google Scholar]
  61. Voineagu I., Wang X., Johnston P., Lowe J.K., Tian Y., Horvath S., Mill J., Cantor R.M., Blencowe B.J., Geschwind D.H. Transcriptomic analysis of autistic brain reveals convergent molecular pathology. Nature. 2011;474:380–384. doi: 10.1038/nature10110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Wang J., Al-Ouran R., Hu Y., Kim S.-Y., Wan Y.-W., Wangler M.F., Yamamoto S., Chao H.-T., Comjean A., Mohr S.E., et al. MARRVEL: Integration of Human and Model Organism Genetic Resources to Facilitate Functional Annotation of the Human Genome. Am. J. Hum. Genet. 2017;100:843–853. doi: 10.1016/j.ajhg.2017.04.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Wang L., Wang C., Moriano J.A., Chen S., Zuo G., Cebrián-Silla A., Zhang S., Mukhtar T., Wang S., Song M., et al. Molecular and cellular dynamics of the developing human neocortex. Nature. 2025 doi: 10.1038/s41586-024-08351-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Warde-Farley D., Donaldson S.L., Comes O., Zuberi K., Badrawi R., Chao P., Franz M., Grouios C., Kazi F., Lopes C.T., et al. The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function. Nucleic Acids Res. 2010;38:W214–W220. doi: 10.1093/nar/gkq537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Wiszniewska J., Bi W., Shaw C., Stankiewicz P., Kang S.-H.L., Pursley A.N., Lalani S., Hixson P., Gambin T., Tsai C.h., et al. Combined array CGH plus SNP genome analyses in a single assay for optimized clinical testing. Eur. J. Hum. Genet. 2014;22:79–87. doi: 10.1038/ejhg.2013.77. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Wojcik M.H., Reuter C.M., Marwaha S., Mahmoud M., Duyzend M.H., Barseghyan H., Yuan B., Boone P.M., Groopman E.E., Délot E.C., et al. Beyond the exome: what’s next in diagnostic testing for Mendelian conditions. Am. J. Hum. Genet. 2023;110:1229–1248. doi: 10.1016/j.ajhg.2023.06.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Wu Y., Korobeynyk V.I., Zamboni M., Waern F., Cole J.D., Mundt S., Greter M., Frisén J., Llorens-Bobadilla E., Jessberger S. Multimodal transcriptomics reveal neurogenic aging trajectories and age-related regional inflammation in the dentate gyrus. Nat. Neurosci. 2025;28:415–430. doi: 10.1038/s41593-024-01848-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Yang Y., Muzny D.M., Xia F., Niu Z., Person R., Ding Y., Ward P., Braxton A., Wang M., Buhay C., et al. Molecular findings among patients referred for clinical whole-exome sequencing. JAMA. 2014;312:1870–1879. doi: 10.1001/jama.2014.14601. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Zhong Y., Liu Z. Gene expression deconvolution in linear space. Nat. Methods. 2011;9:8–9. doi: 10.1038/nmeth.1830. [DOI] [PubMed] [Google Scholar]
  70. Zhong Y., Wan Y.-W., Pang K., Chow L.M.L., Liu Z. Digital sorting of complex tissues for cell type-specific gene expression profiles. BMC Bioinf. 2013;14:89. doi: 10.1186/1471-2105-14-89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Zuckerman N.S., Noam Y., Goldsmith A.J., Lee P.P. A self-directed method for cell-type identification and separation of gene expression microarrays. PLoS Comput. Biol. 2013;9 doi: 10.1371/journal.pcbi.1003189. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Figures S1, S2, and Table S7
mmc1.pdf (37.4MB, pdf)
Data S1. Tables S1–S6
mmc2.xlsx (118.3KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (53.2MB, pdf)

Data Availability Statement

The microarray source data are available at GEO repository. The accession number for this dataset reported in this paper is: GSE303501. The code and the DSA framework are accessible at the GitHub link: https://github.com/geri86Ale/DSA-framework. Our DSA source code is already available in previously published papers (Zhong et al., 2013a; Zhong and Liu, 2011), while the source code for scRNA-seq vs. DSA genes comparison is available at GitHub with the link: https://github.com/zhandong/DSA.


Articles from Stem Cell Reports are provided here courtesy of Elsevier

RESOURCES