Abstract
Background
The integration of single-cell RNA sequencing (scRNA-seq) and high-resolution spatial transcriptomics (ST) could improve our understanding of both tissue architecture and cellular heterogeneity simultaneously. The key to accomplishing this goal mainly relies on effectively co-embedding similar cells with consistent representations from the two types of data.
Methods
In this paper, we construct a conditional variational autoencoder (CVAE) architecture, named SpateCV, to explicitly regularize the embedding alignment of similar cells from scRNA-seq and ST data in a shared latent through a clustering loss.
Results
Benchmark results across twelve datasets demonstrate that SpateCV achieves superior performance in spatial gene imputation and spatial patterns reconstruction. Critically, SpateCV translates this technical accuracy into biological insight. With the imputed genome-wide expression, our method enables the identification of novel spatially differentially expressed genes, such as the astrocyte marker Hepacam, and facilitates the inference of layer-specific intercellular communication networks, identifying corpus callosum cells as key signaling hubs in the mouse visual cortex. Additionally, SpateCV enables the in silico spatial mapping of neuronal subtypes by integrating spatial context into scRNA-seq data.
Conclusion
SpateCV provides a robust framework for extracting biological knowledge from multimodal spatial-omics data.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12967-025-07245-0.
Keywords: Spatial transcriptomics, Single-cell RNA sequencing, Gene imputation, Conditional variational autoencoder (CVAE), Attention mechanism
Introduction
Single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) now enable quantitative dissection of cellular heterogeneity and tissue architecture. These advancements enable the study of dynamic cellular interactions in cancer [1–3]. While scRNA-seq provides high-resolution gene expression profiles at the single-cell level, the loss of spatial information limits its utility for studying tissue architecture and function [4]. In contrast, single-cell spatial transcriptomics (scST) technologies that utilize fluorescence in situ hybridization or in situ sequencing, such as seqFISH [5], osmFISH [6], MERFISH [7], and STARmap [8] can capture both gene expression and spatial information at single-cell resolution [9]. However, scST technologies typically measure only a limited subset of genes (usually a few hundred), necessitating the imputation of unmeasured genes to fully explore biological processes [10, 11]. Integrating scRNA-seq and scST data from the same tissue region offers a promising approach to investigate both cellular identity and tissue architecture simultaneously.
Mapping-based approaches [12, 13] could map scRNA-seq cells onto spatial locations to transfer expression of unmeasured genes. For example, NovoSpaRc [14] assumed that cells in physical proximity have similar gene expression profiles and implicitly maps single cells to spatial locations using optimal transport methods [15]. Tangram [16] assumed each spatial cell as a probabilistic combination of scRNA-seq cells. However, these mapping approaches are hard to deal with the batch effects across platforms, leading to inaccurate gene imputation.
Projecting both scST and scRNA-seq data into a shared latent space could empower the imputation of unmeasured genes for scST data [17–19]. In this space, gene expression can be estimated through averaging the top k similar scRNA-seq cells. While pioneering, many early methods rely on linear assumptions. For instance, approaches like Seurat [20], SpaGE [21], and stPlus [22] estimated unmeasured gene expression of a spatial cell through weighted averages of its k-nearest scRNA-seq cells. Essentially, these methods assumed that the gene expression of a cell can be modelled by a linear combination of its nearest neighbors, which limits their capability to capture the nonlinear patterns inherent in gene expression. To overcome these challenges, deep learning has emerged as a powerful paradigm for single-cell data integration. Deep generative models, in particular, have proven highly effective. Methods such as ENVI [23], gimVI [24], scVI [25], and uniPort [26] use variational autoencoders (VAEs) to project scRNA-seq and ST data into a shared, non-linear latent space where batch effects can be explicitly modeled and corrected. This end-to-end learning approach allows for the direct generation of unmeasured gene expression, providing a more robust solution than linear averaging.
Cross-modal embedding by learning techniques [27–32] is essentially to align similar cells from different modal in a shared latent space. In this paper, we consolidate the alignment in a conditional variational autoencoder (CVAE) [33, 34] framework by adding a clustering loss that aims to explicitly require similar cells having similar embeddings in the latent space. Results across twelve benchmark datasets demonstrate that SpateCV achieves superior performance in gene imputation and spatial expression patterns reconstruction. Moreover, SpateCV outperforms other methods in alleviating batch effects between scRNA-seq and ST data. By imputing unmeasured gene expressions across platforms, SpateCV enables downstream analysis, such as identifying differentially expressed genes in MERFISH data. Applying SpateCV to mouse visual cortex data generates genome-wide ST data, facilitating the detection of critical ligand-receptor interactions. Furthermore, SpateCV facilitates the in silico spatial contextualization of Sst interneuron subtypes from dissociated scRNA-seq data by predicting their local spatial context. Through these distinct biological applications, we demonstrate that SpateCV is not only a robust computational tool but also a powerful engine for generating testable hypotheses and uncovering new biological insights in complex tissues.
Results
Overview of the SpateCV method
SpateCV employs a conditional variational autoencoder (CVAE) backbone to learn a joint latent representation of cells across ST and scRNA-seq data (Fig. 1a). The encoder learns the coherent embeddings of cells in a shared latent space from scRNA-seq and ST gene expression matrices, regularized by KL divergence loss to ensure latent space stability. Two decoders reconstruct gene expression profiles using negative binomial and Poisson loss respectively, while the decoder also reconstructs the spatial covariance (COVET) matrices to preserve local spatial context in the latent space. Here, we highlight that a clustering loss is introduced in the latent space to optimize the alignment of similar cells from two modalities. The multi-head attention in the encoder and decoder serves as feature learning across the two modalities. These learned latent representations could be used for spatial gene expression imputation, identification of spatially variable genes, trajectory inference, and other downstream analytical tasks (Fig. 1b).
Fig. 1.
Overview of SpateCV. (a) SpateCV uses a single CVAE encoder to learn joint latent embeddings of cells for both ST and scRNA-seq data. Two separate decoders reconstruct each data type and COVET matrices. A clustering loss is introduced in the latent space. (b) The learned latent representations could be used for gene imputation, spatial DE genes analysis, cell-cell communication, and other downstream analyses
SpateCV achieves the best gene-expression imputation performance
To evaluate SpateCV’s gene imputation accuracy, we compare its performance against seven baseline methods across 12 benchmark datasets using four metrics to capture both correlation and error. Higher values indicate better performance for PCC, SSIM; lower values are better for RMSE and JS. Figure 2 shows their PCC, SSIM, RMSE, and JS across 12 benchmark datasets, where the values ranked first and second are marked with solid and dashed boxes, respectively (detailed results in Supplementary Tables 1 and 2, Supplementary Fig. 1). Concerning the PCC and RMSE metrics, SpateCV ranks first on seven datasets and second on Dataset 8, which significantly outperforms other baseline methods. For the SSIM metric, SpateCV is second only to stDiff, which ranks first on five datasets and second on one. For the JS metric, SpateCV ranks second to Tangram, which leads on four datasets and ranks second on two. These observations demonstrate that, compared with other methods, SpateCV achieves the best imputation performance on the twelve datasets. This constitutes a critical test, as preserving spatial organization is the ultimate goal of ST analysis. A key consideration here is whether high numerical accuracy in gene expression levels translates into the faithful reconstruction of their underlying spatial patterns. To address this, we next visualized the spatial expression patterns of representative genes to intuitively assess how well SpateCV preserves complex tissue architecture compared to baseline methods.
Fig. 2.
Imputation performance on the 12 benchmark datasets by SpateCV and baseline methods. The values ranked first and second are marked with solid and dashed boxes, respectively
SpateCV could faithfully predict ST spatial patterns
Figure 3 further visualizes the spatial expression patterns by a specific representative gene across 12 datasets. Overall, SpateCV’s predictions generally provide a more balanced representation of high (red) and low (blue) expression regions, often mitigating the tendency for systematic overestimation seen in methods like gimVI and Tangram, or underestimation as observed with ENVI and uniPort in some datasets. SpateCV particularly excels at reconstructing intricate, tissue-specific structures. Further, SpateCV could even preserve the spatial patterns in some tissue-specific dataset. For instance, the expression of Cbln2 in Dataset 6 presents a low-level pattern with ‘Y’ shape around the empty region (white spot). Except for SpateCV, other methods almost fail to reproduce it. SpateCV also successfully presents the unique high-level layer (top region) of Gfap, a marker of type 1 astrocytes in Dataset 9 [35]. These results demonstrate that introducing a spatial covariance loss in the latent space could facilitate SpateCV to predict spatial patterns.
Fig. 3.
Spatial expression patterns of a representative gene predicted by SpateCV and baseline methods across 12 benchmark datasets. The first column displays the ground truth spatial pattern, while subsequent columns show the predicted patterns by all methods
The imputed ST profiles by SpateCV could preserve cellular topological structure
Figure 4 shows AMI, ARI, homogeneity, and NMI of the cell clustering results by the Leiden clustering algorithm [36], using the true and imputed gene expression by SpateCV and other baseline methods (detailed results in Supplementary Tables 3 and 4, Supplementary Fig. 1). Higher values indicate better performance for ARI, AMI, Homo, and NMI. For the ARI metric, SpateCV ranks first on two datasets and second on two others, which performs less well than stDiff, but equally as ENVI and uniport. In terms of AMI and NMI, SpateCV achieves the first on two datasets and ranks second on six. For the homogeneity metric, it ranks first on one dataset and second on seven. In sum, these results indicate that SpateCV accomplishes the best overall performance in preserving cellular topological structure.
Fig. 4.
The clustering performance of cells between true and imputed gene expression by SpateCV and baseline methods. The values ranked first and second are marked with solid and dashed boxes, respectively
SpateCV could mitigate scRNA-seq batch effects
Figure 5 shows the UMAP/t-SNE distribution of cells with blue, green, and orange colors representing scRNA-seq, real ST, and imputed ST expressions respectively, generated by SpateCV and baseline methods. Ideally, imputed ST cells align closely with real ST cells while remaining distinct from scRNA-seq reference cells, indicating effective batch effect suppression. SpateCV’s imputed cells are generally more closely aligned with real ST cells and farther from scRNA-seq cells compared to baseline methods. For instance, the cells imputed by SpaGE, stPlus, and uniPort tend to be closer to scRNA-seq cells than real ST cells in Datasets 1, 3, 8, and 10. The same tendency is also observed by the other four baseline methods in Datasets 2. Additionally, the imputed cells from gimVI, SpaGE, and stPlus exhibit a scattered distribution, which may indicate their poor latent representation capability. Notably, SpateCV performs significantly better than other methods in Dataset 2, although it has a serious sample imbalance, where sparse ST data (175 cells) contrast with abundant scRNA-seq data (9,902 cells). These coherent observations demonstrate that SpateCV effectively minimizes batch effects from scRNA-seq data and aligns imputed cells closely with real ST cells across these datasets.
Fig. 5.
UMAP/t-SNE visualizations of scRNA-seq, real ST, and imputed ST cells by SpateCV and baseline methods. Cells based on gene expression in each sub-figure are colored: scRNA-seq (blue), real ST (green), imputed ST (orange)
SpateCV mitigates batch effects in cross-modal transcriptomic data
To assess SpateCV’s ability to mitigate batch effects, we visualize the alignment of scRNA-seq and ST cells in a shared latent space using UMAP/t-SNE on the mouse hypothalamus dataset. Figure 6a shows the UMAP distribution of the mouse hypothalamus cells from scRNA-seq and MERFISH-based spatial transcriptomics data, while Fig. 6b shows that by SpateCV based on 154 shared genes. In Fig. 6a(left), the UMAP of original scRNA-seq (blue) and MERFISH (orange) data shows clear separation by modality. In contrast, Fig. 6b(left) shows overlapping latent representations from SpateCV, indicating effective modality alignment. The right panels of Fig. 6a–b show cell-type-specific distributions, with SpateCV clustering same-type cells from both modalities together.
Fig. 6.
SpateCV mitigates batch effects between MERFISH and scRNA-seq data. (a) UMAP visualization of MERFISH and scRNA-seq data. The left shows the cell distribution from the two modalities of data (left), and the right shows them by the corresponding cell types. (b) UMAP visualization by SpateCV based on the 154 genes. (c) The aggregated score of F1, ARI, and NMI of SpateCV and the other seven methods. (d) SpateCV achieves a good balance with a large silhouette coefficient and a small batch entropy score
To compare with other methods, Fig. 6c shows the aggregated scores of F1, ARI, and NMI of the 13 cell-type clusters by SpateCV, uniPort, Harmony, Seurat, SCALEX [37], scVI, gimVI, and MultiMAP [38] (UMAPs seen in Supplementary Fig. 2). MultiMAP (2.412) and SpateCV (2.402) achieve almost the same values, which significantly surpass other methods. The similar silhouette coefficients indicate that both SpateCV and MultiMAP effectively cluster cells by type in the latent space.
Further, Fig. 6d reveals that SpateCV could achieve a large silhouette coefficient (Supplementary Table 5). The similar silhouette coefficients of SpateCV and MultiMAP indicate that they effectively cluster cells by type in the latent space. Although SpateCV exhibits a lower batch entropy score compared to MultiMAP (Fig. 6d), this reflects a common trade-off in single-cell data integration, where prioritizing biological variation preservation and cell-type separation (resulting in a high silhouette coefficient) may allow for slightly less aggressive batch mixing [39, 40]. Despite this, SpateCV achieves a similar aggregated score (Fig. 6c), indicating an effective balance between batch effect removal and maintaining cellular heterogeneity. This trade-off ensures that SpateCV avoids over-integration, which could otherwise lead to loss of meaningful biological signals, as observed in other methods [41]. Having established that SpateCV can robustly integrate cross-modal data, the ultimate measure of its utility is its ability to facilitate new biological discoveries. We therefore sought to determine if the genome-wide expression profiles imputed by SpateCV could enhance the identification of biologically meaningful, spatially differentially expressed (DE) genes that are missed in the original sparse data. To this end, we applied SpateCV to a MERFISH dataset of the mouse primary motor cortex (MOp).
SpateCV enables discovery of spatially DE genes and biological pathways in MERFISH data
The MERFISH dataset of mouse primary motor cortex (MOp) [42] includes 5,381 cells with 254 genes. Using paired snRNA-seq data as a reference from the same region, Fig. 7a presents the UMAP distribution of these cells annotated with 20 distinct cell types.
Fig. 7.
SpateCV enhances discovery of spatially DE genes and biological pathways. (a) UMAP visualization of 5,381 cells with 20 annotated cell types in the MOp brain slice. (b) Spatial expression patterns of seven cortical marker genes comparing ground-truth with SpateCV and baseline methods (ENVI, SpaGE, tangram, gimVI, stPlus). (c) Mean absolute error (MAE) for the marker genes in (b), comparing predicted versus ground-truth expression. (d) MAE across all 247 overlapping marker genes. (e) Imputed spatial expression of seven “non-MERFISH” genes (Gfap, Aqp4, Nfia, Hepacam, Nxn, Ptprz1, and Gramd3) by SpateCV and baseline methods, validated against Allen brain Atlas ISH data (rightmost column) (f) Volcano plot of spatially DE genes from the original MERFISH data (247 genes), showing log-fold change (logFC) versus -log10(p-value) from a two-sided z-test; significant genes (FDR < 0.01, fold change > 1.5) are color-coded by cell type. (g) Volcano plot of spatially DE genes from the SpateCV-imputed profile (1,938 genes), analogous to (f). (h) Spatial distribution of imputed Hepacam expression, showing astrocyte-specific localization. (i) Heatmap of the top 20 genes correlated with imputed Hepacam expression. (j) GO enrichment analysis (biological process) for genes co-expressed with Hepacam
To compare with other methods, we select seven well-characterized layer-specific markers: Osr1 (L6 IT), Otof (L2/3 IT), Slc17a6 (L2/3 IT, L5 IT), Fam84b (L5), Opalin (Oligo), Cdh12 (L2-6), and Fezf2 (L6 IT, L6b). Figure 7b shows the spatial expression patterns of these genes by SpateCV, ENVI, stPlus, SpaGE, Tangram, and gimVI. SpateCV effectively denoises and recovers sharp spatial boundaries, such as the L6 IT layer restriction of Osr1. For instance, SpateCV effectively captures the localized high-expression domain of Opalin and Slc17a6-positive glutamatergic populations. In contrast, Tangram and gimVI produce blurred or misaligned patterns. Figure 7c shows the Mean Absolute Error (MAE) for these genes, and Fig. 7d shows the average MAE across 247 shared genes between the two datasets (detailed results in Supplementary Fig. 3a, b, Supplementary Table 6). SpateCV achieves the lowest average MAE.
Further, we select seven unmeasured genes in the original MERFISH panel from glial cells: Gfap (Astro), Aqp4 (Astro), Nfia (Astro, Oligo), Hepacam (Oligo), Nxn (Astro), Ptprz1 (Astro, OPC), and Gramd3 (Astro, Oligo). Figure 7e shows their spatial expression patterns, which align closely with Allen Brain Atlas in situ hybridization (ISH) [43]. In contrast, other methods tend to overestimate or underestimate the expression of some genes (e.g., Nfia, Ptprz1).
SpateCV enhances downstream analysis of spatially differentially expressed (DE) genes. Utilizing the C-SIDE [44] tool, we identify 42 spatially DE genes at FDR < 0.01 (Fig. 7f) in the original MERFISH dataset, while 90 spatially DE genes are identified in the 1,938 SpateCV imputed genes (Fig. 7g). For instance, as an astrocyte marker, Gfap presents a subtle spatial heterogeneity [45, 46] (Fig. 7e).
Additionally, we identify Hepacam as a gene exhibiting astrocyte-specific spatial differential expression, which is not identified in the original data (Fig. 7e, h, Supplementary Fig. 3c, d). Hepacam encodes a cell adhesion molecule predominantly found in glial cells, and plays a critical role in maintaining astrocytic polarity, blood-brain barrier integrity, and regulating cell proliferation and migration [47–49]. Subsequent co-expression analysis reveals the top 20 genes most correlated with imputed Hepacam expression (Fig. 7i), with Gene Ontology (GO) enrichment analysis indicating enrichment in collagen biosynthesis and cell proliferation pathways (Fig. 7j). These results suggest that Hepacam may contribute to glial cell polarity, extracellular matrix remodeling, and astrocyte reactivity during neuroinflammation or injury [50]. By linking Hepacam’s spatial expression to these pathways, SpateCV provides insights into its role in modulating astrocyte-mediated tissue repair and blood-brain barrier integrity. Additionally, SpateCV’s ability to identify spatially DE genes like Hepacam could inform therapeutic targets in diseases such as cancer [51, 52]. Similar analysis of Slc17a7 is provided in Supplementary Fig. 4. Further validation across diverse brain regions and pathological contexts is needed to confirm these associations.
SpateCV reveals layer-specific intercellular communication networks in STARmap data
The STARmap dataset profiles 1,020 genes across 1,207 cells from the mouse primary visual cortex (VISp) [53]. Figure 8a delineates seven cortical cell populations, including hippocampal (HPC), corpus callosum (CC), and cortical layers (L1, L2/3, L4, L5, L6). Figure 8b presents the spatial registration of these cell types onto the Allen Brain Atlas reference.
Fig. 8.
SpateCV reveals layer-specific cellular organization and intercellular communication networks in the mouse VISp. (a) UMAP visualization of 1,207 STARmap cells profiled for 1,020 genes, color-coded by seven cortical cell types: hippocampal (HPC), corpus callosum (CC), and layers L6, L5, L4, L2/3, L1. (b) Spatial registration of these molecularly defined cell types onto the corresponding allen brain atlas reference map of VISp, illustrating the conserved cytoarchitectural organization. (c) Spatial expression patterns of four layer-specific marker genes (mbp for CC, Pcp4 for L6, Nrsn1 for L4, Lamp5 for L2/3) from the original STARmap data, with heatmaps showing mean expression across cell types. (d) SpateCV’s imputed spatial expression patterns for the marker genes in (c), with heatmaps confirming layer-specificity. (e) Imputed spatial expression patterns for four non-STARmap marker genes (Ugt1a7c for CC, Mapk4 for L4/5, Mpped1 for L2/3, and Fcgr2b for L1), with heatmaps of imputed expression. (f) Validation of imputed patterns in (e) against allen brain atlas ISH data. (g-i) Intercellular signaling networks inferred by COMMOT on SpateCV-imputed data, showing interaction strengths (left scatter plots, ligand-sending cell types on dots, receptor-expressing cell types on arrows, color intensity represents interaction strength) and spatial expression of ligands (center, red), receptors (right, green/blue). (g) Wnt11 (ligand) → Fzd5 (receptor) axis, with CC cells targeting HPC and L6 neurons. (h) Wnt5a (ligand) → Fzd4 (receptor) axis, with CC cells targeting HPC. (i) Cxcl1 (ligand) → Ackr1 (receptor) axis, with CC cells targeting L6 neurons
Figure 8c and Fig. 8d respectively show the spatial expression patterns of four layer-specific marker genes (Mbp, Pcp4, Nrsn1, Lamp5) based on the original expressions and imputed expressions by SpateCV. The imputed patterns of these genes are highly coincident with the layer patterns, as shown in Fig. 8a. Except for Mbp, the patterns of other three genes also shows different layer specificities. Figure 8e further shows the expression patterns by four marker genes (Ugt1a7c, Mapk4, Mpped1, Fcgr2b) absent from STARmap. These patterns align closely with Allen Brain Atlas in situ hybridization (ISH) data (Fig. 8f), e.g., capturing Ugt1a7c’s enrichment in CC.
Using the SpateCV’s imputed genome-wide expression atlas, we analyze intercellular communication networks in VISp with the COMMOT [54] tool. Figure 8g–i show ligand-receptor interactions, revealing layer-specific communication axes with CC cells as a key signaling center. Three distinct ligand-receptor interactions underscore this observation: First, Wnt11, predominantly secreted by CC cells, is predicted to signal to both HPC and L6 neurons via the Fzd5 receptor, while also participating in autocrine signaling within other cell types (Fig. 8g). This suggests a role for CC glia-derived Wnt11 in coordinating the maturation and circuit integration of deep-layer cortical and hippocampal neurons. Second, Wnt5a, also originating from CC cells, exhibits preferential signaling to HPC neurons through the Fzd4 receptor, with autocrine loops maintained elsewhere (Fig. 8h), indicative of a targeted non-canonical Wnt pathway potentially modulating synaptic plasticity in the hippocampus. Third, a highly layer-selective chemokine axis is identified, with CC-secreted Cxcl1 specifically targeting L6 neurons via the Ackr1 receptor, whereas other cell types primarily display self-signaling patterns (Fig. 8i). This glia-to-neuron communication may play a role in regulating neuronal excitability or neuroinflammatory states specifically within the deep cortical layers. The spatial co-expression maps of each ligand and its corresponding receptor (Fig. 8g–i) provide visual corroboration for these COMMOT-inferred interactions. Collectively, these findings illustrate that CC cells dispatch distinct molecular signals to neuronal populations (HPC and L6) while engaging in more widespread autocrine signaling. The ability to delineate such high-resolution, layer-specific communication networks, critically enabled by SpateCV’s genome-wide imputation, uncovers previously unappreciated glia-neuron and neuron-neuron cross-talk mechanisms. This provides novel and granular insights into the complex cellular and molecular underpinnings that govern information processing and homeostasis within the visual cortex. Beyond imputing gene expression for ST data, can SpateCV’s framework be leveraged in the reverse direction—to infer the spatial context for cells profiled only with scRNA-seq? Answering this question would address a major limitation of dissociated single-cell studies. To test this capability, we applied SpateCV to the challenging task of resolving the laminar organization of Somatostatin (Sst) interneuron subtypes, whose spatial positioning is critical to their function but is lost during single-cell sequencing.
SpateCV maps the laminar organization of Sst interneuron subtypes by integrating spatial context into scRNA-seq profiles
Somatostatin (Sst)-expressing interneurons are a cardinal class of inhibitory neurons in the cortex [55], involved in regulating dendritic excitability and neuropathological conditions such as Alzheimer’s disease and depression [56]. Their transcriptional subtypes’ spatial organization across cortical laminae remains challenging to study due to the loss of spatial information in scRNA-seq. We apply SpateCV to integrate spatial context to scRNA-seq from the mouse primary visual cortex (VISp).
Fig. 9a shows the UMAP distribution of cells from original scRNA-seq data (left) and embedded in SpateCV’s shared latent space (right). This latent embedding induces more cohesive clustering of cell types compared to the original scRNA-seq embedding. For Sst interneurons (excluding a Chodl outlier population, see Supplementary Fig. 5 for an analysis including this population), we identify eight transcriptional subtypes (Calb2, Crh, Myh4, Crhr2, Cartpt, Hpse, Etv1, and Nmbr), each characterized by its marker genes (Fig. 9b, c). Figure 9d presents the UMAP distribution of cells represented by the COVET matrices predicted by SpateCV.
Fig. 9.
SpateCV resolves laminar organization of sst interneuron subtypes in the mouse VISp. (a) UMAP visualizations for original VISp scRNA-seq data (left, colored by cell types) and embedded in SpateCV’s latent space (right). (b) UMAP visualization of sst-lineage cells, colored by eight subtypes: Calb2, Crh, Myh4, Crhr2, Cartpt, Hpse, Etv1, and Nmbr. (c) Heatmap of eight marker genes defining the sst subtypes. (d) UMAP visualization of the sst-lineage cells. Each cell is represented by the SpateCV imputed COVET matrices. (e) UMAP visualization of sst interneurons, colored by pseudo-depth (DC_COVET1) scores from superficial (purple) to deep (yellow). (f) UMAP visualization of sst interneurons, colored by labeled sst subtype. (g) Violin plots of pseudo-depth for each sst subtype. (h) PAGA graph of subtype centroid covariances in COVET space, with edge thickness indicating spatial coupling strength
We quantify the spatial ordering using the principal diffusion component (DC_COVET1) from the COVET matrix, termed “pseudo-depth.” Figure 9e shows this “pseudo-depth” axis, which correlates linearly with anatomical cortical depth, serving as an in silico proxy for laminar positioning. Figure 9f shows the mapping of Sst subtypes onto this axis. Figure 9g further shows a stratification of these subtypes (see Supplementary Fig. 6 for the stratification of several cortical cell layers). For example, Crh cells localize to the deepest positions, consistent with CRH Martinotti cells in infragranular layers [57], while Calb2 cells occupy superficial positions, aligning with their enrichment in layer I. Intermediate subtypes, such as Hpse (layers 4/5) [53, ]Etv1 (layers 5), and Crhr2 (layers 6), show distinct laminar preferences consistent with prior findings [57]. Furthermore, a Partition-based Graph Abstraction (PAGA) analysis of these subtypes reveals spatial coupling between developmentally related or functionally synergistic lineages (e.g., Etv1 and Cartpt subtypes), underscoring the existence of shared or adjacent laminar niches (Fig. 9h). Together, these results demonstrate that SpateCV’s integration of inferred spatial covariance information not only preserves the fine-grained transcriptional identities of Sst interneuron subtypes but also reconstructs their in situ laminar organization within the cortex. This approach offers novel, high-resolution insights into the cellular and molecular architecture that underpins cortical circuitry and function, providing a valuable strategy for spatially contextualizing dissociated single-cell datasets and bridging the gap between transcriptomic diversity and spatial neuroanatomy.
Discussion
Bridging the gap between the computational challenge of data integration and the biological goal of understanding tissue function is a central theme in spatial genomics. Single-cell spatial transcriptomics captures gene expression with spatial context but is limited by low gene coverage, necessitating methods to impute missing gene expression. In this paper, we present SpateCV, which introduces a clustering loss in a conditional variational autoencoder (CVAE) framework to align embeddings of similar cells from scRNA-seq and ST in a shared latent space. Results across twelve benchmark datasets demonstrate that SpateCV achieves the best imputation performance and effectively reconstructs spatial expression patterns, particularly in tissues with distinct laminar boundaries. The clustering loss promotes SpateCV to obtain consistent cell type representations, enabling it to mitigate batch effects while preserving cellular topology. Further, SpateCV supports diverse downstream analyses across platforms. In MERFISH data, SpateCV enables the identification of spatially differentially expressed genes, including Hepacam, an astrocyte-specific gene linked to collagen biosynthesis and cell proliferation pathways. In STARmap data, genome-wide imputation facilitates the inference of layer-specific intercellular communication networks, identifying corpus callosum cells as key signaling hubs. Furthermore, SpateCV accurately maps Sst interneuron subtypes to their laminar distributions by integrating spatial context into scRNA-seq data in the mouse visual cortex.
Despite these advancements, SpateCV has several limitations that warrant discussion. First, its performance, like other reference-based methods, is intrinsically linked to the quality and coverage of the reference scRNA-seq data. For genes that are rare, lowly expressed, or exhibit high variability in the reference panel, imputation accuracy may be reduced. This likely explains the observed underestimation of genes such as Tead2 and KLHL11 in our benchmarks. Second, there is an inherent trade-off in our CVAE-based model between denoising and preserving sharp expression boundaries. By learning a smooth latent space, SpateCV excels at capturing robust, generalizable patterns, but this can sometimes lead to a slight “smoothing” effect at the edges of distinct expression domains. This may manifest as a minor overestimation in adjacent regions, as noted for certain genes in our analysis. We contend that this is a justifiable trade-off for the model’s strong overall performance in noise reduction and pattern reconstruction. Furthermore, in datasets with small ST sample sizes, the model may risk overfitting due to limited spatial covariance information. SpateCV’s computational complexity, driven by its multi-head attention mechanism and dual decoders, may also pose challenges for very large-scale datasets, requiring significant computational resources. Finally, while validated on imaging-based ST technologies (e.g., seqFISH, MERFISH, STARmap), its application to other ST platforms, such as slide-seq, may require further optimization to account for differences in spatial resolution and gene capture efficiency.
Future developments will focus on addressing these limitations and expanding SpateCV’s capabilities. Integrating multi-omics data, such as proteomics or epigenomics, using SpateCV’s multimodal fusion architecture could enhance its utility for comprehensive cellular profiling. Incorporating graph neural networks (GNNs) to model complex spatial relationships could complement the COVET matrix, improving spatial pattern reconstruction. Additionally, embedding uncertainty quantification for imputed gene expression values will provide confidence measures for downstream analyses. It is also crucial to address computational efficiency issues through model optimization.
Ablation experiments (Supplementary Figs. 7 and 8, Table 7) confirm the effectiveness of the attention mechanism, COVET loss and clustering constraints. The attention mechanism enhances feature learning across modalities, improving imputation accuracy. The clustering loss ensures robust alignment of similar cell types, as evidenced by high ARI and NMI scores. We also test SpateCV’s robustness across cluster numbers and COVET matrix neighborhood sizes. The detailed results are provided in Supplementary Figs. 9 and 10, Table 8.
In conclusion, SpateCV does not rely on any single isolated component. Instead, it lies in the collaborative architecture of its CVAE, attention mechanisms, clustering-guided latent space learning, and dual modality-specific decoders, which collectively achieve comprehensive and robust performance in gene imputation and cross-modal data integration. We believe that SpateCV can serve as a useful tool to generate genome-wide spatial expression maps and facilitate biological discoveries.
Materials and methods
Methods
SpateCV integrates single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) data into a shared latent space to facilitate gene imputation for ST data and spatial context prediction for scRNA-seq data. The model employs a conditional variational autoencoder (CVAE) framework enhanced with multi-head attention to jointly encode both modalities.
SpateCV processes two input matrices: scRNA-seq gene expression data (
),
cells and
genes, and ST gene expression data (
),
spatial cells and
genes. To capture spatial relationships, COVET matrices represent gene expression covariance across each cell’s
nearest neighbors, defined by spatial proximity. For each cell, a niche matrix
is constructed from the gene expression of its
nearest neighbors, forming a tensor
. The COVET matrix for each cell is computed as a covariance matrix:
![]() |
where
is the global average gene expression. This matrix, positive semi-definite, captures the unique microenvironment of each cell relative to the population.
The encoder projects scRNA-seq and ST data into a shared latent space using multi-head attention layers to capture features across modalities. An auxiliary variable
used to indicate the modality:
for scRNA-seq data and
for spatial data. The encoder outputs the mean
and log-standard deviation
of the latent variables
. These parameters describe a Gaussian distribution,
, from which the latent variabl
is sampled using the reparameterization trick:
![]() |
The latent variable
encodes both gene expression and spatial context.
The decoder comprises two components. The expression decoder reconstructs gene expression from
. it models gene counts with a negative binomial distribution, predicting parameters (
) (number of failures) and (
) (success probability) to account for dropout [58].
![]() |
For ST data, it uses a Poisson distribution, predicting the rate parameter
to reflect higher capture efficiency of fluorescence-based techniques.
![]() |
The COVET decoder reconstructs the COVET matrix by parameterizing its lower triangular Cholesky factor, minimizing the
error between the predicted and true matrix square roots.
SpateCV employs a deep k-means clustering loss to promote alignment of similar cells from scRNA-seq and ST in the latent space, defined as:
![]() |
where
is the latent representation of the
-th cell,
is the
-th cluster centroid, and
is the Euclidean distance. The weight
is computed using a Gaussian kernel to smooth optimization:
![]() |
with a temperature parameter
= 0.1. To stabilize optimization, weights are adjusted as:
![]() |
where
is a hyperparameter controlling weight adjustment.
The total loss function combines spatial reconstruction loss
, scRNA-seq reconstruction loss
, covariance loss
, KL divergence loss
, and clustering loss
. The final objective is to minimize this total loss:
![]() |
Each term is weighted by a corresponding coefficient (e.g.,
,
etc.), optionally adjusted adaptively during training to balance objectives.
Once trained, SpateCV facilitates gene imputation for ST data by reconstructing missing genes via the expression decoder. It also predicts spatial context for scRNA-seq data using the COVET decoder, supporting integrated analysis of gene expression and spatial organization in tissues.
Parameter selection and optimization
The SpateCV model employs a conditional variational autoencoder (CVAE) framework with specific design choices tailored to the characteristics of scRNA-seq and ST data. To justify the use of different distributions for data reconstruction, we adopt a Poisson distribution for ST data and a Negative Binomial distribution for scRNA-seq data. ST data, particularly from fluorescence-based technologies like MERFISH and STARmap, exhibit higher capture efficiency and lower dropout rates due to their direct measurement of gene expression in situ [6, 59]. The Poisson distribution is well-suited to model the count-based nature of these data, as it accurately captures the stochasticity of gene expression with minimal overdispersion [25, 58]. In contrast, scRNA-seq data suffer from significant dropout due to low mRNA capture efficiency, leading to zero-inflated gene expression profiles. The Negative Binomial distribution accounts for this overdispersion and dropout by modeling both the mean and variance of gene counts, providing a robust fit for scRNA-seq data.
The encoder and decoders consist of three hidden layers, with neuron counts adjustable based on dataset size. By default, the latent embedding layer includes 512 neurons. The multi-head attention mechanism employs 8 heads with a head dimension of 64, scaled to 128 for datasets with > 20,000 cells to capture complex cross-modal patterns. SpateCV is trained using mini-batch gradient descent with the Adam optimizer at an initial learning rate of
. Each batch contains 1,024 samples, equally split between scRNA-seq and ST data, with alternating batches during training. Early stopping prevents overfitting if the loss plateaus. To reduce computational complexity, the scRNA-seq dataset is subset to the union of the top 2,048 highly variable genes and all ST genes.
The loss weights (
,
,
,
,
) are optimized via grid search on a validation subset, targeting maximization of Pearson Correlation Coefficient (PCC) and minimization of Root Mean Square Error (RMSE) for gene imputation. By default,
,
, and
are set to 1.0 to equally prioritize spatial, scRNA-seq, and COVET matrix reconstruction. Regarding the β coefficient in the model, we clarify that
is set to 0.1 by default to prioritize reconstruction accuracy while maintaining latent space regularization. For small datasets whose total samples size is fewer than 10,000 cells, we recommend increasing the reliance on the prior and set
= 1.0. The
is set to 0.1 to ensure stable latent space learning and cell-type alignment without overfitting.
The number of clusters in the deep k-means clustering loss is determined based on prior knowledge of annotated cell types in each dataset. To assess robustness, we conduct sensitivity experiments on the mouse hypothalamus dataset, varying the number of clusters from 5 to 12. Metrics such as F1, Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), silhouette coefficient, and batch entropy remained stable, with variations < 5% across settings (Supplementary Figs. 9 and 10, Table 8). These results demonstrate SpateCV’s robustness to cluster number selection.
Benchmark datasets
We select 12 pairs of ST and scRNA-seq dataset for benchmark test, which cover multiple protocols, tissues and sizes, with wide ranges in gene and cell counts. Table 1 summarizes all datasets and Supplementary Table 9 summarizes the hyperparameter settings for each dataset.
Table 1.
Summary of the 12 benchmark dataset pairs
| ST data | scRNA-seq data | ||||||
|---|---|---|---|---|---|---|---|
| Data Pair | Tissue | Gene num | Cell num | technology | Gene num | Cell num | technology |
| Dataset1 | mouse gastrulation | 351 | 8425 | seqFISH [60] | 19103 | 4651 | 10X Chromium [60] |
| Dataset2 | mouse embryonic stem cell | 45 | 175 | seqFISH [61] | 16477 | 9902 | Microwell-Seq [62] |
| Dataset3 | mouse hippocampus | 249 | 3582 | seqFISH [63] | 16384 | 8323 | 10X Chromium [64] |
| Dataset4 | mouse cortex | 10000 | 524 | seqFISH+ [65] | 34041 | 14249 | Smart-seq [53] |
| Dataset5 | human osteosarcoma | 915 | 645 | seqFISH+ [66] | 19098 | 9234 | Drop-seq [67] |
| Dataset6 | mouse hypothalamic preoptic region | 154 | 4975 | MERFISH [7] | 18646 | 7891 | 10X Chromium [7] |
| Dataset7 | human osteosarcoma | 12903 | 645 | seqFISH+ [66] | 19098 | 9234 | Drop-seq [67] |
| Dataset8 | mouse VISP | 119 | 2842 | ISS [68] | 34041 | 14249 | Smart-seq [53] |
| Dataset9 | mouse somatosensory cortex | 33 | 3378 | osmFISH [6] | 30527 | 5613 | Smart-seq [53] |
| Dataset10 | Mouse visual cortex | 1020 | 1549 | STARmap [69] | 34041 | 14249 | Smart-seq [53] |
| Dataset11 | mouse prefronatal cortex | 166 | 1380 | STARmap [69] | 14837 | 5694 | 10X Chromium [70] |
| Dataset12 | human middle temporal gyrus | 120 | 1277 | ISS [71] | 48278 | 15928 | Smart-seq [71] |
Baselines
Gene imputation baselines
SpateCV is evaluated alongside established methods for imputing unmeasured genes in spatial transcriptomics (ST) data and integrating scRNA-seq and ST data, using recommended settings from their published pipelines.
Tangram. Implemented per the GitHub repository (https://github.com/broadinstitute/Tangram) with modes = ‘clusters’, density = ‘rna_count_based’.
gimVI. Adopted from scvi-tools documentation (https://docs.scvi-tools.org/en/0.8.0/user_guide/notebooks/gimvi_tutorial.html), using model.get_imputed_values(normalized = False) for spatial gene distributions
SpaGE. Followed the GitHub tutorial (https://github.com/tabdelaal/SpaGE/blob/master/SpaGE_Tutorial.ipynb) with n_pv = gene_num/2.
stPlus. Implemented per the GitHub repository (https://github.com/xy-chen16/stPlus), with tmin = 5, neighbor = 50.
uniPort. Adopted from the official example (https://uniport.readthedocs.io/en/latest/examples/MERFISH/MERFISH_impute.html), using default settings.
ENVI. Followed the tutorial (https://scenvi.readthedocs.io/en/latest/tutorial/MOp_MERFISH_tutorial.html#), with default parameters.
stDiff. Implemented per the GitHub repository (https://github.com/fdu-wangfeilab/stDiff), with default parameters.
Integration baselines
For scRNA-seq and ST data integration, SpateCV is compared to Seurat (v5.2.1), Harmony (v1.2.3), SCALEX (v1.0.2), MultiMAP (v0.0.1), scVI (v0.14.6), gimVI (v0.8.0), and uniPort (v1.2.2). Preprocessing involves log-normalization with a scaling factor of 10,000 and selection of the top 2,000 HVGs, unless specified otherwise. Method-specific settings are:
Seurat: Raw scRNA-seq and MERFISH matrices are log-normalized using NormalizeData, scaled with ScaleData, and reduced to 50 principal components (PCs) via RunPCA. Anchors are identified using IntegrateLayers with CCAIntegration, followed by UMAP visualization on 30 PCs via RunUMAP.
Harmony: Follows Seurat’s preprocessing, with anchors identified using IntegrateLayers and HarmonyIntegration, followed by UMAP on 30 PCs.
scVI: Implemented per GitHub tutorials, with sc.pp.normalize_total and sc.pp.log1p for preprocessing, followed by scvi.model.SCVI.setup_anndata and scvi.model.SCVI for model initialization.
uniPort: Preprocessed with sc.pp.normalize_total, sc.pp.log1p, and sc.pp.highly_variable_genes for 2,000 HVGs, followed by uniport.batch_scale and uniport.train.
SCALEX: Uses the same preprocessed data as uniPort, processed with the SCALEX function and default parameters, bypassing embedded preprocessing.
gimVI: Implemented per scvi-tools tutorials, trained with GIMVI.train for 200 epochs using default parameters.
MultiMAP: Preprocessed with sc.pp.normalize_total and sc.pp.log1p without scaling, followed by sc.pp.scale, sc.pp.pca, and MultiMAP.Integration with default settings.
Evaluation metrics
To assess SpateCV’s performance in gene imputation and data integration compared to baseline methods, we conduct quantitative evaluations from gene and cellular perspectives. Higher values indicate better performance for PCC, SSIM, ARI, AMI, NMI, Homo, and F1; lower values are better for RMSE and JS.
Gene-level metrics
Gene-level evaluation uses five-fold cross-validation to quantify similarity between predicted and ground-truth ST expression, employing four metrics:
Pearson correlation coefficient (PCC). The PCC value was calculated using the following equation:
![]() |
where
and
are the spatial expression vectors of gene
in the ground truth and the predicted result, respectively;
and
denote their corresponding mean values; and
and
represent the standard deviations.
Structural similarity index measure (SSIM). First, the expression matrix is scaled such that the expression value of each gene lies between 0 and 1:
![]() |
where
denotes the expression of gene
at spot
, and
is the total number of spatial spots. Using the scaled gene expression, the SSIM value is calculated as:
![]() |
Here,
,
,
, and
follow the same definitions as in PCC but apply to the scaled data.
= 0.01 and
= 0.03 are constants, and
is the covariance between the scaled ground truth and predicted expression vectors.
Root mean square error (RMSE). To compute RMSE, we first calculate the z-score of the spatial expression of each gene across all spots, and then apply the following equation:
![]() |
where
and
are the z-scores of gene
at spot
in the ground truth and prediction, respectively.
Jensen–Shannon Divergence (JS). JS divergence uses relative entropy (Kullback–Leibler divergence) to quantify the similarity between two distributions. We first compute the spatial probability distribution of each gene as:
![]() |
Then, the JS divergence for gene
is calculated as:
![]() |
with
![]() |
where
and
are the spatial distribution probability vectors of gene
in the ground truth and predicted results, respectively, and
,
represent the probabilities at spot
.
Cell-level metrics
Cell-level metrics uses five-fold cross-validation to evaluate the ability of SpateCV to facilitate cell-type clustering and preserve cell similarity across modalities. We conduct 50 repeated clustering experiments and calculate the mean and standard deviation of four metrics to evaluate the clustering performance:
Adjusted Rand index (ARI)
![]() |
Where
is the total number of samples,
is the number of samples that belong to class
in the ground truth and are assigned to cluster
in the prediction,
is the total number of samples in class
, and
is the total number of samples in predicted cluster
.
Mutual information (MI)
![]() |
Where
and
denote the entropy of clustering
and the conditional entropy of
given
, respectively.
Normalized mutual information (NMI)
![]() |
Adjusted mutual information (AMI)
![]() |
Where
denotes the expected mutual information between
and
under a random model, and
is the average entropy of the two clusters.
Homogeneity (Homo)
![]() |
where
represents the set of true cell labels and
represents the predicted clustering.
Integration metrics
For data integration, we assess cell-type assignment accuracy and cluster quality using:
ARI, NMI, and F1: A k-Nearest-Neighbor (kNN) classifier, trained with sklearn.neighbors.KNeighborsClassifier on UMAP coordinates and cell-type annotations of reference data, predicts MERFISH cell types. The F1 score is:
![]() |
Where
is precision and
is recall of the kNN classifier.
Silhouette coefficient. The Silhouette coefficient is calculated using the mean intra-cluster distance
and the mean nearest-cluster distance
for each sample:
![]() |
The score is scaled between 0 and 1 by adjusting it to
![]() |
Batch entropy score. The Batch Entropy score, derived from SCALAX and inspired by the “entropy of batch mixing” concept, evaluates the sum of regional mixing entropies at randomly chosen cell locations from different datasets. A higher score indicates better mixing of cells from various datasets, calculated as:
![]() |
where
is the proportion of cells in batch
relative to the total number of cells, and
is the proportion of cells from batch
in a given region. However, in data integration, there is often a trade-off between high batch entropy (strong batch mixing) and high silhouette coefficient (strong cell-type separation), as excessive batch correction may erode biological variation. It calculates only for cell types common across different batches. All results are based on UMAP visualization with parameters n_neighbors = 30 and min_dist = 0.1.
Data preprocessing
Three datasets assess SpateCV’s performance in integrating and imputing gene expression for scRNA-seq and ST data. Standard preprocessing includes log-normalization with a median transcript count scaling factor and log(1+x) transformation, with early stopping (patience of 20–30) to reduce training time.
Mouse hypothalamus scRNA-seq and MERFISH dataset
The scRNA-seq (10×) dataset from the preoptic region of the hypothalamus in six mice is obtained from NCBI GEO accession number GSE11357664. The corresponding MERFISH data and annotations are downloaded from the Dryad repository (https://datadryad.org/dataset/doi:10.5061/dryad.8t8s248). For integration, we select Naive female mouse #1 from the MERFISH data, and use mouse #2 to validate the results of online imputation. In the spatial transcriptomics (ST) data, cell type names are processed, and genes containing NaN or labeled as ‘Blank’ are excluded. Cells labeled as ‘Ambiguous’ or ‘Unstable’ are removed from both datasets. As a result, the final MERFISH dataset contains 64,373 cells, 155 genes, and 9 cell types, while the scRNA-seq data comprises 30,370 cells, 27,998 genes, and 12 cell types. Following standard preprocessing, we normalize total counts per cell using the median transcript count and apply a log(1+x) transformation. To minimize training time, an early stopping mechanism is enabled, with patience set to 20. The model ultimately retains 154 genes that overlap between the ST and scRNA-seq datasets, along with 2,189 genes from the scRNA-seq data. Once training is completed, the latent representations for both the ST and scRNA-seq data are extracted from the model.
Mouse MOp snRNA-seq and MERFISH dataset
The mouse brain MOp dataset is analyzed using the image-based spatial transcriptomics method MERFISH at single-cell resolution. The MERFISH data are downloaded from the brain image library (https://doi.brainimagelibrary.org/doi/10.35077/g.8). This dataset includes expression data for 254 genes across 5,551 single cells, which are categorized into 23 cell types. For reference, we utilize a pre-processed paired droplet-based snRNA-seq profile from mouse MOp, capturing gene expression data from 13,516 cells across 20 distinct cell types. They are downloaded from https://drive.google.com/drive/folders/1PXv_brtr-tXshBVEd_HSPIagjX9oF7Kg. Cells corresponding to the cell types ‘SMC’, ‘L6 IT Car3’, and ‘L4/5 IT’ are removed from the MERFISH dataset, leaving the 20 overlapping cell types that match those present in the snRNA-seq reference. SpateCV is then employed to perform the integration and imputation. We select highly variable genes, genes differentially expressed across cell types, and genes of specific interest from the snRNA-seq data. After deduplication, a total of 1,938 genes are identified as targets for prediction. SpateCV is trained on the 247 genes common to both the MERFISH and snRNA-seq datasets to impute the expression of unmeasured genes in the MERFISH dataset. We benchmark SpateCV’s imputation performance against several methods, including Tangram, gimVI, stPlus, SpaGE, uniport, and ENVI, using their respective default parameters (see the ‘Baselines’ subsection for further details).
Mouse visual cortex scRNA-seq and STARmap dataset
The mouse visual cortex STARmap data is accessible at this link (https://www.dropbox.com/scl/fo/s95c5nr3zjjhn4oc2nz72/AOqeae5ZwyF2vZcvpXlxH8I/visual_1020/20180505_BY3_1kgenes?rlkey=herdyoo1sggk4rzleiowfwufe&subfolder_nav_tracking=1&dl=0). The dataset includes 1,207 cells and 1,020 genes, categorized into 7 cell types. The corresponding scRNA-seq dataset is downloaded from the Allen Institute portal (https://portal.brain-map.org/atlases-and-data/rnaseq/mouse-v1-and-alm-smart-seq), which contains data for 14,249 cells, 34,041 genes, and 23 cell types. After standard preprocessing, a log(1+x) transformation is applied.
Downstream analysis
Cell-type specific spatially differential expression (DE) genes
We use C-SIDE to identify cell-type specific spatially DE genes within the MERFISH dataset, utilizing the function run.CSIDE.nonparametric. The analysis follows the guidelines (https://raw.githack.com/dmcable/spacexr/master/vignettes/merfish_nonparametric.html), with the parameters set as follows: gene_threshold = 0.001, cell_type_threshold = 10, and fdr = 0.01.
Cell-cell communication
For the visual cortex STARmap dataset, ligand-receptor (LR) pairs are identified where both the ligand and receptor are expressed in at least 5% of the cells, using the COMMOT (v0.0.3) package. The analysis follows the guidelines (https://commot.readthedocs.io/en/latest/notebooks/visium-mouse_brain.html#Spatial-communication-inference), with a spatial distance limit of 500 for the ligand-receptor pairs.
Gene ontology enrichment analysis
To functionally annotate gene sets, we perform gene set enrichment analysis using Gene Ontology (GO) biological process (BP) terms. All enrichment analyses are performed using the enrichGO function from the clusterProfiler package with the default “BH” method for P-value multiple testing correction and the default significance level at 0.05.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
Not applicable
Author contributions
J.Y. conceived, designed and performed computational modeling, validations and benchmarks. J.Y analyzed data and validated the results. Z.Y. conceived and supervised the project. P.X. validated and analyzed the results. W.L. and J.Y. drafted the manuscript and revised the manuscript. All authors read and approved the final manuscript.
Funding
This work was supported in part by the National Natural Science Foundation of China (62072128 and 62573143), the Natural Science Foundation of Guangdong Province of China (2023A1515011401), the National key R and D Program of China (2019YFA0706338402).
Data availability
All data analyzed in this article are publicly available through online sources. We have presented references or links to all data sources in Table 1 and “Data preprocessing” section.
Code availability
The SpateCV source code are available in Github (https://github.com/SevensLab/SpateCV). We also uploaded all scripts to reproduce all the analyses at the same website. The code can also be used to analyze user’s own datasets.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Jiaqi Yuan and Junhua Yu contributed equally to this work.
Contributor Information
Zheng Ye, Email: zheng_ye@gzhu.edu.cn.
Peng Xu, Email: gdxupeng@gzhu.edu.cn.
Wenbin Liu, Email: wbliu6910@gzhu.edu.cn.
References
- 1.Aval OS, et al. Galectin-9: a double-edged sword in acute myeloid leukemia. Ann Hematol. 2025;1–14 [DOI] [PMC free article] [PubMed]
- 2.Soleimani Samarkhazan H, et al. Unveiling the potential of CLL-1: a promising target for AML therapy. Biomarker Res. 2025;13(28) [DOI] [PMC free article] [PubMed]
- 3.Aghapour SA, et al. Investigating the dynamic interplay between cellular immunity and tumor cells in the fight against cancer: an updated comprehensive review. Iran J Blood and Cancer. 2024;16:84–101 [Google Scholar]
- 4.Wu AR, et al. Quantitative assessment of single-cell RNA-sequencing methods. Nat Methods. 2014;11:41–46 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Cai L. In 53rd European society of human genetics (ESHG) conference
- 6.Simone, et al. Spatial organization of the somatosensory cortex revealed by osmFISH. Nat Methods. 2018 [DOI] [PubMed]
- 7.Moffitt JR, et al. Molecular, spatial, and functional single-cell profiling of the hypothalamic preoptic region. Science. 2018;362:792–792 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Xiao W, et al. Three-dimensional intact-tissue sequencing of single-cell transcriptional states. Science. 2018;361:eaat5691 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Rao A, Barkley D, Frana GS, Yanai I. Exploring tissue architecture using spatial transcriptomics. Nature [DOI] [PMC free article] [PubMed]
- 10.Cassella L, Ephrussi A. Subcellular spatial transcriptomics identifies three mechanistically different classes of localizing RNAs. 2021 [DOI] [PMC free article] [PubMed]
- 11.Tian L, Chen F, Macosko EZ. The expanding vistas of spatial transcriptomics. Nat Biotechnol. 2023;41(10) [DOI] [PMC free article] [PubMed]
- 12.Cang Z, Nie Q. Inferring spatial and signaling relationships between cells from single cell transcriptomic data. Nat Commun. 2020;11:2084. 10.1038/s41467-020-15968-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Welch JD, Kozareva V, Ferreira A, Vanderburg C, Macosko EZ. Single-cell multi-omic integration compares and contrasts features of brain cell identity. Cell. 2019;177:1873–87.e1817 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Moriel N, et al. NovoSpaRc: flexible spatial reconstruction of single-cell gene expression with optimal transport. Nat Protoc. 2021;16:4177–200. 10.1038/s41596-021-00573-7 [DOI] [PubMed] [Google Scholar]
- 15.Santambrogio F. Optimal transport for applied mathematicians. 2015
- 16.Biancalani T, et al. Deep learning and alignment of spatially resolved single-cell transcriptomes with Tangram. Nat Methods. 2021 [DOI] [PMC free article] [PubMed]
- 17.Fan Z, et al. SPASCER: spatial transcriptomics annotation at single-cell resolution. Nucleic Acids Res. 2022;51:D1138–49. 10.1093/nar/gkac889 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Brink SCVD, Alemany A, Batenburg VV, Moris N, Oudenaarden AV. Single-cell and spatial transcriptomics reveal somitogenesis in gastruloids. Nature. 2020;579:1–1 [DOI] [PubMed] [Google Scholar]
- 19.Baccin C, Al-Sabah J, Velten L, Helbling PM, Haas S. Combined single-cell and spatial transcriptomics reveals the molecular, cellular and spatial bone marrow niche organization. Cold Spring Harbor Lab. 2019 [DOI] [PMC free article] [PubMed]
- 20.Stuart T, et al. Comprehensive integration of single-cell data. Cell. 2019;177:1888–902.e1821 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Tamim A, Soufiane M, Ahmed M, Reinders MJT. SpaGE: spatial gene enhancement using scRNA-seq. Nucleic Acids Res. 2020;18 [DOI] [PMC free article] [PubMed]
- 22.Chen S, Zhang B, Chen X, Zhang X, Jiang R. stPlus: a reference-based method for the accurate enhancement of spatial transcriptomics. Bioinformatics. 2021;Supplement_1 [DOI] [PMC free article] [PubMed]
- 23.Haviv D, et al. The covariance environment defines cellular niches for spatial inference. Nat Biotechnol. 2025;43 [DOI] [PMC free article] [PubMed]
- 24.Lopez R, et al. A joint model of unpaired data from scRNA-seq and spatial transcriptomics for imputing missing gene expression measurements. 2019
- 25.Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single-cell transcriptomics. Nat Methods. 2018 [DOI] [PMC free article] [PubMed]
- 26.Cao K, Gong Q, Hong Y, Wan L. A unified computational framework for single-cell data integration with optimal transport. Nat Commun [DOI] [PMC free article] [PubMed]
- 27.Chunman Z, Hao D, Luonan C. Deep cross-omics cycle attention model for joint analysis of single-cell multi-omics data. Bioinformatics. 2021;22 [DOI] [PubMed]
- 28.Gong B, Zhou Y, Purdom E. Cobolt: integrative analysis of multimodal single-cell sequencing data. Genome Biol. 2021;22:351 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Zhang Z, Yang C, Zhang X. Learning latent embedding of multi-modal single cell data and cross-modality relationship simultaneously. Cold Spring Harbor Lab. 2021
- 30.Lotfollahi M, et al. Mapping single-cell data to reference atlases by transfer learning. Nat Biotechnol. 2022;40:121–30. 10.1038/s41587-021-01001-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Zou G, Shen Q, Li L, Zhang S. stAI: a deep learning-based model for missing gene imputation and cell-type annotation of spatial transcriptomics. Nucleic Acids Res. 2025;53:gkaf158 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Li X, Zhu F, Min W. SpaDiT: diffusion transformer for spatial gene expression prediction using scRNA-seq. Briefings Bioinf. 2024;25:bbae571 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sohn K, Yan X, Lee H, Arbor A. International conference on neural information processing systems
- 34.Kingma DP, Welling M. Auto-encoding variational Bayes. arXiv.org. 2014
- 35.Zeisel A, et al. Cell types in the mouse cortex and hippocampus revealed by single-cell RNA-seq. Science. 2015;347:1138–42. 10.1126/science.aaa1934 [DOI] [PubMed] [Google Scholar]
- 36.Traag VA, Waltman L, van Eck NJ. From Louvain to Leiden: guaranteeing well-connected communities. Sci Rep. 2019;9:5233. 10.1038/s41598-019-41695-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Xiong L, et al. Online single-cell data integration through projecting heterogeneous datasets into a common cell-embedding space. Nat Commun. 2022;13(17) [DOI] [PMC free article] [PubMed]
- 38.Jain MS, Conde CD, Polanski K, Chen X, Teichmann SA. MultiMAP: dimensionality reduction and integration of multimodal data. Genome Biol. 2021;22 [DOI] [PMC free article] [PubMed]
- 39.Luecken MD, et al. Benchmarking atlas-level data integration in single-cell genomics. Nat Methods. 2022;19:41–50 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Li H, McCarthy DJ, Shim H, Wei S. Trade-off between conservation of biological variation and batch effect removal in deep generative modeling for single-cell transcriptomics. BMC Bioinf. 2022;23:460. 10.1186/s12859-022-05003-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Rautenstrauch P, Ohler U. Liam tackles complex multimodal single-cell data integration challenges. Nucleic Acids Res. 2024;52:e52–52 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Zhang M, et al. Spatially resolved cell atlas of the mouse primary motor cortex by MERFISH. Nat Publishing Group. 2021 [DOI] [PMC free article] [PubMed]
- 43.Lein ES, Hawrylycz MJ, Ao N, Ayres M, Jones AR. Genome-wide atlas of gene expression in the adult mouse brain. Nature. 2007;445:168–76 [DOI] [PubMed] [Google Scholar]
- 44.Cable DM, et al. Cell type-specific inference of differential expression in spatial transcriptomics. Nat Methods. 2022;19:1076–87. 10.1038/s41592-022-01575-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Lorincz MT, Zawistowski VA. Expanded CAG repeats in the murine Huntington’s disease gene increases neuronal differentiation of embryonic and neural stem cells. Mol Cell Neurosci. 2009;40:1–13. 10.1016/j.mcn.2008.06.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Pugazhenthi S, Wang M, Pham S, Sze CI, Eckman CB. Downregulation of CREB expression in Alzheimer’s brain and in Aβ-treated rat hippocampal neurons. Mol Neurodegener. 2011;6(60). 10.1186/1750-1326-6-60 [DOI] [PMC free article] [PubMed]
- 47.Moh MC, Lee LH, Shen S. Cloning and characterization of hepaCAM, a novel Lg-like cell adhesion molecule suppressed in human hepatocellular carcinoma. J Hepatol. 2005;42:833–41 [DOI] [PubMed] [Google Scholar]
- 48.Baldwin KT, et al. HepaCAM controls astrocyte self-organization and coupling. Neuron. 2021;109:2427–42. e2410 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Tan B, et al. HepaCAM inhibits clear cell renal carcinoma 786–0 cell proliferation via blocking PKCε translocation from cytoplasm to plasma membrane. Mol Cellular Biochem. 2014;391:95–102 [DOI] [PubMed] [Google Scholar]
- 50.Jin S, et al. Inflammatory cytokines disrupt astrocyte exosomal HepaCAM-mediated protection against neuronal excitotoxicity in the SOD1G93A ALS model. Sci Adv. 2024;10:eadq3350 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Bahreiny SS, Ahangarpour A, Rajaei E, Sharifani MS, Aghaei M. Meta-analytical and meta-regression evaluation of subclinical hyperthyroidism’s effect on male reproductive health: hormonal and seminal perspectives. Reprod Sci. 2024;31:2957–71 [DOI] [PubMed] [Google Scholar]
- 52.Sherafat NS, et al. Rationale of using immune checkpoint inhibitors (ICIs) and anti-angiogenic agents in cancer treatment from a molecular perspective. Clin Exp Med. 2025;25:238 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Tasic B, et al. Shared and distinct transcriptomic cell types across neocortical areas. 2017 [DOI] [PMC free article] [PubMed]
- 54.Cang Z, et al. Screening cell-cell communication in spatial transcriptomics via collective optimal transport. Nat Methods. 2023;20:218–28. 10.1038/s41592-022-01728-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Yao Z, Nguyen TN, Velthoven CTJV, Goldy J, Zeng H. A taxonomy of transcriptomic cell types across the isocortex and hippocampal formation. Cold Spring Harbor Lab. 2020 [DOI] [PMC free article] [PubMed]
- 56.Song YH, Yoon J, Lee SH. The role of neuropeptide somatostatin in the brain and its application in treating neurological disorders. Exp Mol Med [DOI] [PMC free article] [PubMed]
- 57.Wu SJ, et al. Cortical somatostatin interneuron subtypes form cell-type-specific circuits. Neuron. 2023;111:2675–92.e2679. 10.1016/j.neuron.2023.05.032 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Hafemeister C, Satija R. Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. Cold Spring Harbor Lab. 2019 [DOI] [PMC free article] [PubMed]
- 59.Cable DM, et al. Robust decomposition of cell type mixtures in spatial transcriptomics. Nat Biotechnol. 2022;40:517–26. 10.1038/s41587-021-00830-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Lohoff T, et al. Integration of spatial and single-cell transcriptomic data elucidates mouse organogenesis. Nat Biotechnol. 2022;40:74–85 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Takei Y, et al. Integrated spatial genomics reveals global architecture of single nuclei. Nature. 2021;590:344–50 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Han X, et al. Mapping the mouse cell atlas by microwell-seq. Cell. 2018;172:1091–107. e1017 [DOI] [PubMed] [Google Scholar]
- 63.Shah S, Lubeck E, Zhou W, Cai L. In situ transcription profiling of single cells reveals spatial organization of cells in the mouse hippocampus. Neuron. 2016;92:342–57 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Joglekar A, et al. Cell-type, single-cell, and spatial signatures of brain-region specific splicing in postnatal development. BioRxiv. 2020;2008.2027 268730
- 65.Eng C-HL, et al. Transcriptome-scale super-resolved imaging in tissues by RNA seqFISH+. Nature. 2019;568:235–39 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Xia C, Fan J, Emanuel G, Hao J, Zhuang X. Spatial transcriptome profiling by MERFISH reveals subcellular RNA compartmentalization and cell cycle-dependent gene expression. Proc Natl Acad Sci. 2019;116:19490–99 [DOI] [PMC free article] [PubMed]
- 67.Zhou Y, et al. Single-cell RNA landscape of intratumoral heterogeneity and immunosuppressive microenvironment in advanced osteosarcoma. Nat Commun. 2020;11:6322 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Li B, et al. Benchmarking spatial and single-cell transcriptomics integration methods for transcript distribution prediction and cell type deconvolution. Nat Methods. 2022;19:662–70 [DOI] [PubMed] [Google Scholar]
- 69.Wang X, et al. Three-dimensional intact-tissue sequencing of single-cell transcriptional states. Science. 2018;361:eaat5691 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Joglekar A, et al. A spatially resolved brain region-and cell type-specific isoform atlas of the postnatal mouse brain. Nat Commun. 2021;12:463 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Hodge RD, et al. Conserved cell types with divergent features in human versus mouse cortex. Nature. 2019;573:61–68 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data analyzed in this article are publicly available through online sources. We have presented references or links to all data sources in Table 1 and “Data preprocessing” section.
The SpateCV source code are available in Github (https://github.com/SevensLab/SpateCV). We also uploaded all scripts to reproduce all the analyses at the same website. The code can also be used to analyze user’s own datasets.

































