Skip to main content
American Journal of Human Genetics logoLink to American Journal of Human Genetics
. 2025 Sep 17;112(11):2739–2750. doi: 10.1016/j.ajhg.2025.08.021

Unveiling tissue heterogeneity through genomic interaction-encoded image representation of RNA-sequencing data

Junyan Liu 1, Zixia Zhou 1, Yizheng Chen 1, Md Tauhidul Islam 1,∗, Lei Xing 1,∗∗
PMCID: PMC12609178  NIHMSID: NIHMS2110666  PMID: 40967222

Summary

Genomic sequencing is essential for both biomedical research and clinical practice. While single-cell RNA sequencing (scRNA-seq) provides insights into biological processes at the cellular level, bulk RNA sequencing remains widely used for its scalability and cost-effectiveness. To explore biological heterogeneity, research efforts have been made toward inferring single-cell-like cellular compositions from bulk samples, i.e., deconvolving bulk samples into multiple cell types. However, existing deconvolution methods face two major limitations: (1) reliance on predefined gene signature matrices without accounting for inter-sample variability and (2) susceptibility to noise within biological systems. Here, we propose a cellular-component analysis (CCA) framework by leveraging a genomic-interaction-encoded image representation of RNA-seq data for substantially improved pattern discovery. The framework incorporates sample-specific gene-expression variability and derives signature patterns by utilizing a convolutional variational autoencoder and Gaussian mixture model. An image-domain linear decomposition of bulk RNA-seq data based on these sample-specific, interpretable gene-signature patterns is then performed for CCA and other downstream tasks, such as cancer subtype classification and biomarker discovery. We demonstrate that the proposed technique improves decomposition accuracy by over 14.1% in average Pearson correlation compared to existing techniques by using both simulation and experimental datasets. This approach offers an effective solution for tissue heterogeneity analysis and lays a foundation for a range of clinical and biological applications.

Keywords: RNA sequencing, decomposition, deep learning

Graphical abstract

graphic file with name fx1.jpg


This study analyzes RNA-sequencing data by converting gene-expression profiles into images that encode gene-gene interactions. This image-based representation, processed through deep learning to extract genetic patterns, significantly improves the accuracy of decomposing cellular compositions in samples while reducing susceptibility to noise compared to traditional methods.

Introduction

Single-cell RNA sequencing (scRNA-seq) has significantly advanced our understanding of cellular heterogeneity, enabling the detection of rare cell types and biomarkers.1,2 Despite these advances, bulk RNA-seq remains indispensable due to its cost-effectiveness and well-established protocols.3 However, bulk RNA-seq provides only average gene-expression profiles (GEPs), obscuring tissue heterogeneity and the contribution of individual cells. To address this limitation, computational models, known as bulk RNA decomposition,4 have been developed to estimate cellular compositions from bulk RNA datasets by assuming that the bulk sample represents a linear combination of the scRNA-seq expression profiles from the constituent cells. This approach is attractive, as it provides valuable insights into cell composition while reducing the need for resource-intensive scRNA-seq experiments. However, the decomposition is challenging due to the absence of sample-specific GEPs for each of the constituent cell types, collectively known as “signature matrices,” as well as the noisy nature of genomic measurements. Currently, predefined gene-signature matrices derived from existing scRNA-seq are used as the basis for decomposition without accounting for inter-sample heterogeneity5 (Figure 1A, workflow of existing approach). Statistical techniques such as support vector regression (SVR)6 or non-negative least squares (NNLS) are then applied to solve the inverse problem and determine the respective cellular proportions in classical algorithms such as CIBERSORTx7 and Bisque.8

Figure 1.

Figure 1

Existing and genoMap-based cell-component analyses

(A) Existing models use known scRNA-seq data from the public domain to derive reference signature matrices by selecting differentially expressed genes for each cell type. These signature matrices representing gene-expression profiles (GEPs) are then applied across all bulk RNA samples for cell-component analysis (CCA).

(B) To derive sample-specific signature genoMaps, we preprocess public scRNA-seq reference cell samples into genoMaps, which encode gene-gene interactions of these cells into the spatial context of images. A convolutional variational autoencoder (VAE) is then used to reduce the dimensionality of the genoMaps into latent space.

(C) The features in latent space are used to fit Gaussian mixture models (GMMs), which learn the distribution of gene expressions of reference cell samples in the feature domain to generate genoMaps with different Gaussian probability. Utilizing the GEPs generated in as decomposition basis, we perform linear image decomposition by multiplying these GEPs with corresponding weights and summing the results, The sample-specific GEPs and the weights are iteratively optimized simultaneously during calibration, with the objective of minimizing the discrepancy between predicted and experimentally measured bulk RNA-seq while maximizing the log likelihood from the GMM latent space.

In this study, both the reference scRNA-seq and input bulk RNA-seq data are generally affected by random noise and measurement uncertainties, which also impact the resultant decompositions. To address this challenge, we investigate a cellular-component analysis (CCA) strategy focusing on groups of genes that exhibit strong correlations, as they are more likely to represent meaningful biological signals. In practice, the observation that two variables are highly correlated often indicates a meaningful relationship rather than random variation. This principle has been widely employed to enhance signal-to-noise ratios in various applications, such as positron emission tomography and two-photon microscopy. In cellular systems, leveraging gene-gene interactions enables the filtering of noise to detect stable, biologically relevant signals9 instead of transient fluctuations that do not contribute to predictive accuracy. Along this line, we take a radically different approach by representing scRNA-seq data in terms of genoMap, which transforms high-dimensional (1D) gene-expression data of individual cells into configured images, with the gene-gene interactions encoded within their spatial context.9 Note that “spatial” in this work refers to the positions that genes occupy within the two-dimensional (2D) images and not to tissue coordinates as in spatial transcriptomics. In our genoMap-based CCA (gCCA), both the search for sample-specific signature genoMaps and CCA are performed in the image domain (Figures 1B–1C). Specifically, we apply a convolutional variational autoencoder (VAE10) (Figure 1B) to the genoMaps constructed from the scRNA-seq reference cells. Gaussian mixture models (GMMs11) are then applied to the feature space in the VAE’s bottleneck layer to identify the GEPs (i.e., signature genoMaps) specific to the sample. In the final step (Figure 1C), CCA determines the cellular proportions by minimizing the difference between the estimated and experimentally measured genoMaps.

In conventional deconvolution methods, such as CIBERSORT, the cellular composition is obtained by deconvolving the bulk sample into a linear combination of the profiles of different cell types/states, which is typically done via least-square optimization by minimizing gene-wise residuals. Because each gene is treated as an independent variable, the inference process is susceptible to noise, dropouts, and outliers in individual gene measurements. In contrast, our approach represents GEPs as 2D images with consideration of inter-relationships among the genes via the use of genoMap, where co-expressed genes are positioned in such a way that their spatial proximity in the image reflects the underlying correlative relationships. A salient feature of the image representation is that rather than relying on individual pixels (genes), this representation emphasizes global, multi-gene spatial patterns. Computationally, the image patterns or deep interactive features can be effectively extracted via the use of convolutional neural networks, enabling the identification of higher-order expression signatures specific to cell types. By leveraging gene-interaction-informed image representations, our framework promotes more robust feature extraction within a deep-learning model, with much reduced susceptibility to technical variation and measurement noise. As a result, the image-based operations reduce the adverse influence of noise and lead to solutions that are more robust against the fluctuations in gene counts. The performance of gCCA is evaluated using five simulations and five experimental datasets. The results demonstrate that the proposed technique achieves an average improvement of 14.1% (from 5.87% up to 51.4% improvement on real datasets) in Pearson correlation compared to existing methods.

Methods

GenoMap visualization

The “genoMap” algorithm9 is an entropy-based cartographic method for visualizing high-dimensional gene-expression data, explicitly incorporating gene-gene interactions into image formats. Initially, genoMap computes a gene-gene interaction matrix derived from scRNA-seq data (as shown in Figure 2A). Genes are then projected onto 2D genoMaps using a projection matrix optimized via an entropy-based transport algorithm (for the detailed methodology of genoMap, please refer to our previous work9). For bulk RNA-expression datasets, the projection matrix learned from scRNA-seq is directly applied to ensure consistent and biologically meaningful gene positioning. The “genoMap” algorithm transforms gene-expression arrays into schematically meaningful images according to gene-gene interactions, thereby encoding biologically relevant “positions” into data, as the gene order in the original count matrix lacks an inherent biological meaning. It positions highly interactive genes toward the center of the images and places less-interactive genes peripherally. In this work, we selected 1,600 differentially expressed genes across annotated cell types from scRNA-seq data to create 40 × 40 genoMaps using the Python package Scanpy via functions “highly_variable_genes” and “rank_genes_groups.” This gene number was chosen to capture sufficient biological information while maintaining computational efficiency, ensuring that genoMaps could be generated in under 5 min.

Figure 2.

Figure 2

genoMap-based batch correction

(A) The genoMap algorithm constructs a gene-gene interaction matrix based on co-expression patterns within the reference single-cell population. Based on the interaction matrix, genes are then spatially organized into a 2D image, with highly interacting genes positioned centrally and weakly interacting genes placed toward the periphery, resulting in a structured image representation of the transcriptome.

(B) genoNet, a convolutional neural network trained on genoMaps, is used to perform image classification. By combining this with Shapley value analysis, the model identifies informative genes from an image-based perspective, highlighting biomarkers that contribute most to the spatial expression patterns.

(C) We reconstruct pseudo-bulk profiles by aggregating reference single cells according to the known proportions for each training sample. An autoencoder is then trained to map real bulk measurements (inputs) to their corresponding pseudo-bulk profiles (outputs). Once trained, the model weights are fixed and the autoencoder is applied to test samples.

(D) The challenge of batch effects in real-world decomposition tasks is highlighted. Before correction, bulk samples (blue and pink dots) fail to align with pseudo-bulks generated from reference cells (purple dots), as shown by sample “1.” In the right panel, non-corrected measurements are plotted against a pseudo-bulk derived by randomly sampling reference cells according to the known ground truth, revealing a violation of the linear assumption. To address this issue, we employ a leave-one-out strategy in which all remaining bulk datasets (blue) and their corresponding pseudo-bulks (green) are used to correct the batch effects in sample “1” (pink). After correction, the linear relationship between bulk and pseudo-bulk data is restored, as demonstrated by sample “2” (orange).

Variational autoencoder and Gaussian mixture models

We employ a standard convolutional VAE model to obtain latent representations for the scRNA-seq reference cells. Both the encoder and decoder consist of three convolutional layers with 3 × 3 kernels. A multivariate Gaussian distribution is imposed on the latent space through the Kullback-Leibler divergence term, in line with the original VAE formulation. Subsequently, we fit a separate GMM to the latent representations of each cell type via Scikit-learn. Once trained, all parameters of the VAE and GMM remain fixed for further analyses. Throughout the text, we refer to the VAE decoder as the “decoder” and the latent variables as “z.”

Bulk RNA decomposition model

The latent space, consisting of latent variables z, can be viewed as a continuous, lower-dimensional manifold that preserves the essential features of the genetic data. Using the VAE decoder, we project z back into the scRNA-seq image space and compute its log likelihood via the GMM. These reconstructed GEPs are then multiplied by weights corresponding to each cell type. Due to the linear assumption (i.e., bulk samples are as linear combinations of the scRNA-seq reference cell samples), we sum the resultant GEPs to obtain the predicted pseudo-bulk.

loglikelihood=∑i=1celltypeslog(GMM(zi)),
pred=∑i=1celltypes(decoder(zi)·ki).

The loss function comprises two components: a log-likelihood term from the GMM that measures the probability of each GEP, and a mean squared error (MSE) term that quantifies the discrepancy between the model prediction and the actual bulk measurements. The hyperparameter γ = 0.01 normalizes these terms to a comparable scale.

Loss(z,k)=−γ·loglikelihood+MSE,
MSE=1imagesize∑(pred−measure)2.

During optimization, the goal is to simultaneously minimize the prediction error (i.e., MSE) and maximize the GMM log likelihood, ensuring accurate data reconstruction while increasing the likelihood of the inferred GEPs. Specifically, we initiate the optimization process by setting each cell type’s latent variables at the center of the corresponding GMM distribution, effectively using population-average GEPs, an approach similar to conventional models. As optimization progresses, if the reconstruction MSE remains large, the model adaptively shifts the latent variables toward the periphery of the GMM distributions, yielding sample-specific GEPs. Ultimately, the process achieves a balance between the MSE term and log-likelihood term. Because of the linear assumption, we can directly derive cell proportions from the fitted parameters k. The fraction of cell type i, denoted as fi, is obtained by normalizing each weight ki such that their sum equals 1.

fi=ki∑ki.

Data and preprocessing

All the datasets used in this study are public and can be accessed as follows. The peripheral blood mononuclear cell (PBMC) scRNA-seq datasets are downloaded from the 10× Genomics official website under the name “6k PBMCs from a Healthy Donor” and “8k PBMCs from a Healthy Donor.” Breast cancer datasets are available in the original study12 and can also be accessed through the Gene Expression Omnibus (GEO) under accession number GEO: GSE176078. PBMC s13 bulk datasets can be accessed under GEO: GSE107011, with ground truth provided in the supplemental material of the original study.13 The raw brain tissue ROSMAP datasets are available on Synapse upon user request (Synapse: syn2580853). The processed data files can be accessed from the GitHub repository provided in the original study.14 The second breast cancer dataset can be accessed under GEO: GSE176078. The COVID-19 PBMC datasets can be accessed under GEO: GSE216529 and GSE132228. Detailed procedures for accessing the bulk datasets and corresponding reference data are provided in supplemental methods: datasets summary.

Real-world datasets that provide matched bulk RNA-seq and cellular composition data from the same samples are scarce, primarily owing to the substantial time and cost associated with additional flow cytometry or scRNA-seq assays. Consequently, decomposition studies predominantly rely on simulation data. In this study, we use a total of ten datasets: three from Splatter simulations, two from PBMC simulations, and five experimentally measured datasets including two PBMC (n = 12 and 85 samples, respectively), two breast cancer (n = 24 and 18 samples, respectively), and one brain tissue dataset (n = 49 individuals).

For scRNA-seq preprocessing and normalization, we employ the Python package Scanpy15 with its default parameters, normalizing each cell’s total counts to 10,000 across all genes. In the non-batch mode, the normalized data are directly input into the convolutional VAE. The batch-mode procedure is described separately below. Detailed information on model runtime and memory usage for each real dataset is provided in supplemental methods: computational time and GPU memory usage.

Shapley value

We employ Shapley value,16 a concept from game theory, to quantify each gene’s contribution (i.e., pixels) to the cell-type-specific patterns. The Shapley value quantifies each input feature’s contribution to the model output, making it possible to evaluate each pixel’s (i.e., gene’s) contribution on the emergence of cell-type-specific image patterns. Specifically to identify potential biomarkers, we applied the genoNet classification model9 (convolutional deep-learning model, as proposed in our original study9 and shown in Figure 2B) to our reference genoMaps, followed by the GradientShap algorithm in the Python package Captum17 to compute the Shapley value of each pixel in the image. Genes associated with the highest Shapley values were designated as candidate biomarkers. This allows us to identify the most informative genes based on image patterns.

Batch correction

In batch mode, we implemented a leave-one-out strategy: for each test sample, cell proportions of all other bulk samples are assumed known and used to correct batch effects in the test sample. This scenario differs from classical batch-correction problems, as it can leverage single-cell references to reconstruct pseudo-bulk expression from known proportions. To implement this, we create 100 pseudo-bulk samples for each training bulk sample by randomly sampling and aggregating single-cell reference data according to their known cell proportions. An autoencoder (as shown in Figure 2C) is trained using the real bulk sample as input and the corresponding pseudo-bulk profiles as outputs, enabling the model to learn batch-effect correction mappings. Subsequently, we fix the autoencoder weights and run the model forward on the test sample to obtain its batch-corrected profile. This design prevents data leakage and mirrors realistic scenarios in which a subset of bulk samples can be reliably characterized and subsequently utilized for batch-effect correction of the remaining samples. This correction step operates independently from the main model, providing flexibility to incorporate alternative batch-correction methods if desired. In real-world experiments, huge batch effects were identified between the reference and bulk samples (Figure 2D), which, if uncorrected, would violate the linear combination assumption (i.e., the bulk samples cannot be modeled as linear combination of reference GEPs).

Comparative model implementation

We evaluated CIBERSORTx using its online platform with default parameters in two modes: B mode, designed for settings with minimal batch effects, and S mode, which corrects for substantial batch effects when corresponding scRNA-seq data are available. For Bisque, Scaden, and MuSiC,18 we trained the models locally using their published packages and default settings. Notably, CIBERSORTx (in B mode and S mode), Bisque, and MuSiC incorporate built-in batch-correction methods. Since the original Scaden model does not specify a batch-correction procedure, we followed its official tutorial and did not implement additional corrections. To enable future reference and comparison, Table S1 presents all model prediction results for the real bulk datasets alongside their corresponding ground-truth values.

Results

GenoMap visualization of reference cell samples

In contrast to conventional techniques that treat genes as isolated entities, our framework captures gene inter-correlations across the entire cell population. This is achieved by strategically encoding gene-gene interactions during the construction of genoMaps (Figure 1B). The resulting spatial arrangement (Figure 3A) reveals unique gene signatures or gene-gene interaction patterns for each cell type. These emergent patterns lay a robust foundation for highly accurate gCCA, as the genoMap representations of each cell type serve as an image basis for decomposition tasks. Moreover, it also offers an intuitive visualization of GEPs as images, thereby overcoming the limitations of one-dimensional count array, where the ordering of genes lacks biological meaning and fails to capture intergenic relationships.

Figure 3.

Figure 3

Data preprocessing and VAE latent space

(A) After converting scRNA-seq datasets into genoMaps, genes with similar interactions are positioned closely in the image domain, forming cell-type patterns.

(B) The VAE generated latent space is projected onto two dimensions for visualization. Here, we present four immune cell types. The axes are labeled “gUMAP” to denote that the uniform manifold approximation and projection (UMAP)19 algorithm has been applied to the genoMap domain rather than to raw data. Note that UMAP is used solely for visualization and plays no role in the actual model fitting or GMM estimation.

(C) Each cell type in the latent space is further quantified by fitting a GMM. We illustrate the GMM for monocytes, where the center (labeled “1”) shows the average GEP corresponding to high log likelihood. Outlying regions depict heterogeneous GEPs of lower log likelihood (labeled “2” and “3”). These latent variables can be mapped back to the genoMap domain via the VAE decoder (right panel), denoted as sig1, sig2, and sig3, corresponding to latent variables 1, 2, and 3, respectively.

After the transformation into genoMaps, VAEs are applied to scRNA-seq reference cells to derive low-dimensional, probabilistic representations. These low-dimensional latent variables effectively capture distinctive cell-type-specific signatures, consistent with prior work of applying VAE in cell clustering.20,21 Specifically, we focus on four immune cell types from a PBMC dataset and visualize the latent space (Figure 3B). This latent representation is further fit by a GMM for each cell type. An enlarged view of the monocyte population (Figure 3C) highlights the GMM center (labeled as “1”), which corresponds to the average monocyte GEP. Anchoring all GEPs at the GMM centers effectively mirrors classical decomposition methods, which rely on average GEPs. Additionally, this latent space captures a broader spectrum of GEPs across the entire distribution, including those low-likelihood profiles that deviate from the population average (e.g., labeled as “2” and “3” in Figure 3C). To assess the necessity of considering the entire distribution instead of relying on the average, we performed an ablation study by replacing the VAE-GMM component with average GEPs only while keeping all other parts of the model unchanged (see supplemental methods: ablation study). The observed decrease in decomposition accuracy confirms that inter-sample heterogeneity can adversely affect model performance when relying only on predefined, average GEPs.

Accurate cell-type estimation

To evaluate model performance, we conducted benchmarking experiments against four established methods: CIBERSORTx, the state-of-the-art SVR model; Bisque, an NNLS regression model; Scaden, an ensemble deep-learning framework combining three fully connected neural networks; and MuSiC, a weighted NNLS-based regression model (Figures 4A and 4B). Pearson correlation between the predicted and ground-truth cellular compositions was used as a benchmarking metric. The gCCA operates in two modes: a “non-batch” mode that directly performs deconvolution on the normalized count data without explicit batch-effect correction, and a “batch” mode that employs a subset of bulk samples with known cell proportions to correct for batch effects. Across all datasets, both modes consistently demonstrated either superior (p < 0.01) or comparable performance. Notably, even in non-batch mode, gCCA demonstrates accuracy comparable to methods that incorporate batch correction (e.g., CIBERSORTx and Bisque), verifying its robustness to batch effects. The decomposition results are presented in Table S1.

Figure 4.

Figure 4

Model performance comparison for each sample

We conducted a comprehensive comparison between genoMap-based CCA (gCCA) (operating in both batch and non-batch modes) and four methods: CIBERSORTx, Scaden, MuSiC, and Bisque. Each point represents one PCC between the predictions and ground truth. Model performance was quantified using PCC on both (A) simulation and (B) real-world bulk RNA-seq datasets using boxplots, with median PCC values labeled in each box. Statistical significance was determined via t test to evaluate whether one model significantly outperformed another.

For the simulation study, we employed the Splatter package to generate scRNA-seq data, creating training (reference) and testing (evaluation) datasets under three conditions: (1) ideal (no dropout or batch effects), (2) dropout only, and (3) batch effect only. The pseudo-bulk samples were then created by random sampling and aggregation of scRNA-seq data. Under ideal conditions, CIBERSORTx marginally outperformed gCCA, but our model’s robustness was evident under dropout and batch-effect conditions—crucial for reliable predictions when data quality is compromised. We extended our simulation to two public PBMC datasets (i.e., PBMC6k and PBMC8k) from 10× Genomics. This analysis assessed the impact of reference selection on model accuracy. Specifically, we conducted two complementary analyses: (1) PBMC6k as the reference and PBMC8k for evaluation; and (2) PBMC8k as the reference and PBMC6k for evaluation. Both gCCA and CIBERSORTx exhibited remarkable consistency in performance, demonstrating stability regardless of which dataset was chosen as the reference. Our approach demonstrates higher accuracy in decomposition tasks, effectively addressing challenges such as dropouts and missing values.22,23 The improved performance suggests that while the dropout or missing values may alter individual gene-expression levels, the broader image patterns formed by gene interactions remain largely preserved.

Finally, we evaluate gCCA using five real-world bulk datasets—two breast cancer datasets,12,24 two PBMC datasets,13,25 and a brain tissue dataset14—selected for their known ground truth determined by flow cytometry, scRNA-seq, or immunohistochemistry in the respective original studies. We implemented the “batch” mode and performed leave-one-out cross-validation, where all samples with known cellular proportions except one were used as training set, to correct the batch effects of the excluded evaluation sample. A detailed comparison of the real experiment predictions versus the ground truth for each model is provided in Figure 5A. Meanwhile, Conventional decomposition models typically yield only point estimates, thereby neglecting the inherent uncertainty in their predictions. To address this limitation, we partitioned the genoMaps into patches comprising correlated genes, ensuring that the convolution of these genes is preserved within each patch. In each iteration, only a randomly selected subset of these patches is activated (Figure 5B), thereby engaging only a fraction of the overall genes while maintaining the structural integrity of gene correlations. Repeating this procedure yields a quantitative measure of model uncertainty, ultimately enabling users to assess the confidence of each prediction.

Figure 5.

Figure 5

Model prediction for each cell type

(A) The Pearson correlation coefficient (PCC) between the model prediction and the corresponding ground truth for three representative experiments on five different models (gCCA batch mode, CIBERSORTx B mode, CIBERSORTx S mode, Scaden, and Bisque). The diagonal line indicates perfect agreement between prediction and ground truth, while the PCC, displayed in the bottom right corner of each plot, quantifies this alignment. A linear regression line, along with its bootstrap confidence intervals shown in purple, further illustrates the relationship. For the PBMC COVID-19 dataset, each bulk sample corresponds to a sorted cell population dominated by a single cell type. Given this binary composition, visualization is not meaningful and therefore not shown in this figure. Notably, gCCA batch mode achieves the highest overall correlation.

(B) We partitioned the genoMaps into patches and randomly activated only a subset of these patches during each run. Predictions from real experimental data (right panel) exhibit larger error bars than those observed in simulations (left panel), likely due to misalignment between the bulk and scRNA-seq reference profiles, even after model correction.

Model interpretability aids biomarker identification and cancer subtype classification

The gCCA is interpretable at both gene and individual levels. We demonstrate interpretability at gene level using the breast cancer dataset. High Shapley values indicate gene biomarkers, and if such values are assigned to non-critical genes, they may indicate model errors and unreliable predictions. This approach performs biomarker discovery from the image domain and allows users to verify whether the model focuses on biologically relevant genes. In Figure 6A, we list the genes with the highest Shapley values for each cell type and examine their functions. For example, ELF3,26 MUC1,27 PPDPF,28 ALDOA,29 MGST1,30 and ASPH31 are either specific to breast cancer or broader oncogenic processes. PFN1,32 CREM,33 IL7R,34 CD2,35 CD96,36 BTG1,37 and CORO1A38 are implicated in T cell activation, regulation, and dynamics.

Figure 6.

Figure 6

Model interpretability

(A) Shapley values for each cell type are plotted, with the genes exhibiting the highest Shapley values listed below.

(B) The ability of our model to predict sample-specific GEPs enables the classification of cancer subtypes from bulk RNA-seq samples. For each of the four individuals analyzed, the model initializes at average GEP (purple dot in the center). The final predictions (black dots) are validated against ground truth derived from experiments (right, top panel). In addition, the GEP for each individual is presented in the bottom panel, where the VAE decoder converts latent variables back into the image domain. These GEPs are sample specific, reflecting individual variability (outlined in the green box) despite sharing a common overall pattern characteristic of cancer cell types.

(C) When decomposing the PBMC dataset from COVID-19 samples, we observed a markedly low prediction accuracy for CD4+ T cells. To interpret the source of this error, we examined the VAE latent space of reference cells (labeled by transparent dots) and found substantial overlap between the prediction for CD4+ and CD8+ T cells (labeled by solid large dots). In addition, we visualized the average genoMaps for CD4+, CD8+, and B cells (right panel) and observed visual similarity between the CD4+ and CD8+ T genoMaps. This resemblance likely causes the model’s difficulty in distinguishing these subsets, thereby explaining the reduced accuracy for CD4+ T cell predictions.

Interpretability at sample level further enables cancer subtype classification. Figure 6B depicts the latent space of breast cancer cells, including three subtypes: luminal A, luminal B, and basal-like.39 During calibration, the model’s weights are initialized at the GMM center (Figure 6B, purple) for all four individuals and are progressively refined by adapting to each person during optimization. Cancer subtypes can be inferred by comparing their final predicted latent variable (Figure 6B, depicted as black circles) against the population average. For instance, cid4290A’s latent variable shifts from the population average toward the luminal A cells, indicating a luminal A subtype. These model-based predictions (Figure 6B, black circles) are compared with experimentally measured ground truths (Figure 6B, right panel). While the model accurately identified the main subtypes in all of four individuals, it fails to detect the dual subtypes such as cid4067, indicating an area for future improvement. In another independent PBMC dataset from COVID-19 samples (Figure 6C), we observed a notably lower decomposition accuracy for CD4+ T cells (Pearson correlation coefficient [PCC] = 47.50%) compared to B cells (PCC = 99.81%) and CD8+ T cells (PCC = 89.69%). To investigate the source of this discrepancy, we visualized the model-fitting results alongside the reference single cells in the VAE latent space. This analysis revealed substantial overlap between CD4+ and CD8+ T cell predictions, suggesting that the model tends to misclassify CD4+ T cells as CD8+ T cells, thereby leading to the reduced accuracy. To further support this interpretation, we plotted the average genoMaps for the three cell types and found that CD4+ and CD8+ T cells exhibit highly similar spatial expression patterns, confirming their GEP similarity.

Discussion

In this study, we present gCCA, a broadly applicable genoMap-based deep-learning model for bulk RNA decomposition, which predicts both cellular proportions and sample-specific signature genoMaps. By encoding the gene-gene interactions in terms of schematically meaningful image representation, this approach greatly enhances the robustness against noise within RNA-seq data, which can be susceptible to dropout and batch effects. Instead of relying on predefined population averages, gCCA generates sample-based GEPs from the distribution of reference cells. In reality, genetic data, similar to high-dimensional medical images,40 contain complex sample-specific patterns often imperceptible to human observers. The convolutional layers used in VAE can capture heterogeneous signatures to generate unique GEPs. This finding aligns with the observation that conventional methods face challenges in accurately decomposing cancer samples,41 where heterogeneity is critical for diagnosis and prognosis.42 Noteworthy, the absence of ground truth when decomposing real-world bulk data poses a significant challenge for model validation. This is particularly problematic for deep-learning models43,44,45 that output only cell proportions, raising doubts about the reliability of the predictions. The gCCA alleviates the issue by introducing interpretability at both gene and person levels, unveiling how the algorithm arrives at its conclusions.

We identify a few directions for future refinement and research. First, manual batch correction between bulk and reference data is currently required before the data are fed into the gCCA, as scRNA-seq reference profiles may not accurately reflect bulk samples. In scenarios where batch effects are minimal, gCCA can be applied directly. However, in cases of substantial batch effects, we introduce a dedicated batch-correction mode that relies on user-provided cellular compositions from a subset of bulk samples. While this approach remains more cost-effective than profiling the entire dataset, users must resort to the non-batch mode or external batch-correction methods if such a subset is unavailable. Recent advances in batch-effect removal46,47,48 may offer promising avenues for exploration. Second, while our current workflow implements the GMM and VAE in two separate steps, an integrated approach combining these techniques (e.g., GMVAE49) presents an intriguing alternative for further investigation. Meanwhile, a disentangling VAE latent space,50 where each latent dimension specifically corresponds to a specific data factor, such as a unique genoMap pattern in this study, could further optimize the cell genoMap generation. Lastly, we note that the proposed method is generalizable to other types of biology data requiring deconvolution, such as spot-based spatial transcriptomics.51,52 However, our current approach treats each sample independently, without incorporating inter-sample spatial dependencies among adjacent spots. The term “spatial” here refers to the physical spatial positioning within tissues as captured by spatial transcriptomics. As demonstrated by methods such as conditional autoregressive-based deconvolution (CARD),53 modeling spatial correlation can significantly improve the accuracy and biological plausibility of deconvolution results. Ignoring these spatial relationships can lead to noisy or fragmented decomposition results, whereas spatial smoothing produces more coherent and anatomically consistent patterns. Since spots in spatial transcriptomics datasets are not independent but often exhibit similar cellular compositions due to tissue continuity, integrating these spatial interactions into genomic data analysis is an intricate problem that we plan to address in future research.

In summary, we present and validate the efficacy of the gCCA framework by using both simulations and real-world bulk RNA-seq datasets, highlighting its versatility across diverse tissue types. This approach offers a cost-effective and accurate solution for deriving cellular compositions and sample-specific GEPs simultaneously without the need for exhaustive scRNA-seq profiling, and its interpretability may facilitate clinical translation and expedite the incorporation of gCCA into clinical care.

Data and code availability

All datasets utilized in this study are publicly available; detailed descriptions and access information are provided in the methods and supplemental methods: datasets summary. The published article also includes all code generated or analyzed during this study. The Python implementation codes and tutorials are available via https://github.com/JunyanReplicant/gCCA_decomposition.

Acknowledgments

The authors would like to acknowledge grant support from Stanford Institute for Human-Centered Intelligence (HAI) and the National Institutes of Health (NIH) (1R01CA223667, 1R01CA176553, 1R01275772, and 1K99LM014309).

Author contributions

M.T.I. and L.X. conceptualized and supervised the study. J.L. developed the computational framework and wrote the manuscript. Z.Z. and Y.C. conducted comparative studies and contributed feedback to improve the manuscript.

Declaration of interests

A patent application based on this work has been submitted (application number: 63/479724) by the Board of Trustees of the Leland Stanford Junior University. The inventors are J.L., M.T.I., and L.X. The patent application covers all the contents of the manuscript.

Published: September 17, 2025

Footnotes

Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2025.08.021.

Contributor Information

Md Tauhidul Islam, Email: tauhid@stanford.edu.

Lei Xing, Email: lei@stanford.edu.

Supplemental information

Document S1. Supplemental methods
mmc1.pdf (241.2KB, pdf)
Table S1. Model prediction results

Complete model prediction results and corresponding ground-truth values for all bulk datasets to facilitate future benchmarking and comparative analyses.

mmc2.xlsx (46.8KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (12.3MB, pdf)

References

  • 1.Jovic D., Liang X., Zeng H., Lin L., Xu F., Luo Y. Single-cell RNA sequencing technologies and applications: A brief overview. Clin. Transl. Med. 2022;12 doi: 10.1002/ctm2.694. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Van de Sande B., Lee J.S., Mutasa-Gottgens E., Naughton B., Bacon W., Manning J., Wang Y., Pollard J., Mendez M., Hill J., et al. Applications of single-cell RNA sequencing in drug discovery and development. Nat. Rev. Drug Discov. 2023;22:496–520. doi: 10.1038/s41573-023-00688-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Janjic A., Wange L.E., Bagnoli J.W., Geuder J., Nguyen P., Richter D., Vieth B., Vick B., Jeremias I., Ziegenhain C., et al. Prime-seq, efficient and powerful bulk RNA sequencing. Genome Biol. 2022;23:88. doi: 10.1186/s13059-022-02660-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Nguyen H., Nguyen H., Tran D., Draghici S., Nguyen T. Fourteen years of cellular deconvolution: methodology, applications, technical evaluation and outstanding challenges. Nucleic Acids Res. 2024;52:4761–4783. doi: 10.1093/nar/gkae267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Newman A.M., Liu C.L., Green M.R., Gentles A.J., Feng W., Xu Y., Hoang C.D., Diehn M., Alizadeh A.A. Robust enumeration of cell subsets from tissue expression profiles. Nat. Methods. 2015;12:453–457. doi: 10.1038/nmeth.3337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Awad M., Khanna R. In: Efficient Learning Machines: Theories, Concepts, and Applications for Engineers and System Designers. Awad M., Khanna R., editors. Apress; 2015. Support Vector Regression; pp. 67–80. [DOI] [Google Scholar]
  • 7.Newman A.M., Steen C.B., Liu C.L., Gentles A.J., Chaudhuri A.A., Scherer F., Khodadoust M.S., Esfahani M.S., Luca B.A., Steiner D., et al. Determining cell type abundance and expression from bulk tissues with digital cytometry. Nat. Biotechnol. 2019;37:773–782. doi: 10.1038/s41587-019-0114-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Jew B., Alvarez M., Rahmani E., Miao Z., Ko A., Garske K.M., Sul J.H., Pietiläinen K.H., Pajukanta P., Halperin E. Accurate estimation of cell composition in bulk expression through robust integration of single-cell information. Nat. Commun. 2020;11:1971. doi: 10.1038/s41467-020-15816-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Islam M.T., Xing L. Cartography of Genomic Interactions Enables Deep Analysis of Single-Cell Expression Data. Nat. Commun. 2023;14:679. doi: 10.1038/s41467-023-36383-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Kingma D.P., Welling M. Auto-Encoding Variational Bayes. arXiv. 2022 doi: 10.48550/arXiv.1312.6114. Preprint at. [DOI] [Google Scholar]
  • 11.Reynolds D. In: Encyclopedia of Biometrics. Li S.Z., Jain A., editors. Springer US; 2009. Gaussian Mixture Models; pp. 659–663. [DOI] [Google Scholar]
  • 12.Wu S.Z., Al-Eryani G., Roden D.L., Junankar S., Harvey K., Andersson A., Thennavan A., Wang C., Torpy J.R., Bartonicek N., et al. A single-cell and spatially resolved atlas of human breast cancers. Nat. Genet. 2021;53:1334–1347. doi: 10.1038/s41588-021-00911-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Monaco G., Lee B., Xu W., Mustafah S., Hwang Y.Y., Carré C., Burdin N., Visan L., Ceccarelli M., Poidinger M., et al. RNA-Seq Signatures Normalized by mRNA Abundance Allow Absolute Deconvolution of Human Immune Cell Types. Cell Rep. 2019;26:1627–1640.e7. doi: 10.1016/j.celrep.2019.01.041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Patrick E., Taga M., Ergun A., Ng B., Casazza W., Cimpean M., Yung C., Schneider J.A., Bennett D.A., Gaiteri C., et al. Deconvolving the contributions of cell-type heterogeneity on cortical gene expression. PLoS Comput. Biol. 2020;16 doi: 10.1371/journal.pcbi.1008120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Wolf F.A., Angerer P., Theis F.J. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol. 2018;19:15. doi: 10.1186/s13059-017-1382-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Fryer D., Strümke I., Nguyen H. Shapley Values for Feature Selection: The Good, the Bad, and the Axioms. IEEE Access. 2021;9:144352–144360. doi: 10.1109/ACCESS.2021.3119110. [DOI] [Google Scholar]
  • 17.Kokhlikyan N., Miglani V., Martin M., Wang E., Alsallakh B., Reynolds J., Melnikov A., Kliushkina N., Araya C., Yan S., et al. Captum: A unified and generic model interpretability library for PyTorch. arXiv. 2020 doi: 10.48550/arXiv.2009.07896. Preprint at. [DOI] [Google Scholar]
  • 18.Wang X., Park J., Susztak K., Zhang N.R., Li M. Bulk tissue cell type deconvolution with multi-subject single-cell expression reference. Nat. Commun. 2019;10:380. doi: 10.1038/s41467-018-08023-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.McInnes L., Healy J., Saul N., Großberger L. UMAP: Uniform Manifold Approximation and Projection. J. Open Source Softw. 2018;3:861. doi: 10.21105/joss.00861. [DOI] [Google Scholar]
  • 20.Grønbech C.H., Vording M.F., Timshel P.N., Sønderby C.K., Pers T.H., Winther O. scVAE: variational auto-encoders for single-cell gene expression data. Bioinformatics. 2020;36:4415–4422. doi: 10.1093/bioinformatics/btaa293. [DOI] [PubMed] [Google Scholar]
  • 21.Yan J., Ma M., Yu Z. bmVAE: a variational autoencoder method for clustering single-cell mutation data. Bioinformatics. 2023;39 doi: 10.1093/bioinformatics/btac790. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Lan T., Hutvagner G., Lan Q., Liu T., Li J. Sequencing dropout-and-batch effect normalization for single-cell mRNA profiles: a survey and comparative analysis. Brief. Bioinform. 2021;22 doi: 10.1093/bib/bbaa248. [DOI] [PubMed] [Google Scholar]
  • 23.Hicks S.C., Townes F.W., Teng M., Irizarry R.A. Missing data and technical variability in single-cell RNA-sequencing experiments. Biostatistics. 2018;19:562–578. doi: 10.1093/biostatistics/kxx053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Cobos F.A., Panah M.J.N., Epps J., Long X., Man T.-K., Chiu H.-S., Chomsky E., Kiner E., Krueger M.J., di Bernardo D., et al. Effective methods for bulk RNA-seq deconvolution using scnRNA-seq transcriptomes. Genome Biol. 2023;24:177. doi: 10.1186/s13059-023-03016-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Sun X., Gao C., Zhao K., Yang Y., Rassadkina Y., Fajnzylber J., Regan J., Li J.Z., Lichterfeld M., Yu X.G. Immune-profiling of SARS-CoV-2 viremic patients reveals dysregulated innate immune responses. Front. Immunol. 2022;13 doi: 10.3389/fimmu.2022.984553. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Park S., Kwon Y. The role of ELF3 in metastasis of triple negative breast cancer. Ann. Oncol. 2017;28 doi: 10.1093/annonc/mdx654.025. [DOI] [Google Scholar]
  • 27.Kufe D.W. MUC1-C oncoprotein as a target in breast cancer: activation of signaling pathways and therapeutic approaches. Oncogene. 2013;32:1073–1081. doi: 10.1038/onc.2012.158. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Xu L., Saunders K., Huang S.-P., Knutsdottir H., Martinez-Algarin K., Terrazas I., Chen K., McArthur H.M., Maués J., Hodgdon C., et al. A comprehensive single-cell breast tumor atlas defines epithelial and immune heterogeneity and interactions predicting anti-PD-1 therapy response. Cell Rep. Med. 2024;5 doi: 10.1016/j.xcrm.2024.101511. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Tian W., Zhou J., Chen M., Qiu L., Li Y., Zhang W., Guo R., Lei N., Chang L. Bioinformatics analysis of the role of aldolase A in tumor prognosis and immunity. Sci. Rep. 2022;12 doi: 10.1038/s41598-022-15866-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Zhang J., Ye Z.W., Morgenstern R., Townsend D.M., Tew K.D. Microsomal glutathione transferase 1 in cancer and the regulation of ferroptosis. Adv. Cancer Res. 2023;160:107–132. doi: 10.1016/bs.acr.2023.05.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Lin Q., Chen X., Meng F., Ogawa K., Li M., Song R., Zhang S., Zhang Z., Kong X., Xu Q., et al. ASPH-notch Axis guided Exosomal delivery of Prometastatic Secretome renders breast Cancer multi-organ metastasis. Mol. Cancer. 2019;18:156. doi: 10.1186/s12943-019-1077-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Schoppmeyer R., Zhao R., Cheng H., Hamed M., Liu C., Zhou X., Schwarz E.C., Zhou Y., Knörck A., Schwär G., et al. Human profilin 1 is a negative regulator of CTL mediated cell-killing and migration. Eur. J. Immunol. 2017;47:1562–1572. doi: 10.1002/eji.201747124. [DOI] [PubMed] [Google Scholar]
  • 33.Yu K., Kuang L., Fu T., Zhang C., Zhou Y., Zhu C., Zhang Q., Zhang Z., Le A. CREM Is Correlated With Immune-Suppressive Microenvironment and Predicts Poor Prognosis in Gastric Adenocarcinoma. Front. Cell Dev. Biol. 2021;9 doi: 10.3389/fcell.2021.697748. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Belarif L., Mary C., Jacquemont L., Mai H.L., Danger R., Hervouet J., Minault D., Thepenier V., Nerrière-Daguin V., Nguyen E., et al. IL-7 receptor blockade blunts antigen-specific memory T cell responses and chronic inflammation in primates. Nat. Commun. 2018;9:4483. doi: 10.1038/s41467-018-06804-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Li B., Lu Y., Zhong M.-C., Qian J., Li R., Davidson D., Tang Z., Zhu K., Argenty J., de Peredo A.G., et al. Cis interactions between CD2 and its ligands on T cells are required for T cell activation. Sci. Immunol. 2022;7 doi: 10.1126/sciimmunol.abn6373. [DOI] [PubMed] [Google Scholar]
  • 36.Feng S., Isayev O., Werner J., Bazhin A.V. CD96 as a Potential Immune Regulator in Cancers. Int. J. Mol. Sci. 2023;24:1303. doi: 10.3390/ijms24021303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Hwang S.S., Lim J., Yu Z., Kong P., Sefik E., Xu H., Harman C.C.D., Kim L.K., Lee G.R., Li H.-B., Flavell R.A. mRNA destabilization by BTG1 and BTG2 maintains T cell quiescence. Science. 2020;367:1255–1260. doi: 10.1126/science.aax0194. [DOI] [PubMed] [Google Scholar]
  • 38.Kaminski S., Hermann-Kleiter N., Meisel M., Thuille N., Cronin S., Hara H., Fresser F., Penninger J.M., Baier G. Coronin 1A is an essential regulator of the TGFβ receptor/SMAD3 signaling pathway in Th17 CD4(+) T cells. J. Autoimmun. 2011;37:198–208. doi: 10.1016/j.jaut.2011.05.018. [DOI] [PubMed] [Google Scholar]
  • 39.Orrantia-Borunda E., Anchondo-Nuñez P., Acuña-Aguilar L.E., Gómez-Valles F.O., Ramírez-Valdespino C.A. In: Breast Cancer. Mayrovitz H.N., editor. Exon Publications; 2022. Subtypes of Breast Cancer. [PubMed] [Google Scholar]
  • 40.Xing L., Giger M.L., Min J.K. Academic Press; 2020. Artificial Intelligence in Medicine: Technical Basis and Clinical Applications. [Google Scholar]
  • 41.Alonso-Moreda N., Berral-González A., De La Rosa E., González-Velasco O., Sánchez-Santos J.M., De Las Rivas J. Comparative Analysis of Cell Mixtures Deconvolution and Gene Signatures Generated for Blood, Immune and Cancer Cells. Int. J. Mol. Sci. 2023;24 doi: 10.3390/ijms241310765. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Dagogo-Jack I., Shaw A.T. Tumour heterogeneity and resistance to cancer therapies. Nat. Rev. Clin. Oncol. 2018;15:81–94. doi: 10.1038/nrclinonc.2017.166. [DOI] [PubMed] [Google Scholar]
  • 43.Dong M., Thennavan A., Urrutia E., Li Y., Perou C.M., Zou F., Jiang Y. SCDC: bulk gene expression deconvolution by multiple single-cell RNA sequencing references. Brief. Bioinform. 2021;22:416–427. doi: 10.1093/bib/bbz166. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Menden K., Marouf M., Oller S., Dalmia A., Magruder D.S., Kloiber K., Heutink P., Bonn S. Deep learning–based cell composition analysis from tissue expression profiles. Sci. Adv. 2020;6 doi: 10.1126/sciadv.aba2619. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Huang J., Du Y., Stucky A., Kelly K.R., Zhong J.F., Sun F. DeepDecon accurately estimates cancer cell fractions in bulk RNA-seq data. Patterns. 2024;5 doi: 10.1016/j.patter.2024.100969. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Li X., Wang K., Lyu Y., Pan H., Zhang J., Stambolian D., Susztak K., Reilly M.P., Hu G., Li M. Deep learning enables accurate clustering with batch effect removal in single-cell RNA-seq analysis. Nat. Commun. 2020;11:2338. doi: 10.1038/s41467-020-15851-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Yu X., Xu X., Zhang J., Li X. Batch alignment of single-cell transcriptomics data using deep metric learning. Nat. Commun. 2023;14:960. doi: 10.1038/s41467-023-36635-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Rahman M.A., Tutul A.A., Sharmin M., Bayzid M.S. BEENE: deep learning-based nonlinear embedding improves batch effect estimation. Bioinformatics. 2023;39 doi: 10.1093/bioinformatics/btad479. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Dilokthanakul N., Mediano P.A.M., Garnelo M., Lee M.C.H., Salimbeni H., Arulkumaran K., Shanahan M. Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders. arXiv. 2017 doi: 10.48550/arXiv.1611.02648. Preprint at. [DOI] [Google Scholar]
  • 50.Burgess C.P., Higgins I., Pal A., Matthey L., Watters N., Desjardins G., Lerchner A. Understanding disentangling in $β$-VAE. arXiv. 2018 doi: 10.48550/arXiv.1804.03599. Preprint at. [DOI] [Google Scholar]
  • 51.Liao J., Qian J., Fang Y., Chen Z., Zhuang X., Zhang N., Shao X., Hu Y., Yang P., Cheng J., et al. De novo analysis of bulk RNA-seq data at spatially resolved single-cell resolution. Nat. Commun. 2022;13:6498. doi: 10.1038/s41467-022-34271-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Zhou Z., Zhong Y., Zhang Z., Ren X. Spatial transcriptomics deconvolution at single-cell resolution using Redeconve. Nat. Commun. 2023;14:7930. doi: 10.1038/s41467-023-43600-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Ma Y., Zhou X. Spatially informed cell-type deconvolution for spatial transcriptomics. Nat. Biotechnol. 2022;40:1349–1359. doi: 10.1038/s41587-022-01273-7. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Supplemental methods
mmc1.pdf (241.2KB, pdf)
Table S1. Model prediction results

Complete model prediction results and corresponding ground-truth values for all bulk datasets to facilitate future benchmarking and comparative analyses.

mmc2.xlsx (46.8KB, xlsx)
Document S2. Article plus supplemental information
mmc3.pdf (12.3MB, pdf)

Data Availability Statement

All datasets utilized in this study are publicly available; detailed descriptions and access information are provided in the methods and supplemental methods: datasets summary. The published article also includes all code generated or analyzed during this study. The Python implementation codes and tutorials are available via https://github.com/JunyanReplicant/gCCA_decomposition.


Articles from American Journal of Human Genetics are provided here courtesy of American Society of Human Genetics

RESOURCES