Abstract
Kidney injury and chronic kidney disease progression are accompanied by spatially heterogeneous activation of programmed cell death (PCD), yet existing approaches have limited ability to jointly infer cell-type composition, death-program activity, and their spatial organization from spatial transcriptomics (ST) data. We present CoDeST (Confidence-weighted Dual-teacher Spatial Training), a multi-task spatial inference framework that predicts, for each ST spot, a 34-class kidney cell-type composition vector together with continuous activity scores for four PCD programs: apoptosis, pyroptosis, necroptosis, and ferroptosis. CoDeST uses a three-stage pseudo-to-real training strategy that combines self-supervised pretraining on real ST slices, supervised deconvolution on donor-matched pseudo-spots, and real-ST domain adaptation with teacher–student distillation, marker-based weak constraints, confidence-weighted AUCell/ssGSEA PCD supervision, and boundary-preserving spatial regularization. In donor-held-out pseudo-spot benchmarks, CoDeST shows competitive recovery of cell-type proportions compared with representative deconvolution methods. On real kidney ST data, marker-consistency analysis and a Visium HD-derived benchmark further support its ability to transfer deconvolution signals from pseudo-spots to real spatial tissue settings. Ablation and sensitivity analyses indicate that the three-stage design, confidence weighting, and spatial graph modeling contribute to stable deconvolution and PCD mapping while balancing spatial coherence with boundary contrast. CoDeST provides a kidney-focused framework for joint spatial mapping of cell composition and PCD-related transcriptional programs.
Keywords: spatial transcriptomics, kidney, deconvolution, programmed cell death, weak supervision, knowledge distillation
Introduction
Kidney injury and chronic kidney disease progression are often accompanied by activation of multiple programmed cell death (PCD) pathways [1]. Loss of parenchymal cells directly compromises nephron function, promotes immune-cell recruitment and inflammatory cascades, and drives tissue remodeling that culminates in fibrosis and irreversible functional decline. PCD is not a single endpoint [2]. Apoptosis, pyroptosis, necroptosis, and ferroptosis differ in molecular triggers and immunological consequences, and their activity is frequently dependent on cell identity and tissue architecture. Within the same kidney specimen, distinct nephron compartments and lesion-associated microenvironments—such as ischemic and inflammatory borders—may be dominated by different cell populations with preferential engagement of specific death programs. Bulk transcriptomic profiling averages these spatially heterogeneous signals [3]. Single-cell sequencing resolves cell-type differences but removes spatial neighborhood context and microenvironmental cues. A central question, therefore, remains unresolved: which cells undergo which death programs, and where, during injury and repair in the kidney. Spatially resolved maps of PCD are needed to connect molecular state, cellular composition, and tissue structure [4–6].
Spatial transcriptomics (ST) provides a direct avenue to address these questions by measuring gene expression together with spatial coordinates at tissue-section resolution, enabling the spatial mapping of transcriptional signals associated with PCD on the same slice. Most current platforms, however, measure at the level of spots rather than single cells. Each spot typically contains multiple cell types, and the observed expression represents a mixture, requiring computational deconvolution to recover cell-type composition and support cell-type-resolved interpretation [7–9]. Supervision for death programs is even more challenging. PCD rarely has directly observable “hard labels,” and spot-matched ground-truth intensities are generally unavailable in tissue samples. Existing studies, therefore, rely on gene-set scoring approaches, such as ssGSEA or AUCell, to convert predefined gene sets into per-spot scores as proxies for death-program activity. These scores are sensitive to gene-set selection, normalization choices, and measurement noise, and they often show substantial method-to-method disagreement and label uncertainty. ST thus supplies the essential data substrate for spatial analyses of PCD, while a robust computational framework remains needed to turn mixed cellular composition and noisy proxy scores into reliable, generalizable spatial inferences [10–12].
Current methods for deconvolving ST data can be broadly grouped into three categories with well-established representative frameworks. The first category comprises probabilistic graphical models and Bayesian inference approaches that explicitly model cell-type-specific expression distributions and measurement noise to estimate the cellular composition of each spot; examples include cell2location [13], Stereoscope [14], and DestVI [15], which decompose mixed spot expression under probabilistic or variational inference formulations. The second category includes regression- and matrix-decomposition-based methods that treat scRNA-seq reference profiles as a dictionary or prior and solve for composition using non-negative regression, regularized regression, or statistical factorization. Representative tools include RCTD [16], SPOTlight [17], and CARD [8], which provide interpretable estimates from the perspectives of robust regression, non-negative matrix factorization, and reference-guided deconvolution. The third category encompasses deep learning, generative, or mapping methods that learn representations linking scRNA-seq and ST data or align embeddings to infer cell-type contributions; examples include Tangram [18] and DECODE [19], which learn mappings to project single cells onto spatial locations, while related deep frameworks often incorporate autoencoders or variational objectives to improve cross-domain adaptation. These approaches constitute the mainstream toolkit for compositional inference, yet extending them to joint inference of cell composition and death-program states, with stable generalization across tissue sections and donors, remains limited by several key gaps.
Supervision for death programs is typically available only through gene-set scoring proxies, with substantial noise and disagreement across methods. AUCell and ssGSEA depend on gene-set definitions, normalization choices, and handling of expression sparsity, producing scores that may differ in magnitude and, in some cases, spatial pattern for the same program within the same section [20–22]. Treating any single score as ground truth risks training on noise and amplifying bias. Simple averaging or selecting one scorer implicitly assumes equal reliability and lacks a sample-level confidence characterization, reducing robustness in regions dominated by uncertain labels [23–25].
Spatial modeling commonly relies on neighborhood smoothing or graph regularization to increase coherence, while naive smoothing can erase genuine tissue boundaries. Many approaches impose consistency constraints over an adjacency graph to reduce noise, yielding predictions that appear more spatially contiguous [26–28]. Lesion margins, nephron compartment interfaces, and inflammatory fronts are often characterized by sharp spatial gradients. Strong L2-type smoothing systematically compresses these differences, turning boundaries into gradual transitions and reducing sensitivity to structural and microenvironmental heterogeneity [29]. Spatial aggregation and regularization that denoise while preserving boundaries are therefore required, rather than equating spatial coherence with global smoothing [30–33].
A unified cross-domain generalization mechanism remains lacking between strongly supervised deconvolution and weakly supervised inference of death programs. Cell-type composition can be trained with strong labels on pseudo-spots derived from reference scRNA-seq data. Pseudo data differ from real ST in distribution, technical bias, and structural constraints, leading to domain shift and degraded performance under direct transfer [34–36]. Death-program supervision is weaker and less stable [37, 38]. Joint training without effective knowledge transfer and in-domain correction can be dominated by the strong task or be pulled by noisy weak supervision. A training strategy is needed to distill compositional knowledge learned from pseudo-spots into real ST and to correct predictions under spatial structure constraints, enabling complementary rather than interfering roles for strong and weak supervision within a single model [39, 40].
In summary, existing approaches struggle to satisfy three requirements simultaneously: reliable supervision under noisy weak labels, boundary-preserving spatial modeling, and cross-donor, cross-section generalization through joint strong–weak task training. These gaps limit the ability to derive cell-type-resolved, spatially coherent, and interpretable maps of PCD in human kidney tissue.
We introduce CoDeST (Confidence-weighted Dual-teacher Spatial Training), a multi-task spatial inference framework for kidney ST that outputs, for each spot, cell-type composition proportions (34 classes) and continuous intensities for four PCD programs (apoptosis, pyroptosis, necroptosis, and ferroptosis).
CoDeST addresses key limitations of existing methods through three design components:
First, CoDeST uses AUCell and ssGSEA as dual weak supervisory signals for PCD prediction, and estimates spot–program-level confidence from their agreement. This confidence is then used to weight the weak-supervision loss, reducing the model’s sensitivity to noisy weak labels.
Second, CoDeST incorporates spatial graph-based information aggregation together with a boundary-preserving constraint, enabling spatially coherent predictions while limiting over-smoothing across tissue boundaries and lesion fronts.
Third, CoDeST adopts a distillation-based pseudo-to-real domain adaptation strategy, in which a deconvolution model trained with strong supervision on pseudo-spots serves as the teacher for a student model trained on real ST data. By further incorporating real-ST spatial structure and marker-based constraints, this stage jointly refines cell-type deconvolution and PCD mapping.
Materials
Kidney Precision Medicine Project single-cell and spatial transcriptomics datasets
Single-cell RNA sequencing (scRNA-seq) and ST data were obtained from the Kidney Precision Medicine Project (KPMP) [41]. The scRNA-seq reference comprised
200 000 cells annotated into 34 kidney cell types, and it was used to define the cell-type label space for downstream analyses. ST data consisted of kidney tissue sections profiled at spot-level resolution and served as the target domain for spatial inference.
To summarize the overall composition and major lineage structure of the scRNA-seq reference, we visualized the cells using UMAP based on the broad cell-class annotation provided with the dataset. Because the 34 fine-grained cell types include several transcriptionally related kidney subtypes that may partially overlap in a 2D UMAP projection, we used five major cell classes in the main figure to improve readability and highlight the global structure of the reference dataset (Fig. 1). The full set of 34 fine-grained cell-type annotations was retained for model training and for defining the deconvolution label space.
Figure 1.

UMAP visualization of the KPMP scRNA-seq reference colored by five broad cell classes.
Programmed cell death gene sets
Gene sets for PCD programs were curated from public resources. Gene sets associated with apoptosis, pyroptosis, and necroptosis were obtained from MSigDB [42], whereas the ferroptosis gene set was curated from FerrDb [43]. Prior to scoring, each gene set was intersected with the gene symbols present in the ST data to ensure platform compatibility. Across programs, the mean overlap between curated gene sets and ST genes was
90%, supporting robust estimation of PCD proxy scores in the spatial domain.
Methods
Overview of CoDeST
CoDeST is a three-stage pseudo-to-real framework designed for spatial cell-type deconvolution and PCD mapping in ST data. The model integrates self-supervised representation learning on real ST data, supervised deconvolution on donor-matched pseudo-spots, and domain adaptation on real ST sections to transfer cell-composition knowledge from simulated mixtures to real tissue contexts (Fig. 2).
Figure 2.
Three-stage pseudo-to-real framework of CoDeST for spatial deconvolution and PCD mapping.
In the first stage, CoDeST performs self-supervised pretraining on real ST slices. Spot-level gene expression and the corresponding spatial graph are used as inputs. Random spot–gene masking and edge dropping are applied as data augmentations, and the encoder and graph module are trained to reconstruct masked gene expression. This stage encourages the model to learn expression representations that are compatible with the spatial structure and noise characteristics of real ST data.
In the second stage, CoDeST is trained on donor-matched pseudo-spots generated from the scRNA-seq reference. Each pseudo-spot is associated with known cell-type proportions, providing strong supervision for deconvolution. This stage trains the encoder and deconvolution decoder to recover the 34-cell-type composition from mixed spot-level expression and produces a pseudo-domain deconvolution teacher for subsequent adaptation.
In the third stage, CoDeST is adapted to real ST data through multi-task training. The Stage 2 deconvolution model provides teacher predictions for confidence-weighted distillation, while marker-derived activity supplies a weak constraint for cell-type abundance prediction. In parallel, AUCell and ssGSEA are used as dual weak supervisory signals for PCD programs, with their agreement used to estimate spot–program-level confidence weights. The final model jointly outputs cell-type proportions and PCD program scores for each ST spot, while spatial graph aggregation and boundary-preserving constraints promote spatially coherent predictions without excessive smoothing across tissue boundaries.
Construction of donor-matched pseudo-spots
To provide strong supervision for cell-type deconvolution, we constructed donor-matched pseudo-spots from the scRNA-seq reference. Each pseudo-spot was generated by sampling and aggregating single cells with known cell-type labels, yielding a spot-level expression profile together with a reference cell-type proportion vector. The pseudo-spot generation procedure was designed to approximate key properties of real ST data while avoiding biologically implausible mixtures. Specifically, pseudo-spots were generated with donor-aware sampling, realistic cell-number and library-size distributions, marker-based plausibility filtering, ST-like noise injection, and calibrated sampling of rare cell types.
Donor-aware cell sampling
Pseudo-spots were generated within donor boundaries. For each pseudo-spot, a donor was first sampled from the scRNA-seq reference, and all component cells were then selected only from the same KPMPID. This donor-aware design avoids mixing cells from unrelated biological backgrounds and reduces the risk of creating artificial combinations driven by donor-specific disease state, technical variation, or batch effects. The donor-level train, validation, and test split was preserved during pseudo-spot generation, such that pseudo-spots in different data splits were derived from non-overlapping donor sets.
Cell-type composition sampling
For each pseudo-spot, a cell-type composition vector was sampled over the 34 reference cell types. The sampling procedure was designed to cover both common tissue compositions and rare-cell-containing mixtures. Rather than drawing all cell types uniformly, we used scenario-based composition sampling to generate a range of mixture patterns, including near-pure spots, common multi-cell-type mixtures, and rare-cell-enriched mixtures. The number of cells per pseudo-spot was sampled from an empirical distribution estimated from real ST data, so that pseudo-spots reflected a realistic range of cell capture rather than a fixed or overly broad cell count.
Rare cell types were deliberately oversampled in a subset of pseudo-spots to ensure sufficient training signal for low-abundance populations. However, to avoid inflating their prevalence during model training, these samples were assigned calibrated sample weights. In practice, the training loss was adjusted using class-balanced or prior-corrected weights, so that rare-cell oversampling improved representation learning without encouraging systematic false-positive predictions on real ST data.
Pseudo-spot expression generation
Given a sampled donor, cell-type composition, and target number of cells, single cells were sampled from the corresponding donor and cell-type pools. The raw count vectors of selected cells were then aggregated to form a pseudo-spot count profile. The reference cell-type proportion of the pseudo-spot was computed as the number of sampled cells from each cell type divided by the total number of sampled cells. This yielded paired training examples consisting of a pseudo-spot expression vector and a 34D cell-type proportion label.
All pseudo-spots were generated in the same gene feature space used by the model. Genes absent from the shared feature set were excluded, and the gene order was fixed to match the reference feature list. For each pseudo-spot, both count-scale and normalized expression representations were retained when applicable: count-scale profiles were used for library-size matching and noise modeling, whereas normalized profiles were used as model inputs.
Library-size matching and noise injection
To better approximate the distribution of real ST measurements, pseudo-spot counts were adjusted to match empirical library-size statistics from real ST sections. For each pseudo-spot, a target library size was sampled from the observed ST library-size distribution, and the aggregated pseudo count vector was downsampled or rescaled accordingly. This step prevents pseudo-spots from having unrealistically high or low total counts relative to real ST spots.
After library-size matching, ST-like technical noise was injected into the pseudo expression profiles. We modeled dropout in a gene-dependent manner: low-expression genes were assigned higher dropout probabilities, whereas highly expressed genes and marker genes were assigned lower dropout probabilities to preserve informative cell-type signals. Additional count noise was introduced using negative binomial sampling or related overdispersed count perturbations, allowing pseudo-spots to better reflect the sparsity and variability of real ST data. These perturbations were applied mildly to avoid destroying the underlying cell-type composition signal.
Marker plausibility filtering and sample weighting
To reduce biologically implausible pseudo-spots, we applied marker-based plausibility checks using the predefined cell-type marker sets. For each generated pseudo-spot, marker activity was computed from the pseudo expression profile and compared with the sampled cell-type composition. Pseudo-spots were filtered or down-weighted if they showed strong marker patterns inconsistent with their assigned cell-type proportions, or if they exhibited clearly conflicting marker combinations suggestive of implausible mixtures.
Marker plausibility was used as a quality-control signal rather than as an additional hard label. Pseudo-spots passing all QC criteria were assigned higher training weights, whereas borderline pseudo-spots were either removed or assigned lower weights depending on the severity of the inconsistency. These weights were combined with the rare-cell calibration weights described above, so that the final sample weight reflected both biological plausibility and composition-prior correction.
Pseudo-spot split
Pseudo-spots were generated separately for the training, validation, and test splits according to the donor-level scRNA-seq split. Cells from donors assigned to the training split were used only to generate pseudo-training spots, while validation and test pseudo-spots were generated from held-out donor sets. This design prevents donor leakage and enables evaluation of deconvolution performance under donor-held-out conditions.
The final pseudo-spot datasets were saved as separate AnnData objects for training, validation, and testing. Each file contained the pseudo-spot expression matrix, count-scale and normalized expression layers when available, the 34D reference cell-type proportion matrix, donor identifiers, cell-count metadata, library-size metadata, QC metrics, and sample weights. These pseudo-spots were used as the supervised training data for Stage 2 deconvolution and as benchmark data for evaluating cell-type proportion recovery.
Unified model architecture
CoDeST uses a shared architecture across the three training stages. The model input is a spot-level expression vector restricted to the shared gene feature space. An expression encoder maps the input vector to a latent representation
for spot
. When the spatial graph module is enabled, graph-based message passing aggregates information from neighboring spots to produce a spatially informed representation:
![]() |
(1) |
where
denotes the slice-specific adjacency matrix, and
represents the graph module. The residual formulation preserves a direct expression pathway and stabilizes training when switching between graph-free and graph-enabled settings.
Two task-specific decoders operate on the shared representation. The deconvolution decoder maps the representation to a 34D probability simplex through a softmax layer, yielding predicted cell-type proportions. The PCD decoder maps the same representation to four continuous outputs corresponding to apoptosis, pyroptosis, necroptosis, and ferroptosis. In Stage 2, the graph module is bypassed by setting
, producing a graph-free deconvolution teacher. In Stages 1 and 3, the graph module is enabled for real ST slices to incorporate spatial context.
Three-stage training objectives
Stage 1: real-spatial transcriptomics self-supervised pretraining
Stage 1 trains the expression encoder and graph module on real ST slices using masked gene reconstruction. For each training batch, a subset of spot–gene entries is randomly masked in the expression matrix. In addition, a fraction of spatial graph edges is randomly dropped to improve robustness to neighborhood variability. The corrupted expression matrix and perturbed graph are passed through the encoder and graph module, and the model is optimized to reconstruct the original expression values at masked entries. This objective encourages the model to learn real-ST-compatible expression and spatial representations before supervised deconvolution training.
Let
be the expression vector of spot
, and let
denotes the set of masked gene entries. The Stage 1 reconstruction loss is computed only over masked entries:
![]() |
(2) |
where
is the reconstructed expression value, and
is a robust penalty such as the Huber loss.
Stage 2: pseudo-spot supervised deconvolution
Stage 2 trains the expression encoder and deconvolution decoder on donor-matched pseudo-spots. Each pseudo-spot has a known cell-type proportion vector
, where
denotes the 34D probability simplex. The model predicts
from pseudo-spot expression, with the graph module bypassed. The supervised deconvolution loss is defined as the Kullback–Leibler divergence from the reference composition to the predicted composition:
![]() |
(3) |
where
is a small constant for numerical stability. Sample weights from pseudo-spot QC and rare-cell calibration were incorporated during optimization when applicable. The trained Stage 2 model was used as the deconvolution teacher for real ST adaptation.
Stage 3: pseudo-to-real domain adaptation and multi-task training
Stage 3 adapts the model to real ST slices and jointly trains deconvolution and PCD prediction. The Stage 2 deconvolution model is frozen and used to generate teacher soft labels
for real ST spots. The Stage 3 student predicts
with the graph module enabled. Teacher–student distillation is enforced using a KL divergence:
![]() |
(4) |
To further constrain real-ST deconvolution, marker-derived activity was computed for each cell type using marker sets selected from the scRNA-seq reference. The marker constraint encourages the predicted abundance of each cell type to be consistent with its corresponding marker-derived activity across spots. This constraint was used as a weak target-domain signal rather than as a ground-truth cell-type proportion label.
For PCD prediction, AUCell and ssGSEA were used as dual weak supervisory signals. Let
and
denote the AUCell and ssGSEA scores for spot
and PCD program
. Scores were standardized within each program:
![]() |
(5) |
A spot–program-level confidence weight was computed from the agreement between the two standardized teachers:
![]() |
(6) |
where
controls sensitivity to teacher disagreement, and
prevents vanishing weights. The PCD target was defined as the average of the two standardized teachers,
![]() |
(7) |
Given the student prediction
, the confidence-weighted PCD loss was
![]() |
(8) |
where
denotes the Huber penalty:
![]() |
(9) |
To promote spatial coherence while preserving boundaries, we applied a robust spatial regularizer on the slice-specific neighborhood graph. Let
denotes the undirected edge set, and let
represents a model output at spot
, such as deconvolution proportions or PCD scores. The spatial regularization term was defined as
![]() |
(10) |
The robust penalty suppresses small local discrepancies while reducing the influence of large gradients across potential tissue boundaries, thereby limiting excessive smoothing.
Results
Simulated benchmarks across abundance regimes demonstrate the robustness of CoDeST
To systematically evaluate deconvolution performance, we constructed donor-matched pseudo-spot benchmarks by mixing cells sampled from 34 reference cell types. Under a unified input feature set and evaluation protocol, we compared CoDeST with representative deconvolution methods, including Cell2location [13], RCTD [16], DestVI, and Stereoscope [14] (Fig. 3a). On this standard benchmark, CoDeST showed competitive performance across MAE, Pearson correlation, and Spearman correlation, indicating its ability to recover cell-type composition from mixed spot-level expression.
Figure 3.
Deconvolution benchmarking of CoDeST on simulated pseudo-spots (a) Pseudo-spots were generated by sampling and mixing cells from 34 reference cell types within the same donor under predefined proportions, forming a standard benchmark dataset. CoDeST was compared with representative SOTA deconvolution methods using MAE, Pearson correlation, and Spearman correlation. (b) Additional benchmark datasets were constructed under multiple abundance and availability regimes using the same 34-cell-type reference. Two scenarios were considered, Ubiquitous and Regional, each evaluated under low and high cell abundance conditions to capture different degrees of mixture complexity and rare-cell prevalence.
We further evaluated model robustness under different abundance regimes by designing two composition patterns, Ubiquitous and Regional, each assessed under low- and high-abundance conditions to capture variation in rare-cell prevalence and mixture complexity (Fig. 3b). Across these settings, CoDeST maintained reasonable error and correlation profiles in certain more challenging scenarios, such as the Regional low-abundance setting, supporting its capacity to handle non-uniform cell-type compositions and rare cell types, and providing a basis for subsequent analyses on real spatial tissue sections.
Validation of deconvolution performance on real spatial transcriptomics data
Beyond the pseudo-spot benchmarks, we further evaluated the deconvolution performance of CoDeST on real ST data. Because real ST sections generally lack spot-level ground-truth cell-type proportions, we adopted two complementary validation strategies. First, we used high-resolution Visium HD data to construct an external benchmark that more closely reflects the spatial organization of real tissue, thereby evaluating deconvolution performance on a near-single-cell-resolution spatial benchmark. Second, we derived marker-based cell-type activity scores from cell-type marker expression and assessed whether predicted cell-type abundances were consistent with the corresponding marker signals across real ST spots.
To construct the Visium HD-derived benchmark, we used the cell segmentation-derived expression matrix and spatial coordinates from a public Visium HD dataset. Segmented cells were annotated by reference mapping from the scRNA-seq reference, assigning them to the same 34-cell-type label space used by CoDeST. Spatially neighboring segmented cells were then aggregated into spot-like regions, and the cell-type composition within each aggregation window was used as the segmentation-derived reference proportion. This procedure yielded a deconvolution benchmark that better preserves the spatial organization of real tissue.
On this Visium HD-derived benchmark, we compared CoDeST with SOTA deconvolution methods and evaluated the agreement between predicted cell-type proportions and segmentation-derived reference proportions using MAE, Pearson correlation, and Spearman correlation (Fig. 4a). Unlike random pseudo-spot mixing, this analysis retains local spatial neighborhood relationships, cell aggregation patterns, and platform-specific expression noise present in real tissue. CoDeST maintained competitive error and correlation profiles on this near-single-cell-resolution spatial benchmark, suggesting that its deconvolution ability can transfer to spatial data settings that more closely resemble real tissue organization.
Figure 4.

Real-ST validation of cell-type deconvolution. (a) Comparison of deconvolution methods on a Visium HD-derived benchmark. (b) Marker-consistency analysis for six representative cell types in real ST sections, comparing predicted cell-type abundance with marker-derived activity.
We next performed marker consistency analysis on real ST sections to complement the lack of spot-level cell-type proportion ground truth. Marker genes for each cell type were first selected from the scRNA-seq reference. Specifically, differential expression analysis was performed between each cell type and all remaining cells, and candidate markers were evaluated by adjusted statistical significance, effect size, expression prevalence in the target cell type, and specificity relative to non-target cell types. Candidate marker genes were required to pass predefined statistical and expression filters and to be present in the shared gene feature space used by the model. For each cell type, high-specificity marker genes passing these criteria were retained for downstream marker activity calculation in real ST sections. Marker gene lists for representative cell types are provided in Table 1.
Table 1.
Marker genes used for marker consistency analysis in representative cell types
| Cell type | Marker genes |
|---|---|
| T cell | HLA-F, PAXX, MBP, OSTF1, GTF3A |
| Schwann cell | LINC02240, NUP210L, MYO16, MARK4, MYO9B, GTF2IRD1, MTA3 |
| Podocyte | LRRC2, PARD6G, MPP5, MAP6, ILDR2, NECTIN2, MFSD6, PARVA, MIB1, NDRG2, PAK1, GRK5, NDN, LRP11, NOTCH2, MYLIP, MYO1C, MCC, GSTA4, MEGF9, MED21, PAPSS1, NCSTN, NAPEPLD, PAQR7, HEXIM1, OPTN, MTRES1, MAGT1, P4HTM |
| Parietal epithelial cell | GUCY1B1, MYOF, NSMF, MAOB, NOTCH2, NECTIN2, GTDC1, PARVA, NREP, MYEF2, MICAL3, MAP3K5, MEGF9, OPTN, ILK, PAPSS1, IFT57, NOVA1, MXRA7, MAZ, HIVEP2, IFNGR2, IFT43, NME7, MYO1C, MAP4K4 |
| Conventional DC | IFNGR1, HLA-DOA, HACD4, PAK1, OTULINL, NAIP, MIS18BP1, NAGK, NUDT1, MAN2B1, LPXN, NSUN6, OSTF1, ZYX, MANBA, NOTCH2, MFSD1, MAP3K5, MGAT1, IFNAR2, IFNGR2, NFKB1, LTA4H, LSM10, MYO5A, PABPC4 |
| MAST | IL18R1, LINC02752, P2RX1, MLPH, LMO4, NDST2, LYL1, MAOB, GSE1, MANBA, MYO9B, GRK2, MBOAT7, NSMCE1, MYO15B, PARP4 |
Using the selected marker gene sets, we computed marker-derived activity scores for each cell type in each real ST spot. These activity scores represent the relative strength of the corresponding cell-type marker signal within each spot. For each evaluated cell type, we then compared the predicted cell-type proportion with the corresponding marker-derived activity across all spots. Importantly, marker-derived activity was not treated as a ground-truth cell-type proportion, but rather as an independent proxy for the spatial distribution of a given cell type in real ST data. Thus, if the marker signal for a cell type is high in a subset of spots, a biologically consistent deconvolution result should assign relatively higher abundance to that cell type in those spots.
We used Spearman correlation, Pearson correlation, and MAE as overall metrics for marker consistency. Spearman correlation primarily captures rank agreement between predicted abundance and marker-derived activity across spots, Pearson correlation measures linear agreement, and MAE quantifies the deviation between predicted abundance and the marker-derived signal. Figure 4b shows representative cell-type-level comparisons, covering immune cells, glomerular-associated cells, and relatively rare or spatially restricted cell populations. Across these representative cell types, CoDeST maintained stable marker agreement, suggesting that it can recover not only the spatial distribution of major cell populations but also signals from rare or locally enriched cell types.
Together, the Visium HD-derived benchmark and real-ST marker consistency analysis support the pseudo-to-real generalization ability of CoDeST from two complementary perspectives. These results indicate that CoDeST can recover cell-type composition in simulated mixtures and can also generate biologically consistent deconvolution results on real ST sections.
CoDeST preserves spatial coherence and boundary contrast while supporting robust deconvolution and programmed cell death mapping
Beyond cell-type deconvolution accuracy, we further evaluated whether CoDeST could produce spatially coherent PCD maps while preserving boundary contrast, and whether its key training stages and model components contributed to this behavior. To this end, we combined PCD spatial coherence analysis, stage-wise ablation, PCD weak-supervision ablation, and spatial graph sensitivity analysis to assess model stability in both deconvolution and PCD mapping tasks.
We first quantified the spatial behavior of model predictions using Moran’s I and Boundary
(Fig. 5a). Moran’s I measures spatial autocorrelation and was used to assess whether predicted PCD scores formed coherent spatial domains. Boundary
measures the difference in prediction changes between expression-defined boundary edges and non-boundary edges, and was used to evaluate whether the model preserved tissue boundaries and lesion fronts while smoothing local noise. These two metrics were, therefore, interpreted jointly to distinguish improved spatial coherence from excessive smoothing. The results showed that CoDeST maintained coherent PCD program predictions while preserving boundary contrast, suggesting that spatial graph aggregation and boundary-preserving constraints help generate PCD maps that better respect tissue structure.
Figure 5.

Ablation analysis of key training stages and weak-supervision modules in CoDeST. (a) Stage-wise ablation analysis. (b) Ablation analysis of PCD weak supervision and the confidence-weighting module.
We next performed stage-wise ablation to assess the contribution of each component in the three-stage training pipeline (Fig. 5a). Specifically, we removed Stage 1, Stage 2, or the real-ST domain adaptation stage, and evaluated the resulting models using pseudo-spot deconvolution performance, real-ST marker agreement, the PCD high-confidence benchmark, Moran’s I, and Boundary
. Removing Stage 2 led to the most pronounced decline in pseudo-spot deconvolution performance, indicating that pseudo-spot strong supervision is critical for learning the mapping to 34-cell-type compositions. Removing Stage 1 mainly affected real-ST marker agreement and spatial metrics, suggesting that self-supervised pretraining on real ST data provides a useful initialization for target-domain expression and spatial structure. Without real-ST domain adaptation, the model retained part of its pseudo-domain deconvolution ability, but marker agreement on real ST data decreased, indicating that the final pseudo-to-real adaptation stage is needed to reduce the domain gap in real ST sections.
We further evaluated the contribution of PCD weak supervision and confidence weighting (Fig. 5b). In the full CoDeST model, AUCell and ssGSEA were used as dual weak supervisory signals, and their agreement was used to estimate spot–program-level confidence weights. Compared with AUCell-only, ssGSEA-only, or unweighted dual-teacher supervision, the full model achieved higher Spearman correlation and lower MAE on the high-confidence AUCell–ssGSEA consensus benchmark. These results indicate that confidence weighting is not merely an auxiliary design choice, but helps reduce the influence of spot–program pairs for which AUCell and ssGSEA disagree, thereby improving the stability of PCD weak-supervision learning.
Finally, we performed a sensitivity analysis by varying the neighborhood size
of the spatial kNN graph to determine whether CoDeST depended on a specific graph construction parameter. All other training settings were kept unchanged. We assessed the effect of different
values on two types of outputs: Fig. 6a evaluates Moran’s I and Boundary
for the 34-cell-type deconvolution outputs, whereas Fig. 6b evaluates the same metrics for the four PCD program scores. Across the tested range of
values, CoDeST maintained relatively stable Moran’s I and Boundary
for both deconvolution and PCD outputs, indicating that its spatial modeling is robust to reasonable variation in graph density.
Figure 6.

Sensitivity analysis of the effect of spatial graph neighborhood size on CoDeST spatial modeling. (a) Moran’s I and Boundary
of deconvolution outputs across different
values. (b) Moran’s I and Boundary
of PCD program outputs across different
values.
Together, these analyses show that the spatial modeling and training strategy of CoDeST jointly support robust cell-type deconvolution and boundary-preserving PCD mapping. The stage-wise ablations demonstrate the complementary roles of self-supervised pretraining, pseudo-spot strong supervision, and real-ST domain adaptation. The PCD weak-supervision ablations support the utility of AUCell–ssGSEA dual-teacher signals and confidence weighting. The spatial graph sensitivity analysis further indicates that the model maintains a favorable balance between spatial coherence and boundary preservation across different neighborhood settings.
CoDeST reveals spatially heterogeneous programmed cell death programs in real spatial transcriptomics sections
On a representative KPMP kidney ST slice, CoDeST produces spatial maps of continuous intensities for four PCD programs with interpretable structure (Fig. 7). The four programs exhibit distinct spatial enrichment patterns on the same tissue background. Some regions show concordant, band-like high-score distributions, whereas others display program-specific hotspot and coldspot partitioning, consistent with differential regulation of death pathways across local microenvironments. Compared with visualizations based solely on raw gene-set scoring or naive graph smoothing, CoDeST yields more coherent within-region patterns while retaining sharp transitions at compartment interfaces and local structural changes, producing hotspot morphologies that better align with tissue partitioning. Comparison across the four maps further suggests partial co-occurrence in certain anatomical contexts and complementary distributions in others, providing a basis for subsequent analyses that integrate deconvolution estimates and marker evidence to separate compositional and state effects.
Figure 7.
Spatial maps of four PCD programs in a real ST section. (a) Spatial distribution of the apoptosis program. (b) Spatial distribution of the pyroptosis program. (c) Spatial distribution of the necroptosis program. (d) Spatial distribution of the ferroptosis program. Colors indicate the relative predicted activity score of each PCD program in each spot.
Discussion
In this study, we developed CoDeST, a three-stage pseudo-to-real framework for cell-type deconvolution and PCD mapping in ST data. Using kidney ST data as the primary application context, CoDeST integrates cell-type composition knowledge from an scRNA-seq reference, spatial structure from real ST sections, and PCD gene-set activity inferred from AUCell and ssGSEA. Through self-supervised pretraining on real ST data, supervised learning on donor-matched pseudo-spots, and real-ST domain adaptation, CoDeST aims to reduce the domain gap between simulated mixtures and real spatial tissue sections while jointly predicting cell-type proportions and PCD program activity.
Recent advances in spatial omics have highlighted the need for computational methods that jointly model molecular expression, tissue architecture, and spatially organized biological programs [44]. In parallel, graph-based and context-aware representation learning has been increasingly applied to biomedical data integration, including multi-omic disease association prediction and disease-contextualized regulatory modeling [45, 46]. CoDeST follows this broader direction by combining spatial graph aggregation, teacher–student transfer, weak biological supervision, and confidence weighting within an ST framework. Unlike general biomedical graph-learning studies, CoDeST is specifically designed for pseudo-to-real spatial deconvolution and PCD program mapping in tissue sections.
The three training stages of CoDeST play complementary roles. Stage 1 uses masked gene reconstruction and spatial graph perturbation to learn representations adapted to real ST expression sparsity and spatial structure. Stage 2 uses donor-matched pseudo-spots with known cell-type proportions to provide strong supervision for learning a 34-cell-type deconvolution mapping. Stage 3 then adapts the model to real ST data using teacher–student distillation, marker-based weak constraints, AUCell–ssGSEA dual-teacher PCD weak supervision, and boundary-preserving spatial regularization. The ablation analyses support the contribution of each stage: pseudo-spot supervision is important for cell-composition learning, real-ST pretraining improves target-domain representation, and final pseudo-to-real adaptation improves marker consistency and spatial stability on real ST sections.
A key feature of CoDeST is its use of weak supervision rather than requiring unavailable ground-truth labels in real ST data. Marker activity, AUCell, and ssGSEA were not treated as strict ground truth, but as proxy signals for target-domain learning. For PCD prediction, AUCell and ssGSEA provide complementary weak labels, and their agreement is used to assign spot–program-level confidence weights. This confidence-weighted dual-teacher design helps down-weight inconsistent proxy labels and reduces sensitivity to noisy weak supervision. Although this study focuses on apoptosis, pyroptosis, necroptosis, and ferroptosis, the same strategy could be extended to other spatial cell-state programs, such as proliferation, senescence, differentiation, inflammatory activation, fibrotic response, or metabolic reprogramming, provided that reliable gene sets, marker signatures, or external weak labels are available.
Several limitations should be noted. First, CoDeST is constrained by the 34-cell-type label space defined by the scRNA-seq reference. This ensures consistency across pseudo-spots, model outputs, and evaluation metrics, but prevents explicit identification of unknown, novel, or extremely rare cell types absent from the reference. In practice, such out-of-reference populations may be assigned to the most similar known cell type or may increase prediction uncertainty. Future work could incorporate open-set recognition, unknown cell-type detection, or reference refinement to flag potential out-of-reference signals.
Second, pseudo-spots remain an approximation of real spot composition. Although donor-aware mixing, realistic cell-number and library-size matching, noise injection, marker plausibility filtering, and rare-cell prior correction were used to improve realism, true tissue-level cell adjacency, cell-size variation, capture efficiency, regional injury states, and platform-specific noise cannot be fully simulated. The Visium HD-derived benchmark and real-ST marker consistency analysis provide complementary validation, but the Visium HD reference proportions still depend on segmentation and label transfer and should be interpreted as segmentation-derived references rather than absolute ground truth.
Third, CoDeST-predicted PCD scores should be interpreted as transcriptional program activity rather than direct measurements of cell death events. PCD-related gene sets may overlap with stress, inflammation, hypoxia, injury repair, or cell-composition signals. Thus, apoptosis, pyroptosis, necroptosis, and ferroptosis maps generated by CoDeST should be viewed as spatial PCD-related transcriptional activity profiles. Future validation using immunostaining, in situ imaging, pathology scores, or functional assays will be important for linking these predicted programs to specific biological mechanisms.
Finally, CoDeST spatial modeling depends on graph construction and boundary-preserving regularization. Sensitivity analysis showed stable Moran’s I and Boundary
across a reasonable range of kNN neighborhood sizes, but graph construction may still be affected by tissue morphology, spot density, section quality, and platform resolution. Future work could incorporate histology, tissue compartments, pathological boundaries, or adaptive neighborhood selection to better align the spatial graph with tissue architecture.
Overall, CoDeST provides an extensible framework for integrating deconvolution, weakly supervised cell-state inference, and spatial tissue-structure modeling in ST data. Despite limitations related to reference coverage, weak-label noise, and spatial graph construction, our results suggest that CoDeST can recover cell-type composition in simulated and Visium HD-derived benchmarks and generate biologically consistent deconvolution and PCD maps in real kidney ST sections.
Key Points
CoDeST (Confidence-weighted Dual-teacher Spatial Training) jointly infers spot-level cell-type composition across 34 kidney cell types and predicts continuous activity scores for four programmed cell death (PCD)-related transcriptional programs: apoptosis, pyroptosis, necroptosis, and ferroptosis.
The framework uses a three-stage pseudo-to-real training strategy, combining real-spatial transcriptomics (ST) self-supervised pretraining, supervised deconvolution on donor-matched pseudo-spots, and real-ST domain adaptation.
AUCell and ssGSEA provide dual weak supervisory signals for PCD mapping, with agreement-based confidence weighting to reduce the influence of inconsistent proxy labels.
Boundary-preserving spatial regularization supports spatially coherent PCD maps while limiting excessive smoothing across tissue boundaries.
Donor-held-out pseudo-spot benchmarks, marker consistency on real ST sections, and a Visium HD-derived benchmark support CoDeST’s utility within the evaluated kidney ST datasets.
Contributor Information
Chunling Wu, Department of Rheumatology and Immunology, The First Hospital of China Medical University, No.155 Nanjing Street of Heping District, 110001 Liaoning Province, China.
Xiaomeng Luo, Department of Rheumatology and Immunology, The First Hospital of China Medical University, No.155 Nanjing Street of Heping District, 110001 Liaoning Province, China.
Yuansong Zhao, The University of Texas Health Science Center at Houston, 7000 Fannin Street, Houston, TX 77030, United States.
Ying Chen, Department of Nephrology, The First Hospital of China Medical University, No. 155 Nanjing Street of Heping District, 110001 Liaoning Province, China.
Nan Zuo, Department of Nephrology, The First Hospital of China Medical University, No. 155 Nanjing Street of Heping District, 110001 Liaoning Province, China.
Conflicts of interest
None declared.
Funding
None declared.
Data availability
The code and data for CoDeST are publicly available at the following GitHub repository: https://github.com/Peiczn/CoDeST.
References
- 1. Li C, Yu Y, Zhu S et al. The emerging role of regulated cell death in ischemia and reperfusion-induced acute kidney injury: current evidence and future perspectives. Cell Death Discov 2024;10:216. 10.1038/s41420-024-01979-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Sanz AB, Sanchez-Niño MD, Ramos AM et al. Regulated cell death pathways in kidney disease. Nat Rev Nephrol 2023;19:281–99. 10.1038/s41581-023-00694-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Mu J, Yao Y, Chen Q et al. GATCL: graph attention network meets contrastive learning for spatial domain identification. Brief Bioinform 2026;27:bbag043. 10.1093/bib/bbag043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Xuanyuan Q, Wu H, Humphreys BD. Emerging high-resolution spatial transcriptomic technologies in kidney research. Nephrol. Dial. Transplant. 2024;39:1747–50. 10.1093/ndt/gfae125 [DOI] [PubMed] [Google Scholar]
- 5. Li X, Wang CY. From bulk, single-cell to spatial RNA sequencing. Int J Oral Sci 2021;13:36. 10.1038/s41368-021-00146-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Lake BB, Menon R, Winfree S et al. An atlas of healthy and injured cell states and niches in the human kidney. Nature 2023;619:585–94. 10.1038/s41586-023-05769-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Ouologuem S, Martens LD, Schaar AC et al. Spatial transcriptomics deconvolution methods generalize well to spatial chromatin accessibility data. Bioinformatics 2025;41:i314–22. 10.1093/bioinformatics/btaf268 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Ma Y, Zhou X. Spatially informed cell-type deconvolution for spatial transcriptomics. Nat Biotechnol 2022;40:1349–59. 10.1038/s41587-022-01273-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Jia Y, Dong H, Li L et al. xQTLatlas: a comprehensive resource for human cellular-resolution multi-omics genetic regulatory landscape. Nucleic Acids Res 2025;53:D1270–7. 10.1093/nar/gkae837 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Liang Q, Huang Y, He S et al. Pathway centric analysis for single-cell RNA-seq and spatial transcriptomics data with GSDensity. Nat Commun 2023;14:8416. 10.1038/s41467-023-44206-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Wang R, Thakar J. Comparative analysis of single-cell pathway scoring methods and a novel approach. NAR Genom Bioinform 2024;6:lqae124. 10.1093/nargab/lqae124 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Toro-Domínguez D, Wang C, Ellson-Lancho I et al. Benchmarking single-sample gene set scoring methods for application in precision medicine. Brief Bioinform 2025;26:bbaf684. 10.1093/bib/bbaf684 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Kleshchevnikov V, Shmatko A, Dann E et al. Cell2location maps fine-grained cell types in spatial transcriptomics. Nat Biotechnol 2022;40:661–71. 10.1038/s41587-021-01139-4 [DOI] [PubMed] [Google Scholar]
- 14. Andersson A, Bergenstråhle J, Asp M et al. Single-cell and spatial transcriptomics enables probabilistic inference of cell type topography. Commun Biol 2020;3:565. 10.1038/s42003-020-01247-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Lopez R, Li B, Keren-Shaul H et al. DestVI identifies continuums of cell types in spatial transcriptomics data. Nat Biotechnol 2022;40:1360–9. 10.1038/s41587-022-01272-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Cable DM, Murray E, Zou LS et al. Robust decomposition of cell type mixtures in spatial transcriptomics. Nat Biotechnol 2022;40:517–26. 10.1038/s41587-021-00830-w [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Elosua-Bayes M, Nieto P, Mereu E et al. SPOTlight: seeded NMF regression to deconvolute spatial transcriptomics spots with single-cell transcriptomes. Nucleic Acids Res 2021;49:e50–0. 10.1093/nar/gkab043 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Biancalani T, Scalia G, Buffoni L et al. Deep learning and alignment of spatially resolved single-cell transcriptomes with tangram. Nat Methods 2021;18:1352–62. 10.1038/s41592-021-01264-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Zhao T, Liu R, Sun Y et al. DECODE: Deep learning-based common deconvolution framework for various omics data. Nat Methods 2026;23:596–608. 10.1038/s41592-026-03007-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Lu S, Yang J, Yan L et al. Transcriptome size matters for single-cell RNA-seq normalization and bulk deconvolution. Nat Commun 2025;16:1246. 10.1038/s41467-025-56623-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Fan C, Chen F, Chen Y et al. irGSEA: the integration of single-cell rank-based gene set enrichment analysis. Brief Bioinform 2024;25:bbae243. 10.1093/bib/bbae243 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22. Noureen N, Ye Z, Chen Y et al. Signature-scoring methods developed for bulk samples are not adequate for cancer single-cell RNA sequencing data. eLife 2022;11:e71994. 10.7554/eLife.71994 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Wright D, Augenstein I. Aggregating soft labels from crowd annotations improves uncertainty estimation under distribution shift. PLoS One 2025;20:e0323064. 10.1371/journal.pone.0323064 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Song H, Kim M, Park D et al. Learning from noisy labels with deep neural networks: a survey. IEEE Trans Neural Netw Learn Syst 2023;34:8135–53. 10.1109/TNNLS.2022.3152527 [DOI] [PubMed] [Google Scholar]
- 25. Kurup AR, Kar MK, Hashunao S et al. A truth inference scheme for crowdsourcing using NLP and swin transformers. Sci Rep 2025;15:28338. 10.1038/s41598-025-10942-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Qin X, Liu C, Gu F et al. Probabilistic-graph-based spatial context-aware framework for interpretable spatial omics denoising and augmentation. Brief Bioinform 2025;26:bbaf674. 10.1093/bib/bbaf674 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Singhal V, Chou N, Lee J et al. BANKSY unifies cell typing and tissue domain segmentation for scalable spatial omics data analysis. Nat Genet 2024;56:431–41. 10.1038/s41588-024-01664-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Sun Y, Liu R, Zhang L et al. GE-MFAT: Heterogeneous graph feature transfer with focusing attention for protein-metabolite interaction prediction. In: Liu J, Huang J, Wang X, Zhang F, Zou X, Tian T, Hu X, Hu B, Xiong Y (eds.), 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 7719–26. Wuhan, China: IEEE, 2025. https://ieeexplore.ieee.org/document/11356109/ [Google Scholar]
- 29. Dong M, Su DG, Kluger H et al. SIMVI disentangles intrinsic and spatial-induced cellular states in spatial omics data. Nat Commun 2025;16:2990. 10.1038/s41467-025-58089-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Dong K, Zhang S. Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder. Nat Commun 2022;13:1739. 10.1038/s41467-022-29439-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Zhang L, Sun Y, Liu R et al. GDTGO: advancing protein function prediction via graph convolutional network and iterative optimization. In: Liu J, Huang J, Wang X, Zhang F, Zou X, Tian T, Hu X, Hu B, Xiong Y (eds.), 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 7743–9. Wuhan, China: IEEE, 2025. https://ieeexplore.ieee.org/document/11356544/ [Google Scholar]
- 32. Hu J, Li X, Coleman K et al. SpaGCN: integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network. Nat Methods 2021;18:1342–51. 10.1038/s41592-021-01255-8 [DOI] [PubMed] [Google Scholar]
- 33. Shang L, Zhou X. Spatially aware dimension reduction for spatial transcriptomics. Nat Commun 2022;13:7203. 10.1038/s41467-022-34879-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Li Y, Luo Y. STdGCN: spatial transcriptomic cell-type deconvolution using graph convolutional networks. Genome Biol 2024;25:206. 10.1186/s13059-024-03353-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Bae S, Choi H, Lee DS. spSeudoMap: cell type mapping of spatial transcriptomics using unmatched single-cell RNA-seq data. Genome Med 2023;15:19. 10.1186/s13073-023-01168-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36. Dai S, Li J, Xia Z et al. ST-deconv: an accurate deconvolution approach for spatial transcriptome data utilizing self-encoding and contrastive learning. NAR Genom Bioinform 2025;7:lqaf109. 10.1093/nargab/lqaf109 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37. Cao K, Wei S, Ma T et al. Integrating bulk, single-cell, and spatial transcriptomics to identify and functionally validate novel targets to enhance immunotherapy in NSCLC. NPJ Precis Oncol 2025;9:112. 10.1038/s41698-025-00893-x [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38. Wang X, Lian Q, Dong H et al. Benchmarking algorithms for gene set scoring of single-cell ATAC-seq data. Genom Proteom Bioinform 2024;22:qzae014. 10.1093/gpbjnl/qzae014 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39. Zhang X, Wu Z, Wang T et al. DANST enables cell-type deconvolution in spatial transcriptomics using deep domain adversarial neural networks. Commun Biol 2026;9:388. 10.1038/s42003-026-09659-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Wang L, Hu Y, Gao L. Adjustment of scRNA-seq data to improve cell-type decomposition of spatial transcriptomics. Brief Bioinform 2024;25:bbae063. 10.1093/bib/bbae063 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. De Boer IH, Alpers CE, Azeloglu EU et al. Rationale and design of the kidney precision medicine project. Kidney Int 2021;99:498–510. 10.1016/j.kint.2020.08.039 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Liberzon A, Birger C, Thorvaldsdóttir H et al. The molecular signatures database Hallmark gene set collection. Cell Syst 2015;1:417–25. 10.1016/j.cels.2015.12.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Zhou N, Bao J. FerrDb: a manually curated resource for regulators and markers of ferroptosis and ferroptosis-disease associations. Database 2020;2020:baaa021. 10.1093/database/baaa021 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Zhang L, Xiong Z, Xiao M. A review of the application of spatial Transcriptomics in neuroscience. Interdiscip Sci: Comput Life Sci 2024;16:243–60. 10.1007/s12539-024-00603-4 [DOI] [PubMed] [Google Scholar]
- 45. Wang Y, Shen W, Shen Y et al. Integrative graph-based framework for predicting circRNA drug resistance using disease contextualization and deep learning. IEEE J Biomed Health Inform 2025;29:7932–44. 10.1109/JBHI.2024.3457271 [DOI] [PubMed] [Google Scholar]
- 46. Wang Y, Wang Z, Wang T et al. HiGLDP: a hierarchical graph neural network for predicting lncRNA-disease associations through multi-omic integration. BMC Biol 2026;24:80. 10.1186/s12915-026-02557-z [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The code and data for CoDeST are publicly available at the following GitHub repository: https://github.com/Peiczn/CoDeST.













