Abstract
Spatially resolved genomics and transcriptomics are reshaping our understanding of tumor evolution and therapeutic resistance, yet benchmarking spatial copy number variation (CNV), single-nucleotide variant (SNV), and spatial mutation-burden proxy analyses is constrained by the scarcity of datasets with known ground truth. Existing simulators often produce only count matrices, lack matched DNA–RNA outputs, or do not propagate genomic variation to sequencing-level signals, limiting end-to-end benchmarking of multi-omics pipelines, including virtual cell-oriented multimodal benchmarks. Here, we present STGBench, a sequencing-level spatial DNA–RNA simulator that generates paired DNA-seq alignments (BAM files) and matched gene expression matrices on a user-defined 2D tissue grid. STGBench builds tissue masks from geometric templates or image-derived masks, overlays spatial CNV landscapes and SNV/VAF fields in boundary, gradient, and nested modes, and synthesizes DNA and RNA readouts by coupling copy number states to expression under a negative binomial model with spatially correlated technical effects; outputs are directly consumable by downstream tools. Using AneuFinder on simulated DNA data, spatial CNV profiles are recovered with Pearson r up to 0.855. Applying InferCNV to simulated transcriptomes, CNV-driven expression signatures are reproduced (r up to 0.996) and support unsupervised structure consistent with clonal organization; DNA-derived and RNA-inferred CNVs show concordance (r ≈ 0.77). CellSNP recovers diverse spatial SNV patterns from simulated reads, and IGV inspection confirms realistic allelic balance and CNV-associated coverage shifts at nucleotide resolution. Collectively, STGBench provides a controllable benchmark generator for spatial CNV/SNV and mutation-burden analyses with explicit ground truth across paired DNA–RNA modalities. STGBench is open source at https://github.com/Icarus200110/STGBench.
Keywords: tumor mutation burden, spatial clonal heterogeneity, multimodal benchmarking, sequencing-level simulation, probabilistic generative model
Introduction
Cancer is inherently a spatial disease. Clonal expansion, immune evasion, and therapeutic resistance emerge from local interactions between genetically diverse tumor cells and their microenvironment, unfolding in space and time rather than in well-mixed bulk [1–3]. Recent advances in spatially resolved DNA and RNA sequencing have made it possible to map copy number variation (CNV), single-nucleotide variants (SNVs), tumor mutation burden (TMB, indicated by the spatial mutation-burden proxy), and gene expression at near-cellular resolution across intact tissue sections [4–9]. These technologies are beginning to reveal how spatially structured genomic alterations shape transcriptional programs, microenvironmental niches, and clinical trajectories in ways that are invisible to bulk or nonspatial single-cell assays [10–12].
In parallel, the analysis toolbox for spatial omics has expanded rapidly. Methods for CNV detection, SNV calling, TMB quantification, clone deconvolution, and multimodal integration are now being adapted or newly developed for spatial data, often by extending tools originally designed for bulk or single-cell DNA and RNA sequencing [13–16]. For example, CNV callers such as AneuFinder are being applied to spatial DNA-seq, while tools like InferCNV infer large-scale CNVs from spatial transcriptomics; CopyKAT provides complementary CNV-from-RNA capabilities [9, 13–15]. Downstream, these genomic layers are increasingly combined with spatially resolved expression to dissect clonal architecture, quantify microenvironment-specific TMB, and prioritize therapeutic targets [17–20].
Yet, despite this methodological growth, systematic benchmarking and stress-testing of spatial mutation analysis pipelines remain extremely challenging. At the core of this difficulty lies the absence of spatial datasets with complete, cell-level ground truth. Real tumors provide rich biological complexity but only partial truth labels, often restricted to curated loci or coarse region-level annotations. Public spatial omics datasets are heterogeneous in protocols, coverage, and quality, making it hard to disentangle method performance from dataset-specific artefacts [12, 16, 21]. As a result, new algorithms are frequently evaluated on a small number of real datasets, with limited ability to probe failure modes, sensitivity to spatial patterns, or cross-modal consistency between DNA and RNA-derived readouts. This situation mirrors earlier stages in single-cell RNA-seq and variant-calling research, where the lack of realistic benchmark data limited the interpretability and reproducibility of performance claims [16].
Simulation offers a principled way to address these issues: by generating synthetic data under explicit generative models, one can embed known ground-truth mutation patterns and systematically vary biological and technical factors [22, 23]. However, existing simulators are ill-suited for the specific demands of spatial mutation analysis. Many tools focus on nonspatial single-cell RNA-seq, modeling expression counts without explicit spatial coordinates or matched DNA. Others simulate tumor evolution at the genotype level but do not output realistic sequencing reads, or treat spatial structure only implicitly. Critically, very few frameworks model DNA and RNA jointly on a single spatial scaffold, propagate copy number and SNV events into both sequencing-level DNA outputs and transcriptomes, and produce data that can be fed directly into standard analysis pipelines. This gap increasingly limits not only conventional method benchmarking, but also multimodal evaluation settings in which spatial DNA and RNA are jointly exploited to recover CNVs, quantify mutation burden and reconstruct clonal structure. More broadly, as virtual cell-oriented initiatives and multimodal modeling workflows place growing emphasis on paired modalities, standardized interfaces and reproducible benchmarks, the lack of controllable, sequencing-level spatial DNA–RNA simulators with explicit ground truth becomes a key bottleneck for rigorous stress-testing and fair comparison across computational pipelines [17, 18].
To fill this gap, we developed STGBench, a sequencing-level simulation platform that synthesizes spatially resolved DNA sequencing data and matched gene expression matrices through a probabilistic, multilayer architecture. STGBench models a tissue slice as a discrete 2D grid, with a flexible tissue mask that can be specified as idealized geometric shapes or derived from real histology via binary images and morphological processing. On this spatial scaffold, the platform defines configurable CNV and SNV probability fields using a small set of interpretable pattern types—boundary, gradient, and nested—that capture common clonal topologies in solid tumors, from sharp tumor–normal interfaces to continuous instability gradients and concentric clonal expansions. These fields determine the local copy number state and variant allele frequency (VAF) at each position, forming a ground-truth mutational landscape that explicitly encodes spatial heterogeneity. In the final layer of the generative process, STGBench translates these latent genomic fields into observable DNA and RNA readouts. For DNA, it generates spatially indexed sequencing reads whose coverage and allelic composition reflect the underlying CNVs and SNVs, producing BAM files that are indistinguishable in structure from experimental spatial DNA-seq data and directly compatible with off-the-shelf variant callers. For RNA, STGBench starts from a reference expression matrix and modulates the expected expression of each gene as a function of local copy number, thereby embedding genome-transcriptome coupling directly into the simulator. Counts are then sampled from a negative binomial distribution with spatially correlated mean parameters, capturing both overdispersed single-cell–like expression and realistic technical effects such as broad expression waves and spatially structured noise. The result is a spatial gene expression matrix that responds coherently to the simulated CNV landscape while maintaining realistic variability at the single-gene and single-cell levels. By design, STGBench is not tied to any particular downstream method. Instead, it aims to provide an end-to-end benchmark generator on which diverse pipelines can be deployed under controlled conditions.
In this work, we demonstrate its utility through a multi-tier evaluation focusing on CNV, SNV, and cross-modal consistency. First, we show that standard CNV callers applied to STGBench-generated DNA-seq data can recover user-specified spatial CNV patterns with high fidelity: AneuFinder accurately reconstructs complex CNV gradients and nested structures across the simulated tissue, with genome-wide correlations between ground truth and inferred copy number reaching 0.855 [13]. Second, we use STGBench’s RNA output to benchmark CNV inference from gene expression, demonstrating that InferCNV recapitulates the imposed CNV landscapes with near-perfect correlation in some scenarios (r up to 0.996) and that unsupervised clustering based solely on expression separates clonal regions in line with underlying genomic structure [15, 16]. Third, we directly quantify cross-modal concordance between CNVs derived from DNA-seq and those inferred from RNA, observing strong agreement (r ≈ 0.77) despite expected discrepancies due to biological regulation and algorithmic noise; we also relate these analyses to recent spatial DNA–RNA integration frameworks [17, 18]. Further analyses show that CellSNP can robustly recover diverse spatial SNV patterns from STGBench data, and IGV inspection confirms that simulated reads exhibit realistic allelic balance and coverage shifts at the nucleotide level [24, 25]. Collectively, these results establish STGBench as a versatile platform for benchmarking spatial CNV, SNV, and spatial mutation-burden proxy analysis in a multimodal setting. By linking tissue morphology, spatially structured mutational fields and DNA–RNA readouts in a unified generative model, STGBench allows developers to design targeted stress tests, explore the limits of existing methods and rationally optimize new algorithms before deployment on precious clinical samples. We anticipate that STGBench will serve as a common reference for the spatial omics community, facilitating more transparent, reproducible, and biologically grounded evaluation of computational pipelines that seek to unlock the full potential of spatial genomics and transcriptomics in cancer research and precision oncology [1–3, 10, 21].
Method
Overview of the methods
STGBench is a spot-level sequencing-level simulation platform tailored for mainstream spot-based spatial multi-omics capture technologies, which generates paired spatial DNA–RNA outputs—including alignment-ready DNA-seq BAM files and matched gene expression matrices—under an explicit, controllable ground-truth model. The framework represents a tissue section as a spatial stochastic process with spatially indexed capture spots as the core basic simulation unit, and propagates user-defined genomic alterations, specifically CNV and SNVs, into both read-level DNA measurements and transcriptional consequences while preserving the spatial coordinates of each spot. STGBench is organized as a probabilistic multilayer architecture with three sequential modules: (i) tissue morphology construction to define the spatial scaffold, (ii) spatial mutation-field specification to generate CNV landscapes and SNV/VAF patterns, and (iii) observation-layer synthesis to produce pipeline-compatible DNA-seq alignments and transcriptomic count matrices from the shared latent mutational state. This modular design enables controlled stress tests and end-to-end benchmarking of spatial mutation-analysis workflows across DNA, RNA, and their cross-modal consistency. An overview of the STGBench workflow is shown in Fig. 1.
Figure 1.
Overview of the STGBench probabilistic framework for joint spatial DNA–RNA simulation. STGBench consists of three sequential modules. Module 1 constructs a tissue mask on a discrete 2D grid from geometric templates or an image-derived mask. Module 2 generates spatially structured ground-truth mutation fields, including CNV copy number landscapes and SNV/VAF probability fields under boundary, gradient, or nested modes. Module 3 synthesizes pipeline-ready spatial DNA-seq outputs (reads/aligned BAM) and matched spatial gene expression matrices by propagating the mutation fields to read-level variation and transcriptional consequences, enabling end-to-end benchmarking with explicit ground truth.
Spatial scaffold and tissue mask construction
STGBench first defines the spatial support on which all downstream genomic and transcriptomic processes are simulated. Specifically, we represent tissue as a discrete 2D capture grid and construct a binary tissue mask that marks valid spot-capture locations. This module provides a flexible interface to simulate both idealized geometries and realistic tissue shapes extracted from images, thereby decoupling morphological structure from downstream mutation and sequencing generation. The resulting mask serves as a common spatial scaffold for CNV/SNV field construction and for generating paired DNA–RNA outputs at matched coordinates.
The spatial domain is defined as a discrete 2D grid:
![]() |
where R and C are user-specified grid dimensions.
The tissue mask identifies valid spot capture positions:
![]() |
The tissue mask identifies valid spot capture positions through a mode parameter
∈{circle, square, image} as follows:
![]() |
In circle mode,
represents the center coordinates and r represents the radius; in square mode,
represents the center coordinates and s represents the half side length; in image mode, the mask is generated from a user-provided binary image I. The generation process includes: edge detection using the Canny algorithm, morphological dilation with a 3 × 3 convolution kernel, and hole filling to create continuous regions. This flexible approach can simulate both idealized geometric tissue shapes and complex, biologically realistic tissue morphologies extracted from real histology images.
This module outputs a reproducible binary tissue mask and the corresponding list of valid spatial coordinates that define where simulated spots exist. All downstream CNV/SNV fields and multi-omics outputs are generated only on these valid locations, ensuring that tissue morphology acts as an explicit boundary condition rather than an implicit artifact of sampling. Using the same parameters (circle/square) or the same image preprocessing procedure (image mode) yields identical masks, which is essential for fair benchmarking across repeated simulations and across methods evaluated on the same spatial support.
Spatial CNV landscape generation
To model spatial clonal structure at the copy number level, STGBench generates a ground-truth CNV field over the tissue grid, parameterized by a small set of interpretable spatial modes. The CNV module outputs an integer-valued copy number matrix that captures sharp boundaries, continuous gradients, or concentric zone structures, thereby reflecting common tumor evolution scenarios such as distinct clones, genome instability gradients, or core–margin compartmentalization. Importantly, CNV generation is defined independently of sequencing depth and caller assumptions: the CNV field is later propagated to read-level coverage and to transcriptional consequences, enabling end-to-end benchmarking with explicit truth.
The CNV field is defined through a unified function with mode parameter θ ∈ {boundary, gradient, nested}:
![]() |
In the above formula (the hardcoded value 2 is set based on the diploid genomic background of the human hg38 reference genome, representing the baseline copy number of normal somatic spots without copy number variation),
represents the normalized Euclidean distance from a preset spatial focus (e.g. tumor center) to the given position, ensuring scale invariance of the pattern. The boundary mode creates a sharp clonal boundary along the diagonal, simulating two distinct spot populations; the gradient mode models continuous CNV variation, recapitulating effects of clonal expansion or genome instability gradients (where k controls gradient steepness); the nested mode defines concentric zones with different CNV levels, determined by radius parameters 
The final CNV matrix is generated by masking with tissue geometry:
![]() |
This module produces an integer-valued ground-truth CNV landscape over the tissue grid, with non-tissue locations explicitly excluded. Because the CNV field is controlled by a small set of interpretable parameters (mode choice, copy number bounds, and geometry controls), users can generate families of datasets with systematically varied difficulty (e.g. sharper versus smoother transitions) while retaining known truth for quantitative evaluation. The CNV landscape is subsequently propagated to both DNA-seq read-depth patterns and transcriptional consequences, providing a shared latent genomic state that supports downstream cross-modal consistency testing.
Spatial SNV/VAF Field and Mutation Burden (spatial mutation-burden proxy) Specification
Complementary to CNVs, STGBench defines an SNV probability field that specifies the spatial distribution of VAF across the same tissue scaffold. Using the same mode family as the CNV module ensures that SNV patterns can be coordinated with CNV-driven clonal architecture when desired, while remaining fully controllable through a compact parameterization. This module outputs a spatial VAF map that serves as ground truth for allele sampling during read generation, enabling locus-level validation (SNV detectability) as well as aggregated burden-level analyses.
The SNV field defines VAF across the spatial grid, employing the same spatial framework and mode types as the CNV field to ensure coordination of genomic variations:
![]() |
These patterns simulate key mutational processes: boundary can model trunk mutations present in one major clone; gradient can represent subclonal mutations expanding from a founder spot; nested is particularly effective for simulating mutations across different clonal states, such as ubiquitous mutations in the tumor core
, absent mutations in the normal margin (implicitly VAF = 0), and heterozygous states in intermediate regions (simulating loss of heterozygosity or cellular mixing).
The SNV module outputs a spatial VAF truth map that governs allele sampling during read generation at each locus and spatial location. The shared mode family with the CNV module enables coordinated specification of SNV and CNV heterogeneity under a unified spatial scaffold, while still allowing independent control when desired. In downstream evaluation, this output supports both locus-level validation (recovery of spatial SNV patterns by standard callers) and burden-level aggregation across loci (spatially heterogeneous local mutation burden as a proxy for spatial mutation-burden proxy), while maintaining explicit truth for each simulated variant. To rigorously quantify this aggregated local burden, the spatial mutation-burden proxy (
) at a given spatial spot
is calculated by aggregating the estimated VAF across a predefined panel of
SNV loci. Mathematically, it is defined as the sum of variant allele frequencies within the region:
![]() |
Alternatively, if a threshold-based discrete mutation count is applied, the proxy is calculated as follows:
![]() |
where
is an indicator function that equals
if the mutation’s VAF exceeds a defined detection threshold
, and
otherwise. This formulation strictly defines our proxy as a relative quantitative metric characterizing spatial mutational heterogeneity within a fixed panel, structurally distinguishing it from clinical TMB evaluated across the callable genome.
CNV-coupled spatial expression simulation
To generate matched spatial transcriptomes that are biologically coherent with the simulated genomic states, STGBench propagates gene-specific CNV dosage effects into gene expression and then samples counts under an overdispersed distribution. Starting from a user-provided reference expression matrix, the model adjusts the expected expression of each gene individually according to the copy number state of the genomic interval where the gene is located, and introduces spatially correlated technical effects to emulate structured noise commonly observed in spatial transcriptomics experiments.
The core implementation of the CNV-expression coupling module relies on two mandatory input files: (i) a gene annotation file containing the chromosome coordinates (chromosome number, start and end genomic positions) of each target gene; (ii) predefined CNV events with explicit genomic interval boundaries and spatial distribution patterns. For each gene, the model first retrieves its complete genomic location from the annotation file, matches it to the corresponding CNV interval in the pre-constructed spatial CNV landscape, and obtains the gene- and spot-specific copy number value for each spatial spot. This matching process ensures that only genes located within the genomic interval of a CNV event undergo expression adjustment, avoiding uniform expression shifts across the entire transcriptome. The module outputs a spot-by-gene count matrix whose expression signal is coupled to the underlying gene-resolved CNV landscape, enabling cross-modal consistency benchmarking between DNA-derived CNVs and expression-inferred CNVs.
Let
be the reference expression matrix, where N is the number of spots and G is the number of target genes. For gene g in a spatial spot at position (i,j):
Retrieve the genomic coordinate interval of gene g from the input gene annotation file and match it to the predefined CNV genomic intervals to obtain the gene- and spot-specific copy number value
;
Set
=2 as the baseline copy number of normal diploid somatic spots, consistent with the diploid background of the human reference genome used in this study;
Calculate the CNV-adjusted mean expression of gene g at spot (i,j) as:
![]() |
To simulate spatially correlated technical noise, a negative binomial distribution is introduced with its complete probability mass function:
![]() |
Here, the spatially varying mean
is:
![]() |
where ρ controls the magnitude of spatial correlation; λ is spatial frequency determining the size of high/low expression patches; ϕ is the dispersion parameter of the negative binomial distribution (with τ = 1/ϕ representing overdispersion intensity). This formulation accurately captures both the overdispersed count nature of spot gene expression data and introduces realistic, globally continuous spatial technical effects.
This module generates a simulated spot-by-gene expression count matrix whose expected expression levels are coupled to local CNV states and whose variability reflects realistic overdispersion and spatially structured technical effects. The parameterization separates biological signal (CNV-driven mean shifts) from technical structure (count dispersion and spatial correlation), enabling controlled stress-testing of downstream methods under different noise regimes without changing the underlying CNV truth. The resulting transcriptomes can therefore be used to benchmark expression-derived CNV inference and multimodal integration methods in a reproducible setting anchored to DNA-level ground truth.
Spatial sequencing read synthesis and base-level probabilistic modeling
To address the synthesis of sequencing reads and translate the latent genomic fields into observable data, STGBench employs a custom probabilistic read synthesis engine rather than wrapping opaque external tools. This engine operates independently for each valid spatial capture location
defined by the tissue mask
, providing explicit control over base qualities, sequencing errors, and spatial barcoding.
Reference genome and targeted region definition
All simulations are performed using the GRCh38 (hg38) human reference genome assembly. For targeted sequencing simulations, genomic regions of interest are defined via a standard Browser Extensible Data (BED) file, with each entry specifying the chromosome, 0-based start coordinate, and 1-based end coordinate of the target interval. The BED file is preprocessed to merge overlapping intervals and remove redundant regions. The total size of the target region (
) is calculated as the sum of the lengths of all non-overlapping target intervals:
![]() |
where
is the number of non-overlapping target intervals. To map random positions within the target region to absolute genomic coordinates, we precompute a cumulative length index of the sorted target intervals and use binary search to achieve
time complexity for coordinate conversion, ensuring efficient position sampling for large target panels.
Spatially resolved NGS library and sequencing simulation framework
The framework models the full process of next-generation sequencing (NGS) library preparation and sequencing for each valid spot
, with parameters configurable to match Illumina short-read platforms (e.g. read length, insert size distribution, Polymerase Chain Reaction (PCR) duplication rate, and sequencing mode). For each spatial location
, the theoretical number of template DNA fragments required to achieve the baseline sequencing depth is calculated as follows:
![]() |
where
is the target average sequencing depth,
is the mean template fragment length,
is the PCR duplication rate, and the factor 10 is empirically calibrated to match the observed depth distribution. Template fragments are randomly sampled from the target region, with fragment lengths sampled uniformly from the user-specified insert size range. PCR duplicates are simulated by generating multiple identical read sets from the same template fragment at a rate matching the user-defined duplication rate.
Genomic variant simulation modules (SNV and CNV)
The framework introduces SNVs and CNVs by dynamically fetching parameters from the previously defined spatial mutation fields.
Copy number variation simulation
For a spot at
, the local copy number state is governed by the ground-truth CNV landscape
. CNVs are simulated by modulating the number of template fragments sampled from a given genomic region. For any genomic position
, the effective copy number multiplier
is determined by
. To accommodate non-integer copy numbers (e.g. for subclonal mixing), we use a stochastic rounding strategy where
. The actual amplification multiplier
is sampled as
with probability
. The number of amplification loops for a fragment is then defined as
, ensuring the sequencing depth is linearly proportional to the local spatial copy number.
Single-nucleotide variant simulation
SNVs are introduced via site-specific base substitutions in template DNA fragments. For a spot at
, the VAF (rate) is dynamically fetched from the spatial SNV/VAF field
. For each generated template fragment spanning the SNV locus, a random number
is generated. The variant is introduced if
. Formally, let
denote the
-th generated template fragment. The mutated sequence
is defined as follows:
![]() |
where
denotes the site-specific sequence modification operator,
is the indicator function,
, and
denotes the base substitution replacing the reference base with the alternative base.
Coordinated modeling
When an SNV locus falls within a CNV region at spot
, the observed allele frequency is modulated by the local copy number
to ensure biological plausibility (e.g. a heterozygous SNV in a deletion region yielding an expected AF of 1.0). Variants are introduced in a fixed biologically informed order (insertions
SNVs
inversions
gene fusions) to avoid coordinate conflicts.
Sequencing error and base quality profile modeling
To recapitulate technical noise without relying on external wrappers, the framework incorporates a custom probabilistic per-base sequencing error model. For each base in the generated reads, an error is introduced with probability
. Base quality scores (
) are modeled to match the Phred-scale distribution, where
. Separate uniform distributions are used to sample quality scores for error-free bases (
) and erroneous bases (
).
Parallelization and spatial BAM output
The synthesis engine implements a multi-process, lock-free parallel architecture. A global pseudo-random number generator seed is used to ensure full reproducibility. Crucially, to preserve the spatial origin of each read for downstream analysis, every generated read is tagged with a spatial cellular barcode (e.g. standard CB tag in BAM format) corresponding to its grid coordinate
. Reads are then aligned to the hg38 reference genome using BWA-MEM, sorted, and deduplicated using SAMtools to generate the final spatially indexed, alignment-ready BAM file (and BAI index), enabling seamless integration with spatial downstream callers like CellSNP.
Results
We developed a sequencing-level spatial multi-omics simulator that generates spatially indexed DNA-seq outputs (including alignment-ready BAM files) together with matched gene expression matrices, while embedding explicit ground-truth CNV and SNV/VAF landscapes on a shared spatial scaffold. To assess both fidelity and pipeline readiness of the simulated data, we performed a tiered evaluation spanning genomic, cross-modal, and read-level properties.
At the CNV tier, we injected predefined segmental copy number states and quantified (i) recovery from simulated DNA-seq using standard CNV callers and (ii) concordance with CNV signals inferred from the matched transcriptomes. At the SNV and mutation-burden tier, we simulated canonical spatial SNV patterns and tested locus-level recoverability with CellSNP, and further summarized spatial variation in local mutation burden as a proxy for spatial mutation-burden proxy heterogeneity under controlled ground truth. At the read/base tier, we inspected BAM-level evidence to verify that simulated alignments exhibit expected signatures of variant allele sampling and CNV-associated coverage shifts, supporting end-to-end benchmarking from reads to calls. Finally, we conducted a fair, tool-appropriate comparison against representative spatial transcriptomics simulators using published, directly available outputs, to clarify differences in capability, output granularity, and benchmarking scope rather than forcing non-equivalent evaluations.
To rigorously evaluate the performance of downstream callers and provide clear context for our metrics, we established a standardized set of input parameters and quality baselines for the simulated spatial multi-omics datasets used in our subsequent experiments. The simulations were conducted on a discrete spatial tissue grid comprising 50 rows by 50 columns. For the spatial SNV modeling, we introduced a targeted mutation on chromosome 12 at position 25 398 284, defined by a reference base cytosine (C) mutating to thymine (T), with the maximum VAF across the spatial field set to 0.8. Concurrently, for the spatial CNV landscapes, we simulated a regional amplification event on chromosome 1, spanning from genomic coordinate 173 864 146 to 173 867 745, and assigned a maximum copy number of 4.7 to the core affected region. To comprehensively assess the robustness of pattern recovery, we systematically tested all three supported spatial mutation architectures: the sharp boundary mode, the continuous gradient mode, and the nested clonal structure mode. Furthermore, within the continuous gradient configurations specifically, we generated and evaluated multiple directional variations—including radial (center-to-periphery), horizontal, and vertical spatial gradients—to ensure that the simulated mutational fields accurately reflect diverse topological layouts and provide rigorously controlled input quality for subsequent benchmarking.
Spatial genomics recapitulates predefined copy number patterns
We first evaluated whether STGBench can faithfully encode user-specified spatial CNV landscapes into DNA-seq outputs that remain recoverable by standard CNV callers. We constructed a ground-truth CNV field on a 50 × 50 spatial grid, where copy number states vary smoothly across the tissue slice with a high-copy core region and gradual transitions toward the periphery (Fig. 2a). STGBench then generated spatially indexed DNA-seq alignments, which were processed by AneuFinder to infer CNV profiles from the simulated data.
Figure 2.
Recovery of predefined spatial CNV landscapes from STGBench-simulated DNA-seq data. (a) User-defined ground-truth spatial copy number field on a 50 × 50 tissue grid. (b) Spatial CNV profile inferred by AneuFinder from STGBench-simulated DNA-seq alignments, shown on the same coordinate system and color scale as in (a). (c) Genome-wide concordance between ground-truth copy numbers and AneuFinder estimates across spatial locations, summarized by a scatter plot with the corresponding Pearson correlation.
As shown in Fig. 2b, the AneuFinder-inferred spatial CNV map closely matches the ground-truth CNV field, capturing both the major high-copy region and the continuous transition zones across the tissue. To quantify fidelity beyond visual agreement, we performed genome-wide concordance analysis between the ground-truth copy numbers and the AneuFinder estimates. The scatter plot in Fig. 2c shows a strong linear relationship, with points concentrated near the identity trend and an overall Pearson correlation of ~0.86 (as indicated in the figure). Together, these results demonstrate that STGBench generates DNA-seq data in which predefined spatial CNV structures are preserved and remain detectable by an off-the-shelf CNV caller.
Spatial transcriptome data accurately reflects underlying genomic copy number landscape
After validating the authenticity of our spatial DNA simulation results, we next assessed the simulator’s performance in generating gene expression data for spatially resolved transcriptomes. A key feature of real tumors is that copy number variations profoundly impact the transcriptome, typically manifesting as broad expression changes across affected genomic regions. Our simulator explicitly models this relationship, enabling DNA-based CNV landscapes to drive corresponding expression changes in RNA data.
To rigorously validate this, we designed two distinct spatial CNV scenarios (Scenario A and Scenario B), each with unique clonal architecture and mutation magnitude. For each scenario, we generated corresponding spatial gene expression matrices and analyzed them using InferCNV.
In Scenario A (nested CNV mode), the heatmap inferred by InferCNV (Fig. 3a-ii) successfully recapitulated the major spatial patterns of the ground-truth map (Fig. 3a-i). Quantitative correlation analysis confirmed strong concordance (Pearson correlation coefficient r = 0.967; Fig. 3a-iii). Notably, unsupervised clustering performed directly on the gene expression matrix (without any CNV information introduced) yielded clustering results (Fig. 3a-iv) that highly matched the spatial structure of underlying CNVs. This result indicates that the CNV-driven expression signal is sufficiently strong to dominate the clustering process.
Figure 3.

Results of copy number variation inference from simulated spatial transcriptome data under two distinct scenarios: Scenario A (Nested CNV mode): (i) ground-truth CNV map; (ii) CNV map inferred by InferCNV from simulated gene expression matrix; (iii) correlation scatter plot: demonstrating correlation between ground-truth CNV values and InferCNV-inferred values; (iv) unsupervised clustering results of spots based on gene expression patterns: showing structure consistent with predefined mutation regions; (v) InferCNV detection plot, confirming consistency of all spot mutation positions with settings. Scenario B (Boundary CNV mode): panels (i)–(v) correspond to those in (a), demonstrating simulator performance under different instability conditions. Correlation between ground-truth and inferred values is higher, and expression-based clustering similarly delineates clonal boundaries clearly.
In Scenario B (boundary CNV mode), the simulator’s performance was even more remarkable. The results inferred by InferCNV (Fig. 3b-ii) visually matched the ground-truth map (Fig. 3b-i) almost perfectly. The correlation coefficient was higher (r = 0.996; Fig. 3b-iii), indicating that more pronounced CNVs produce stronger and more readily detectable transcriptomic signals. Consequently, gene expression-based clustering (Fig. 3b-iv) achieved equally clear separation of major mutations and clearly demarcated spatial boundaries.
Cross-modal consistency validation between spatial genomics and spatial transcriptomics
A core challenge in the field of spatial omics is the integrative analysis of multimodal data (e.g. DNA and RNA) collected from the same tissue slice. A reliable simulator must not only generate accurate single-modality data, but also ensure biological consistency across these modalities to reflect their shared origin from the same spot population. To validate the cross-modal consistency of our simulator, we performed a direct comparison between the CNV map obtained from simulated DNA sequencing data and the CNV map inferred from simulated spatial transcriptomics data.
Visual inspection reveals spatial consistency between the two: both results obtained through direct DNA measurement and those inferred from RNA consistently recapitulate the major mutation structures, although both software tools show insensitivity to boundary detection. The transcriptomic data recapitulated the genomic structure set in the DNA simulation, consistent with known biological principles that large-scale copy number variations have significant and detectable effects on regional gene expression levels.
Subsequently, we performed quantitative analysis of this consistency. The correlation analysis in Fig. 4c shows strong correlation between DNA-based CNVs and gene expression matrix-inferred CNVs (Pearson correlation coefficient r = 0.7695). This indicates that signals in the gene expression matrix can genuinely reflect the underlying copy number variations from genomic DNA sequencing. Notably, the failure to achieve perfect correlation (r = 1.0) is both expected and biologically meaningful: because both CNV detection from DNA sequencing data and inference from gene expression matrices are subject to technical limitations of inference algorithms, noise, gene-specific regulation, and other factors.
Figure 4.
Cross-modal consistency between spatial DNA-derived CNVs and expression-inferred CNVs in STGBench-simulated data. (a) Spatial CNV states detected from STGBench-simulated spatial DNA-seq data using AneuFinder, shown as a categorical map (deletion/normal/amplification) on the tissue grid. (b) CNV states inferred from the matched simulated spatial transcriptomics expression matrix using InferCNV, displayed on the same coordinate system and with the same state definitions as in (a). (c) Confusion matrix comparing CNV state calls between AneuFinder (DNA-based) and InferCNV (RNA-inferred) across spatial locations, with overall concordance summarized by Pearson correlation (r = 0.7695).
Experiments demonstrate that our simulator can generate biologically coherent multi-omics profiles. Strong consistency in both spatial distribution and quantitative results between DNA-determined CNVs and expression matrix-inferred CNVs indicates that different modality data generated by our platform are not isolated outputs but are intrinsically linked, accurately simulating the fundamental relationship between genome and transcriptome. This capability provides our simulator with a unique advantage for developing and validating various computational methods aimed at achieving unified understanding of tissue biological features and spatial mutations through integration of spatial DNA and expression matrix-based data.
Spatial genomics: spatial SNV patterns and local mutation burden (spatial mutation-burden proxy)
Beyond copy number alterations, spatially structured SNVs contribute to fine-scale genetic heterogeneity and shape the spatial mutational landscape. Here, we assess whether the simulator can generate recoverable spatial SNV patterns and whether aggregating SNV signals across loci induces a spatial mutation-burden proxy, which serves as a practical proxy for spatial TMB heterogeneity at spot/region resolution.
We designed a panel of SNVs under four canonical spatial mutation modes (Fig. 5): region-enriched (ring-like), left-to-right gradient, center-to-periphery radial gradient, and top-to-bottom gradient. For each mode, the simulator specifies a ground-truth VAF field over the tissue grid, which governs allele sampling during read generation.
Figure 5.
Spatial SNV mutation patterns and local mutation-burden induced by an SNV panel. For each spatial mutation mode, the simulator specifies a ground-truth VAF field over the tissue grid and generates sequencing reads accordingly. In each pair, the left panel shows the user-defined ground-truth VAF heatmap, and the right panel shows the corresponding CellSNP-derived SNV signal map computed from the simulated data. Four canonical spatial modes are shown: region-enriched (ring-like), left-to-right gradient, center-to-periphery radial gradient, and top-to-bottom gradient. The strong visual concordance between ground-truth VAF fields and CellSNP-derived maps indicates robust recovery of locus-level spatial SNV structure; because each mode is instantiated using a panel of SNVs governed by the same VAF field, these patterns collectively imply a spatially structured local mutation-burden landscape (a practical proxy for spatial TMB heterogeneity).
In Fig. 5, the left subpanel in each pair shows the user-defined ground-truth VAF heatmap, and the right subpanel shows the corresponding CellSNP-derived SNV signal map obtained by processing the simulated sequencing data.
Across all modes, CellSNP recovers spatial SNV patterns that are visually consistent with the ground-truth VAF fields (Fig. 5). In particular, the simulated data preserve (i) sharp spatial boundaries for region-enriched mutations and (ii) smooth VAF transitions for continuous gradient modes. This indicates that the simulator reproduces not only static spatial templates but also the underlying stochastic allele sampling process that determines read-level allele balance and downstream caller behavior.
To connect locus-level SNV recovery to burden-level validation, we note that when multiple SNVs are governed by the same spatial VAF field, the expected local mutation burden at each spatial location is a monotone function of the aggregated mutant evidence across the SNV panel. Concretely, a simple burden proxy can be defined as an aggregate of per-locus SNV signals at each spot, e.g.
![]() |
where
denotes a spatial spot,
is the number of SNV loci in the panel, and
is the VAF estimated by CellSNP for locus
at spot
. Under this design, successful recovery of the spatial VAF fields implies that the simulator generates a spatially structured burden landscape—i.e. a spatial TMB proxy—because spots with higher ground-truth VAF are expected to yield higher aggregated mutant evidence across loci.
Overall, this experiment provides dual validation: (i) locus-level SNV detectability and spatial pattern fidelity using an off-the-shelf tool (CellSNP), and (ii) burden-level spatial heterogeneity induced by aggregating SNV signals across loci, which is essential for benchmarking spatial mutation burden (TMB-oriented) analyses.
Nucleotide-level validation confirms authenticity of simulated sequencing data
While spatial heatmaps and correlation-based analyses assess macroscopic fidelity, a sequencing simulator must ultimately produce read-level evidence that is consistent with expected NGS signatures. To validate nucleotide- and alignment-level realism, we inspected the STGBench-simulated BAM files using the Integrative Genomics Viewer (IGV) and examined whether the simulated variants manifest as canonical patterns in aligned reads (Fig. 6).
Figure 6.
Nucleotide-level validation of simulated sequencing reads using IGV. (a) Representative predefined SNV locus in a STGBench-simulated BAM file. A subset of aligned reads supports the alternative allele at the expected position, demonstrating nucleotide-level visibility of the simulated SNV after alignment. (b) Representative simulated deletion region. IGV shows an abrupt reduction of read depth within the deleted segment compared with flanking regions, consistent with the expected coverage signature of copy number loss.
At a representative predefined SNV locus, IGV visualization shows that a subset of aligned reads carries the alternative allele at the expected genomic position, while the remaining reads support the reference base (Fig. 6a). This pattern is consistent with allele sampling under a non-unity VAF setting and indicates that the simulated SNV signal is preserved through alignment and can be directly observed at nucleotide resolution.
For a representative copy number deletion region, IGV displays a clear and localized reduction of read depth within the deleted segment, with comparatively higher and more uniform coverage in flanking diploid regions (Fig. 6b). The abrupt coverage shift confined to the deletion boundaries is a characteristic read-level signature of copy number loss in NGS data, supporting that the simulator translates CNV dosage changes into realistic coverage patterns in the generated alignments.
Together, these IGV inspections provide direct nucleotide-level evidence that STGBench produces pipeline-ready sequencing reads with variant signatures that are observable in BAM files and compatible with standard downstream workflows (alignment visualization, SNV inspection, and coverage-based CNV assessment). This read-level authenticity complements the matrix- and region-level validations and supports the use of STGBench for end-to-end benchmarking of spatial genomics pipelines.
Comparison with existing simulators
To contextualize the contribution of STGBench relative to existing simulators, we compared STGBench with two representative and widely used reference-based simulation frameworks for spatial transcriptomics: SRTsim and scDesign3.
Fair comparison statement. Because these tools were designed for different primary tasks and output modalities, we adopt a fair comparison principle: we restrict comparisons to evaluation dimensions that are (i) explicitly defined in the corresponding original publications and (ii) computable from the outputs each tool is described to generate, without assuming unreported functionality or extrapolating missing results. Under this principle, we focus on (i) spatial pattern preservation, (ii) gene–gene correlation structure, (iii) reported computational resource usage, and (iv) end-to-end output modality and pipeline compatibility. The overall differences in output modalities, benchmarking scope, and tool-reported evaluation targets are summarized in Fig. 7.
Figure 7.
Functional scope and output granularity of STGBench versus representative spatial transcriptomics simulators and established read-level variant simulators. (a) Capability matrix comparing supported features, output modalities, and primary benchmarking targets across STGBench, SRTsim, scDesign3, BAMSurgeon, and ART. Symbols indicate native/core support (✓) and not supported (—). (b) Conceptual benchmarking workflows enabled by each simulator class. Expression-based spatial simulators generate synthetic expression count matrices and coordinates from reference data to perform expression evaluations (statistics and correlation) and downstream expression tasks (clustering, deconvolution). In contrast, STGBench (end-to-end spatial mutation simulation) utilizes a tissue scaffold/spatial grid and ground-truth spatial CNV/SNV fields to generate spatial DNA-seq reads (aligned BAM) and a matched spatial expression matrix for evaluations of genomic alterations. Meanwhile, established read-level variant simulators focus on modifying existing real BAM files or reference genomes through variant/error spike-in to generate variant-containing synthetic BAMs for evaluating variant call accuracy.
Output targets and benchmarking scope
SRTsim is an SRT-specific simulator primarily aimed at generating spatial expression count matrices while preserving spatial expression patterns, with a key mechanism that assigns simulated expression counts to locations to mimic the reference spatial ranking of expression.
scDesign3 is a unified probabilistic simulator for single-cell and spatial omics count-level data that explicitly models joint feature distributions via copula-based approaches and evaluates resemblance using summary statistics, including spatial-pattern similarity and feature–feature correlation structure.
In contrast, STGBench is designed to benchmark spatial mutation analysis workflows by generating paired spatial DNA-seq reads and a matched spatial gene expression matrix from a shared latent CNV/SNV landscape, with outputs intended to be ingested by standard downstream pipelines.
Spatial pattern preservation and spatial autocorrelation
SRTsim explicitly quantified spatial pattern preservation using Moran’s I: it reports that Moran’s I statistics computed in SRTsim-generated synthetic data are highly consistent with those in reference data and are more consistent than several compared simulators.
scDesign3 evaluates spatial pattern resemblance by comparing spatial relationships learned from real versus synthetic data; specifically, it defines a per-feature spatial-pattern similarity metric using the Pearson correlation between predictions from learners trained on real and synthetic data.
While these metrics target spatial expression realism, STGBench’s primary validation target is different: it tests whether simulated spatial multi-omics data can be used to recover known ground-truth genomic perturbations with standard analysis tools. In our experiments, spatial CNV patterns embedded in simulated DNA-seq reads were recovered by AneuFinder with high genome-wide concordance (Pearson r = 0.855), and CNV-driven expression signals in simulated transcriptomes were recovered by InferCNV with r up to 0.996.
Together, these results position STGBench as a simulator emphasizing mutation-ground-truth-driven benchmarking (CNV/SNV/spatial mutation-burden proxy), whereas SRTsim and scDesign3 emphasize expression-pattern realism for spatial transcriptomics.
Niche advantages and spatial autocorrelation preservation
Although STGBench provides a systematic framework for mutation-centric benchmarking, based on the principle of fair comparison, it must be objectively acknowledged that existing spatial transcriptomics simulators still possess significant performance advantages in certain specific niches. For example, SRTsim is specifically designed to preserve spatial expression patterns from reference spatial transcriptomics data, performing highly realistically in maintaining spatial autocorrelation (such as Moran’s I statistics). When the core objective of benchmarking is to evaluate downstream algorithms that are highly sensitive to subtle spatial expression gradients, SRTsim’s gene expression simulation mechanism, which relies on spatial ranking, can generate more precise local spatial count features than STGBench’s CNV-driven expression model. Clarifying this performance boundary helps strengthen the objective positioning of STGBench: it is not intended to replace existing simulators that dominate in spatial expression fidelity, but rather to complement the existing tool ecosystem by filling the benchmarking gap for mutation analysis tasks.
Gene–gene correlation structure
scDesign3 explicitly targets preservation of feature dependencies through copula models, noting that Gaussian copula is fast but yields only a gene correlation matrix, whereas vine copula is slower but interpretable.
For evaluation, scDesign3 defines similarity of real versus synthetic correlation matrices as the Pearson correlation across all off-diagonal entries in the upper triangle of the feature–feature correlation matrices (computed on highly expressed features).
STGBench’s RNA layer is driven by CNV-to-expression coupling under a negative binomial model with spatially correlated technical noise, and its validation emphasizes whether CNV structure is preserved strongly enough to be recovered by CNV-from-RNA tools and unsupervised clustering.
Accordingly, scDesign3’s feature-correlation-centric evaluation and STGBench’s CNV-driven expression consistency evaluation are complementary rather than redundant.
End-to-end outputs and pipeline compatibility
A key differentiator of STGBench is that it produces spatially indexed DNA-seq reads in BAM format and a matched gene expression matrix under the same ground-truth CNV/SNV fields, enabling end-to-end benchmarking from alignment to variant calling, CNV inference, and integrative DNA–RNA consistency checks.
This capability is central for benchmarking spatial mutation pipelines where failure modes may arise from alignment artifacts, allelic balance, coverage shifts, and downstream caller heuristics—phenomena that cannot be evaluated using count matrices alone. Consistent with this goal, we further validated nucleotide-level realism by IGV inspection and demonstrated recovery of diverse spatial SNV patterns using CellSNP.
By comparison, SRTsim’s described generative procedure centers on simulating and assigning expression counts to preserve spatial patterns, without describing alignment-level DNA read outputs.
scDesign3 similarly formalizes generation of synthetic cell-by-feature matrices for single-cell and spatial omics; while it notes that certain modalities can be coupled to external tools (e.g. scReadSim) for read-level signal representations in specific contexts, its core simulator output is count-level synthetic data.
Comparison with established read-level variant simulation frameworks
The current comparison primarily focuses on count-level simulators such as SRTsim and scDesign3, which is appropriate for validating expression realism. However, to further support STGBench’s novelty claims regarding the generation of DNA BAM files and mutation benchmarking, it is necessary to compare it against established read-level variant simulation frameworks. For instance, tools like BAMSurgeon can precisely spike-in SNVs, insertions/deletions (INDELs), and structural variants into real BAM files, providing rigorous testing standards for somatic mutation calling algorithms in nonspatial contexts. Similarly, simulators like ART (a sequencing read simulator) generate synthetic reads with high sequencing-characteristic fidelity through empirical error models. Although these established tools excel in read-level variant simulation, they generally lack a spatial dimension and cannot generate paired RNA data. In contrast, STGBench advances the paradigm of variant simulation by introducing a spatial tissue scaffold and synchronously generating spatially indexed paired DNA-seq reads (BAM format) alongside matched transcriptome expression matrices based on the same explicit CNV/SNV ground truth. This end-to-end mapping from the underlying mutation field to multi-omics reads is unachievable by conventional nonspatial read-level simulators, thereby providing a dedicated benchmarking platform for spatial mutation analysis workflows that jointly exploit DNA and RNA.
Overall, these comparisons support that STGBench fills a distinct benchmarking gap: it links spatial tissue architecture to explicit CNV/SNV ground truth and produces pipeline-ready DNA-seq BAM plus matched transcriptomes, enabling evaluation of spatial mutation analysis workflows that jointly exploit DNA and RNA. In particular, the capability gap of pipeline-ready DNA BAM output is highlighted in Fig. 7a, and the corresponding benchmarking workflow difference is illustrated in Fig. 7b.
Benchmark of CNV detection robustness under realistic sequencing scenarios
To address the concern that single best-case correlation metrics may overestimate the benchmark utility of the simulator, and to verify that STGBench can generate biologically realistic CNV signals rather than only unusually clean data aligned to tool assumptions, we performed a systematic set of gradient experiments to quantify CNV detection performance across common confounding scenarios in spatial transcriptomics research. We used the Pearson correlation coefficient (r) between the ground-truth copy number values predefined by STGBench and the CNV values inferred by InferCNV as the core evaluation metric, where a value closer to 1 indicates higher concordance between the inferred result and the ground truth, representing better CNV detection accuracy.
We designed three independent experimental scenarios to simulate the most prevalent technical and biological confounders in spatial transcriptomics sequencing, with four gradient levels set for each scenario to capture the dose-dependent effect of each factor. First, the coverage experiment was designed to simulate signal loss caused by decreased sequencing depth: we performed binomial distribution-based sampling to retain original gene counts at a set probability, with four coverage gradients set at 100%, 80%, 60%, and 40%. Second, the tumor purity experiment simulated the common clinical scenario of normal cell contamination in tumor samples: we mixed the expression profile of tumor cells with the average expression profile of normal diploid cells according to the preset tumor purity ratio, and generated the final observed expression matrix via Poisson distribution, with four purity gradients set at 100%, 80%, 60%, and 40%. Third, the noise experiment simulated technical variation and biological noise in spatial transcriptomics data: we introduced overdispersed noise using a negative binomial distribution consistent with the core implementation of STGBench, with four noise levels including 0 noise (no additional noise), low noise (dispersion = 1.0), medium noise (dispersion = 0.5), and high noise (dispersion = 0.2). To further clarify the impact of CNV amplitude (the distance between the altered copy number and the normal diploid baseline of 2) on detection robustness, we set four parallel CNV test groups across all experiments: high-level amplification (copy number = 6, distance from baseline Δ = 4), low-level amplification (copy number = 3, distance from baseline Δ = 1), homozygous deletion (copy number = 0, distance from baseline Δ = 2), and heterozygous deletion (copy number = 1, distance from baseline Δ = 1).
The results of all experiments are summarized in Fig. 8. In the coverage experiment, the CNV detection correlation coefficient of all test groups showed a continuous upward trend as sequencing coverage increased from 40% to 100%, with all groups reaching near-optimal performance with r values close to 1.0 at 100% coverage. CNV events with larger distance from the diploid baseline exhibited significantly stronger robustness to reduced coverage: even at 40% low coverage, the homozygous deletion group and high-level amplification group maintained high correlation coefficients of 0.97 and 0.92, respectively. In contrast, the detection performance of CNV events closer to the normal baseline was more severely affected by coverage reduction, with the correlation coefficient of the low-level amplification group dropping to 0.70 and the heterozygous deletion group dropping to 0.82 at 40% coverage. In the tumor purity experiment, the detection performance of all test groups improved continuously with increasing tumor purity, with all groups achieving r values above 0.98 in 100% pure tumor samples. Under the challenging scenario of 40% low purity, high-amplitude CNV events still maintained measurable detection performance, with the high-level amplification group reaching an r value of 0.93 and the homozygous deletion group reaching 0.51, while the performance of low-amplitude CNV events decreased sharply: the correlation coefficient of the heterozygous deletion group dropped to 0.21, and the low-level amplification group dropped to 0.79. This pattern is fully consistent with real-world clinical observations, where normal cell contamination dilutes tumor-derived CNV signals, with a more pronounced effect on variants with smaller signal amplitude. In the noise experiment, the detection performance of all test groups showed a gradient downward trend with increasing noise level, with all groups achieving r values close to 1.0 under noise-free conditions. Even under high noise conditions, the homozygous deletion group and high-level amplification group remained highly accurate with r values of 0.99 and 0.96, respectively, while low-amplitude CNV events were more significantly affected by noise interference, with the correlation coefficient of the low-level amplification group dropping to 0.75 and the heterozygous deletion group dropping to 0.77 under high noise.
Figure 8.
CNV detection performance across simulated sequencing and biological confounding scenarios. (a) Pearson correlation between ground-truth copy numbers and InferCNV-inferred values across sequencing coverage gradients ranging from 40% to 100%. (b) Pearson correlation between ground-truth copy numbers and InferCNV-inferred values across tumor purity gradients ranging from 40% to 100%. (c) Pearson correlation between ground-truth copy numbers and InferCNV-inferred values across increasing levels of technical noise, including 0 noise, low noise, medium noise, and high noise. For all panels, the y-axis represents the Pearson correlation coefficient (r), the core metric quantifying the concordance between predefined ground truth and inferred CNV results, with higher values indicating better detection accuracy. Four CNV test groups are included in all experiments: high-level amplification (copy number = 6, blue line), low-level amplification (copy number = 3, orange line), homozygous deletion (copy number = 0, green line), and heterozygous deletion (copy number = 1, red line).
Across all three experimental scenarios, the distance between the CNV event and the normal diploid baseline was identified as the core determinant of detection robustness: under all confounding conditions, high-amplitude CNV events consistently showed significantly higher detection accuracy and stronger resistance to technical and biological interference than low-amplitude CNV events. More importantly, these results confirm that STGBench can realistically simulate the impact of reduced sequencing coverage, decreased tumor purity, and elevated technical noise on CNV signals, rather than only generating overly idealized clean data. This set of gradient experiments not only strengthens the credibility and benchmark utility of STGBench, but also provides a standardized reference framework for the development, optimization, and performance evaluation of CNV detection methods for spatial transcriptomics data.
Discussion
Clarification of core metric definition, functional positioning, and applicable boundaries of STGBench
To clarify the functional positioning, core metric definition, and applicable boundaries of our STGBench software, we hereby explicitly elaborate the design rationale of our core output indicator, its essential differences from clinically standardized TMB, and the clear applicable scope and limitations of this tool.
First, we give a precise and unambiguous definition of the core quantitative output of STGBench: the spatial mutation-burden proxy is a relative quantitative indicator derived from a prespecified fixed panel of SNV loci, calculated by aggregating the VAF of target SNVs across spatial spots or contiguous tissue regions. The core design rationale of this proxy is to characterize the spatial heterogeneity and continuous distribution patterns of somatic mutations in tissue sections to support the development and benchmark verification of spatial genomics and transcriptomics analysis methods focusing on the spatial landscape of tumor mutations. It is not designed for clinical TMB quantification, nor is it equivalent to the clinically standardized TMB metric recognized in oncology practice.
Clinically, TMB is strictly and uniformly defined as the number of somatic non-synonymous mutations per megabase (mut/Mb) of the callable genomic territory. Its standardized calculation relies on three core premises: clear definition of callable genome regions (such as whole-exome sequencing capture regions or targeted panel BED files), strict somatic mutation filtering pipelines to exclude germline variants, sequencing artifacts and low-confidence variants, and unified statistical calibration of eligible mutation counts normalized by genome length. In contrast, our spatial mutation-burden proxy is a pattern-focused indicator based on a user-defined fixed SNV panel, which does not involve the calculation of mutation density per unit genome length, nor does it integrate the standardized filtering and calibration processes required for clinical TMB quantification. This essential difference determines that the two indicators have distinct application scenarios and cannot be used interchangeably, and we have avoided any over-claiming of clinical TMB equivalence throughout the revised manuscript.
Based on the above definition, we explicitly clarify the applicable boundaries of the STGBench software. This tool is suitable for the following scenarios: (i) simulation of spatial genomic and transcriptomic joint data with heterogeneous mutation distribution patterns; (ii) development and benchmark verification of spatial multi-omics analysis methods focusing on the spatial heterogeneity of tumor somatic mutations; (iii) reconstruction and validation of tumor subclonal spatial structure inference algorithms. Meanwhile, we clearly state that the current version of STGBench is not suitable for: (i) performance evaluation of clinical TMB detection pipelines for patient-derived samples; (ii) verification of clinical TMB threshold and prognostic predictive value; (iii) benchmarking of companion diagnostic-related TMB quantification methods.
We also acknowledge the corresponding limitations of the current version of the tool. Consistent with its core design goal of spatial mutation pattern simulation, the current version of STGBench does not yet support the functional modules required for clinical TMB benchmarking as suggested by the reviewer, including the specification of callable genomic territory via BED files, simulation of per-megabase normalized mutation counts with explicit ground-truth labels, and generation of paired ground-truth and inferred TMB maps for accuracy evaluation across different sequencing coverage and tumor purity regimes. These limitations will be addressed in the subsequent iterative updates of the STGBench software to expand its applicability while maintaining its core advantages in high-fidelity spatial multi-omics data simulation.
Conclusion
STGBench addresses a central bottleneck in spatial mutation analysis: the lack of benchmark datasets that are simultaneously spatially structured, multimodal, and equipped with complete ground truth. We present STGBench as a sequencing-level spatial DNA–RNA simulator that models tissue morphology, spatially organized CNV and SNV/VAF landscapes, and their propagation into paired DNA-seq and transcriptomic readouts within a single probabilistic framework. By explicitly separating a spatial scaffold (tissue mask), latent mutational fields, and observation models that generate pipeline-ready outputs, STGBench provides a controllable and interpretable substrate for stress-testing computational methods under conditions that emulate key features of real spatial genomics experiments while preserving full access to truth labels at cellular and locus resolution.
Across multiple tiers of evaluation, STGBench generates data that are both biologically coherent and analytically challenging. At the DNA level, standard CNV callers recover user-defined spatial copy number architectures—including sharp boundaries, continuous gradients, and nested clonal structures—from simulated sequencing data, demonstrating that the simulator reproduces the signal structures exploited by coverage-based CNV inference. At the RNA level, copy number-driven expression consequences are sufficiently strong and spatially consistent to support CNV inference and to yield unsupervised structure aligned with the underlying clonal design. Importantly, STGBench enables explicit cross-modal validation: CNV landscapes derived from DNA and inferred from RNA show substantial concordance, reflecting realistic genome–transcriptome coupling while retaining method-relevant discrepancies introduced by regulatory variability and inference noise. Beyond CNVs, diverse spatial SNV patterns are recovered reliably by standard SNV workflows, and nucleotide-level inspection confirms that simulated reads encode expected allele balances and CNV-associated coverage shifts, supporting end-to-end benchmarking from alignments through variant calls.
STGBench differs from existing simulators in three defining ways. First, it treats space as a primary object rather than a post hoc annotation, allowing users to specify tissue geometry from parametric templates or image-derived masks and to impose interpretable mutational fields directly on that scaffold. Second, it is explicitly multimodal: DNA and RNA outputs are generated from a shared latent mutational state, enabling principled evaluation of workflows that integrate spatial DNA and transcriptomes for mutation burden analysis, clonal reconstruction, and cross-modal consistency. Third—and most importantly for practical benchmarking—STGBench is sequencing-level. By producing alignment-compatible DNA outputs (including BAM files), it allows benchmarking of complete pipelines that include read-level artifacts, allele sampling, coverage distortions, and caller-specific heuristics. This capability enables fairer and more actionable assessment than count-matrix–only simulations when the research question concerns mutation calling, CNV inference, and spatial mutation-burden proxy-oriented workflows.
STGBench is intentionally scoped to remain tractable and parameterizable rather than to model every dimension of tumor biology or every platform-specific artifact. The current implementation focuses on 2D tissue sections, a core set of spatial mutation templates, and a gene-expression model that captures overdispersed counts and spatially correlated technical effects. Future extensions can systematically increase realism where it matters most for sequencing-level benchmarks, including improved modeling of coverage and GC biases, mapping ambiguity and mappability constraints, base-quality and error profiles, duplicate formation, and library-preparation effects. Methodologically, richer spatial mutation templates inspired by emerging datasets, additional omics layers (e.g. chromatin accessibility, proteomics), and imaging-derived covariates would broaden the scope of multimodal benchmarks while preserving explicit ground truth.
Despite these planned extensions, STGBench provides immediate value as an open, configurable benchmark generator for spatial mutation analysis. For method developers, it enables reproducible stress tests, controlled ablation studies, and failure-mode discovery under tunable perturbations of tissue architecture, sequencing depth, noise structure, and clonal complexity. For practitioners, it offers a practical sandbox to prototype end-to-end workflows, calibrate thresholds, and understand the behavior of CNV, SNV, and mutation-burden analyses before deployment on limited clinical material. More broadly, as the field moves toward standardized multimodal benchmarks—including efforts aligned with virtual cell-oriented modeling ecosystems—STGBench provides a sequencing-level testbed that supports transparent, pipeline-aware evaluation of mutation-centric components in spatial precision oncology.
Key Points
STGBench provides a sequencing-level spatial DNA–RNA simulation framework that generates alignment-ready BAM files together with matched gene expression matrices, enabling end-to-end benchmarking from reads to calls under explicit ground truth.
By modeling tissue morphology and spatially structured CNV and SNV/VAF fields on a shared spatial scaffold, STGBench supports controllable clonal architectures (boundary, gradient, and nested patterns) that emulate common tumor spatial topologies.
STGBench is explicitly multimodal, jointly simulating DNA-derived CNVs and CNV-driven transcriptional consequences to quantify cross-modal DNA–RNA concordance, and thereby providing a benchmark substrate aligned with virtual cell-oriented evaluation settings that require paired modalities and reproducible ground truth.
The simulator produces pipeline-compatible outputs that can be directly analyzed with standard tools (e.g. AneuFinder, InferCNV, CellSNP), allowing fair stress-testing of real-world analysis pipelines rather than matrix-only surrogate evaluations.
Through tiered validation at CNV, SNV/mutation-burden, and read-level signatures, STGBench establishes a reproducible benchmark substrate for assessing robustness, failure modes, and performance trade-offs in spatial mutation and spatial mutation-burden proxy-related methods.
Acknowledgements
We thank all of the faculty members and graduate students who discussed the mathematical and statistical issues in seminars.
Contributor Information
Shenjie Wang, Department of Respiratory Medicine, The Second Affiliated Hospital of Xi'an Jiaotong University, No. 157, Xiwu Road, Xincheng District, Xi'an 710049, China; School of Computer Science and Technology, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China; Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China.
Yuhang Li, School of Computer Science and Technology, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China; Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China.
Xiaonan Wang, Nanjing Geneseeq Technology Inc., 128 Huakang Road, Pukou, Nanjing 211800, China.
Xuwen Wang, Department of Respiratory Medicine, The Second Affiliated Hospital of Xi'an Jiaotong University, No. 157, Xiwu Road, Xincheng District, Xi'an 710049, China; School of Computer Science and Technology, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China; Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China.
Tianci Wang, Department of Respiratory Medicine, The Second Affiliated Hospital of Xi'an Jiaotong University, No. 157, Xiwu Road, Xincheng District, Xi'an 710049, China; School of Computer Science and Technology, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China; Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China.
Shuanying Yang, Department of Respiratory Medicine, The Second Affiliated Hospital of Xi'an Jiaotong University, No. 157, Xiwu Road, Xincheng District, Xi'an 710049, China.
Jiayin Wang, Department of Respiratory Medicine, The Second Affiliated Hospital of Xi'an Jiaotong University, No. 157, Xiwu Road, Xincheng District, Xi'an 710049, China; School of Computer Science and Technology, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China; Shaanxi Engineering Research Center of Medical and Health Big Data, Xi'an Jiaotong University, 28 Xianning West Road, Beilin, Xi'an 710049, China.
Author contributions
J.W.: Conceptualization. S.W. and Y.L.: Methodology. Y.L.: Software. T.W. and X.N.W.: Investigation, Validation, and Formal analysis. S.W., S.Y., Y.L., and X.W.W.: Writing—original draft; Writing—review and editing. All authors have read and agreed to the published version of the manuscript.
Conflicts of interest
The authors declare that they have no conflicts of interest.
Funding
This work was funded by the National Natural Science Foundation of China, grant numbers T2541083,62572389,62402376. This work was also supported by China Postdoctoral Science Foundation, grant number 2025M771554.
Data availability
STGBench is an open-source package. Source code, documentation, and example configurations are available at GitHub: https://github.com/Icarus200110/STGBench.
References
- 1. Seferbekova Z, Lomakin A, Yates LR et al. Spatial biology of cancer evolution. Nat Rev Genet 2023;24:295–313. 10.1038/s41576-022-00553-x [DOI] [PubMed] [Google Scholar]
- 2. Walker C, Angelo M. Toward clinical applications of spatial-omics in cancer research. Nat Cancer 2024;5:1771–3. 10.1038/s43018-024-00868-0 [DOI] [PubMed] [Google Scholar]
- 3. Gong D, Arbesfeld-Qiu JM, Perrault E et al. Spatial oncology: translating contextual biology to the clinic. Cancer Cell 2024;42:1653–75. 10.1016/j.ccell.2024.09.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Rodriques SG, Stickels RR, Goeva A et al. Slide-seq: a scalable technology for measuring genome-wide expression at high spatial resolution. Science 2019;363:1463–7. 10.1126/science.aaw1219 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Vickovic S, Eraslan G, Salmén F et al. High-definition spatial transcriptomics for in situ tissue profiling. Nat Methods 2019;16:987–90. 10.1038/s41592-019-0548-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Stickels RR, Murray E, Kumar P et al. Highly sensitive spatial transcriptomics at near-cellular resolution with slide-seqV2. Nat Biotechnol 2020;39:313–9. 10.1038/s41587-020-0739-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Chen A, Liao S, Cheng M et al. Spatiotemporal transcriptomic atlas of mouse organogenesis using DNA nanoball-patterned arrays. Cell 2022;185:1777–1792.e21. 10.1016/j.cell.2022.04.003 [DOI] [PubMed] [Google Scholar]
- 8. Liu Y, Yang M, Deng Y et al. High-spatial-resolution multi-omics sequencing via deterministic barcoding in tissue. Cell 2020;183:1665–1681.e18. 10.1016/j.cell.2020.10.026 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Zhao T, Chiang ZD, Morriss JW et al. Spatial genomics enables multi-modal study of clonal heterogeneity in tissues. Nature 2022;601:85–91. 10.1038/s41586-021-04217-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Lomakin A, Svedlund J, Strell C et al. Spatial genomics maps the structure, nature and evolution of cancer clones. Nature 2022;611:594–602. 10.1038/s41586-022-05425-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11. Greenwald AC, Darnell NG, Hoefflin R et al. Integrative spatial analysis reveals a multi-layered organization of glioblastoma. Cell 2024;187:2485–2501.e26. 10.1016/j.cell.2024.03.029 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Zhou W, Su M, Jiang T et al. SORC: an integrated spatial omics resource in cancer. Nucleic Acids Res 2024;52:D1429–37. 10.1093/nar/gkad820 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Bakker B, Taudt A, Belderbos ME et al. Single-cell sequencing reveals karyotype heterogeneity in murine and human malignancies. Genome Biol 2016;17:115. 10.1186/s13059-016-0971-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Gao R, Bai S, Henderson YC et al. Delineating copy number and clonal substructure in human tumors from single-cell transcriptomes. Nat Biotechnol 2021;39:599–608. 10.1038/s41587-020-00795-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15. Broad Institute . InferCNV: Inferring CNV from Single-Cell RNA-Seq. San Francisco, CA, USA: GitHub, Inc., n.d. https://github.com/broadinstitute/infercnv
- 16. Song M, Ma S, Wang G et al. Benchmarking copy number aberrations inference tools using single-cell multi-omics datasets. Brief Bioinform 2025;26:bbaf076. 10.1093/bib/bbaf076 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Ma C, Balaban M, Liu J et al. Inferring allele-specific copy number aberrations and tumor phylogeography from spatially resolved transcriptomics. Nat Methods 2024;21:2239–47. 10.1038/s41592-024-02438-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Shafighi S, Geras A, Jurzysta B et al. Integrative spatial and genomic analysis of tumor heterogeneity with Tumoroscope. Nat Commun 2024;15:9343. 10.1038/s41467-024-53374-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Kazdal D, Endris V, Allgäuer M et al. Spatial and temporal heterogeneity of panel-based tumor mutational burden in pulmonary adenocarcinoma: separating biology from technical artifacts. J Thorac Oncol 2019;14:1935–47. 10.1016/j.jtho.2019.07.006 [DOI] [PubMed] [Google Scholar]
- 20. Xu H, Clemenceau JR, Park S et al. Spatial heterogeneity and organization of tumor mutation burden with immune infiltrates within tumors based on whole slide images correlated with patient survival in bladder cancer. J Pathol Inform 2022;13:100105. 10.1016/j.jpi.2022.100105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21. Chen J, Larsson L, Swarbrick A et al. Spatial landscapes of cancers: insights and opportunities. Nat Rev Clin Oncol 2024;21:660–74. 10.1038/s41571-024-00926-7 [DOI] [PubMed] [Google Scholar]
- 22. Zhu J, Shang L, Zhou X. SRTsim: spatial pattern preserving simulations for spatially resolved transcriptomics. Genome Biol 2023;24:39. 10.1186/s13059-023-02879-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Song D, Wang Q, Yan G et al. scDesign3 generates realistic in silico data for multimodal single-cell and spatial omics. Nat Biotechnol 2024;42:247–52. 10.1038/s41587-023-01772-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24. Huang X, Huang Y. Cellsnp-lite: an efficient tool for genotyping single cells. Bioinformatics 2021;37:4569–71. 10.1093/bioinformatics/btab358 [DOI] [PubMed] [Google Scholar]
- 25. Thorvaldsdóttir H, Robinson JT, Mesirov JP. Integrative genomics viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform 2013;14:178–92. 10.1093/bib/bbs017 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
STGBench is an open-source package. Source code, documentation, and example configurations are available at GitHub: https://github.com/Icarus200110/STGBench.






















