Abstract
Pathological events often impact tissue regions in a spatial variable manner, making it challenging to identify therapeutic targets. Spatial transcriptomics (ST) is a powerful technology to map spatially variable molecular mechanisms, yet suitable analytical methods have been lacking. We introduce SPaSE (Spatially-resolved Pathology ScorE), an optimal transport-based algorithm to compare ST data from diseased and control tissues. SPaSE computes a “pathology score” for each spot in the diseased sample, quantifying the pathological impact at that spot. In post-MI (myocardial infarction) mouse hearts, these scores delineated zones that matched independent expert annotations. Modeling pathology scores from gene expression revealed signatures predictive of varying pathological severity. The scoring model learned from mouse data showed accurate predictions on human post-MI data. We also demonstrated SPaSE’s efficacy on additional simulated and real ST data from traumatic brain injury and Duchenne muscular dystrophy mouse models. SPaSE is a useful addition to the existing ST algorithms. A record of this paper’s Transparent Peer Review process is included in the Supplemental Information.
eTOC Blurb
SPaSE, an optimal transport-based tool, computes spatially-resolved pathology scores to quantify pathological impact in spatial transcriptomics (ST) data. Modeling and analysis of these pathology scores identified key genes, cell types, and biological processes activated in pathological regions. Pathology scoring learned from mouse ST generalized well to human data.
Introduction
Pathological events, such as diseases and injuries, often impact different tissue regions by different extents. Thus, restoring tissue physiology relies on resolving two challenges: (a) quantifying the extent of the pathological effects across the affected tissue and (b) understanding the mechanisms behind these spatial variations in pathological impact. While histology image analysis can provide some help, it leaves the spatially variable mechanisms mostly undefined. Spatial transcriptomics (ST) now offers an opportunity to address these challenges, but the appropriate methods are still lacking.
Recent tools have shown major advances in demarcating spatial “domains” in ST data, i.e., regions where gene expression values remain consistent1–15. In some analyses, the domains coincided with diseased tissue regions, such as tumors or plaque- dense regions in Alzheimer’s brain1,8. Unfortunately, these methods cannot be used to rank spatial domains based on the severity of pathological effects or to resolve spatial variation of pathological effects within a given domain. To address this gap, we propose SPaSE (Spatially-resolved Pathology ScorE), an optimal transport (OT) based algorithm to compare ST samples from diseased and control tissues and assign real-valued “pathology scores” to each location of the diseased sample. Using a null distribution derived from the control sample, SPaSE also determines p-values of the pathology scores. As we outline below, downstream modeling and analyses of pathology scores can further reveal the transcriptional mechanisms underlying spatial variation of pathological effects. Overall, we demonstrate that SPaSE can quantify pathological effects throughout an injured tissue and define the mechanisms underlying spatial variation of these pathological effects.
In simulated data, SPaSE correctly assigned higher pathology scores to ground truth pathological spots, suggesting its applicability to real datasets. We applied SPaSE to ST data from post-MI (myocardial infarction) mouse hearts collected at four time points spanning immediate inflammation to fibrosis development. This dataset, generated using Visium ST technology, consists of spatially resolved circular spots. SPaSE’s pathology scores were as effective as orthogonal single-nucleus data and prior biological knowledge in delineating zonation (infarct, border, and remote regions) in post-MI hearts. Further analysis of these pathology scores revealed key biological processes and cell types driving spatial variations.
To link gene expression with pathology, we trained support vector regression (SVR) models to predict pathology scores in mouse post-MI samples from gene expression data. These models accurately applied to human heart samples (27 samples from 23 individuals), confirming the known relationship between pathology, cell types, and pathways involved in human MI. We tested SPaSE on two additional ST datasets from Duchenne muscular dystrophy (DMD) and traumatic brain injury mouse models. In both cases, spots in pathological regions showed higher pathology scores, confirming SPaSE’s broad applicability. Overall, our results highlight SPaSE’s efficacy in quantifying pathological impacts and suggest it can reduce the need for additional data generation and manual intervention in modeling workflows.
Results
SPaSE: an OT-based tool to delineate pathologically impacted regions in diseased or injured tissues
SPaSE assigns a “pathology score” to each spot of the ST data of a diseased or injured tissue. A spot’s pathology score quantifies how severely the disease or injury has impacted transcriptional expression at that region. SPaSE also computes a null distribution of pathology scores, which it uses to demarcate the regions with significantly high pathology scores (Figure 1a). SPaSE accomplishes these goals using optimal transport (OT) (Figure 1b), a classical method for quantifying the difference between two probability distributions. In doing so, OT establishes a minimum-cost transportation plan, , for moving masses from a source distribution to a target distribution. SPaSE considers the control (healthy) and diseased ST datasets as probability distributions by assigning a “uniform mass” to each spot (Methods). The entry of denotes the amount of mass transferred from the -th spot of the healthy sample to the -th spot of the diseased sample. The plan is obtained by maximizing: (a) for spot pairs and that have high transcriptional similarity and (b) and when the distances between spot pairs , and , are similar (Equation 1). The pathology score of a spot s in the diseased tissue is a weighted summation of the transcriptional dissimilarities between and the spots of the healthy tissue that mapped to (Equation 2). To determine if the value at spot is significantly high, SPaSE derives an empirical null distribution of values using the healthy sample (Methods).
Figure 1.
Schematic diagram of SPaSE. a. Overview of how SPaSE identifies pathologically remodeled regions of a diseased ST sample with respect to a healthy ST sample by calculating pathology score for each of the spots in . To determine which spots are to be labeled as pathologically remodeled, SPaSE employs a threshold selection method. This threshold is obtained from a null distribution of pathology score values generated using two synthetic healthy samples, and . and are synthesized from the original healthy sample . b. SPaSE employs an FGW optimal transport plan with entropic regularization, integrating a feature dissimilarity measure based on the JSD. This method is used to establish a mapping from the spots of to the spots of . Subsequently, for each spot within , SPaSE computes the pathology score as follows. First, it normalizes the transport plan column-wise (represented as ). In , each cell quantifies the proportion of mass from spot in that was transported to spot in . Next, SPaSE performs element-wise multiplication with the pairwise JSD values derived from the gene expressions between spots in and . Finally, the resulting matrix is summed column-wise to calculate the pathology score for each spot in .
SPaSE’s use of an OT formulation is inspired by a recent OT-based tool, PASTE16, which aligns and integrates ST data collected from adjacent tissue slices. However, SPaSE goes beyond alignment, and complements current tools with a rigorous approach to compare ST datasets from two contrasting conditions. In particular, SPaSE incorporates entropic regularization with the fused Gromov-Wasserstein (FGW) optimal transport plan. This transforms the problem into one that is -strongly convex, differentiable, and less susceptible to the curse of dimensionality17–19. Secondly, SPaSE features a hyperparameter selection technique that introduces perturbations to specific regions of the diseased ST sample, allowing the hyperparameter selection strategy to be sample-specific. Third, SPaSE utilizes the Jensen-Shannon Distance (JSD) as the feature dissimilarity measure, specifically because it satisfies the necessary criterion of being a true metric within the feature space, as outlined in20. Fourth, in contrast to most other methods that assign discrete labels to each spot, SPaSE provides each spot with a real-valued pathology score, which exhibits smooth variation within spatial domains. Finally, leveraging pathology scores, SPaSE uses a p-value cutoff (based on the empirical null distribution) to identify the regions with significantly high pathological scores.
To assess SPaSE’s ability to detect pathological regions and quantify variations in pathological impacts, we simulated pathological samples by perturbing gene expressions in randomly selected areas of a healthy mouse heart sample (Methods). Specifically, we randomly selected a spot and applied a 3-layer perturbation, forming three concentric regions with increasing radii centered on the chosen spot. The intensity of the perturbation decreased as the radius expanded. Using this approach, we generated three distinct simulated pathological samples, each with a different center spot (Figure 2a). SPaSE successfully assigned higher pathological scores to spots with perturbed gene expressions (Figure 2b), and pathology score increased with the extent of the simulated pathology (Figure 2c). These results demonstrated SPaSE’s ability to identify pathological regions. With these findings, we proceeded to apply SPaSE to real datasets.
Figure 2.

SPaSE’s findings in simulated data. a. Simulated pathological regions in a mouse heart sample (Sham_1), created by introducing perturbations in three distinct areas. b. SPaSE effectively assigns high pathology scores to the simulated pathological spots across all three samples. c. Box plots comparing pathology scores between regions with simulated pathology and regions without any perturbation.
SPaSE computes spatially resolved pathology scores in infarcted mouse hearts
To demonstrate the utility of SPaSE on real data, we applied it to Calcagno et al.’s21 recent ST data of mouse hearts collected at four time points post-MI: 1 hour (1hr), 4 hour (4hr), 3 day (3d), and 7 day (7d). We reasoned that this dataset could test SPaSE’s ability to elucidate how the extent of pathological effects changed between infarct-proximal and distal regions. Post-MI, the heart is characterized by three distinct regions: the infarct zone (IZ), where myocardial cells die; the border zone (BZ), a transitional region between necrotic and viable myocardium; and the remote zone (RZ), comprising the largely unaffected myocardial tissue22–24. The BZ is of special interest as it is not a mere demarcation, but it actively participates in infarct expansion, ventricular remodeling, arrhythmogenesis, and cardiomyocyte (CM) proliferation and repair23,25–32. Over time, the BZ undergoes dynamic changes, leading to potential ventricular dilatation, wall thinning, or even rupture, culminating in exacerbated cardiac symptoms and heart failure progression29,33,34.
Despite its significance, molecular characterization of the BZ remains elusive21. Traditional histological approaches offer limited precision in distinguishing it from neighboring zones21. In their study, Calcagno et al. derived gene signatures of homeostatic and structurally stressed CMs from a separate single-nucleus RNA-seq dataset. They used these signatures to define the IZ, RZ, and two BZs (BZ1 and BZ2), which separated the RZ and IZ in their ST data.
In contrast, SPaSE computes pathology scores solely by comparing control and infarcted hearts, without leveraging additional knowledge such as gene signatures of CM states derived from additional data. Nevertheless, SPaSE’s pathology scores showed high concordance with Calcagno et al.’s findings and, in some cases, provided a more nuanced understanding of post-MI tissue responses. For example, Calcagno et al. reported that the IZ does not begin to form at the 1hr and 4hr time points. SPaSE broadly supports this observation, showing that the majority of spots at 1hr and 4hr have pathology scores that were not statistically significant. However, SPaSE also identifies that, even as early as 1hr and 4hr post-MI, spots immediately adjacent to the ligature site exhibit significantly high pathology scores (Supplementary Figures S1a–S1b).
At the 3d and 7d time points, SPaSE’s pathology scores (Figure 3a) aligned remarkably well with Calcagno et al.’s annotations (Figure 3b). The pathology score distributions differ significantly between each pair of zones and displayed a clear increase from the RZ to BZ to IZ, indicating that pathology scores can effectively demarcate the zones formed post-MI. Notably, the interquartile ranges (IQR) of pathology scores in the RZ are the smallest among all zones in the 3d and 7d samples (Figure 3c), consistent with the expectation that RZ tissues are transcriptionally similar to healthy tissues. In contrast, the BZ2 pathology score IQRs are consistently the largest, likely reflecting the diverse processes activated within the BZs post-MI, such as infarct expansion, ventricular remodeling, and cardiomyocyte proliferation.
Figure 3.
SPaSE’s findings in injured mouse heart samples. a. The distribution of pathology scores generated by SPaSE (****: p-value ≤ 0.0001, independent t-test). b. ST data clustering by Calcagno et al. into four domains (IZ, BZ1, BZ2, and RZ) c. The boxplots show the pathology scores of the spots for each of the domains identified by Calcagno et al. in the ST samples of the two time points 3d and 7d.
These results demonstrate that SPaSE is a promising tool for investigating pathologically impacted tissues without the need for intricate microdissections or additional molecular data. Furthermore, since pathology scores are continuous values, they capture subtle changes in pathological impact across tissues (Figure 3a). Below, we illustrate how modeling and analyzing pathology scores can shed light on the spatially variable biological processes activated post-MI.
Biological insights from pathology scores
SPaSE’s pathology score provided spatial and temporal insight into various cellular processes and cell types in the post-MI niche. Our downstream analysis revealed that most post-MI cellular activity occurs in the days following the infarct event, which is consistent with the literature35. We elucidated the impact of different cellular processes and cell types on the remodeling process, both spatially and temporally, by examining the correlation between pathology scores and spatial abundance scores of multiple gene sets associated with various cellular activities (Methods). We identified apoptosis and senescence as driving cellular processes at 3d and 7d post-MI while autophagy did not appear to play a crucial role (Figure 4 and Supplementary Figure S2). Apoptosis is known to play an important role in the post-myocardial infarction process of left ventricular remodeling and may lead to the downstream development of symptomatic heart failure36. Similarly, senescence is known to be accelerated in the post-infarcted heart37. Furthermore, we applied our correlation study to four fibroblast populations from Kuppe et al.38 (Supplementary Table S1). We identified a positive correlation for two fibroblast populations (fibroblast 1 and 2) involved in remodeling and biosynthetic processes at 7d while fibroblast 4, which is thought to support cardiomyocytes38, is negatively correlated (Figure 4 and Supplementary Figure S3). Studies have previously demonstrated that cardiac fibroblasts are the primary source of extracellular matrix (ECM) molecules, like collagen, in the post-MI heart39, supporting the roles of the positively correlated fibroblast populations. Genes marking the fibroblast 4 population, on the other hand, are enriched for cardiac muscle contraction and heart contraction according to Enrichr40, suggesting that this population of fibroblasts may play a role in supporting cardiomyocytes. This explains the negative correlation between the spatial abundance score of the fibroblast 4 population and pathology scores. Additionally, we recapitulated the role of immune cells in the post-MI niche as several immune cell populations, e.g. monocytes and dendritic cells (Supplementary Table S1), demonstrate a positive correlation with the pathology score (Figure 4 and Supplementary Figure S5). Monocytes, for example, are known to help protect against further cardiac injury in the post-MI heart according to studies in adult mice41 while dendritic cells may play an anti-inflammatory role in the post-MI heart42. We also studied various cardiomyocyte populations, vCM1–438 (Supplementary Table S1). Our analysis demonstrated that vCM1–3 are anticorrelated with the pathology score while vCM4 is positively correlated with the pathology score at 7d post-MI (Figure 4 and Supplementary Figure S5). vCM1–3 appears to be involved in the conduction and muscular functions of the heart. Enrichr pathway enrichment analysis of vCM1 suggested that this population of cells is involved in the regulation of cardiac conduction while vCM2 and vCM3 are enriched in cardiomyocytes and myoblasts. The observed anticorrelation is expected as contractile and conductive function in the post-MI heart infarct area will experience a loss of cardiomyocytes concurrently with remodeling of the infarct zone35. On the other hand, vCM4 is enriched in chondrocytes, which are cells responsible for collagen synthesis and formation. Our previous work43 has demonstrated that Hippo-deficient cardiac fibroblasts may differentiate into osteochondroprogenitors, thereby decreasing cardiac fibroblast plasticity and potentially contributing to fibrosis in the post-MI niche. Accordingly, it is expected that vCM4 cell populations would be positively correlated with the pathology score in the post-MI heart and may play a role in post-MI cardiac fibrosis44. Taken together, our pathology score provides an unbiased insight into various post-MI processes and cell types, providing a hypothesis-generating tool for researchers interested in post-MI biology and a hypothesis-driven tool for researchers with specific questions.
Figure 4.
Connected scatter plots depict the spatial abundance scores for gene sets associated with senescence, fibroblast populations 2 and 4, monocyte, vCM2, and vCM4 across all time points. The vertical axis represents the spatial abundance score, while the horizontal axis represents spots organized in ascending order based on their pathology score. A 100-value moving average is represented by the red line.
SPaSE demonstrates broad applicability in additional tissues
We applied SPaSE to two additional ST datasets (Visium, 10x Genomics) collected from mouse models of Duchenne muscular dystrophy45 and traumatic brain injury (TBI)3. The and mouse models are widely used for Duchenne muscular dystrophy research. The mdx model carries a nonsense mutation in exon 23 of the DMD gene, while the more severe model has the same mutation on a genetic background. Preclinical studies have shown that both models exhibit necrosis, inflammation, regeneration, central nucleation, fiber size variability, fibrosis, and tissue calcification46–48. In ST data, the pathological sample contains 1880 spots, with 2341 spots in its healthy counterpart (C57BL10), while the sample has 1367 spots compared to 1425 in its healthy counterpart . Clustering based on gene expression, marker gene profiles, and histological observations by Heezen et al.45 provided detailed annotations of pathological regions. Using SPaSE, we identified pathological regions in both models, with higher pathology scores assigned to affected areas annotated by Heezen et al. In the sample, higher pathology scores correspond to regenerating fibers and inflamed patches, necrotic fibers and macrophages (Figure 5a). In the sample, higher pathology scores highlight inflamed, calcified, and necrotic fibers (Figure 5b). Notably, the model, associated with greater disease severity47, also displayed higher mean pathology scores (50.53) than the sample (45.73), as visualized in the heatmaps. These results confirm SPaSE’s effectiveness in identifying and quantifying pathological regions across different models.
Figure 5.
SPaSE’s findings in mouse DMD and mouse brain TBI samples. a. Annotation by Heezen et al. of mouse DMD sample (), and the distribution of pathology scores calculated by SPaSE on the same sample. b. Annotation by Heezen et al. of mouse DMD sample (), and the distribution of pathology scores calculated by SPaSE on the same sample. c. Annotation by Pham et al. of mouse brain TBI sample (this figure of clusters annotated by Pham et al. is taken from their published manuscript53 which is licensed under a Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/). The image of the tissue and the spot annotations are taken only.), and the distribution of pathology scores calculated by SPaSE on the same sample.
For the TBI model, characterized by gradients of microglia activation49, we analyzed one healthy and one TBI brain sample generated by Pham et al.3, consisting of 2968 and 2442 spots, respectively. Under normal conditions, microglia exhibit minimal variation across brain regions and maintain a ramified structure50,51. Following injury, however, they undergo significant changes in gene expression and morphology3. Using PCA and Louvain clustering on the top 50 PCs, the authors identified 15 clusters across the TBI sample, annotated using the Allen Mouse Brain Atlas52. SPaSE effectively identified pathological regions, assigning higher pathology scores to the lesion core/impact site, with a gradual decrease toward the lesion margin (Figure 5c). It also highlighted injury around the fiber tract region (Figure 5c).
These results demonstrate SPaSE’s broad applicability across diverse ST datasets for identifying pathological regions.
Regression models using gene expression data reveals gene sets defining pathology scores
We asked if pathology scores can be modeled from gene expression values in an injured tissue. This kind of predictive model can reveal the ‘signature genes’ defining pathology scores. In turn, the signature genes can be used more broadly to score cells in non-spatial single-cell RNA-seq data or samples of bulk RNA-seq data. Furthermore, when pathological responses are largely similar between humans and model organisms, such as mice, the signature gene sets can be first derived from model organism data where control samples are easily available. Then, the gene sets can be used to analyze diseased or injured human tissues where control samples are difficult to obtain. Thus, we applied support vector regression (SVR) to fit pathology scores from gene expression data of the mouse heart post-MI samples. For each time point, the model was trained using data from one replicate (chosen arbitrarily), and validated on a different replicate from the same time point. The model showed high predictive accuracy (Spearman correlation coefficient of ~0.90; Supplementary Table S2), indicating that gene expressions in ST spots are accurate predictors of pathology scores. In addition, we applied SVR models trained with mouse gene expression data to 27 human heart biopsy samples. These samples were collected from 23 individuals, including four non-transplanted donor hearts as controls, and samples from tissues with necrotic areas (ischemic zone and border zone) and the unaffected left ventricular myocardium (remote zone) of patients with acute myocardial infarction38. In the following section, we will show that these SVR models can effectively predict pathology in human samples as well.
In addition to achieving high accuracy, the SVR models also revealed the genes strongly associated with pathology scores at each time point (Figure 5a). In particular, we assigned a predictive importance score to each gene (Methods; Supplementary Table S3) and took the top 50 genes per time point for functional gene set enrichment analysis. The gene sets from the two time points were largely independent (only four genes in common; Supplementary Table S3), immediately indicating that distinct biological processes characterize the pathologically impacted regions in the 3d and 7d samples. Enrichr analysis shows that genes from the 3d sample are enriched in pathways such as myogenesis and inflammatory response, while genes from the 7d sample are enriched in pathways like myogenesis and apoptosis. This suggests that the muscle generation process (myogenesis) continues at later time points post-MI, whereas processes like apoptosis may not initiate immediately after MI but require some time to begin. Several crucial genes were identified, including Actb, Csrp3, Cryab at 3d, as well as Clu, Itih4, Tbxas1 at 7d. The beta-actin gene (Actb) is involved in vascular remodeling processes, which can contribute to cardiovascular diseases53. Additionally, it is significant that Actb shows differential expression across various cardiovascular conditions, including hypertrophic cardiomyopathy (HCM)54. Csrp3 is also a well-known gene associated with dilated cardiomyopathy (DCM) and HCM55. Alpha B-crystallin (Cryab), a member of the small heat shock protein (HSP) 20 family, is highly expressed in cardiomyocytes and is upregulated in response to stressors such as heat shock, inflammation, and oxidative stress56. CLU is a promising biomarker for prognostic evaluation in patients with chronic heart failure57. Itih4 has been linked to coronary artery disease through genome-wide association studies58. TBXAS1 is another gene that is implicated in cardiovascular functions and diseases59. Overall, prior literature consistently suggested roles for SPaSE-prioritized genes in cardiac diseases and functions.
The earlier interpretation strategy yielded interesting genes crucial for predicting pathology scores across the entire sample. However, for a more nuanced understanding of spatial variability in pathological effects, a finer-grained interpretation at the cardiac zone level would be immensely valuable. Therefore, to identify important genes specific to each cardiac zone, we employed SHAP (SHapley Additive exPlanations)60. SHAP, a computationally intensive but widely used tool, employs a game-theoretic approach to elucidate the output of machine learning models. Due to its computational constraints, we constructed an SVR model similar to the previous one but trained on the top 100 common Spatially Variable Genes (SVG). We employed SHAP to interpret this SVR model at the spot level. The explanation of the predicted pathology score for each spot involves assigning SHAP values to individual genes associated with that spot. SHAP quantifies the contribution of genes to deviations in pathological scores from the tissue-wide average. At each spot, the SHAP values of all genes sum to the difference between the predicted pathology score and the average pathology score. The SHAP value of a gene represents its individual contribution to this deviation. Therefore, a negative SHAP value indicates a negative contribution to the pathology score, whereas a positive SHAP value indicates a positive contribution to the pathology score. We scrutinized these SHAP values within distinct cardiac zones, as illustrated in Figure 6b. Noticeably, the distribution of SHAP values for the top-ranked genes varied across different zones. We illustrate this variability through heatmaps displaying SHAP values for each cardiac zone in 3d (Figure 6c). In these heatmaps, the vertical axis represents genes, ordered in descending importance from top to bottom and the horizontal axis corresponds to spots within that particular zone. The importance of a gene is determined by summing the SHAP values across all spots within the designated zone, as visualized by a vertical histogram positioned adjacent to the heatmaps (the row sum of the heatmap). In alignment with our anticipated outcomes, the SHAP values of genes associated with RZ spots exhibit notably low values, represented by a bluish heatmap matrix, signifying a negative impact on the pathology score. Moving from BZ1 and BZ2 towards the IZ, the heatmap cells progressively transition to a redder hue, indicating increased SHAP values and a positive impact on the pathology score. A noteworthy observation is the presence of a gene with lower SHAP values in the RZ and simultaneously higher SHAP values in the IZ. This emphasizes that a gene may exhibit a dual nature, influencing both positive and negative effects on the pathology score. The top genes include Myl3, Trdn, and Ctsb, all of which are associated with sudden cardiac death (SCD). Homozygous loss-of-function variants in Myl3 can cause recessive cardiomyopathy and may lead to SCD61. Common variants in Trdn are associated with an increased risk of SCD in patients with chronic heart failure62. Ctsb was found to be upregulated in SCD samples, with functional enrichment analysis indicating its role in the lysosome pathway. Furthermore, Ctsb has been identified as an autophagy-related gene in the Human Autophagy Database63.
Figure 6.

SPaSE identifies sample-specific and zone-specific important genes. a. Sample-specific important genes are identified by training an SVR model with a linear kernel on gene expressions to predict pathology scores. The training set comprised the top 2000 HVGs. Post-training, genes are ranked based on the model-learned coefficients. b. Identification of cardiac zone-specific genes is carried out using an SVR model, trained with the top 100 HVGs. SHAP values are calculated for each gene in every spot. Gene ranks are determined within specific cardiac zones based on the sum of absolute SHAP values. c. Heatmaps illustrating SHAP values of the spots for each cardiac zone of 3d sample. Genes are ranked according to their importance from top to bottom, with the importance values depicted as a vertical histogram on the right side of the heatmap. d. Expression distributions of some important genes found using SHAP analysis from 3d sample.
We performed a similar analysis on the 7d sample. The heatmaps exhibited similar trends (Supplementary Figure S6). However, the genes identified as important were entirely different from those in the 3d case, suggesting distinct underlying mechanisms at these time points. The genes Bgn, Lox, and Tmsb4x were among the top genes prioritized by SHAP.
Bgn has been identified as a potential target for arrhythmogenic right ventricular cardiomyopathy64. Lox is crucial for extracellular matrix (ECM) assembly, a vital process for tissue maintenance, repair, and cellular communication. Its roles span extracellular, intracellular, and nuclear functions, impacting normal tissue physiology and contributing to various disorders, including vascular and cardiac conditions65. Tmsb4x was found to be enriched in heart valves and ventricular walls at later stages of development, as revealed by spatial RNA-seq analysis, emphasizing its importance in heart structure and function66. Notably, Tmsb4x has been identified as a gene involved in later stages of both cardiogenesis and post-MI, suggesting that gene responses post-MI may reflect their roles during developmental stages.
We anticipated discovering distinct sets of genes contributing to the pathology scores in each respective zone. However, we observed that in the 3d sample, 16 of the top 20 important genes identified through SHAP analysis were common across the four zones, while in the 7d sample, 18 genes were commonly identified. This suggests that nearly identical sets of genes are governing the pathology scores across the entire sample. Overall, we observed that the top genes contributing to pathology showed high overlap across different zones within the same time point but varied substantially between the two time points.
SPaSE predicts pathological scores of human ST samples
We apply SVR models trained on mouse heart data to human heart samples from the study by Kuppe et al.38. Their dataset includes 31 samples from 23 individuals, comprising four non-transplanted donor hearts as controls. The study focuses on tissue from necrotic areas (IZ, BZ) and unaffected left ventricular myocardium (RZ) in patients with acute myocardial infarction. Samples were collected at various stages post-symptom onset (chest pain) and prior to receiving artificial hearts or left-ventricular assist devices as a bridge to transplantation. Additionally, nine specimens from later-stage myocardial infarction (fibrotic zone, FZ) were obtained from heart transplant recipients. In total, we analyzed 27 Visium ST human heart samples, each containing approximately 3285 spots and 15055 genes on average. For each pair of mouse and human samples, we trained an SVR model using shared genes as features, with pathology scores in mouse samples as the target. The SVR model trained on mouse samples predicts pathology scores for the corresponding human samples (Methods). The predicted pathology scores in human data align with expected trends, showing a clear progression: control samples exhibit lower pathology scores, which increase through the RZ, FZ, and IZ (Figure 7a). Regions with elevated activity in disease-related pathways, such as TGFβ and (as reported in Figures 1h and 6i of Kuppe et al.), correspond to higher pathology scores (Figure 7b). Additionally, we observed reduced pathology scores in CM regions (Figure 7c–e). These findings demonstrated that SVR models trained on mouse heart data with SPaSE-derived pathology scores can provide valuable insights into human hearts, supporting the broad applicability of SPaSE.
Figure 7.
SPaSE predicts pathology scores of injured human heart samples. a. Boxplots of predicted pathological scores for human heart samples, including control, RZ, FZ, and IZ zones, reveal an increasing trend (****: p-value ≤ 0.0001, independent t-test). b. Predicted pathological scores of IZ sample from P3 is higher in the regions where TGFβ or pathways are active (Figure 1h of Kuppe et al.’s work). Moreover, higher predicted pathological scores are observed in IZ sample of P10 where TGFβ pathway is active (Figure 6i of Kuppe et al.’s work). c-e. SPaSE’s predicted pathological scores are anti-correlated with cardiomyocyte abundance scores calculated by Kuppe et al. in IZ samples from P9 and P15.
Discussion
We presented SPaSE, a tool designed to quantify the pathological impact on tissue regions in a diseased ST sample compared to a corresponding healthy ST sample. This quantification leverages an optimal transport plan constructed based on differences in gene expression and the spatial locations of ST spots in both samples. This approach enables the assessment of spatial variability in pathological effects.
We used SPaSE to compute pathology scores for mouse heart post-MI samples. The pathology scores align effectively with Calcagno et al.’s annotations, increasing from RZ to IZ as anticipated. Notably, the interquartile ranges (IQRs) of pathology scores in BZ2 are consistently the widest, likely reflecting the diverse processes activated post-MI within the BZ. The real-valued nature of the pathology scores allowed us to capture the variability of pathological effects specific to each region. These findings suggested it was inappropriate to treat all spots within a specific cardiac zone uniformly.
To investigate the mechanisms underlying this spatial variability, we trained an SVR model on gene expression data to predict pathology score based on the expression profile of each spot. This model identified crucial genes that influenced pathology score predictions for a given sample, serving as interpretative features for the scores across the sample. Our analysis revealed that the top 50 most important genes for 3d and 7d samples were largely independent, indicating that distinct biological processes characterized the pathology at these two time points.
Through enrichment analysis, we uncovered a prolonged state of myogenesis post-MI. At 3d, we found no association with apoptosis regulation, while at 7d, we observed a clear link to apoptosis. This suggested that during the early time point, the immune system may have attempted to counteract apoptosis, but this response did not persist with continued inflammation. Additionally, some important genes, when examined in their human counterparts, showed associations with cardiac injury. This indicates that impactful genes related to human cardiac injury fan be identified through model organisms like mice.
To explore the spatial variability of pathological effects within a tissue sample, we employed SHAP, a model interpretation tool that identified region-specific critical features (genes). SHAP analysis revealed that at a specific time point, most genes crucial for predicting pathology score remained consistently important across various cardiac zones. However, some genes exhibited dual roles, contributing both positively and negatively to pathology score. This duality suggested that targeting a single gene for therapeutic intervention can be challenging, as it may have differing effects in pathologically affected and unaffected zones. Furthermore, the most critical region-specific genes differed between 3d and 7d samples, underscoring the need for distinct therapeutic targets at different time points.
We evaluated SPaSE on two additional mouse datasets (mouse brain and mouse DMD) and one human heart dataset to assess its broad applicability. In mouse DMD samples, SPaSE effectively identifies pathological regions across two samples, assigning high pathological scores to features such as regenerating fibers, calcified fibers, necrotic fibers, and inflamed patches. For the mouse TBI sample, SPaSE highlights high pathology scores at key pathological areas, including the impact site and lesion margin, with moderate pathology scores observed in regions like the lateral ventricle, choroid plexus, and fiber tracts. These findings align with expected pathological responses in these regions. In human heart samples, SPaSE reveals mechanistic insights, such as an inverse relationship between cardiomyocyte activity and pathology score, suggesting a defensive role of cardiomyocytes post-MI. Additionally, higher pathology scores are observed in disease-associated pathways, including TGFβ and NFκB. These results demonstrate SPaSE ‘s capability to provide meaningful insights into pathological responses across various diseases and injuries.
There are several potential directions to enhance SPaSE. The current objective function does not incorporate cell-type information, which can provide valuable insights if available. The mapping strategy used for computing the optimal transport plan treated both healthy and diseased samples as discrete uniform probability distributions. Future work can adopt a more systematic approach to assigning initial mass to each spot, incorporating biological or technical factors. While this study applied SPaSE to Visium data, these changes would make SPaSE more broadly applicable to other spatial transcriptomics technologies such as MERFISH67, STARmap68, and Slide-Seq269. Another important direction would be incorporating H&E images into the SPaSE pipeline. This would introduce additional terms into the cost function, and to gain the type of mechanistic understanding reported in this study, it would be essential to quantify how the two modalities (gene expression and H&E image) influence pathology score. This, in turn, would require datasets with carefully curated annotations for pathologically impacted regions. As more such data become available, incorporating H&E images would be a high-priority extension of SPaSE.
Resource availability
Lead contact
Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Md. Abul Hassan Samee.
Materials availability
This study did not generate new materials.
Data and code availability
ST data of mouse hearts can be found in the Gene Expression Omnibus (GEO) under accession number GSE214611. The raw data files for the mouse DMD samples can be accessed in the GEO under accession number GSE199659. The processed data files are available on Zenodo at https://doi.org/10.5281/zenodo.7401196. The raw and processed sequencing data of the mouse TBI sample have been deposited in the GEO under accession number GSE236171 and are publicly available. The spatial transcriptomics of human heart samples can be accessed on CellxGene at https://cellxgene.cziscience.com/collections/8191c283–0816-424b-9b61-c3e1d6258a77 and in the Zenodo data archive at https://zenodo.org/record/6578047.
All original code has been deposited at https://zenodo.org/records/14941956 and is publicly available as of the date of publication. DOIs are listed in the key resources table.
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Deposited data | ||
| Spatial transcriptomics (Visium, 10x Genomics) data of healthy and injured mouse heart. Injuries include MI (permanent ligation of the LAD coronary artery), I/R-30 (30 minutes of ischemia followed by reperfusion), transaortic constriction, isoproterenol injection, and traumatic needle injection. | Calcagno et al. | GEO: GSE214611 |
| Spatial transcriptomics (Visium, 10x Genomics) data of Duchenne mouse models | Heezen et al. | GEO: GSE199659 |
| Spatial transcriptomics data of mouse traumatic brain injury | Pham et al. | GEO: GSE236171 |
| Spatial multi-omic map of human myocardial infarction | Kuppe et al. | DOI: 10.5281/zenodo.6578046 |
| Software and algorithms | ||
| Loupe Browser v7.0.1 | 10x Genomics | https://www.10xgenomics.com/support/software/loupebrowser/latest |
| Python v3.8.13 | Python Software Foundation | https://www.python.org/downloads/release/python3813/ |
| POT v0.8.2 | Flamary et al. | https://PythonOT.github.io/ |
| SHAP v0.42.0 | Su-In Lee’s lab at the University of Washington, and Microsoft Research. | https://shap.readthedocs.io/en/latest/ |
| Original codes | This manuscript | DOI: 10.5281/zenodo.14941956 |
STAR Methods
Method details
Dataset preprocessing
Within the dataset, certain spots were identified as not positioned over the tissue. To streamline the Visium data and exclude these non-relevant spots, we employed the CLOUPE files through the Loupe Browser. By aligning the provided image of the spots and detecting those requiring removal, we effectively filtered out specific data points, ensuring a refined dataset for subsequent analysis.
SPaSE Algorithm
The challenges encountered in ST experiments frequently result in variations in slice orientations, which can complicate analyses that involve multiple ST slices. A potent solution, such as optimal transport, has emerged to address this issue. It achieves orientation invariance by formulating spatial constraints and incorporating data features, thus enhancing the analysis of the data. Optimal transport is a technique that takes as input two probability distributions and a cost matrix defined between the respective support sets of these distributions. It then generates an optimal mapping or transport plan between the points within these distributions. This transport plan effectively moves the masses from the source distribution to match the target distribution, all while minimizing an objective function that incorporates the given cost matrix. SPaSE successfully uses the theory of optimal transport in a unique way to identify pathologically remodeled regions of a diseased ST slice, by mapping the spots of a healthy ST slice, to the spots of . To the best of our knowledge, there are no other approaches that deduce afflicted regions or spatial domains within a diseased ST sample through a comparison with a healthy ST sample. ST data consists of the transcript counts of a set of genes and the spatial locations of all the spots of an ST slice. Let and be the number of spots of and respectively. The number of genes commonly expressed in both and is . We will consider the expressions of these genes as the features of the ST spots for both and . Let and be the transcript count matrix and and be the Euclidean distance matrices for the samples and respectively.
SPaSE identifies pathologically remodeled regions of with respect to by constructing a transcriptionally and spatially aware joint probabilistic map, from the spots of to the spots of with entropic regularization, which minimizes the FGW distance between the spots of and spots of using the theory of optimal transport. To apply the method to find an optimal transport plan, SPaSE treats and as probability distributions, and respectively where each of the spots is assigned a probability mass. We assigned a mass of and to spots of and respectively making and discrete uniform probability distributions. The FGW distance jointly takes into account the gene expression dissimilarity and the structural dissimilarity of the spots. The calculation of structural dissimilarity relies on computing the pairwise Euclidean distances among the spots within a single sample. Therefore, the necessity for and to be aligned initially is eliminated. To represent the gene expression dissimilarity between two spots, SPaSE leverages the JSD between the lists of gene expression values of those spots, as JSD is a metric, and when addressing the optimal transport problem with the FGW distance, it is assumed that the distance within the feature space conforms to the metric properties20. Formally, the dissimilarity between the spot of and the spot of is given by,
where is the Kullback-Leibler divergence of distribution from distribution . SPaSE represents the structural dissimilarity of and as the difference in the Euclidean distance between a pair of spots in and that between its corresponding mapped spots in , summed over all possible mappings for all pairs of spots of and . Formally, the structural dissimilarity component is given by,
Combining the gene expression dissimilarity and the structural dissimilarity and adding the entropic regularization, SPaSE minimizes the following cost function,
| (1) |
where and are hyperparameters. is the entropic regularization term. To calculate the transport plan we used the Python OT’s70 generalized conditional gradient procedure. We chose the values of the hyperparameters and using our hyperparameter selection method (described in Hyperparameter selection subsection).
After calculating the transport plan, , SPaSE assigns a pathology score to each spot, of denoted by . The column of represents the transport plan for the spot of , i.e., the amount of mass transferred to the spot from each of the spots of . To make the values compatible with other transport plans obtained using a different healthy-disease sample pair, SPaSE normalizes each of the columns of by dividing each value of that column by the initially assigned mass of the spot corresponding to that column ( in this case).
Hence, cell of this normalized transport plan, say , represents the percentage of the initial mass of the spot of which is transported to the spot of . The pathology score of the spot of is given by,
| (2) |
Labeling a spot as pathologically remodeled
SPaSE labels tissue regions of pathologically remodeled if the spots of that region consist of spots having pathology scores having p-values less than 0.05 with respect to a null distribution which was generated in the following approach. As the pathology score generation procedure requires a mapping between two ST samples, two healthy ST samples are required for the generation of the null distribution. Hence, SPaSE decomposes into two synthetic healthy ST samples, and . The spots of can be thought of as rows of spots stacked together (Figure 1a). consists of the odd rows and consists of the even rows of . SPaSE computes the pathology scores for spots of and with respect to and respectively. Subsequently, it combined these two distributions. Then this distribution is considered to be the null distribution and the pathology scores of the spots that lie over the 95th percentile of this distribution were labeled as pathologically remodeled.
Generation of simulated pathological regions
The strategy for generating perturbed regions in a sample involves randomly selecting a spot and applying perturbations within a specified radius centered at in . The radius is defined as: , where and denote the maximum and minimum x-coordinate values across all spots, respectively, and and denote the maximum and minimum y-coordinate values across all spots, respectively, and is chosen between 3 and 6 randomly. For each spot within this region, the transcript count of a gene is replaced with a value sampled from a normal distribution, , where and denote the mean and standard deviation of the transcript count of across all spots in the sample, respectively. The scalar parameter controls the extent of gene expression perturbation.
Hyperparameter selection
To determine the optimal combination of hyperparameters, we perturbed the diseased sample, , to generate multiple perturbed diseased samples. Specifically, we created three perturbed diseased samples using the previously described perturbation strategy. The spots subjected to perturbation were labeled as “pathological”, while the remaining spots in were labeled as “non-pathological”. These labels served as the ground truth for optimizing hyperparameters. SPaSE was then applied to identify spots within exhibiting pathological remodeling, assigning each spot a label of either “pathological” or “non-pathological”. To assess SPaSE’s performance, we calculated the F1-score based on these labeling results. We employed a grid search to identify the optimal hyperparameters, varying within 0.0001, 0.001, 0.01, 0.1 and within 0.001, 0.01, 0.1. The average F1-score across the three perturbed diseased samples was used as the criterion for selecting the optimal hyperparameter values.
Gene sets defining pathology score
To identify the important genes that define pathology scores of the spots of an ST sample, we trained SVR models with linear kernels that take gene expressions of the spots as inputs and outputs pathology score. Firstly, we identified genes that are specific to a sample and hold predictive value for the pathology scores associated with the spots in that particular sample. Given that we have two replicates for each of the two time points (3d and 7d), our approach involved training an SVR model on one replicate and testing it on the other to check the model’s accuracy. Dealing with spots characterized by an array of gene expression values (approximately 32,000), training a model using the entire dataset becomes computationally prohibitive. As a practical solution, we leveraged the top common 2000 SVGs between 3d and 7d as the feature set for these spots. These SVGs were identified using SpatialDE71, which helped streamline the computational process. After the model training, genes were prioritized based on their respective coefficient values. A higher squared coefficient value indicates a stronger predictive capacity for the pathology scores. Following the sorting of genes from two replicates of a specific sample (either 3d or 7d) according to the coefficients, we merged and arranged the genes based on their average ranks. Next, we conducted enrichment analysis using Enrichr on the top 50 genes to identify shared genes across various pathways, cell types, and other categories. Secondly, in characterizing pathology score variability within ST samples, we identified cardiac zone-specific important genes by training an SVR model with 100 common SVGs between 3d and 7d and analyzed the model for each cardiac zone’s spots using SHAP python module72. The choice of 100 SVGs was made to optimize computational efficiency and memory usage while maintaining a reasonable level of model accuracy. We used the ‘KernelExplainer’ of the SHAP module as the explainer for our SVR model. This explainer was used to determine the SHAP values for every gene for each of the spots of a particular cardiac zone. These SHAP values were then used to rank the genes. The ranking criteria involved the summation of absolute SHAP values for a gene across the spots. A higher value indicates greater importance.
Calculating gene set spatial abundance score
Given the gene sets associated with various cell types and processes (Supplementary Table S1), we computed the spatial abundance score at each spot for a particular cell type or process. The calculation of for a specific set of genes at a spot is as follows:
Let represent a set of genes. The spatial abundance score at spot for the gene set is defined as:
where is the expression value of gene at spot . The term represents the background score of the gene set at spot and is calculated as follows: The genes were sorted in ascending order based on their average expression across all spots in the sample. This sorted gene list was divided into 49 bins. For each gene in , a corresponding background gene was selected by randomly sampling a gene from the same bin, excluding the gene itself, to construct a background gene set. This process was repeated 1000 times to generate 1000 background gene sets. The background score was computed as the average of the mean expressions of these 1000 background gene sets at spot .
Pathological scores for human samples
To calculate the pathology scores for human samples, we utilized SVR models derived from mouse samples. Let denote the set of human samples and denote the set of mouse samples. For each mouse sample , an SVR model was already trained to predict pathology scores for each spot based on 2000 spatially variable genes (SVGs). Let represent the set of genes commonly expressed in the human sample and the 2000 SVGs used to train the SVR model for the mouse sample . Using these common genes, we trained an SVR model for the human sample , where the pathology scores derived from the mouse model served as the target variable. In this manner, we trained SVR models, enabling us to compute pathology scores for human samples while leveraging information from mouse samples.
Supplementary Material
Supplementary table S1: Gene sets corresponding to various cell-types, counts of each of the cell-types, and translation between human and mouse genes, related to Figure 4.
Supplementary table S2: SVR model’s (trained on top 2000 highly variable genes (HVG)) performance measured using Spearman’s correlation coefficient, related to Figure 6.
Supplementary table S3: Average ranks of the genes based on SVR coefficients, related to Figure 6.
Highlights.
SPaSE quantifies spatially variable pathological effects using optimal transport.
Pathology scores delineate infarct, border, and remote zones in post-MI mouse hearts.
Gene expression-based models predict pathology severity and generalize to human MI data.
Pathology scores help identify key genes, cell-types, and biological processes.
Acknowledgement
M.N.R. was supported by the ICT Fellowship at the Masters level in the fiscal year 2022–2023 for research in ICT sector, administered by ICT Division, Government of People's Republic of Bangladesh. M.S.R. was partially supported by Basic Research Grant from BUET. J.F.M. was supported by the National Institutes of Health (HL 169511, HL 171574, and HL 118761 to J.F.M.) and the Vivian L. Smith Foundation. M.A.H.S. was supported by the National Institutes of Health (AG 081192).
Footnotes
Author contributions
Conceptualization: M.A.H.S.;
Algorithm design (the SPaSE algorithm, the SVR model, and downstream analyses): M.N.R., M.S.R., M.A.H.S.;
Implementation (the SPaSE algorithm, the SVR model, and downstream analyses): M.N.R.;
Result interpretation: M.N.R., J.F.M., M.S.R., M.A.H.S.;
Gene set modeling: M.A.A.;
Interpreting results of gene set modeling: V.R.S.;
Manuscript writing: MNR with inputs from M.A.A., V.R.S., M.S.R., J.F.M., and
M.A.H.S.; All authors edited the manuscript.
Declaration of Interests
J.F.M. is a cofounder and owns shares in YAP Therapeutics.
Declaration of generative AI and AI-assisted technologies in the writing process
During the preparation of this work, we used ChatGPT to improve the readability and language of the manuscript. After using this tool, we reviewed and edited the content as necessary and took full responsibility for the final content of the published article.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- 1.Zhao E et al. (2021). Spatial transcriptomics at subspot resolution with BayesSpace. Nat Biotechnol 39, 1375–1384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Tan X, Su A, Tran M & Nguyen Q. (2020). Spacell: integrating tissue morphology and spatial gene expression to predict disease cells. Bioinformatics 36, 2293–2294. [DOI] [PubMed] [Google Scholar]
- 3.Pham D et al. (2020). stLearn: integrating spatial location, tissue morphology and gene expression to find cell types, cell-cell interactions and spatial trajectories within undissociated tissues. bioRxiv 10.1101/2020.05.31.125658. https://www.biorxiv.org/content/early/2020/05/31/2020.05.31.125658.full.pdf. [DOI]
- 4.Fu H et al. Unsupervised spatially embedded deep representation of spatial transcriptomics. Biorxiv 2021–06 (2021). [DOI] [PMC free article] [PubMed]
- 5.Shan Y et al. (2022). Tist: transcriptome and histopathological image integrative analysis for spatial transcriptomics. Genomics, Proteomics & Bioinforma. 20, 974–988. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Hu J et al. (2021) SpaGCN: Integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network. Nat. Methods 18, 1342–1351, 10.1038/s41592-021-01255-8. [DOI] [PubMed] [Google Scholar]
- 7.Yang Y et al. (2021). SC-MEB: spatial clustering with hidden Markov random field using empirical Bayes. Briefings Bioinforma. 23, 10.1093/bib/bbab466. Bbab466, https://academic.oup.com/bib/article-pdf/23/1/bbab466/42230744/bbab466.pdf. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zhang X, Wang X, Shivashankar G & Uhler C. (2022). Graph-based autoencoder integrates spatial transcriptomics with chromatin images and identifies joint biomarkers for alzheimer’s disease. Nat. Commun. 13, 1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Xu C et al. (2022). Deepst: identifying spatial domains in spatial transcriptomics by deep learning. Nucleic Acids Res. 50, e131–e131. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zong Y et al. (2022). const: an interpretable multi-modal contrastive learning framework for spatial transcriptomics. bioRxiv 2022–01.
- 11.Hou S et al. (2022). St2cell: Reconstruction of in situ single-cell spatial transcriptomics by integrating high-resolution histological image. bioRxiv 2022–10.
- 12.Dong K & Zhang S. (2022). Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder. Nat. communications 13, 1739. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Zeng Y et al. (2022). Deciphering spatial domains by integrating histopathological image and transcriptomics via contrastive learning. bioRxiv 2022–09.
- 14.Long Y et al. (2023). Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with graphst. Nat. Commun. 14, 1155. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Rahman MN et al. (2023). Scribbledom: using scribble-annotated histology images to identify domains in spatial transcriptomics data. Bioinformatics 39, btad594. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Zeira R, Land M, Strzalkowski A & Raphael BJ. (2022). Alignment and integration of spatial transcriptomics data. Nat. Methods 19, 567–575, 10.1038/s41592-022-01459-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Peyre G, Cuturi M et al. (2019). Computational optimal transport: With applications to data science. Foundations Trends Mach. Learn. 11, 355–607. [Google Scholar]
- 18.Cuturi M. (2013). Sinkhorn distances: Lightspeed computation of optimal transport. Adv. neural information processing systems 26. [Google Scholar]
- 19.Genevay A, Chizat L, Bach F, Cuturi M & Peyré G. (2019). Sample complexity of sinkhorn divergences. In The 22nd international conference on artificial intelligence and statistics, 1574–1583 (PMLR; ). [Google Scholar]
- 20.Titouan V, Courty N, Tavenard R & Flamary R. (2019). Optimal transport for structured data with application on graphs. In International Conference on Machine Learning, 6275–6284 (PMLR; ). [Google Scholar]
- 21.Calcagno D et al. (2022). Single-cell and spatial transcriptomics of the infarcted heart define the dynamic onset of the border zone in response to mechanical destabilization. Nat. Cardiovasc. Res. 1, 1039–1055. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Factor S, Sonnenblick E & Kirk E. (1978). The histologic border zone of acute myocardial infarction–islands or peninsulas? The Am. journal pathology 92, 111. [PMC free article] [PubMed] [Google Scholar]
- 23.Janse MJ et al. (1979). The” border zone” in myocardial ischemia. an electrophysiological, metabolic, and histochemical correlation in the pig heart. Circ. Res. 44, 576–588. [DOI] [PubMed] [Google Scholar]
- 24.van Duijvenboden K et al. (2019). Conserved nppb+ border zone switches from mef2-to ap-1–driven gene program. Circulation 140, 864–879. [DOI] [PubMed] [Google Scholar]
- 25.Vivien CJ, Hudson JE & Porrello ER. (2016). Evolution, comparative biology and ontogeny of vertebrate heart regeneration. NPJ Regen. medicine 1, 1–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Smith RM, Black AJ, Velamakanni SS, Akkin T & Tolkacheva EG. (2012). Visualizing the complex 3d geometry of the perfusion border zone in isolated rabbit heart. Appl. optics 51, 2713–2721. [DOI] [PubMed] [Google Scholar]
- 27.Pirolo JS, Hutchins GM & Moore GW. (1986). Infarct expansion: pathologic analysis of 204 patients with a single myocardial infarct. J. Am. Coll. Cardiol. 7, 349–354. [DOI] [PubMed] [Google Scholar]
- 28.Erlebacher J, Richter R, Alonso D, DEVEREUX R & Gay W. (1984). Early infarct expansion-structural or functional. In JOURNAL OF THE AMERICAN COLLEGE OF CARDIOLOGY, vol. 3, 593–593 (ELSEVIER SCIENCE INC 655 AVENUE OF THE AMERICAS, NEW YORK, NY 10010, 1984). [DOI] [PubMed] [Google Scholar]
- 29.Gao X-M, Xu Q, Kiriazis H, Dart AM & Du X-J. (2005). Mouse model of post-infarct ventricular rupture: time course, strain-and gender-dependency, tensile strength, and histopathology. Cardiovasc. research 65, 469–477. [DOI] [PubMed] [Google Scholar]
- 30.Mounier-Vehier F et al. (1994). Borderzone infarct subtypes: preliminary study of the presumed mechanism. Eur. neurology 34, 11–15. [DOI] [PubMed] [Google Scholar]
- 31.Wetstein L, Michelson EL, Simson MB, Moore EN & Harken AH. (1982). Increased normoxic-to-ischemic tissue borderzone as the cause for reentrant ventricular tachyarrhythmias. J. Surg. Res. 32, 526–534. [DOI] [PubMed] [Google Scholar]
- 32.Hearse D & Yellon D. (1981). The ‘border zone’and myocardial protection: a time for reassessment? Acta Medica Scand. 210, 37–46. [PubMed] [Google Scholar]
- 33.Yang Y et al. (2008). Age-related differences in postinfarct left ventricular rupture and remodeling. Am. J. Physiol. Circ. Physiol. 294, H1815–H1822. [DOI] [PubMed] [Google Scholar]
- 34.Cavasin MA, Tao Z, Menon S & Yang X-P. (2004). Gender differences in cardiac function during early remodeling after acute myocardial infarction in mice. Life sciences 75, 2181–2192. [DOI] [PubMed] [Google Scholar]
- 35.Frangogiannis NG. (2011). Pathophysiology of myocardial infarction. Compr. Physiol. 5, 1841–1875. [DOI] [PubMed] [Google Scholar]
- 36.Teringova E & Tousek P. (2017). Apoptosis in ischemic heart disease. J. translational medicine 15, 1–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Redgrave RE et al. (2023). Senescent cardiomyocytes contribute to cardiac dysfunction following myocardial infarction. npj Aging 9, 15. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Kuppe C et al. (2022). Spatial multi-omic map of human myocardial infarction. Nature 608, 766–777. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Burke RM, Villar KNB & Small EM. (2021). Fibroblast contributions to ischemic cardiac remodeling. Cell. signalling 77, 109824. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Kuleshov MV et al. (2016). Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic acids research 44, W90–W97. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Wang Y et al. (2020). Mydgf promotes cardiomyocyte proliferation and neonatal heart regeneration. Theranostics 10, 9100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Kologrivova I, Shtatolkina M, Suslova T & Ryabov V. (2021). Cells of the immune system in cardiac remodeling: main players in resolution of inflammation and repair after myocardial infarction. Front. immunology 12, 664457. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Tsai C-R et al. (2023). Hippo-deficient cardiac fibroblasts differentiate into osteochondroprogenitors. bioRxiv 2023–09.
- 44.Frangogiannis NG. (2021). Cardiac fibrosis. Cardiovasc. research 117, 1450–1488. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Heezen L et al. (2023). Spatial transcriptomics reveal markers of histopathological changes in duchenne muscular dystrophy mouse models. Nat. Commun. 14, 4909. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Coley WD et al. (2016). Effect of genetic background on the dystrophic phenotype in mdx mice. Hum. molecular genetics 25, 130–145. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.van Putten M et al. (2019). Natural disease history of the d2-mdx mouse model for duchenne muscular dystrophy. The FASEB J. 33, 8110. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Gordish-Dressman H et al. (2018). “of mice and measures”: a project to improve how we advance duchenne muscular dystrophy therapies to the clinic. J. neuromuscular diseases 5, 407–417. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Willis EF et al. (2020). Repopulating microglia promote brain repair in an il-6-dependent manner. Cell 180, 833–846. [DOI] [PubMed] [Google Scholar]
- 50.Masuda T et al. (2019). Spatial and temporal heterogeneity of mouse and human microglia at single-cell resolution. Nature 566, 388–392. [DOI] [PubMed] [Google Scholar]
- 51.Li Q et al. (2019). Developmental heterogeneity of microglia and brain myeloid cells revealed by deep single-cell rna sequencing. Neuron 101, 207–223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Lein ES et al. (2007). Genome-wide atlas of gene expression in the adult mouse brain. Nature 445, 168–176. [DOI] [PubMed] [Google Scholar]
- 53.Yang S et al. (2020). The actb variants and alcohol drinking confer joint effect to ischemic stroke in chinese han population. J. Atheroscler. Thromb. 27, 226–244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Xue J et al. (2018). Exploring mirna-mrna regulatory network in cardiac pathology in na+/h+ exchanger isoform 1 transgenic mice. Physiol. Genomics 50, 846–861. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Giri P, Jain D, Kumar A & Mohapatra B. (2023). Identification and in silico characterization of csrp3 synonymous variants in dilated cardiomyopathy. Mol. Biol. Reports 50, 4105–4117. [DOI] [PubMed] [Google Scholar]
- 56.Kappé G et al. (2003). The human genome encodes 10 α-crystallin–related small heat shock proteins: Hspb1–10. Cell stress & chaperones 8, 53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Koller L et al. (2017). Clusterin/apolipoprotein j is independently associated with survival in patients with chronic heart failure. J. clinical lipidology 11, 178–184. [DOI] [PubMed] [Google Scholar]
- 58.Ravindran A et al. (2024). Translatome profiling reveals itih4 as a novel smooth muscle cell–specific gene in atherosclerosis. Cardiovasc. Res. cvae028. [DOI] [PMC free article] [PubMed]
- 59.Beccacece L et al. (2023). A genome-wide analysis of a sudden cardiac death cohort: Identifying novel target variants in the era of molecular autopsy. Genes 14, 1265. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Lundberg SM & Lee S-I. (2017). A unified approach to interpreting model predictions. Adv. neural information processing systems 30. [Google Scholar]
- 61.Osborn DPS et al. (2021). Autosomal recessive cardiomyopathy and sudden cardiac death associated with variants in myl3. Genet. Medicine 23, 787–792. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Liu Z et al. (2015). Common variants in trdn and calm1 are associated with risk of sudden cardiac death in chronic heart failure patients in chinese han population. PLoS One 10, e0132459. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Dai J et al. (2021). Significances of viable synergistic autophagy-associated cathepsin b and cathepsin d (ctsb/ctsd) as potential biomarkers for sudden cardiac death. BMC Cardiovasc. Disord. 21, 233. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Li H, Song S, Shi A & Hu S. (2024). Identification of potential lncrna-mirna-mrna regulatory network contributing to arrhythmogenic right ventricular cardiomyopathy. J. Cardiovasc. Dev. Dis. 11, 168. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Laczko R & Csiszar K. (2024). Lysyl oxidase (lox): functional contributions to signaling pathways. Biomolecules 10, 1093. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Mantri M et al. (2021). Spatiotemporal single-cell rna sequencing of developing chicken hearts identifies interplay between cellular differentiation and morphogenesis. Nat. communications 12, 1771. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Chen KH, Boettiger AN, Moffitt JR, Wang S & Zhuang X. (2015). Spatially resolved, highly multiplexed rna profiling in single cells. Science 348, aaa6090. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Wang X et al. (2018). Three-dimensional intact-tissue sequencing of single-cell transcriptional states. Science 361, eaat5691. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Stickels RR et al. (2021). Highly sensitive spatial transcriptomics at near-cellular resolution with slide-seqv2. Nat. biotechnology 39, 313–319. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Flamary R et al. (2021). Pot: Python optimal transport. J. Mach. Learn. Res. 22, 1–8. [Google Scholar]
- 71.Svensson V, Teichmann SA & Stegle O. (2018). Spatialde: identification of spatially variable genes. Nat. methods 15, 343–346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Shap. https://shap.readthedocs.io/en/latest/. Accessed: 2024-01-08.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary table S1: Gene sets corresponding to various cell-types, counts of each of the cell-types, and translation between human and mouse genes, related to Figure 4.
Supplementary table S2: SVR model’s (trained on top 2000 highly variable genes (HVG)) performance measured using Spearman’s correlation coefficient, related to Figure 6.
Supplementary table S3: Average ranks of the genes based on SVR coefficients, related to Figure 6.
Data Availability Statement
ST data of mouse hearts can be found in the Gene Expression Omnibus (GEO) under accession number GSE214611. The raw data files for the mouse DMD samples can be accessed in the GEO under accession number GSE199659. The processed data files are available on Zenodo at https://doi.org/10.5281/zenodo.7401196. The raw and processed sequencing data of the mouse TBI sample have been deposited in the GEO under accession number GSE236171 and are publicly available. The spatial transcriptomics of human heart samples can be accessed on CellxGene at https://cellxgene.cziscience.com/collections/8191c283–0816-424b-9b61-c3e1d6258a77 and in the Zenodo data archive at https://zenodo.org/record/6578047.
All original code has been deposited at https://zenodo.org/records/14941956 and is publicly available as of the date of publication. DOIs are listed in the key resources table.
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Deposited data | ||
| Spatial transcriptomics (Visium, 10x Genomics) data of healthy and injured mouse heart. Injuries include MI (permanent ligation of the LAD coronary artery), I/R-30 (30 minutes of ischemia followed by reperfusion), transaortic constriction, isoproterenol injection, and traumatic needle injection. | Calcagno et al. | GEO: GSE214611 |
| Spatial transcriptomics (Visium, 10x Genomics) data of Duchenne mouse models | Heezen et al. | GEO: GSE199659 |
| Spatial transcriptomics data of mouse traumatic brain injury | Pham et al. | GEO: GSE236171 |
| Spatial multi-omic map of human myocardial infarction | Kuppe et al. | DOI: 10.5281/zenodo.6578046 |
| Software and algorithms | ||
| Loupe Browser v7.0.1 | 10x Genomics | https://www.10xgenomics.com/support/software/loupebrowser/latest |
| Python v3.8.13 | Python Software Foundation | https://www.python.org/downloads/release/python3813/ |
| POT v0.8.2 | Flamary et al. | https://PythonOT.github.io/ |
| SHAP v0.42.0 | Su-In Lee’s lab at the University of Washington, and Microsoft Research. | https://shap.readthedocs.io/en/latest/ |
| Original codes | This manuscript | DOI: 10.5281/zenodo.14941956 |





