Skip to main content
Frontiers in Immunology logoLink to Frontiers in Immunology
. 2026 Jul 28;17:1856896. doi: 10.3389/fimmu.2026.1856896

Optimal transport analysis of high-dimensional flow cytometry data in immuno-oncology

Abida Sanjana Shemonti 1,†, Justin C Wang 2,†, Albert D Donnenberg 3,4,*, Bartek Rajwa 5, Patrick L Wagner 3,4,6, David L Bartlett 3,4,6, Bosko Popov 7, Evan T Alicuben 2,7,8, Vera S Donnenberg 2,7,8,*
PMCID: PMC13457739  PMID: 42582472

Abstract

Introduction

Advances in single-cell and spatial profiling have enabled detailed characterization of heterogeneous samples, but analyzing this data remains challenging in settings involving multiple comparisons. While tools like UMAP and t-SNE are valuable for visualization, their stochastic, parameter-sensitive nature limits their use in longitudinal comparisons, treatment group analysis, and multicenter trials. Although OT was first described in the 19th century, the Sinkhorn algorithm makes it computationally tractable for high-dimensional data. By directly comparing distributions of cellular states, OT provides reproducible measures of change in high-dimensional space. This framework is amenable to integration with machine learning, including deep generative models.

Methods

OT was applied to longitudinal data from a phase I trial of tocilizumab for cavitary malignancies (NCT 06016179). The current implementation makes use of expert-guided phenotypic population definitions and their relationships. An OT-based graph representation was created for baseline and follow-up samples. The graph layout was fixed across samples and computed from phenotypic relationships. In this implementation vertex radii are proportional to their relative abundance, allowing for rapid visual assessment of population-level increases and decreases. Graph edge thickness and color encode inter-population similarity based on the optimal transport (Sinkhorn) distance between marker expression distributions.

Results

This representation enabled rapid identification of populations undergoing substantial change, such as the CD8+/IFNɣ+ population, which decreased from 63% to 17% of CD8+ T cells following treatment. Population changes across all fluorescence parameters were encoded in the graph edit distance (GED), which captures changes in population abundance and phenotypic shifts in marker space.

Discussion

Future implementations can combine this expert-guided approach with unbiased clustering algorithms to enhance scalability and cross-platform harmonization. In our recently initiated clinical trials, we will apply OT to identify key shifts in tumor, immune, and stromal cell states, summarizing patient trajectories and quantitatively supporting predictive models of treatment response. Potential applications include quantifying residual disease after chemotherapy, tracking immune activation during immunotherapy, and linking host–microbiome interactions to disease progression. This approach overcomes the limitations of traditional, local-structure-optimized tools (UMAP or t-SNE) to provide a comprehensive, longitudinal view of tumor evolution and treatment response.

Keywords: clinical trial endpoints, immuno-oncology, malignant pleural effusion, omics, optimal transport, peritoneal ascites, Sinkhorn algorithm

1. Introduction

1.1. Visualization of intra-cavitary heterogeneity of malignant pleural effusions and peritoneal ascites

The pleural and peritoneal cavities are common sites for metastasis for epithelial cancers. Metastasis is often accompanied by accumulation of pleural or peritoneal fluids and carries a dire prognosis (1). Because there is currently no effective therapy, treatment consists of routine drainage to relieve symptoms of dyspnea and discomfort. Malignant pleural effusions (MPE) and malignant peritoneal ascites (MPA) are marked by striking heterogeneity (1–3), arising from a variety of primary tumors and shaped by an evolving and ultimately self-amplifying process in which tumor cells, immune cells and stromal cells interact to create a maladaptive and tumor promoting environment (1, 2). The availability of drained MPE or MPA, which are discarded as medical waste, has opened a window onto the biology of metastasis (1, 4–6), with ready access to tumor, infiltrating immune cells, as well as the cytokine-rich fluid that conditions the behavior of all populations.

Single-cell and spatial omics are capable of addressing MPE and MPA heterogeneity, revealing tumor cell states, T-cell functional capacity (7, 8) and myeloid polarization as they evolve during cytoreduction, regional therapies such as decortication and hyperthermic therapies (HIPEC), and systemic therapies (9–11). In this rapidly expanding data landscape, visualization tools, such as UMAP and t-SNE, that map high-dimensional data in 2-dimensional space, have become widely accepted tools for data exploration. Their utility has been demonstrated for visual hypothesis generation, quality control, and discerning patterns in complex datasets (12–14).

As useful as these visualization tools are for revealing hidden structures in high-dimensional data, inherent limitations preclude their use when the goal is to evaluate repeated sampling and longitudinal comparisons as is required when serially collected high-dimensional data sets are generated in clinical trials (15, 16). Here, we discuss a new alternative approach that focuses on capturing biological information as well as previously known domain information about the visualized data. This recently introduced methodology is manifold agnostic, which means it can use any notion of underlying similarity on the assumed manifold as input. Unlike earlier dimensionality reduction techniques, in which the visualization is directly driven by the supposedly recovered manifold, our implementation of optimal transport (OT) communicates biological notions of cellular heterogeneity, and how this heterogeneity is manifested in the phenotypic outcomes.

The approach utilizes the Sinkhorn algorithm, a computationally tractable, entropically regularized method that approximates earlier, less applicable OT methods. We assert that when integrated with contemporary machine learning (ML), these methodologies offer a systematic approach to the reliable, comparable quantification of clinically significant biological alterations (17–19). In the present work, we expand this perspective by combining OT-based quantification with a biologically interpretable, phenotype-aware graph framework that embeds expert-defined population relationships directly into the visualization. This framework draws on recently described methodology from Shemonti et al. (20), who demonstrated that the Sinkhorn distance provides a mathematically rigorous yet computationally tractable measure of inter-population similarity in high-parameter flow cytometry data. Building on this foundation, we pair these computational tools with an experimentally validated flow-cytometry panel and apply them to malignant pleural effusions and ascites collected before and after therapeutic intervention in an ongoing clinical trial.

Our aim is to illustrate how an OT-driven, domain-informed visualization approach can serve not only as an alternative to manifold learning, but as a reproducible and clinically meaningful method for summarizing immune and phenotypic structure in longitudinal samples. Specifically, this report (i) outlines conceptual limitations of manifold-learning visualizations for longitudinal clinical cytometry, (ii) describes an OT/Sinkhorn-based, phenotype-aware graph framework adapted from Shemonti et al. (20), and (iii) demonstrates its feasibility using high-parameter flow cytometry from malignant ascites collected during an ongoing intracavitary cytokine-blockade trial.

1.2. Limitations of UMAP and t-SNE in biological data analysis

Both UMAP and t-SNE methods are stochastic and, are sensitive to parameter configurations when applied in flow cytometry contexts. Empirically, they both excel in revealing local structure, inherent groupings and relationships (e.g., immune cell phenotypic populations). In t-SNE, inter-cluster distances are meaningless. Although UMAP can preserve some global structure, it is more suited to finding local neighborhoods than global arrangements (16, 21). In t-SNE, parameters like perplexity and early exaggeration dictate the extent of neighborhood preservation (13). In UMAP, parameters such as the number of nearest neighbors and the minimum distance affect the density of local structures formed (12). Both methods are sensitive to batch effects and minor alterations in preprocessing, normalization, batch correction, or the incorporation of cells from a new time point, can yield markedly different outcomes (16, 22, 23).

In both t-SNE and UPMAP, graph axes and cluster size are without meaning, and cluster formations can shift or elongate, even if the underlying biology remains the same (16). Most importantly, comparison of data files requires that they be concatenated and analyzed together to achieve a consistent visual representation. For multiple comparisons across time and between subjects, this becomes computationally unfeasible.

As a result, while these methods remain useful for exploration and hypothesis generation, they are fundamentally unsuited to multi-subject longitudinal settings (15, 16). Clinical trials necessitate methods that can produce consistent results across multiple sites and time points while remaining dependable even when the underlying sampling varies (24). In this context, a visually consistent plot may be useful, but it does not always reflect a quantitative measure of true biological change (15, 16).

Flow and mass cytometry data present distinct challenges for UMAP because the measurement process itself contravenes the algorithm’s fundamental assumptions. In cytometry, fluorescent or metal-tagged antibodies attach to cellular markers, and detectors quantify photons or ions during short acquisition periods. The photon-counting process adheres to super-Poissonian distributions (such as the negative binomial) and is affected by technical variables including antibody affinity, excitation intensity, and detector sensitivity. The resultant data exhibit overdispersion, are frequently discrete, and display heteroskedasticity, rendering them not directly comparable, even locally, using Euclidean distances, thereby undermining the geometric assumptions foundational to UMAP.

Cytometry panels often include markers that differ in abundance by several orders of magnitude. Highly expressed surface proteins such as CD45 on leukocytes may yield hundreds of thousands of photon counts, whereas low-abundance transcription factors generate only tens of counts above background. When UMAP applies Euclidean distance to such data, high-intensity markers dominate distance calculations, rendering low-intensity markers nearly invisible. This distortion violates UMAP’s assumption of a locally constant Riemannian metric because the effective geometry of the space changes dramatically across markers. The binary search that UMAP uses to optimize local σ (sigma) values cannot compensate for this scale disparity.

The super-Poissonian distributions naturally create highly non-uniform densities in measurement space. Cells lacking expression for multiple markers cluster densely near the origin, whereas rare phenotypes occupy sparse peripheral regions. UMAP’s Lemma 1 assumes uniform sampling on the underlying manifold to justify its k-nearest-neighbor-based metric learning. When this assumption fails, σ values vary widely between dense and sparse regions, producing a “fuzzy simplicial set” with inconsistent geometric meaning. Neighborhood size, therefore, reflects sampling density rather than biological similarity.

These mathematical violations result in reproducibility failures. When the same cytometry dataset is processed with different random initializations or preprocessing pipelines, the resulting embeddings have radically different topologies. Rare populations can merge or vanish, and adjusting the number of nearest neighbors parameter does not smooth out resolution but instead causes abrupt cluster fragmentation or fusion. Technical artifacts, such as batch effects or compensation (demultiplexing) errors, can further distort the embedding because UMAP lacks a mechanism to distinguish measurement noise from biological variation when its geometric assumptions fail. The resulting visualizations may appear interpretable, but they have no stable mathematical relationship to the underlying biological structure.

1.3. Optimal transport as a framework to interpreting high-dimensional biological data

Data visualization through optimal transport (OT) offers an alternative method for encoding and visualizing high-dimensional biological data. Rather than attempting to “discover” the underlying manifold in order to reduce apparent dimensionality, it relies on domain knowledge and prior identification of specific functional cell populations. This identification can be accomplished through any technique, including supervised learning or even manual annotation. This method assumes that the data alone are insufficient, that they do not convey all of the necessary information about the biological context, and that prior domain knowledge, which is absent from the raw measurements, must also be incorporated into the visualization.

The visualization of heterogeneous phenotypes is thus achieved not by simply estimating the manifold and projecting the high-dimensional data to a two-dimensional plot regardless of domain knowledge, but by computing a notion of dissimilarity between the identified functional (and biologically meaningful and interpretable) cell populations. The concept of similarity is encoded in the amount of “work” required to transform one distribution of cellular states into another (25, 26). In omics applications, the Wasserstein distance, also known as the Earth Mover’s distance, captures the minimal cost of reallocating probability mass between samples such as pre- and post-treatment single-cell profiles (25). This calculation uses a biologically meaningful ground metric, which may be based on gene-expression distance, latent space distances from a variational autoencoder, or multimodal kernels.

The result is a single value that reflects the magnitude of change, where a larger transport cost indicates a greater shift at the population level (25, 26). Importantly, optimal transport works within the original or learned high-dimensional space rather than in a potentially unstable two-dimensional embedding, and it compares entire distributions instead of trying to match individual cells one by one (25, 27).

Shemonti et al. (20) describe the computational framework that underpins this perspective and introduce a graph-based visualization method based on optimal transport theory. In that study, cell populations are user-defined by marker-expression profiles, and inter-population similarity is measured using the Sinkhorn distance. These distances are embedded in phenotype-aware graphs, which enable both visual and quantitative comparisons between samples. The framework allows for robust tracking of biological changes in high-dimensional flow cytometry data, including longitudinal shifts in immune cell states during therapy, and has been validated in clinical and public datasets.

Our work builds directly on this computational foundation but does not introduce a new optimal-transport algorithm. Instead, we adapt the Sinkhorn-based framework of Shemonti et al. (20) to a new biological setting of malignant ascites in peritoneal carcinomatosis and integrate it with a rigorously validated flow-cytometry pipeline. This provides a practical demonstration of how OT-driven, phenotype-aware graphs can be incorporated into longitudinal clinical studies where reproducibility and interpretability are essential.

2. Methods

2.1. Patients and sample collection

Flow cytometry data were obtained from a single patient enrolled in the Regional Immuno-Oncology Trial 2 (RIOT-2), a Phase I study evaluating intraperitoneal delivery of the IL-6 receptor antagonist tocilizumab (NCT06016179) (3). The patient had malignant ascites due to metastatic appendiceal carcinoma and underwent therapeutic drainage via an indwelling catheter placed as part of standard care. Ascitic fluid was collected either at the pretreatment baseline or at a follow-up visit after initiation of therapy. All samples were assigned accession numbers and de-identified prior to analysis.

Ascites were drained aseptically into 1000 mL vacuum bottles (Becton Dickinson, SKU/REF 50-7700). Upon receipt at the Allegheny Health Network Cellular Therapy Laboratory, preservative-free sodium heparin was added to a final concentration of 10 U/mL. After aliquots were removed for bacterial and fungal cultures, penicillin G (40 µg/mL) and gentamicin (100 U/mL) were added. Samples were filtered through a 170–260 μm filter (Baxter, 2C8750 Blood/Solution Set), and a de-identified aliquot was transferred to the UPMC Hillman Cancer Center research laboratory for downstream flow cytometry (MTA00011955). This clinical setting is an ideal test case for OT-based visualization, as repeated sampling of the peritoneal compartment is feasible, yet traditional manifold-learning approaches fail to provide stable, interpretable longitudinal comparisons.

2.2. Institutional review board protocol and informed consent

This study was conducted in accordance with the Declaration of Helsinki and approved by the Allegheny Health Network Cancer Institute Institutional Review Board for the Regional Immuno-Oncology Trial 2 (RIOT-2), a Phase I study of intrapleural and intraperitoneal tocilizumab (IRB RC# 2023-137; NCT06016179). Flow cytometry analyses in this report were performed on de-identified malignant ascites specimens collected under this approved protocol. All samples were anonymized prior to research use in accordance with institutional guidelines.

Written informed consent was obtained from the patient prior to enrollment in the RIOT-2 clinical trial and prior to collection of malignant ascites for research purposes, including exploratory flow-cytometric and computational analyses. All samples used in this study were de-identified before analysis in accordance with the approved protocol.

2.3. Data availability statement

Data will be made available to qualified investigators upon request to the corresponding author.

2.4. Flow cytometry for surface marker expression

Sample preparation and staining for surface markers followed protocols previously established by Donnenberg and colleagues for analysis of immune subsets and tumor-associated stromal populations (28, 29). Malignant ascites cells were washed in staining buffer (4% calf serum, 2 mM EDTA in Ca2+/Mg2+-free PBS, pH 7.2, 4 °C). Cell pellets were incubated with neat decomplemented mouse serum (10 µL, 5 min, 4 °C) to block Fc-mediated nonspecific binding, followed by staining with 2 µL each of fluorescently conjugated antibodies targeting surface markers. A complete list of antibodies, fluorochromes, catalog numbers, and vendors appears in Table 1.

Table 1.

Antibodies used for flow cytometry of surface proteins and intracellular cytokines.

Analyte Fluorochrome Catalog # Vendor
CD3 BV711 563725 BD Horizon
CD4 BV785 317442 BioLegend
CD8 ECD 660478 Beckman Coulter
CD14 PerCP Cy5.5 325622 BioLegend
CD45 BUV395 563792 BD Horizon
CD56 PE Cy5 IM2654U Beckman Coulter
CD19 APC-A700 A78837 Beckman Coulter
CD326 APC Cy7 324234 BioLegend
DNA DAPI D1306 Invitrogen
TNFα FITC 502915 BioLegend
IL-10 PE 506804 BioLegend
lFNγ PECy7 506518 BioLegend
IL-2 APC 500310 BioLegend

CD, cluster of differentiation; BV711 and BV785, Brilliant Violet 711 and Brilliant Violet 785; BD, Becton Dickinson; ECD, phycoerythrin–Texas Red tandem fluorochrome; PerCP-Cy5.5, peridinin–chlorophyll-protein–Cyanine5.5 tandem fluorochrome; BUV395, Brilliant Ultraviolet 395; PE-Cy5, phycoerythrin–Cyanine5 tandem fluorochrome; APC-A700, allophycocyanin–Alexa Fluor 700 tandem fluorochrome; APC-Cy7, allophycocyanin–Cyanine7 tandem fluorochrome; DNA, deoxyribonucleic acid; DAPI, 4′,6-diamidino-2-phenylindole; TNF, tumor necrosis factor; FITC, fluorescein isothiocyanate; IL, interleukin; PE, phycoerythrin; IFNγ, interferon gamma; PECy7, phycoerythrin–Cyanine7 tandem fluorochrome; APC, allophycocyanin.

After staining, cells were fixed in 2% methanol-free formaldehyde (Polysciences, Cat. No. 0401A) for 20 minutes, permeabilized with 0.05% saponin (Coulter) in staining buffer, washed, and resuspended to 5 × 106 cells/mL. Prior to acquisition, DAPI (Invitrogen D1306, 10 µg/mL) was added for DNA content assessment and identification of cycling CD3+ cells.

All samples were acquired on a Fortessa SORP flow cytometer (BD Biosciences, Hillman Cancer Center Cytometry Core). Daily instrument calibration was performed using CS&T beads (BD Biosciences, Cat. No. 650621). PMT voltages were adjusted to predetermined target channels using the seventh peak of 8-peak Rainbow Calibration Particles (Spherotech RCP-30-5A). Spectral compensation was derived from FITC, PE (BD 349502), and APC (BD 340487) Calibrite beads, single-stained BD CompBeads (anti-mouse Igκ; BD 51-90-9001229) for tandem dyes, and unstained DAPI-only cells.

2.5. Flow cytometry for intracellular cytokine profiling

Intracellular staining for IFNγ, TNFα, IL-10, and IL-2 followed methods previously described for cytokine profiling in pleural infiltrating T cells (1). Cells were washed and resuspended in complete medium (RPMI-1640 supplemented with 10% FBS) at 2 × 106 cells/mL. A 4 mL suspension was plated in 15 mL polypropylene tubes and incubated for 1 hour at 37 °C with tetradecanoylphorbol-13-acetate (TPA, 0.05 µM) and ionomycin (1 µM) (30). Cells were then washed, resuspended in complete medium, and treated with Brefeldin A (1 µg/mL) for two additional hours at 37 °C to inhibit cytokine secretion.

Following stimulation, cells were washed and blocked with decomplemented mouse serum (10 µL, 5 min, 4 °C), stained for surface markers as above, fixed with 2% methanol-free formaldehyde, permeabilized with 0.05% saponin, and washed. Cells were split into two equal aliquots, one of which was stained with cytokine-specific antibodies (listed in Table 1). Samples were then treated with PureLink RNase A (Invitrogen, 1:400, 10 min, 37 °C), washed, resuspended to 5 × 106 cells/mL, and stained with DAPI prior to acquisition on the Fortessa SORP cytometer.

Cytokine gating thresholds were defined using samples stained only for extracellular markers (no intracellular cytokines). Data analysis was performed using VenturiOne software (Applied Cytometry Systems, V7.6.0.47.X64). DAPI staining enabled estimation of the proportion of CD3+ cells in cycle (DNA >2N). The full gating strategy is shown in Figure 1.

Figure 1.

Flow cytometry gating strategy diagram showing a sequence of scatter plots for cell population classification. Top row displays sequential gating from ungated to specific populations labeled A to F using parameters like FSC-A, FSC-H, SSC-A, DAPI Log, CD3 BV711, and CD45 BUV395, with population percentages indicated in each gate. Arrows connect to lower plots splitting cells into CD3 positive and CD3 negative groups. Left lower section (CD3+) shows plots for CD4 BV785 versus CD8 ECD with cell frequency percentages in colored boxes. Right lower section (CD3-) shows plots for CD19 APC Cy7 versus CD14 PerCPCY5_5 and CD8 ECD versus CD56 PE Cy5, with percentages marked for each subpopulation.

Manual Gating Strategy. From left to right, top to bottom: Forward light scatter (FSC) area versus time is used to detect fluidics perturbations that might result in artifact (here none were detected). (A) FSC pulse area versus FSC pulse height is used to identify and gate out events with pulse area too great for the pulse height (Outside gate B, doublets and multiples and small debris). (B) FSC area versus side light scatter (SSC) area is used to gate our small debris (Outside gate C). (C) DAPI versus FSC is used to gate out events with less then 2n DNA (note that both 2n and 4n populations are visible within the D gate). (D) CD45 and CD3 are used to distinguish between T cells (E gate) and non-T leukocytes (F gate). T cells are further gated on CD8+ T cells (G gate) and CD4+ T cells (H gate). CD14 vs CD19 is used to identify B cells (I gate), macrophages (J gate) and NK cells (K gate) among non-T leukocytes. Cytokines were measured as “outcomes” on these “classifier” populations. This gating strategy was used for all baseline and follow-up samples.

2.6. Antibody reagents

All antibody reagents are summarized in Table 1, including marker, fluorochrome, catalog number, and vendor. Surface markers included lineage and functional proteins (CD3, CD4, CD8, CD14, CD19, CD45, CD56, CD326), and intracellular cytokine antibodies targeted IFNγ, TNFα, IL-10, and IL-2. DAPI was used for DNA content.

2.7. Data preprocessing and cell population identification

Computational preprocessing drew on established approaches for flow cytometry data annotation and semi-automated gating, consistent with previously published pipelines. Data were imported from FCS files and quality-controlled by removing debris and margin events. Surface and intracellular markers were compensated and transformed as described above.

Population identification relied on a semi-automated, knowledge-driven gating schema derived from the explicit fluorescence-intensity boundaries defined during manual gating of representative samples. This approach parallels the strategy used in Shemonti et al. (31), where population labels, whether manually gated, heuristically defined, or produced by automated methods, serve as the required input for downstream optimal transport computations. Each cell was assigned to a biologically interpretable phenotype defined by marker expression (e.g., CD3+CD4+ T cells, CD3+CD8+ T cells, NK cells, CD14+ monocytes). These population definitions form the categorical structure on which optimal transport distances are computed.

As described by Shemonti et al. (31), this framework is fully agnostic to the upstream labeling method: the core requirement is that each cell belongs to a defined population. This preserves interpretability and ensures compatibility with domain knowledge and prior cytometric expertise.

2.8. Optimal transport computation and graph-based visualization

Inter-population dissimilarity was quantified using the entropically regularized optimal transport (OT) framework described by Shemonti et al. (20) and based on the Sinkhorn distance. A plain-language description of the OT approach to high-dimensional flow cytometry data is provided in Supplementary Materials: Optimal Transport Basics.

For any two populations C1 and C2, each cell was treated as a unit mass normalized to total population size (1/n1 and 1/n2). The ground cost matrix was computed as the squared Euclidean distance between marker-expression vectors, consistent with the formulation presented in the preprint.

The Sinkhorn–Knopp algorithm was used to compute regularized transport plans, enabling efficient scaling to large cell counts. The resulting Sinkhorn distances reflect high-dimensional distributional differences between cell populations and form the quantitative backbone of the visualization framework.

To construct interpretable sample-level visualizations, we adopted the graph-based approach described in Shemonti et al. (20). Each cell population corresponds to a vertex, with vertex size proportional to its frequency within the sample. Edges encode inter-population Sinkhorn distances, visualized through thickness and color intensity. Vertex positions were determined by fixed, phenotype-aware layouts using Hamming distances between expert-defined phenotype strings, ensuring reproducible and biologically meaningful geometry.

This fixed layout allows multiple samples to be compared directly, and changes in edge weights reflect shifts in marker-expression distributions across populations. As in the preprint, these graphs serve as compact, interpretable “fingerprints” of sample structure. Graph edit distance (GED), with biologically informed vertex and edge substitution costs, can be used to quantify dissimilarity between samples and is well-suited for longitudinal comparisons in future multi-timepoint analyses.

2.9. The Sinkhorn algorithm and machine learning integration

While our present application focuses on a single-patient proof-of-concept, the same computational framework scales to multi-patient cohorts and supports machine-learning workflows designed for longitudinal PC trials. The Sinkhorn algorithm makes optimal transport feasible at the scale required for processing large datasets, such as those generated in clinical trials, where serial observations are made to track responses (19, 32, 33). By adding entropic regularization, the classical transport problem becomes a convex, well-conditioned objective that can be solved efficiently via parallel matrix scaling (19, 33). This approach performs well on GPUs and scales to very large datasets with appropriate batching (19, 32).

Sinkhorn divergences also help correct for bias introduced by regularization, which allows meaningful comparisons even when sample sizes are unequal, as is often the case in cells taken from serial tumor or ascites sampling (17, 18). In addition, optimal transport barycenters can summarize groups of patients or define reference atlases of peritoneal carcinomatosis microenvironments, such as immune-active, myeloid-dominant, or mesothelial-fibrotic states (27, 34). Each patient’s samples can then be measured against these references, turning endpoints into shifts in transport distance rather than movements of clusters in a two-dimensional plot (25, 27).

Machine learning fits naturally into this framework. Deep generative models such as variational autoencoders provide denoised, batch-corrected latent spaces that preserve biological structure (35). Optimal transport then supplies geometry-aware distances, alignments, and trajectories within these spaces (36). This view also supports domain adaptation, where transport minimizes differences between sites or platforms, such as CyTOF, single-cell RNA sequencing, and spatial proteomics, thereby improving comparability across centers in multicenter trials (27, 36). Transport plans can identify which cell states change most after therapy, highlighting, for example, shifts in proliferative EpCAM-positive tumor cells, interferon-responsive macrophages (37, 38), CXCL13-positive Tfh-like cells, or mesothelial cells with high mesothelin expression (9, 25).

These features can then be used as inputs to predictive models of treatment response, progression-free survival, or the transition from an immune-cold to an immune-hot microenvironment (39). This approach allows machine learning models to learn from distributions of cell states rather than from potentially unstable two-dimensional representations (35).

3. Results

3.1. Applications for sequential drainage of malignant pleural effusions and ascites: an example from peritoneal carcinomatosis MPA

In peritoneal carcinomatosis trials, optimal transport and machine learning can be applied directly to clinically meaningful endpoints (31, 40, 41). In the context of immunotherapy or cytokine-targeted treatments, optimal transport can define mechanistic distributional shifts (42–44), such as movement from a patient’s baseline toward an atlas enriched for activated T cells, transitions into or out of macrophage programs linked to immunosuppression (45), or redistribution among epithelial EMT states that drive peritoneal spread (39, 46–48). For cytoreduction, intrapleural and hyperthermic intraperitoneal chemotherapy studies, transport distances can quantify how residual disease evolves after surgery, capturing shifts toward quiescent or stress-response phenotypes (9, 49, 50).

When the peritoneal microbiome is incorporated, multimodal transport, or machine learning co-embedding followed by transport, can identify coordinated changes between microbial communities in ascites and host immune states such as myeloid polarization or cytokine signaling (27). Unlike two-dimensional embeddings, these measurements can be applied more consistently across time points and trial centers, providing a stable and reproducible way to evaluate biological change (25, 27). These analyses can quantify distributional shifts, but links to clinical response require prospective validation (25, 51). Together, these examples illustrate how transport-based approaches provide a practical and reproducible framework for clinical research in PC (25, 27, 36). Although these applications extend beyond the scope of the present demonstration, they outline realistic future directions for integrating OT-based cytometry endpoints into clinical trials.

In this demonstration, we have applied t-SNE, UMAP and OT to malignant ascites samples collected before and after an immunotherapeutic intervention. Cells recovered from the ascitic fluid were stimulated in short-term culture with TPA plus ionomycin (TPA+I) and stained to measure intracellular cytokines in different immune cell populations. Figure 1 (Baseline gating strategy) shows the gating strategy used for conventional flow cytometric analysis, creating classifier populations (52) (CD4+ and CD8+ T cells, B cells, macrophages, NK cells) in which cytokine outcomes (IL-2, IL-10, interferon-γ, and TNF-α) are measured. Figure 2 is gated on the 5 classifier populations and shows the fluorescence intensity of the channels that are used to measure the four cytokines (fluorochromes minus outcomes controls (52)).

Figure 2.

Twenty flow cytometry scatter plots arranged in a five-by-four grid, each labeled for specific immune cell populations including CD8+ T cells, CD4+ T cells, CD19+ B cells, CD14+ TAM, and CD3-CD56+ NK cells. Axes are labeled SSC-A and with various fluorophore combinations. Colored gates, percentages, and cell subset codes are shown per plot. The fifth column in each row displays bivariate plots for polyfunctional cells characterized by IL10 and TNFα expression. Color gradients differentiate cell types for visual clarity.

Baseline fluorescence minus outcomes plots. A separate tube, stained for the “classifier” populations, but not cytokine “outcomes”, was used to determine the boundaries between cytokine positive and negative populations. CD4+ T cells are shown in the top histograms (red border.) and CD8+ T cells are shown in the bottom histograms (blue borders). Note autofluorescence or non-specific staining in the high SSC population. The same strategy was used for B-cell, macrophage and NK-cell classifier populations.

The regions were drawn to include ≤ 1% of events and were applied to the data files stained for both classifier and outcome analytes (Figures 3, 4). Figure 3 shows intracellular cytokines in the ascites sample collected at baseline. At baseline, TPA+I stimulation elicited modest TNF-α responses and robust IFN-γ responses in CD4+ and CD8+ T cells. Macrophages and NK cells also produced IFN-γ. Following tocilizumab treatment IL-2 appeared to be upregulated in CD4+ and CD8+ T cells, TNFα was upregulated in macrophages, and IFN was downregulated in T- and NK cells but upregulated in macrophages (Figure 4).

Figure 3.

Multipanel scientific figure showing flow cytometry dot plots of immune cell populations, each labeled by cell type and functional marker (IL10, TNFa, IL2, IFNg), with quantified percentages and multicolor backgrounds for easy distinction. Polyfunctional cell plots identify IL10 and TNFa double-positive populations for CD8+ T cells, CD4+ T cells, B cells, TAM (tumor-associated macrophages), and NK (natural killer) cells, with quadrants showing distribution percentages. Each panel is annotated with cell marker and cytokine, assisting comparison of cytokine production across cell types.

Baseline cytokine outcomes. Histograms are in the same order as the fluorescence minus outcomes control. Rows from top to bottom: CD4+ T cells (red border); bottom row: CD8+ T cells (blue border), B cells (pink border), macrophages (gold border), NK cells (light blue border). Histograms to the far right are gated on the IL-10 negative, TNF-alpha positive population and are used to determine polyfunctionality. Note that significant populations of CD4+ and CD8+ IL-10-/TNF+ T cells also secrete IFN-gamma.

Figure 4.

Grid of twenty representative flow cytometry dot plots displays cell populations gated for IL10, TNFa, IL2, and IFNg cytokine expression in CD8+ T, CD4+ T, CD19+ B, CD14+ TAM, and CD3-CD56+ NK cells, with corresponding percentages annotated for each subset. Polyfunctionality plots on the right summarize dual and triple cytokine expression for each cell type.

Follow-up cytokine staining. Histograms are arranged in the same order as for baseline cytokine staining. IFN-gamma secretion was markedly down-regulated and IL-2 secretion is increased in CD4+ and CD8+ populations compared to baseline, while TNF-alpha and IFN-gamma secretion increased in macrophages.

Figure 5 shows the data from Figures 3, 4, projected as t-SNE plots. Data were gated on CD45+ events. Events are colored according to their analyte expression by conventional flow cytometric analysis. The plots are dominated by uncategorized (i.e., cytokine negative) CD8+ (purple), CD4+ (green) T cells, and B cells (hot pink). Positional and size changes of these clusters after therapy are without meaning. Most cytokine+ populations are inapparent, with the exception of CD8+/IFN-γ+ cells (light purple), which comprised 63% and 17% of CD8+ T cells at baseline and follow-up, respectively. This decrement in IFN-γ+ T cells after therapy was not captured in the t-SNE plot, because cluster area does not represent population prevalence.

Figure 5.

tSNE plots show high-dimensional immune cell data at baseline and follow-up, with colored clusters representing different immune cell types and subtypes. Smaller tSNE plots below display distribution changes for CD8+ cells, CD4+ cells, B-cells, macrophages, and other cells at both time points, highlighting differences by cell type across conditions. A color-coded legend at the bottom identifies each immune cell subtype shown in the plots.

T-SNE visualization of baseline and follow-up flow cytometry samples. T-SNE embeddings of baseline and follow-up flow cytometry datasets are shown. A shared low-dimensional reference embedding was first computed from a combined, cell population density-aware subsample of both datasets to ensure representation of all annotated cell populations while enabling scalable computation. All cells from the full baseline and follow-up datasets were then projected into this common T-SNE space. T-SNE was generated in R (version 4.4.1) using the Rtsne package on macOS Sequoia (version 15.4.1), with a perplexity of 7. (A) T-SNE embedding of the full baseline dataset, with all cells projected into the shared reference space. (B) T-SNE embedding of the full follow-up dataset, projected into the same reference space. (C) Baseline T-SNE embeddings restricted to selected major immune populations (CD8+ T cells, CD4+ T cells, B cells, macrophages, and uncategorized cells), shown separately to highlight population-specific structure within the baseline sample. (D) Follow-up T-SNE embeddings of the same selected immune populations (CD8+ T cells, CD4+ T cells, B cells, macrophages, and uncategorized cells), as in (C), facilitating qualitative comparison of population distributions between baseline and follow-up conditions.

Figure 6 shows the same data in UMAP plots. As in the t-TNE plots, the graphs are dominated by CD4+, CD8+ T cells and B cells, with the minor cytokine+ populations obscured. However, homologous populations from the two samples are greatly displaced in 2D space (e.g., the B-cell clusters, shown in hot pink). This artifact is due to UMAP’s susceptibility to subtle batch effects. Within samples, some global structure is preserved, as CD4+/IFN-γ+ cells (gold) are largely clustered within CD4+ cells, and CD8+/IFN-γ+ cells (light purple) are clustered within CD8+ cells. Overall, these plots entirely fail to capture quantitative therapy-related changes in cytokine production and are insensitive to minor cytokine+ populations.

Figure 6.

Figure contains UMAP plots showing cell cluster distributions from baseline and follow-up data, with each point colored by cell type. Lower panels display subsets for CD8+, CD4+, B-cells, macrophages, and other cells for both time points, further divided by functional markers. Legends indicate marker and cell type by color.

UMAP visualization of baseline and follow-up flow cytometry samples. UMAP embeddings of baseline and follow-up flow cytometry datasets are shown using a shared reference-based embedding strategy (Figures 3, 4). A combined, cell population density-aware subsample from both datasets was used to compute a joint UMAP reference, preserving representation of all annotated cell populations while reducing computational complexity. The complete baseline and follow-up datasets were subsequently projected into this shared UMAP space. UMAP was computed in R (version 4.4.1) using the uwot package on macOS Sequoia (version 15.4.1), with n_neighbors = 15. (A) UMAP embedding of the full baseline dataset, with all cells projected into the shared reference space. (B) UMAP embedding of the full follow-up dataset, projected into the same reference space. (C) Baseline UMAP embeddings restricted to selected immune populations (CD8+ T cells, CD4+ T cells, B cells, macrophages, and uncategorized cells), shown separately to emphasize population-specific organization within the baseline sample. (D) Follow-up UMAP embeddings of the same immune populations (CD8+ T cells, CD4+ T cells, B cells, macrophages, and uncategorized cells), as in (C), enabling visual assessment of changes in population structure and distribution between baseline and follow-up samples.

Figure 7 illustrates the optimal transport–based graph representation of the same baseline (Figure 3) and follow-up (Figure 4) samples. Our implementation of OT analysis begins with manual gating, as in Figures 1–3. The graph layout is fixed across samples and computed from phenotypic relationships among cell populations, ensuring positional consistency and enabling direct visual comparison between time points. Vertices represent cell populations, with vertex radii proportional to their relative abundance with respect to the immediate parent population defined during manual gating. This encoding allows rapid visual assessment of population-level increases and decreases between baseline and follow-up samples. In addition, the shared layout enables direct graph overlay: in the third panel of Figure 7, baseline populations are shown with dotted outlines and follow-up populations with solid outlines, facilitating intuitive identification of treatment-associated changes.

Figure 7.

Three network graphs compare immune cell cluster relationships using Sinkhorn distance for baseline data, follow-up data, and their differences. Nodes represent specific cell types with percentages, node size indicates cluster size, and edge colors show edge weight. A legend and color scale denote edge weights and timepoints, while node labels reflect cell type changes and proportions.

OT analysis visualization of data provided in Figures 3, 4. Sinkhorn distance-based visual summary of changes in classifier populations and outcomes before (baseline) and after (follow-up) intraperitoneal administration of tocilizumab. Illustrations of the flow cytometry sample graphs at baseline (A) and one week after completion of 4 therapeutic doses (B). The vertices (circles) represent individual classifier populations identified manual gating. Their radius is proportional to their relative proportion among total valid events (CD45+ positive events identified by the gating strategy). Vertex positions are determined by pairwise phenotype distances, so phenotypically similar populations lie closer together. Edges connect the vertices. Their weights (but not lengths), indicated by color and line thickness, encode the inter-population Sinkhorn distances: smaller distances indicating greater similarity in marker expression produce thicker, more vividly colored edges, while larger distances yield thinner, paler edges. Bottom panel: Vertices from the two samples are distinguished by their outlines: The baseline sample’s vertices appear as dotted circles and the follow-up’s as solid circles. When a population’s proportion increases, the solid circle encloses the dotted one; when it decreases, the dotted circle is larger. Numeric labels at each vertex indicate the percentage change in that cell population. Edges encode the absolute change in Sinkhorn distance between the two samples. A positive edge weight denotes increased similarity (i.e., a decrease in Sinkhorn distance), whereas a negative weight denotes decreased similarity (i.e., an increase in Sinkhorn distance).

Edge thickness and color encode inter-population similarity based on the optimal transport (Sinkhorn) distance between marker expression distributions. Thicker and more saturated edges indicate greater phenotypic similarity, whereas thinner and paler edges correspond to increased dissimilarity. In the overlay graph, edges represent differences in Sinkhorn distance between baseline and follow-up samples, highlighting population pairs exhibiting the most pronounced phenotypic shifts.

This representation enables rapid identification of populations undergoing substantial change, such as the CD8+/IFN-γ+ population (light purple), which decreased from 63% to 17% of CD8+ T cells following treatment. Beyond visualization, encoding flow cytometry samples as optimal transport–based graphs allows quantitative comparison using graph-theoretic metrics. In particular, graph edit distance (GED) captures both changes in population abundance (via vertex size) and phenotypic shifts in marker space (via Sinkhorn distances). While individual GED values are not directly interpretable in isolation, pairwise GEDs across a patient cohort enable downstream statistical analyses, such as pseudo-ANOVA, to assess treatment effects and compare therapeutic strategies.

4. Discussion

UMAP and t-SNE remain useful tools for data visualization and exploratory data analysis. They can be used to perform rapid dimensionality reduction, which is helpful in the identification of unique populations that can only be visualized in multidimensional space, for revealing integration artifacts, and generating biologically grounded hypotheses. In multiparameter flow cytometry, such the data set provided here, they are useful for exploratory analysis of the tumor microenvironment, immune responses, and host–microbiome interactions (12, 13). However, employing these methods as a preprocessing step prior to unsupervised clustering is methodologically problematic. UMAP’s theoretical framework assumes data is locally uniformly distributed on the underlying manifold (12), which is fundamentally incompatible with biological single-cell data where cells form discrete, locally clustered populations. When this assumption is violated (as it inherently is with clustered cell types) UMAP introduces substantial distortions: it fails to preserve majority of nearest neighbors, distorts global cell-type relationships (correlation ≤0.4 with high-dimensional space), and inflates distance ratios 4–200 fold (53). These distortions make the 2D coordinates unreliable for clustering algorithms that depend on accurate distance metrics. For these reasons, it has been argued that clustering should be performed in the original high-dimensional space or in a dimensionality-reduced space that preserves global structure (e.g., PCA), with UMAP/t-SNE used solely for visualization of the resulting clusters (12).

As a result, while manifold learning’s limitations do not affect data exploration, they can limit their usefulness for longitudinal clinical studies in the flow cytometry field, where the objective is to quantify biologically meaningful changes occurring over time (16, 31, 54). In contrast, optimal transport-based analysis and visualization, which are efficiently implemented with the Sinkhorn algorithm and combined with machine learning, provide the stability and comparability that trials require (27, 36). As cancer research becomes more multi-omic and multicenter, with longitudinal follow-up, transport-based endpoints can transform the complexity of single-cell and spatial data into measures that are clinically interpretable (25, 27, 36). This approach emphasizes the underlying biological changes represented by data visualizations rather than their appearance (25, 27).

In this manuscript, we provide a proof-of-concept example demonstrating how integrating OT with phenotype-aware graph construction yields a reproducible, biologically interpretable representation of flow-cytometry samples collected from a patient undergoing intraperitoneal cytokine blockade (1, 2). This approach is directly informed by the methodological advances described by Shemonti et al. (20), particularly the use of the Sinkhorn distance as a robust statistical measure of inter-population similarity. By anchoring the visualization in explicit phenotype definitions rather than stochastic manifold embeddings, we recover a structure that faithfully reflects both proportional shifts in cell populations and distributional changes in marker expression.

Because the framework accommodates both longitudinal sampling and heterogeneous clinical datasets, it provides a foundation for trial-ready cytometry endpoints capable of capturing immune modulation, treatment response, and evolving tumor–immune interactions in PC. As OT-based methodologies continue to mature, they offer a path toward unified, reproducible, and quantitatively interpretable visualization standards for high-parameter cytometry in clinical research.

In this proof-of-concept analysis, we demonstrate that optimal-transport–based, phenotype-aware graph representations can reliably summarize high-parameter flow-cytometry profiles from malignant effusions and ascites and reveal biologically interpretable shifts that are obscured in conventional manifold-learning visualizations. We are currently applying this methodology to our longitudinal clinical trial data in malignant effusions and ascites (1, 2) and anticipate further refinements facilitated by integration with existing flow cytometry packages. The small dataset presented here represents one panel among four currently in use to monitor two Phase I clinical trials of an adoptive cellular therapeutic (RIOT4A, NCT07192900, RIOT4B, NCT07443020), the others being T-cell activation, immune checkpoint expression, and T-cell differentiation. Future applications include expanding the OT framework to these panels in the RIOT2 patient cohorts, comparing it to conventional analysis, and evaluating the extent to which OT-based endpoints track therapeutic response.

Acknowledgments

We would like to thank the patients who participate in our research studies and clinical trials.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by grants BC210533, BC211396, CA230972 and BC240179 from the Congressionally Directed Medical Research Program of the Department of Defense. The Hillman Cancer Center Flow Cytometry Core is supported by Cancer Center Support Grant P30CA047904.

Edited by: Michal Chovanec, Comenius University, Slovakia

Reviewed by: Yingkun Xu, Shandong University, China

Dominika Rychtáriková, National Cancer Institute, Slovakia

The following abbreviations are used in this manuscript:

MPE, Malignant pleural effusions; MPA, Malignant peritoneal ascites; PC, Peritoneal carcinomatosis; HIPEC, Hyperthermic intraperitoneal chemotherapy; UMAP, Uniform Manifold Approximation and Projection; t-SNE, t-distributed Stochastic Neighbor Embedding; OT, Optimal Transport; ML, Machine Learning; GPU, Graphics Processing Unit; CyTOF, Cytometry by Time of Flight; RNA, Ribonucleic acid; EpCAM, Epithelial Cell Adhesion Molecule; CXCL13, C-X-C Motif Chemokine Ligand 13; Tfh, T follicular helper; EMT, Epithelial–mesenchymal transition.

Data availability statement

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding author.

Ethics statement

The studies involving humans were approved by Allegheny Health Network Cancer Institute Institutional Review Board. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

JW: Writing – original draft, Writing – review & editing. AS: Methodology, Software, Writing – original draft, Writing – review & editing. AD: Writing – original draft, Writing – review & editing, Conceptualization, Data curation, Formal analysis. BR: Writing – original draft, Writing – review & editing. PW: Writing – review & editing, Investigation, Resources. DB: Investigation, Resources, Writing – review & editing. EA: Investigation, Resources, Writing – review & editing. BP: Writing – original draft, Resources. VD: Investigation, Resources, Writing – review & editing, Conceptualization, Formal analysis, Funding acquisition, Methodology, Project administration, Supervision, Visualization, Writing – original draft, Data curation.

Conflict of interest

Author AS was employed by the company Miftek Corporation.

The remaining author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fimmu.2026.1856896/full#supplementary-material

DataSheet1.pdf (146.7KB, pdf)

References

  • 1. Donnenberg VS, Luketich JD, Sultan I, Lister J, Bartlett DL, Ghosh S, et al. A maladaptive pleural environment suppresses preexisting anti-tumor activity of pleural infiltrating T cells. Front Immunol. (2023) 14:1157697. doi:  10.3389/fimmu.2023.1157697 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Donnenberg VS, Luketich JD, Popov B, Bartlett DL, Donnenberg AD. A common secretomic signature across epithelial cancers metastatic to the pleura supports IL-6 axis therapeutic targeting. Front Immunol. (2024) 15:1404373. doi:  10.3389/fimmu.2024.1404373 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Park H, Lewis C, Dadgar N, Sherry C, Evans S, Ziobert S, et al. Intra-pleural and intra-peritoneal tocilizumab therapy for managing Malignant pleural effusions and ascites: The Regional Immuno-Oncology Trial (RIOT)-2 study protocol. Surg Oncol Insight. (2024) 1:100045. doi:  10.1016/j.soi.2024.100045 38826717 [DOI] [Google Scholar]
  • 4. Donnenberg ADDVS. The secretome of Malignant pleural effusions: Clues to targets of therapy. Leukemia Res. (2019) 85:S11–2. doi:  10.1016/s0145-2126(19)30222-x [DOI] [Google Scholar]
  • 5. Donnenberg VS, Wagner PL, Luketich JD, Bartlett DL, Donnenberg AD. Localized intra-cavitary therapy to drive systemic anti-tumor immunity. Front Immunol. (2022) 13:846235. doi:  10.3389/fimmu.2022.846235 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Wagner PL, Knotts CM, Donneberg VS, Dadgar N, Pico CC, Xiao K, et al. Characterizing the immune environment in peritoneal carcinomatosis: Insights for novel immunotherapy strategies. Ann Surg Oncol. (2023) 31(3):2069–77. doi:  10.1245/s10434-023-14553-6 [DOI] [PubMed] [Google Scholar]
  • 7. Liang R, Zhu X, Lan T, Ding D, Zheng Z, Chen T, et al. TIGIT promotes CD8+T cells exhaustion and predicts poor prognosis of colorectal cancer. Cancer Immunol Immunother: CII. (2021) 70:2781–93. doi:  10.1007/s00262-021-02886-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Ostroumov D, Duong S, Wingerath J, Woller N, Manns MP, Timrott K, et al. Transcriptome profiling identifies TIGIT as a marker of T-cell exhaustion in liver cancer. Hepatol (Baltimore Md). (2021) 73:1399–418. doi:  10.1002/hep.31466 [DOI] [PubMed] [Google Scholar]
  • 9. Huang X-Z, Pang M-J, Li J-Y, Chen H-Y, Sun J-X, Song Y-X, et al. Single-cell sequencing of ascites fluid illustrates heterogeneity and therapy-induced evolution during gastric cancer peritoneal metastasis. Nat Commun. (2023) 14:822. doi:  10.1038/s41467-023-36310-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Launonen I-M, Vähärautio A, Färkkilä A. The emerging role of the single-cell and spatial tumor microenvironment in high-grade serous ovarian cancer. Cold Spring Harbor Perspect Med. (2023) 13:a041314. doi:  10.1101/cshperspect.a041314 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Schoenfeld AJ, Lee SM, Doger de Spéville B, Gettinger SN, Häfliger S, Sukari A, et al. Lifileucel, an autologous tumor-infiltrating lymphocyte monotherapy, in patients with advanced non-small cell lung cancer resistant to immune checkpoint inhibitors. Cancer Discov. (2024) 14:1389–402. doi:  10.1158/2159-8290.cd-23-1334 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Becht E, McInnes L, Healy J, Dutertre C-A, Kwok IWH, Ng LG, et al. Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol. (2018). doi:  10.1038/nbt.4314 [DOI] [PubMed] [Google Scholar]
  • 13. Kobak D, Berens P. The art of using t-SNE for single-cell transcriptomics. Nat Commun. (2019) 10:5416. doi:  10.1038/s41467-019-13056-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Zacharakis N, Huq LM, Seitter SJ, Kim SP, Gartner JJ, Sindiri S, et al. Breast cancers are immunogenic: Immunologic analyses and a phase II pilot clinical trial using mutation-reactive autologous lymphocytes. J Clin Oncol Off J Am Soc Clin Oncol. (2022) 40:1741–54. doi:  10.1200/jco.21.02170 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Hackenberg M, Canal Guitart L, Backofen R, Binder H. Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data. Briefings Bioinf. (2025) 26:bbaf287. doi:  10.1093/bib/bbaf287 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Liu Z, Ma R, Zhong Y. Assessing and improving reliability of neighbor embedding methods: a map-continuity perspective. Nat Commun. (2025) 16:5037. doi:  10.1038/s41467-025-60434-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Bao H, Sakaue S. Sparse regularized optimal transport with deformed q-entropy. Entropy (Basel Switzerland). (2022) 24:1634. doi:  10.3390/e24111634 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Huizing G-J, Peyré G, Cantini L. Optimal transport improves cell-cell similarity inference in single-cell omics data. Bioinf (Oxford England). (2022) 38:2169–77. doi:  10.1101/2021.03.19.436159 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Khamis A, Tsuchida R, Tarek M, Rolland V, Petersson L. Scalable optimal transport methods in machine learning: A contemporary survey. IEEE Trans Pattern Anal Mach Intell. (2024) 10(2):123–45. doi:  10.1109/tpami.2024.3379571 [DOI] [PubMed] [Google Scholar]
  • 20. Shemonti AS, Gmyrek GB, Quintelier KLA, Van Gassen S, Saeys Y, Willemsen M, et al. CARGO: A cytometry analysis framework via regularized graph optimal-transport. PloS Comput Biol. (2026). doi:  10.1371/journal.pcbi.1014358 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Xia L, Lee C, Li JJ. Statistical method scDEED for detecting dubious 2D single-cell embeddings and optimizing t-SNE and UMAP hyperparameters. Nat Commun. (2024) 15:1753. doi:  10.1038/s41467-024-45891-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Wang Y-MC, Wang J, Hon YY, Zhou L, Fang L, Ahn HY. Evaluating and reporting the immunogenicity impacts for biological products—a clinical pharmacology perspective. AAPS J. (2016) 18:395–403. doi:  10.1208/s12248-015-9857-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Zhang W, Jordan KR, Schulte B, Purev E. Characterization of clinical grade CD19 chimeric antigen receptor T cells produced using automated CliniMACS Prodigy system. Drug Des Dev Ther. (2018) 12:3343–56. doi:  10.2147/dddt.s175113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Smith KD, MacDonald JW, Li X, Beirne E, Stewart G, Bammler TK, et al. Rigor and reproducibility of spatial transcriptomics performed on clinically sourced human tissues. Lab Investigation; A J Tech Methods Pathol. (2025) 105:104190. doi:  10.1016/j.labinv.2025.104190 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Joodaki M, Shaigan M, Parra V, Bülow RD, Kuppe C, Hölscher DL, et al. Detection of patient-level distances from single cell genomics and pathomics data with optimal transport (PILOT). Mol Syst Biol. (2024) 20:57–74. doi:  10.1038/s44320-023-00003-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Orlova DY, Zimmerman N, Meehan S, Meehan C, Waters J, Ghosn EEB, et al. Earth mover's distance (EMD): A true metric for comparing biomarker expression levels in cell populations. PloS One. (2016) 11:e0151859. doi:  10.1371/journal.pone.0151859 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Klein D, Palla G, Lange M, Klein M, Piran Z, Gander M, et al. Mapping cells through time and space with moscot. Nature. (2025) 638:1065–75. doi:  10.1038/s41586-024-08453-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Jin J, Sabatino M, Somerville R, Wilson JR, Dudley ME, Stroncek DF, et al. Simplified method of the growth of human tumor infiltrating lymphocytes in gas-permeable flasks to numbers needed for patient treatment. J Immunother (Hagerstown Md: 1997). (2012) 35:283–92. doi:  10.1097/cji.0b013e31824e801f [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Zimmerlin L, Donnenberg VS, Donnenberg AD. Rare event detection and analysis in flow cytometry: bone marrow mesenchymal stem cells, breast cancer stem/progenitor cells in Malignant effusions, and pericytes in disaggregated adipose tissue. Methods Mol Biol (Clifton NJ). (2011) 699:251–73. doi:  10.1007/978-1-61737-950-5_12 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Kamat R, Henney CS. Studies on T cell clonal expansion. I. Suppression of killer T cell production in vivo. J Immunol (Baltimore Md: 1950). (1975) 115:1592–8. [PubMed] [Google Scholar]
  • 31. Shemonti AS, Gmyrek GB, Quintelier KLA, Gassen SV, Saeys Y, Willemsen M, et al. An optimal transport-based low-dimensional visualization framework for high-parameter flow cytometry. (2025). doi:  10.1101/2025.08.19.670604 [DOI] [Google Scholar]
  • 32. Chen Y, Hu Z, Chen W, Huang H. Fast and scalable Wasserstein-1 neural optimal transport solver for single-cell perturbation prediction. Bioinf (Oxford England). (2025) 41:i513–22. doi:  10.1093/bioinformatics/btaf253 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Cuturi M. Sinkhorn distances: Lightspeed computation of optimal transportation distances. (2013). [Google Scholar]
  • 34. Cao Q, Zhao J, Wang H, Guan Q, Zheng C. An integrated method based on Wasserstein distance and graph for cancer subtype discovery. IEEE/ACM Trans Comput Biol Bioinf. (2023) 20:3499–510. doi:  10.1109/tcbb.2023.3293472 [DOI] [PubMed] [Google Scholar]
  • 35. Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. Deep generative modeling for single-cell transcriptomics. Nat Methods. (2018) 15:1053–8. doi:  10.1038/s41592-018-0229-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Cao K, Gong Q, Hong Y, Wan L. A unified computational framework for single-cell data integration with optimal transport. Nat Commun. (2022) 13:7419. doi:  10.1038/s41467-022-35094-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Jorgovanovic D, Song M, Wang L, Zhang Y. Roles of IFN-γ in tumor progression and regression: a review. biomark Res. (2020) 8:49. doi:  10.1186/s40364-020-00228-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Li M, Xia P, Du Y, Liu S, Huang G, Chen J, et al. T-cell immunoglobulin and ITIM domain (TIGIT) receptor/poliovirus receptor (PVR) ligand engagement suppresses interferon-γ production of natural killer cells via β-arrestin 2-mediated negative signaling. J Biol Chem. (2014) 289:17647–57. doi:  10.1074/jbc.m114.572420 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Bagaev A, Kotlov N, Nomie K, Svekolkin V, Gafurov A, Isaeva O, et al. Conserved pan-cancer microenvironment subtypes predict response to immunotherapy. Cancer Cell. (2021) 39:845–865.e847. doi:  10.1016/j.ccell.2021.04.014 [DOI] [PubMed] [Google Scholar]
  • 40. Adusumilli PS, Zauderer MG, Rivière I, Solomon SB, Rusch VW, O'Cearbhaill RE, et al. A phase I trial of regional mesothelin-targeted CAR T-cell therapy in patients with Malignant pleural disease, in combination with the anti–PD-1 agent pembrolizumab. Cancer Discov. (2021) 11:2748–63. doi:  10.1158/2159-8290.cd-21-0407 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Gohil SH, Iorgulescu JB, Braun DA, Keskin DB, Livak KJ. Applying high-dimensional single-cell technologies to the analysis of cancer immunotherapy. Nat Rev Clin Oncol. (2021) 18:244–56. doi:  10.1038/s41571-020-00449-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Hall MS, Mullinax JE, Cox CA, Hall AM, Beatty MS, Blauvelt J, et al. Combination nivolumab, CD137 agonism, and adoptive cell therapy with tumor-infiltrating lymphocytes for patients with metastatic melanoma. Clin Cancer Research: Off J Am Assoc For Cancer Res. (2022) 28:5317–29. doi:  10.1158/1078-0432.ccr-22-2103 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. L’Orphelin J-M, Lancien U, Nguyen J-M, Coronilla FJS, Saiagh S, Cassecuel J, et al. NIVO-TIL: combination anti-PD-1 therapy and adoptive T-cell transfer in untreated metastatic melanoma: an exploratory open-label phase I trial. Acta Oncol. (2024) 63:40495. doi:  10.1080/0284186X.2023.2289417 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Wu L, Mao L, Liu J-F, Chen L, Yu G-T, Yang L-L, et al. Blockade of TIGIT/CD155 signaling reverses T-cell exhaustion and enhances antitumor capability in head and neck squamous cell carcinoma. Cancer Immunol Res. (2019) 7:1700–13. doi:  10.1158/2326-6066.cir-18-0725 [DOI] [PubMed] [Google Scholar]
  • 45. Joller N, Hafler JP, Brynedal B, Kassam N, Spoerl S, Levin SD, et al. Cutting edge: TIGIT has T cell-intrinsic inhibitory functions. J Immunol (Baltimore Md: 1950). (2011) 186:1338–42. doi:  10.4049/jimmunol.1003081 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Bunne C, Stark SG, Gut G, Del Castillo JS, Levesque M, Lehmann K-V, et al. Learning single-cell perturbation responses using neural optimal transport. Nat Methods. (2023) 20:1759–68. doi:  10.1038/s41592-023-01969-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Kang S, Mansurov A, Kurtanich T, Chun HR, Slezak AJ, Volpatti LR, et al. Engineered IL-7 synergizes with IL-12 immunotherapy to prevent T cell exhaustion and promote memory without exacerbating toxicity. Sci Adv. (2023) 9:eadh9879. doi:  10.1126/sciadv.adh9879 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Oviedo-Orta E, Bermudez-Fajardo A, Karanam S, Benbow U, Newby AC. Comparison of MMP-2 and MMP-9 secretion from T helper 0, 1 and 2 lymphocytes alone and in coculture with macrophages. Immunology. (2008) 124:42–50. doi:  10.1111/j.1365-2567.2007.02728.x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Aronson SL, Walker C, Thijssen B, van de Vijver KK, Horlings HM, Sanders J, et al. Tumour microenvironment characterisation to stratify patients for hyperthermic intraperitoneal chemotherapy in high-grade serous ovarian cancer (OVHIPEC-1). Br J Cancer. (2024) 131:565–76. doi:  10.1038/s41416-024-02731-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Goey SH, Eggermont AMM, Punt CJA, Slingerland R, Gratama JW, Oosterom R, et al. Intrapleural administration of interleukin 2 in pleural mesothelioma: a phase I-II study. Br J Cancer. (1995) 72:1283–8. doi:  10.1038/bjc.1995.501 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Kielsen K, Oostenbrink LVE, von Asmuth EGJ, Jansen-Hoogendijk AM, van Ostaijen-Ten Dam MM, Ifversen M, et al. Il-7 and il-15 levels reflect the degree of t cell depletion during lymphopenia and are associated with an expansion of effector memory t cells after pediatric hematopoietic stem cell transplantation. J Immunol (Baltimore Md: 1950). (2021) 206:2828–38. doi:  10.4049/jimmunol.2001077 [DOI] [PubMed] [Google Scholar]
  • 52. Donnenberg VS, Donnenberg AD. Coping with artifact in the analysis of flow cytometric data. Methods (San Diego Calif). (2015) 82:3–11. doi:  10.1016/j.ymeth.2015.03.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53. Chari T, Pachter L. The specious art of single-cell genomics. PloS Comput Biol. (2023) 19:e1011288. doi:  10.1371/journal.pcbi.1011288 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Thommen DS, Schumacher TN. T cell dysfunction in cancer. Cancer Cell. (2018) 33:547–62. doi:  10.1016/j.ccell.2018.03.012 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

DataSheet1.pdf (146.7KB, pdf)

Data Availability Statement

Data will be made available to qualified investigators upon request to the corresponding author.

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding author.


Articles from Frontiers in Immunology are provided here courtesy of Frontiers Media SA

RESOURCES