Abstract
Motivation
A crucial component of intuitive data visualization is presenting a hierarchical tree structure with interactive functions. For example, single-cell transcriptomics studies may generate gene expression values with developmental trajectories or cell lineage structures. Two common visualization methods, t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP), require two separate figures to depict the distribution of cell types and gene expression data, with low-dimension projections that may not capture the hierarchical structures among cells.
Results
Here, we present a JavaScript framework and an interactive web app named Collapsible Tree, which presents values jointly with interactive, expandable, and collapsible lineage structures. For example, the Collapsible Tree presents cellular states and gene expression from single-cell transcriptomics within a single hierarchical plot, enabling comparisons of gene expression across lineages and subtle patterns between sub-lineages. Our framework can facilitate the exploration of complicated value patterns that are not evident in UMAP or t-SNE plots.
Availability and implementation
The Collapsible Tree web interface is available at https://collapsibletree.data2in.net. The JavaScript library source code is available at https://github.com/data2intelligence/collapsible_tree.
1 Introduction
Many biological data can be abstracted into hierarchical tree structures. For example, recent advancements in single-cell RNA sequencing (scRNA-seq) have enabled researchers to unravel transcriptional hierarchies among individual cells to explore differentiation trajectories or cell lineages. The widely used visualization methods for single-cell analysis include Uniform Manifold Approximation and Projection (UMAP) and t-Distributed Stochastic Neighbor Embedding (t-SNE) (Van der Maaten and Hinton 2008, McInnes et al. 2018), which transform high-dimensional data into low-dimension (typically two) representations while attempting to preserve underlying structures in the data. Despite the widespread application of UMAP and t-SNE, the substantial reduction in dimensionality can distort the distances between cell clusters (Chari and Pachter 2023) and cannot faithfully reflect the cell lineage structures. Thus, there is an unmet need for an interactive framework to visualize gene or pathway expression in lineage structures from scRNA-seq data to complement t-SNE/UMAP projections.
Besides serving as a visualization scheme, tree-based structures, representing variants within or deviating from the consensus root, may work as a better nomenclature reference in cell-type classification than current cell-type atlas references (Domcke and Shendure 2023). Beyond single-cell data, many data types can be summarized into tree-like structures (Kapli et al. 2020). Therefore, the tree-structure visualization framework can be broadly useful in diverse biological studies.
There are several tools utilizing hierarchical methods to visualize the relationship between cell clusters. One notable example is TooManyCells (Schwartz et al. 2020), which recursively bi-partition cells into clades to form a tree-like structure. However, TooManyCells lacks customizability since users cannot input their own structures, as nodes and sub-clades are algorithmically generated without direct biological interpretability (i.e. without the ability to provide supervised input). Additionally, it does not provide interactive functions for users to tailor the lineage depth to show the right detail level.
Here, we introduce a JavaScript framework named Collapsible Tree and an example web interface that presents cell lineage structures in scRNA-seq data. With the hierarchical structure and interactive functions, users can observe gene expression data across major and sub-cell types at various resolutions. These cell groups are collapsible and expandable with mouse clicks, enabling a detailed or broad view as needed. For single-cell data, the Collapsible Tree can help identify subtle transcriptomics patterns that may not be apparent from t-SNE or UMAP plots.
2 Software description
2.1 Input formats
The input includes a predefined tree structure and a value profile on tree nodes. For example, in single-cell transcriptomics, the tree structure could be knowledge-based lineage annotations as parent–child relationships. The value profile on tree nodes could be the average gene expression across all cells for each lineage. Users can provide a collection of icons depicting each tree node for the input tree structure. We provided a collection of cell-shape icons for common immune and stroma cells in tumors. Users also have the flexibility to use their icons. Although our sample input is prepared with the lineage structure for scRNA-seq data, the web application can accommodate various other data types, such as pathway enrichment scores. Additionally, while we used a knowledge-based lineage structure as an example, the input tree can be any tree-like structure derived by algorithms.
2.2 Functionalities
In this section, we focus on examples of scRNA-seq data visualization. However, all features discussed are broadly applicable to tree structures. Unlike t-SNE or UMAP, which embed the data in a low-dimensional space based on the nearest-neighbor graph, our approach utilizes a cell lineage structure to depict the relationships between gene expression (or any other variables) and annotated cell lineages. In our visualization, nodes represent distinct cell types. Leaf nodes denote sub-cell types, and their parent nodes represent the major cell types encompassing these subgroups.
The incoming edges are color-coded proportional to the value of the node. The width of the edge represents any weight input by users. However, if users do not assign values and weights for a tree node, which may happen for a non-leaf node, our framework will determine value using child nodes as follows:
Value of parent node =
Weight of parent node =
For scRNA-seq data, the and are the number of cells and average expression at each cell lineage node, respectively.
Traditional methods, such as t-SNE or UMAP, require separate plots to display cell type mappings and gene expression levels. One plot typically uses color coding to represent cell-type clusters without hierarchical structures, while the other illustrates the gene expression intensity. In contrast, our interface integrates cell type and gene expression information into a single graph, harnessing the location of each node to represent the cell type. Of course, the trade-off is that, unlike t-SNE and UMAP, our approach cannot visualize the “proximity” relationship of individual cells.
Our interface offers several interactive features. By clicking on the node icon, users can expand or collapse certain sub-lineages and focus on specific cell types of interest. Expression values are hidden by default and displayed upon mouseover, maintaining a clean interface while providing ample information through mouse actions. In addition, we offer two layout options: horizontal and radial trees. There are three color themes to choose from: Brewer Blues, Red, and Viridis.
2.3 Implementation
For users lacking programming expertise, we provide an easy-to-use web interface (https://collapsibletree.data2in.net) that enables researchers to upload a predefined CSV file of any tree structures and quantitative value (e.g. gene expression values in a scRNA-seq study). The generated image can be downloaded as an SVG, which can be opened with a web browser such as Google Chrome and printed as a PDF. The PDF file can then be edited using Inkscape or Adobe Illustrator as a vector graphic.
For programmers, we have packed all functions into a general-purpose library on GitHub (https://github.com/data2intelligence/collapsible_tree) based on the JavaScript D3 library. The D3 library allows the plotting of hierarchical information, enabling the stratification of cell lineages and automatic node positioning. The D3 library also enables interactive functions, including dynamic layout rearrangement when expanding or hiding branches, as well as various mouse-related events. Our source code is highly customizable and can be easily tailored to accommodate different structures or purposes, making it an ideal choice for integration into existing websites.
An example website utilizing the JavaScript framework, enabling gene expression queries in common cell lineages with the pre-processed data from our SpaCET study (Ru et al. 2023), is available at our Tres framework (Zhang et al. 2022) (https://resilience.ccr.cancer.gov).
3 Evaluation
To demonstrate the hierarchical structure’s strength in comparing gene expression across subcellular types, we utilize non-small cell lung cancer (NSCLC) data as an example, showcasing the IFI6 expression pattern (Fig. 1). The NSCLC scRNA-seq data was from the EBI ArrayExpress website with the ID E-MTAB-6149 (Lambrechts et al. 2018), containing 19 samples of human NSCLC tumors and covering distinct immune and stromal cell types and sub-lineages. IFI6 is widely expressed across multiple lineages and thus serves as an ideal example of genes with complicated expression patterns. We visualized the data using t-SNE plots and UMAP plots, generated by Seurat (Hao et al. 2024) with default parameters.
Figure 1.
Visualization of IFI6 expression across all cell types. (a) Gene expression visualized by t-SNE. (b) Cell type annotations for panel a. (c) Gene expression visualized by UMAP. (d) Cell type annotations for panel c. (e) Collapsible tree.
The t-SNE and UMAP visualizations demonstrated IFI6 expression across multiple cell types without being enriched in specific lineages (Fig. 1a–d). In contrast, the Collapsible Tree presents IFI6 expression in two main branches: myeloid and lymphoid progenitors (Fig. 1e). Specifically, IFI6 exhibits high expression in several middle nodes, such as Macrophage, T CD4, and T CD8. Within T CD8 groups, a notable enrichment of expression was observed in exhausted T cells, as indicated by color. This example indicates that our structure provides an intuitive and informative approach to identifying distinctions within and across different lineages, which may be hidden in t-SNE or UMAP visualizations.
4 Discussion
We present the Collapsible Tree app to visualize values in interactive tree structures, thus providing new visualizations to complement traditional methods like UMAP and t-SNE. A limitation of our framework in single-cell data visualization is that it presents the average value in each branch, thus losing the resolution of individual cells. A future direction is to combine lineage trees and low-dimension presentations, such as UMAP or t-SNE, to visualize single-cell values organized in hierarchical structures. Nonetheless, we foresee that the Collapsible Tree will greatly facilitate biological discoveries and therapeutic developments through intuitive and interactive data analysis.
Acknowledgements
We sincerely thank Beibei Ru and Dongze He for their valuable discussions.
Contributor Information
Yuan Gao, Cancer Data Science Lab, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD 20892, United States; Center for Bioinformatics and Computational Biology, University of Maryland, College Park, MD 20742, United States.
Rob Patro, Center for Bioinformatics and Computational Biology, University of Maryland, College Park, MD 20742, United States.
Peng Jiang, Cancer Data Science Lab, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD 20892, United States.
Conflict of interest
R.P. is a co-founder of Ocean Genomics Inc.
Funding
This work was supported by the intramural budget allocation and the FLEX Synergy award of National Cancer Institute (NCI), National Institutes of Health (NIH), Technology Impact Award of Cancer Research Institute (CRI) awarded to P.J.; and NIH R01HG009937 and National Science Foundation (NSF) awards CCF-1750472 and CNS-1763680 to R.P. This project has been made possible in part by grant number 2022-252586 from the Chan Zuckerberg Initiative DAF, an advised fund of Silicon Valley Community Foundation.
Data availability
The data underlying this article are available in Github at https://github.com/data2intelligence/collapsible_tree/tree/main/static/sample_data.
References
- Chari T, Pachter L.. The specious art of single-cell genomics. PLoS Comput Biol 2023;19:e1011288. 10.1371/journal.pcbi.1011288 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Domcke S, Shendure J.. A reference cell tree will serve science better than a reference cell atlas. Cell 2023;186:1103–14. 10.1016/j.cell.2023.02.016 [DOI] [PubMed] [Google Scholar]
- Hao Y, Stuart T, Kowalski MH. et al. Dictionary learning for integrative, multimodal and scalable single-cell analysis. Nat Biotechnol 2024;42:293–304. 10.1038/s41587-023-01767-y [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kapli P, Yang Z, Telford MJ.. Phylogenetic tree building in the genomic age. Nat Rev Genet 2020;21:428–44. 10.1038/s41576-020-0233-0 [DOI] [PubMed] [Google Scholar]
- Lambrechts D, Wauters E, Boeckx B. et al. Phenotype molding of stromal cells in the lung tumor microenvironment. Nat Med 2018;24:1277–89. 10.1038/s41591-018-0096-5 [DOI] [PubMed] [Google Scholar]
- McInnes L, Healy J, Saul N. et al. UMAP: uniform manifold approximation and projection. JOSS 2018;3:861. 10.21105/joss.00861 [DOI] [Google Scholar]
- Ru B, Huang J, Zhang Y. et al. Estimation of cell lineages in tumors from spatial transcriptomics data. Nat Commun 2023;14:568. 10.1038/s41467-023-36062-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schwartz GW, Zhou Y, Petrovic J. et al. TooManyCells identifies and visualizes relationships of single-cell clades. Nat Methods 2020;17:405–13. 10.1038/s41592-020-0748-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Van der Maaten L, Hinton G.. Visualizing data using t-SNE. J Mach Learn Res 2008;9:2579–605. [Google Scholar]
- Zhang Y, Vu T, Palmer DC. et al. A T cell resilience model associated with response to immunotherapy in multiple tumor types. Nat Med 2022;28:1421–31. 10.1038/s41591-022-01799-y [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data underlying this article are available in Github at https://github.com/data2intelligence/collapsible_tree/tree/main/static/sample_data.

