Summary
Single-cell-resolved systems biology methods, including omics- and imaging-based measurement modalities, generate a wealth of high-dimensional data characterizing the heterogeneity of cell populations. Representation learning methods are routinely used to analyze these complex, high-dimensional data by projecting them into lower-dimensional embeddings. This facilitates the interpretation and interrogation of the structures, dynamics, and regulation of cell heterogeneity. Reflecting their central role in analyzing diverse single-cell data types, a myriad of representation learning methods exist, with new approaches continually emerging. Here, we contrast general features of representation learning methods spanning statistical, manifold learning, and neural network approaches. We consider key steps involved in representation learning with single-cell data, including data pre-processing, hyperparameter optimization, downstream analysis, and biological validation. Interdependencies and contingencies linking these steps are also highlighted. This overview is intended to guide researchers in the selection, application, and optimization of representation learning strategies for current and future single-cell research applications.
Keywords: deep learning, manifold learning, dimension reduction, hyperparameter, omics, systems microscopy
High-dimensional data generated by single-cell systems biology (omics) methods require powerful representation learning approaches to enable interpretation and interrogation of the structures, dynamics, and regulation of cell heterogeneity. In this perspective, Gunawan et al. elucidate key steps involved in representation learning to guide optimized method application by researchers analyzing diverse single-cell data modalities.
Introduction
Dramatic advances in single-cell analysis technologies have brought cellular heterogeneity into sharp focus. The combination of quantitative, multi-parametric systems biology techniques, termed “omics,” with single-cell-resolved measurement capabilities has permitted the delineation of cellular states and their transitional trajectories, as well as inference of mechanisms controlling cell-state dynamics.1 Single-cell omics provide insight by leveraging information inherent within cellular heterogeneity: in the quantitative variation and co-variation of cellular features,2 accessible whether induced by experimental perturbations, driven by disease processes, or naturally occurring. Insights into cell subpopulation subtypes and spectra, state dynamics, and molecular control, as a result of single-cell data analysis, have been influential in diverse fields, including but not limited to cancer research,3 immunology,4 developmental biology,5 neuroscience,6 and drug discovery.7 Given these advantages, single-cell omics technologies have rapidly diversified to encompass proteomics,8 metabolomics,9 genomics,10 epigenetics,10 and especially transcriptomics, where single-cell RNA sequencing (scRNA-seq) technologies are reaching maturity.11 Maturing in parallel, high-throughput, high-content single-cell microscopy, also termed systems microscopy,12,13 is becoming increasingly powerful for the systematic quantification of cell phenotypes in time and space.14,15 Now, emerging interfaces between single-cell omics and systems microscopy has begun to permit interrogation of cellular heterogeneity defined by spatiotemporally resolved measures of both molecular and cellular states. This fusion will illuminate the diversity and the control of cellular phenotypes as never before.16
The wide and diverse deployment of single-cell measurement technologies has produced a deluge of data requiring corresponding advances in computational methods for data analysis.17 A significant challenge faced by researchers analyzing omics and imaging-based single-cell data is high data dimensionality. With numerous (100s–10,000s) features simultaneously measured per cell, a vast data landscape of potential cell states is created, requiring (theoretically) a correspondingly vast dataset to achieve a well-sampled state space. Real-world experimental constraints often lead to relative sparsity in the true data landscape, obscuring relationships between cells and between features.18,19,20 Additionally, intrinsic biological and technical noise can further convolute meaningful signals in an under-sampled space.21 In single-cell data, cell atlas initiatives have produced huge omics datasets, some reaching hundreds of thousands of cells, creating the potential for powerful analyses.22 Yet, increases in the number of measurable features impose ever-greater computational and scaling challenges, often resulting in intractable analyses.21
To address these challenges, dimension-rich datasets are now routinely interrogated via some form of representation learning, a term encompassing a suite of analytical techniques that project high-dimensional data into a lower-dimensional space, or “representation.” Such low-dimensional embedding is achieved by stripping out redundancies and noise in the data that occupy dimensions beyond the intrinsic structure of the phenomena under investigation, i.e., cell diversity. Representation learning is not only statistically motivated but is also biologically intuitive. For example, regulatory modules formed by genes are expressed in a coordinated manner, and thus the dimensionality needed to represent the expression of highly correlated genes can naturally be compressed,23 better revealing the underlying metaparameters driving biological phenomena. Generally, representation learning is used to transform data into configurations amenable to various types of downstream analysis. This coupling of representation learning with downstream analytical tasks is exemplified by software tools such as semisoft clustering with pure cells (SOUP),24 scPred,25 and Monocle,26 which apply representation learning for dimension reduction prior to downstream clustering (e.g., subpopulation detection), classification (e.g., supervised cell-type labeling), or pseudo-time construction (e.g., inferring the main trajectories of dynamic state changes), respectively (Figure 1A).
Figure 1.
Two-dimensional representation learning from three-dimensional data
(A) Data are sampled from an underlying biological structure or state space. Single-cell experiments often measure thousands of features simultaneously, resulting in a high-dimensional space, here exemplified in only three dimensions for visualization (left). Degrees of freedom in high dimensions create a vast state-space volume, leading to sparsity, which makes identifying relationships challenging (center). Projection by representation learning to a lower dimension reduces sparsity to reveal relationships in the data for further analysis, but fidelity to the original state-space structure is not guaranteed (right). The choice of representation learning method is critical to determining analytical outcomes. Alternate analytical goals (e.g., clustering and/or classification or lineage inference) may encourage selection of distinct, task-matched representation learning methods.
(B) Overview of common stages in a general single-cell analysis pipeline. Data require preparation and pre-processing before they can undergo representation learning to reveal relationships and compress dimensions. The application of representation learning requires hyperparameter optimization to achieve optimal results. Further downstream analyses then leverage learned embeddings. Result interpretation requires an understanding of the quality of the analysis and, often, further biological validation. Iterative refinement of each stage, or the entire pipeline, may be required.
In practice across the biological sciences, representation learning is often a central component in analytical processes comprising several common steps: (1) data preparation and pre-processing, (2) selection of representation learning methods, (3) hyperparameter optimization, (4) downstream analyses, and (5) evaluation and interpretation of results. Such pipelines are typically iterative, with analyses often requiring cycles of refinement.
Here, we provide an end-to-end overview of techniques and options in the composition of single-cell analysis pipelines centered on representation learning. We highlight key considerations aligned to the common steps of such pipelines (Figure 1B), exploring influences on the nature and quality of learned embeddings and how these impact downstream analysis outcomes. Constituting a guide for those less familiar with the field, we step through each stage of a typical single-cell investigation. We first discuss data preparation and pre-processing and give an overview of representation learning approaches and methods including their optimization via hyperparameter tuning. We highlight how downstream analytical goals should influence upstream representation learning strategies and, finally, emphasize the challenges inherent to evaluating and interpreting results.
Common steps in representation learning-centered single-cell analysis pipelines
Upstream data preparation and pre-processing
Upstream pre-processing prepares a dataset so that it can be embedded optimally for further analysis. Pre-processing of data, including transformations/filtration, denoising, imputation, and integration, occurs prior to representation learning to improve the quality of embeddings. At times overlooked, pre-processing can significantly influence analysis. While each data modality has its own specific concerns and methods, we highlight common themes and examples.
Feature transformation, filtration, and denoising
Raw data are often unsuitable for embedding and further analysis. Measured features may be at different scales and between different ranges, requiring normalization or standardization to make features comparable and to prevent domination by a subset of features. Additionally, confounding factors can be minimized by transformations or feature extraction. For example, log transformations are often applied to scRNA-seq data to remove mean-variance dependencies that can be problematic for some representation learning methods including principal-component analysis (PCA).27 Other transformations may highlight information, such as contextual relationships emphasized by fish-eye transformation in single-cell microscopy,28 that is otherwise insufficiently represented in learned embeddings. Feature filtration is also commonly used in single-cell microscopy, removing features that contain (in many cases) irrelevant spatial information, such as a cell’s angular orientation within the image field, while retaining features such as mean intensity or nuclear-cytoplasmic intensity ratios, associated with marker (e.g., protein) levels and/or localization, respectively. Notably, single-cell image segmentation and analysis methods are routinely used to transform cell images into high-dimensional quantitative cell descriptors.29,30 This permits scalable and robust interpretation, as well as the use of representation learning techniques that would otherwise be restricted to omics-style numerical matrix data.31,32 Interestingly, while image data are typically transformed into quantitative data for analysis, omics researchers have begun experimenting with “imagification” of omics data, bringing potent image modeling methods, especially convolutional neural networks, to bear on the analysis of omics data.33,34
Removal of background variation and biological noise with respect to the phenomena under investigation can also improve embeddings and subsequent analysis. This is demonstrated in the comparison of methods to filter out “nuisance genes” in scRNA-seq data from Drosophila wing disc myoblast cells35 and by Satjia et al., who exclude known housekeeping genes in scRNA-seq data to study spatial patterning of gene expression in zebrafish embryos.36 Technical noise is also unavoidable and is specific to any measurement system. Such noise obfuscates signals of interest and is exacerbated by high-dimensional spaces. Hence, raw data generally require some form of denoising before further analysis. Imaging artifacts such as skews caused by optical objectives create systematic aberrations requiring corrections. Random noise is also common, exemplified by salt and pepper noise and imaging backgrounds, which can also be minimized37,38 (Figure 2A). In scRNA-seq measurements, random noise is exemplified by the spurious zero counts (“dropouts”) that are commonly accounted for by imputation methods.39,40,41 However, denoising can be particularly challenging in rare cell populations or where scRNA-seq sequencing depth is low, as noted by Wang et al., who use transfer learning to improve data quality.42 Denoising and imputation strategies for scRNA-seq data have been reviewed by Patruno et al., identifying versatile tools as well as context-dependent performers.43
Figure 2.
Influence of pre-processing on representation learning
(A) Imaging artifacts obscuring cells in label-free imaging (top; original data) are removed by background flattening through Gaussian smoothing and subtraction, creating a more consistent image (bottom; original data).
(B) Batch-effect correction performed by CytoNorm in mouse bone marrow cytometry data to form a unified dataset while maintaining informative cell-type clusters (adapted from Ashhurst et al. with permission44).
While not all representation learning methods require the same pre-processing steps, it is necessary to carefully consider these in order to suppress spurious or irrelevant variation and make data more amenable to effective representation learning and subsequent analyses.45 Because pre-processing impacts the performance of any subsequent representation learning tools, iterative refinement of pre-processing on the basis of downstream outcomes may be important as an investigation matures.
Integrating heterogeneous data from multiple batches and measurement modes
Even with single omics, heterogeneity exists between experiments, which could be due to different measurement systems, different cells, and/or experimental procedures. For a unified analysis, corrections between batches are required (“horizontal integration”) (Figure 2B). Several approaches have been proposed for this including alignment through data alignment and empirical Bayesian modeling.46 This is demonstrated by Lotfollahi et al., who use transfer learning with conditional variational autoencoders to produce joint embeddings between query datasets and reference atlases, allowing large cell numbers in such atlases to be leveraged. However, even though individual omics technologies have brought deep insights, each separate omics method unveils a distinct perspective on the cell. Only by bringing these views together can we achieve a more comprehensive understanding of cellular states, machinery, and their interplay. Such multi-modal data integration is generally made possible either through experimental methods enabling concurrent measurement of alternate data modalities in the same individual cells (“vertical integration”) or through emerging computational methods to integrate multi-modal data derived in parallel from different cell populations47 (“diagonal integration”).
Measurement of multiple omics data types in the same cells has so far included single-cell methylome and transcriptome sequencing (scM&T-seq), as well as single-cell nucleosome, methylation, and transcription sequencing (scNMT-seq).16,48 The capacity to diversify data modalities concurrently accessed per cell is strongly limited by the distinct and often incompatible experimental requirements for each measurement modality. Thus, more diverse (less constrained) data integration relies on computational approaches for multi-modal data integration, which commonly leverage representation learning methods. For example, Yang et al. performed manifold alignment using deep learning autoencoders to integrate unmatched scRNA-seq and chromatin images of naive CD4+ T cells.49 Other computational tools include manifold alignment to characterize experimental relationships (MATCHER),50 maximum mean discrepancy manifold alignment (MMD-MA),51 and UnionCom52; however, their general use and interpretation remain challenging.53 We direct the reader to a review by Argelaguet et al.47 for further details.
Finally, we highlight multi-modal data integration as a key research opportunity as well as a data preparation/pre-processing challenge, with several integration methods inherently using representation learning approaches to output joint embeddings.54 Undoubtedly, optimized methods for computational data integration will permit cogent interrogation of ever more diverse data sources irrespective of inflexible experimental requirements, providing a pathway for increasingly comprehensive analyses of cellular processes.
State-of-the-art representation learning methods and their characteristics
Representation learning projects data into spaces amenable to further analysis (often lower dimensional). The emphasis on the development and comparison of representation learning methods has increased dramatically in recent years due to the pivotal role of these methods in analyzing increasingly complex single-cell datasets. Here, we group the diverse selection of popular representation learning methods (Table 1) into three generalized categories wherein methods broadly share similar characteristics: (1) statistical/probabilistic approaches, including matrix decomposition, (2) manifold learning, and (3) neural networks. Each method has advantages and limitations highly dependent on their circumstance of use and the dataset of interest. Below, we discuss the characteristics of each group in turn, seeking to provide the reader with some starting rationale and intuition for method selection and comparison.
Table 1.
Selective summary of representation learning methods
| Method | Example usage in single-cell analyses | Linearity | |
|---|---|---|---|
| Statistical methods | Principle component analysis (PCA)55 | sc-qPCR,56 microscopy features57 | linear |
| Non-negative matrix factorization (NMF)58 | scRNA-seq59 | linear | |
| Similarly weighted non-negative embedding (SWNE)60 | scRNA-seq60 | non-linear | |
| Singular value decomposition (SVD)61 | scRNA-seq45 | linear | |
| Latent semantic analysis (LSI)62 | scATAC-seq63,64 | linear | |
| Independent component analysis (ICA)65 | scRNA-seq26 | linear | |
| Zero inflated factor analysis (ZIFA)66 | scRNA-seq66 | linear | |
| Latent dirichlet allocation (LDA)67 | scRNA-seq,68 scATAC-seq69 | non-linear | |
| Gaussian process latent variable models (GPLVM)70 | scRNA-seq71 | non-linear | |
| Manifold learning | t-distributed stochastic neighbor embedding (t-SNE)72 | scRNA-seq,73 microscopy features74 | non-linear |
| Isomap75 | scRNA-seq76 | non-linear | |
| Locally linear embedding (LLE)77 | scRNA-Seq78 | non-linear | |
| Uniform manifold approximation and projection (UMAP)79 | scRNA-seq,80 microscopy features81 | non-linear | |
| Potential of heat-diffusion affinity-based transition embedding (PHATE)82 | CyTOF, scRNA-Seq82 | non-linear | |
| Laplacian eigenmaps83 | scRNA-seq84 | non-linear | |
| Diffusion maps85 | sc-qPCR, scRNA-seq85 | non-linear | |
| Local tangent space alignment (LTSA)86 | scRNA-seq87 | non-linear | |
| PhenoGraph88 | mass cytometry88 | non-linear | |
| Markov affinity-based graph imputation of cells (MAGIC)40 | scRNA-seq40 | non-linear | |
| Neural networks | Sparse autoencoder for unsupervised clustering, imputation, and embedding (SAUCIE)89 | CyTOF, scRNA-seq89 | non-linear |
| Deep variational autoencoder for scRNA-seq data (VASC)90 | scRNA-seq90 | non-linear | |
| Scvis91 | scRNA-seq91 | non-linear | |
| Two-stage variational autoencoder92 | microscopy92 | non-linear | |
| Batch-adversarial variational auto-encoder (BAVARIA)93 | scATAC-seq93 | non-linear | |
| Ivis94 | CyTOF, scRNA-seq94 | non-linear | |
| Cellular in-painting95 | microscopy95 | non-linear | |
| ScSemiGAN96 | scRNA-seq96 | non-linear | |
| scNym97 | scATAC-seq97 | non-linear | |
| Treatment convolutional neural network98 | microscopy98 | non-linear | |
| Cytoself99 | microscopy99 | non-linear |
Statistical/probabilistic methods
Representation learning has strong roots in statistical modeling and can be traced back to Pearson’s 1901 PCA,55 which remains relevant and widely used today. Statistical methods often make strong assumptions, such as the relationships between features being linear. This assumption generates methods that are more robust to noise, outliers in data, and hyperparameter choice. Moreover, the algorithms themselves (and their logical motivations) are often highly interpretable. As an example, probabilistically motivated linear discriminant analysis (LDA) is able to distinguish “recurrent cellular neighborhoods” in skin samples with pre-cancer or cancer histology using multiplexed immunofluorescent imaging, aiding in the identification of discrete cell states whose frequency and proximity varied throughout cancer progression100 (Figure 3A). However, these same assumptions limit the modeling of more complex (non-linear) relationships that are commonly present in single-cell data.82 Recognition of this limitation has motivated the development of non-linear approaches more suited to learning the underlying structure of single-cell datasets. These include both manifold learning methods and those based on deep neural networks.
Figure 3.
Application of example representation learning methods
(A) LDA latent space shows 10 recurring neighborhoods with different cell-type composition signatures in skin cancer samples (adapted from Nirmal et al.100).
(B) UMAP embedding of scRNA-seq data of CD45+ cells shows four clusters of major cell types with a further 23 subgroupings (adapted from Zhao et al.101).
(C) Ivis embedding of scRNA-seq from vascular myeloid cells reveals 12 clusters for hypothesis generation and further investigation of Clec4a2 macrophage regulation (adapted from Park et al.102).
Manifold learning
A common and often useful assumption is that cellular states, and ideally empirical cellular data, lie on a “manifold,” defined as a locally smooth landscape that has relatively lower dimensionality compared to the ambient data space. Such a landscape was famously conceptualized in the 1950s by Conrad Waddington103 as a way to conceive the relationship between cellular state diversity and underlying regulatory mechanisms. Formal validation of the manifold assumption in relation to single-cell data remains challenging. Yet, it is both rational and intuitive that cellular states move in incremental steps within a state space intrinsically compressed through coordinated gene expression.104 Successful applications of manifold learning further empirically support manifold modeling of single-cell data.40,88 For example, Zhao et al. has demonstrated the use of manifold learning using uniform manifold approximation and projection (UMAP) to visualize the top 50 principal components of scRNA-seq data of CD45+ cells, showing clustering of major cell types and further subgroupings in the population101 (Figure 3B). In general, the manifold assumption is less restrictive than assumptions made by statistical methods; hence, manifold learning is thought to better capture true cell state-space structures. However, this flexibility can make the more complex projections learned less interpretable.104
Neural networks
In the single-cell context, neural networks make the least assumptions about underlying data relationships. This implies broad application generalizability, as exemplified by cross-domain use of neural network-based natural language processing techniques in biological sequence data.105 A vast range of different neural network training paradigms and architectures have been proposed, including autoencoders, contrastive learning models, and semi-supervised models (Table 1). As an example, Park et al. demonstrated contrastive learning using ivis on scRNA-seq of vascular myeloid cells, identifying 12 clusters for further analysis, with three sharing resident macrophage genes. Further characterization of these three macrophage clusters revealed that Clec4a2 potentially had a role in modulating macrophage activation and homeostasis, generating a hypothesis for further investigation by the authors102 (Figure 3C). Though the flexibility given by fewer assumptions make them powerful tools capable of learning highly complex relationships, neural networks often require demanding hyperparameter optimization and are more sensitive to variations in data, thus demanding larger datasets for successful model training. What precisely is learned by neural networks may be even less interpretable (than manifold learning methods). As such, neural networks are often considered “black box” models, a limitation widely recognized by the field.
Hyperparameter optimization
Representation methods typically contain adjustable parameters that govern the model fitting process, referred to as hyperparameters. Hyperparameter optimization depends on the chosen tools and can have substantial impacts on the resultant embedding (Figure 4A) Neglecting to optimize hyperparameters will impact downstream performance and can even lead to nonsensical outputs, especially when using hyperparameter-sensitive methods.106 Though the importance of hyperparameters is recognized by some expert users and method developers,107,108 many studies applying existing methods—and even some detailing new methods—describe little hyperparameter optimization. This suggests a limited focus on the key role such optimization may play.109 In many circumstances, such as multi-method benchmarking, researchers rely on default hyperparameter values.45,87 Here, we briefly describe the generalized hyperparameter optimization process and note how different challenges are minimized or exacerbated by different representation learning methods, potentially influencing method suitability for a given user or analysis task.
Figure 4.
Hyperparameter tuning in representation learning
(A) Independent-component analysis (ICA) embeddings of scRNA-seq data. Each embedding reflects a different number of input dimensions, revealing how the dimension hyperparameter influences embedding. With 20, 30, or 50 dimensions, clustering is relatively straightforward compared to using 100, 200, or 300 dimensions. This highlights hyperparameter influence on downstream (in this case, clustering) performance (adapted from Feng et al.106).
(B) The hyperparameter k (number of nearest neighbors; x axes) in the context of different data subsample sizes alters identification accuracy for manual gates in the neural network ivis. Ivis is relatively robust to data subsampling, especially for small values of k. However, the value choice of k has a considerable impact when less data are available (adapted from Szubert et al.94).
(C) Dataset influences on representation learning. Embeddings of scRNA-seq of the mouse neocortex by PCA and t-SNE. In comparison to PCA, t-SNE captures local structure at the expense of global structure (adapted from Kobak and Berens73).
Hyperparameter optimization is often framed as the search for a good, or even the best, set of hyperparameter values from an array of all possible value combinations. This constitutes an exploration of the so-called hyperparameter “search space,” typically done by repeated applications of a representation learning method configured with different hyperparameter values sampled either randomly or exhaustively or by some more sophisticated schemes.110 State-space representations from each model instance undergo comparison to find a set of hyperparameter values producing the most desirable embedding according to available criteria.
Key challenges in optimization include practical restrictions relating to computational costs for hyperparameter search space exploration. Methods with larger hyperparameter numbers or extent define search spaces with many more value combinations. This is especially compounded by interdependencies between hyperparameters, causing the search space to quickly undergo a combinatorial explosion. This can be compounded by the inherent speed and resource demands of a given method, such as memory limits (video RAM) associated with graphical processor unit (GPU) hardware used in neural network approaches. Another challenge reflects the difficulty in assessing and interpreting results to find the most desirable embedding. This is particularly difficult when ground-truth annotations are inaccessible, as further discussed below.
Where researchers have limited capacity for hyperparameter optimization, or a limited basis to evaluate results, it may be advisable to favor methods with limited hyperparameter sensitivity and/or to apply hyperparameter values that are relatively robust to changes in the dataset (e.g., cell population). Such robust characteristics have, for example, been validated for PhenoGraph, markov affinity-based graph imputation of cells (MAGIC), and ivis (Figure 4B), which reduce risks of overfitting and lack of generalizability.40,88,94 In general, the use of less complex and more robust representation learning methods is a practical strategy to achieve reasonable outcomes with a lower risk of defective embeddings, which may not be readily discernible to the user. This robust approach may nonetheless cost peak performance.
Downstream analyses
Ultimately, in the context of single-cell analysis, representation learning is used to derive a data embedding that facilitates one or more downstream analysis tasks. Given that downstream analysis tasks may be substantially different, it follows that the functional utility of any individual embedding may be highly variable depending on the downstream task. For example, clustering and lineage analysis are common downstream tasks performed in single-cell analyses. In the case of clustering, splitting the dataset into informative cell subpopulations is a desirable outcome. Conversely, when viewed through the lens of lineage analysis, an embedding with this discretized character may be undesirable because continuity and smoothness in the embedding are required to map continuous cell-state changes.
It is important to recognize that computational tools are often developed for a specific purpose that will significantly inform (or inhibit) their broader utility. For instance, PCA excels at revealing structures in global cell states but has been noted to struggle to maintain complex local structural details. This is exemplified in the embedding of embryoid body differentiation.82 Similarly, t-distributed stochastic neighbor embedding (t-SNE) was originally intended for visualization of high-dimensional data and hypothesis generation rather than deep analysis72 and, unlike PCA, fragments the global structure of the embedding of mouse neocortex scRNA-seq data (Figure 4C). Repurposing tools for different downstream analysis tasks require consideration and understanding of their original rationale, data-dependent performance, and behavior.
As a general example of this, numerous benchmarking studies on scRNA-seq datasets have sought to define optimal representation learning methods with respect to particular downstream tasks.45,107,87,111,112 Yet despite substantial efforts and rigor, outcomes have been mixed, depending significantly on the underlying data used. This is highlighted by Heisner and Lau, who stress that benchmarking results are highly dataset specific, ultimately questioning the premise of benchmarking efforts purposed to find a “best” method since result generalizability appears minimal.112 Instead, benchmarking typically provides a starting point for further evaluation and validation rather than crowning a winning, one-size-fits-all tool. For the practitioner, methods validated to perform well on similar datasets might best be prioritized where possible. Best practice dictates testing of multiple methods, with careful evaluation and validation of results for each dataset and analysis task.
Evaluation and biological validation
Ground-truth annotations for the biology of interest are often absent in single-cell investigations. Yet, interpreting an analysis depends on understanding the quality of the results. Hence, evaluating learned representations is a crucial, though often challenging, step. Though computational tools may reveal phenomena of interest, biological validation is required to validate indications. We next consider the role of technical metrics and biological validation in such result evaluation.
In some cases, it is possible for domain experts to qualitatively assess embeddings. For instance, single-cell imaging data can sometimes be visually assessed by experts (e.g., using a depiction of images within the learned representation).113 This is demonstrated by Bryce et al., who, after mapping cellular phenotypes in a t-SNE space, examine phenotypic differences in each cluster by inspection of microscopy images7 (Figure 5A) and use this approach to match learned representations to visual perception and thus avoid model over-/under-fitting. For other data modalities, such as scRNA-seq, direct human assessment of high-dimensional RNA expression data is impractical. Validation is thus limited to qualitative expectations, such as those that might pertain to the relative proximity of known cell types114. Qualitative evaluations can be informative but are generally not scalable and lack the quantitative comparative power required for robust optimization. Hence, quantitative metrics are often preferentially used.
Figure 5.
Evaluating embeddings and metrics
(A) Example actin phenotypes induced by chemical compound library screening; representative of numbered clusters in the t-SNE map in (C) (adapted from Bryce et al.7).
(B) FICA and UMAP embeddings of scRNA-seq datasets are shown, with colors representing different cell types. The UMAP embedding compresses and fractures the data considerably more than FICA (adapted from Koch et al.45).
(C) t-SNE phenotypic embedding of 114,400 compounds, with distinct phenotypic clusters numbered. The total number of actin phenotypes (see A) and positive control compound-induced phenotype self-clustering metrics were evaluated against t-SNE perplexity, showing combinatorial use of supervised and unsupervised metrics to evaluate the embedding (adapted from Bryce et al.7).
Quantitative metrics often score embeddings based on specific criteria associated with downstream analysis tasks, including clustering (e.g., cluster compactness, separability, dispersion, and cell-type numbers),115 classification (e.g., accuracy and robustness),87 and trajectory inference analyses (topological similarity, branch assignment, and correlation).116 Thus, metric selection may depend primarily on the expectations of the data. Moreover, metrics can be unsupervised—based solely on the embedding irrespective of any additional annotations—or supervised, assessing the embedding through the lens of available labels like cell type or treatment condition. It is important to keep in mind that although annotations can help interpret an embedding, over-reliance on them can also introduce a bias toward established dogmas. This is especially true when annotations are only weakly related to the phenomenon under investigation.95 For example, Koch et al. show embeddings of scRNA-seq data across multiple datasets using fast independent-component analysis (FICA) and UMAP.45 Though the UMAP embedding is well clustered with respect to the cell-type labels used—yielding a high clustering score—the embedding is severely fractured and compresses the nuance and variability among other morphological properties not measured during the establishment of the cell-type definition (Figure 5B). Ideally, some combination of supervised and unsupervised metrics should be applied, with the former leveraging validation of available ground-truth knowledge to support novel inferences illuminated by the latter. Such an approach was exemplified in a systems microscopy analysis of drug-response phenotypes wherein the number and coherence of known drug phenotypes were used as part of an adversarial strategy to optimize representation-learned embeddings of a library of 114,400 compounds with unknown effects7 (Figure 5C). Through such combinatorial approaches counterposing both supervised and unsupervised metrics, researchers deploying representation learning may more confidently extrapolate from the known to the unknown.
Finally, we mention subsampling evaluation schemes, such as bootstrapping and cross-validation, as a method to evaluate the generalizability of any findings. The consistency of embeddings on subsamples of the data suggests robustness to noise and spurious variations. This provides a heuristic for model reliability, as demonstrated by Moon et al., who show the robustness of potential of heat-diffusion affinity-based transition embedding (PHATE) in subsampling of Splatter-simulated scRNA-seq data.82
Conclusion and future outlook
Single-cell technologies continue to mature, enriching and transforming our understanding of the cell. As increasingly powerful biological data measurement methods are applied, the role of computational tools for enabling complex data interpretation becomes ever more central. Representation learning is already widely used in the biomedical community for this purpose, yet it remains a complex challenge to navigate the selection, optimization, application, and comparison of representation learning methods for single-cell analyses. Ever-expanding applications are driving an explosion of new methods encompassing more nuanced understandings of the cell and increasingly sophisticated computational processes. Thus, the complexity involved in applying representation learning for single-cell analysis continues to escalate. Here, we hope to have provided a broad roadmap covering the end-to-end process of representation learning, helping to orient new practitioners in the field.
In this exciting era, new perspectives on, and aspects of, the cell are being revealed by proliferating modalities of data measurement. Representation learning lies at the intersection of these data modalities, providing the means to achieve a mature and integrated characterization of the cell. In particular, very diverse omics and imaging data types require merging in a way that is interpretable for further analyses. Excitingly, representation learning techniques have recently been applied to graph representations of single-cell relationships, such as molecular similarity graphs or gene regulatory networks.117 The ubiquity of graph-style relationships and the ability to integrate them in single-cell investigations suggest a powerful capability. Thus, active research continues to develop tools that generalize and scale well,53 helping biomedical researchers obtain systemic views incorporating unique, unbiased, coherent perspectives. With the stage set for transformational advances in our understanding of the cell, the need for biomedical researchers to understand, select, and effectively deploy representation learning methods becomes an ever more significant priority.
Acknowledgments
I.G. is supported by an Australian Government Research Training Program (RTP) Scholarship. J.G.L. is supported by a Ramaciotti Biomedical Research Award, an ARC Development Project grant (DP170103599), NHMRC Ideas Grants (GNT1184009 and GNT2012848), and a Tour de Cure Pioneering Grant (RSP-547-FY2023).
Author contributions
All authors were involved in paper concept generation. Design of the manuscript was led by I.G. and J.G.L. This work was written primarily by I.G., with the final manuscript edited and approved by all authors.
Declaration of interests
F.V. declares a relationship with OmniOmics.AI Pty., Ltd., which may involve financial interests or other forms of cooperative ventures.
Inclusion and diversity
We support inclusive, diverse, and equitable conduct of research.
References
- 1.Burkhardt D.B., San Juan B.P., Lock J.G., Krishnaswamy S., Chaffer C.L. Mapping Phenotypic Plasticity upon the Cancer Cell State Landscape Using Manifold Learning. Cancer Discov. 2022;12:1847–1859. doi: 10.1158/2159-8290.CD-21-0282. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Altschuler S.J., Wu L.F. Cellular heterogeneity: do differences make a difference? Cell. 2010;141:559–563. doi: 10.1016/j.cell.2010.04.033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Chung W., Eum H.H., Lee H.-O., Lee K.-M., Lee H.-B., Kim K.-T., Ryu H.S., Kim S., Lee J.E., Park Y.H., et al. Single-cell RNA-seq enables comprehensive tumour and immune cell profiling in primary breast cancer. Nat. Commun. 2017;8:15081. doi: 10.1038/ncomms15081. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Chattopadhyay P.K., Gierahn T.M., Roederer M., Love J.C. Single-cell technologies for monitoring immune systems. Nat. Immunol. 2014;15:128–135. doi: 10.1038/ni.2796. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Marioni J.C., Arendt D. How single-cell genomics is changing evolutionary and developmental biology. Annu. Rev. Cell Dev. Biol. 2017;33:537–553. doi: 10.1146/annurev-cellbio-100616-060818. [DOI] [PubMed] [Google Scholar]
- 6.Tasic B. Single cell transcriptomics in neuroscience: cell classification and beyond. Curr. Opin. Neurobiol. 2018;50:242–249. doi: 10.1016/j.conb.2018.04.021. [DOI] [PubMed] [Google Scholar]
- 7.Bryce N.S., Failes T.W., Stehn J.R., Baker K., Zahler S., Arzhaeva Y., Bischof L., Lyons C., Dedova I., Arndt G.M., et al. High-Content Imaging of Unbiased Chemical Perturbations Reveals that the Phenotypic Plasticity of the Actin Cytoskeleton Is Constrained. Cell Syst. 2019;9:496–507.e5. doi: 10.1016/j.cels.2019.09.002. [DOI] [PubMed] [Google Scholar]
- 8.Marx V. A dream of single-cell proteomics. Nat. Methods. 2019;16:809–812. doi: 10.1038/s41592-019-0540-6. [DOI] [PubMed] [Google Scholar]
- 9.Duncan K.D., Fyrestam J., Lanekoff I. Advances in mass spectrometry based single-cell metabolomics. Analyst. 2019;144:782–793. doi: 10.1039/c8an01581c. [DOI] [PubMed] [Google Scholar]
- 10.Ziffra R.S., Kim C.N., Ross J.M., Wilfert A., Turner T.N., Haeussler M., Casella A.M., Przytycki P.F., Keough K.C., Shin D., et al. Single-cell epigenomics reveals mechanisms of human cortical development. Nature. 2021;598:205–213. doi: 10.1038/s41586-021-03209-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Aldridge S., Teichmann S.A. Single cell transcriptomics comes of age. Nat. Commun. 2020;11:4307. doi: 10.1038/s41467-020-18158-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Lock J.G., Strömblad S. Systems microscopy: an emerging strategy for the life sciences. Exp. Cell Res. 2010;316:1438–1444. doi: 10.1016/j.yexcr.2010.04.001. [DOI] [PubMed] [Google Scholar]
- 13.Hériché J.K., Alexander S., Ellenberg J. Integrating imaging and omics: Computational methods and challenges. Annu. Rev. Biomed. Data Sci. 2019;2:175–197. [Google Scholar]
- 14.Karacosta L.G. From imaging a single cell to implementing precision medicine: an exciting new era. Emerg. Top. Life Sci. 2021;5:837–847. doi: 10.1042/ETLS20210219. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Antonelli L., Guarracino M.R., Maddalena L., Sangiovanni M. Integrating imaging and omics data: a review. Biomed. Signal Process Control. 2019;52:264–280. [Google Scholar]
- 16.Watson E.R., Taherian Fard A., Mar J.C. Computational methods for single-cell imaging and omics data integration. Front. Mol. Biosci. 2021;8:768106. doi: 10.3389/fmolb.2021.768106. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Hie B., Peters J., Nyquist S.K., Shalek A.K., Berger B., Bryson B.D. Computational methods for single-cell RNA sequencing. Annu. Rev. Biomed. Data Sci. 2020;3:339. [Google Scholar]
- 18.Newell E.W., Cheng Y. Mass cytometry: blessed with the curse of dimensionality. Nat. Immunol. 2016;17:890–895. doi: 10.1038/ni.3485. [DOI] [PubMed] [Google Scholar]
- 19.Wu Y., Zhang K. Tools for the analysis of high-dimensional single-cell RNA sequencing data. Nat. Rev. Nephrol. 2020;16:408–421. doi: 10.1038/s41581-020-0262-0. [DOI] [PubMed] [Google Scholar]
- 20.Kulkarni A., Anderson A.G., Merullo D.P., Konopka G. Beyond bulk: a review of single cell transcriptomics methodologies and applications. Curr. Opin. Biotechnol. 2019;58:129–136. doi: 10.1016/j.copbio.2019.03.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Domingos P. A few useful things to know about machine learning. Commun. ACM. 2012;55:78–87. [Google Scholar]
- 22.Regev A., Teichmann S.A., Lander E.S., Amit I., Benoist C., Birney E., Bodenmiller B., Campbell P., Carninci P., Clatworthy M., et al. The human cell atlas. Elife. 2017;6 doi: 10.7554/eLife.27041. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Segal E., Shapira M., Regev A., Pe’er D., Botstein D., Koller D., Friedman N. Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data. Nat. Genet. 2003;34:166–176. doi: 10.1038/ng1165. [DOI] [PubMed] [Google Scholar]
- 24.Zhu L., Lei J., Klei L., Devlin B., Roeder K. Semisoft clustering of single-cell data. Proc. Natl. Acad. Sci. USA. 2019;116:466–471. doi: 10.1073/pnas.1817715116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Alquicira-Hernandez J., Sathe A., Ji H.P., Nguyen Q., Powell J.E. scPred: accurate supervised method for cell-type classification from single-cell RNA-seq data. Genome Biol. 2019;20:264. doi: 10.1186/s13059-019-1862-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Trapnell C., Cacchiarelli D., Grimsby J., Pokharel P., Li S., Morse M., Lennon N.J., Livak K.J., Mikkelsen T.S., Rinn J.L. The dynamics and regulators of cell fate decisions are revealed by pseudotemporal ordering of single cells. Nat. Biotechnol. 2014;32:381–386. doi: 10.1038/nbt.2859. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Linderman G.C. Dimensionality reduction of single-cell RNA-seq data. Methods Mol. Biol. 2021;2284:331–342. doi: 10.1007/978-1-0716-1307-8_18. [DOI] [PubMed] [Google Scholar]
- 28.Toth T., Bauer D., Sukosd F., Horvath P. Fisheye transformation enhances deep-learning-based single-cell phenotyping by including cellular microenvironment. Cell Rep. Methods. 2022;2 doi: 10.1016/j.crmeth.2022.100339. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Carpenter A.E., Jones T.R., Lamprecht M.R., Clarke C., Kang I.H., Friman O., Guertin D.A., Chang J.H., Lindquist R.A., Moffat J., et al. CellProfiler: image analysis software for identifying and quantifying cell phenotypes. Genome Biol. 2006;7:R100. doi: 10.1186/gb-2006-7-10-r100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Ljosa V., Caie P.D., ter Horst R., Sokolnicki K.L., Jenkins E.L., Daya S., Roberts M.E., Jones T.R., Singh S., Genovesio A., et al. Comparison of Methods for Image-Based Profiling of Cellular Morphological Responses to Small-Molecule Treatment. J. Biomol. Screen. 2013;18:1321–1329. doi: 10.1177/1087057113503553. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Shafqat-Abbasi H., Kowalewski J.M., Kiss A., Gong X., Hernandez-Varas P., Berge U., Jafari-Mamaghani M., Lock J.G., Strömblad S. An analysis toolbox to explore mesenchymal migration heterogeneity reveals adaptive switching between distinct modes. Elife. 2016;5 doi: 10.7554/eLife.11384. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Kowalewski J.M., Shafqat-Abbasi H., Jafari-Mamaghani M., Endrias Ganebo B., Gong X., Strömblad S., Lock J.G. Disentangling Membrane Dynamics and Cell Migration; Differential Influences of F-actin and Cell-Matrix Adhesions. PLoS One. 2015;10 doi: 10.1371/journal.pone.0135204. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Zandavi S.M., Liu D., Chung V., Anaissi A., Vafaee F. Fotomics: fourier transform-based omics imagification for deep learning-based cell-identity mapping using single-cell omics profiles. Artif. Intell. Rev. 2022;56:7263–7278. doi: 10.1007/s10462-022-10357-4. [DOI] [Google Scholar]
- 34.Sharma A., Vans E., Shigemizu D., Boroevich K.A., Tsunoda T. DeepInsight: A methodology to transform a non-image data to an image for convolution neural network architecture. Sci. Rep. 2019;9 doi: 10.1038/s41598-019-47765-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Gayoso A., Lopez R., Xing G., Boyeau P., Valiollah Pour Amiri V., Hong J., Wu K., Jayasuriya M., Mehlman E., Langevin M., et al. A Python library for probabilistic analysis of single-cell omics data. Nat. Biotechnol. 2022;40:163–166. doi: 10.1038/s41587-021-01206-w. [DOI] [PubMed] [Google Scholar]
- 36.Satija R., Farrell J.A., Gennert D., Schier A.F., Regev A. Spatial reconstruction of single-cell gene expression data. Nat. Biotechnol. 2015;33:495–502. doi: 10.1038/nbt.3192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Smith K., Li Y., Piccinini F., Csucs G., Balazs C., Bevilacqua A., Horvath P. CIDRE: an illumination-correction method for optical microscopy. Nat. Methods. 2015;12:404–406. doi: 10.1038/nmeth.3323. [DOI] [PubMed] [Google Scholar]
- 38.Yin Z., Kanade T., Chen M. Understanding the phase contrast optics to restore artifact-free microscopy images for segmentation. Med. Image Anal. 2012;16:1047–1062. doi: 10.1016/j.media.2011.12.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Jiang R., Sun T., Song D., Li J.J. Statistics or biology: the zero-inflation controversy about scRNA-seq data. Genome Biol. 2022;23 doi: 10.1186/s13059-022-02601-5. 31–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Van Dijk D., Sharma R., Nainys J., Yim K., Kathail P., Carr A.J., Burdziak C., Moon K.R., Chaffer C.L., Pattabiraman D., et al. Recovering gene interactions from single-cell data using data diffusion. Cell. 2018;174:716–729.e27. doi: 10.1016/j.cell.2018.05.061. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Eraslan G., Simon L.M., Mircea M., Mueller N.S., Theis F.J. Single-cell RNA-seq denoising using a deep count autoencoder. Nat. Commun. 2019;10:390. doi: 10.1038/s41467-018-07931-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Wang J., Agarwal D., Huang M., Hu G., Zhou Z., Ye C., Zhang N.R. Data denoising with transfer learning in single-cell transcriptomics. Nat. Methods. 2019;16:875–878. doi: 10.1038/s41592-019-0537-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Patruno L., Maspero D., Craighero F., Angaroni F., Antoniotti M., Graudenzi A. A review of computational strategies for denoising and imputation of single-cell transcriptomic data. Briefings Bioinf. 2021;22 doi: 10.1093/bib/bbaa222. [DOI] [PubMed] [Google Scholar]
- 44.Ashhurst T.M., Marsh-Wakefield F., Putri G.H., Spiteri A.G., Shinko D., Read M.N., Smith A.L., King N.J.C. Integration, exploration, and analysis of high-dimensional single-cell cytometry data using Spectre. Cytometry A. 2022;101:237–253. doi: 10.1002/cyto.a.24350. [DOI] [PubMed] [Google Scholar]
- 45.Koch F.C., Sutton G.J., Voineagu I., Vafaee F. Supervised application of internal validation measures to benchmark dimensionality reduction methods in scRNA-seq data. Briefings Bioinf. 2021;22:bbab304. doi: 10.1093/bib/bbab304. [DOI] [PubMed] [Google Scholar]
- 46.Zhang Y., Parmigiani G., Johnson W.E. ComBat-seq: batch effect adjustment for RNA-seq count data. NAR Genom. Bioinform. 2020;2:lqaa078. doi: 10.1093/nargab/lqaa078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Argelaguet R., Cuomo A.S.E., Stegle O., Marioni J.C. Computational principles and challenges in single-cell data integration. Nat. Biotechnol. 2021;39:1202–1215. doi: 10.1038/s41587-021-00895-7. [DOI] [PubMed] [Google Scholar]
- 48.Lee J., Hyeon D.Y., Hwang D. Single-cell multiomics: technologies and data analysis methods. Exp. Mol. Med. 2020;52:1428–1442. doi: 10.1038/s12276-020-0420-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Yang K.D., Belyaeva A., Venkatachalapathy S., Damodaran K., Katcoff A., Radhakrishnan A., Shivashankar G. v, Uhler C. Multi-domain translation between single-cell imaging and sequencing data using autoencoders. Nat. Commun. 2021;12:31. doi: 10.1038/s41467-020-20249-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Welch J.D., Hartemink A.J., Prins J.F. MATCHER: manifold alignment reveals correspondence between single cell transcriptome and epigenome dynamics. Genome Biol. 2017;18:138. doi: 10.1186/s13059-017-1269-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Liu J., Huang Y., Singh R., Vert J.-P., Noble W.S. Algorithms in bioinformatics:… International Workshop, WABI…, proceedings. WABI (Workshop) (NIH Public Access) 2019. Jointly embedding multiple single-cell omics measurements. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Singh A., Reinders M., Mahfouz A., Abdelaal T. TopoGAN: Unsupervised manifold alignment of single-cell data. bioRxiv. 2022 doi: 10.1101/2022.04.27.489829. Prepint at. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Xu Y., McCord R.P. Diagonal integration of multimodal single-cell data: potential pitfalls and paths forward. Nat. Commun. 2022;13:3505. doi: 10.1038/s41467-022-31104-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Luecken M.D., Büttner M., Chaichoompu K., Danese A., Interlandi M., Mueller M.F., Strobl D.C., Zappia L., Dugas M., Colomé-Tatché M., Theis F.J. Benchmarking atlas-level data integration in single-cell genomics. Nat. Methods. 2022;19:41–50. doi: 10.1038/s41592-021-01336-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Pearson K. LIII. On lines and planes of closest fit to systems of points in space. London, Edinburgh Dublin Phil. Mag. J. Sci. 1901;2:559–572. [Google Scholar]
- 56.Buettner F., Moignard V., Göttgens B., Theis F.J. Probabilistic PCA of censored data: accounting for uncertainties in the visualization of high-throughput single-cell qPCR data. Bioinformatics. 2014;30:1867–1875. doi: 10.1093/bioinformatics/btu134. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.MoradiAmin M., Samadzadehaghdam N., Kermani S., Talebi A. Enhanced Recognition of Acute Lymphoblastic Leukemia Cells in Microscopic Images based on Feature Reduction using Principal Component Analysis. Frontiers in Biomedical Technologies. 2015;2 [Google Scholar]
- 58.Wang Y.-X., Zhang Y.-J. Nonnegative matrix factorization: A comprehensive review. IEEE Trans. Knowl. Data Eng. 2013;25:1336–1353. [Google Scholar]
- 59.Shao C., Höfer T. Robust classification of single-cell transcriptome data by nonnegative matrix factorization. Bioinformatics. 2017;33:235–242. doi: 10.1093/bioinformatics/btw607. [DOI] [PubMed] [Google Scholar]
- 60.Wu Y., Tamayo P., Zhang K. Visualizing and interpreting single-cell gene expression datasets with similarity weighted nonnegative embedding. Cell Syst. 2018;7:656–666.e4. doi: 10.1016/j.cels.2018.10.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Klema V., Laub A. The singular value decomposition: Its computation and some applications. IEEE Trans. Automat. Control. 1980;25:164–176. [Google Scholar]
- 62.Dumais S.T. Latent semantic analysis. Annu. Rev. Inf. Sci. Technol. 2005;38:188–230. doi: 10.1002/aris.1440380105. [DOI] [Google Scholar]
- 63.Granja J.M., Corces M.R., Pierce S.E., Bagdatli S.T., Choudhry H., Chang H.Y., Greenleaf W.J. ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis. Nat. Genet. 2021;53:403–411. doi: 10.1038/s41588-021-00790-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Satpathy A.T., Granja J.M., Yost K.E., Qi Y., Meschi F., McDermott G.P., Olsen B.N., Mumbach M.R., Pierce S.E., Corces M.R., et al. Massively parallel single-cell chromatin landscapes of human immune cell development and intratumoral T cell exhaustion. Nat. Biotechnol. 2019;37:925–936. doi: 10.1038/s41587-019-0206-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Hyvärinen A. Independent component analysis: recent advances. Philos. Trans. A Math. Phys. Eng. Sci. 2013;371 doi: 10.1098/rsta.2011.0534. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Pierson E., Yau C. ZIFA: Dimensionality reduction for zero-inflated single-cell gene expression analysis. Genome Biol. 2015;16:241. doi: 10.1186/s13059-015-0805-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Blei D.M., Ng A.Y., Jordan M.I. Latent dirichlet allocation. J. Mach. Learn. Res. 2003;3:993–1022. [Google Scholar]
- 68.Wu X., Wu H., Wu Z. Penalized Latent Dirichlet Allocation Model in Single-Cell RNA Sequencing. Stat. Biosci. 2021;13:543–562. [Google Scholar]
- 69.Bravo González-Blas C., Minnoye L., Papasokrati D., Aibar S., Hulselmans G., Christiaens V., Davie K., Wouters J., Aerts S. cisTopic: cis-regulatory topic modeling on single-cell ATAC-seq data. Nat. Methods. 2019;16:397–400. doi: 10.1038/s41592-019-0367-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Lawrence N. Gaussian process latent variable models for visualisation of high dimensional data. Adv. Neural Inf. Process. Syst. 2003;16 [Google Scholar]
- 71.Lönnberg T., Svensson V., James K.R., Fernandez-Ruiz D., Sebina I., Montandon R., Soon M.S.F., Fogg L.G., Nair A.S., Liligeto U., et al. Single-cell RNA-seq and computational analysis using temporal mixture modeling resolves TH1/TFH fate bifurcation in malaria. Sci. Immunol. 2017;2 doi: 10.1126/sciimmunol.aal2192. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.van der Maaten L., Hinton G. Visualizing data using t-SNE. Journal of machine learning research. 2008;9 [Google Scholar]
- 73.Kobak D., Berens P. The art of using t-SNE for single-cell transcriptomics. Nat. Commun. 2019;10:5416. doi: 10.1038/s41467-019-13056-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Wang S., Zhou Y., Qin X., Nair S., Huang X., Liu Y. Label-free detection of rare circulating tumor cells by image analysis and machine learning. Sci. Rep. 2020;10:12226. doi: 10.1038/s41598-020-69056-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Tenenbaum J.B., de Silva V., Langford J.C. A global geometric framework for nonlinear dimensionality reduction. Science. 2000;290:2319–2323. doi: 10.1126/science.290.5500.2319. [DOI] [PubMed] [Google Scholar]
- 76.Chen Y., Zhang Y., Ouyang Z. BIOCOMPUTING 2019: Proceedings of the Pacific Symposium. World Scientific); 2018. LISA: accurate reconstruction of cell trajectory and pseudo-time for massive single cell RNA-seq data; pp. 338–349. [PMC free article] [PubMed] [Google Scholar]
- 77.Roweis S.T., Saul L.K. Nonlinear dimensionality reduction by locally linear embedding. Science. 2000;290:2323–2326. doi: 10.1126/science.290.5500.2323. [DOI] [PubMed] [Google Scholar]
- 78.Welch J.D., Hartemink A.J., Prins J.F. SLICER: inferring branched, nonlinear cellular trajectories from single cell RNA-seq data. Genome Biol. 2016;17:106–115. doi: 10.1186/s13059-016-0975-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.McInnes L., Healy J., Melville J. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv. 2018 doi: 10.48550/arXiv.1802.03426. Preprint at. [DOI] [Google Scholar]
- 80.Becht E., McInnes L., Healy J., Dutertre C.-A., Kwok I.W.H., Ng L.G., Ginhoux F., Newell E.W. Dimensionality reduction for visualizing single-cell data using UMAP. Nat. Biotechnol. 2019;37:38–44. doi: 10.1038/nbt.4314. [DOI] [PubMed] [Google Scholar]
- 81.Hillsley A., Santoso M.S., Engels S.M., Halwachs K.N., Contreras L.M., Rosales A.M. A strategy to quantify myofibroblast activation on a continuous spectrum. Sci. Rep. 2022;12:12239. doi: 10.1038/s41598-022-16158-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Moon K.R., van Dijk D., Wang Z., Gigante S., Burkhardt D.B., Chen W.S., Yim K., Elzen A.v.d., Hirn M.J., Coifman R.R., et al. Visualizing structure and transitions in high-dimensional biological data. Nat. Biotechnol. 2019;37:1482–1492. doi: 10.1038/s41587-019-0336-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Belkin M., Niyogi P. Laplacian eigenmaps for dimensionality reduction and data representation. Neural Comput. 2003;15:1373–1396. [Google Scholar]
- 84.Campbell K., Ponting C.P., Webber C. Laplacian eigenmaps and principal curves for high resolution pseudotemporal ordering of single-cell RNA-seq profiles. bioRxiv. 2015 doi: 10.1101/027219. Preprint at. [DOI] [Google Scholar]
- 85.Haghverdi L., Buettner F., Theis F.J. Diffusion maps for high-dimensional single-cell analysis of differentiation data. Bioinformatics. 2015;31:2989–2998. doi: 10.1093/bioinformatics/btv325. [DOI] [PubMed] [Google Scholar]
- 86.Zhang Z.y., Zha H.y. Principal manifolds and nonlinear dimensionality reduction via tangent space alignment. J. Shanghai Univ. 2004;8:406–424. [Google Scholar]
- 87.Sun S., Zhu J., Ma Y., Zhou X. Accuracy, robustness and scalability of dimensionality reduction methods for single-cell RNA-seq analysis. Genome Biol. 2019;20:269. doi: 10.1186/s13059-019-1898-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Levine J.H., Simonds E.F., Bendall S.C., Davis K.L., Amir E.a.D., Tadmor M.D., Litvin O., Fienberg H.G., Jager A., Zunder E.R., et al. Data-driven phenotypic dissection of AML reveals progenitor-like cells that correlate with prognosis. Cell. 2015;162:184–197. doi: 10.1016/j.cell.2015.05.047. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Amodio M., van Dijk D., Srinivasan K., Chen W.S., Mohsen H., Moon K.R., Campbell A., Zhao Y., Wang X., Venkataswamy M., et al. Exploring single-cell data with deep multitasking neural networks. Nat. Methods. 2019;16:1139–1145. doi: 10.1038/s41592-019-0576-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Wang D., Gu J. VASC: dimension reduction and visualization of single-cell RNA-seq data by deep variational autoencoder. Dev. Reprod. Biol. 2018;16:320–331. doi: 10.1016/j.gpb.2018.08.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Ding J., Condon A., Shah S.P. Interpretable dimensionality reduction of single cell transcriptome data with deep generative models. Nat. Commun. 2018;9:2002–2013. doi: 10.1038/s41467-018-04368-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Wei W., Haidinger S., Lock J., Meijering E. International Workshop on Machine Learning in Medical Imaging. Springer; 2021. Deep Representation Learning for Image-Based Cell Profiling; pp. 487–497. [Google Scholar]
- 93.Kopp W., Akalin A., Ohler U. Simultaneous dimensionality reduction and integration for single-cell ATAC-seq data using deep learning. Nat. Mach. Intell. 2022;4:162–168. [Google Scholar]
- 94.Szubert B., Cole J.E., Monaco C., Drozdov I. Structure-preserving visualisation of high dimensional single-cell datasets. Sci. Rep. 2019;9:8914. doi: 10.1038/s41598-019-45301-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Lu A.X., Kraus O.Z., Cooper S., Moses A.M. Learning unsupervised feature representations for single cell microscopy images with paired cell inpainting. PLoS Comput. Biol. 2019;15 doi: 10.1371/journal.pcbi.1007348. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 96.Xu Z., Luo J., Xiong Z. scSemiGAN: a single-cell semi-supervised annotation and dimensionality reduction framework based on generative adversarial network. Bioinformatics. 2022;38:5042–5048. doi: 10.1093/bioinformatics/btac652. [DOI] [PubMed] [Google Scholar]
- 97.Kimmel J.C., Kelley D.R. Semisupervised adversarial neural networks for single-cell classification. Genome Res. 2021;31:1781–1793. doi: 10.1101/gr.268581.120. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 98.Caicedo J.C., McQuin C., Goodman A., Singh S., Carpenter A.E. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. Weakly supervised learning of single-cell feature embeddings; pp. 9309–9318. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Kobayashi H., Cheveralls K.C., Leonetti M.D., Royer L.A. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nat. Methods. 2022;19:995–1003. doi: 10.1038/s41592-022-01541-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Nirmal A.J., Maliga Z., Vallius T., Quattrochi B., Chen A.A., Jacobson C.A., Pelletier R.J., Yapp C., Arias-Camison R., Chen Y.-A., et al. The spatial landscape of progression and immunoediting in primary melanoma at single cell resolution. Cancer Discov. 2022;12:1518–1541. doi: 10.1158/2159-8290.CD-21-1357. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 101.Zhao J., Zhang S., Liu Y., He X., Qu M., Xu G., Wang H., Huang M., Pan J., Liu Z., et al. Single-cell RNA sequencing reveals the heterogeneity of liver-resident immune cells in human. Cell Discov. 2020;6:22. doi: 10.1038/s41421-020-0157-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Park I., Goddard M.E., Cole J.E., Zanin N., Lyytikäinen L.P., Lehtimäki T., Andreakos E., Feldmann M., Udalova I., Drozdov I., Monaco C. C-type lectin receptor CLEC4A2 promotes tissue adaptation of macrophages and protects against atherosclerosis. Nat. Commun. 2022;13:215. doi: 10.1038/s41467-021-27862-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Waddington C.H. Canalization of development and the inheritance of acquired characters. Nature. 1942;150:563–565. [Google Scholar]
- 104.Moon K.R., Stanley J.S., III, Burkhardt D., van Dijk D., Wolf G., Krishnaswamy S. Manifold learning-based methods for analyzing single-cell RNA-sequencing data. Curr. Opin. Struct. Biol. 2018;7:36–46. [Google Scholar]
- 105.Iuchi H., Matsutani T., Yamada K., Iwano N., Sumi S., Hosoda S., Zhao S., Fukunaga T., Hamada M. Representation learning applications in biological sequence analysis. Comput. Struct. Biotechnol. J. 2021;19:3198–3208. doi: 10.1016/j.csbj.2021.05.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 106.Feng C., Liu S., Zhang H., Guan R., Li D., Zhou F., Liang Y., Feng X. Dimension reduction and clustering models for single-cell RNA sequencing data: a comparative study. Int. J. Mol. Sci. 2020;21:2181. doi: 10.3390/ijms21062181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 107.Xiang R., Wang W., Yang L., Wang S., Xu C., Chen X. A comparison for dimensionality reduction methods of single-cell RNA-seq data. Front. Genet. 2021;12 doi: 10.3389/fgene.2021.646936. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 108.Wang Y., Huang H., Rudin C., Shaposhnik Y. Understanding How Dimension Reduction Tools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMap, and PaCMAP for Data Visualization. J. Mach. Learn. Res. 2021;22:1–73. [Google Scholar]
- 109.Hu Q., Greene C.S. BIOCOMPUTING 2019: Proceedings of the Pacific Symposium. World Scientific); 2018. Parameter tuning is a key part of dimensionality reduction via deep variational autoencoders for single cell RNA transcriptomics; pp. 362–373. [PMC free article] [PubMed] [Google Scholar]
- 110.Yang L., Shami A. On hyperparameter optimization of machine learning algorithms: Theory and practice. Neurocomputing. 2020;415:295–316. [Google Scholar]
- 111.Tsuyuzaki K., Sato H., Sato K., Nikaido I. Benchmarking principal component analysis for large-scale single-cell RNA-sequencing. Genome Biol. 2020;21:9–17. doi: 10.1186/s13059-019-1900-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 112.Heiser C.N., Lau K.S. A Quantitative Framework for Evaluating Single-Cell Data Structure Preservation by Dimensionality Reduction Techniques. Cell Rep. 2020;31 doi: 10.1016/j.celrep.2020.107576. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 113.Lock J.G., Filonik D., Lawther R., Pather N., Gaus K., Kenderdine S., Bednarz T. Proceedings of the 16th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and its Applications in Industry. 2018. Visual analytics of single cell microscopy data using a collaborative immersive environment; pp. 1–4. [Google Scholar]
- 114.Zandavi S.M., Koch F.C., Vijayan A., Zanini F., Mora F.V., Ortega D.G., Vafaee F. Disentangling single-cell omics representation with a power spectral density-based feature extraction. Nucleic Acids Res. 2022;50:5482–5492. doi: 10.1093/nar/gkac436. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 115.Yu L., Cao Y., Yang J.Y.H., Yang P. Benchmarking clustering algorithms on estimating the number of cell types from single-cell RNA-sequencing data. Genome Biol. 2022;23 doi: 10.1186/s13059-022-02622-0. 49–21. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 116.Saelens W., Cannoodt R., Todorov H., Saeys Y. A comparison of single-cell trajectory inference methods. Nat. Biotechnol. 2019;37:547–554. doi: 10.1038/s41587-019-0071-9. [DOI] [PubMed] [Google Scholar]
- 117.Hetzel L., Fischer D.S., Günnemann S., Theis F.J. Graph representation learning for single-cell biology. Curr. Opin. Struct. Biol. 2021;28 doi: 10.1016/j.coisb.2021.05.008. [DOI] [Google Scholar]





