Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2022 Aug 5.
Published in final edited form as: Methods Mol Biol. 2022;2399:21–60. doi: 10.1007/978-1-0716-1831-8_3

Single-Cell Analysis of the Transcriptome and Epigenome

Krystyna Mazan-Mamczarz, Jisu Ha, Supriyo De, Payel Sen
PMCID: PMC9352558  NIHMSID: NIHMS1825971  PMID: 35604552

Abstract

Epigenome regulation has emerged as an important mechanism for the maintenance of organ function in health and disease. Dissecting epigenomic alterations and resultant gene expression changes in single cells provides unprecedented resolution and insight into cellular diversity, modes of gene regulation, transcription factor dynamics and 3D genome organization. In this chapter, we summarize the transformative single-cell epigenomic technologies that have deepened our understanding of the fundamental principles of gene regulation. We provide a historical perspective of these methods, brief procedural outline with emphasis on the computational tools used to meaningfully dissect information. Our overall goal is to aid scientists using these technologies in their favorite system of interest.

Keywords: Single-cell, Epigenome, Transcriptome, Multiomics

1. Introduction

With the completion of the human genome in 2003 [1], it was anticipated that in the following decade, scientists would have solutions to most major diseases afflicting humans. Although much was gleaned from the sequencing information, the returns on investment were relatively modest. This was primarily due to the faulty assumption that common genetic variants caused all human diseases. Nevertheless, it ushered in the search for single nucleotide polymorphisms (SNPs) in individuals, accelerated genome-wide association studies (GWAS) which attributed SNPs to disease risk, and whole-exome sequencing that sequenced the coding exons of all genes. Unfortunately, despite the assembly of large SNP atlases such as the SNP Consortium [2], the International HapMap Project [3] and the 1000 Genomes Project [4], only a few SNPs could be associated causally with human disease with <10% explaining heritability and almost none the underlying biology. Furthermore, SNPs largely localized to intergenic regions, many kilobases away from a gene and almost nothing at the time was known about the function of nongenic “junk” DNA. It became clear that human diseases were complex.

To explain this apparent disconnect between genetic variants and disease causation, work during the early/mid 2000s, also focused on epigenetic regulation of genetic information, a way to link environmental exposures to phenotypic traits. Epigenetics comprised the study of chemical modifications on DNA and histone proteins, chromatin accessibility, and RNA transcript measurements. Epigenetic mechanisms were found to be transgenerational, lending some credence to Lamarckism. The NIH Roadmap Epigenomics Mapping Consortium [5], the Encyclopedia of DNA Elements (ENCODE) project [6], the related model organism ENCODE (modENCODE) project [7] and the International Human Epigenome Consortium (IHEC) [8] were thus launched to provide reference epigenomes generated using standardized protocols.

The challenge to deciphering useful information from ENCODE/modENCODE studies was to relate data from cell lines or model organisms to actual human tissues. Given the interindividual variability in human populations, it was difficult to match timing, staging and cell types. Furthermore, the ENCODE project constituted bulk or average measurements and it was realized that ensemble averages can distort the understanding of individual cell responses. Human tissues themselves present a highly heterogenous environment that is further altered with disease. These gaps in understanding fueled the launch of the Human Cell Atlas (HCA) Consortium in 2017 [9], that in the next few years will provide comprehensive reference maps of all cells of the human body. In part, revolutionary advances in single cell preparation protocols, sensitive detection of biomolecules, sophisticated fluidics for single cell capture, clever multiplexing strategies and importantly, computational pipelines for rapid parsing of sparse data has greatly catalyzed this initiative. For single cell preparation methods, we direct the readers to a comprehensive review by the Regev lab [10]. In this chapter, we focus on summarizing the various methods for single-cell transcriptomic and epigenomic approaches and outline the computational workflow and tools for analysis. See Table 1 for a glossary of terms used in this chapter.

Table 1.

Glossary of terms

Term Description
10 × Genomics A biotechnology company that designs and manufactures single-cell fluidics and kits for genomics, transcriptomics, epigenomics etc.
3C Chromosome Conformation Capture.
Alevin A pseudoalignment-based RNA-seq quantification program.
ArchR An end-to-end R package for analysis of scATAC-seq data.
ArchR projects Objects in ArchR to associate Arrow files together into a single analytical framework in R.
Arrow files The base unit of an analytical project in ArchR that stores all of the metadata associated with an individual sample.
ASAP-seq Select Antigen Profiling by sequencing.
ATAC-RNA-seq Simultaneous profiling of chromatin accessibility by ATAC-seq and transcriptome by scRNA-seq in single cells.
ATAC-seq Assay for Transposase Accessible Chromatin sequencing.
Autoencoder A type of artificial neural network-based learning.
AutoImpute Python package for analysis and implementation of imputation methods.
BAM Binary Alignment Map, a compressed binary format for storing sequence data.
BRIE Bayesian Regression for Isoform Estimation.
CCA Canonical Correlation Analysis.
CCAN Cis-coaccessibility networks which are hubs of coaccessibility identified utilizing Cicero.
Cell Hashing A sample multiplexing method with oligo tagged antibodies directed against cell surface proteins based on the concept of hash functions in computer science to index datasets with specific features.
Cell Ranger A set of analysis pipelines for processing 10× Genomics data.
CEL-seq Cell Expression by Linear amplification and sequencing.
ChromVAR Chromatin Variation Across Regions, a method to assess TF dynamics from scATAC-seq data.
Cicero An algorithm that identifies coaccessible pairs of DNA elements using scATAC-seq data and makes predictions about promoter–enhancer pairs.
cisTopic A probabilistic framework used to simultaneously discover coaccessible cis-regulatory elements and derive cell states using TOPIC modeling.
CiteFuse A streamlined package consisting tools for preprocessing analysis and web-based visualization of CITE-seq data.
CITE-seq Cellular Indexing of Transcriptomes and Epitopes by sequencing.
Clonealign A method that assigns gene expression states to cancer clones using single-cell data.
coupleNMF coupled Nonnegative Matrix Factorizations.
CRISP-seq A reverse genetics method that allows for analysis of thousands of CRISPR-mediated perturbations within a single experiment by combining pooled CRISPR screen to single-cell RNA sequencing. Similar techniques are Perturb-seq and CROP-seq.
CROP-seq CRISPR droplet-sequencing, a reverse genetics method that allows for analysis of thousands of CRISPR-mediated perturbations within a single experiment by combining pooled CRISPR screen to single-cell RNA sequencing. Similar techniques are Perturb-seq and CRISP-seq.
CSV A Comma-Separated Values file is a delimited text file format.
CUT&RUN Cleavage Under Targets and Release Using Nuclease.
CUT&Tag Cleavage Under Targets and Tagmentation.
CytoSeq Gene expression cytometry.
DBSCAN Density-Based Spatial Clustering of Applications with Noise.
DecontX A Bayesian method to estimate and remove contamination in individual cells.
Dip-C Method to obtain high-resolution contact maps in single diploid cells.
DoubletDecon R package that uses deconvolution to identify and remove doublets in scRNA-seq data.
DoubletFinder R package that can interface with Seurat and can predict and remove doublets in scRNA-seq data.
Drop-ChIP Chromatin-immunoprecipitation in droplets.
Drop-seq Method to profile mRNA transcripts from nanoliter-sized droplets of individual cells.
Dynamo A tool that predicts cell states over time periods (RNA velocity), and incorporates not only splicing information but also promoter state switching, translation and RNA/protein degradation by taking advantage of scRNA-seq and combined transcriptomics and proteomics.
ENCODE Encyclopedia of DNA Elements.
ENCODE blacklist regions A comprehensive set of regions in the human, mouse, worm, and fly genomes that have artifactual high signal in next-generation sequencing experiments.
Expedition A computational framework consisting of outrigger, anchor and bonvoyage, algorithms to detect alternative splicing, assign modalities and visualize results, respectively.
Feature matrix A matrix listing the number of UMIs associated with a feature (row) and a barcode (column) created for analysis of scRNA-seq or scATAC-seq data.
Fluidigm C1 Commercially available, automated single-cell isolation and preparation system for genomic analysis from Fluidigm.
FRiP FRaction of fragments in Peaks
fuzzy C-means A clustering algorithm similar to K-means clustering that computes centroids and clusters cells.
GBR Gradient Boosting Regression.
GEM Gel Beads in Emulsion.
GFA Group Factor Analysis.
Graph-based clustering An unsupervised classification algorithm that first identifies read similarities by performing pairwise comparisons and then constructs a graph in which the vertices correspond to sequence reads, the edges are overlapping reads and their similarity score is an edge weight.
Graphical LASSO Least Absolute Shrinkage and Selection Operator
GWAS Genome-wide Association Studies.
Harmony A unified framework for data integration, visualization, analysis, and interpretation of single-cell genomics data across discrete timepoints.
HCA Human Cell Atlas.
HDF5 Hierarchical Data Format version 5 is a file format that stores and organizes large data.
Hi-C all-to-all 3C method.
Hierarchical clustering Clustering based on similarities in objects.
HTML HyperText Markup Language is a standard markup language for documents designed to be displayed in a web browser.
ICA Independent Component Analysis.
iCELL8 A single-cell system from Takara Bio with open platforms that enable the processing of hundreds of single cells or nuclei.
IHEC International Human Epigenome Consortium.
inDrop indexing Droplets.
iNMF Nonnegative Matrix Factorization.
IVT In vitro transcription.
Kallisto/Bustools A pseudoalignment-based RNA-seq quantification program.
kBET k-nearest-neighbor Batch-Effect Test.
K-means clustering A clustering algorithm that computes the centroids and clusters cells.
Knee plot A standard single-cell RNA-seq ranked-ordered UMI plot that is used to determine a threshold for considering cells valid for analysis in an experiment.
LDA Latent Dirichlet Allocation, an example of TOPIC modeling.
LIGER Linked Inference of Genomic Experimental Relationships, a computational pipeline for integrating and analyzing multiple single-cell datasets.
Louvain algorithm An algorithm to extract communities from large networks utilizing the greedy optimization method.
LSI An indexing and retrieval method that uses SVD to identify patterns in the relationships between the terms and concepts.
MACS Model-based Analysis of ChIP-Seq, a popular peak calling algorithm.
MAGIC Markov Affinity-based Graph Imputation of Cells.
Mandalorion A tool to detect isoforms in scRNA-seq data.
MARS-seq MAssively parallel RNA Single-cell sequencing.
MATCHER Manifold Alignment to CHaracterize Experimental Relationships.
mcImpute Matrix Completion based Imputation for scRNA-seq data.
MDS MultiDimensional Scaling.
MEX Market EXchange (MEX) is a format that is used to represent the gene-barcode matrix output by Cell Ranger.
MIMOSCA Multiple Input Multiple Output Single Cell Analysis, a method to analyze single-cell Perturb-seq data.
MISC Missing Imputation for Single-Cell RNA-seq.
MISO Mixture-of-Isoforms, a model that estimates expression of alternatively spliced exons and isoforms.
MNN Mutual Nearest Neighbors.
modENCODE Model Organism Encyclopedia of DNA Elements.
Monocle An analysis toolkit for single-cell data that performs clustering, differential analysis and trajectory construction.
mtscATAC-seq Mitochondrial single-cell Assay for Transposase-Accessible Chromatin with sequencing.
MuSiC Multisubject single cell deconvolution method that utilizes scRNA-seq to estimate cell type proportions in bulk RNA-seq data.
netImpute Imputation of cell types from scRNA-seq data by integrating multiple types of biological networks.
Nucleosome banding pattern A specific DNA fragment size banding pattern produced by transposition/digestion on chromatin that corresponds to subnucleosome, mononucleosome, dinucleosome, and so on.
Nucleus hashing A sample multiplexing method (like cell hashing) with oligo tagged antibodies directed against nuclear pore complex proteins based on the concept of hash functions in computer science to index datasets with specific features.
Nystrom A sampling technique that generates the low rank embedding for large-scale dataset. It first builds a low-dimensional embedding with a subset of cells and then projects the rest of the data to the embedding structure.
Pacbio Pacific Biosciences, a biotechnology company that pioneered long-read sequencing.
PAGA PArtition-based Graph Abstraction.
PBAT Post-Bisulfite Adaptor Tagging.
PBMC Peripheral Blood Mononuclear Cells.
PCA Principal Component Analysis.
Perturb-seq A reverse genetics method that allows for analysis of thousands of CRISPR-mediated perturbations within a single experiment by combining pooled CRISPR screen to single-cell RNA sequencing. Similar techniques are CRISP-seq and CROP-seq.
Pseudobulk A sum of counts from single-cell data to imitate bulk data.
Pseudotime A measure of how far cells have progressed along a biological process.
REAP-seq RNA Expression And Protein sequencing.
RNA velocity A high-dimensional vector that predicts the future state of individual cells based on time-dependent relationship between the abundance of precursor and mature mRNA.
Salmon A pseudoalignment-based RNA-seq quantification program.
SC3 Single Cell Consensus Clustering.
Scanpy/EpiScanpy Single-Cell ANalysis in Python, a toolkit for end-to-end analysis of single-cell datasets.
Scasat Single Cell ATAC-Seq Analysis Tool, an end-to-end pipeline for analysis of scATAC-seq data.
scATAC-pro An end-to-end pipeline for analysis of scATAC-seq data.
scATAC-seq Single-cell ATAC-seq.
Scater A Bioconductor package for analyses of single-cell RNA-seq gene expression data, with a focus on quality control and visualization.
scBS-seq Single-cell bisulfite sequencing.
scCAT-seq Single-cell Chromatin Accessibility and Transcriptome sequencing.
scCOOL-seq Single-cell Chromatin Overall Omic-scale Landscape sequencing.
scGAIN scRNA-seq data imputation using Generative Adversarial Networks.
scHi-C Single-cell Hi-C.
scImpute A statistical method to impute the dropouts in scRNA-seq data.
scLVM/f-scLVM (factorial) single-cell latent variable model, a modeling framework for unraveling sources of heterogeneity and removing confounding factors for downstream analysis.
scM&T-seq Single-cell Methylome and Transcriptome sequencing.
scMethyl-HiC Single-cell Methyl-HiC.
scMT-seq Single-cell methylome and transcriptome sequencing.
scNOMeRe-seq Single-cell nucleosome, methylome, and transcriptome sequencing.
scNOMe-seq Single-cell Nucleosome Occupancy and Methylome sequencing.
scPipe A bioconductor package for single cell RNA-seq data analysis.
scRRBS-seq Single cell Reduced Representation Bisulfite Sequencing.
Scrublet Single-Cell Remover of DoUBLETs, a python code for identifying doublets in scRNA-seq data.
Scruff Single Cell RNA-Seq UMI Filtering Facilitator, a Bioconductor package that is used to preprocess scRNA-seq data.
scTrio-seq Single-cell genome, methylome, and transcriptome (trio) sequencing.
scVelo An improved estimate of RNA velocity that utilizes a likelihood-based dynamical model.
scWGBS Single-Cell Whole Genome Bisulfite Sequencing.
Seurat An end-to-end R package for analyses of scRNA-seq data.
Signac An extension of Seurat for analysis of single-cell chromatin data.
SingleCellExperiment object A lightweight Bioconductor container for storing and manipulating single-cell genomics data
SingleSplice An algorithm for detecting alternative splicing in a population of single cells.
Slingshot A method for inferring cell lineages and pseudotimes from scRNA-seq data.
SMART-seq Switching Mechanism At 5′ end of RNA Template sequencing.
SnapATAC An end-to-end pipeline for analysis of scATAC-seq data.
snm3C-seq Single-nucleus methyl Chromatin Conformation Capture sequencing.
snmCT-seq Single-nucleus methylome and transcriptome sequencing.
SNP Single Nucleotide Polymorphism.
SOLiD Small Oligonucleotide Ligation and Detection System.
Solo A neural network framework to classify doublets in single-cell data.
souporcell A toolkit for robust clustering, doublet detection and ambient RNA estimation in scRNA-seq data.
SoupX An R package for the estimation and removal of ambient mRNA contamination in droplet-based scRNA-seq data.
STAR Spliced Transcripts Alignment to a Reference, a popular alignment tool.
STREAM Single-cell Trajectories Reconstruction, Exploration And Mapping, a comprehensive single-cell trajectory analysis pipeline.
STRT-seq Single-cell Tagged Reverse Transcription sequencing.
SVD Singular Value Decomposition
Tagmentation Transposon-mediated cleaving and simultaneous tagging of DNA.
TF-IDF Term Frequency-Inverse Document Frequency
TOPIC modeling A type of statistical modeling for discovering the abstract “topics” that occur in a collection of documents.
t-SNE t-distributed Stochastic Neighbor Embedding.
TSO Template Switch Oligo.
TSS enrichment score Ratio of fragments centered at the TSS to fragments in TSS-flanking regions.
UMAP Uniform Manifold Approximation and Projection.
UMI Unique Molecular Identifier.
Variational Bayes A family of ensemble learning techniques used in machine learning.
Velocyto A tool to estimate RNA velocities in R or python.
VIPER Variability-Preserving ImPutation for Expression Recovery.

1.1. Single-Cell Transcriptomic Approaches

The single-cell revolution started with strategies to analyze the transcriptome in individual cells. The first reports of mRNA sequencing in single cells came from the profiling of individual mouse oocytes and blastomeres manually picked under a microscope. Individual cells were lysed, the RNA reverse-transcribed, the cDNA amplified by PCR, and sequenced on the SOLiD platform [11, 12].

Then came single-cell tagged reverse transcription sequencing (STRT-seq), which used an oligo-dT primer for first strand synthesis followed by incorporation of a cytosine string at the 3′ end of the first strand. A template switch oligo (TSO) that had complementary Gs allowed for barcoding of individual cells. The cDNA was then amplified using a single primer PCR, bound to beads, fragmented, A-tailed, and annealed to Illumina P2 adapter. The P1 adapter was introduced during library PCR step for short-read sequencing from the P1 side on an Illumina platform [13]. STRT-seq introduced one of the first multiplexing approaches for single-cell RNA-sequencing and was later adapted to the Fluidigm C1 microfluidics system for a more high-throughput version of the protocol (STRT-seq C1) [14]. Around the same time, unique molecular identifiers (UMIs) were introduced to serve as molecular barcodes during the reverse transcription step enabling absolute quantification of transcripts and distinguishing molecules from notorious PCR duplicates, a problemin low-input amplification protocols [15]. It also became routine to introduce exogenous spike-in controls to monitor technical reproducibility.

These early approaches, however, resulted in transcript read bias in the 3′ end [11, 12] due to their use of oligo-dT primers or 5′ end (STRT-seq) to incorporate barcodes for multiplexing. SMART-seq (Switching Mechanism At 5′ end of RNA Template sequencing) was introduced as a method for full length cDNA sequencing that allowed for improved read coverage across transcripts, significantly enhancing alternative transcript isoform discovery, dissection of allele-specific expression and identification of SNPs. SMART-seq used an oligo dT primer, a special TSO primer and a reverse transcriptase with terminal transferase activity. The first strand synthesis and template switch generated a full-length cDNA with an anchor for second strand synthesis which was followed by cDNA amplification. Libraries were generated by either ultrasonication followed by Illumina adapter ligation or direct tagmentation [16]. The SMART-seq kit was commercialized by Takara Clontech as the SMARTer Ultra Low RNA Kit for Illumina sequencing. SMART-seq2 was developed to further refine the reverse transcription, template switching, and preamplification steps and increase yield and length of cDNA libraries in a cost-effective way [17]. However, the SMART-seq/SMART-seq2 library preparation protocols lacked the ability to incorporate UMIs. Recently, SMART-seq3 was developed and can be used to construct UMI barcoded RNA libraries with the same 5′ end but different 3′ ends. The full-length transcript can be confidently reconstructed in silico as verified with PacBio long-read sequencing [18].

Right around the time the first SMART-seq protocol was developed, another method called CEL-seq (Cell Expression by Linear amplification and sequencing) was demonstrated to work on single cells. CEL-seq’s protocol made use of linear amplification of pooled barcoded RNA samples by in vitro transcription (IVT) made possible by the use of an oligo-dT primer that also had a T7 promoter sequence [19]. CEL-seq2 was modified to increase sensitivity and throughput by primer redesign, ligation-free library preparation and use of the Fluidigm platform [20]. CEL-seq2 but not CEL-seq also incorporated UMI molecular barcoding.

Massively parallel RNA single-cell sequencing (MARS-seq) is a method similar to CEL-seq2, and uses a three-tier barcoding system at the molecular (via UMIs), cellular and plate level prior to pooling and IVT amplification of RNA from single cells in a plate format [21]. MARS-seq2.0 builds upon MARS-seq improving throughput, reducing noise and contamination, increasing yield, and providing a user-friendly computational pipeline for data analysis [22].

CytoSeq was introduced as a highly scalable, automation-free method to capture any transcribed RNA from a single cell. Individual cells are deposited in picoliter wells, lysed, and hybridized to beads containing cellular barcodes and UMIs. The diversity of available cellular barcodes and UMIs ensured that each cell and each mRNA molecule was assigned a unique tag. The cells from each well are then pooled for reverse transcription, amplification, and sequencing [23].

Drop-seq [24], inDrop (indexing Droplets) [25] and 10× Genomics [26] are newer methods that use microfluidic partitioning to generate droplets that capture single cells. They also operate on similar principles where bead embedded primers containing an oligo-dT stretch, unique cell barcodes, UMIs, and PCR handles are cocaptured in the same droplet and hybridize to the mRNAs from a single cell. A few key differences include the composition of the bead itself and the location of the reverse transcription (RT) reaction. In Drop-seq, the bead is a functionalized microparticle and the primers are synthesized on the surface. The RT reaction for Drop-seq occurs after the droplets are broken by addition of perfluorooctanol. In inDrop, the bead is a hydrogel carrying photocleavable combinatorially barcoded primers. 10× Genomics uses a Gel Beads in EMulsion (GEM) technology. For both inDrop and 10× Genomics, the RT reaction occurs inside droplets and the products are released either by photocleavage (inDrop) or dissolution of beads (10× Genomics). The 10× Genomics method with ~65% capture efficiency far outruns the other two methods which are 5–7%. Thus, 10× Genomics has become the method of choice for most laboratories although the kits and instrument are costly.

1.2. Single-Cell Epigenomic Approaches

Sometimes, steady state transcript levels alone do not reveal the nuances of gene regulation within individual cells. Thus, recent advances in single-cell technology have evolved to analyze the epigenome in detail and include DNA accessibility, transcription factor (TF) binding, DNA and histone modifications, enhancer–promoter contact maps and 3D genome organization. DNA methylation analysis for example can provide insights into promoter regulation of gene expression being anticorrelated to transcription. Due to the incredible stability of DNA methyl marks and the ability to perform absolute quantification makes methylation analysis invaluable in aging studies. For example, methylation at some CpG sites serve as epigenetic clocks that predict biological age [27]. Aged somatic cells undergo global hypomethylation at repeat regions and enhancers, and focal hypermethylation at genes that direct age-related decline and disease [28]. Chromatin accessibility patterns also offer deep insights into biological function. By measuring cell-to-cell variability of TF binding under various perturbing conditions, it is possible to identify the key TFs responsive to those perturbations [29]. Furthermore, coaccessibility patterns at enhancer promoter pairs in conjunction with single-cell chromosome conformation studies can reveal 3D genome organization in unprecedented detail.

The first single-cell epigenome technology called scRRBS-seq was a method detecting DNA methylation in single cells using a reduced representation bisulfite sequencing (RRBS) which digests DNA with restriction enzymes to enrich for CpGs. In this method, cells are lysed, the genomic DNA digested, and ligated to adapters before bisulfite conversion, all in a single tube. This is followed by DNA purification and PCR amplification for sequence-ready libraries [30]. While RRBS significantly reduces sequencing cost, it suffers from poor coverage and exclusion of many regulatory regions. Furthermore, because bisulfite conversion which fragments DNA is performed after adapter ligation, this leads to further reduction in representation.

To counter this deficiency, a method called Post-Bisulfite Adaptor Tagging (PBAT) was developed where bisulfite treatment precedes adaptor tagging [31]. Single-cell bisulfite sequencing (scBSseq) adopted the PBAT strategy with additional PCR amplification for the absolute quantification of 5-methyl cytosine (5mC) in single cells [32]. In a related method called scWGBS (single-cell Whole Genome Bisulfite Sequencing), PBAT was used with all steps from lysis to library cleanup being performed in a single tube [33].

The first single-cell chromatin immunoprecipitation sequencing (ChIP-seq) method called Drop-ChIP was published by the Bernstein lab [34] and used an in-house microfluidic device which formed coflowing drops of cells, micrococcal nuclease (MNase) and detergent on one side, and barcoded adapters on the other side. Cell lysis, chromatin fragmentation and unique adapter ligation were done inside the drops. The drops were then broken, indexed chromatin fragments pooled, and ChIP performed with carrier chromatin before restriction digestion to remove excess adapter followed by PCR amplification. A second restriction digest created overhangs to ligate yet another sample barcode allowing for further multiplexing and Illumina sequencing-ready libraries. Given that labs had to build their own microfluidic device, this method did not gain as much popularity. Single-cell ChIP-seq remained a challenge due to the small amounts of immunoprecipitated material and low signal to noise ratio. Recently, major advancements in ChIP technology including CUT&RUN (Cleavage Under Targets and Release Using Nuclease) and CUT&Tag (Cleavage Under Targets and Tagmentation) have improved the signal-to-noise ratio by using tethered enzymes that specifically release antibody bound chromatin fragments. Indeed, CUT&Tag was used to profile H3K27me3 and H3K4me2 in single K562 and H1 cells using Takara’s iCELL8 nanodispenser system and the SMARTer iCELL8 350v chip [35]. Subsequently, CUT&Tag was adapted to the 10× Genomics platform to profile histone modifications in multiple complex tissues [36, 37].

The most widely used epigenomic investigation in single cells is a technique called ATAC-seq (Assay for Transposase Accessible Chromatin sequencing). ATAC-seq works on the principle that open regions of chromatin (generally promoters, enhancers, DNA flanking TF bound sites, etc.) are more accessible for Tn5 transposome-mediated tagmentation. ATAC-seq at single cell resolution (scATAC-seq) was first performed using a highly scalable combinatorial indexing method where nuclei are first distributed in wells with a unique barcode, pooled, diluted and then distributed to another set of wells with a different set of barcodes [38]. The sequential and randomized passage of cells through two different 96-well plates at low concentration minimizes the probability of two cells receiving the same set of barcodes. The first barcode is introduced during tagmentation using uniquely barcoded transposome complexes. In the second well, the cells are lysed and the released tagmented DNA is PCR-amplified by priming off the adapters thereby introducing the second barcode. Although this was a highly scalable method, the median number of reads per cell was ~2500. In an advancement of the scATAC-seq technology, the Fluidigm platform was adopted with an automated workflow to tagment individual cells, PCR-amplify with dual index primers and pool for sequencing [29]. This method had the same throughput but increased the median number of fragments per cell to ~73,000. Currently, scATAC-seq has been adopted by the commercial 10× Genomics platform which performs the tagmentation in bulk followed by droplet capture of single cells [39].

A relatively new frontier of epigenomics is the study of the spatial organization of chromatin in 3D. Chromosome Conformation Capture (3C) techniques assess millions of contacts in the genome and together with statistical modeling can give rise to high-resolution structural views of chromosomes [40]. scHi-C (single-cell Hi-C) [41] is an “all-to-all” 3C method that performs chromatin cross-linking, restriction enzyme digestion, biotin fill-in, and proximity ligation in nuclei (also known as in situ Hi-C [42]). Individual nuclei are then placed in wells, cross-links reversed, and ligation products purified on streptavidin beads. The products are then further digested with another restriction enzyme, ligated to adapters, PCR-amplified, and sequenced. A variant method called Dip-C [43], obtained high-resolution contact maps in single diploid cells (which have many more contacts) by omitting the biotin pulldown and coupling to a whole genome amplification and sequencing.

1.3. Single-Cell Multiomics Approaches

Interrogation of single modalities such as transcriptome or accessibility or DNA/histone modifications provides only a partial picture of biological events in single cells. With technological advances, several clever modifications have enabled the simultaneous profiling of multiple modalities from a single cell.

scM&T-seq (single-cell methylome and transcriptome sequencing) is made possible by the physical separation of poly-A RNA by magnetic beads with oligo-dT and DNA from the same cell. This is followed by parallel processing of RNA using the SMART-seq2 strategy and DNA using the scBS-seq strategy [44]. Related methods include scMT-seq (methylome and transcriptome) [45], snmCT-seq (single-nucleus methylome and transcriptome) [46], and scTrio-seq (genome, methylome, and transcriptome) [47].

In addition to joint profiling of genome, methylome, and transcriptome, nucleosome occupancy mapping was adopted in single cells. scNOMe-seq (single-cell nucleosome occupancy and methylome sequencing) is performed by incubating cells with a GpC methyltransferase that methylates accessible regions while being unable to methylate nucleosome protected regions and then FACS sorting cells into individual wells. The DNA is then subjected to bisulfite conversion and sequencing enabling the measurement of accessibility, TF binding, nucleosome occupancy, and phasing [48]. This nucleosome mapping was combined with transcriptome profiling in scNMT-seq (nucleosome, methylome, and transcriptome) [49] and scNOMeRe-seq (nucleosome, methylome, and transcriptome) [50]. scCOOL-seq (single-cell Chromatin Overall Omic-scale Landscape sequencing) is a method that simultaneously measures nucleosome positioning, DNA methylation, copy number variation and ploidy from single cells by combining scNOMe-seq, PBAT-seq, and spike-ins [51].

Perhaps the most popular and commercialized multiomics approach is the simultaneous assessment of chromatin accessibility and transcriptome. Two methods that accomplish this are scCAT-seq (single-cell chromatin accessibility and transcriptome sequencing) [52] and ATAC-RNA-seq [53]. In scCAT-seq, after single cell isolation, a mild lysis method separates the cytoplasmic components (mRNA) from the nuclei (DNA). The SMART-seq2 strategy is used to prepare gene expression libraries while the nucleus is subjected to Tn5 transposition and ATAC-seq library preparation. In ATAC-RNA seq, cells are first fixed with dithiobis(succinimidyl propionate), permeabilized and transposed in bulk. The cells are then stained with DAPI and individual cells sorted into wells with lysis buffer. The mRNA component undergoes SMART-seq2 library preparation in a separate plate while the genomic DNA undergoes indexed PCR to amplify the transposed fragments. Paired scRNA-seq and scATAC-seq from the same cell has recently been adapted to the 10× Genomics platform (Chromium single-cell multiome ATAC and gene expression) enabling unprecedented scalability in this kind of multi-omics approach. In this method, nuclei are extracted, permeabilized, transposed in bulk, and captured in droplets. The nuclear mRNA and DNA components are captured by the gel bead oligos and get tagged with the same barcode identifying the cell of origin. The gene expression and ATAC components then get separated and undergo their separate library preparations based on 10× Genomics protocols.

Finally, scMethyl-HiC (single-cell Methyl-HiC) [54] and snm3C-seq (single-nucleus methyl Chromatin Conformation Capture sequencing) [55] are two methods that combine methylomics and 3D genomics in single cells. Both protocols combine in situ Hi-C followed by bisulfite conversion, library construction, and paired-end sequencing.

1.4. Single-Cell Multiplexing Approaches

RNA Expression And Protein sequencing (REAP-seq) and Cellular Indexing of Transcriptomes and Epitopes by sequencing (CITE-seq) were two methods developed independently for multimodal single cell immunophenotyping and transcriptome analysis [56, 57]. These methods used oligo-tagged antibodies against cell surface markers that were captured in most oligo-dT-based scRNA-seq protocols thus receiving the same cell-barcode and helping to identify the cell type. REAP-seq and CITE-seq differed in how the oligo was conjugated to the antibody. In REAP-seq the aminated oligo is covalently bound to the antibody while in CITE-seq, the biotinylated oligo is bound to the streptavidin tagged antibody. Following cDNA synthesis, a size-selection separates the cDNA from cell transcripts and antibody oligos for parallel library preparations. A variant of CITE-seq applied to scATAC-seq data is ATAC with Select Antigen Profiling by sequencing (ASAP-seq) [58]. The introduction of a single bridge oligo makes ASAP-seq compatible with simultaneous phenotyping and chromatin accessibility profiling. Importantly, ASAP-seq (unlike CITE-seq) can be used to detect intracellular epitopes. On the sample preparation side, ASAP-seq utilizes the newly developed mitochondrial single-cell assay for transposase-accessible chromatin with sequencing (mtscATAC-seq) in which Tn5 is added to fixed, permeabilized whole cells rather than nuclei and both nuclear DNA and mitochondrial DNA are transposed [59].

A clever modification of the CITE-seq protocol allows for multiplexing and superloading of scRNA-seq samples in a method called Cell Hashing [60]. This multiplexing strategy dramatically reduces experimental costs. Cell Hashing uses barcoded oligo-tagged antibodies (hashtags) directed against ubiquitously expressed proteins. Every sample is uniquely tagged and then pooled before droplet capture and scRNA-seq library preparation. Downstream deconvolution of the oligo sequences can identify individual samples. In addition to multiplexing samples, Cell Hashing can be used to overload microfluidic channels and then remove intersample cell multiplets computationally. Similar multiplexing strategies have also been applied to scATAC-seq preparations. The first, called Nucleus Hashing, uses hashtags directed to the nuclear pore complex [61]. The second, is an expanded utility of ASAP-seq with the bridge oligo capturing hashtags against nuclear epitopes [58].

1.5. Single-Cell Functional Genomics Approaches

An evolving new area of single-cell technology is the development of CRISPR-based functional genomic screens combined with single-cell transcriptomics readout. Perturb-seq [62] is one such method where pools of lentiviruses containing designer barcoded guide RNA (gRNA) constructs are transduced into cells. While the gRNA component knocks out target genes, the associated barcode is captured with the rest of the cellular mRNA during droplet generation thus acting as a molecular identifier of the perturbation. The readout is simply the gene expression profile of the perturbed cell. By fine-tuning the multiplicity of infection, it is possible to study either single gene deletions or epistatic relationships between multiple genes. CRISP-seq [63] and CROP-seq (CRISPR dropletsequencing [64]) are two alternative versions of CRISPR-screening in single cells with different vector designs. Furthermore, it was possible to use Perturb-seq and related methods with other CRISPR manipulations such as CRISPR inhibition rather than gene knockout [65, 66] as well as multiplex several gRNAs in the same vector [65]. Perturb-seq was further improved by developing a method to directly sequence the gRNA eliminating the need for barcodes [67].

A timeline of the key method developments in single-cell genomics is depicted in Fig. 1. Also, see Notes 1 and 2 for some key considerations when performing the wet-lab procedures to generate single-cell genomics data.

Fig.1.

Fig.1

A timeline of the evolution of single-cell genomics technologies. The transformative technological evolution of single-cell methods to probe the transcriptome and epigenome are outlined. The methods are colored based on the precise ”omics“ used

2. Methods

2.1. Computational Methods to Analyze Single-Cell Transcriptomics Data

Rapid progress in single-cell experimental protocols (described in Subheading 1), and the great complexity of the obtained results, prompted development of new analytical tools and resources for accurate interpretation of scRNA-seq data. Although, there is no standardized computational pipeline for scRNA-seq data analysis, multiple bioinformatics tools are continuously evolving [68]. The overall workflow entails preprocessing and alignment of raw sequencing data to create count matrices of mRNA molecules. The count matrices are further evaluated for quality, normalized, and corrected to remove any technical variation and elucidated for biological function. Finally, the top highly variable genes are applied for dimensionality reduction and clustering to visualize cell heterogeneity and run desired downstream analysis. See Note 3 for software version control when reporting data from single-cell analysis.

A practical example of scRNA-seq analysis covering the fundamental principles with 1K peripheral blood mononuclear cells (PBMCs) from a healthy human donor (dataset available for download from the 10× Genomics website) is outlined in Fig. 2 and discussed in detail below. Figure 2a illustrates the basic workflow.

Fig.2.

Fig.2

Preliminary analysis of 10× Genomics scRNA-seq data on 1000 human PBMCs. (A) Typical workflow and tools used for analysis of scRNA-seq data. (B) Knee plot showing the distinction between barcodes identified as cells or noncells in Cell Ranger. (C) nFeature (number of genes with at least one UMI), nCount (number of UMI counts taken across all cells) and percent.mt (percent mitochondrial genes) before (top) and after (bottom) filtering in Seurat. (D) Top 2000 variable genes identified by mean-variance modeling in Seurat is shown in red. (E) PCA (left) and UMAP (right) representation of clusters. (F) Visualization of marker gene expression on a UMAP

2.1.1. Data Preprocessing

Raw scRNA-seq data generated by a sequencer is first preprocessed for cell barcode identification and deconvolution of UMIs to demultiplex cell-specific reads. Subsequently, reads from each cell are aligned to a reference genome with implementation of bioinformatics tools developed for bulk RNA-seq. The most commonly used software for initial scRNA-seq data alignment is STAR [69] that exploits feature counts, or pseudoalignment-based programs such as Kallisto/Bustools [70], Salmon [71], or Alevin [72]. 10× Genomics’s Cell Ranger pipeline automates many of these preprocessing steps including STAR alignment and generating output files in BAM (Binary Alignment Map), MEX (Market Exchange), CSV (Comma Separated Values), HDF5 (Hierarchical Data Format version 5), and HTML (HyperText Markup Language) formats. A typical plot generated by Cell Ranger is an UMI vs barcode “knee plot” (Fig. 2b) that provides visual cues of a good separation between cell-containing or empty droplets. A sharp drop in UMI (knee) signals a good separation. In addition, Bioconductor offers a broad range of freely available R-packages that are continuously being enriched by the scientific community [73]. Among them, scPipe [74] and scruff [75] can be used for scRNA-seq preprocessing.

2.1.2. Quality Control

A fundamental component of scRNA-seq data analysis is a cell quality control (QC) procedure that aims to accurately identify and remove outlier barcodes corresponding to either low-quality cells or doublets, thus ensuring investigation of real biological differences in cells [76, 77]. Multiple computational tools for scRNA-seq QC use counts and gene expression features for cell quality analysis (such as Scater [78] and Seurat [79]), while other methods apply machine learning techniques [80] or use gene coverage distribution [77].

Metrics that are commonly used for assessment of cell quality involve number of counts per cell, number of genes per cell and the percent of reads in cells that comes from mitochondrial genome. Barcodes that are characterized by a low number of counts, few detected genes, and a high fraction of mitochondrial counts, are qualified as low-quality cells. These can include damaged cells with leakage of cytoplasmic RNA or cells undergoing apoptosis. On the other hand, when two or more cells (doublets or multiplets) are captured and labeled in the same droplet, the resulting barcode will contain extremely large number of detected genes and very high counts. A typical example of how a scRNA-seq dataset might look before and after filtering for empty droplets, doublets or high mitochondrial reads is illustrated in Fig. 2c. Setting a universal QC threshold applicable for various types of scRNA-seq data is challenging since thresholds can depend on cell-type, cell-state and other variables. Additionally, each of the metrics used for cell filtering can illustrate other cellular characteristics, such as, lower counts can indicate smaller or quiescent cells, and a high proportion of mitochondrial genes can represent cells with high energy needs such as cardiomyocytes. Therefore, assessment of cell quality should consider all QC metrics in parallel, as using these parameters alone may be misinformative for downstream data analysis and results. Computational packages such as Seurat or Scater [78] apply a manual strategy that allows for visual evaluation of the sample before and after filtering. This provides a quick and effective way for setting dataset-specific thresholds, particularly if the results obtained from an initial setting are not clear. In addition, for more rigid cell filtering, many other computational tools are available, such as SoupX [81], souporcell [82], DecontX [83] for removal of ambient RNA contamination; and DoubletFinder [84], DoubletDecon [85], Scrublet [86], and Solo [87] for detection of doublet barcodes. See Note 4 highlighting the importance of QC steps.

2.1.3. Batch Correction

Batch effect is a major driver of data variation that results from differences in experimental processing of sample sets (see Note 5). Consequently, by interfering with true biological variation, it can confound interpretation of results. Methods such as Seurat Canonical Correlation Analysis (CCA) [88] and mutual nearest neighbors (MNN) [89] are considered the most successful algorithms for removing batch effect from scRNA-seq datasets [90]. MNN algorithm has been incorporated into the Cell Ranger pipeline to enable correction of different 10× Genomics chemistries. Other widely accepted algorithms for quantification of batch effects in scRNAseq data are Harmony [91], LIGER [92], and the recently developed k-nearest-neighbor Batch-Effect Test (kBET) [93].

2.1.4. Data Normalization

The gene counts of individual cells in a scRNA-seq dataset may show differences that are associated with technical variation created by sample processing and sequencing depth. This technical noise can confound true sample heterogeneity and must be corrected before downstream analysis. Thus, the count data are normalized to make the counts comparable across cells, and to ensure elimination of cell-specific biases. Over the years, several methods of normalization have been adapted from bulk RNA-seq analysis pipelines or developed specifically for single cells [94]. The most popular Seurat [95] and Scanpy [96] platforms employ global scaling normalization method, in which each gene expression in a cell is divided by the total gene expression in that cell, multiplied by scale factor 10,000, and finally natural log transformed. This ensures that highly expressed genes do not dominate the genes with low expression, thereby masking real biological differences. Additional data correction can be performed for regressing out genes related to unwanted biological processes such as cell cycle phase, mitochondrial, or ribosomal genes. These corrections are available in some scRNA-seq analysis platforms, such as Scanpy or Seurat but also can be performed separately with more specified tools such as scLVM [97] or f-scLVM [98]. However, given that cell processes are tightly connected, removing genes associated to specific biological functions can bias data interpretation and therefore such corrections need to be monitored with caution.

2.1.5. Dimensionality Reduction

The main goal of scRNA-seq is to uncover cell-to-cell variability within a sample, thus focusing on genes that drive heterogeneity inside the cell population. For this reason, only a small subset of genes, with the highest variation in expression across the dataset is informative and must be selected and used for downstream analysis. Most software packages implement algorithms for selection of highly variable genes [99] (usually top 500–2000 genes), that depend on dataset complexity [100], and ultimately generate a high-dimensional expression profile for each cell (Fig. 2d). Therefore, this high-dimensional data complexity requires further reduction to allow data visualization in a 2D plot and clustering of cells for biological interpretation. Several dimensional reduction algorithms have been developed [101, 102] that, in combination with different clustering models, can effectively extract biological insights of cells. This includes investigation of (1) clusters of cells with similar gene expression profile to uncover cell states or types, (2) groups of cells with coexpressing genes to study coregulatory relationships between genes, or (3) cells demonstrating continuous gene expression patterns for tracing cellular dynamic processes.

The most common, unsupervised linear dimensionality reduction technique, principal component analysis (PCA), is widely applied in many scRNA-seq software, including Seurat and Scanpy, as an internal algorithm that contributes for multiple analyses with different informational use, such as assessment of data QC and detection of batch effects (Fig. 2e, left). In addition, scores of principal components (PCs) can be exploited for other nonlinear dimensionality reduction methods, cell clustering or trajectory inferences [101, 103]. Besides PCA, other dimensionality reduction methods are often used for targeting different biological approaches in scRNA-seq analysis in combination with multiple clustering and projection models [101]. Among them, the most popular, are nonlinear dimensionality reduction methods such as multidimensional scaling (MDS) that is implemented in Scater, t-distributed stochastic neighbor embedding (t-SNE) [104], and uniform manifold approximation and projection (UMAP; Fig. 2e, right) [105]. They are used to generate 2D or 3D plots for dimension-reducing visualization of relationships between cells. For exploratory data visualization in Cell Ranger, both t-SNE and UMAP are implemented.

2.1.6. Cell Clustering, Find Marker Genes

Clustering of cells with similar gene expression into separate groups is the main strategy that elucidates different cell states or types. Clustering algorithms use intercell distance matrices to assemble subgroups of cells with similar gene expression patterns. Among the variety of clustering methods, hierarchical clustering, K-means, fuzzy C-means, DBSCAN, and Louvain algorithm are the models often used for single cell clusters establishment [106]. The most common K-means method, that selects k-clusters and assigns cells to closest cluster, is used by the SC3 package [107] and Monocle 2 [108]. Alternatively, community detection-based Louvain algorithm was shown to outperform other clustering algorithms [106, 109, 110]. Among the commonly implemented platforms, Scanpy and Seurat perform clustering by Louvain community on a single-cell K-Nearest Neighbor (KNN) graph, while Monocle 3 [111] uses Louvain/Leiden algorithm [112]. Once the cell clusters are established, they are further characterized by identification of gene markers that are uniquely expressed in each cluster, therefore defining cell state or function. This can be very challenging for single-cell data due to the low counts, high dropout rate and amplification biases in mRNA. However, several programs address these challenges and report marker genes. For example, Seurat identifies marker genes based on differential expression in a cluster relative to average expression in all other clusters (see Note 6 for additional considerations). In Fig. 2f, marker genes identified were granulocytes/monocytes (cluster 0), T cells (clusters 1, 2, 5, 7 and 8), B cells and Dendritic cells (cluster 3), NK cells (cluster 4) and Plasma cells (cluster 6). Since this was a dataset of only 1000 cells, rarer populations such as hematopoietic stem cells were not detectable.

2.1.7. Trajectory Analysis

Another type of downstream analysis, that relies on results from dimensional reduction methods, is a machine-learning approach called cell trajectory inference. This analysis aims to order cells along a trajectory as a function of pseudotime (an abstract unit of progress), to trace divergence of cell lineages that can indicate multiple dynamic cellular processes such as development, differentiation, activation, or tumor progression. Among the multiple tools developed for trajectory extrapolation [113], Slingshot [114], Monocle 3 [111], and PAGA [115] are popular methods of choice. Trajectory analysis can be unsupervised, that is, it uses no prior information on genes that are important for the dynamic biological process or semisupervised, where some important marker genes are part of the input. A significant challenge of trajectory extrapolation is a necessity for supplementary validation with supportive evidence, to ensure that created trajectories truly reflect biological processes. An example of a trajectory confirmation method is RNA velocity [116] in which single-cell RNA transcription and splicing information is integrated into a trajectory, to verify directions of the pseudotime lineages. Recently developed, Velocyto [116], scVelo [117], and Dynamo [118] algorithms can be applied to recapitulate transient cell states and the dynamics of different biological processes.

2.1.8. Splice-Variant Analysis Using SMART-Seq

Given that alternative splicing is a key driver of protein diversity, scRNA-seq is a target technology for uncovering splice variant expression and transcriptomic variation that characterize individual cell functionality in physiology and disease. However, most microfluidic or droplet-based scRNA-seq platforms (such as 10× Genomics, CEL-seq, and STRT-seq) generate short UMI tagged transcripts. Since only a fragment of the transcript is sequenced, the sensitivity is too low for investigation of transcript isoforms [119]. Instead, plate-based, low-throughput methods such as SMART-seq and SMART-seq2, allow full length transcriptome sequencing that is sufficiently sensitive for gene, isoform and allele expression detection and can be applicable for single-cell alternative splicing analysis. Moreover, recently introduced SMART-seq3 method, with improved sensitivity and UMI incorporation promises to accelerate discovery of transcript splice isoforms in a high-throughput manner.

Currently, several computational strategies have been developed to perform splice variant detection and quantitation, using data obtained from existing short-read scRNA-seq technologies [120, 121]. MISO [122], BRIE [123], and Expedition [124] packages can be used to assess transcript splice variant expression at the exon level, using reads aligned to splice junctions. Single-Splice [125] applies statistical model to detect genes with different isoform usage across a set of single cells. In addition, the integrated informatics pipeline Mandalorion [126], has been introduced for single-cell isoform detection from long-read Nanopore sequencing data. However, this technology awaits further improvement for wide application to single-cell data.

2.1.9. CITE-seq and Cell-Hashing

As mentioned in Subheading 1.4, CITE-seq enables parallel quantification of RNA profiles and surface protein expression at the single cell level [56] while Cell Hashing is an application of CITE-seq that enables sample indexing for multiplexing [60]. CITE-seq analysis can be done using Seurat and Scanpy pipelines. The recently developed CiteFuse [127] package offers a more comprehensive set of tools that include doublet identification, cell hashing information, and transcriptome analysis.

2.2. Computational Methods to Analyze Single-Cell ATAC-seq Data

Compared to scRNA-seq data analysis, tools for scATAC-seq analysis are relatively limited but rapidly evolving. The workflow involves similar preprocessing and QC steps as scRNA-seq data with the parsing of cellular barcodes, generation of fragment files and a feature matrix (usually peaks). An important difference between scRNA-seq and scATAC-seq data is the inherent sparsity of the latter. The total number of accessible sites in the genome far exceed the total number of reads obtained per cell. Additionally, scATAC-seq data is binary, with each locus having mostly 0 or 1 reads and maximally two reads [29]. Thus, standard computational tools cannot accurately reconstruct cellular activities. As a remedy, most scATAC-seq-based workflows include some statistical framework that measures activities in ensemble either from bulk or single-cell (pseudobulk) data. Downstream analysis then focuses on cell-type specific peaks, gene activity deduction, underlying TF motif identification, estimation of TF dynamics, enhancer function, and so on. See Note 3 for software version control when reporting data from single-cell analysis.

A practical example of scATAC-seq analysis covering the fundamental principles with 1K PBMCs from a healthy human donor (dataset available for download from the 10× Genomics website) is outlined in Fig. 3 and discussed in detail below. Figure 3a illustrates the basic workflow.

Fig.3.

Fig.3

Preliminary analysis of 10× Genomics scATAC-seq data on 1000 human PBMCs. (A) Typical workflow and tools used for analysis of scATAC-seq data. (B) Knee plot showing the distinction between barcodes identified as cells or noncells in Cell Ranger. (C) QC metrics such as percent reads in peaks (fraction of all fragments that fall within ATAC-seq peaks), peak region fragments (measure of cellular sequencing depth/complexity), TSS enrichment (ratio of fragments centered at the TSS to fragments in TSS-flanking regions as defined by ENCODE), blacklist ratio (proportion of reads mapping to ENCODE blacklist regions), and nucleosome signal (ratio of mononucleosomal to nucleosome-free fragments) before and after filtering in Signac. (D) Histogram of DNA fragment sizes from the paired-end sequencing reads showing strong nucleosome banding pattern determined in ArchR. (E) ArchR derived TSS enrichment scores which is the average accessibility in a 50bp region centered at the TSS divided by the average accessibility of the TSS flanking positions 2 kb. (F) Doublet inference with ArchR projected to a UMAP plot. Cells with high doublet enrichment (purple) are computationally removed for downstream analysis. (G) UMAP representation of clusters. (H) Compartment analysis of high-quality peaks in each cluster. (I) Heatmap representing marker genes identified in each cluster. (J) ChromVAR analysis showing the top TFs with high variability

2.2.1. Data Preprocessing

The upstream data preprocessing steps such as alignment to reference genome and peak-by-cell count matrix generation from BAM files are similar to scRNA-seq data analysis and utilizes the same tools packaged into specialized scATAC-seq software. One additional step in scATAC analysis is peak calling and most software use MACS2 [128]. 10× Genomics’s Cell Ranger ATAC pipeline automates these preprocessing steps and generates fragment files and peak matrices. It also outputs basic plots such as the knee plot to separate cells and empty droplets (Fig. 3b). Alternatively, Bioconductor packages such as scPipe-ATAC can preprocess data and generate SingleCellExperiment objects that can then be used as input to the many other R packages in Bioconductor. Other handy end-to-end scATAC analysis software include Scasat [129], scATAC-pro [130], EpiScanpy [131], SnapATAC [132], and ArchR [133], and can be used directly. Among these, SnapATAC and ArchR are two methods that are fast and use less memory due to optimization and parallelization of methods. Instead of using a predetermined peak set from bulk/pseudobulk data, SnapATAC and ArchR tile the genome into 5kb (SnapATAC) or 500bp (ArchR) bins. This allows for a more unbiased element representation that is otherwise skewed toward most abundant cell types. Furthermore, a higher resolution 500bp binning in ArchR ensures proper capture of singular regulatory elements that tend to be 300–500bp in size. ArchR attains an additional dimension of computational efficiency by parsing fragment files into small chunks per chromosome and then storing them in compressed randomaccess HDF5 files. These HDF5 files form the constituent pieces of “Arrow” files which are then grouped into “ArchR Projects” for specific downstream analysis.

2.2.2. Quality Control

The QC for scATAC data determines several quantitative parameters to determine success of transposition. Qualitative plots generated from Cell Ranger ATAC pipeline include nucleosome banding pattern, TSS enrichment score, FRiP (FRaction of fragments in Peaks), and ratio of reads in ENCODE blacklist regions, and can be used to ascertain whether majority of reads originate from expected accessible regions in the genome. Alternatively, Signac, an extension of the Seurat R toolkit, takes fragment files as input and also offers well organized QC metrics to assess quality [134]. Figure 3c illustrates how the 1K PBMC scATAC-seq data might look like before and after filtering out poor quality nuclei by Signac. Figure 3d and e shows fragment distribution and TSS enrichment of reads as determined by ArchR. ArchR also uses an innovative doublet removal technique that synthesizes synthetic doublets by mixing cells in silico in thousands of combinations and then embedding in an UMAP. A nearest neighbor analysis is then used to identify nuclei that behave like doublets and are removed (Fig. 3f, purple dots). See Note 4 highlighting the importance of QC steps.

2.2.3. Batch Correction

Batch correction of scATAC-seq data is often carried out as part of preprocessing steps or dimensionality reduction, sometimes using tools already used for bulk or scRNA-seq data (see Subheading 2.1.3). These tools (such as Harmony used in SnapATAC and ArchR) scan for variable peaks (or noise) in the data and effectively eliminate them. See Note 5 highlighting the importance of batch correction steps.

2.2.4. Data Normalization

Cell Ranger, Signac and ArchR perform Latent Semantic Indexing (LSI), a method originally used for language processing and first applied to scATAC-seq data by Cusanovich et al. [38]. LSI performs term frequency-inverse document frequency (TF-IDF) transformation where a “term” is a peak and the “document” is a sample. TF-IDF normalizes across cells to correct for sequencing depth, and then across peaks to give more weight to rare peaks. Finally, a singular value decomposition (SVD) is applied so that the most valuable information across samples is identified and represented in a lower dimensional space. Another algorithm derived from text mining is Latent Dirichlet Allocation (LDA) which is used by cisTopic to group cells by similarities in accessibility profiles [135]. SnapATAC uses a slightly different approach: chromatin accessibility profiles of every cell are represented as binary vectors, the lengths of which correspond to the number of 5 kb bins used to segment the genome. Normalized Jaccard Indices are then calculated where the value of each element corresponds to the fraction of overlapping bins between every pair of cells. An eigenvector decomposition is then performed on this normalized similarity matrix for dimensionality reduction.

2.2.5. Dimensionality Reduction Visuals

Linear dimensionality reduction methods like PCA are generally less preferred to visualize scATAC-seq data because the sparsity results in high cell-to-cell similarity. Instead, after LSI/LDA normalization, UMAP (Fig. 3g) or t-SNE embeddings are used.

2.2.6. Cell Clustering, Find Marker Genes

Clustering of scATAC-seq data is performed using the same tools as scRNA-seq. For example, Cell Ranger uses either a k-means or a graph-based clustering method, the latter also used by Seurat and ArchR and utilizing Louvain community detection algorithms. SnapATAC utilizes a sampling method called Nystrom which first performs a low dimension embedding from a subsample of cells and then projects the remainder cells to the embedding structure. However, unlike Seurat/ArchR’s deterministic clustering algorithm, Nystrom based clustering can be stochastic. Thus, to improve the robustness of the clustering method, SnapATAC uses an ensemble Nystrom, which generates a mixture of Nystrom approximations which tend to provide better clustering stability. Each cluster is then annotated based on marker gene identification techniques founded on gene score models, a variety of which are used. For example, in ArchR, a gene score model based on accessibility within gene bodies, activity of nearby regulatory elements and strict gene boundaries is used to identify marker genes for each cluster (see Note 6 for additional considerations). In Fig. 3i, marker genes identified were granulocytes (cluster 1), B cells (cluster 2), T/NK/Dendritic cells (cluster 3) while cluster 4 did not show any prominent markers. Compared to scRNA-seq, marker gene identification was less informative for scATAC-seq data which can be due to the inherent challenges of clustering from sparse data, or the fact that chromatin accessibility does not fully correlate to gene expression. Finally, peak distribution may be assessed in each cluster to detect where most of the biological changes are occurring. In Fig. 3h, the PBMC dataset shows that majority of high-quality peaks are located in introns or intergenic (distal) regions.

2.2.7. Trajectory Analysis

As with scRNA-seq analysis, cellular trajectories can also be constructed with scATAC-seq data to identify nondiscrete progressive changes in chromatin accessibility along developmental or differentiation pseudotimes. As before, if there are multiple outcomes, the trajectory will show a branched structure, indicative of critical cellular decision-making points. Once the cells have been ordered, differential peak analysis can be used to identify the exact genomic regions that drive the progressive changes. Commonly used tools for trajectory generation with scATAC-seq data include Cicero (which uses Monocle 3), SnapATAC (which uses Slingshot), ArchR, and STREAM [136].

2.2.8. Chromatin Variation Across Regions

Chromatin accessibility data is highly variable across individual cells due to heterogenous cell distribution in a sample, and cell type-, time-, and sex-specific TF activities within samples. ChromVAR (Chromatin Variation Across Regions) is an R package that assesses the functionality of trans-acting factors and cis-regulatory elements from this highly variable scATAC-seq data by measuring gain and loss of accessibility within peaks [29, 137]. ChromVAR takes three input files: aligned sequencing reads, pseudobulk/bulk peak information, and a set of chromatin features representing either position weight matrices (PWMs) of TF motifs or other user-determined genomic features such as enhancer locations, ChIP-seq peaks, and GWAS annotations. ChromVAR outputs include bias-corrected “deviation” values which reflect the difference between the observed number of fragments that map to peaks containing a particular motif and the expected number of mapped fragments based on bulk/pseudobulk data and a “z-score.” Variability, which is the standard deviation of the z-scores is then calculated, the expected value of which is 1 if the motif peak sets are no more variable than the background peak sets for that motif. High variability is correlated to TF activity in a biological state or in response to a perturbation. In the example PBMC dataset, a number of TFs known to correlate with immune cell activity and division show high variability by ChromVAR analysis (Fig. 3j). ChromVAR thus provides biologically relevant, comprehensive views of TF activity in cell subsets from scATAC-seq data.

2.2.9. Enhancer–Promoter Looping Predictions by Cicero

The goal of Cicero is to detect connections between regulatory DNA elements and their putative target genes, using chromatin coaccessibility patterns from scATAC-seq data [138]. For accurate estimation, cells are first aggregated (in groups of 50 based on clustering or trajectory space) using a KNN algorithm to obtain high density, high depth counts. Nearby peaks are then assessed for their coaccessibility structure using covariance analysis. Regularized covariance matrices are generated by Graphical LASSO estimation with distance penalties. Finally, the coaccessibility scores (between −1 and 1) between each pair of accessible sites within a user-defined distance is reported. Additionally, positive coaccessibility scores can detect larger cis-coaccessibility networks (CCANs or chromatin accessibility hubs) by establishing graph structures. The nodes are the peaks of accessibility, and the edges of the graph are coaccessibility scores above a user defined threshold. Communities within this genome-wide graph identify CCANs and can be found using the Louvain algorithm. Cicero also provides an extension toolkit for differential peak analysis, clustering, visualization, trajectory analysis, and so on of scATAC-seq experiments using Monocle 3.

2.3. Computational Methods to Analyze Single-Cell DNA Multiomics Data

Rapid development of single-cell multiomics experimental protocols has prompted coevolution of innovative computational methods to integrate heterogeneous data from multiple datasets. This heterogeneity is compounded by missing data from the various technologies. In theory, all measurements should come from the same cell (paired or matched data) but to take advantage of already published datasets, more often than not, unpaired data from different cells from the same sample or different samples are used. One of the approaches for integration uses a common latent space to generate a combined reference [139]. In many cases, scRNA-seq is used as the common reference as it is the most developed single-cell application and given the reputation of RNA in the central dogma. To fill in the missing values, imputing algorithms have been developed based on heuristic model (MAGIC [140]), regression model (MISC [141], scImpute [142], VIPER [143]), matrix completion (mcImpute [144]), network-based (netImpute [145], scGAIN [146]) or using autoencoder (AutoImpute [147]).

The initial scRNA-seq preprocessing steps remain as mentioned before (Subheading 2.1.1) and the data is further integrated by CCA using latent space projection. This can be effectively used for integrating expression with epigenome data (scM&T-seq, scMT-seq, snmCT-seq, etc.). The Variational Bayes (VB) method is based on Bayesian modeling and is best suited for integrating single-cell genome sequencing with expression. LASSO (Least Absolute Shrinkage and Selection Operator) and GBR (Gradient Boosting Regression) are regression-based methods used mainly for expression and chromatin accessibility data (scCAT-seq, ATAC-RNA-seq, etc.). scRNA and scATAC-seq data have also been analyzed using matrix factorization methods such as integrative Nonnegative Matrix Factorization (iNMF), coupled Nonnegative Matrix Factorizations (coupleNMF), Group Factor Analysis (GFA), and Independent Component Analysis (ICA). TOPIC modeling has been used for integrating scRNA-seq with single-cell CRISPR. For integrating >2 single-cell modalities, Deep Learning (Autoencoder) has been used. Several published software tools (e.g., Seurat, MOFA+ [148], MATCHER [149], MuSiC [150], MIMOSCA [151], LIGER, clonealign [152]) are used to handle single-cell multiomics data [153].

2.4. Conclusions

The advent of single-cell genomics has revolutionized the world of big data, adding volume, complexity, resolution, and dimension. While this revolution promises conceptual insights into biological processes and disease development, it presents significant analytical and statistical challenges. In this chapter, we summarize the variety of single-cell methodologies available till date focusing on the transcriptome and epigenome. We also systematically outline and exemplify how scientists embarking on the single-cell journey might approach the analysis using a small publicly available PBMC dataset for each category of single-cell experiment. We discuss the pros and cons of popularly used packages and suggest alternatives when necessary. Overall, we hope this chapter will encourage scientists to consider single-cell applications to answer their favorite biological questions.

3. Notes

  1. Sample quality and dead cell removal: sample quality is one of the key factors that determines experimental success. For scRNA-seq, it is necessary to ensure that the single-cell suspension has high viability and if not, dead cells must be removed prior to droplet production. Poor cell viability may result in increased ambient RNA that makes it difficult to distinguish cells from empty droplets. For scATAC-seq, it is important to use intact nuclei which can be especially challenging to obtain from frozen tissues.

  2. Number of cells and read depth: the number of cells queried, and individual cell read depth greatly influences biological interpretations. For example, enough cells per sample must be targeted for droplet generation to ensure proper sampling of both abundant and rare cell populations (tissue-resident stem cells, cancer subclones, etc.). Similarly, sufficient read depth is required to detect low-abundance transcripts and for proper detection of differentially expressed genes and cell annotation.

  3. Software versions: due to the constant evolution of software dedicated to single-cell data analysis, and some stochasticity in the performance of dimensionality reduction, clustering or imputation algorithms, it is imperative that researchers record the software versions for every experiment and adequately report in publications.

  4. Filtering: filtering out low quality cells from the analysis is perhaps the most important upstream step in single-cell analysis as it greatly influences biological interpretations. QC steps outlined in Subheadings 2.1.2 and 2.2.2 are used to remove empty droplets, multiplets, apoptotic or lysing cells with high mitochondrial reads, cells with poor transposition or other technical artifacts.

  5. Batch effects: systemic variations in single-cell datasets are produced due to differences in the source/lab generating the data, technologies used to make the single-cell libraries, sequencing platforms, scale of the data, and so on. Additionally, researchers may choose to integrate multiple modalities produced by different labs. For these purposes, efficacious batch correction techniques (see Subheadings 2.1.3 and 2.2.3) need to be applied to dissect true biological variation.

  6. Cell annotation: marker gene identification of cell clusters give clues to cell identity. While several automatic cell classification methods have been identified, we propose that researchers always check several marker genes manually for correct identification of cell types. This is because many annotation tools perform poorly for rare cell populations. Additionally, clustering itself may not be optimal and may have mixed populations of cells.

References

  • 1.International Human Genome Sequencing C (2004) Finishing the euchromatic sequence of the human genome. Nature 431(7011): 931–945. 10.1038/nature03001 [DOI] [PubMed] [Google Scholar]
  • 2.Sachidanandam R, Weissman D, Schmidt SC, Kakol JM, Stein LD, Marth G, Sherry S, Mullikin JC, Mortimore BJ, Willey DL, Hunt SE, Cole CG, Coggill PC, Rice CM, Ning Z, Rogers J, Bentley DR, Kwok PY, Mardis ER, Yeh RT, Schultz B, Cook L, Davenport R, Dante M, Fulton L, Hillier L, Waterston RH, JD MP, Gilman B, Schaffner S, Van Etten WJ, Reich D, Higgins J, Daly MJ, Blumenstiel B, Baldwin J, Stange-Thomann N, Zody MC, Linton L, Lander ES, Altshuler D, International SNPMWG (2001) A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature 409(6822): 928–933. 10.1038/35057149 [DOI] [PubMed] [Google Scholar]
  • 3.International HapMap C (2005) A haplotype map of the human genome. Nature 437(7063):1299–1320. 10.1038/nature04226 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Genomes Project C, Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA, Hurles ME, GA MV (2010) A map of human genome variation from population-scale sequencing. Nature 467(7319):1061–1073. 10.1038/nature09534 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Bernstein BE, Stamatoyannopoulos JA, Costello JF, Ren B, Milosavljevic A, Meissner A, Kellis M, Marra MA, Beaudet AL, Ecker JR, Farnham PJ, Hirst M, Lander ES, Mikkelsen TS, Thomson JA (2010) The NIH roadmap epigenomics mapping consortium. Nat Biotechnol 28(10):1045–1048. 10.1038/nbt1010-1045 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Consortium EP, Birney E, Stamatoyannopoulos JA, Dutta A, Guigo R, Gingeras TR, Margulies EH, Weng Z, Snyder M, Dermitzakis ET,Thurman RE, Kuehn MS, Taylor CM, Neph S, Koch CM, Asthana S, Malhotra A, Adzhubei I, Greenbaum JA, Andrews RM, Flicek P, Boyle PJ, Cao H, Carter NP, Clelland GK, Davis S, Day N, Dhami P, Dillon SC, Dorschner MO, Fiegler H, Giresi PG, Goldy J, Hawrylycz M, Haydock A, Humbert R, James KD, Johnson BE, Johnson EM, Frum TT, Rosenzweig ER, Karnani N, Lee K, Lefebvre GC, Navas PA, Neri F, Parker SC, Sabo PJ, Sandstrom R, Shafer A, Vetrie D, Weaver M, Wilcox S, Yu M, Collins FS, Dekker J, Lieb JD, Tullius TD, Crawford GE,Sunyaev S, Noble WS, Dunham I, Denoeud F, Reymond A, Kapranov P, Rozowsky J, Zheng D, Castelo R, Frankish A, Harrow J, Ghosh S, Sandelin A, Hofacker IL, Baertsch R, Keefe D, Dike S, Cheng J, Hirsch HA, Sekinger EA, Lagarde J, Abril JF, Shahab A, Flamm C, Fried C, Hackermuller J, Hertel J, Lindemeyer M, Missal K, Tanzer A, Washietl S, Korbel J, Emanuelsson O, Pedersen JS, Holroyd N, Taylor R, Swarbreck D, Matthews N, Dickson MC, Thomas DJ, Weirauch MT, Gilbert J, Drenkow J, Bell I, Zhao X, Srinivasan KG, Sung WK, Ooi HS, Chiu KP, Foissac S, Alioto T, Brent M, Pachter L, Tress ML, Valencia A, Choo SW, Choo CY, Ucla C, Manzano C, Wyss C, Cheung E, Clark TG, Brown JB, Ganesh M, Patel S, Tammana H, Chrast J, Henrichsen CN,Kai C, Kawai J, Nagalakshmi U, Wu J, Lian Z, Lian J, Newburger P, Zhang X, Bickel P, Mattick JS, Carninci P, Hayashizaki Y, Weissman S, Hubbard T, Myers RM, Rogers J, Stadler PF, Lowe TM, Wei CL, Ruan Y, Struhl K, Gerstein M, Antonarakis SE, Fu Y, Green ED, Karaoz U, Siepel A, Taylor J, Liefer LA, Wetterstrand KA,Good PJ, Feingold EA, Guyer MS, Cooper GM, Asimenos G, Dewey CN, Hou M, Nikolaev S, Montoya-Burgos JI, Loytynoja A, Whelan S, Pardi F, Massingham T, Huang H, Zhang NR, Holmes I, Mullikin JC, Ureta-Vidal A, Paten B, Seringhaus M, Church D, Rosenbloom K, Kent WJ, Stone EA, Program NCS,Baylor College of Medicine Human Genome Sequencing C, Washington University Genome Sequencing C, Broad I, Children’s Hospital Oakland Research I, Batzoglou S, Goldman N, Hardison RC, Haussler D, Miller W, Sidow A, Trinklein ND, Zhang ZD, Barrera L, Stuart R, King DC, Ameur A, Enroth S, Bieda MC, Kim J, Bhinge AA, Jiang N, Liu J, Yao F, Vega VB, Lee CW, Ng P, Shahab A, Yang A, Moqtaderi Z, Zhu Z, Xu X, Squazzo S, Oberley MJ, Inman D, Singer MA, Richmond TA, Munn KJ, Rada-Iglesias A, Wallerman O, Komorowski J, Fowler JC, Couttet P, Bruce AW,Dovey OM, Ellis PD, Langford CF, Nix DA,Euskirchen G, Hartman S, Urban AE, Kraus P, Van Calcar S, Heintzman N, Kim TH, Wang K, Qu C, Hon G, Luna R, Glass CK, Rosenfeld MG, Aldred SF, Cooper SJ, Halees A, Lin JM, Shulha HP, Zhang X, Xu M, Haidar JN, Yu Y, Ruan Y, Iyer VR, Green RD, Wadelius C, Farnham PJ, Ren B, Harte RA, Hinrichs AS, Trumbower H, Clawson H, Hillman-Jackson J, Zweig AS, Smith K, Thakkapallayil A, Barber G, Kuhn RM, Karolchik D, Armengol L, Bird CP, de Bakker PI, Kern AD, Lopez-Bigas N, Martin JD, Stranger BE, Woodroffe A, Davydov E, Dimas A, Eyras E, Hallgrimsdottir IB, Huppert J, Zody MC, Abecasis GR, Estivill X, Bouffard GG, Guan X, Hansen NF, Idol JR, Maduro VV, Maskeri B, McDowell JC, Park M, Thomas PJ, Young AC, Blakesley RW, Muzny DM, Sodergren E, Wheeler DA, Worley KC, Jiang H, Weinstock GM, Gibbs RA, Graves T, Fulton R, Mardis ER, Wilson RK, Clamp M, Cuff J, Gnerre S, Jaffe DB, Chang JL, Lindblad-Toh K, Lander ES, Koriabine M, Nefedov M, Osoegawa K, Yoshinaga Y, Zhu B, de Jong PJ (2007) Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project. Nature 447 (7146):799–816. 10.1038/nature05874 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Celniker SE, Dillon LA, Gerstein MB, Gunsalus KC, Henikoff S, Karpen GH, Kellis M, Lai EC,Lieb JD, MacAlpine DM, Micklem G, Piano F, Snyder M, Stein L, White KP, Waterston RH, modENCODE Consortium (2009) Unlocking the secrets of the genome. Nature 459(7249):927–930. 10.1038/459927a [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Stunnenberg HG, International Human Epigenome C, Hirst M (2016) The International Human Epigenome Consortium: a Blueprint for scientific collaboration and discovery. Cell 167 (5):1145–1149. doi: 10.1016/j.cell.2016.11.007 [DOI] [PubMed] [Google Scholar]
  • 9.Rozenblatt-Rosen O, Stubbington MJT, Regev A, Teichmann SA (2017) The Human Cell Atlas: from vision to reality. Nature 550(7677):451–453. 10.1038/550451a [DOI] [PubMed] [Google Scholar]
  • 10.Slyper M, Porter CBM, Ashenberg O, Waldman J, Drokhlyansky E, Wakiro I, Smillie C, Smith-Rosario G, Wu J, Dionne D, Vigneau S, Jane-Valbuena J, Tickle TL, Napolitano S, Su MJ, Patel AG, Karlstrom A, Gritsch S, Nomura M, Waghray A, Gohil SH, Tsankov AM, Jerby-Arnon L, Cohen O, Klughammer J, Rosen Y, Gould J, Nguyen L, Hofree M, Tramontozzi PJ, Li B, Wu CJ, Izar B, Haq R, Hodi FS, Yoon CH, Hata AN, Baker SJ, Suva ML, Bueno R, Stover EH, Clay MR, Dyer MA, Collins NB, Matulonis UA, Wagle N, Johnson BE, Rotem A, Rozenblatt-Rosen-O, Regev A (2020) A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors. Nat Med 26(5):792–802. 10.1038/s41591-0200844-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Tang F, Barbacioru C, Nordman E, Li B, Xu N, Bashkirov VI, Lao K, Surani MA (2010) RNA-Seq analysis to capture the transcriptome landscape of a single cell. Nat Protoc 5(3):516–535. 10.1038/nprot.2009.236 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Tang F, Barbacioru C, Wang Y, Nordman E, Lee C, Xu N, Wang X, Bodeau J, Tuch BB, Siddiqui A, Lao K, Surani MA (2009) mRNA-Seq whole-transcriptome analysis of a single cell. Nat Methods 6(5):377–382. 10.1038/nmeth.1315 [DOI] [PubMed] [Google Scholar]
  • 13.Islam S, Kjallquist U, Moliner A, Zajac P, Fan JB, Lonnerberg P, Linnarsson S (2011) Characterization of the single-cell transcriptional landscape by highly multiplex RNA-seq. Genome Res 21(7):1160–1167. 10.1101/gr.110882.110 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Pollen AA, Nowakowski TJ, Shuga J, Wang X, Leyrat AA, Lui JH, Li N, Szpankowski L, Fowler B, Chen P, Ramalingam N, Sun G, Thu M, Norris M, Lebofsky R, Toppani D, Kemp DW 2nd, Wong M, Clerkson B, Jones BN, Wu S, Knutsson L, Alvarado B, Wang J, Weaver LS, May AP, Jones RC, Unger MA, Kriegstein AR, West JA (2014) Low-coverage single-cell mRNA sequencing reveals cellular heterogeneity and activated signaling pathways in developing cerebral cortex. Nat Biotechnol 32(10):1053–1058. 10.1038/nbt.2967 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Islam S, Zeisel A, Joost S, La Manno G, Zajac P, Kasper M, Lonnerberg P, Linnarsson S (2014) Quantitative single-cell RNA-seq with unique molecular identifiers. Nat Methods 11(2):163–166. 10.1038/nmeth.2772 [DOI] [PubMed] [Google Scholar]
  • 16.Ramskold D, Luo S, Wang YC, Li R, Deng Q, Faridani OR, Daniels GA, Khrebtukova I, Loring JF, Laurent LC, Schroth GP, Sandberg R (2012) Full-length mRNA-Seq from single-cell levels of RNA and individual circulating tumor cells. Nat Biotechnol 30(8): 777–782. 10.1038/nbt.2282 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Picelli S, Faridani OR, Bjorklund AK, Winberg G, Sagasser S, Sandberg R (2014) Full-length RNA-seq from single cells using Smart-seq2. Nat Protoc 9(1):171–181. 10.1038/nprot.2014.006 [DOI] [PubMed] [Google Scholar]
  • 18.Hagemann-Jensen M, Ziegenhain C, Chen P, Ramskold D, Hendriks GJ, Larsson AJM, Faridani OR, Sandberg R (2020) Single-cell RNA counting at allele and isoform resolution using Smart-seq3. Nat Biotechnol 38(6): 708–714. 10.1038/s41587-020-0497-0 [DOI] [PubMed] [Google Scholar]
  • 19.Hashimshony T, Wagner F, Sher N, Yanai I (2012) CEL-Seq: single-cell RNA-Seq by multiplexed linear amplification. Cell Rep 2(3):666–673. 10.1016/j.celrep.2012.08.003 [DOI] [PubMed] [Google Scholar]
  • 20.Hashimshony T, Senderovich N, Avital G, Klochendler A, de Leeuw Y, Anavy L, Gennert D, Li S, Livak KJ, Rozenblatt-Rosen O, Dor Y, Regev A, Yanai I (2016) CEL-Seq2: sensitive highly-multiplexed single-cell RNA-Seq. Genome Biol 17:77. 10.1186/s13059-0160938-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Jaitin DA, Kenigsberg E, Keren-Shaul H, Elefant N, Paul F, Zaretsky I, Mildner A, Cohen N, Jung S, Tanay A, Amit I (2014) Massively parallel single-cell RNA-seq for marker-free decomposition of tissues into cell types. Science 343(6172):776–779. 10.1126/science.1247651 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Keren-Shaul H, Kenigsberg E, Jaitin DA, David E, Paul F, Tanay A, Amit I (2019) MARS-seq2.0: an experimental and analytical pipeline for indexed sorting combined with single-cell RNA sequencing. Nat Protoc 14(6):1841–1862. 10.1038/s41596-019-0164-4 [DOI] [PubMed] [Google Scholar]
  • 23.Fan HC, Fu GK, Fodor SP (2015) Expression profiling. Combinatorial labeling of single cells for gene expression cytometry. Science 347(6222):1258367. 10.1126/science.1258367 [DOI] [PubMed] [Google Scholar]
  • 24.Macosko EZ, Basu A, Satija R, Nemesh J, Shekhar K, Goldman M, Tirosh I, Bialas AR, Kamitaki N, Martersteck EM, Trombetta JJ, Weitz DA, Sanes JR, Shalek AK, Regev A, McCarroll SA (2015) Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets. Cell 161(5):1202–1214. 10.1016/j.cell.2015.05.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Klein AM, Mazutis L, Akartuna I, Tallapragada N, Veres A, Li V, Peshkin L, Weitz DA, Kirschner MW (2015) Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells. Cell 161(5): 1187–1201. 10.1016/j.cell.2015.04.044 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Zheng GX, Terry JM, Belgrader P, Ryvkin P, Bent ZW, Wilson R, Ziraldo SB, Wheeler TD, McDermott GP, Zhu J, Gregory MT, Shuga J, Montesclaros L, Underwood JG, Masquelier DA, Nishimura SY, Schnall-Levin-M,Wyatt PW, Hindson CM, Bharadwaj R, Wong A, Ness KD, Beppu LW, Deeg HJ, McFarland C, Loeb KR, Valente WJ, Ericson NG, Stevens EA, Radich JP, Mikkelsen TS, Hindson BJ, Bielas JH (2017) Massively parallel digital transcriptional profiling of single cells. Nat Commun 8:14049. 10.1038/ncomms14049 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Horvath S, Raj K (2018) DNA methylation-based biomarkers and the epigenetic clock theory of ageing. Nat Rev Genet 19(6):371–384. 10.1038/s41576-018-0004-3 [DOI] [PubMed] [Google Scholar]
  • 28.Sen P, Shah PP, Nativio R, Berger SL (2016) Epigenetic mechanisms of longevity and aging. Cell 166(4):822–839. 10.1016/j.cell.2016.07.050 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Buenrostro JD, Wu B, Litzenburger UM, Ruff D, Gonzales ML, Snyder MP, Chang HY, Greenleaf WJ (2015) Single-cell chromatin accessibility reveals principles of regulatory variation. Nature 523(7561):486–490. 10.1038/nature14590 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Guo H, Zhu P, Wu X, Li X, Wen L, Tang F (2013) Single-cell methylome landscapes of mouse embryonic stem cells and early embryos analyzed using reduced representation bisulfite sequencing. Genome Res 23(12):2126–2135. 10.1101/gr.161679.113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Miura F, Enomoto Y, Dairiki R, Ito T (2012) Amplification-free whole-genome bisulfite sequencing by post-bisulfite adaptor tagging. Nucleic Acids Res 40(17):e136. 10.1093/nar/gks454 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Smallwood SA, Lee HJ, Angermueller C, Krueger F, Saadeh H, Peat J, Andrews SR, Stegle O, Reik W, Kelsey G (2014) Single-cell genome-wide bisulfite sequencing for assessing epigenetic heterogeneity. Nat Methods 11(8):817–820. 10.1038/nmeth.3035 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Farlik M, Sheffield NC, Nuzzo A, Datlinger P, Schonegger A, Klughammer J, Bock C (2015) Single-cell DNA methylome sequencing and bioinformatic inference of epigenomic cell-state dynamics. Cell Rep 10(8):1386–1397. 10.1016/j.celrep.2015.02.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Rotem A, Ram O, Shoresh N, Sperling RA, Goren A, Weitz DA, Bernstein BE (2015) Single-cell ChIP-seq reveals cell subpopulations defined by chromatin state. Nat Biotechnol 33(11):1165–1172. 10.1038/nbt.3383 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kaya-Okur HS, Wu SJ, Codomo CA, Pledger ES,Bryson TD, Henikoff JG, Ahmad K, Henikoff S (2019) CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Commun 10(1):1930. 10.1038/s41467-019-09982-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Wu SJ, Furlan SN, Mihalas AB, Kaya-Okur H, Feroze AH, Emerson SN, Zheng Y, Carson K, Cimino PJ, Keene CD, Holland EC, Sarthy JF,Gottardo R, Ahmad K, Henikoff S, Patel AP (2020) Single-cell analysis of chromatin silencing programs in developmental and tumor progression. bioRxiv:2020.2009.2004.282418 10.1101/2020.09.04.282418 [DOI] [Google Scholar]
  • 37.Bartosovic M, Kabbe M, Castelo-Branco G (2020) Single-cell profiling of histone modifications in the mouse brain. bioRxiv:2020.2009.2002.279703 10.1101/2020.09.02.279703 [DOI] [Google Scholar]
  • 38.Cusanovich DA, Daza R, Adey A, Pliner HA, Christiansen L, Gunderson KL, Steemers FJ, Trapnell C, Shendure J (2015) Multiplex single cell profiling of chromatin accessibility by combinatorial cellular indexing. Science 348(6237):910–914. 10.1126/science.aab1601 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Satpathy AT, Granja JM, Yost KE, Qi Y, Meschi F, McDermott GP, Olsen BN, Mumbach MR, Pierce SE, Corces MR, Shah P, Bell JC, Jhutty D, Nemec CM, Wang J, Wang L, Yin Y, Giresi PG, Chang ALS, Zheng GXY, Greenleaf WJ, Chang HY (2019) Massively parallel single-cell chromatin landscapes of human immune cell development and intratumoral T cell exhaustion. Nat Biotechnol 37(8):925–936. 10.1038/s41587-019-0206-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Bonev B, Cavalli G (2016) Organization and function of the 3D genome. Nat Rev Genet 17(12):772. 10.1038/nrg.2016.147 [DOI] [PubMed] [Google Scholar]
  • 41.Nagano T, Lubling Y, Stevens TJ, Schoenfelder S, Yaffe E, Dean W, Laue ED, Tanay A, Fraser P (2013) Single-cell Hi-C reveals cell-to-cell variability in chromosome structure. Nature 502(7469):59–64. 10.1038/nature12593 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Rao SS, Huntley MH, Durand NC, Stamenova EK, Bochkov ID, Robinson JT, Sanborn AL,Machol I, Omer AD, Lander ES, Aiden EL (2014) A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell 159(7):1665–1680. 10.1016/j.cell.2014.11.021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Tan L, Xing D, Chang CH, Li H, Xie XS (2018) Three-dimensional genome structures of single diploid human cells. Science 361(6405):924–928. 10.1126/science.aat5641 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Angermueller C, Clark SJ, Lee HJ, Macaulay IC,Teng MJ, Hu TX, Krueger F, Smallwood S, Ponting CP, Voet T, Kelsey G, Stegle O, Reik W (2016) Parallel single-cell sequencing links transcriptional and epigenetic heterogeneity. Nat Methods 13(3):229–232. 10.1038/nmeth.3728 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Hu Y, Huang K, An Q, Du G, Hu G, Xue J, Zhu X, Wang CY, Xue Z, Fan G (2016) Simultaneous profiling of transcriptome and DNA methylome from a single cell. Genome Biol 17:88. 10.1186/s13059-016-0950-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Luo C, Liu H, Wang B-A, Bartlett A, Rivkin A, Nery JR, Ecker JR (2018) Multiomic profiling of transcriptome and DNA methylome in single nuclei with molecular partitioning. bioRxiv:434845. 10.1101/434845 [DOI] [Google Scholar]
  • 47.Hou Y, Guo H, Cao C, Li X, Hu B, Zhu P, Wu X, Wen L, Tang F, Huang Y, Peng J (2016) Single-cell triple omics sequencing reveals genetic, epigenetic, and transcriptomic heterogeneity in hepatocellular carcinomas. Cell Res 26(3):304–319. 10.1038/cr.2016.23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Pott S (2017) Simultaneous measurement of chromatin accessibility, DNA methylation, and nucleosome phasing in single cells. Elife 6. 10.7554/eLife.23203 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Clark SJ, Argelaguet R, Kapourani CA, Stubbs TM, Lee HJ, Alda-Catalinas C, Krueger F, Sanguinetti G, Kelsey G, Marioni JC,Stegle O, Reik W (2018) scNMT-seq enables joint profiling of chromatin accessibility DNA methylation and transcription in single cells. Nat Commun 9(1):781. 10.1038/s41467-018-03149-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Wang Y, Yuan P, Yan Z, Yang M, Huo Y, Nie Y, Zhu X, Yan L, Qiao J (2019) Single-cell multiomics sequencing reveals the functional regulatory landscape of early embryos. bioRxiv:803890. 10.1101/803890 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Guo F, Li L, Li J, Wu X, Hu B, Zhu P, Wen L, Tang F (2017) Single-cell multi-omics sequencing of mouse early embryos and embryonic stem cells. Cell Res 27(8): 967–988. 10.1038/cr.2017.82 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Liu L, Liu C, Quintero A, Wu L, Yuan Y, Wang M, Cheng M, Leng L, Xu L, Dong G, Li R, Liu Y, Wei X, Xu J, Chen X, Lu H, Chen D, Wang Q, Zhou Q, Lin X, Li G, Liu S, Wang Q, Wang H, Fink JL, Gao Z, Liu X, Hou Y, Zhu S, Yang H, Ye Y, Lin G, Chen F, Herrmann C, Eils R, Shang Z, Xu X (2019) Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity. Nat Commun 10(1):470. 10.1038/s41467-018-08205-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Reyes M, Billman K, Hacohen N, Blainey PC (2019)Simultaneous profiling of gene expression and chromatin accessibility in single cells. Adv Biosyst 3(11). 10.1002/adbi.201900065 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Li G, Liu Y, Zhang Y, Kubo N, Yu M, Fang R, Kellis M, Ren B (2019) Joint profiling of DNA methylation and chromatin architecture in single cells. Nat Methods 16(10):991–993. 10.1038/s41592-0190502-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Lee DS, Luo C, Zhou J, Chandran S, Rivkin A, Bartlett A, Nery JR, Fitzpatrick C, O’Connor C, Dixon JR, Ecker JR (2019) Simultaneous profiling of 3D genome structure and DNA methylation in single human cells. Nat Methods 16(10):999–1006. 10.1038/s41592-0190547-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Stoeckius M, Hafemeister C, Stephenson W, Houck-Loomis B, Chattopadhyay PK, Swerdlow H, Satija R, Smibert P (2017) Simultaneous epitope and transcriptome measurement in single cells. Nat Methods 14(9): 865–868. 10.1038/nmeth.4380 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Peterson VM, Zhang KX, Kumar N, Wong J, Li L, Wilson DC, Moore R, McClanahan TK, Sadekova S, Klappenbach JA (2017) Multiplexed quantification of proteins and transcripts in single cells. Nat Biotechnol 35(10): 936–939. 10.1038/nbt.3973 [DOI] [PubMed] [Google Scholar]
  • 58.Mimitou EP, Lareau CA, Chen KY, Zorzetto-Fernandes AL, Takeshima Y, Luo W, Huang T-S,Yeung B, Thakore PI, Wing JB, Nazor KL,Sakaguchi S, Ludwig LS, Sankaran VG, Regev A, Smibert P (2020) Scalable, multi-modal profiling of chromatin accessibility and protein levels in single cells. bioRxiv:2020.2009.2008.286914 10.1101/2020.09.08.286914 [DOI] [Google Scholar]
  • 59.Lareau CA, Ludwig LS, Muus C, Gohil SH, Zhao T, Chiang Z, Pelka K, Verboon JM, Luo W, Christian E, Rosebrock D, Getz G, Boland GM, Chen F, Buenrostro JD, Hacohen N, Wu CJ, Aryee MJ, Regev A, Sankaran VG (2021) Massively parallel single-cell mitochondrial DNA genotyping and chromatin profiling. Nat Biotechnol 39(4):451–461. 10.1038/s41587-0200645-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Stoeckius M, Zheng S, Houck-Loomis B, Hao S, Yeung BZ, Mauck WM 3rd, Smibert P, Satija R (2018) Cell Hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics. Genome Biol 19(1):224. 10.1186/s13059-018-1603-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Gaublomme JT, Li B, McCabe C, Knecht A, Drokhlyansky E, Wittenberghe NV, Waldman J, Dionne D, Nguyen L, Jager PD, Yeung B, Zhao X, Habib N, Rozenblatt-Rosen O, Regev A (2018) Nuclei multiplexing with barcoded antibodies for single nucleus genomics. bioRxiv:476036. 10.1101/476036 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Dixit A, Parnas O, Li B, Chen J, Fulco CP, Jerby-Arnon L, Marjanovic ND, Dionne D, Burks T, Raychowdhury R, Adamson B, Norman TM, Lander ES, Weissman JS, Friedman N, Regev A (2016) Perturb-Seq: dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell 167(7):1853–1866.e1817. 10.1016/j.cell.2016.11.038 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Jaitin DA, Weiner A, Yofe I, Lara-Astiaso D, Keren-Shaul H, David E, Salame TM, Tanay A, van Oudenaarden A, Amit I (2016) Dissecting immune circuits by linking CRISPR-pooled screens with single-cell RNA-Seq. Cell 167(7):1883–1896.e1815. 10.1016/j.cell.2016.11.039 [DOI] [PubMed] [Google Scholar]
  • 64.Datlinger P, Rendeiro AF, Schmidl C, Krausgruber T, Traxler P, Klughammer J, Schuster LC, Kuchler A, Alpar D, Bock C (2017) Pooled CRISPR screening with single-cell transcriptome readout. Nat Methods 14(3):297–301. 10.1038/nmeth.4177 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Adamson B, Norman TM, Jost M, Cho MY, Nunez JK, Chen Y, Villalta JE, Gilbert LA, Horlbeck MA, Hein MY, Pak RA, Gray AN, Gross CA, Dixit A, Parnas O, Regev A, Weissman JS (2016) A multiplexed single-cell CRISPR screening platform enables systematic dissection of the unfolded protein response. Cell 167(7):1867–1882.e1821. 10.1016/j.cell.2016.11.048 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Xie S, Duan J, Li B, Zhou P, Hon GC (2017) Multiplexed engineering and analysis of combinatorial enhancer activity in single cells. Mol Cell 66(2):285–299.e285. 10.1016/j.molcel.2017.03.007 [DOI] [PubMed] [Google Scholar]
  • 67.Replogle JM, Norman TM, Xu A, Hussmann JA,Chen J, Cogan JZ, Meer EJ, Terry JM, Riordan DP, Srinivas N, Fiddes IT, Arthur JG, Alvarado LJ, Pfeiffer KA, Mikkelsen TS, Weissman JS, Adamson B (2020) Combinatorial single-cell CRISPR screens by direct guide RNA capture and targeted sequencing. Nat Biotechnol 38(8):954–961. 10.1038/s41587-020-0470-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Rostom R, Svensson V, Teichmann SA, Kar G (2017) Computational approaches for interpreting scRNA-seq data. FEBS Lett 591(15):2213–2225. 10.1002/1873-3468.12684 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR (2013) STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29(1):15–21. 10.1093/bioinformatics/bts635 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Melsted P, Booeshaghi AS, Gao F, Beltrame E, Lu L, Hjorleifsson KE, Gehring J, Pachter L (2019) Modular and efficient pre-processing of single-cell RNA--seq. bioRxiv:673285. 10.1101/673285 [DOI] [Google Scholar]
  • 71.Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C (2017) Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods 14(4):417–419. 10.1038/nmeth.4197 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Srivastava A, Malik L, Smith T, Sudbery I, Patro R (2019) Alevin efficiently estimates accurate gene abundances from dscRNA-seq data. Genome Biol 20(1):65. 10.1186/s13059-019-1670-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Amezquita RA, Lun ATL, Becht E, Carey VJ, Carpp LN, Geistlinger L, Marini F, Rue-Albrecht K, Risso D, Soneson C, Waldron L, Pages H, Smith ML, Huber W, Morgan M, Gottardo R, Hicks SC (2020) Orchestrating single-cell analysis with bioconductor. Nat Methods 17(2):137–145. 10.1038/s41592-0190654-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Tian L, Su S, Dong X, Amann-Zalcenstein D, Biben C, Seidi A, Hilton DJ, Naik SH, Ritchie ME (2018) scPipe: a flexible R/Bioconductor preprocessing pipeline for single-cell RNA-sequencing data. PLoS Comput Biol 14(8): e1006361. 10.1371/journal.pcbi.1006361 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Wang Z, Hu J, Johnson WE, Campbell JD (2019) scruff: an R/bioconductor package for preprocessing single-cell RNA-sequencing data. BMC Bioinformatics 20(1):222. 10.1186/s12859-019-2797-2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Jiang P (2019) Quality control of single-cell RNA-seq. Methods Mol Biol 1935:1–9. 10.1007/978-1-4939-9057-3_1 [DOI] [PubMed] [Google Scholar]
  • 77.Abugessaisa I, Noguchi S, Cardon M, Hasegawa A, Watanabe K, Takahashi M, Suzuki H, Katayama S, Kere J, Kasukawa T (2020) Quality assessment of single-cell RNA sequencing data by coverage skewness analysis. bioRxiv:2019.2012.2031.890269 10.1101/2019.12.31.890269 [DOI] [Google Scholar]
  • 78.McCarthy DJ, Campbell KR, Lun AT, Wills QF (2017) Scater: pre-processing, quality control, normalization and visualization of single-cell RNA-seq data in R. Bioinformatics 33(8):1179–1186. 10.1093/bioinformatics/btw777 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Butler A, Hoffman P, Smibert P, Papalexi E, Satija R (2018) Integrating single-cell transcriptomic data across different conditions, technologies, and species. Nat Biotechnol 36(5):411–420. 10.1038/nbt.4096 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Ilicic T, Kim JK, Kolodziejczyk AA, Bagger FO, McCarthy DJ, Marioni JC, Teichmann SA (2016) Classification of low quality cells from single-cell RNA-seq data. Genome Biol 17:29. 10.1186/s13059-016-0888-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Young MD, Behjati S (2020) SoupX removes ambient RNA contamination from droplet based single-cell RNA sequencing data. bioRxiv:303727. 10.1101/303727 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Heaton H, Talman AM, Knights A, Imaz M, Gaffney D, Durbin R, Hemberg M, Lawniczak M (2019) Souporcell: robust clustering of single cell RNAseq by genotype and ambient RNA inference without reference genotypes. bioRxiv:699637. 10.1101/699637 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Yang S, Corbett SE, Koga Y, Wang Z, Johnson WE, Yajima M, Campbell JD (2020) Decontamination of ambient RNA in single-cell RNA-seq with DecontX. Genome Biol 21(1):57. 10.1186/s13059-020-1950-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.McGinnis CS, Murrow LM, Gartner ZJ (2019) DoubletFinder: doublet detection in single-cell RNA sequencing data using artificial nearest neighbors. Cell Syst 8(4): 329–337.e324. 10.1016/j.cels.2019.03.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.DePasquale EAK, Schnell DJ, Van Camp PJ, Valiente-Alandi I, Blaxall BC, Grimes HL, Singh H, Salomonis N (2019) DoubletDecon: deconvoluting doublets from single-cell RNA-sequencing data. Cell Rep 29(6): 1718–1727.e1718. 10.1016/j.celrep.2019.09.082 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Wolock SL, Lopez R, Klein AM (2019) Scrublet: computational identification of cell doublets in single-cell transcriptomic data. Cell Syst 8(4):281–291.e289. 10.1016/j.cels.2018.11.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Bernstein NJ, Fong NL, Lam I, Roy MA, Hendrickson DG, Kelley DR (2020) Solo: doublet identification in single-cell RNA-Seq via semi-supervised deep learning. Cell Syst 11(1):95–101.e105. 10.1016/j.cels.2020.05.010 [DOI] [PubMed] [Google Scholar]
  • 88.Hardoon DR, Szedmak S, Shawe-Taylor J (2004) Canonical correlation analysis: an overview with application to learning methods. Neural Comput 16(12):2639–2664. 10.1162/0899766042321814 [DOI] [PubMed] [Google Scholar]
  • 89.Haghverdi L, Lun ATL, Morgan MD, Marioni JC (2018) Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors. Nat Biotechnol 36(5):421–427. 10.1038/nbt.4091 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Tran HTN, Ang KS, Chevrier M, Zhang X, Lee NYS, Goh M, Chen J (2020) A benchmark of batch-effect correction methods for single-cell RNA sequencing data. Genome Biol 21(1):12. 10.1186/s13059-019-1850-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Korsunsky I, Millard N, Fan J, Slowikowski K, Zhang F, Wei K, Baglaenko Y, Brenner M, Loh PR, Raychaudhuri S (2019) Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods 16(12): 1289–1296. 10.1038/s41592-019-0619-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Welch JD, Kozareva V, Ferreira A, Vanderburg C, Martin C, Macosko EZ (2019) Single-cell multi-omic integration compares and contrasts features of brain cell identity. Cell 177(7):1873–1887.e1817. 10.1016/j.cell.2019.05.006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Buttner M, Miao Z, Wolf FA, Teichmann SA, Theis FJ (2019) A test metric for assessing single-cell RNA-seq batch correction. Nat Methods 16(1):43–49. 10.1038/s41592-018-0254-1 [DOI] [PubMed] [Google Scholar]
  • 94.Lytal N, Ran D, An L (2020) Normalization methods on single-cell RNA-seq data: an empirical survey. Front Genet 11:41. 10.3389/fgene.2020.00041 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Hafemeister C, Satija R (2019) Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. Genome Biol 20(1): 296. 10.1186/s13059-019-1874-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Wolf FA, Angerer P, Theis FJ (2018) SCANPY: large-scale single-cell gene expression data analysis. Genome Biol 19(1):15. 10.1186/s13059-0171382-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Buettner F, Natarajan KN, Casale FP, Proserpio V, Scialdone A, Theis FJ, Teichmann SA, Marioni JC, Stegle O (2015) Computational analysis of cell-to-cell heterogeneity in single-cell RNA-sequencing data reveals hidden subpopulations of cells. Nat Biotechnol 33(2):155–160. 10.1038/nbt.3102 [DOI] [PubMed] [Google Scholar]
  • 98.Buettner F, Pratanwanich N, McCarthy DJ, Marioni JC, Stegle O (2017) f-scLVM: scalable and versatile factor analysis for single-cell RNA-seq. Genome Biol 18(1):212. 10.1186/s13059-017-1334-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Yip SH, Sham PC, Wang J (2019) Evaluation of tools for highly variable gene discovery from single-cell RNA-seq data. Brief Bioinformat 20(4):1583–1589. 10.1093/bib/bby011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Townes FW, Hicks SC, Aryee MJ, Irizarry RA (2019) Feature selection and dimension reduction for single-cell RNA-Seq based on a multinomial model. Genome Biol 20(1): 295. 10.1186/s13059-019-1861-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Sun S, Zhu J, Ma Y, Zhou X (2019) Accuracy, robustness and scalability of dimensionality reduction methods for single-cell RNA-seq analysis. Genome Biol 20(1):269. 10.1186/s13059-019-1898-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Heiser CN, Lau KS (2020) A quantitative framework for evaluating single-cell data structure preservation by dimensionality reduction techniques. Cell Rep 31(5): 107576. 10.1016/j.celrep.2020.107576 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Tsuyuzaki K, Sato H, Sato K, Nikaido I (2020) Benchmarking principal component analysis for large-scale single-cell RNA-sequencing. Genome Biol 21(1):9. 10.1186/s13059-019-1900-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.vanDerMaaten L, Hinton G (2008) Visualizing data using t-SNE. J Machine Learning Res 9:2579–2605 [Google Scholar]
  • 105.McInnes L, Healy J, Melville J (2018) UMAP: uniform manifold approximation and projection for dimension reduction. arXiv:1802.03426 [Google Scholar]
  • 106.Feng C, Liu S, Zhang H, Guan R, Li D, Zhou F, Liang Y, Feng X (2020) Dimension reduction and clustering models for single-cell RNA sequencing data: a comparative study. Int J Mol Sci 21(60):2181. 10.3390/ijms21062181 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Kiselev VY, Kirschner K, Schaub MT, Andrews T, Yiu A, Chandra T, Natarajan KN,Reik W, Barahona M, Green AR, Hemberg M (2017) SC3: consensus clustering of single-cell RNA-seq data. Nat Methods 14(5):483–486. 10.1038/nmeth.4236 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Qiu X, Mao Q, Tang Y, Wang L, Chawla R, Pliner HA, Trapnell C (2017) Reversed graph embedding resolves complex single-cell trajectories. Nat Methods 14(10):979–982. 10.1038/nmeth.4402 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Duo A, Robinson MD, Soneson C (2018) A systematic performance evaluation of clustering methods for single-cell RNA-seq data. F1000Res 7:1141. 10.12688/f1000research.15666.2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Freytag S, Tian L, Lonnstedt I, Ng M, Bahlo M (2018) Comparison of clustering tools in R for medium-sized 10x Genomics single-cell RNA-sequencing data. F1000Res 7:1297. 10.12688/f1000research.15809.2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L, Steemers FJ, Trapnell C, Shendure J (2019) The single-cell transcriptional landscape of mammalian organogenesis. Nature 566(7745):496–502. 10.1038/s41586-019-0969-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Traag VA, Waltman L, van Eck NJ (2019) From Louvain to Leiden: guaranteeing well-connected communities. Sci Rep 9(1):5233. 10.1038/s41598-01941695-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 113.Saelens W, Cannoodt R, Todorov H, Saeys Y(2019) A comparison of single-cell trajectory inference methods. Nat Biotechnol 37(5): 547–554. 10.1038/s41587-019-0071-9 [DOI] [PubMed] [Google Scholar]
  • 114.Street K, Risso D, Fletcher RB, Das D, Ngai J, Yosef N, Purdom E, Dudoit S (2018) Slingshot: cell lineage and pseudotime inference for single-cell transcriptomics. BMC Genomics 19(1):477. 10.1186/s12864-018-4772-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Wolf FA, Hamey FK, Plass M, Solana J, Dahlin JS, Gottgens B, Rajewsky N, Simon L, Theis FJ (2019) PAGA: graph abstraction reconciles clustering with trajectory inference through a topology preserving map of single cells. Genome Biol 20(1):59. 10.1186/s13059-019-1663-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.La Manno G, Soldatov R, Zeisel A, Braun E, Hochgerner H, Petukhov V, Lidschreiber K, Kastriti ME, Lonnerberg P, Furlan A, Fan J, Borm LE, Liu Z, van Bruggen D, Guo J, He X, Barker R, Sundstrom E, Castelo-Branco G, Cramer P, Adameyko I, Linnarsson S, Kharchenko PV (2018) RNA velocity of single cells. Nature 560(7719): 494–498. 10.1038/s41586-018-0414-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Bergen V, Lange M, Peidli S, Wolf FA, Theis FJ (2020) Generalizing RNA velocity to transient cell states through dynamical modeling. Nat Biotechnol. 10.1038/s41587-020-0591-3 [DOI] [PubMed] [Google Scholar]
  • 118.Qiu X, Zhang Y, Yang D, Hosseinzadeh S, Wang L, Yuan R, Xu S, Ma Y, Replogle J, Darmanis S, Xing J, Weissman JS (2019) Mapping vector field of single cells. bioRxiv:696724. 10.1101/696724 [DOI] [Google Scholar]
  • 119.Mereu E, Lafzi A, Moutinho C, Ziegenhain C, McCarthy DJ, Alvarez-Varela-A,Batlle E, Sagar GD, Lau JK, Boutet SC, Sanada C, Ooi A, Jones RC, Kaihara K, Brampton C, Talaga Y, Sasagawa Y, Tanaka K, Hayashi T, Braeuning C, Fischer C, Sauer S, Trefzer T, Conrad C, Adiconis X, Nguyen LT, Regev A, Levin JZ, Parekh S, Janjic A, Wange LE, Bagnoli JW, Enard W, Gut M, Sandberg R, Nikaido I, Gut I, Stegle O, Heyn H (2020) Benchmarking single-cell RNA-sequencing protocols for cell atlas projects. Nat Biotechnol 38(6): 747–755. 10.1038/s41587-020-0469-4 [DOI] [PubMed] [Google Scholar]
  • 120.Wen WX, Mead AJ, Thongjuea S (2020) Technological advances and computational approaches for alternative splicing analysis in single cells. Comput Struct Biotechnol J 18: 332–343. 10.1016/j.csbj.2020.01.009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Arzalluz-Luque A, Conesa A (2018) Single-cell RNAseq for the study of isoforms-how is that possible? Genome Biol 19(1):110. 10.1186/s13059-0181496-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Katz Y, Wang ET, Airoldi EM, Burge CB (2010) Analysis and design of RNA sequencing experiments for identifying isoform regulation. Nat Methods 7(12):1009–1015. 10.1038/nmeth.1528 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Huang Y, Sanguinetti G (2017) BRIE: transcriptome-wide splicing quantification in single-cells. Genome Biol 18(1):123. 10.1186/s13059-017-1248-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Song Y, Botvinnik OB, Lovci MT, Kakaradov B, Liu P, Xu JL, Yeo GW (2017) Single-cell alternative splicing analysis with expedition reveals splicing dynamics during neuron differentiation. Mol Cell 67(1): 148–161.e145. 10.1016/j.molcel.2017.06.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Welch JD, Hu Y, Prins JF (2016) Robust detection of alternative splicing in a population of single cells. Nucleic Acids Res 44(8): e73. 10.1093/nar/gkv1525 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Byrne A, Beaudin AE, Olsen HE, Jain M, Cole C, Palmer T, DuBois RM, Forsberg EC, Akeson M, Vollmers C (2017) Nanopore long-read RNAseq reveals widespread transcriptional variation among the surface receptors of individual B cells. Nat Commun 8: 16027. 10.1038/ncomms16027 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 127.Kim HJ, Lin Y, Geddes TA, Yang JYH, Yang P (2020) CiteFuse enables multi-modal analysis of CITE-seq data. Bioinformatics 36(14): 4137–4143. 10.1093/bioinformatics/btaa282 [DOI] [PubMed] [Google Scholar]
  • 128.Zhang Y, Liu T, Meyer CA, Eeckhoute J, Johnson DS, Bernstein BE, Nusbaum C, Myers RM, Brown M, Li W, Liu XS (2008) Model-based analysis of ChIP-Seq (MACS). Genome Biol 9(9):R137. 10.1186/gb-2008-9-9-r137 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Baker SM, Rogerson C, Hayes A, Sharrocks AD,Rattray M (2019) Classifying cells with Scasat, a single-cell ATAC-seq analysis tool. Nucleic Acids Res 47(2):e10. 10.1093/nar/gky950 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 130.Yu W, Uzun Y, Zhu Q, Chen C, Tan K (2020) scATAC-pro: a comprehensive workbench for single-cell chromatin accessibility sequencing data. Genome Biol 21(1):94. 10.1186/s13059-020-02008-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Danese A, Richter ML, Fischer DS, Theis FJ, Colomé-Tatché M (2019) EpiScanpy: integrated single-cell epigenomic analysis. bioRxiv:648097. 10.1101/648097 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Fang R, Preissl S, Li Y, Hou X, Lucero J, Wang X, Motamedi A, Shiau AK, Zhou X, Xie F, Mukamel EA, Zhang K, Zhang Y, Behrens MM, Ecker JR, Ren B (2020) SnapATAC: a comprehensive analysis package for single cell ATAC-seq. bioRxiv:615179. 10.1101/615179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Granja JM, Corces MR, Pierce SE, Bagdatli ST,Choudhry H, Chang HY, Greenleaf WJ (2020) ArchR: an integrative and scalable software package for single-cell chromatin accessibility analysis. bioRxiv:2020.2004.2028.066498 10.1101/2020.04.28.066498 [DOI] [Google Scholar]
  • 134.Stuart T, Srivastava A, Lareau C, Satija R (2020) Multimodal single-cell chromatin analysis with Signac. bioRxiv:2020.2011.2009.373613 10.1101/2020.11.09.373613 [DOI] [Google Scholar]
  • 135.Bravo Gonzalez-Blas C, Minnoye L, Papasokrati D, Aibar S, Hulselmans G, Christiaens V, Davie K, Wouters J, Aerts S (2019) cisTopic: cis-regulatory topic modeling on single-cell ATAC-seq data. Nat Methods 16(5):397–400. 10.1038/s41592-019-0367-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Chen H, Albergante L, Hsu JY, Lareau CA, Lo Bosco G, Guan J, Zhou S, Gorban AN, Bauer DE, Aryee MJ, Langenau DM, Zinovyev A, Buenrostro JD, Yuan GC, Pinello L (2019) Single-cell trajectories reconstruction, exploration and mapping of omics data with STREAM. Nat Commun 10(1): 1903. 10.1038/s41467-019-09670-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Schep AN, Wu B, Buenrostro JD, Greenleaf WJ (2017) chromVAR: inferring transcription-factor-associated accessibility from single-cell epigenomic data. Nat Methods 14(10):975–978. 10.1038/nmeth.4401 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Pliner HA, Packer JS, McFaline-Figueroa JL, Cusanovich DA, Daza RM, Aghamirzaie D, Srivatsan S, Qiu X, Jackson D, Minkina A, Adey AC, Steemers FJ, Shendure J, Trapnell C (2018) Cicero Predicts cis-Regulatory DNA interactions from single-cell chromatin accessibility data. Mol Cell 71(5):858–871. e858. 10.1016/j.molcel.2018.06.044 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Efremova M, Teichmann SA (2020) Computational methods for single-cell omics across modalities. Nat Methods 17(1):14–17. 10.1038/s41592-0190692-4 [DOI] [PubMed] [Google Scholar]
  • 140.van Dijk D, Sharma R, Nainys J, Yim K, Kathail P, Carr AJ, Burdziak C, Moon KR, Chaffer CL, Pattabiraman D, Bierie B, Mazutis L, Wolf G, Krishnaswamy S, Pe’er D (2018) Recovering gene interactions from single-cell data using data diffusion. Cell 174(3):716–729.e727. 10.1016/j.cell.2018.05.061 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Yang MQ, Weissman SM, Yang W, Zhang J, Canaann A, Guan R (2018) MISC: missing imputation for single-cell RNA sequencing data. BMC Syst Biol 12 (Suppl 7):114. 10.1186/s12918-0180638-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Li WV, Li JJ (2018) An accurate and robust imputation method scImpute for single-cell RNA-seq data. Nat Commun 9(1):997. 10.1038/s41467-018-03405-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Chen M, Zhou X (2018) VIPER: variability-preserving imputation for accurate gene expression recovery in single-cell RNA sequencing studies. Genome Biol 19(1):196. 10.1186/s13059-0181575-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Mongia A, Sengupta D, Majumdar A (2019) McImpute: matrix completion based imputation for single cell RNA-seq data. Front Genet 10:9. 10.3389/fgene.2019.00009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Qi Y, Guo Y, Jiao H, Shang X (2020) A flexible network-based imputing-and-fusing approach towards the identification of cell types from single-cell RNA-seq data. BMC Bioinformatics 21(1):240. 10.1186/s12859-020-03547-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Gunady MK, Kancherla J, Bravo HC, Feizi S (2019) scGAIN: single cell RNA-seq data imputation using generative adversarial networks. bioRxiv:837302. 10.1101/837302 [DOI] [Google Scholar]
  • 147.Talwar D, Mongia A, Sengupta D, Majumdar A(2018) AutoImpute: autoencoder based imputation of single-cell RNA-seq data. Sci Rep 8(1):16329. 10.1038/s41598-018-34688-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 148.Argelaguet R, Arnol D, Bredikhin D, Deloro Y, Velten B, Marioni JC, Stegle O (2020) MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol 21(1):111. 10.1186/s13059-020-02015-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 149.Welch JD, Hartemink AJ, Prins JF (2017) MATCHER: manifold alignment reveals correspondence between single cell transcriptome and epigenome dynamics. Genome Biol 18(1):138. 10.1186/s13059-017-1269-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 150.Wang X, Park J, Susztak K, Zhang NR, Li M (2019)Bulk tissue cell type deconvolution with multi-subject single-cell expression reference. Nat Commun 10(1):380. 10.1038/s41467-018-08023-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 151.Duan B, Zhou C, Zhu C, Yu Y, Li G, Zhang S, Zhang C, Ye X, Ma H, Qu S, Zhang Z, Wang P, Sun S, Liu Q (2019) Model-based understanding of single-cell CRISPR screening. Nat Commun 10(1): 2233. 10.1038/s41467019-10216-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 152.Campbell KR, Steif A, Laks E, Zahn H, Lai D, McPherson A, Farahani H, Kabeer F, O’Flanagan C, Biele J, Brimhall J, Wang B, Walters P, Consortium I, Bouchard-Cote A, Aparicio S, Shah SP (2019) clonealign: statistical integration of independent single-cell RNA and DNA sequencing data from human cancers. Genome Biol 20(1):54. 10.1186/s13059-0191645-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153.Ma A, McDermaid A, Xu J, Chang Y, Ma Q (2020) Integrative methods and practical challenges for single-cell multi-omics. Trends Biotechnol 38(9):1007–1022. 10.1016/j.tibtech.2020.02.013 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES