Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Nov 30.
Published in final edited form as: Trends Genet. 2024 Apr 24;40(7):587–600. doi: 10.1016/j.tig.2024.03.010

Discovering mechanisms of human genetic variation and controlling cell states at scale

Max Frenkel 1,2,3,*, Srivatsan Raman 3,4,5,*
PMCID: PMC11607914  NIHMSID: NIHMS2036276  PMID: 38658256

Abstract

Population-scale sequencing efforts have cataloged substantial genetic variation in humans such that variant discovery dramatically outpaces interpretation. We discuss how single-cell sequencing is poised to reveal genetic mechanisms at a rate that may soon approach that of variant discovery. The functional genomics toolkit is modular enough to systematically profile almost any type of variation within increasingly diverse contexts and with molecularly comprehensive and unbiased readouts. As a result, we can construct deep phenotypic atlases of variant effects spanning the entire regulatory cascade. The same conceptual approach to interpreting genetic variation should be applied to engineering therapeutic cell states. In this way, variant mechanism discovery and cell state engineering will become reciprocating and iterative processes towards genomic medicine.

Keywords: Functional genomics, single-cell sequencing, variant effect interpretation, cell-state engineering, genomic medicine

The daunting scale of human genetic variation

There is staggering diversity of naturally occurring DNA sequence variation among humans. Over the last two decades, the near million-fold decrease in cost to sequence a human genome has allowed researchers to sequence populations and cells from disease states at an unprecedented scale and pace. As a result, researchers have amassed hundreds of millions of genetic variants across increasingly diverse populations and conditions16. On the one hand, ever-increasing sequencing efforts (particularly of populations with diverse ancestries79) leads to the discovery of large effect-size rare variants that report on important biological processes. On the other hand, it necessarily increases the burden of variants with unclear biological and clinical significance.

A principal barrier to the scalable implementation of genomic medicine (see Glossary) is our relative inability to interpret arbitrary sets of genetic variants. Genome-wide association studies (GWAS) are useful for understanding the genetic architecture of complex traits and diseases. While the power to detect associations scales with sample size, so too does the discovery of rare variants for which interpretation is particularly difficult. Further, most heritability and disease-risk is widely distributed among hundreds or thousands of variants10,11 many of which may operate through distinct mechanisms despite being associated with the same trait or even occurring within the same gene. Complicating variant interpretation, most GWAS hits are non-coding and there is extensive linkage disequilibrium making it difficult to identify the causal gene or variant12,13. Even for protein-coding variation, variant interpretation can be difficult. This is evidenced by the fact that a plurality of missense variants in the human population are annotated as uncertain significance1416. Worse still, even for bona fide causal variants exerting certain pathogenicity, the molecular mechanisms are rarely completely understood which severely limits the promise of genomic medicine. Sequencing naturally occurring variation and performing association testing alone is therefore inadequate for interpreting the full extent of genetic diversity and for generating mechanistic insight into causal variants.

Recent advances in functional genomics methods stand to redefine how we interpret genomic variation and intervene on complex biological processes. Several excellent reviews exist describing high-throughput (but low-resolution) assays for variant effects15,16 and the more recent coupling of pooled CRISPR-based perturbations to high-resolution single-cell readouts for dissecting biological networks17. Here we emphasize the application of deep-phenotypic readouts at scale to a wider range of possible types of genetic variation. On the one hand, coupling single-cell sequencing to libraries of clinically relevant genetic variants will change the way in which variant interpretation is performed. The scale of pooled single-cell sequencing will allow the pace of variant interpretation to approach that of discovery. Simultaneously, the high-content nature of these methods is such that variant interpretation will occur based on the full mechanistic complexity by which each variant exerts its effects. When coupled with functional assays, we will be able to trace the mechanisms of genetic variants throughout their entire network of effects and learn the instructions for complex cell functions, phenotypes, and diseases (Figure 1a). On the other hand, the same methods will also be useful for intervening on disease processes. High-content measures of cell state will become targets for therapeutic cell engineering. These two applications of high-throughput deep phenotyping are synergistic. We propose that researchers apply this growing functional genomics toolkit to variant interpretation and cellular engineering as interdependent processes towards producing a reciprocating genomic medicine pipeline.

Figure 1: The complex network of variant effects and methods for their interpretation.

A diagram of a cell structure

(A) Genetic variants instruct cell behaviors and traits by differentially regulating complex genomic features like chromatin organization and gene expression. Unique genotypes can converge on similar higher-order genomic and cellular phenotypes. Individual genotypes may also regulate distinct cell states and behaviors depending upon the environment or temporal dynamics. Many of these relationships are non-intuitive without mapping the entire network. These mappings can be provided at scale by single-cell sequencing. (B) Most methods for interpreting genetic variation compete throughput (number of variants tested per experiment) with resolution (depth of genomic and cellular information retrieved for each variant). New-age functional genomics techniques are achieving both scale and resolution.

Shortcomings of current experimental methods for inferring variant effects

Improvements in both DNA reading and writing have resulted in an explosion of technologies for mapping genotype to high-dimensional genomic and cellular phenotypes. Key to most of these advances is the ability to resolve the phenotypes of single cells within heterogeneous pools of both genotypes and cell states. By uniquely tagging the contents of single cells (or single nuclei; “cells” is used to stand in for both throughout), these methods distinguish the activities of individual cells from the bulk population. Simultaneous measurement of genotypes and high-dimensional measures of cell state (e.g. accessible chromatin, gene expression etc.) allows for association testing between genotypes and complex phenotypes. The first major demarcation among these scalable methods for understanding variant effects is whether they capture naturally occurring genetic variation or engineer their own variant libraries.

Methods that leverage naturally occurring variation are attractive because the samples profiled are the least contrived and therefore most directly clinically relevant. In its most primitive form, this task is accomplished by detecting naturally occurring variation within expressed transcripts of single-cell libraries either with18 or without19 targeted amplification. This approach is limited to detecting highly expressed coding variants proximal to the transcripts’ end. Several methods of increasing complexity have been contrived to retrieve a wider range of genotypes and those from non-transcriptomic or multi-omic single-cell datasets2026. Despite these exciting improvements, there are at least three disadvantages to profiling naturally occurring variation. First, it is difficult to scale given that patient or sample selection is arrayed. Second, naturally occurring variation is non-exhaustive and therefore insufficient for comprehensively understanding sequence-function landscapes. It is also a necessarily reactive process whereas the availability of genomic medicine largely depends on our ability to prospectively define variant mechanisms and drug targets. Third, naturally derived samples (and the cells within them) usually differ by more than one variable. While this variation is attractive in that it represents essential to understand genomic realities, it also necessarily confounds causal inferences. Unambiguously mapping genotype to complex phenotype is difficult within the complex milieu of natural samples.

Reverse genetics experiments control for all but one genetic variable and clearly establish causal mechanisms. Unlike methods that detect naturally occurring variation within a complex background, reverse genetics experiments introduce one variant or perturb a single variable at a time thereby eliminating the effects of genetic and cellular confounders. For most such assays, there is a direct tradeoff between throughput and resolution (Figure 1b). Traditional reverse genetics techniques dissect complex phenotypes caused by genetic variants with exquisite mechanistic resolution (Figure 1b, left side); however, by virtue of being arrayed (the very feature that enables mechanistic resolution) these methods do not scale to the millions of genetic variants that demand interpretation. Conversely, pooled assays are now staples for testing the effects of upwards of ~106 variants per experiment2731 (Figure 1b, bottom right). The unifying feature of high-throughput experiments is that variant function is coupled to a change in either cellular abundance or that of a reporter that can be sequenced, sorted, or selected. Next-generation sequencing enumerates variant abundances from which thousands of variant effects are simultaneously inferred.

Reliance on sequencing counts as a proxy for variant effect severely limits most pooled approaches. For abundance differences to represent function, variants must directly alter cell proliferation27,32,33, regulate transcription at a single (supposedly representative) reporter28,3436, or a custom (i.e. not generalizable) assay must be designed to translate gene-specific mechanisms into differences that can be measured by sequencing35,37,38. In turn, these highly abstract readouts limit the resolution of information obtained for each variant. Proliferation, for example, is the cumulative effect of all molecular effects of each variant, which are impossible to disentangle with a proxy readout like sequencing counts. In cases where variants alter cellular processes without inducing fitness consequences, functionally relevant variants are entirely missed. Similarly, while a single reporter transcript might be sufficient for understanding variants that act exclusively in cis (e.g. 3’ untranslated region variants that alter transcript stability29,39), they are insufficient for understanding variants that exert their effects differentially across the genome. For instance, different mutations within the same transcription factor can cause different genome-wide binding profiles, chromatin states, or gene expression patterns that underlie clinically relevant phenotypes4042. Even some non-coding variants have been implicated in regulating high-dimensional measures of cell state (e.g. cell state quantitative trait loci, csQTLs)43, and many enhancers regulate transcription factors which in turn alter global gene expression patterns44. Cell state QTLs and the like will become more frequently encountered as researchers increasingly rely on high resolution methods to sample diverse biological systems. Further, pathogenic variants can cause diseases or alter complex traits by a diverse range of molecular mechanisms including dominant negative and highly variable neomorphic molecular functions45. A human genetics instantiation of the Anna Karenina principle: all benign variants are alike; each pathogenic variant is function-altering in its own way. The diverse mechanisms of pathogenic (or otherwise clinically relevant) variants are hidden from single gene reporter assays, the output of which can belie genome-wide and mechanistic differences.

Deep phenotyping of human genetic variation with single-cell sequencing

Single-cell sequencing is a powerful set of tools for marrying the throughput of pooled screens with the mechanistic resolution of traditional genomics methods. Unlike reporter assays, single-cell sequencing generates information-dense, high-dimensional measures of cell state. Genomic fingerprints—consisting of the quantitative expression of tens of thousands of transcripts46 or tens of thousands of transposase-accessible loci47,48—can be constructed for hundreds of thousands of cells simultaneously. Just as single-cell sequencing can be co-opted to characterize naturally occurring variation, so too can it be used to profile libraries of user-defined perturbations (Figure 1b, top right). In Perturb-seq (and its variations), CRISPR/Cas9 guide RNAs (gRNAs), or barcode proxies thereof, are captured and tagged alongside each cell’s mRNAs; this results in rich classifications of genetic perturbations (i.e. gRNAs) based on their resulting global transcriptomes and networks4952. Similarly, gRNAs can be captured alongside single-cell profiles of chromatin accessibility (as in Spear-ATAC) to learn mechanisms of epigenetic regulation53,54. CRISPR-based gene disruption was a natural starting point for this technology for several reasons. Chief among these is that gRNAs are amenable to high-throughput oligonucleotide synthesis, easily introduced to cells in a pooled fashion, and are readily recovered with few modifications to most single-cell sequencing workflows. Leveraging these benefits, libraries with thousands of gRNAs can be routinely synthesized for custom Perturb-seq experiments, and indeed the first genome-scale Perturb-seq experiment was recently performed characterizing the transcriptional programs altered by more than 12,000 gRNAs in one experiment55. While a powerful approach, most Perturb-seq experiments are limited by their inability to determine which genetic variant is present in each cell; gRNAs are instead used as proxies. Particularly in base editing screens, many gRNAs can introduce multiple variants which can have different transcriptional effects, and this heterogeneity is computationally averaged over as cells are grouped by their gRNAs and not their variants. Further, CRISPR-based assays are limited in the repertoire of mutations they can install either by the distribution of PAM sites or by the base editor chemistries allowed. Expanding the conceptual approach of Perturb-seq beyond CRISPR-mediated gene disruption (e.g. to comprehensively mutagenic libraries of well-defined variants) stands to redefine how researchers approach the discovery of disease mechanisms, genetic modifiers, and cell-state engineering.

Before single-cell sequencing, high-throughput variant effect characterization required laborious, customized experimental schema that depended upon a priori knowledge of how a disease gene or set of variants is likely to function. From this knowledge, custom assays of gene-specific function can be constructed to ascertain variant effects. The change from bespoke assays to high-throughput, generalizable variant effect platforms (like single-cell sequencing) is illustrated best by the example case of human Ras variants. Prior to single-cell platforms, researchers analyzed all possible single missense variants in human Ras with a custom assay that linked Ras function to bacterial cell survival56. Knowing that Ras signaling depends on Ras-Raf binding, Bandaru et al. engineered a two-hybrid system wherein functional Ras variants bind Raf, transcribe an antibiotic resistance gene, and become enriched in the setting of the appropriate antibiotic selection. The assay is clever but is inextricably linked to Ras biology. It requires the inducible expression of a Ras library covalently attached to a subunit of E. coli RNA polymerase, the Raf DNA binding domain covalently attached to the λ repressor, the expression of two cofactors for Ras nucleotide exchange, and the presence of an antibiotic resistance reporter under control of the λ operator. This assay could not have existed without extensive prior Ras-specific research, and the assay cannot be quickly ported to profile any protein other than Ras. What happens when the mechanism of a putative disease gene or causal variant is unknown? What happens when our understanding of a candidate protein’s function is incomplete, context-dependent, or cannot be recapitulated in a selectable or sequencing-based assay? Could we have arrived at the same conclusions without the laborious task of creating a Ras-specific assay? It would be useful to have general platforms that are equally capable of revealing the effects of any disease-relevant gene or process and require little to no upfront knowledge or customization.

Single-cell sequencing provides a generalizable platform for deeply phenotyping human genetic variation without depending on prior knowledge. Because single-cell sequencing retrieves rich cellular information without requiring variant or gene-specific customization, it poses few constraints on the types of libraries that can be studied (Figure 2, left box). Ursu et al. proved this point by using single-cell RNA sequencing (scRNA-seq) to determine how hundreds of missense variants in KRAS and TP53 alter global transcriptional regulation57. Whereas profiling H-Ras missense variants previously required a laborious, customized bacterial two-hybrid assay that depended on extensive prior understanding of Ras’ mechanism56, Ursu et al. discovered continuums of dysfunction induced by different KRAS mutations without a priori specification of the mechanisms involved. Importantly, single gene reporter assays would have never revealed global transcriptional dysregulation, heterogeneity of effects caused by a single variant (i.e. true single-cell resolution and not pseudobulk effects when variants of a particular identity are aggregated), or continuums of function all of which were revealed by scRNA-seq. Further, that this work simultaneously reported on variant effects within KRAS and TP53—two unrelated proteins with different mechanisms—demonstrated scRNA-seq’s generality. A logical extension is that multi-omic single-cell sequencing can be used to generate comprehensive atlases of missense variant effects across multiple levels of cellular regulation.

Figure 2: A modular functional genomics toolkit for understanding variant mechanisms and controlling cell states at scale.

A diagram of a puzzle with text

Diverse libraries of genetic and environmental perturbations (top left) can be profiled in heterogeneous contexts (top middle) and with diverse high-content readouts (top right). The combinations across this parameter space will yield information-dense insight into variant mechanisms and the epistatic interactions between genetic variants (bottom). From this, we can construct rich genotype-phenotype maps of variant effects. This toolkit also serves as a foundation for therapy discovery based on controlling high-dimensional measures of cell state by either defining perturbations that revert cells towards healthy states or by finding drugs that target specific variant mechanisms.

Encouraged by this success, there is now a growing toolkit of multi-omic, scalable methods for understanding variant effects. While scRNA-seq was the first and still most frequently used single-cell readout for mapping protein-coding variants to complex cell-state dysregulation22,57,58, additional methods are following suit. Single-cell epigenomics methods (e.g. scATAC-seq and scCUT&Tag) lag behind scRNA-seq in their applications to missense variant interpretation because of the increased difficulty of retrieving each cell’s perturbation. Gene expression is a natural signal amplification process that epigenomic methods cannot leverage. To address this and to map CRISPR perturbations to genome-wide chromatin accessibility, Spear-ATAC was designed to specifically amplify gRNAs from genomic DNA during scATAC-seq library preparation53. Few further modifications generalized Spear-ATAC to arbitrarily sized protein-coding perturbations59. This generalization, PROD-ATAC, was used to understand how arbitrary sets of protein-coding variants alter genome-wide chromatin accessibility; however, there is no theoretical barrier to PROD-ATAC’s future coupling to complementary high-dimensional measures of epigenomic cell state like scCUT&Tag60. This new combination (single-cell perturb CUT&Tag) would yield genome-wide histone modification tracks for thousands of variants simultaneously. Integrating multiplexed histone modifications with accessibility landscapes would allow researchers to understand how each genetic variant or perturbation within the library disrupts regulatory logic in a relatively unbiased way. Even more complete mechanistic understandings of variant effects could be achieved with the coupling of perturbations to single-cell Hi-C (a readout for 3-dimensional chromatin contacts)6163 and methods that simultaneously capture protein expression information too64,65. Single-cell readouts spanning large swaths of the central dogma and cellular phenotypes either exist or are imminently on the horizon (Figure 2, right most box). This impressive repertoire of techniques will allow researchers to reconstruct comprehensive regulatory networks perturbed by genetic variation (Box 1).

Box 1: A case study in genetic mechanism discovery at scale.

There are several recurrently mutated proteins for which we have limited understanding of variant effects. For instance, there are thousands of mutations within the mammalian SWI/SNF (mSWI/SNF) chromatin remodeling complex that cause significant molecular and phenotypic heterogeneity. Mutations catalogued include deletions, truncations, missense variants, and fusions underlying heterogeneous syndromes from cancer to neurodevelopmental disorders86,87. A microcosm of human genetic disorders, many of these variants are of uncertain significance. Those with mechanistic insight imply daunting diversity given that mSWI/SNF differentially regulates tens of thousands of sites across the genome, mutations can have both loss and gain of function, and the phenotypes caused by mSWI/SNF variants are variable with complex inheritance patterns. Applying the full repertoire of single-cell sequencing techniques to libraries with thousands of mSWI/SNF variants would reveal the diverse ways in which these mutations alter histone modifications, chromatin accessibility, and in turn disease-causing gene expression programs (Figure 2). This would produce troves of data for thousands of variants and their impacts across the central dogma. While rich in information content, it is far from trivial to use these data to establish the underlying regulatory networks. Several reviews have discussed integrating multimodal single-cell sequencing datasets to this end88,89. Not only have these approaches been useful for defining regulatory networks that underly endogenous processes like differentiation and development, but they have already proved useful for learning the mechanisms of genetic variants90,91. Applying network approaches to datasets collected from deep mutational scans (e.g. all possible single mSWI/SNF variants) will allow researchers to generate rich sequence-function landscapes from which we could learn the sequence determinants of genomic regulation. Network models might also allow us to learn non-intuitive convergent mechanisms of genomic dysregulation between seemingly distinct mutations (e.g. mutations in distal parts of a protein or mutations across different proteins causing similar networks of dysfunction). In cases where the genes profiled cause disease (e.g. mSWI/SNF), these networks might represent useful mechanistic classifications, and perturbed networks may point to therapeutic targets. Finally, this data will provide the ideal basis on which researchers can train machine learning models for variant effect prediction. Existing pathogenicity prediction tools are largely based on evolutionary sequence conservation and limited biochemical features9294. Models based on multimodal sequencing data will likely prove useful for predicting the mechanisms of untested genetic variation. In this way, single-cell sequencing provides the basis for discovering genetic mechanisms at scale which stands to redefine how clinical variation is interpreted, classified, and intervened upon.

The growing single-cell arsenal (and their multi-omic combinations) is well-positioned for mapping the impacts of protein-coding variants at population scale. Extrapolating from existing protein-coding variant effect sizes (which dictates the throughput and sequencing depth needed to characterize variant effects)57,59, it is reasonable to imagine that multimodal single-cell sequencing could be used to profile all possible single missense variants in clinically actionable genes within a few years’ time. This will provide a first pass draft of a high-dimensional variant effect atlas. The current variant classification system (e.g. that proposed by ACMG66 and used by ClinVar) is conceptually tied to a single-plexed conceptualization of variant effects (i.e. a one-dimensional scale from benign to pathogenic). A new single-cell sequencing derived atlas should accommodate many axes of genomic variant effects as they all report on important biological processes and possible therapeutic targets.

Systematically dissecting context-dependent variant effects

Future drafts of a high-dimensional variant effect atlas should include even greater complexity, the most important of which is likely cellular context. Existing single-cell sequencing studies of missense variation were performed in easy to engineer cell lines (293T22,59 and A549 cells57) which are poor models of context-specific disease mechanisms. On the one hand, these studies have shown surprising recovery of disease biology across diverse mechanisms and diseases. To this end, we believe that profiling large libraries of variants in genetically tractable (if not disease relevant) contexts will be informative particularly for classifying variant effects in a common context. That is partly because we believe that the overall classifications (i.e. clusters) of variants (when based on high-dimensional genomic measurements) will likely be more robust to the context than individual disease-gene targets. On the other hand, it is well established that certain variant effects can be exquisitely cell-type and context-dependent. Systematically exploring the effects of genetic variation conditional on genetic and cellular context is difficult given the combinatorial explosive number of possible pairings. However, studying context dependence has been possible by porting single-gene reporter experiments across a small set of cell lines29. This will be a useful approach for future single-cell sequencing too (Figure 2, middle box). Even if some variant effects are not robust to context switching (and therefore are difficult to interpret clinically), learning that certain effects are (or are not) context-sensitive is itself a meaningful measure of variant effect.

Performing single-cell variant effect studies in genetically tractable cell lines also presents the opportunity to systematically dissect epistasis and context specificity. Just as sequencing naturally occurring samples can make causal inference difficult, so too will it be difficult to disentangle background effects when porting assays between cell lines. Instead, researchers should make well-controlled background perturbations—in the presence of libraries of causal variants—to specifically test for epistasis and the presence of genetic modifiers. To this end, libraries of protein-coding variants could be combined with libraries of background perturbations to unambiguously ascertain context-dependent variant effects and their mechanisms. Such combinations of libraries have been exploited to discover allosteric networks within proteins and learn the rules of protein-protein interactions6769. In these cases, low-resolution selection-based schemes were used to infer variant effects. It is likely that different allosteric networks and epistatic interactions can alter the targeting of global genome regulators differently across the genome. If true, then just as single gene reporter assays belie genomic complexity for understanding the effects of individual mutations, so too would it be necessary to use high-dimensional multi-omic methods to study allostery and epistasis. This has been attempted once for combinations of gRNAs with scRNA-seq70, but has yet to be applied to protein-coding perturbations. Libraries of missense variants ought to be combined with genome-scale CRISPR perturbations or combined with libraries of variants in downstream effector proteins and binding partners. Missense libraries could also be coupled with CRISPR perturbations specifically targeting druggable proteins. In this case, those background CRISPR perturbations that mitigate the influences of causal variants would be immediately suggestive of precision (i.e. genotype informed) therapies. We anticipate this conceptual approach to epistasis and allostery will become a staple for learning context dependency of mutational effects and identifying genetic modifiers of cell state.

Finally, there are many more types of genetic variation that ought to be profiled with deep phenotypic readouts. Structural variation has been difficult to engineer and specifically perturb. While writing this article, the first instance of a pooled library of structural variants was generated and their effects were profiled with scRNA-seq71. A CRISPR-based method has recently been created for engineering aneuploidy72, which if multiplexed could readily be coupled with Perturb-seq to understand the global effects of diverse sets of aneuploidies. From such data, we could begin to dissect the rules of dosage compensation, identify instances where aneuploidy might be causal (versus the consequence) of diseases, and learn non-intuitive instances of trans regulation of gene expression. Similarly, there are several outstanding questions regarding the length and content dependence of tandem repeats and their cell-type specific effects on gene expression73,74. Single-cell sequencing of libraries with user-defined tandem repeat profiles (i.e. systematically altering repeat length and motif content) across multiple genes and within many cell contexts will help discover the rules by which they alter gene expression and cause disease. Once again, the near unbounded diversity of variation that can be assayed is a function of the generality of single-cell sequencing readouts.

Taken together, single-cell sequencing not only provides a high-resolution measurement for variant interpretation—spanning the entire regulatory cascade—but it is rapidly scalable to both the number of clinical variants demanding interpretation and the myriad contexts that they differentially execute function within. Modular combinations of diverse library types, cellular contexts, and multi-omic phenotypic readouts (Figure 2, top row) will produce rich, mechanistic variant effect atlases (Figure 2, bottom).

A new target for cell-state engineering and genomic medicine

The most common use case of single-cell sequencing is the measurement of cell state diversity in naturally occurring tissues. From this, we’ve learned of profoundly heterogeneous cell types and regulatory mechanisms within health and disease7578. A natural corollary to these cataloging efforts is that we have simultaneously discovered high-dimensional targets for cell-state engineering and new goals for genomic medicine. For centuries (and with good reason) biochemical and pharmacologic research has focused on producing therapies that narrowly target single genes, individual biomolecular interactions, and single biochemical steps. Just as single-reporter or selection-based assays belie genome-wide complexity, so too might targeting individual pathways or interactions miss the opportunity to more widely control and tune cell state.

An exciting application for high-throughput perturbational experiments is the control of genome-wide regulatory logic and cell state. While certain diseases mechanisms are caused by dysregulation of one pathway, there are many biological processes the instructions for which are determined by global gene regulatory patterns (i.e. dozens to hundreds of genes working in concert). Similarly, the fact that most traits and diseases are profoundly polygenic implies that controlling multiple genes and pathways at once will be therapeutically valuable. As a result, an earnest implementation of genomic medicine entails control not just of individual pathways but instead of entire genetic networks. Perhaps the clearest example of this is the regulation of stem-like states via induced pluripotency or the directed differentiation of specific cell types therefrom. The discovery of four transcription factors that can revert fibroblasts into a pluripotent state was a striking feat of cell engineering79. This relied on laborious low-throughput experiments to characterize a small fraction of the possibility space of transcription factor combinations. Since then, high-throughput screens have been used to test the effects of thousands of putative transcription factors on differentiation80. Like most high-throughput assays, this depended on the selection of a single pluripotency marker which is an imperfect measure of cell fate. Similarly, others have created libraries of variants within a single transcription factor and determined their effects on cell fate, but this was performed in an arrayed fashion and was therefore was not comprehensive81. Joung et al. demonstrated the value of a high-throughput, high-content readout for studying cell fate control by using scRNA-seq to profile the effects of thousands of transcription factors82. In just one experiment they discovered transcription factors that shift cell state distributions of embryonic stem cells towards multiple distinct differentiated cell types. Future multi-omic work should seek to understand how libraries of transcription factors alter cell state regulation and should strive to comprehensively map combinations of transcription factors as well as mutational scans within each. This approach equally applies to synthetic transcription factors and genetic circuits. Most genetic circuits are designed to regulate one or a few target genes and many are orthogonal to endogenous gene regulation83. Single-cell multiomic sequencing sets the stage for researchers to design and test genetic circuits and synthetic transcription factors that dynamically interact with endogenous pathways. Most obviously this will be useful for directed differentiation (e.g. for differentiation therapy targeted at cancer stem cells or for cell engineering) and for controlling engineered immune cell functions (e.g. designer CAR T-cells with circuits for executing complex logic towards the recognition of malignant cells and subsequent activation of therapeutic programs). However, this conceptual approach will become broadly useful for biasing cells towards healthy or therapeutic states. In this way, deep phenotyping of arbitrary variant types at scale will encourage researchers to control increasingly complex genomic networks: a new paradigm for genomic medicine.

Perhaps most interestingly is that high-content, scalable functional genomics assays allow the combination of variant effect interpretation with cell state control. This should become a foundational approach for ushering in true genotype-specific precision medicine. We previously argued that multimodal single-cell sequencing can map the mechanisms of complex cell state changes induced by genetic variation. In cases where this models disease-causing variation, it therefore also represents a platform with which to test therapies. Most obviously, libraries of genetic variants ought to be combined with drugs to reveal variant-specific drug mechanisms that can revert cells towards healthy states. This concept has been partially proved as hundreds of drug effects on cancer cell lines have been profiled with scRNA-seq84. This has been extended to map the interactions between >14,000 combinations of CRISPR and chemical perturbations85. It will be fascinating to combine libraries of many more types of genetic variation with libraries of drugs to determine how each pairing affects global genome regulation (Figure 2, bottom). Whereas high-throughput screens that rely on one reporter could easily miss off-target effects or miss compensatory mechanisms that underly therapy escape, single-cell sequencing will reveal the global impacts of drug-by-variant effects. We hypothesize that high-dimensional measures of cell state will become one of the most robust readouts for drug discovery and one of the most predictive for both efficacy and off-target effects. Disease cells are likely more able to evade therapies directed at modulating one pathway than they are new-age therapies that modulate entire regulatory networks.

Conclusions

The confluence of recent technological advances in DNA writing and reading have set the stage for scalable genomic medicine. On the one hand, combining high-throughput genetic manipulation (i.e. writing) with single-cell readouts (i.e. reading) will allow researchers to create high-content genotype to phenotype atlases (see Outstanding Questions). This combination represents a scalable mechanism discovery pipeline that can approach the pace of population-scale variant discovery. From this effort, we anticipate the discovery of (i) mechanisms of causal variants across the entire regulatory cascade (ii) mechanisms even for exceptionally rare or patient-specific variants which traditional reverse genetic experiments largely ignore out of necessity (iii) non-intuitive convergent mechanisms across seemingly disparate variants which will inform mechanistic variant classifications and (iv) genetic modifiers of causal variants which might explain context specificity, variable penetrance, and even imply possible therapeutic targets to restore healthy cell states. Unlike traditional methods which inject bias when researchers chose which variants to study, these efforts will be significantly less biased. The comprehensive and less biased nature of the libraries studied with single-cell sequencing will likely prove useful in the construction of machine learning models for predicting variant effects. That is, studying larger and more diverse samples of sequence space will allow models to learn increasingly complex features for predicting variant effects and improve the prediction of genotype-informed therapies. Of course, form does not necessarily imply function. Coupling single-cell sequencing to measurements of high-level cell functions (e.g. proliferation, migration, immune cell activation and targeting etc.) will allow researchers to determine the genomic instructions of complex cellular behaviors. On the other hand, the same high-content phenotyping is being rapidly deployed to profile naturally occurring genetic variation and cell state heterogeneity in healthy and disease tissues. The resulting atlases represent high-dimensional cell state targets for future cell engineering. In lieu of targeting individual pathways or enzymatic reactions, researchers can target high-dimensional measures of cell state which we hypothesize will result in more efficacious therapies. For instance, a therapy which redirects a cancer cell towards a healthy cell state (as measured by thousands of dimensions) will likely be less prone to resistance than one that targets a single pathway. These two conceptual approaches (1. perturbing cell states with high-throughput reverse genetics and 2. discovering naturally occurring cell states) ought to become part of a reciprocating process. Modeling causal variant effects in cell culture and manipulating them towards high-dimensional therapeutic end goals is now in sight. Iterating through these interdependent processes will allow us to make more sophisticated versions of genotype-informed therapies and clinical choices.

Outstanding questions.

  • How do arbitrary sets of genetic variants alter genome-wide regulatory logic? Can we create high-content maps (based on chromatin conformation, histone modifications, gene and protein expression) for all possible single mutations in all disease-causing genes? Will there be convergent mechanisms across seemingly disparate mutations?

  • What is the best way to incorporate high-dimensional, mechanistic insight derived from single-cell sequencing into population-scale variant effect atlases? Can existing paradigms of variant interpretation (e.g. benign vs pathogenic) accommodate thousands of axes of variation that change over space, context, and time?

  • To what extent are causal variant mechanisms context dependent? Is there a minimal set of cellular contexts or set of high-throughput epistasis experiments from which we can extrapolate to unseen variants in new contexts?

  • Do high-dimensional measures of cell state serve as more robust targets for therapeutic cell engineering than do measures of individual protein functions or pathways?

Highlights.

  • There is remarkable genetic diversity throughout the human population, but variant interpretation lags far behind discovery. As a result, most genetic variation is not clinically actionable.

  • Single-cell sequencing provides a high-throughput, disease- and pathway-agnostic means for profiling variant effects with mechanistic detail across the entire regulatory cascade.

  • Many single-cell genomics methods have been coupled to CRISPR-based gene disruption, but these techniques are rapidly expanding to accommodate increasingly diverse types of perturbations including missense variants, protein fusions, structural variants, genetic circuits, chemical perturbations, and increasingly their combinations.

  • High-throughput, deep phenotyping of genetic variation will allow researchers to define causal mechanisms and nominate precision therapeutics for controlling cell states.

Acknowledgments

We would like to thank Silas Miller and Margaux Hujoel for their reviews of the manuscript. M.F. was supported by an NIH T32 fellowship (T32HG2760-17).

Glossary

Barcode proxies

short sequences of DNA that are associated with perturbations within a pooled library. Sequencing the barcode in each cell or nuclei therefore serves as a proxy for which perturbation was present.

Base editing

base editing enzymes attached to enzymatically dead Cas9 can be directed to make substitutions at specific sites by guide RNAs. Cytosine base editors convert C•G to T•A whereas adenine base editors convert A•T to G•C. When used in single-cell sequencing studies, base editors suffer from two major demerits. First, the sites that can be mutated are limited by gRNA accessibility (via PAM dependence) which is not universal throughout the genome. Second, even at targetable loci, not all possible mutations are achievable. Lastly, while not inherent to base editors per se, these experiments frequently infer mutations based on gRNAs rather than reading out the mutations directly, which results in pseudobulk samples that average over likely important heterogeneity.

Cell state

cells adopt different states across space and time which are defined by differential execution of genetically encoded programs. The spatiotemporal execution of these programs is regulated mostly at the level of chromatin organization, changes in gene expression, and changes in protein expression and modifications. Cell states can also refer to phenotypic behaviors of cells (e.g. pluripotency versus terminal differentiation; immune cell activation versus exhaustion etc.). Here we generally mean the former as these are representations of states that can be readout in high-throughput with single-cell sequencing. Phenotypic readouts of cell state are often low throughput.

Chromatin accessibility

the availability of DNA wrapped around histones to effector proteins. This is a dynamic feature of chromatin, accessibility is correlated with transcriptional and regulatory activity, and the patterns of accessibility across the genome are a measure of a particular cell type’s state.

CRISPR-based gene disruption

guide-RNA directed CRISPR disruption of an endogenous gene. This can take the form of transcriptional disruption (i.e. activation or repression) or gene editing (e.g. random errors or homology directed recombination of a user-defined template).

Deep mutational scan

experiments wherein every amino acid within a protein is mutated to every other possible amino acid. These are useful for (i) generating prospective atlases of variant effect within disease-causing genes and (ii) for understanding the sequence determinants of the protein’s function in a relatively unbiased way.

Deep phenotyping

single-cell genomics like scRNA-seq and scATAC-seq provide quantitative measures of genomic cell state across thousands to tens of thousands of dimensions for each cell. Compared to single-gene reporter assays, these provide a deeper phenotypic understanding of cellular processes.

Dominant negative

some mutant proteins can inhibit the function of the wild-type copy. This most often occurs in multimeric complexes where the mutant protein binds to and “poisons” the activity of the wild-type protein within the complex.

Functional genomics

the study of functional elements of the genome and the effects of their alterations.

Genomic medicine

the use of genomic information to guide clinical decision-making such as classification of diseases, prediction of disease risk, and prediction of therapy response or adverse events. Here we use genomic medicine as an umbrella term to also encompass precision medicine which seeks to make clinically actionable decisions based on genomic information even for N-of-1 or rare causal variants or diseases. The typical case study of precision medicine is the tyrosine kinase inhibitor imatinib for chronic myelogenous leukemia. Here we use the term more broadly to indicate precise control of cell states and fates (rather than just control over individual pathways and enzymatic reactions).

Guide RNA (gRNA)

a short single stranded RNA molecule that directs Cas9 based on complementary base pairing to genetically alter or transcriptionally regulate a specific locus. gRNAs are frequently used in pooled screens to generate variants for high-throughput deep phenotyping. In traditional gRNA-based single-cell sequencing assays, cell perturbations are defined based on which gRNA(s) the cell possesses. In cases where the gRNAs introduce genetic changes (e.g. by base editors or by error-prone nonhomologous end joining), pseudobulk grouping by gRNA identity averages over heterogeneity that could be functionally important. Namely, different genetic variants introduced by the same gRNA may cause different genomic profiles which is lost in many of these assays.

GWAS

genome-wide association studies test the association between naturally occurring genetic variants and phenotypes at population scale.

Missense variant

a genetic change in a coding region that results in an amino acid change. Some missense variants have no impact on the protein’s function (i.e. benign) whereas others can exert a wide range of influences from loss of function to gain of function or dominant negative effects.

Neomorphic

a description of the effect a mutation has on a protein’s function. While most protein-coding mutations influence the protein’s functions native to the wild-type copy (e.g. loss of function mutations), others can create de novo functions (i.e. neomorphic mutations). Neomorphic functions can arise from extremely diverse mechanisms which are often hidden from single gene reporter assays.

Perturb-seq

this method relies on single-cell RNA sequencing to reveal the gene expression patterns that are altered by genetic perturbations (and usually CRISPR-mediated gene disruption). Usually, CRISPR/Cas9 guide RNAs are introduced to a cell line in a pooled fashion. The gene expression programs and the guides present in individual cells are retrieved such that high-dimensional maps between perturbation and transcriptional programs can be learned.

Pooled assays

unlike arrayed assays where cell lines with different genotypes or perturbations are separated, these assays combine multiple (often thousands of) genotypes or perturbations into a single pool. These pools are often referred to as libraries.

PROD-ATAC

protein-coding single-cell ATAC sequencing. This method modified Spear-ATAC (itself a version of scATAC-seq) to perturb cells with arbitrary lengthy libraries of protein-coding variants

Pseudobulk

single-cell sequencing retrieves complex genomic information for individual cells. If every cell were treated as its own experiment, it would be nearly impossible to learn anything of value (i) because the number of hypothesis tests would be so large—for experiments of typical throughput—that the barrier to meet statistical significance would be astronomical after multiple testing correction (ii) because the information content for individual cells is usually sparse (e.g. only a small fraction of the transcriptome from each cell is sampled with scRNA-seq). One way to combat this is to computationally group similar cells into pseudobulk samples such that the data is less sparse and fewer hypotheses require testing. Pseudobulk groupings can be made based on which perturbation or genetic variant each cell contains. Importantly, grouping based on ground truth labels (e.g. perturbation identity) does not suffer from statistical double dipping in the way that grouping based on the signals themselves does (e.g. as in typical single-cell sequencing of naturally occurring cell populations without known labels).

Reporter

the basis for most high-throughput variant effect assays. A transcript that either encodes for a selectable phenotype (e.g. a fluorophore or antibiotic resistance gene) or a sequenceable transcript (e.g. contains the variant studied or a proxy thereof) to be used for variant interpretation within a pooled assay.

scATAC-seq

single-cell assay for transposase accessibility and sequencing. This method maps chromatin accessibility across the genome at single-cell resolution. Tn5 transposase preferentially inserts DNA sequences in accessible chromatin. Here, Tn5 is loaded with adapters amenable for next-generation sequencing so that sequencing counts are a proxy for chromatin accessibility.

scCUT&Tag

single-cell cleavage under targets and tagmentation. These methods distinguish genome-wide deposition of histone modifications or protein-nucleic acids at single-cell resolution. Like ATAC, RNA etc. this yields a high dimensional representation of cell state and when combined with other modalities can create mechanistic understanding of regulatory networks.

scHi-C

single-cell high-throughput chromatin conformation capture. This reveals the three-dimensional architecture of chromatin within single-cells by sequencing regions of the genome in contact with each other.

scRNA-seq

single-cell RNA sequencing. This method quantifies the transcription of all coding genes within the genome at single-cell resolution. These gene expression programs are representative of cell state and can reveal mechanisms of disease.

Single-cell sequencing

genomic sequencing methods that can distinguish individual cell states within heterogeneous pools of cells (or nuclei). These methods often rely on physically compartmentalizing cells in droplets and uniquely tagging its molecular contents (e.g. transcripts, accessible chromatin etc.) with a unique molecular barcode. Sequencing the barcodes allows computational demultiplexing such that molecular contents can be grouped based on the cells from which they came.

Spear-ATAC

single-cell perturbations with an accessibility read-out using scATAC-seq. Cells are treated with CRISPR-based gRNA perturbations as in Perturb-seq. gRNAs are flanked by sequencing adapters and a modified version of 10X Genomics single-cell ATAC sequencing is used to increase the likelihood of capturing each nuclei’s encoded gRNA perturbation alongside chromatin accessibility information.

Footnotes

Declaration of interests

The author declares no conflicts of interest.

References

  • 1.Tam V et al. Benefits and limitations of genome-wide association studies. Nat. Rev. Genet 20, 467–484 (2019). [DOI] [PubMed] [Google Scholar]
  • 2.Auton A et al. A global reference for human genetic variation. Nature 526, 68–74 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Hawkes G et al. Whole genome association testing in 333,100 individuals across three biobanks identifies rare non-coding single variant and genomic aggregate associations with height. 2023.11.19.566520 Preprint at 10.1101/2023.11.19.566520 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Karczewski KJ et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Backman JD et al. Exome sequencing and analysis of 454,787 UK Biobank participants. Nature 1–10 (2021) doi: 10.1038/s41586-021-04103-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Aaltonen LA et al. Pan-cancer analysis of whole genomes. Nature 578, 82–93 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Graham SE et al. The power of genetic diversity in genome-wide association studies of lipids. Nature 600, 675–679 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Peterson RE et al. Genome-wide association studies in ancestrally diverse populations: opportunities, methods, pitfalls, and recommendations. Cell 179, 589–603 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Mahajan A et al. Multi-ancestry genetic study of type 2 diabetes highlights the power of diverse populations for discovery and translation. Nat. Genet 54, 560–572 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Sinnott-Armstrong N, Naqvi S, Rivas M & Pritchard JK GWAS of three molecular traits highlights core genes and pathways alongside a highly polygenic background. eLife 10, e58615 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Boyle EA, Li YI & Pritchard JK An expanded view of complex traits: from polygenic to omnigenic. Cell 169, 1177–1186 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Cano-Gamez E & Trynka G From GWAS to Function: Using Functional Genomics to Identify the Mechanisms Underlying Complex Diseases. Front. Genet 11, (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Ward LD & Kellis M Interpreting noncoding genetic variation in complex traits and human disease. Nat. Biotechnol 30, 1095–1106 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Fowler DM & Rehm HL Will variants of uncertain significance still exist in 2030? Am. J. Hum. Genet 111, 5–10 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Findlay GM Linking genome variants to disease: scalable approaches to test the functional impact of human mutations. Hum. Mol. Genet 30, R187–R197 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Tabet D, Parikh V, Mali P, Roth FP & Claussnitzer M Scalable Functional Assays for the Interpretation of Human Genetic Variation. Annu. Rev. Genet 56, 441–465 (2022). [DOI] [PubMed] [Google Scholar]
  • 17.Morris JA, Sun JS & Sanjana NE Next-generation forward genetic screens: uniting high-throughput perturbations with single-cell analysis. Trends Genet. TIG S0168–9525(23)00240–8 (2023) doi: 10.1016/j.tig.2023.10.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Nam AS et al. Somatic mutations and cell identity linked by Genotyping of Transcriptomes. Nature 571, 355–360 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Poirion O, Zhu X, Ching T & Garmire LX Using single nucleotide variations in single-cell RNA-seq to identify subpopulations and genotype-phenotype linkage. Nat. Commun 9, 4892 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Cooper SE et al. scSNV-seq: high-throughput phenotyping of single nucleotide variants by coupled single-cell genotyping and transcriptomics. Genome Biol. 25, 20 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Muyas F et al. De novo detection of somatic mutations in high-throughput single-cell profiling data sets. Nat. Biotechnol 1–10 (2023) doi: 10.1038/s41587-023-01863-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Kim HS et al. Direct measurement of engineered cancer mutations and their transcriptional phenotypes in single cells. Nat. Biotechnol 1–9 (2023) doi: 10.1038/s41587-023-01949-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Wells MF et al. Natural variation in gene expression and viral susceptibility revealed by neural progenitor cell villages. Cell Stem Cell 30, 312–332.e13 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Rodriguez-Meira A et al. Unravelling Intratumoral Heterogeneity through High-Sensitivity Single-Cell Mutational Analysis and Parallel RNA Sequencing. Mol. Cell 73, 1292–1305.e8 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Myers RM et al. Integrated Single-Cell Genotyping and Chromatin Accessibility Charts JAK2V617F Human Hematopoietic Differentiation. 2022.05.11.491515 Preprint at 10.1101/2022.05.11.491515 (2022). [DOI] [Google Scholar]
  • 26.Macaulay IC, Ponting CP & Voet T Single-Cell Multiomics: Multiple Measurements from Single Cells. Trends Genet. TIG 33, 155–168 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Findlay GM et al. Accurate classification of BRCA1 variants with saturation genome editing. Nature 562, 217–222 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Tewhey R et al. Direct Identification of Hundreds of Expression-Modulating Variants using a Multiplexed Reporter Assay. Cell 165, 1519–1529 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Griesemer D et al. Genome-wide functional screen of 3’UTR variants uncovers causal variants for human disease and evolution. Cell 184, 5247–5260.e19 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.van Arensbergen J et al. High-throughput identification of human SNPs affecting regulatory element activity. Nat. Genet 51, 1160–1169 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Przybyla L & Gilbert LA A new era in functional genomics screens. Nat. Rev. Genet 1–15 (2021) doi: 10.1038/s41576-021-00409-w. [DOI] [PubMed] [Google Scholar]
  • 32.Hanna RE et al. Massively parallel assessment of human variants with base editor screens. Cell 184, 1064–1080.e20 (2021). [DOI] [PubMed] [Google Scholar]
  • 33.Jin X et al. A metastasis map of human cancer cell lines. Nature 588, 331–336 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.van Arensbergen J et al. High-throughput identification of human SNPs affecting regulatory element activity. Nat. Genet 51, 1160–1169 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Jones EM et al. A Scalable, Multiplexed Assay for Decoding GPCR-Ligand Interactions with RNA Sequencing. Cell Syst. 8, 254–260.e6 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Fulco CP et al. Activity-by-contact model of enhancer–promoter regulation from thousands of CRISPR perturbations. Nat. Genet 51, 1664–1669 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Gray VE et al. Elucidating the Molecular Determinants of Aβ Aggregation with Deep Mutational Scanning. G3 GenesGenomesGenetics 9, 3683–3689 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Chiasson MA et al. Multiplexed measurement of variant abundance and activity reveals VKOR topology, active site and human variant impact. eLife 9, e58026 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Siegel DA, Le Tonqueze O, Biton A, Zaitlen N & Erle DJ Massively parallel analysis of human 3′ UTRs reveals that AU-rich element length and registration predict mRNA destabilization. G3 GenesGenomesGenetics 12, jkab404 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Adams EJ et al. FOXA1 mutations alter pioneering activity, differentiation and prostate cancer phenotypes. Nature 571, 408–412 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Arruabarrena-Aristorena A et al. FOXA1 Mutations Reveal Distinct Chromatin Profiles and Influence Therapeutic Response in Breast Cancer. Cancer Cell 38, 534–550.e9 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Harrod A et al. Genome engineering for estrogen receptor mutations reveals differential responses to anti-estrogens and new prognostic gene signatures for breast cancer. Oncogene 41, 4905–4915 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Alice K, Marioni JC & Morgan MD Milo2.0 unlocks population genetic analyses of cell state abundance using a count-based mixed model 2023.11.08.566176 Preprint at 10.1101/2023.11.08.566176 (2023). [DOI] [Google Scholar]
  • 44.Xie S, Armendariz D, Zhou P, Duan J & Hon GC Global Analysis of Enhancer Targets Reveals Convergent Enhancer-Driven Regulatory Modules. Cell Rep. 29, 2570–2578.e5 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Backwell L & Marsh JA Diverse Molecular Mechanisms Underlying Pathogenic Protein Mutations: Beyond the Loss-of-Function Paradigm. Annu. Rev. Genomics Hum. Genet 23, 475–498 (2022). [DOI] [PubMed] [Google Scholar]
  • 46.Zheng GXY et al. Massively parallel digital transcriptional profiling of single cells. Nat. Commun 8, 14049 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Cusanovich DA et al. Multiplex single-cell profiling of chromatin accessibility by combinatorial cellular indexing. Science 348, 910–914 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Buenrostro JD et al. Single-cell chromatin accessibility reveals principles of regulatory variation. Nature 523, 486–490 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Adamson B et al. A Multiplexed Single-Cell CRISPR Screening Platform Enables Systematic Dissection of the Unfolded Protein Response. Cell 167, 1867–1882.e21 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Dixit A et al. Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens. Cell 167, 1853–1866.e17 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Replogle JM et al. Combinatorial single-cell CRISPR screens by direct guide RNA capture and targeted sequencing. Nat. Biotechnol 1–8 (2020) doi: 10.1038/s41587-020-0470-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Datlinger P et al. Pooled CRISPR screening with single-cell transcriptome readout. Nat. Methods 14, 297–301 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Pierce SE, Granja JM & Greenleaf WJ High-throughput single-cell chromatin accessibility CRISPR screens enable unbiased identification of regulatory networks in cancer. Nat. Commun 12, 2969 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Liscovitch-Brauer N et al. Scalable pooled CRISPR screens with single-cell chromatin accessibility profiling. bioRxiv 2020.11.20.390971 (2020) doi: 10.1101/2020.11.20.390971. [DOI] [Google Scholar]
  • 55.Replogle JM et al. Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell 185, 2559–2575.e28 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Bandaru P et al. Deconstruction of the Ras switching cycle through saturation mutagenesis. eLife 6, e27810 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Ursu O et al. Massively parallel phenotyping of coding variants in cancer with Perturb-seq. Nat. Biotechnol 40, 896–905 (2022). [DOI] [PubMed] [Google Scholar]
  • 58.Martin-Rufino JD et al. Massively parallel base editing to map variant effects in human hematopoiesis. Cell 186, 2456–2474.e24 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Frenkel M, Hujoel MLA, Morris Z & Raman S Discovering chromatin dysregulation induced by protein-coding perturbations at scale. BioRxiv Prepr. Serv. Biol 2023.09.20.555752 (2023) doi: 10.1101/2023.09.20.555752. [DOI] [Google Scholar]
  • 60.Bartosovic M, Kabbe M & Castelo-Branco G Single-cell CUT&Tag profiles histone modifications and transcription factors in complex tissues. Nat. Biotechnol 39, 825–835 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Tan L, Xing D, Chang C-H, Li H & Xie XS Three-dimensional genome structures of single diploid human cells. Science 361, 924–928 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Ramani V et al. Massively multiplex single-cell Hi-C. Nat. Methods 14, 263–266 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Nagano T et al. Single-cell Hi-C reveals cell-to-cell variability in chromosome structure. Nature 502, 59–64 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Stoeckius M et al. Cell Hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics. Genome Biol. 19, 224 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Chen AF et al. NEAT-seq: simultaneous profiling of intra-nuclear proteins, chromatin accessibility and gene expression in single cells. Nat. Methods 19, 547–553 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Richards S et al. Standards and Guidelines for the Interpretation of Sequence Variants: A Joint Consensus Recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med. Off. J. Am. Coll. Med. Genet 17, 405–424 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Faure AJ et al. Mapping the energetic and allosteric landscapes of protein binding domains. Nature 604, 175–183 (2022). [DOI] [PubMed] [Google Scholar]
  • 68.Weng C, Faure AJ, Escobedo A & Lehner B The energetic and allosteric landscape for KRAS inhibition. Nature 1–10 (2023) doi: 10.1038/s41586-023-06954-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Boldridge WC et al. A multiplexed bacterial two-hybrid for rapid characterization of protein–protein interactions and iterative protein design. Nat. Commun 14, 4636 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Yao D et al. Scalable genetic screening for regulatory circuits using compressed Perturb-seq. Nat. Biotechnol 1–14 (2023) doi: 10.1038/s41587-023-01964-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Pinglay S et al. Multiplex generation and single cell analysis of structural variants in a mammalian genome. 2024.01.22.576756 Preprint at 10.1101/2024.01.22.576756 (2024). [DOI] [Google Scholar]
  • 72.Bosco N et al. KaryoCreate: A CRISPR-based technology to study chromosome-specific aneuploidy by targeting human centromeres. Cell 186, 1985–2001.e19 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Mätlik K et al. Cell Type Specific CAG Repeat Expansions and Toxicity of Mutant Huntingtin in Human Striatum and Cerebellum. BioRxiv Prepr. Serv. Biol 2023.04.24.538082 (2023) doi: 10.1101/2023.04.24.538082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Lu T-Y, Smaruj PN, Fudenberg G, Mancuso N & Chaisson MJP The motif composition of variable number tandem repeats impacts gene expression. Genome Res. 33, 511–524 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Tabula Sapiens Consortium* et al. The Tabula Sapiens: A multiple-organ, single-cell transcriptomic atlas of humans. Science 376, eabl4896 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Emont MP et al. A single cell atlas of human and mouse white adipose tissue. Nature 603, 926–933 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Eraslan G et al. Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function. Science 376, eabl4290 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Travaglini KJ et al. A molecular cell atlas of the human lung from single-cell RNA sequencing. Nature 587, 619–625 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 79.Takahashi K & Yamanaka S Induction of pluripotent stem cells from mouse embryonic and adult fibroblast cultures by defined factors. Cell 126, 663–676 (2006). [DOI] [PubMed] [Google Scholar]
  • 80.Ng AHM et al. A comprehensive library of human transcription factors for cell fate engineering. Nat. Biotechnol 39, 510–519 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Roberts GA et al. Dissecting OCT4 defines the role of nucleosome binding in pluripotency. Nat. Cell Biol 23, 834–845 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Joung J et al. A transcription factor atlas of directed differentiation. Cell 186, 209–229.e26 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Xie M & Fussenegger M Designing cell function: assembly of synthetic gene circuits for cell biology applications. Nat. Rev. Mol. Cell Biol 19, 507–525 (2018). [DOI] [PubMed] [Google Scholar]
  • 84.Srivatsan SR et al. Massively multiplex chemical transcriptomics at single-cell resolution. Science 367, 45–51 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.McFaline-Figueroa JL et al. Multiplex single-cell chemical genomics reveals the kinase dependence of the response to targeted therapy. Cell Genomics 0, (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Valencia AM et al. Landscape of mSWI/SNF chromatin remodeling complex perturbations in neurodevelopmental disorders. Nat. Genet 55, 1400–1412 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Kadoch C & Crabtree GR Mammalian SWI/SNF chromatin remodeling complexes and cancer: Mechanistic insights gained from human genomics. Sci. Adv 1, e1500447 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Badia-i-Mompel P et al. Gene regulatory network inference in the era of single-cell multi-omics. Nat. Rev. Genet 24, 739–754 (2023). [DOI] [PubMed] [Google Scholar]
  • 89.Kim D et al. Gene regulatory network reconstruction: harnessing the power of single-cell multi-omic data. Npj Syst. Biol. Appl 9, 1–13 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 90.Bravo González-Blas C et al. SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks. Nat. Methods 20, 1355–1367 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Kamimoto K et al. Dissecting cell identity via network inference and in silico gene perturbation. Nature 614, 742–751 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Frazer J et al. Disease variant prediction with deep generative models of evolutionary data. Nature 599, 91–95 (2021). [DOI] [PubMed] [Google Scholar]
  • 93.Brandes N, Goldman G, Wang CH, Ye CJ & Ntranos V Genome-wide prediction of disease variant effects with a deep protein language model. Nat. Genet 55, 1512–1522 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Adzhubei IA et al. A method and server for predicting damaging missense mutations. Nat. Methods 7, 248–249 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES