Skip to main content
Bioinformatics Advances logoLink to Bioinformatics Advances
. 2025 May 9;5(1):vbaf106. doi: 10.1093/bioadv/vbaf106

Perspective on recent developments and challenges in regulatory and systems genomics

Julia Zeitlinger 1,2, Sushmita Roy 3,4, Ferhat Ay 5,6,7, Anthony Mathelier 8,9,10, Alejandra Medina-Rivera 11, Shaun Mahony 12, Saurabh Sinha 13,14, Jason Ernst 15,16,17,18,19,20,
Editor: Thomas Lengauer
PMCID: PMC12233095  PMID: 40626041

Abstract

Summary: Predicting how genetic variation affects phenotypic outcomes at the organismal, cellular, and molecular levels requires deciphering the cis-regulatory code, the sequence rules by which non-coding regions regulate genes. In this perspective, we discuss recent computational progress and challenges toward solving this fundamental problem. We describe how cis-regulatory elements are mapped with various genomics assays and how studies of the 3D chromatin organization could help identifying long-range regulatory effects. We discuss how the cis-regulatory sequence rules can be learned and interpreted with sequence-to-function neural networks, with the goal of identifying genetic variants in human disease. We also describe current methods for mapping gene regulatory networks to describe biological processes. We point out current gaps in knowledge along with technical limitations and benchmarking challenges of computational methods. Finally, we discuss newly emerging technologies, such as spatial transcriptomics, and outline strategies for creating a more general model of the cis-regulatory code that is more broadly applicable across cell types and individuals.

1 The fundamental problem of the cis-regulatory code

Predicting how genetic variation affects phenotypic outcomes at the organismal, cellular, and molecular levels is a key challenge in biology. This is especially difficult for variants found in the non-coding portion of the genome, which regulates when, where, and at which level genes are transcribed in each cell type. Gene regulatory instructions are encoded in units of 100 bp- to 1 kb-long DNA sequences called cis-regulatory elements (CREs). CREs such as enhancers and promoters contain binding sites for transcription factors (TFs), which function together with various transcriptional regulators and complexes to set the desired gene expression levels. This cis-regulatory code, the set of rules by which CRE sequences collectively control gene expression in a cell type, is incompletely understood, which makes it extremely challenging to predict how genetic variation alters gene regulation.

A comprehensive understanding of the cis-regulatory code would provide a blueprint of how cells differentiate into the various cell types during embryonic development, predict how genetic variants influence development and health, and identify the molecular mechanisms altered by disease-associated genetic differences. The resulting knowledge may also allow us to develop therapeutic interventions, including engineering enhancer variants with highly specific activities, to direct cells toward favorable gene expression programs that restore and maintain cellular function.

Deciphering the cis-regulatory code is an extraordinarily complex problem. Each cell type has a unique combination of TFs, expressed at specific levels and whose activity is sometimes under the control of extracellular signals. Given the activity of all TFs in a given cell type, the transcription of all genes should be predictable from the DNA sequence alone. However, such predictions are challenging since TFs can act combinatorially and influence multiple regulatory layers: TFs cooperate to bind and access CREs, change the chromatin environment, recruit additional regulatory proteins, and, together with TFs at other CREs, regulate the transcription of target genes (Slattery et al. 2014, Zabidi and Stark 2016, Zeitlinger 2020, Kim and Wysocka 2023, Preissl et al. 2023). Furthermore, many CREs are specific for TF combinations that may only be present in specialized cell types. Thus, complexity arises from the many regulatory layers and the cell-type-specific nature by which TFs regulate these layers.

Although deciphering the cis-regulatory complexity is a daunting challenge, there has been tremendous progress (Fig. 1). Thanks to decades of work with model organisms, we have a much-improved mechanistic understanding of gene regulation (Zabidi and Stark 2016, Zeitlinger 2020, Kim and Wysocka 2023, Preissl et al. 2023). With the genomics revolutions, we have a vast and growing number of genomics datasets from multiple experimental assays over many cell types, conditions, and organisms. Importantly, advances in computational methods, especially deep learning (Eraslan et al. 2019), have shown that complex cis-regulatory rules can be learned from such datasets in a given cell type. These approaches increasingly reveal mechanistic insights into TF cooperativity, make experimental predictions, and allow the effect of genetic variants to be studied.

Figure 1.

Figure 1.

Regulatory genomics studies the intricate relationships between transcription factor activities and target gene expression, mediated by cis-regulatory elements. Researchers seek to identify general principles of gene regulation, such as how enhancer activities, motif syntax rules, 3D genome organization, and chromatin states relate to each other and ultimately to gene expression. A major goal is to build accurate and generalizable sequence-function models that can not only reveal underlying mechanisms but make predictions of variant effects. Another major theme is the reconstruction of gene regulatory networks to describe biological processes of interest. Research into computational methods and models in regulatory genomics strives to make best use of diverse and rapidly advancing experimental technologies.

In this perspective, we delve into the current computational challenges facing the field of regulatory genomics, with a specific focus on the cis-regulatory code that regulates transcription [for post-transcriptional mechanisms, see Keene (2007) and Zhao et al. (2017)]. We discuss state-of-the-art computational methods that aim to map CREs and the 3D chromatin organization, learn the cis-regulatory rules of specific cell types, predict the effect of genetic variants, and map gene regulatory networks (GRNs) during specific biological processes (Fig. 1). We provide our perspective on the current gaps in knowledge, limitations of current methods and how to benchmark them, and opportunities for developing methods that analyze and integrate newly emerging data, such as those from spatial omics technologies. Finally, we outline possible strategies for closing the remaining gaps and creating a path toward a more general model of the cis-regulatory code that is more broadly applicable across cell types and individuals.

2 Leveraging experimental progress: chromatin-based annotations of CREs

The advent of high-throughput genomic technologies has made it possible to comprehensively map the regulatory landscape across the genome in many cell types, conditions, and organisms (Hawkins et al. 2010, Zhou et al. 2011), but this is a very complex task. While the basic building blocks of the cis-regulatory code, the TFs and the specific sequence motifs they bind, have been experimentally determined for the majority of human TFs (Lambert et al. 2018, Rauluseviciute et al. 2024), TF binding in vivo is highly cooperative and cell-type specific. TFs can be mapped genome-wide inside cells using chromatin immunoprecipitation coupled to sequencing (ChIP-seq) or related assays (Nakato and Sakata 2021) (Table 1), but extensive ChIP experiments are only available for a few human cell types (Moore et al. 2020) and challenging to perform comprehensively across many cell types and conditions. Our understanding of the combinatorial landscape by which TFs cooperate to access and regulate different CREs across cell types is therefore still limited.

Table 1.

Genomics assays.

Assay type Examples Readout References
TF binding and chromatin state at CREs ChIP-seq, CUT&Tag, ChIP-DIP TF binding, histone modifications (H3K27ac, H3K4me3, H3Kme1), Pol II, co-factors (p300, Brd4, mediator, TBP), loop extrusion factors (Cohesin) Zabidi and Stark (2016), Moore et al. (2020), Henikoff et al. (2021), Perez et al. (2024)
High-resolution TF binding ChIP-exo, ChIP-nexus Binding of TFs, Pol II, initiation factors (TBP, TFIIA, TFIIB, TFIID, TFIIF) He et al. (2015), Lai et al. (2021)
Chromatin accessibility ATAC-seq, DNase-seq Accessible CREs, nucleosomes, and TF footprints Buenrostro et al. (2013), Vierstra et al. (2020)
Reporter assays MPRA, STARR-seq Enhancer or silencer activity, defined as the amount by which sequences boost or repress the expression of a reporter gene Neumayr et al. (2019), La Fleur et al. (2024)
Nascent transcripts PRO-cap, NET-seq, NET-CAGE, TT-seq 5′ end of promoter or enhancer RNA, i.e. capped, short, or has a newly incorporated special nucleotide, depending on protocol Mahat et al. (2016), Mayer and Churchman (2016), Schwalb et al. (2016), Hirabayashi et al. (2019)
Transcripts RNA-seq Depending on protocol, transcripts may be preferentially nuclear or polyadenylated. Long-read sequencing improves isoform identification Belchikov et al. (2024)
3D contact maps Hi-C, Hi-ChIP, Micro-C DNA fragments in close proximity are crosslinked and ligated after DNA digestion with restriction enzymes, MNase or exonucleases. Chimeric linear fragments are enriched and sequenced at high coverage Lieberman-Aiden et al. (2009), Rao et al. (2014), Hsieh et al. (2015), Mumbach et al. (2016)
Single-molecule footprinting SMF, Fiber-seq Nucleosomes and TF footprints at each individual DNA fragment since they protect from exogenous DNA methylation. Methylation is detected by bisulfite conversion and Illumina sequencing, or directly by Nanopore/PacBio sequencing Krebs (2021), Tullius et al. (2024)
Nucleosomes MNase-seq, Chemical-Cleavage-seq Nucleosome-sized DNA fragments generated by MNase or by site-directed chemical cleavage of nucleosomes that carry modified histones Voong et al. (2017)

To facilitate the discovery of CREs and their TF binding sites, a popular approach is to identify the genomic sequences that are accessible in chromatin. CREs have long been known to be hypersensitive to DNase digestion (Burch and Weintraub 1983), and DNase-seq provides comprehensive and quantitative information on chromatin accessibility genome-wide, while also revealing TF binding footprints with deeper sequencing (Song and Crawford 2010, Thurman et al. 2012, Vierstra et al. 2020). A similar chromatin accessibility readout including footprints is obtained with ATAC-seq, a simple assay that requires less input material (Buenrostro et al. 2013, Bentsen et al. 2020, Pampari et al. 2025). However, while these assays allow the comprehensive identification of candidate CREs in a cell type, which of them contribute to gene regulation under the examined conditions, and what class of CREs they might represent is unclear.

We discuss here the following broad classes of CREs: promoters, architectural elements, enhancers, and the not uniquely defined class of silencers (Ernst and Kellis 2013). Promoters, the regions around the transcription start sites that initiate transcription, as well as architectural elements that help organize chromatin in 3D (e.g., CTCF-bound regions), represent the backbone of gene regulation and tend to be constitutively open across cell types (Phillips and Corces 2009). Enhancers on the other hand, which activate promoters from various distances away from the promoter, are often highly cell-type-specific. They tend to be regulated at the level of accessibility and can be identified through their differential accessibility across cell types. However, enhancers may not necessarily activate transcription when accessible. This is because TFs can activate or repress transcription, and thus an accessible enhancer that is not active may even function as a silencer when bound by repressive TFs (Berest et al. 2019, Pang and Snyder 2020, Segert et al. 2021, Zhu et al. 2025).

Determining when TFs activate or repress is not straightforward. TFs typically harbor effector domains that mediate activation and repression in functional assays, but the net effect on transcription can be context-dependent (Tycko et al. 2024). An activating TF that helps recruit a repressor to a nearby motif can have a net repressive effect (Papagianni et al. 2018, Brennan et al. 2023). Furthermore, repressors are often deployed to make the activity of enhancers more stimulus or cell-type dependent, and this type of repression is difficult to detect in genomics assays (Pang et al. 2023). Thus, repression is often found at CREs classified and functioning as enhancers, which means that silencers do not represent a separate class of CREs, although it is possible that some CREs are dedicated silencers.

To better classify the function of candidate CREs, additional datasets that measure aspects of the chromatin state are helpful, especially histone modifications measured by ChIP-seq. Active enhancers tend to be flanked by high histone acetylation levels and specific forms of histone methylation (e.g. H3K27ac, H3K4me1). Other markers for active enhancers are the presence of transcriptional co-activators (e.g. p300, Brd4, and Mediator) (Kagey et al. 2010, Zabidi and Stark 2016, Malik and Roeder 2023) and enhancer transcription (e.g. measured by CAGE-NET, PRO-seq, or NET-seq) (Table 1) (Policastro and Zentner 2021). Repressed chromatin regions can be marked by H3K27me3 or H3K9me3, but this type of repression tends to be more constitutive and long-range and is not a good marker for repression by sequence-specific TFs at enhancers and silencers (Koenecke et al. 2017, Zhu et al. 2025). Obtaining comprehensive experimental data for characterizing chromatin states has been a major goal for the ENCODE (Moore et al. 2020) and Roadmap Epigenomics (Roadmap Epigenomics Consortium et al. 2015) projects and has opened the door to systematically analyzing CREs across cell types.

As data for multiple chromatin marks are often collected in the same cell type, a common strategy for systematically annotating candidate CREs is to integrate information using multivariate hidden Markov models such as ChromHMM, or related probabilistic models, to define chromatin states (Ernst and Kellis 2010, 2012, Hoffman et al. 2012, Libbrecht et al. 2021). This strategy has been used for CRE annotations across a wide range of cell types and conditions (Roadmap Epigenomics Consortium et al. 2015, Stunnenberg et al. 2016, Moore et al. 2020). In addition to directly using observed data, CREs have been annotated in cell types with incompletely observed data using imputed epigenomic datasets (Ernst and Kellis 2015, Schreiber et al. 2020). The availability of data from hundreds of cell types and conditions has also spurred analysis approaches that can directly categorize different classes of cell-type restricted or constitutively active CREs (Meuleman et al. 2020, Vu and Ernst 2022).

While such chromatin-based approaches have increased the number of identified putative CREs into the millions (Meuleman et al. 2020, Moore et al. 2020), many additional CREs are likely encoded in the genome. For example, evolutionarily conserved sequence analyses suggest that a substantial portion of conserved non-coding bases are not well captured by large compendiums of annotations (Grujic et al. 2020, Christmas et al. 2023). Therefore, expanding these datasets by profiling rare or hard-to-access cell types and conditions is critical.

Single-cell chromatin assays are ideal for capturing CREs from rare cell types. Due to their sparsity and dimensionality, these datasets add additional computational challenges. The most widely applied single-cell chromatin assay is ATAC-seq, sometimes jointly performed with gene expression (Preissl et al. 2023). Single-cell assays for measuring other epigenetic features, including histone modifications and DNA methylation, have also been developed and continue to mature, along with computational methods to integrate the resulting data (Shema et al. 2019, Preissl et al. 2023). To inform mechanistic models of gene regulation, another useful chromatin assay is single-molecule footprinting. This technology uses exogenous DNA methylation to capture the footprints of bound TFs and nucleosomes on the same DNA molecule, allowing absolute measurements of bound fractions and co-occurrence events in a population of cells (Sönmezer et al. 2021, Kreibich et al. 2023). Finally, assays that measure how chromatin accessibility, TF binding, and histone marks are spatially organized inside cells (Deng et al. 2022, Lu et al. 2022) provide additional cell-type-specific information on gene regulation.

A continuing challenge is to benchmark CRE annotations. Traditionally, this is performed by comparing the results with established genome annotations and other experimental data not used during model learning (Ernst and Kellis 2010, Vu and Ernst 2022). However, this approach does not directly validate novel annotations. Additional high-throughput functional assays, such as massively parallel reporter assays (MPRA) and non-coding CRISPR-based screens, provide a promising avenue to evaluate and characterize CRE annotations (Gasperini et al. 2020, Yao et al. 2024).

3 Learning from 3D genome organization: linking CREs to target genes

While CREs provide the building blocks of the cis-regulatory code, one of the outstanding problems is to understand how multiple CREs, proximal and distal, affect the expression of a gene. This is challenging because enhancers can act over large genomic distances of over 1 Mb to modulate the expression levels of their target promoters (Lettice et al. 2003, Long et al. 2020, Chandra et al. 2021). Since these regulatory connections are thought to require physical proximity in 3D space, mapping the 3D organization of DNA in the context of chromatin has been an important goal in the field.

The 3D chromatin structure can be captured as genome-wide contact maps using Chromosome Conformation Capture (3C) technologies such as Hi-C (Lieberman-Aiden et al. 2009, Rao et al. 2014). These methods can characterize multiple layers of genome organization, including cell-type-specific aspects (Dekker et al. 2023). However, the limited resolution of Hi-C (typically 5–40 kb) and the requirement for very high sequencing coverage (e.g. 1 billion for 5-kb resolution) coupled with underrepresentation of distal interactions make it challenging to assign enhancers to target genes. This led to the development of additional steps in the assay that increase the coverage of the relevant regions (Fullwood et al. 2009, Mifsud et al. 2015, Fang et al., 2016, Mumbach et al. 2016) or increase the resolution by which contacts are detected. For example, Micro-C can achieve kilobase resolution genome-wide (Hsieh et al. 2015, Krietenstein et al. 2020, Harris et al. 2023) or sub-kilobase resolution for targeted regions (Goel et al. 2023). High-resolution contact maps in primary cell types have helped predict target genes for distally located disease-associated genetic variants (Javierre et al. 2016, Chandra et al. 2021, Hamley et al. 2023). However, they are difficult to obtain for a large collection of primary cell types, and the genome-wide resolution is still not at the level of individual CREs to obtain generalizable insights into enhancer–promoter interactions.

Computational methods that detect patterns in these contact maps have revealed multiple levels of organization that could influence gene expression: chromatin compartments, topologically associating domains (TADs), and chromatin interactions or loops (Zhang et al. 2024). Multi-megabase chromatin compartments correspond to the large-scale division between transcriptionally active euchromatin (A compartment) and inactive heterochromatin (B compartment). TADs spanning 100 kb–1 Mb regions are often considered as regulatory units of coordinated gene expression within which CREs interact more frequently with one another (Beagan and Phillips-Cremins 2020). Two convergent CTCF motifs, which stop cohesin-mediated loop extrusion (Fudenberg et al. 2016), often demarcate their boundaries. Another pattern is preferential interactions among CREs, whether it is mediated by loop extrusion to specifically bring two distal CREs in close proximity (e.g. chromatin loop between an enhancer and a promoter) or broader co-localization of CREs in the 3D space potentially through their homotypic interactions.

With the increasing resolution of contact maps, recent methods have detected additional patterns that provide insights into how CREs influence loop extrusion and, in turn, gene expression (Vian et al. 2018, Guo et al. 2022, Yoon et al. 2022). For instance, stripes form when a loop anchor interacts with a long stretch of chromatin at high frequency (Yoon et al. 2022). Stripe anchors generally mark clusters of enhancers that regulate multiple genes throughout the domain (Vian et al. 2018, Guo et al. 2022, Yoon et al. 2022), e.g. at immunoglobulin loci (Vian et al. 2018, Hu et al. 2023) or developmentally regulated genes (Kraft et al. 2019). However, how these emerging contact patterns can be generalized to model gene regulation is not yet clear.

The role of chromatin contacts in determining the effect of an enhancer on gene expression has also been analyzed more explicitly (Fulco et al. 2019). This led to the Activity-by-Contact (ABC) model, which has been applied across a large collection of cell types to better link non-coding risk variants to disease genes (Nasser et al. 2021). The ABC model assumes that the influence of an enhancer depends on its activity, multiplied by the intensity of contact with the promoter, and that multiple enhancers contribute to gene expression independently. This analysis demonstrated, in part, the utility of cell-type-specific contact maps in determining functional enhancer–promoter interactions but also showed that genomic distance is the strongest determinant of enhancer–promoter contacts and has a large effect on gene expression levels. This distance dependence agrees with experimental data (Zuin et al. 2022).

While the ABC model serves as a good baseline model, it cannot predict the effect of some validated enhancers, suggesting that additional unknown mechanisms are at play. While the majority of closely spaced enhancers appear to function multiplicatively without interactions, there is evidence for synergistic or redundant interactions, and specific promoters may not be as responsive to enhancers (Liu et al. 2019, Gschwind et al. 2023, Zhou et al. 2024a). Most notably, it is unclear how some enhancers can find their target genes with high specificity over hundreds of kilobases, while others cannot. Studies in the fruit fly suggest a new class of architectural CREs that enables enhancers to do so by mediating long-distance chromatin interactions (Batut et al. 2022). Such “extender elements” have recently been identified in mice: when located next to an enhancer, they allow the enhancer to regulate target genes over hundreds of kilobases of distance (Bower et al. 2024).

Taken together, these results suggest that enhancer–promoter interactions are a complex layer of the cis-regulatory code. Enhancer–promoter interactions depend on genomic distance, the 3D organization created by architectural CREs, and long-range contacts enabled by newly emerging CREs. Predictive features might be identified and characterized more precisely by additional experimental efforts, including generating high-resolution contact maps for more cell types, devising experimental techniques with improved temporal and spatial resolution (“4D”) (Sekelja et al. 2016, Dekker et al. 2023), measuring multiple modalities such as chromatin organization and gene expression simultaneously (Su et al. 2020, Liu et al. 2021, 2023, Zhou et al. 2024b), as well as mapping multi-way contact of chromatin (Beagrie et al. 2017, Quinodoz et al. 2018, Tavares-Cadete et al. 2020, Oudelaar et al. 2022, Deshpande et al. 2022b, Chen et al. 2023a). Significant computational innovation will be required to leverage these additional data to identify general mechanistic principles and extract the relevant sequence features that mediate long-distance interactions and the interplay of multiple CREs in gene regulation.

4 The solution to complexity: sequence-to-function models

The ultimate goal of understanding CRE function is to decipher the cis-regulatory code embedded within these sequences. Each assay measures specific regulatory activities across the genome in a given cell type, and these activities should be driven by TF binding motifs and possibly other sequence patterns within CRE sequences. However, determining the exact relationship between sequence and a given functional readout is difficult. The traditional approach is to select the regions with high TF binding, chromatin accessibility, or enhancer activity and to identify TF motifs that are statistically overrepresented (Bailey 2021). While this provides a set of TF motifs, it does not capture how the affinity and syntax of the motifs in their genomic context affect the readout (Crocker et al. 2008, Farley et al. 2016). Since each genomic region is very different, discovering genome-wide rules by which motifs interact and predicting the experimental readout is challenging. Fortunately, it is now possible to train neural networks (“sequence-to-function models”) to perform this task (Table 2).

Table 2.

(A) Sequence-to-function deep learning models and (B) interpretation tools for the cis-regulatory code of transcription.

(A) Model type Examples DNA region size Training data References
CNN that predicts binary features of genomics data Sei Sequences of 4 kb Genomic assay data across selected regions Chen et al. (2022b)
CNN that learns to predict 128 bp binned profiles Basenji Sequences of ∼131 kb to 1 Mb Genomic assay data across selected regions Kelley et al. (2018)
CNN that learns to predict counts and base-resolution profiles BPNet, ChromBPNet Sequences of ∼1 to 20 kb Genomic assay data across selected regions, option for bias correction Avsec et al. (2021b), Pampari et al. (2025)
CNN that learns to predict counts or bins DeepSTARR, LegNet Sequences from STARR-seq/MPRA (∼150 to 500 bp) STARR-seq/MPRA-seq counts de Almeida et al. (2022), Penzar et al. (2023)
Transformer that learns to predict long-range binned profiles Enformer, Borzoi Sequences of 200–500 kb as 32–128 bp bins Genomic assays, including expression data, across tiled or selected regions Avsec et al. (2021a), Linder et al. (2025)
CNN that learns to predict binned 2D matrices Akita, ORCA Sequences of 1–256 Mb Hi-C data Fudenberg et al. (2020)  Zhou (2022)
(B) Interpretation task Example Input Method Output References
Sequence attribution through backpropagation DeepSHAP, gradient-based saliency maps Trained model Shapley or gradients Attribution scores for predicted regions Lundberg and Lee (2017), Shrikumar et al. (2017), Majdandzic et al. (2023)
Sequence attribution by in silico saturation mutagenesis (ISM) ISM, fast ISM Trained model Systematically mutate and predict Attribution scores for predicted regions Zhou and Troyanskaya (2015), Nair et al. (2022)
Motif identification among sequences with high attribution TF-MoDISco Sequences with attribution Jaccard similarity Motif representations as PWM and CWM Shrikumar et al. (2020)
Mapping motifs in input regions CWM scanning Attribution and CWMs of motifs Match by sequence and attribution Regions with mapped motifs instances Avsec et al. (2021b)
Motif strength measured in isolation Affinity distillation, global importance Trained model, mapped motifs Predict motif in randomized sequences over control Relative predicted motif strength for every sequence variant of the motif Koo et al. (2021), Alexandari et al. (2023)
Motif cooperativity Systemtatic in silico perturbations Trained model, mapped motif instances Mutate motifs pairwise and predict effect Cooperativity for each motif pair as a function of distance Avsec et al. (2021b), Koo et al. (2021), de Almeida et al. (2022)

During training, sequence-to-function models learn to predict an experimental readout across a large number of CREs directly from the underlying genomic sequence. By optimizing the prediction accuracy, the model learns sequence rules inside a “black box” without prior biological assumptions. When the predictive performance is high on withheld data not seen by the model during training, this suggests that the learned rules apply genome-wide. To train on different experimental data types, models are typically optimized to the biological problem of interest. They may differ in their DNA input size, the selection of regions, model type (e.g. convolutional neural networks or transformers), architecture and model size (e.g. filters, layers, receptive field), and loss function. In general, training a model to predict high-resolution, high-coverage data quantitatively at base resolution produces the most nuanced sequence features (Toneyan et al. 2022, Avsec et al. 2021b). For some purposes, models predict binary or categorical data or data that are averaged across genomic bins, which reduces the computing requirements. For example, by binning to 128 bp and using a transformer architecture, Enformer predicts data across 200 kb (Kelley et al. 2018, Kelley 2020, Avsec et al. 2021a, Chen et al. 2022b).

The high prediction accuracy of these models is useful on their own, e.g. to test the effect of genetic variants, but considerable power of sequence-to-function models lies in the interpretation. This is counterintuitive as “black box” models are traditionally considered as uninterpretable because features and relationships are learned in a distributed manner inside the neural network. However, with DNA sequence being a relatively simple input, post hoc interpretation approaches that query the model as a whole have been very successful in revealing sequence motifs and their syntax rules (Table 2) (Alipanahi et al. 2015, Avsec et al. 2021b, de Almeida et al. 2022, Novakovsky et al. 2023). An alternative is to incorporate a priori biological knowledge into the neural network architecture, but these constraints come at the expense of learning more complex phenomena and cannot capture unknown cis-regulatory sequence rules (Agarwal et al. 2021, Balcı et al. 2023, Novakovsky et al. 2023, Toneyan and Koo 2024). Therefore, interpreting the learned sequence features of “black box” models is currently the best approach to uncover new sequence rules in the cell type of interest.

There are several complementary approaches by which models can be interpreted (Table 2). The most common first step is to use an attribution method, e.g. gradient-based saliency maps (Sundararajan et al. 2017, Jha et al. 2020, Majdandzic et al. 2023), or Shapley-based methods such as DeepSHAP (Lundberg and Lee 2017, Shrikumar et al. 2017). These methods assign scores for how much each feature in the input (i.e. each base) contributes to the output prediction. Motifs highlighted in genomic regions by high scores can be summarized by tools like TF-MoDisco (Shrikumar et al. 2020). These TF motif representations can then label high-scoring motif instances in the genome (Avsec et al. 2021b). This approach outperforms traditional position weight matrix (PWM) methods because the mapped motif instances have high attributions and thus were considered to be important by the model in the specific sequence context.

To extract additional sequence rules, such as how TF motifs interact, trained models can be queried with in silico sequence designs, e.g. by perturbing motifs in genomic sequences or injecting motifs into randomized sequences (Alipanahi et al. 2015, Zhou and Troyanskaya 2015, Avsec et al. 2021b, Koo et al. 2021, Trevino et al. 2021, de Almeida et al. 2022, Nair et al. 2022). Analyzing predictions with systematic sequence designs allows the extraction of relative motif affinities (Alexandari et al. 2023, Brennan et al. 2023) and specific syntax rules by which motif pairs interact (Koo et al. 2021, Avsec et al. 2021b, de Almeida et al. 2022).

In this manner, the genome-wide cis-regulatory sequence rules that underlie various data modalities have been characterized and linked to molecular mechanisms (Novakovsky et al. 2023). For example, the syntax rules by which TFs cooperate based on in vivo binding data show that some TFs preferentially interact when the motifs are within nucleosome distance, while others may physically cooperate on DNA when the motifs are spaced at a fixed distance (Avsec et al. 2021b). These rules match prior mechanistic studies (Long et al. 2016, Smith et al. 2023), showing that neural networks can learn accurate biological representations without a priori knowledge of the underlying biophysical properties and mechanistic principles.

Sequence-to-function models can also predict or interpret how TF binding relates to cell-type-specific chromatin environments. Some approaches incorporate chromatin features alongside sequence as model inputs, allowing the models to distinguish between direct sequence predictors of TF-DNA binding and a more generalized dependency on chromatin state (Srivastava et al. 2021, Arora et al. 2023). With a sufficiently complex sequence-to-function model, chromatin accessibility and other chromatin features can themselves be predicted from DNA, revealing the sequence rules by which TFs shape the chromatin landscape. Such approaches have revealed that TFs drive chromatin accessibility proportional to the motif affinity, that some TFs have a repressive effect, and that TFs often function synergistically in making chromatin accessible (Kim et al. 2021a, Brennan et al. 2023, Bravo González-Blas et al. 2024). Presumably, these sequence rules reflect how TFs act on nucleosomes (Brennan et al. 2023), but the mechanisms for this type of TF cooperativity are not well understood. Here is therefore an opportunity for sequence models to inspire mechanistic studies.

Another challenge is understanding how sequence rules specify enhancer activity and target gene activation. The rules of enhancer activity differ from those of chromatin accessibility (Brennan et al. 2023), but how exactly is not well understood. Models trained on large-scale reporter assays have shown that motif syntax and repressive motifs are important for enhancer activity (Movva et al. 2019, de Almeida et al. 2022). These assays are however typically episomal and use short DNA regions, thus how the measured activity translates into a genomic context is not entirely clear (Inoue et al. 2017, Sahu et al. 2022). For example, active enhancers in the genome are typically flanked by ChIP-seq signals of histone modifications. This signal can be predicted from DNA sequence (Kelley et al. 2018, Kelley 2020, Avsec et al. 2021a, Chen et al. 2022b), but how the underlying motifs combinatorially instruct histone modifications and enhancer activity remains to be extracted from the models and experimentally validated.

To understand how enhancers interact with promoters, sequence-to-function models have been trained to predict Hi-C contact maps (Table 2) (Fudenberg et al. 2020, Schwessinger et al. 2020, Zhou 2022, Zhang et al. 2024). These models have learned that the backbone of 3D organization is established by motifs of architectural proteins, such as CTCF in vertebrates (Piecyk et al. 2022). However, the backbone tends to be cell-type invariant, and the sequence features that promote enhancer–promoter interactions in specific cell types are difficult to identify. It is possible that further improvements, e.g. using chromatin accessibility and ChIP-seq data as additional input during training (Tan et al. 2023), better data coverage and resolution, or more extensive model interpretation, could reveal cell-type-specific features that help predict gene expression.

Ultimately, one would like to directly predict gene expression, either steady-state RNA levels or nascent transcription data, from DNA sequence. This is possible and yields highly accurate predictions (Kelley et al. 2018, Avsec et al. 2021a, Linder et al. 2025) and promoter syntax (Dudnyk et al. 2024, Cochran et al. 2024, He and Danko 2024). However, the input from distal enhancers is not well captured, suggesting that the models are missing some cell-type-specific sequence rules, perhaps related to enhancer activation or long-distance enhancer-promoter interactions (Karollus et al. 2023, Kathail et al. 2024).

This shows that our understanding of the cis-regulatory code and the molecular mechanisms by which TFs mediate enhancer activation and target gene expression is still incomplete. Current models specifically learn the data on which they were trained and are thus specific for a data modality and cell type. This is still true when a single model is trained on many datasets at the same time as a multi-task model (Kelley et al. 2018, Kelley 2020, Avsec et al. 2021a, Chen et al. 2022b). In this case, features learned from different datasets might be shared inside the model, but this does not necessarily mean that the learned sequence rules are more coherent and better represent biology.

A solution for handling different data modalities is to learn assay biases in separate deep learning models (e.g. Tn5 insertion bias in ATAC-seq data) such that specific features of the cis-regulatory code can be learned more explicitly (Brennan et al. 2023, Pampari et al. 2025). If the biophysical properties of these rules are known, secondary surrogate models can be trained to fit these properties, e.g. TF binding and cooperativity (Seitz et al. 2024). This could eventually lead to biophysical models of the cis-regulatory code, but this would require extensive knowledge of the molecular mechanisms beyond TF binding, which currently does not exist.

To improve existing models or develop new models, a better mechanistic understanding of the cis-regulatory code through systematic interpretation of various models would be highly beneficial. While interpreting models to uncover new biology, it may also be important to examine the model’s limitations, e.g. using data simulation models (Chen and Capra 2020, Prakash et al. 2022, Chen et al. 2024), and to analyze the limitations of the experimental data, e.g. assay biases, experimental artifacts, or the effect of low resolution or low coverage. If this is done for many datasets and modalities, the sequence rules should overlap and provide clearer expectations of the underlying biophysical constraints and molecular mechanisms that lead to the activation of enhancers and their target genes. The ultimate challenge will be to train models that generalize cis-regulatory rules and can predict data for cell types not trained on, a goal that may require significant computational innovation involving domain adaptation.

Meanwhile, experimental validation is needed to ensure the learned sequence rules are accurate. While strongly dependent on the experimental accessibility of the model system, one approach is to predict and test the effect of targeted perturbations, e.g. mutating genomic regions by CRISPR (Avsec et al. 2021b) or knocking down a TF (Brennan et al. 2023). Large-scale MPRA reporter assays can be used to validate the learned sequence rules at higher throughput (Kim et al. 2021a). A powerful validation is to use the trained model to create synthetic enhancers and test them in vivo with a reporter assay. Synthetic designs can be generated from random sequences, manual manipulation or trimming of existing enhancers, or through de novo design of enhancers (de Almeida et al. 2024, Taskiran et al. 2024). In the future, such synthetic enhancer designs could be used to create enhancers with increased or altered function, with the potential of using such designs for therapeutic treatments.

5 From genotype to phenotype: predicting the effect of regulatory variants

A promising application of sequence-to-function models is the prediction and interpretation of regulatory variants involved in the predisposition, onset or progression of complex diseases. Most SNPs identified in genome-wide association studies (GWAS) fall in the non-coding portion of the human genome (Buniello et al. 2019). However, due to linkage disequilibrium, the SNPs identified in GWAS are often not the causal variants but are located nearby (Visscher et al. 2012). This necessitates fine-mapping to pinpoint the causal variants, most of which are expected to alter gene expression. Sequence-to-function models can be used to identify and interpret regulatory variants by quantifying their predicted effect on expression variation or other molecular features.

Different types of gene expression variation are however not equally amenable to modeling. Predicting gene-to-gene variation within the same cell type is an easier task because the levels vary widely and can be predicted from promoter-proximal sequences without a comprehensive understanding of the cis-regulatory code of distal enhancers (Karollus et al. 2023). Predicting the variation between cell types across an organism is more challenging because the cis-regulatory code is highly diverse across cell types and often driven by distal enhancers far away from promoters. Predicting variation in gene expression across individuals in a population is the most challenging task (Huang et al. 2023, Sasse et al. 2023, Tang et al. 2023). Not only does it require predicting cell-type-specific gene expression of individual genetic variants, whose effects are often small, but also how they affect specific gene expression programs and phenotypes of cells, leading to disease susceptibility (Manolio et al. 2009). The few examples of genetic variants that are well characterized include fetal hemoglobin variants at the BCL11A locus (Bauer et al. 2013) and an obesity variant at the FTO locus (Claussnitzer et al. 2015).

A starting point for this challenging task is to assess which genetic variants differentially affect TF binding. The simplest models use PWMs to assess how TF motif affinities differ between the alternate and the reference alleles (Kumar et al. 2017, Fornes et al. 2018, Santana-Garcia et al. 2019). Predictions from such methods can correlate well with observed allele-specific binding events derived from ChIP-seq experiments (Fornes et al. 2018). More sophisticated models of TF binding specificity are trained on high-throughput data, such as in vitro TF binding data (Martin et al. 2019). Another category of methods combines multiple features, such as evolutionary conservation, chromatin states, and enhancer–promoter interactions, to predict causal variants. Such models are trained or evaluated on known genetic variants from the Human Gene Mutation Database or ClinVar (Huang et al. 2017, Gao et al. 2018, Rogers et al. 2018) and can be combined into ensemble models (Zhang et al. 2019).

If sequence-to-function models are trained on molecular genomic datasets, they can directly predict how genetic variants alter the experimental outcome (Sokolova et al. 2024). This is because these models learn genome-wide rules and make accurate predictions across diverse genomic sequences, including variants not directly trained on. Models that predict TF binding, chromatin accessibility, and histone modifications have had reported successes for predicting individual variant effects (Zhou and Troyanskaya 2015, Kelley et al. 2016, Trevino et al. 2021, Chen et al. 2023b). However, it is not always clear how these effects influence enhancer activation and specific changes in nearby promoter activity. Therefore, sequence-to-function models that predict MPRA reporter activity (Movva et al. 2019) and expression data are appealing and have shown promise (Zhou et al. 2018, Avsec et al. 2021a, Linder et al. 2025).

The additional information gained from sequence-to-function models has so far been limited when identifying genetic variants from GWAS studies (Dey et al. 2020). A fundamental limitation is that cell-type-specific effects can only be predicted when trained on appropriate experimental data, and currently existing datasets only cover limited cell types. Furthermore, current expression models often struggle to learn long-range enhancer–promoter interactions (Karollus et al. 2023) or whether a given variant will boost or diminish expression (Huang et al. 2023, Sasse et al. 2023). This may change with the availability of more data, more advanced sequence-to-function models, and better strategies for integrating the model predictions with human GWAS data. A potential bridge between genetic variants and disease traits is gene expression data or other genomics data profiled across individuals, ideally from a tissue of interest for the trait (Drusinsky et al. 2024). These data can be used to infer molecular quantitative trait loci (QTLs) (e.g. Ramdas et al. 2022). While there has been limited overlap between expression QTLs and GWAS hits (Mostafavi et al. 2023), additional molecular QTLs have shown greater though still partial overlap (Wu et al. 2023). Overall, these results underscore the difficulty of predicting variation in gene expression across individuals.

A potentially lower-hanging fruit is to use sequence-to-function models to identify genetic variants that cause rare diseases (Critical Assessment of Genome Interpretation Consortium 2024, Sokolova et al. 2024). Rare variants are likely purged from the population by purifying selection and thus can have larger effect sizes that are more easily detected. Since they are rare in the population, GWAS studies lack the power to discover them, but they can still be predicted by sequence-to-function models.

Moving forward, a major bottleneck is the availability of uniformly processed and validated genomics data, as well as high-quality QTL data for expression and other genomics datasets for less characterized cell types (IGVF Consortium 2024). But even for well-studied cell types, identifying the key genes that affect the disease phenotype in the presence of multiple regulatory variants and secondary transcription effects is challenging (Manolio et al. 2009, Liu et al. 2019, Li and Ritchie 2021). Models that predict the effect of multiple regulatory variants and consider coding and non-coding epistasis in predicting disease outcomes would be useful (Monti and Ohler 2023). We note that while models are generally able to make accurate predictions for unseen variants, it is nevertheless important to include more ancestry-diverse sequences during training and to benchmark variants from the entire population so that all humans can maximally benefit in eventual clinical applications (Martin et al. 2017, Taylor et al. 2024). Finally, it will be critical to further develop and apply more experimental techniques such as CRISPR editing to validate causal genetic variants (Chen et al. 2023b, Pihlajamaa et al. 2023).

6 Assembling the parts: GRNs

Ultimately, cis-regulatory sequences are not only key for predicting expression levels, but also how cells change their gene expression program dynamically during embryonic development, exposure to stress, or disease pathogenesis. The methods described so far aimed to predict gene expression given a fixed steady-state cellular state defined by a specific set of active TFs. To predict how cells change their gene expression program, we need to understand how TF activities change over time. The activity of some TFs is regulated by signal transduction pathways in response to extracellular stimuli. However, most TFs, and the regulators they depend upon, are, to some extent, regulated at the expression level. Thus, TF activities are themselves the target of CRE regulation, which creates a dynamic system that allows cells to transition along specific cellular trajectories depending on their extracellular environment. The changing interactions between TFs, CRE activity, and the expression of target genes over time are called GRNs.

GRNs play an important role during embryonic development, where cells transition through multiple states to eventually acquire a specific cell fate with a characteristic cell-type-specific expression program. Indeed, the concept of GRNs was pioneered in sea urchin and Drosophila embryos by studying how key regulators identified through developmental genetics are themselves regulated (Levine and Davidson 2005). This led to the discovery of enhancers, which each drive the expression of the target gene in a specific spatio-temporal manner and are controlled by the combinatorial input from TFs active at that time. By tracing back how these TFs are regulated, coherent descriptions of developmental processes were obtained. However, such top-down models were restricted to key enhancers and TFs, required a laborious iterative experimental process, and the identified cis-regulatory sequence rules did not generalize to allow genome-wide predictions of gene expression from sequence alone.

Methods aimed at building GRNs from genome-wide data by explicitly modeling sequence motifs and their TFs as regulators are called cis-GRN methods (Table 3). These methods model gene expression over time or across cell types as a function of TF motifs found in candidate CRE regions, identified by histone marks, chromatin accessibility, or TF ChIP-seq data taken from developmental model systems (González et al. 2015, Ding et al. 2018b, Siahpirani et al. 2022). The initial candidate TF motifs are typically identified by scanning CREs with a library of known motifs (Sherwood et al. 2014, Chen et al. 2017, Bentsen et al. 2020) or by de novo motif finding (Setty and Leslie 2015). Relevant TF motif features are then discovered through their association with co-regulated target genes or modules. A drawback of this modeling approach is that it strongly depends on the choice of candidate CRE regions, TF motifs, and target gene modules. Initial methods have focused on proximal enhancers near promoters, but recent approaches also model or incorporate distal enhancers, either by their high accessibility when the target is active (e.g. Bravo González-Blas et al. 2023) or by using 3C data (e.g. Karbalayghareh et al. 2022). An advantage of cis-GRN methods is that CREs and the TF motifs involved in a particular process can be identified de novo and followed up with experiments.

Table 3.

GRN inference methods.

Example Model Input Output References
cis-GRN inference DREM Hidden Markov models TF-gene associations, gene expression (bulk) Dynamic regulatory maps between TF and gene sets Ernst et al. (2007)
DRMN Probabilistic models Chromatin assays, TF ChIP-seq, gene expression (bulk) CRE, TF, gene interactions Siahpirani et al. (2022)
trans-GRN inference Inferelator Linear regression Gene expression (bulk), TF-target prior interactions Regulator–gene interactions Skok Gibbs et al. (2022)
MERLIN Linear regression Gene expression (bulk), TF–target prior interactions Regulator–gene interactions Siahpirani and Roy (2017)
GENIE3 Random forests regression Gene expression (bulk) Regulator–gene interactions Huynh-Thu et al. (2010)
SCENIC Random forests regression Gene expression (scRNA-seq) Regulator–gene interactions Aibar et al. (2017)
Cell Oracle Linear regression Gene expression (scRNA-seq), accessibility (scATAC-seq) Regulator–gene interactions Kamimoto et al. (2023)
scMTNI Linear regression, multi-task learning Gene expression (scRNA-seq), accessibility (sc or bulk) Regulator–gene interactions Zhang et al. (2023)

An alternative set of approaches for building GRNs are trans-GRN methods, which infer the role of TFs and other “trans” regulators through their co-expression with their target genes (Kim et al. 2009, Amit et al. 2011). These methods require large sample sizes and use probabilistic graphical models such as Bayesian networks and their extensions (Friedman et al. 2000, Segal et al. 2003) or dependency networks with linear (Greenfield et al. 2013, Siahpirani and Roy 2017) or non-linear regression models (Huynh-Thu et al. 2010, Baran et al. 2012) (Table 3). Several of these methods have integrated known TF motifs located near promoters as secondary features to inform the network structure (Glass et al. 2013, Greenfield et al. 2013, Petralia et al. 2015, Siahpirani and Roy 2017). TF motifs can also inform TF activity levels when the expression levels are a poor predictor of its activity, e.g. because the TF is regulated by signaling (Wang et al. 2018, Miraldi et al. 2019).

The rise in single-cell genomics technology has greatly benefited GRN inference, especially trans-GRNs, which can leverage the large amounts of single-cell expression experiments (i.e. scRNA-seq) from normal and disease conditions. Single-cell resolution data better distinguish cell types and states among heterogeneous biological samples, thereby improving the discovery of the TFs that define each cell type. The larger sample sizes also allow for capturing additional non-linearities between TF and target expression with the help of deep learning models (Shu et al. 2021, Luo et al. 2022). Furthermore, cellular dynamics can be inferred using measured time (e.g. Ding et al. 2018a), velocity (Bocci et al. 2022, Burdziak et al. 2023), or pseudotime (Wang et al. 2023b, Kim et al. 2021b, Deshpande et al. 2022a), which can be used to inform GRN inference algorithms to capture fine-grained dynamics. These inferred networks can be analyzed to identify regulators that may mediate the transition between specific cell states, e.g. by determining rewiring scores of its local network topology (e.g. Zhang et al. 2023, Wang et al. 2023b) or using in silico perturbation analysis (Fleck et al. 2023, Kamimoto et al. 2023). These predictions can be used to engineer improved in vitro differentiation or trans-differentiation models.

Additional power for both cis- and trans-GRNs comes from the combination of scRNA-seq with single-cell ATAC-seq (scATAC-seq) data, ideally from a multi-omics assay where these data modalities are measured simultaneously in the same cells (Badia-I-Mompel et al. 2023). Single-cell chromatin accessibility data improve the quality of the inferred networks, resolution of cell types, and allows the analysis of TF motifs within CREs (Bravo González-Blas et al. 2023, Zhang et al. 2023, Wang et al. 2023b). Especially for cis-GRNs, cell-type-specific accessibility changes enable the prediction of long-range enhancer–promoter interactions without requiring Hi-C experiments (Pliner et al. 2018, Mitra et al. 2024, Sakaue et al. 2024).

Through such recent advances, trans-GRN and cis-GRN approaches are increasingly being combined, but current approaches still primarily leverage one approach and are limited in the information they capture. For trans-GRN methods, CREs and their sequence information are a secondary feature, which limits their ability to predict the effect of genetic variation and reveal the molecular underpinnings of the cis-regulatory code. But without the “trans” information, cis-GRNs cannot directly infer which TFs and additional cell-type-specific regulators shape the TF activities that regulate the identified TF motifs. Furthermore, both approaches typically lack the TF motif interaction rules and predictive accuracy that sequence-to-function models provide. Combining different approaches in a more seamless way could therefore enable an improved understanding of the cis-regulatory code.

A hurdle toward this goal is the benchmarking of cell-type-specific GRN models beyond well-studied systems. For example, comprehensive gold standards are lacking when studying human cell types not covered by ENCODE (Chen and Mar 2018, Pratapa et al. 2020, McCalla et al. 2023). To confirm the identity of a regulator and establish a relationship with target genes as causal therefore requires experimental perturbations (Seçilmiş et al. 2022, Choo et al. 2024). Since such follow-up experiments can be time-consuming, the availability of high-throughput perturbation screens that measure scRNA-seq in response to various regulator perturbations in individual cells (Dixit et al. 2016, Replogle et al. 2022, Schraivogel et al. 2023) would accelerate this validation step. Such data could also improve the models, including their ability to predict how perturbations disrupt cellular function. Finally, centralized repositories of datasets and methods (Ben Guebila et al. 2023, Chevalley et al. 2025, Wen et al. 2023) will be useful to advance methodological development and assess their practical utility for inferring the cis-regulatory code.

7 Multi-scale integration: spatial transcriptomics and beyond

As we improve our ability to integrate different data modalities into coherent computational frameworks, an exciting new frontier will be the integration of data across multiple scales, from molecule to cell to tissue and organs. A particularly interesting aspect of such multi-scale integration is the spatial organization of cells, which determines which signals a cell receives from its neighboring cells, an element of GRNs that can currently only be inferred indirectly. Spatial aspects of gene regulation have long been studied at a small scale during embryonic development (Dubuis et al. 2013) and are particularly relevant in the brain (Piwecka et al. 2023). Such spatial organization can now be captured in a high-throughput manner for virtually any tissue using recently developed spatial transcriptomics technologies.

Spatial transcriptomics quantifies the expression distributions of large numbers of genes in a tissue at the resolution of single (or few) cells or individual transcripts (Moses and Pachter 2022). These technologies rely either on sequencing with spatial barcoding (Chen et al. 2022a, Rodriques et al. 2019, Schott et al. 2024, Oliveira et al. 2024) or single-molecule-fluorescent in situ hybridization (Lubeck et al. 2014, Chen et al. 2015, Janesick et al. 2023, Khafizov et al. 2024). All methods aim to provide a comprehensive description of gene expression patterns across cells and tissues, without requiring prior hypotheses on which genes might be differentially expressed. However, current methods have limitations and represent different tradeoffs between the number of genes measured, spatial resolution, and tissue area coverage (Moses and Pachter 2022).

While these methods are still in their infancy, they have created new opportunities for developing analytical tools that can extract different types of biological patterns. Such computational tools are capable of (1) identifying genes that exhibit interesting spatial patterns (Svensson et al. 2018), (2) distilling a large number of spatial expression patterns into a smaller, representative set of patterns (Townes and Engelhardt 2023), (3) inferring prominent spatial regions in the tissue (Hu et al. 2021, Dong and Zhang 2022), (4) characterizing cell–cell interactions in terms of involved cell types or genes (Arnol et al. 2019, Dries et al. 2021, Cang et al. 2023), (5) identifying gene–gene interactions, e.g. ligand–receptor pairs, related to spatial expression (Yuan and Bar-Joseph 2020, Tanevski et al. 2022), and (6) detecting transcript localization (Xia et al. 2019, Mah et al. 2024) or transcript co-localization (Kumar et al. 2024) at subcellular resolution. With these promising developments, it will be important to systematically evaluate and benchmark these tools (Moses and Pachter 2022).

Spatial transcriptomics promises to provide novel insights into how the cellular dynamics and organization of tissues influence chromatin organization and gene regulation, but their potential has so far remained largely untapped. Glimpses into what is possible can be seen in some existing approaches, e.g. those for detecting cell–cell communication while incorporating signaling and regulatory networks (Browaeys et al. 2020) or identifying tissue-level variations in RNA localization events that hint at post-transcriptional regulatory processes (Kumar et al. 2024).

In the future, integrating spatial transcriptomics with additional data could improve the identification of signaling events and their impact on gene regulation in specific tissues and cell types. For example, combining spatial transcriptomics with spatial chromatin accessibility assays (Lu et al. 2022) could help understand gene regulatory mechanisms. Furthermore, tools for mapping non-spatial single-cell data to spatial data from the same tissue (Biancalani et al. 2021) could be highly informative, as they can lead to a common analytical framework for analyzing different single-cell measurements, e.g. transcripts and chromatin accessibility (Fang et al. 2021), contact maps (Rappoport et al. 2023), and proteomics (Bennett et al. 2023).

8 Future outlook for deciphering the cis-regulatory code

While there is rapid progress, major bottlenecks still exist in four areas. (1) We need to fill the remaining gaps in our mechanistic understanding of the cis-regulatory code to better model all steps of the cis-regulatory code (Fig. 1, top-left). (2) We need experimental data for more cell types and more comprehensive multiomic datasets, including perturbation experiments (Fig. 1, top-right). (3) We need better GRN methods that more seamlessly combine cis-GRN, trans-GRN, and sequence-to-function approaches to model cell state changes (Fig. 1, bottom-right). (4) We need major advances in sequence-to-function models that leverage the innovations above into more integrated and generalizable frameworks (Fig. 1, bottom-left).

Since we have an incomplete mechanistic understanding of how the cis-regulatory code is executed from sequence all the way to gene expression, it is currently difficult to pinpoint which regulatory steps are not well captured by current sequence-to-function expression models. Most notably, it is unclear how multiple enhancers activate specific target genes, which could depend on their biochemical activities, chromatin environment, relative distances, 3D organization, and other nearby CREs. Approaches that integrate 3C data, single-cell multiomic data, single-molecule footprinting, and perturbation experiments could provide the much-needed insights. The steps before, the local activities produced by TFs at enhancers, leading to chromatin accessibility, histone modifications, and other activating or repressing biochemical properties, are also not well understood. Deciphering the steps and general mechanistic principles will require an iterative process of model interpretation and experimental testing. Altogether, the mechanistic insights will help benchmark current models and enable more focused efforts to improve them.

Another challenge will be to predict gene expression across a much larger number of cell types. Each cell type has a unique combination and TF activities, and the exact sequence rules by which the TFs read out CREs cannot easily be predicted based on the TFs’ individual binding specificities (Jolma et al. 2015) and will require more high-resolution in vivo TF binding data, perhaps by adopting large-scale approaches (Perez et al. 2024). Increasing the experimental coverage to hard-to-access cell types and conditions during developmental processes and in heterogeneous adult tissues will likely occur through more general methods such as single-cell multi-omics data. While some missing data can be imputed when appropriate training data exist (Ernst and Kellis 2015, Schreiber et al. 2020), very unique combinatorial TF binding specificities are difficult to discover without sufficient experimental data.

One way to obtain missing cis-regulatory sequence information without experimental data is to directly leverage the large number of sequenced genomes across species. The specific combinations of TFs that specify a cell type tend to be evolutionarily conserved (Tarashansky et al. 2021, Kuderna et al. 2024), allowing CREs to be studied across evolution with sequence-to-function models (Minnoye et al. 2020, Kaplow et al. 2023). Moreover, DNA large language models trained to predict masked genome sequences may detect combinatorial TF motif patterns in some cases (Karollus et al. 2024, da Silva et al. 2024). However, this approach has limitations since these models currently struggle to learn cis-regulatory information when trained on genome-wide mammalian sequence data (Tang and Koo 2024) and not all features of the cis-regulatory code are conserved across larger evolutionary distances (Cavalheiro et al. 2023, Magri et al. 2024).

Another important gap is GRN methods that more fully describe how cells dynamically change their TF repertoire and gene expression program. The increasing number and quality of single-cell spatial and temporal multi-omics data create an opportunity to develop methods that better integrate cis-GRN, trans-GRN, and sequence-to-function approaches. This includes trans-GRN methods that link expression programs to the cell-type-specific distal and proximal enhancers that regulate the genes (Mitra et al. 2024, Sakaue et al. 2024), sequence-to-function models that identify the effect of TF motifs (Kim et al. 2021a, Bravo González-Blas et al. 2024), and cis-GRN methods that more explicitly identify which TFs might bind these motifs based on expression data (Yang and Pe’er 2024). Such methods can describe cellular changes across time and tissues and point to key TFs and CREs (Maslova et al. 2020, Janssens et al. 2022, Özel et al. 2022).

As with sequence-to-function models, the insights obtained from modeling GRNs are however specific to the system of interest. Although the principles should conceptually extend to other biological systems, there is currently no generalizable framework that enables the effective prediction of expression changes de novo, e.g. based on only TF activities or accessible CREs. Such domain adaptation will require significant computational innovation. Intermediate steps toward this goal may involve developing sequence-to-function methods that learn more directly how enhancers change their activities as a function of changing TF activities.

In the long term, fully deciphering the cis-regulatory code will be important to predict how genetic variation affects multi-scale behavior at the organismal, cellular, and molecular level. We will need major innovations in computational models to accurately predict cell-type-specific expression variation and the effect of genetic variants. Knowledge about the mechanistic steps and structural constraints of the cis-regulatory code could be introduced, e.g. through geometric deep learning (AlQuraishi and Sorger 2021, Bronstein et al. 2021, Wang et al. 2023a), while foundation models that can learn from large amounts of non-coding genome sequences could improve generalization to new cell types and systems (Simon et al. 2024, Szałata et al. 2024). Finally, it is expected that more experimental data, including high-throughput perturbation screens, will increase the performance of models (de Boer and Taipale 2024). To enable such breakthroughs, having the best possible benchmarks in the form of gold-standard datasets, and tools for mechanistic interpretation will be paramount.

Finally, having a strong community that promotes collaboration and communication will accelerate the pace by which progress is made. For this reason, we, as authors, are part of the Regulatory and Systems Genomics community of special interest (RegSys COSI), which organizes scientific sessions at the annual ISMB meeting. We encourage more participation from people of all backgrounds and career stages.

Glossary

  • Cis-regulatory element (CRE): A DNA region in the genome bound by transcription factors or other proteins and contributes to gene regulation. It regulates a target gene in cis, thus its proximity to the gene on DNA is important. The most commonly studied CREs are enhancers, promoters, and architectural elements.

  • Cis-regulatory code: The set of rules by which CREs are read by transcription factors and control gene expression.

  • Enhancer: A CRE harboring motifs for one or more transcription factors with the ability to become active in one or more cellular conditions, resulting typically in the enhancement of transcription of a nearby gene. When not active, enhancers may be bound by repressive transcription factors.

  • Silencer: A CRE that represses transcription of a nearby gene, typically by harboring motifs for repressive transcription factors. Repression often occurs at enhancers to ensure that activation occurs in a context-specific and stimulus-specific fashion. Thus, silencers are not a unique class of CREs but may be a state in which enhancers function under certain conditions. More comprehensive analyses of silencers are needed to better understand their role in gene regulation.

  • Transcriptional activator and repressor: In addition to the DNA-binding domain, transcription factors often have separate domains that mediate transcriptional activation or repression, e.g. by recruiting co-activators or co-repressors.

  • Transcription factor motif: Short DNA sequence pattern recognized by a transcription factor through protein–DNA interactions.

  • Position Weight Matrix (PWM): A mathematical model for representing a transcription factor binding motif, where the frequencies of each base are summarized for each position.

  • Contribution Weight Matrix (CWM): A representation of a transcription factor binding motif that was extracted from a trained sequence-to-function model and summarizes the average predicted contribution of each base.

  • Chromatin state: The state in which genomic regions are found in vivo in the context of nucleosomes. Of specific interest are the chromatin states that change dynamically depending on the cellular conditions, such as DNA accessibility, histone modifications, and other bound proteins, since they are often the cause or effect of ongoing regulatory processes of transcriptional regulation.

  • ChIP-seq: Chromatin immunoprecipitation sequencing. An assay for genome-wide profiling of transcription factor binding, histone modifications, and other features of chromatin state that can be specifically targeted by an antibody.

  • ATAC-seq: Assay for transposase-accessible chromatin with sequencing, an experimental method for the genome-wide profiling of chromatin accessibility. scATAC-seq refers to the single-cell version.

  • scRNA-seq: Genome-wide assay that measures RNA abundance in single cells.

  • Multi-omics assay: Assays that measure multiple data modalities simultaneously, ideally as a single-cell assay. The most common example is the combination of RNA-seq and ATAC-seq.

  • Hi-C: A high-throughput technique for mapping the 3D structure of chromatin by mapping the pairwise contact frequencies between genomic regions.

  • Contact map: A description of the 3D chromatin structure as mapped by Hi-C or related techniques, which is useful for understanding regulatory interactions among CREs and genes.

  • CTCF: CCCTC-binding factor, a protein with a major role in regulating the 3D chromatin structure.

  • Loop extrusion: A model that proposes that long-range cis-interactions within a DNA molecule are generated by loop extrusion factors (e.g. cohesin) that bind to DNA and reel flanking regions of the same DNA molecule into a loop, i.e. demarcated by insulating elements (e.g. CTCF binding in convergent orientation).

  • Neural network: A type of machine learning model, which can be trained to make accurate predictions from large amounts of complex data, typically by allowing many flexible parameters without specifying specific variables, features, or their relationships (black box model).

  • Deep learning model: A neural network model with many layers (is “deep”), typically used to learn complex features of the cis-regulatory code.

  • Sequence-to-function model: A neural network trained to predict experimental data (“function”) from DNA sequences. The model architecture and sequence input length can vary depending on whether transcription factor binding, chromatin accessibility, histone modification data, gene expression, or genome contact maps are predicted. Examples are convolutional neural networks and transformers.

  • Interpreting a neural network/deep learning model: Using specific interpretation tools to open a black box model to understand what features and rules a neural network model learned during training.

  • Single-nucleotide polymorphisms (SNPs): Genomic sequences in which specific bases (A, C, T, or G) differ between individuals.

  • Genome-wide association studies (GWAS): Statistical analysis of how genetic variants (usually SNPs) in individuals are associated with traits or diseases.

  • Gene regulatory network (GRN): A collection of direct regulatory relationships between transcription factors, CREs, and target genes, often used as a model for expression changes between cellular conditions.

  • cis-GRN: GRNs reconstructed based on analyzing cis-regulatory elements and transcription factor motifs.

  • trans-GRN: GRNs reconstructed based on analyzing co-expression between transcription factors and target genes.

Acknowledgements

We thank Anshul Kundaje, Žiga Avsec, Neşet Özel, Alireza Karbalayghareh, and Christina Leslie for feedback on the article. We also would like to thank all past organizers and participants of the RegSys COSI sessions at ISMB, which provided the basis for much of the work described in this perspective. Since the breadth of the topic is huge, there are many important contributions, and we apologize to those whose work we missed.

Contributor Information

Julia Zeitlinger, Stowers Institute for Medical Research, Kansas City, MO 64112, United States; Department of Pathology & Laboratory Medicine, The University of Kansas Medical Center, Kansas City, KS 66160, United States.

Sushmita Roy, Department of Biostatistics and Medical Informatics, University of Wisconsin-Madison, Madison, WI 53715, United States; Wisconsin Institute for Discovery, University of Wisconsin-Madison, Madison, WI 53715, United States.

Ferhat Ay, Centers for Autoimmunity, Inflammation and Cancer Immunotherapy, La Jolla Institute for Immunology, La Jolla, CA 92037, United States; Bioinformatics and Systems Biology Program, University of California, San Diego, La Jolla, CA 92093, United States; Department of Pediatrics, University of California, San Diego, La Jolla, CA 92093, United States.

Anthony Mathelier, Norwegian Centre for Molecular Biosciences and Medicine (NCMBM), Nordic EMBL Partnership, University of Oslo, Oslo 0318, Norway; Department of Medical Genetics, Institute of Clinical Medicine, University of Oslo and Oslo University Hospital, Oslo 0450, Norway; Bioinformatics in Life Science (BiLS) initiative, Department of Pharmacy, University of Oslo, Oslo 0371, Norway.

Alejandra Medina-Rivera, Laboratorio Internacional de Investigación sobre el Genoma Humano, Universidad Nacional Autónoma de México, Santiago de Querétaro 76230, Mexico.

Shaun Mahony, Center for Eukaryotic Gene Regulation, Department of Biochemistry and Molecular Biology, The Pennsylvania State University, University Park, PA 16802, United States.

Saurabh Sinha, Walter H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA 30332, United States; H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA 30332, United States.

Jason Ernst, Department of Biological Chemistry, University of California, Los Angeles, Los Angeles, CA 90095, United States; Computer Science Department, University of California, Los Angeles, Los Angeles, CA 90095, United States; Department of Computational Medicine, University of California, Los Angeles, Los Angeles, CA 90095, United States; Eli and Edythe Broad Center of Regenerative Medicine and Stem Cell Research at University of California, Los Angeles, Los Angeles, CA 90095, United States; Jonsson Comprehensive Cancer Center, University of California, Los Angeles, Los Angeles, CA 90095, United States; Molecular Biology Institute, University of California, Los Angeles, Los Angeles, CA 90095, United States.

Author Contributions

Julia Zeitlinger (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Sushmita Roy (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Ferhat Ay (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Anthony Mathelier (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Alejandra Medina-Rivera (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Shaun Mahony (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), Saurabh Sinha (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal]), and Jason Ernst (Conceptualization [equal], Writing—original draft [equal], Writing—review & editing [equal])

Conflict of interest

J.Z. owns a patent on ChIP-nexus (no. 10287628). The other authors have no conflicts of interest to declare.

Funding

This work was supported by the Stowers Institute for Medical Research [to J.Z.]; the National Institutes of Health [R35-GM128938 to F.A., DP1DA044371 to J.E., U01MH130995 to J.E., U01HG012079 to J.E., R01-GM144708-03 to S.R., R35-GM131819 to S.S., R35GM144135 to S.M.]; the National Science Foundation [CAREER 2045500 to S.M.]; the Research Council of Norway [187615 to A.M.]; the Helse Sør-Øst [to A.M.]; the University of Oslo through the Norwegian Centre for Molecular Biosciences and Medicine (NCMBM) [to A.M.]; the Norwegian Cancer Society [215027, 272930 to A.M.]; the Chan Zuckerberg Initiative Ancestry Network [2021—240438 to A.M.-R.]; the CONACYT-FORDECYT-PRONACES [11311 and 6390 to A.M.-R.]; and the Programa de Apoyo a Proyectos de Investigación e InnovaciónTecnológica—Universidad Nacional Autónoma de México (PAPIIT-UNAM) [IN218023 to A.M.-R.].

Data availability

As this is a review article there is no associated data.

References

  1. Agarwal R, Melnick L, Frosst N  et al.  Neural additive models: interpretable machine learning with neural nets. In: Advances in Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates, Inc., 2021, 4699–4711. [Google Scholar]
  2. Aibar S, González-Blas CB, Moerman T  et al.  SCENIC: single-cell regulatory network inference and clustering. Nat Methods  2017;14:1083–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Alexandari AM, Horton CA, Shrikumar A  et al. De novo distillation of thermodynamic affinity from deep learning regulatory sequence models of in vivo protein–DNA binding. bioRxiv Prepr. Serv. Biol., 10.1101/2023.05.11.540401, 2023, preprint: not peer reviewed. [DOI]
  4. Alipanahi B, Delong A, Weirauch MT  et al.  Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning. Nat Biotechnol  2015;33:831–8. [DOI] [PubMed] [Google Scholar]
  5. AlQuraishi M, Sorger PK.  Differentiable biology: using deep learning for biophysics-based and data-driven modeling of molecular mechanisms. Nat. Methods  2021;18:1169–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Amit I, Regev A, Hacohen N  et al.  Strategies to discover regulatory circuits of the mammalian immune system. Nat Rev Immunol  2011;11:873–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Arnol D, Schapiro D, Bodenmiller B  et al.  Modeling cell–cell interactions from spatial molecular data with spatial variance component analysis. Cell Rep  2019;29:202–11.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Arora S, Yang J, Akiyama T  et al. Joint sequence & chromatin neural networks characterize the differential abilities of Forkhead transcription factors to engage inaccessible chromatin. bioRxiv, Prepr. Serv. Biol., 10.1101/2023.10.06.561228, 2023, preprint: not peer reviewed. [DOI]
  9. Avsec Ž, Agarwal V, Visentin D  et al.  Effective gene expression prediction from sequence by integrating long-range interactions. Nat Methods  2021a;18:1196–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Avsec Ž, Weilert M, Shrikumar A  et al.  Base-resolution models of transcription-factor binding reveal soft motif syntax. Nat Genet  2021b;53:354–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Badia-I-Mompel P, Wessels L, Müller-Dott S  et al.  Gene regulatory network inference in the era of single-cell multi-omics. Nat Rev Genet  2023;24:739–54. [DOI] [PubMed] [Google Scholar]
  12. Bailey TL.  STREME: accurate and versatile sequence motif discovery. Bioinformatics  2021;37:2834–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Balcı AT, Ebeid MM, Benos PV  et al.  An intrinsically interpretable neural network architecture for sequence-to-function learning. Bioinformatics  2023;39:i413–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Baran Y, Pasaniuc B, Sankararaman S  et al.  Fast and accurate inference of local ancestry in Latino populations. Bioinformatics  2012;28:1359–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Batut PJ, Bing XY, Sisco Z  et al.  Genome organization controls transcriptional dynamics during development. Science  2022;375:566–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Bauer DE, Kamran SC, Lessard S  et al.  An erythroid enhancer of BCL11A subject to genetic variation determines fetal hemoglobin level. Science  2013;342:253–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Beagan JA, Phillips-Cremins JE.  On the existence and functionality of topologically associating domains. Nat Genet  2020;52:8–16. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Beagrie RA, Scialdone A, Schueler M  et al.  Complex multi-enhancer contacts captured by genome architecture mapping. Nature  2017;543:519–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Belchikov N, Hsu J, Li XJ  et al.  Understanding isoform expression by pairing long-read sequencing with single-cell and spatial transcriptomics. Genome Res  2024;34:1735–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Ben Guebila M, Wang T, Lopes-Ramos CM  et al.  The Network zoo: a multilingual package for the inference and analysis of gene regulatory networks. Genome Biol  2023;24:45. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Bennett HM, Stephenson W, Rose CM  et al.  Single-cell proteomics enabled by next-generation sequencing or mass spectrometry. Nat Methods  2023;20:363–74. [DOI] [PubMed] [Google Scholar]
  22. Bentsen M, Goymann P, Schultheis H  et al.  ATAC-seq footprinting unravels kinetics of transcription factor binding during zygotic genome activation. Nat Commun  2020;11:4267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Berest I, Arnold C, Reyes-Palomares A  et al.  Quantification of differential transcription factor activity and multiomics-based classification into activators and repressors: diffTF. Cell Rep  2019;29:3147–59.e12. [DOI] [PubMed] [Google Scholar]
  24. Biancalani T, Scalia G, Buffoni L  et al.  Deep learning and alignment of spatially resolved single-cell transcriptomes with Tangram. Nat Methods  2021;18:1352–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Bocci F, Zhou P, Nie Q  et al.  spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data. Mol Syst Biol  2022;18:e11176. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Bower G, Hollingsworth EW, Jacinto S et al. Conserved cis-acting range extender element mediates extreme long-range enhancer activity in mammals. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.05.26.595809, 2024, preprint: not peer reviewed. [DOI]
  27. Bravo González-Blas C, De Winter S, Hulselmans G  et al.  SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks. Nat Methods  2023;20:1355–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Bravo González-Blas C, Matetovici I, Hillen H  et al.  Single-cell spatial multi-omics and deep learning dissect enhancer-driven gene regulatory networks in liver zonation. Nat Cell Biol  2024;26:153–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Brennan KJ, Weilert M, Krueger S  et al.  Chromatin accessibility in the Drosophila embryo is determined by transcription factor pioneering and enhancer activation. Dev Cell  2023;58:1898–916.e9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Bronstein MM, Bruna J, Cohen T  et al. Geometric deep learning: grids, groups, graphs, geodesics, and gauges, arXiv, arXiv:2104.13478, 2021, preprint: not peer reviewed.
  31. Browaeys R, Saelens W, Saeys Y  et al.  NicheNet: modeling intercellular communication by linking ligands to target genes. Nat Methods  2020;17:159–62. [DOI] [PubMed] [Google Scholar]
  32. Buenrostro JD, Giresi PG, Zaba LC  et al.  Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods  2013;10:1213–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Buniello A, MacArthur JAL, Cerezo M  et al.  The NHGRI-EBI GWAS catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res  2019;47:D1005–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Burch JB, Weintraub H.  Temporal order of chromatin structural changes associated with activation of the major chicken vitellogenin gene. Cell  1983;33:65–76. [DOI] [PubMed] [Google Scholar]
  35. Burdziak C, Zhao CJ, Haviv D  et al.  scKINETICS: inference of regulatory velocity with single-cell transcriptomics data. Bioinformatics  2023;39:i394–403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Cang Z, Zhao Y, Almet AA  et al.  Screening cell–cell communication in spatial transcriptomics via collective optimal transport. Nat Methods  2023;20:218–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Cavalheiro GR, Girardot C, Viales RR  et al.  CTCF, BEAF-32, and CP190 are not required for the establishment of TADs in early Drosophila embryos but have locus-specific roles. Sci Adv  2023;9:eade1085. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Chandra V, Bhattacharyya S, Schmiedel BJ  et al.  Promoter-interacting expression quantitative trait loci are enriched for functional genetic variants. Nat Genet  2021;53:110–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Chen A, Liao S, Cheng M  et al.  Spatiotemporal transcriptomic atlas of mouse organogenesis using DNA nanoball-patterned arrays. Cell  2022a;185:1777–92.e21. [DOI] [PubMed] [Google Scholar]
  40. Chen KH, Boettiger AN, Moffitt JR  et al.  RNA imaging. Spatially resolved, highly multiplexed RNA profiling in single cells. Science  2015;348:aaa6090. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Chen KM, Wong AK, Troyanskaya OG  et al.  A sequence-based global map of regulatory activity for deciphering human genetics. Nat Genet  2022b;54:940–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Chen L, Capra JA.  Learning and interpreting the gene regulatory grammar in a deep learning framework. PloS Comput Biol  2020;16:e1008334. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Chen L-F, Lee J, Boettiger A  et al.  Recent progress and challenges in single-cell imaging of enhancer–promoter interaction. Curr Opin Genet Dev  2023a;79:102023. [DOI] [PubMed] [Google Scholar]
  44. Chen S, Mar JC.  Evaluating methods of inferring gene regulatory networks highlights their lack of performance for single cell gene expression data. BMC Bioinformatics  2018;19:232. [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Chen V, Yang M, Cui W  et al.  Applying interpretable machine learning in computational biology-pitfalls, recommendations and opportunities for new developments. Nat Methods  2024;21:1454–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Chen X, Yu B, Carriero N  et al.  Mocap: large-scale inference of transcription factor binding sites from chromatin accessibility. Nucleic Acids Res  2017;45:4315–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Chen Z, Javed N, Moore M  et al.  Integrative dissection of gene regulatory elements at base resolution. Cell Genom  2023b;3:100318. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Chevalley M, Roohani YH, Mehrjou A et al. A large-scale benchmark for network inference from single-cell perturbation data. Commun Biol 2025;8:412. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Choo D, Shiragur K, Uhler C. Causal discovery under off-target interventions. In: Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, Valencia, Spain. PMLR, 2024, 1621–9.
  50. Christmas MJ, Kaplow IM, Genereux DP  et al. ; Zoonomia Consortium. Evolutionary constraint and innovation across hundreds of placental mammals. Science  2023;380:eabn3943. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Claussnitzer M, Dankel SN, Kim K-H  et al.  FTO obesity variant circuitry and adipocyte browning in humans. N Engl J Med  2015;373:895–907. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Cochran K, Yin M, Mantripragada A  et al. Dissecting the cis-regulatory syntax of transcription initiation with deep learning. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.05.28.596138, 2024, preprint: not peer reviewed. [DOI]
  53. Critical Assessment of Genome Interpretation Consortium. CAGI, the critical assessment of genome interpretation, establishes progress and prospects for computational genetic variant interpretation methods. Genome Biol  2024;25:53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  54. Crocker J, Tamori Y, Erives A  et al.  Evolution acts on enhancer organization to fine-tune gradient threshold readouts. PloS Biol  2008;6:e263. [DOI] [PMC free article] [PubMed] [Google Scholar]
  55. da Silva PT, Karollus A, Hingerl J  et al. Nucleotide dependency analysis of DNA language models reveals genomic functional elements. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.07.27.605418, 2024, preprint: not peer reviewed. [DOI]
  56. de Almeida BP, Reiter F, Pagani M  et al.  DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers. Nat Genet  2022;54:613–24. [DOI] [PubMed] [Google Scholar]
  57. de Almeida BP, Schaub C, Pagani M  et al.  Targeted design of synthetic enhancers for selected tissues in the Drosophila embryo. Nature  2024;626:207–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  58. de Boer CG, Taipale J.  Hold out the genome: a roadmap to solving the cis-regulatory code. Nature  2024;625:41–50. [DOI] [PubMed] [Google Scholar]
  59. Dekker J, Alber F, Aufmkolk S  et al.  Spatial and temporal organization of the genome: current state and future aims of the 4D nucleome project. Mol Cell  2023;83:2624–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
  60. Deng Y, Bartosovic M, Ma S  et al.  Spatial profiling of chromatin accessibility in mouse and human tissues. Nature  2022;609:375–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Deshpande A, Chu L-F, Stewart R  et al.  Network inference with Granger causality ensembles on single-cell transcriptomics. Cell Rep  2022a;38:110333. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Deshpande AS, Ulahannan N, Pendleton M  et al.  Identifying synergistic high-order 3D chromatin conformations from genome-scale nanopore concatemer sequencing. Nat Biotechnol  2022b;40:1488–99. [DOI] [PubMed] [Google Scholar]
  63. Dey KK, van de Geijn B, Kim SS  et al.  Evaluating the informativeness of deep learning annotations for human complex diseases. Nat Commun  2020;11:4703. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Ding J, Aronow BJ, Kaminski N  et al.  Reconstructing differentiation networks and their regulation from time series single-cell expression data. Genome Res  2018a;28:383–95. [DOI] [PMC free article] [PubMed] [Google Scholar]
  65. Ding J, Hagood JS, Ambalavanan N  et al.  iDREM: interactive visualization of dynamic regulatory networks. PloS Comput Biol  2018b;14:e1006019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  66. Dixit A, Parnas O, Li B  et al.  Perturb-Seq: dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell  2016;167:1853–66.e17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Dong K, Zhang S.  Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder. Nat. Commun  2022;13:1739. [DOI] [PMC free article] [PubMed] [Google Scholar]
  68. Dries R, Zhu Q, Dong R  et al.  Giotto: a toolbox for integrative analysis and visualization of spatial expression data. Genome Biol  2021;22:78. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Drusinsky S, Whalen S, Pollard KS. Deep-learning prediction of gene expression from personal genomes. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.07.27.605449, 2024, preprint: not peer reviewed. [DOI]
  70. Dubuis JO, Tkacik G, Wieschaus EF  et al.  Positional information, in bits. Proc Natl Acad Sci U S A  2013;110:16301–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Dudnyk K, Cai D, Shi C  et al.  Sequence basis of transcription initiation in the human genome. Science  2024;384:eadj0116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  72. Eraslan G, Avsec Ž, Gagneur J  et al.  Deep learning: new computational modelling techniques for genomics. Nat Rev Genet  2019;20:389–403. [DOI] [PubMed] [Google Scholar]
  73. Ernst J, Kellis M.  Discovery and characterization of chromatin states for systematic annotation of the human genome. Nat Biotechnol  2010;28:817–25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  74. Ernst J, Kellis M.  ChromHMM: automating chromatin-state discovery and characterization. Nat Methods  2012;9:215–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  75. Ernst J, Kellis M.  Interplay between chromatin state, regulator binding, and regulatory motifs in six human cell types. Genome Res  2013;23:1142–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  76. Ernst J, Kellis M.  Large-scale imputation of epigenomic datasets for systematic annotation of diverse human tissues. Nat Biotechnol  2015;33:364–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  77. Ernst J, Vainas O, Harbison CT  et al.  Reconstructing dynamic regulatory maps. Mol Syst Biol  2007;3:74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  78. Fang R, Preissl S, Li Y  et al.  Comprehensive analysis of single cell ATAC-seq data with SnapATAC. Nat Commun  2021;12:1337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  79. Fang R, Yu M, Li G  et al.  Mapping of long-range chromatin interactions by proximity ligation-assisted ChIP-seq. Cell Res  2016;26:1345–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  80. Farley EK, Olson KM, Zhang W  et al.  Syntax compensates for poor binding sites to encode tissue specificity of developmental enhancers. Proc Natl Acad Sci USA  2016;113:6508–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  81. Fleck JS, Jansen SMJ, Wollny D  et al.  Inferring and perturbing cell fate regulomes in human brain organoids. Nature  2023;621:365–72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Fornes O, Gheorghe M, Richmond PA  et al.  MANTA2, update of the mongo database for the analysis of transcription factor binding site alterations. Sci Data  2018;5:180141. [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Friedman N, Linial M, Nachman I  et al.  Using Bayesian networks to analyze expression data. J Comput Biol  2000;7:601–20. [DOI] [PubMed] [Google Scholar]
  84. Fudenberg G, Imakaev M, Lu C  et al.  Formation of chromosomal domains by loop extrusion. Cell Rep  2016;15:2038–49. [DOI] [PMC free article] [PubMed] [Google Scholar]
  85. Fudenberg G, Kelley DR, Pollard KS  et al.  Predicting 3D genome folding from DNA sequence with Akita. Nat Methods  2020;17:1111–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  86. Fulco CP, Nasser J, Jones TR  et al.  Activity-by-contact model of enhancer–promoter regulation from thousands of CRISPR perturbations. Nat Genet  2019;51:1664–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  87. Fullwood MJ, Liu MH, Pan YF  et al.  An oestrogen-receptor-alpha-bound human chromatin interactome. Nature  2009;462:58–64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  88. Gao L, Uzun Y, Gao P  et al.  Identifying noncoding risk variants using disease-relevant gene regulatory networks. Nat Commun  2018;9:702. [DOI] [PMC free article] [PubMed] [Google Scholar]
  89. Gasperini M, Tome JM, Shendure J  et al.  Towards a comprehensive catalogue of validated and target-linked human enhancers. Nat Rev Genet  2020;21:292–310. [DOI] [PMC free article] [PubMed] [Google Scholar]
  90. Glass K, Huttenhower C, Quackenbush J  et al.  Passing messages between biological networks to refine predicted interactions. PloS One  2013;8:e64832. [DOI] [PMC free article] [PubMed] [Google Scholar]
  91. Goel VY, Huseyin MK, Hansen AS  et al.  Region Capture Micro-C reveals coalescence of enhancers and promoters into nested microcompartments. Nat Genet  2023;55:1048–56. [DOI] [PMC free article] [PubMed] [Google Scholar]
  92. González AJ, Setty M, Leslie CS  et al.  Early enhancer establishment and regulatory locus complexity shape transcriptional programs in hematopoietic differentiation. Nat Genet  2015;47:1249–59. [DOI] [PMC free article] [PubMed] [Google Scholar]
  93. Greenfield A, Hafemeister C, Bonneau R  et al.  Robust data-driven incorporation of prior knowledge into the inference of dynamic regulatory networks. Bioinformatics  2013;29:1060–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  94. Grujic O, Phung TN, Kwon SB  et al.  Identification and characterization of constrained non-exonic bases lacking predictive epigenomic and transcription factor binding annotations. Nat Commun  2020;11:6168. [DOI] [PMC free article] [PubMed] [Google Scholar]
  95. Gschwind AR, Mualim KS, Karbalayghareh A et al. An encyclopedia of enhancer-gene regulatory interactions in the human genome. bioRxiv, Prepr. Serv. Biol., 10.1101/2023.11.09.563812, 2023, preprint: not peer reviewed. [DOI]
  96. Guo Y, Al-Jibury E, Garcia-Millan R  et al.  Chromatin jets define the properties of cohesin-driven in vivo loop extrusion. Mol Cell  2022;82:3769–80.e5. [DOI] [PubMed] [Google Scholar]
  97. Hamley JC, Li H, Denny N  et al.  Determining chromatin architecture with Micro Capture-C. Nat Protoc  2023;18:1687–711. [DOI] [PubMed] [Google Scholar]
  98. Harris HL, Gu H, Olshansky M  et al.  Chromatin alternates between A and B compartments at kilobase scale for subgenic organization. Nat Commun  2023;14:3303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  99. Hawkins RD, Hon GC, Ren B  et al.  Next-generation genomics: an integrative approach. Nat Rev Genet  2010;11:476–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
  100. He AY, Danko CG. Dissection of core promoter syntax through single nucleotide resolution modeling of transcription initiation. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.03.13.583868, 2024, preprint: not peer reviewed. [DOI]
  101. He Q, Johnston J, Zeitlinger J  et al.  ChIP-nexus enables improved detection of in vivo transcription factor binding footprints. Nat Biotechnol  2015;33:395–401. [DOI] [PMC free article] [PubMed] [Google Scholar]
  102. Henikoff S, Henikoff JG, Ahmad K  et al.  Simplified epigenome profiling using antibody-tethered tagmentation. Bio Protoc  2021;11:e4043. [DOI] [PMC free article] [PubMed] [Google Scholar]
  103. Hirabayashi S, Bhagat S, Matsuki Y  et al.  NET-CAGE characterizes the dynamics and topology of human transcribed cis-regulatory elements. Nat Genet  2019;51:1369–79. [DOI] [PubMed] [Google Scholar]
  104. Hoffman MM, Buske OJ, Wang J  et al.  Unsupervised pattern discovery in human chromatin structure through genomic segmentation. Nat Methods  2012;9:473–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  105. Hsieh T-HS, Weiner A, Lajoie B  et al.  Mapping nucleosome resolution chromosome folding in yeast by Micro-C. Cell  2015;162:108–19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  106. Hu J, Li X, Coleman K  et al.  SpaGCN: integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network. Nat Methods  2021;18:1342–51. [DOI] [PubMed] [Google Scholar]
  107. Hu Y, Salgado Figueroa D, Zhang Z  et al.  Lineage-specific 3D genome organization is assembled at multiple scales by IKAROS. Cell  2023;186:5269–89.e22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  108. Huang C, Shuai RW, Baokar P  et al.  Personal transcriptome variation is poorly explained by current genomic deep learning models. Nat Genet  2023;55:2056–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  109. Huang Y-F, Gulko B, Siepel A  et al.  Fast, scalable prediction of deleterious noncoding variants from functional and population genomic data. Nat Genet  2017;49:618–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  110. Huynh-Thu VA, Irrthum A, Wehenkel L  et al.  Inferring regulatory networks from expression data using tree-based methods. PloS One  2010;5:e12776. [DOI] [PMC free article] [PubMed] [Google Scholar]
  111. IGVF Consortium. Deciphering the impact of genomic variation on function. Nature  2024;633:47–57. [DOI] [PMC free article] [PubMed] [Google Scholar]
  112. Inoue F, Kircher M, Martin B  et al.  A systematic comparison reveals substantial differences in chromosomal versus episomal encoding of enhancer activity. Genome Res  2017;27:38–52. [DOI] [PMC free article] [PubMed] [Google Scholar]
  113. Janesick A, Shelansky R, Gottscho AD  et al. ; 10x Development Teams. High resolution mapping of the tumor microenvironment using integrated single-cell, spatial and in situ analysis. Nat Commun  2023;14:8353. [DOI] [PMC free article] [PubMed] [Google Scholar]
  114. Janssens J, Aibar S, Taskiran II  et al.  Decoding gene regulation in the fly brain. Nature  2022;601:630–6. [DOI] [PubMed] [Google Scholar]
  115. Javierre BM, Burren OS, Wilder SP  et al. ; BLUEPRINT Consortium. Lineage-specific genome architecture links enhancers and non-coding disease variants to target gene promoters. Cell  2016;167:1369–84.e19. [DOI] [PMC free article] [PubMed] [Google Scholar]
  116. Jha A, K Aicher J, R Gazzara M  et al.  Enhanced integrated gradients: improving interpretability of deep learning models using splicing codes as a case study. Genome Biol  2020;21:149. [DOI] [PMC free article] [PubMed] [Google Scholar]
  117. Jolma A, Yin Y, Nitta KR  et al.  DNA-dependent formation of transcription factor pairs alters their binding specificity. Nature  2015;527:384–8. [DOI] [PubMed] [Google Scholar]
  118. Kagey MH, Newman JJ, Bilodeau S  et al.  Mediator and cohesin connect gene expression and chromatin architecture. Nature  2010;467:430–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  119. Kamimoto K, Stringa B, Hoffmann CM  et al.  Dissecting cell identity via network inference and in silico gene perturbation. Nature  2023;614:742–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  120. Kaplow IM, Lawler AJ, Schäffer DE  et al. ; Zoonomia Consortium. Relating enhancer genetic variation across mammals to complex phenotypes using machine learning. Science  2023;380:eabm7993. [DOI] [PMC free article] [PubMed] [Google Scholar]
  121. Karbalayghareh A, Sahin M, Leslie CS  et al.  Chromatin interaction-aware gene regulatory modeling with graph attention networks. Genome Res  2022;32:930–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  122. Karollus A, Hingerl J, Gankin D  et al.  Species-aware DNA language models capture regulatory elements and their evolution. Genome Biol  2024;25:83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  123. Karollus A, Mauermeier T, Gagneur J  et al.  Current sequence-based models capture gene expression determinants in promoters but mostly ignore distal enhancers. Genome Biol  2023;24:56. [DOI] [PMC free article] [PubMed] [Google Scholar]
  124. Kathail P, Shuai RW, Chung R  et al.  Current genomic deep learning models display decreased performance in cell type-specific accessible regions. Genome Biol  2024;25:202. [DOI] [PMC free article] [PubMed] [Google Scholar]
  125. Keene JD.  RNA regulons: coordination of post-transcriptional events. Nat Rev Genet  2007;8:533–43. [DOI] [PubMed] [Google Scholar]
  126. Kelley DR.  Cross-species regulatory sequence activity prediction. PloS Comput Biol  2020;16:e1008050. [DOI] [PMC free article] [PubMed] [Google Scholar]
  127. Kelley DR, Reshef YA, Bileschi M  et al.  Sequential regulatory activity prediction across chromosomes with convolutional neural networks. Genome Res  2018;28:739–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  128. Kelley DR, Snoek J, Rinn JL  et al.  Basset: learning the regulatory code of the accessible genome with deep convolutional neural networks. Genome Res  2016;26:990–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  129. Khafizov R, Piazza E, Cui Y et al. Sub-cellular imaging of the entire protein-coding human transcriptome (18933-plex) on FFPE tissue using spatial molecular imaging. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.11.27.625536, 2024, preprint: not peer reviewed. [DOI]
  130. Kim DS, Risca VI, Reynolds DL  et al.  The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation. Nat Genet  2021a;53:1564–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  131. Kim HD, Shay T, O’Shea EK  et al.  Transcriptional regulatory circuits: predicting numbers from alphabets. Science  2009;325:429–32. [DOI] [PMC free article] [PubMed] [Google Scholar]
  132. Kim J, Jakobsen ST, Natarajan KN  et al.  TENET: gene network reconstruction using transfer entropy reveals key regulatory factors from single cell transcriptomic data. Nucleic Acids Res  2021b;49:e1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  133. Kim S, Wysocka J.  Deciphering the multi-scale, quantitative cis-regulatory code. Mol Cell  2023;83:373–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
  134. Koenecke N, Johnston J, He Q  et al.  Drosophila poised enhancers are generated during tissue patterning with the help of repression. Genome Res  2017;27:64–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  135. Koo PK, Majdandzic A, Ploenzke M  et al.  Global importance analysis: an interpretability method to quantify importance of genomic features in deep neural networks. PloS Comput Biol  2021;17:e1008925. [DOI] [PMC free article] [PubMed] [Google Scholar]
  136. Kraft K, Magg A, Heinrich V  et al.  Serial genomic inversions induce tissue-specific architectural stripes, gene misexpression and congenital malformations. Nat Cell Biol  2019;21:305–10. [DOI] [PubMed] [Google Scholar]
  137. Krebs AR.  Studying transcription factor function in the genome at molecular resolution. Trends Genet TIG  2021;37:798–806. [DOI] [PubMed] [Google Scholar]
  138. Kreibich E, Kleinendorst R, Barzaghi G  et al.  Single-molecule footprinting identifies context-dependent regulation of enhancers by DNA methylation. Mol Cell  2023;83:787–802.e9. [DOI] [PubMed] [Google Scholar]
  139. Krietenstein N, Abraham S, Venev SV  et al.  Ultrastructural details of mammalian chromosome architecture. Mol Cell  2020;78:554–65.e7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  140. Kuderna LFK, Ulirsch JC, Rashid S  et al.  Identification of constrained sequence elements across 239 primate genomes. Nature  2024;625:735–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  141. Kumar A, Schrader AW, Aggarwal B  et al.  Intracellular spatial transcriptomic analysis toolkit (InSTAnT). Nat Commun  2024;15:7794. [DOI] [PMC free article] [PubMed] [Google Scholar]
  142. Kumar S, Ambrosini G, Bucher P  et al.  SNP2TFBS—a database of regulatory SNPs affecting predicted transcription factor binding site affinity. Nucleic Acids Res  2017;45:D139–44. [DOI] [PMC free article] [PubMed] [Google Scholar]
  143. La Fleur A, Shi Y, Seelig G  et al.  Decoding biology with massively parallel reporter assays and machine learning. Genes Dev  2024;38:843–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  144. Lai WKM, Mariani L, Rothschild G  et al.  A ChIP-exo screen of 887 Protein Capture Reagents Program transcription factor antibodies in human cells. Genome Res  2021;31:1663–79. [DOI] [PMC free article] [PubMed] [Google Scholar]
  145. Lambert SA, Jolma A, Campitelli LF  et al.  The human transcription factors. Cell  2018;172:650–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  146. Lettice LA, Heaney SJH, Purdie LA  et al.  A long-range Shh enhancer regulates expression in the developing limb and fin and is associated with preaxial polydactyly. Hum Mol Genet  2003;12:1725–35. [DOI] [PubMed] [Google Scholar]
  147. Levine M, Davidson EH.  Gene regulatory networks for development. Proc Natl Acad Sci USA  2005;102:4936–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  148. Li B, Ritchie MD.  From GWAS to gene: transcriptome-wide association studies and other methods to functionally understand GWAS discoveries. Front Genet  2021;12:713230. [DOI] [PMC free article] [PubMed] [Google Scholar]
  149. Libbrecht MW, Chan RCW, Hoffman MM  et al.  Segmentation and genome annotation algorithms for identifying chromatin state and other genomic patterns. PloS Comput Biol  2021;17:e1009423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  150. Lieberman-Aiden E, van Berkum NL, Williams L  et al.  Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science  2009;326:289–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
  151. Linder J, Srivastava D, Yuan H  et al.  Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation. Nat Genet  2025;57:949–61. [DOI] [PMC free article] [PubMed] [Google Scholar]
  152. Liu M, Yang B, Hu M  et al.  Chromatin tracing and multiplexed imaging of nucleome architectures (MINA) and RNAs in single mammalian cells and tissue. Nat Protoc  2021;16:2667–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  153. Liu X, Li YI, Pritchard JK  et al.  Trans effects on gene expression can drive omnigenic inheritance. Cell  2019;177:1022–34.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  154. Liu Z, Chen Y, Xia Q  et al.  Linking genome structures to functions by simultaneous single-cell Hi-C and RNA-seq. Science  2023;380:1070–6. [DOI] [PubMed] [Google Scholar]
  155. Long HK, Osterwalder M, Welsh IC  et al.  Loss of extreme long-range enhancers in human neural crest drives a craniofacial disorder. Cell Stem Cell  2020;27:765–83.e14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  156. Long HK, Prescott SL, Wysocka J  et al.  Ever-changing landscapes: transcriptional enhancers in development and evolution. Cell  2016;167:1170–87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  157. Lu T, Ang CE, Zhuang X  et al.  Spatially resolved epigenomic profiling of single cells in complex tissues. Cell  2022;185:4448–64.e17. [DOI] [PMC free article] [PubMed] [Google Scholar]
  158. Lubeck E, Coskun AF, Zhiyentayev T  et al.  Single-cell in situ RNA profiling by sequential hybridization. Nat Methods  2014;11:360–1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  159. Lundberg SM, Lee S-I.  A unified approach to interpreting model predictions. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17. Red Hook, NY: Curran Associates Inc., 2017, 4768–4777. [Google Scholar]
  160. Luo Q, Yu Y, Lan X  et al.  SIGNET: single-cell RNA-seq-based gene regulatory network prediction using multiple-layer perceptron bagging. Brief Bioinform  2022;23:bbab547. [DOI] [PMC free article] [PubMed] [Google Scholar]
  161. Magri MS, Voronov D, Foley S et al. Deep conservation of cis-regulatory elements and chromatin organization in echinoderms uncover ancestral regulatory features of animal genomes. bioRxiv, 10.1101/2024.11.30.626178, 2024, preprint: not peer reviewed. [DOI]
  162. Mah CK, Ahmed N, Lopez NA  et al.  Bento: a toolkit for subcellular analysis of spatial transcriptomics data. Genome Biol  2024;25:82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  163. Mahat DB, Kwak H, Booth GT  et al.  Base-pair-resolution genome-wide mapping of active RNA polymerases using precision nuclear run-on (PRO-seq). Nat Protoc  2016;11:1455–76. [DOI] [PMC free article] [PubMed] [Google Scholar]
  164. Majdandzic A, Rajesh C, Koo PK  et al.  Correcting gradient-based interpretations of deep neural networks for genomics. Genome Biol  2023;24:109. [DOI] [PMC free article] [PubMed] [Google Scholar]
  165. Malik S, Roeder RG.  Regulation of the RNA polymerase II pre-initiation complex by its associated coactivators. Nat Rev Genet  2023;24:767–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  166. Manolio TA, Collins FS, Cox NJ  et al.  Finding the missing heritability of complex diseases. Nature  2009;461:747–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  167. Martin AR, Gignoux CR, Walters RK  et al.  Human demographic history impacts genetic risk prediction across diverse populations. Am J Hum Genet  2017;100:635–49. [DOI] [PMC free article] [PubMed] [Google Scholar]
  168. Martin V, Zhao J, Afek A  et al.  QbiC-Pred: quantitative predictions of transcription factor binding changes due to sequence variants. Nucleic Acids Res  2019;47:W127–35. [DOI] [PMC free article] [PubMed] [Google Scholar]
  169. Maslova A, Ramirez RN, Ma K  et al. ; Immunological Genome Project. Deep learning of immune cell differentiation. Proc Natl Acad Sci U S A  2020;117:25655–66. [DOI] [PMC free article] [PubMed] [Google Scholar]
  170. Mayer A, Churchman LS.  Genome-wide profiling of RNA polymerase transcription at nucleotide resolution in human cells with native elongating transcript sequencing. Nat Protoc  2016;11:813–33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  171. McCalla SG, Fotuhi Siahpirani A, Li J  et al.  Identifying strengths and weaknesses of methods for computational network inference from single-cell RNA-seq data. G3 Bethesda Md  2023;13:jkad004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  172. Meuleman W, Muratov A, Rynes E  et al.  Index and biological spectrum of human Dnase I hypersensitive sites. Nature  2020;584:244–51. [DOI] [PMC free article] [PubMed] [Google Scholar]
  173. Mifsud B, Tavares-Cadete F, Young AN  et al.  Mapping long-range promoter contacts in human cells with high-resolution capture Hi-C. Nat Genet  2015;47:598–606. [DOI] [PubMed] [Google Scholar]
  174. Minnoye L, Taskiran II, Mauduit D  et al.  Cross-species analysis of enhancer logic using deep learning. Genome Res  2020;30:1815–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  175. Miraldi ER, Pokrovskii M, Watters A  et al.  Leveraging chromatin accessibility for transcriptional regulatory network inference in T helper 17 cells. Genome Res  2019;29:449–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  176. Mitra S, Malik R, Wong W  et al.  Single-cell multi-ome regression models identify functional and disease-associated enhancers and enable chromatin potential analysis. Nat Genet  2024;56:627–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  177. Monti R, Ohler U.  Toward identification of functional sequences and variants in noncoding DNA. Annu Rev Biomed Data Sci  2023;6:191–210. [DOI] [PubMed] [Google Scholar]
  178. Moore JE, Purcaro MJ, Pratt HE  et al. ; ENCODE Project Consortium. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature  2020;583:699–710. [DOI] [PMC free article] [PubMed] [Google Scholar]
  179. Moses L, Pachter L.  Museum of spatial transcriptomics. Nat Methods  2022;19:534–46. [DOI] [PubMed] [Google Scholar]
  180. Mostafavi H, Spence JP, Naqvi S  et al.  Systematic differences in discovery of genetic effects on gene expression and complex traits. Nat Genet  2023;55:1866–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  181. Movva R, Greenside P, Marinov GK  et al.  Deciphering regulatory DNA sequences and noncoding genetic variants using neural network models of massively parallel reporter assays. PloS One  2019;14:e0218073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  182. Mumbach MR, Rubin AJ, Flynn RA  et al.  HiChIP: efficient and sensitive analysis of protein-directed genome architecture. Nat Methods  2016;13:919–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  183. Nair S, Shrikumar A, Schreiber J  et al.  fastISM: performant in silico saturation mutagenesis for convolutional neural networks. Bioinformatics  2022;38:2397–403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  184. Nakato R, Sakata T.  Methods for ChIP-seq analysis: a practical workflow and advanced applications. Methods San Diego Calif  2021;187:44–53. [DOI] [PubMed] [Google Scholar]
  185. Nasser J, Bergman DT, Fulco CP  et al.  Genome-wide enhancer maps link risk variants to disease genes. Nature  2021;593:238–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  186. Neumayr C, Pagani M, Stark A  et al.  STARR-seq and UMI-STARR-seq: assessing enhancer activities for genome-wide-, high-, and low-complexity candidate libraries. Curr Protoc Mol Biol  2019;128:e105. [DOI] [PMC free article] [PubMed] [Google Scholar]
  187. Novakovsky G, Dexter N, Libbrecht MW  et al.  Obtaining genetics insights from deep learning via explainable artificial intelligence. Nat Rev Genet  2023;24:125–37. [DOI] [PubMed] [Google Scholar]
  188. Oliveira MF, Romero JP, Chung M  et al. Characterization of immune cell populations in the tumor microenvironment of colorectal cancer using high definition spatial profiling. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.06.04.597233, 2024, preprint: not peer reviewed. [DOI]
  189. Oudelaar AM, Downes DJ, Hughes JR.  Assessment of multiway interactions with Tri-C. Methods Mol Biol  2022;2532:95–112. [DOI] [PubMed] [Google Scholar]
  190. Özel MN, Gibbs CS, Holguera I  et al.  Coordinated control of neuronal differentiation and wiring by sustained transcription factors. Science  2022;378:eadd1884. [DOI] [PMC free article] [PubMed] [Google Scholar]
  191. Pampari A, Shcherbina A, Kvon EZ. ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.12.25.630221, 2025, preprint: not peer reviewed. [DOI]
  192. Pang B, Snyder MP.  Systematic identification of silencers in human cells. Nat Genet  2020;52:254–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  193. Pang B, van Weerd JH, Hamoen FL  et al.  Identification of non-coding silencer elements and their regulation of gene expression. Nat Rev Mol Cell Biol  2023;24:383–95. [DOI] [PubMed] [Google Scholar]
  194. Papagianni A, Forés M, Shao W  et al.  Capicua controls toll/IL-1 signaling targets independently of RTK regulation. Proc Natl Acad Sci USA  2018;115:1807–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  195. Penzar D, Nogina D, Noskova E  et al.  LegNet: a best-in-class deep learning model for short DNA regulatory regions. Bioinformatics  2023;39:btad457. [DOI] [PMC free article] [PubMed] [Google Scholar]
  196. Perez AA, Goronzy IN, Blanco MR  et al.  ChIP-DIP maps binding of hundreds of proteins to DNA simultaneously and identifies diverse gene regulatory elements. Nat Genet  2024;56:2827–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  197. Petralia F, Wang P, Yang J  et al.  Integrative random forest for gene regulatory network inference. Bioinformatics  2015;31:i197–205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  198. Phillips JE, Corces VG.  CTCF: master weaver of the genome. Cell  2009;137:1194–211. [DOI] [PMC free article] [PubMed] [Google Scholar]
  199. Piecyk RS, Schlegel L, Johannes F  et al.  Predicting 3D chromatin interactions from DNA sequence using deep learning. Comput Struct Biotechnol J  2022;20:3439–48. [DOI] [PMC free article] [PubMed] [Google Scholar]
  200. Pihlajamaa P, Kauko O, Sahu B  et al.  A competitive precision CRISPR method to identify the fitness effects of transcription factor binding sites. Nat Biotechnol  2023;41:197–203. [DOI] [PMC free article] [PubMed] [Google Scholar]
  201. Piwecka M, Rajewsky N, Rybak-Wolf A  et al.  Single-cell and spatial transcriptomics: deciphering brain complexity in health and disease. Nat Rev Neurol  2023;19:346–62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  202. Pliner HA, Packer JS, McFaline-Figueroa JL  et al.  Cicero predicts cis-regulatory DNA interactions from single-cell chromatin accessibility data. Mol Cell  2018;71:858–71.e8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  203. Policastro RA, Zentner GE.  Global approaches for profiling transcription initiation. Cell Rep Methods  2021;1:100081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  204. Prakash EI, Shrikumar A, Kundaje A. Towards more realistic simulated datasets for benchmarking deep learning models in regulatory genomics. In: Proceedings of the 16th Machine Learning in Computational Biology meeting, online. PMLR, 2022, 58–77.
  205. Pratapa A, Jalihal AP, Law JN  et al.  Benchmarking algorithms for gene regulatory network inference from single-cell transcriptomic data. Nat Methods  2020;17:147–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  206. Preissl S, Gaulton KJ, Ren B  et al.  Characterizing cis-regulatory elements using single-cell epigenomics. Nat Rev Genet  2023;24:21–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
  207. Quinodoz SA, Ollikainen N, Tabak B  et al.  Higher-order inter-chromosomal hubs shape 3D genome organization in the nucleus. Cell  2018;174:744–57.e24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  208. Ramdas S, Judd J, Graham SE  et al. ; Global Lipids Genetics Consortium. A multi-layer functional genomic analysis to understand noncoding genetic variation in lipids. Am J Hum Genet  2022;109:1366–87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  209. Rao SSP, Huntley MH, Durand NC  et al.  A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell  2014;159:1665–80. [DOI] [PMC free article] [PubMed] [Google Scholar]
  210. Rappoport N, Chomsky E, Nagano T  et al.  Single cell Hi-C identifies plastic chromosome conformations underlying the gastrulation enhancer landscape. Nat Commun  2023;14:3844. [DOI] [PMC free article] [PubMed] [Google Scholar]
  211. Rauluseviciute I, Riudavets-Puig R, Blanc-Mathieu R  et al.  JASPAR 2024: 20th anniversary of the open-access database of transcription factor binding profiles. Nucleic Acids Res  2024;52:D174–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  212. Replogle JM, Saunders RA, Pogson AN  et al.  Mapping information-rich genotype–phenotype landscapes with genome-scale perturb-seq. Cell  2022;185:2559–75.e28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  213. Roadmap Epigenomics Consortium, Kundaje A, Meuleman W  et al.  Integrative analysis of 111 reference human epigenomes. Nature  2015;518:317–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  214. Rodriques SG, Stickels RR, Goeva A  et al.  Slide-seq: a scalable technology for measuring genome-wide expression at high spatial resolution. Science  2019;363:1463–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  215. Rogers MF, Shihab HA, Mort M  et al.  FATHMM-XF: accurate prediction of pathogenic point mutations via extended features. Bioinformatics  2018;34:511–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  216. Sahu B, Hartonen T, Pihlajamaa P  et al.  Sequence determinants of human gene regulatory elements. Nat Genet  2022;54:283–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  217. Sakaue S, Weinand K, Isaac S  et al.  Tissue-specific enhancer–gene maps from multimodal single-cell data identify causal disease alleles. Nat Genet  2024;56:615–26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  218. Santana-Garcia W, Rocha-Acevedo M, Ramirez-Navarro L  et al.  RSAT variation-tools: an accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding. Comput Struct Biotechnol J  2019;17:1415–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  219. Sasse A, Ng B, Spiro AE  et al.  Benchmarking of deep neural networks for predicting personal gene expression from DNA sequence highlights shortcomings. Nat Genet  2023;55:2060–4. [DOI] [PubMed] [Google Scholar]
  220. Schott M, León-Periñán D, Splendiani E  et al.  Open-ST: high-resolution spatial transcriptomics in 3D. Cell  2024;187:3953–72.e26. [DOI] [PubMed] [Google Scholar]
  221. Schraivogel D, Steinmetz LM, Parts L  et al.  Pooled genome-scale CRISPR screens in single cells. Annu Rev Genet  2023;57:223–44. [DOI] [PubMed] [Google Scholar]
  222. Schreiber J, Durham T, Bilmes J  et al.  Avocado: a multi-scale deep tensor factorization method learns a latent representation of the human epigenome. Genome Biol  2020;21:81. [DOI] [PMC free article] [PubMed] [Google Scholar]
  223. Schwalb B, Michel M, Zacher B  et al.  TT-seq maps the human transient transcriptome. Science  2016;352:1225–8. [DOI] [PubMed] [Google Scholar]
  224. Schwessinger R, Gosden M, Downes D  et al.  DeepC: predicting 3D genome folding using megabase-scale transfer learning. Nat Methods  2020;17:1118–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  225. Seçilmiş D, Hillerton T, Tjärnberg A  et al.  Knowledge of the perturbation design is essential for accurate gene regulatory network inference. Sci Rep  2022;12:16531. [DOI] [PMC free article] [PubMed] [Google Scholar]
  226. Segal E, Shapira M, Regev A  et al.  Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data. Nat Genet  2003;34:166–76. [DOI] [PubMed] [Google Scholar]
  227. Segert JA, Gisselbrecht SS, Bulyk ML  et al.  Transcriptional silencers: driving gene expression with the brakes on. Trends Genet  2021;37:514–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  228. Seitz EE, McCandlish DM, Kinney JB  et al.  Interpreting cis-regulatory mechanisms from genomic deep neural networks using surrogate models. Nat Mach Intell  2024;6:701–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  229. Sekelja M, Paulsen J, Collas P  et al.  4D nucleomes in single cells: what can computational modeling reveal about spatial chromatin conformation?  Genome Biol  2016;17:54. [DOI] [PMC free article] [PubMed] [Google Scholar]
  230. Setty M, Leslie CS.  SeqGL identifies context-dependent binding signals in genome-wide regulatory element maps. PloS Comput Biol  2015;11:e1004271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  231. Shema E, Bernstein BE, Buenrostro JD  et al.  Single-cell and single-molecule epigenomics to uncover genome regulation at unprecedented resolution. Nat Genet  2019;51:19–25. [DOI] [PubMed] [Google Scholar]
  232. Sherwood RI, Hashimoto T, O’Donnell CW  et al.  Discovery of directional and nondirectional pioneer transcription factors by modeling Dnase profile magnitude and shape. Nat Biotechnol  2014;32:171–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  233. Shrikumar A, Greenside P, Kundaje A. Learning important features through propagating activation differences. In: Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia. PMLR, 2017, 3145–53.
  234. Shrikumar A, Tian K, Avsec Ž  et al. Technical note on transcription factor motif discovery from importance scores (TF-MoDISco) version 0.5.6.5. arXiv Prepr. arXiv: 1811.00416v5, 2020, preprint: not peer reviewed.
  235. Shu H, Zhou J, Lian Q  et al.  Modeling gene regulatory networks using neural network architectures. Nat Comput Sci  2021;1:491–501. [DOI] [PubMed] [Google Scholar]
  236. Siahpirani AF, Knaack S, Chasman D  et al.  Dynamic regulatory module networks for inference of cell type-specific transcriptional networks. Genome Res  2022;32:1367–84. [DOI] [PMC free article] [PubMed] [Google Scholar]
  237. Siahpirani AF, Roy S.  A prior-based integrative framework for functional transcriptional regulatory network inference. Nucleic Acids Res  2017;45:e21. [DOI] [PMC free article] [PubMed] [Google Scholar]
  238. Simon E, Swanson K, Zou J  et al.  Language models for biological research: a primer. Nat Methods  2024;21:1422–9. [DOI] [PubMed] [Google Scholar]
  239. Skok Gibbs C, Jackson CA, Saldi G-A  et al.  High-performance single-cell gene regulatory network inference at scale: the inferelator 3.0. Bioinformatics  2022;38:2519–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
  240. Slattery M, Zhou T, Yang L  et al.  Absence of a simple code: how transcription factors read the genome. Trends Biochem Sci  2014;39:381–99. [DOI] [PMC free article] [PubMed] [Google Scholar]
  241. Smith GD, Ching WH, Cornejo-Páramo P  et al.  Decoding enhancer complexity with machine learning and high-throughput discovery. Genome Biol  2023;24:116. [DOI] [PMC free article] [PubMed] [Google Scholar]
  242. Sokolova K, Chen KM, Hao Y  et al.  Deep learning sequence models for transcriptional regulation. Annu Rev Genomics Hum Genet  2024;25:105–22. [DOI] [PubMed] [Google Scholar]
  243. Song L, Crawford GE.  Dnase-seq: a high-resolution technique for mapping active gene regulatory elements across the genome from mammalian cells. Cold Spring Harb Protoc  2010;2010:pdb.prot5384. [DOI] [PMC free article] [PubMed] [Google Scholar]
  244. Sönmezer C, Kleinendorst R, Imanci D  et al.  Molecular co-occupancy identifies transcription factor binding cooperativity in vivo. Mol Cell  2021;81:255–67.e6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  245. Srivastava D, Aydin B, Mazzoni EO  et al.  An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced transcription factor binding. Genome Biol  2021;22:20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  246. Stunnenberg HG, Hirst M; International Human Epigenome Consortium. The international human epigenome consortium: a blueprint for scientific collaboration and discovery. Cell  2016;167:1145–9. [DOI] [PubMed] [Google Scholar]
  247. Su J-H, Zheng P, Kinrot SS  et al.  Genome-scale imaging of the 3D organization and transcriptional activity of chromatin. Cell  2020;182:1641–59.e26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  248. Sundararajan M, Taly A, Yan Q. Axiomatic attribution for deep networks. In: Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia. PMLR, 2017, 3319–28.
  249. Svensson V, Teichmann SA, Stegle O  et al.  SpatialDE: identification of spatially variable genes. Nat Methods  2018;15:343–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  250. Szałata A, Hrovatin K, Becker S  et al.  Transformers in single-cell omics: a review and new perspectives. Nat Methods  2024;21:1430–43. [DOI] [PubMed] [Google Scholar]
  251. Tan J, Shenker-Tauris N, Rodriguez-Hernaez J  et al.  Cell-type-specific prediction of 3D chromatin organization enables high-throughput in silico genetic screening. Nat Biotechnol  2023;41:1140–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  252. Tanevski J, Flores ROR, Gabor A  et al.  Explainable multiview framework for dissecting spatial relationships from highly multiplexed data. Genome Biol  2022;23:97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  253. Tang Z, Somia N, Yu Y, et al. Evaluating the representational power of pre-trained DNA language models for regulatory genomics. bioRxiv, Prepr. Serv. Biol., 10.1101/2024.02.29.582810, 2024, preprint: not peer reviewed. [DOI]
  254. Tang Z, Toneyan S, Koo PK  et al.  Current approaches to genomic deep learning struggle to fully capture human genetic variation. Nat Genet  2023;55:2021–2. [DOI] [PubMed] [Google Scholar]
  255. Tarashansky AJ, Musser JM, Khariton M  et al.  Mapping single-cell atlases throughout metazoa unravels cell type evolution. Elife  2021;10:e66747. [DOI] [PMC free article] [PubMed] [Google Scholar]
  256. Taskiran II, Spanier KI, Dickmänken H  et al.  Cell-type-directed design of synthetic enhancers. Nature  2024;626:212–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  257. Tavares-Cadete F, Norouzi D, Dekker B  et al.  Multi-contact 3C reveals that the human genome during interphase is largely not entangled. Nat Struct Mol Biol  2020;27:1105–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  258. Taylor DJ, Chhetri SB, Tassia MG  et al.  Sources of gene expression variation in a globally diverse human cohort. Nature  2024;632:122–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  259. Thurman RE, Rynes E, Humbert R  et al.  The accessible chromatin landscape of the human genome. Nature  2012;489:75–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  260. Toneyan S, Koo PK.  Interpreting cis-regulatory interactions from large-scale deep neural networks. Nat Genet  2024;56:2517–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  261. Toneyan S, Tang Z, Koo PK  et al.  Evaluating deep learning for predicting epigenomic profiles. Nat Mach Intell  2022;4:1088–100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  262. Townes FW, Engelhardt BE.  Nonnegative spatial factorization applied to spatial genomics. Nat Methods  2023;20:229–38. [DOI] [PMC free article] [PubMed] [Google Scholar]
  263. Trevino AE, Müller F, Andersen J  et al.  Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution. Cell  2021;184:5053–69.e23. [DOI] [PubMed] [Google Scholar]
  264. Tullius TW, Isaac RS, Dubocanin D  et al.  RNA polymerases reshape chromatin architecture and couple transcription on individual fibers. Mol Cell  2024;84:3209–22.e5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  265. Tycko J, Van MV, DelRosso N  et al.  Development of compact transcriptional effectors using high-throughput measurements in diverse contexts. Nat Biotechnol  2024;Epub ahead of print. [DOI] [PMC free article] [PubMed] [Google Scholar]
  266. Vian L, Pękowska A, Rao SSP  et al.  The energetics and physiological impact of cohesin extrusion. Cell  2018;173:1165–78.e20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  267. Vierstra J, Lazar J, Sandstrom R  et al.  Global reference mapping of human transcription factor footprints. Nature  2020;583:729–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  268. Visscher PM, Brown MA, McCarthy MI  et al.  Five years of GWAS discovery. Am J Hum Genet  2012;90:7–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  269. Voong LN, Xi L, Wang J-P  et al.  Genome-wide mapping of the nucleosome landscape by micrococcal nuclease and chemical mapping. Trends Genet  2017;33:495–507. [DOI] [PMC free article] [PubMed] [Google Scholar]
  270. Vu H, Ernst J.  Universal annotation of the human genome through integration of over a thousand epigenomic datasets. Genome Biol  2022;23:9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  271. Wang H, Fu T, Du Y  et al.  Scientific discovery in the age of artificial intelligence. Nature  2023a;620:47–60. [DOI] [PubMed] [Google Scholar]
  272. Wang L, Trasanidis N, Wu T  et al.  Dictys: dynamic gene regulatory network dissects developmental continuum with single-cell multiomics. Nat Methods  2023b;20:1368–78. [DOI] [PubMed] [Google Scholar]
  273. Wang Y, Cho D-Y, Lee H  et al.  Reprogramming of regulatory network using expression uncovers sex-specific gene regulation in Drosophila. Nat Commun  2018;9:4061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  274. Wen Y, Huang J, Guo S  et al.  Applying causal discovery to single-cell analyses using CausalCell. Elife  2023;12:e81464. [DOI] [PMC free article] [PubMed] [Google Scholar]
  275. Wu Y, Qi T, Wray NR  et al.  Joint analysis of GWAS and multi-omics QTL summary statistics reveals a large fraction of GWAS signals shared with molecular phenotypes. Cell Genom  2023;3:100344. [DOI] [PMC free article] [PubMed] [Google Scholar]
  276. Xia C, Fan J, Emanuel G  et al.  Spatial transcriptome profiling by MERFISH reveals subcellular RNA compartmentalization and cell cycle-dependent gene expression. Proc Natl Acad Sci U S A  2019;116:19490–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  277. Yang Y, Pe’er D.  REUNION: transcription factor binding prediction and regulatory association inference from single-cell multi-omics data. Bioinformatics  2024;40:i567–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  278. Yao D, Tycko J, Oh JW  et al.  Multicenter integrated analysis of noncoding CRISPRi screens. Nat Methods  2024;21:723–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  279. Yoon S, Chandra A, Vahedi G  et al.  Stripenn detects architectural stripes from chromatin conformation data using computer vision. Nat Commun  2022;13:1602. [DOI] [PMC free article] [PubMed] [Google Scholar]
  280. Yuan Y, Bar-Joseph Z.  GCNG: graph convolutional networks for inferring gene interaction from spatial transcriptomics data. Genome Biol  2020;21:300. [DOI] [PMC free article] [PubMed] [Google Scholar]
  281. Zabidi MA, Stark A.  Regulatory enhancer–core-promoter communication via transcription factors and cofactors. Trends Genet  2016;32:801–14. [DOI] [PMC free article] [PubMed] [Google Scholar]
  282. Zeitlinger J.  Seven myths of how transcription factors read the cis-regulatory code. Curr Opin Syst Biol  2020;23:22–31. [DOI] [PMC free article] [PubMed] [Google Scholar]
  283. Zhang S, He Y, Liu H  et al.  regBase: whole genome base-wise aggregation and functional prediction for human non-coding regulatory variants. Nucleic Acids Res  2019;47:e134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  284. Zhang S, Pyne S, Pietrzak S  et al.  Inference of cell type-specific gene regulatory networks on cell lineages from single cell omic datasets. Nat Commun  2023;14:3064. [DOI] [PMC free article] [PubMed] [Google Scholar]
  285. Zhang Y, Boninsegna L, Yang M  et al.  Computational methods for analyzing multiscale 3D genome organization. Nat Rev Genet  2024;25:123–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
  286. Zhao BS, Roundtree IA, He C  et al.  Post-transcriptional gene regulation by mRNA modifications. Nat Rev Mol Cell Biol  2017;18:31–42. [DOI] [PMC free article] [PubMed] [Google Scholar]
  287. Zhou J.  Sequence-based modeling of three-dimensional genome architecture from kilobase to chromosome scale. Nat Genet  2022;54:725–34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  288. Zhou J, Theesfeld CL, Yao K  et al.  Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk. Nat Genet  2018;50:1171–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  289. Zhou J, Troyanskaya OG.  Predicting effects of noncoding variants with deep learning-based sequence model. Nat Methods  2015;12:931–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  290. Zhou JL, Guruvayurappan K, Toneyan S  et al.  Analysis of single-cell CRISPR perturbations indicates that enhancers predominantly act multiplicatively. Cell Genom  2024a;4:100672. [DOI] [PMC free article] [PubMed] [Google Scholar]
  291. Zhou T, Zhang R, Jia D  et al.  GAGE-seq concurrently profiles multiscale 3D genome organization and gene expression in single cells. Nat Genet  2024b;56:1701–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  292. Zhou VW, Goren A, Bernstein BE  et al.  Charting histone modifications and the functional organization of mammalian genomes. Nat Rev Genet  2011;12:7–18. [DOI] [PubMed] [Google Scholar]
  293. Zhu X, Huang L, Wang C  et al.  Uncovering the whole genome silencers of human cells via Ss-STARR-seq. Nat Commun  2025;16:723. [DOI] [PMC free article] [PubMed] [Google Scholar]
  294. Zuin J, Roth G, Zhan Y  et al.  Nonlinear control of transcription through enhancer–promoter interactions. Nature  2022;604:571–7. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

As this is a review article there is no associated data.


Articles from Bioinformatics Advances are provided here courtesy of Oxford University Press

RESOURCES