Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 May 13.
Published in final edited form as: Annu Rev Genomics Hum Genet. 2026 May 11;27(1):159–182. doi: 10.1146/annurev-genom-120324-123022

Synthetic Regulatory Genomics

Matthew T Maurano 1
PMCID: PMC13166085  NIHMSID: NIHMS2149839  PMID: 42113869

Abstract

The genomics era has yielded high-quality genome assemblies, comprehensive atlases of biochemical signatures of gene regulation, and genetic associations for thousands of common human diseases and traits. These dramatic advances in observational approaches have not been matched by perturbational genetic tools to facilitate direct and systematic hypothesis testing. Enabled by advances in DNA synthesis and assembly, genome engineering tools, and genomic readouts, synthetic regulatory genomics now promises access to a new scale of genomic manipulation to study the function of cohesive genomic units. Synthetic regulatory genomics is distinguished by the breadth of the genetic manipulations and their divergence from the reference sequence. These new tools enable an expanded focus to encompass sufficiency in addition to necessity and to enable a new era of perturbation analysis.

Keywords: Gene regulation, Genome engineering, Perturbation, Synthetic regulatory genomics

INTRODUCTION

The consensus from genome-wide association studies (GWAS) is that the overwhelming majority of common disease- and trait-associated variants localize within noncoding regions of the human genome (77, 128). Connecting these genetic sequence variants with specific downstream biological functions promises to illuminate the mechanistic basis of human disease and quantitative phenotypes, facilitate disease risk assessment on both an individual and population levels, enable mechanistic cellular and organismal disease models, and guide the development of new therapies.

Realizing these goals requires solving the genomic sequence-to-function problem: to predict the effects of genomic sequence changes on gene regulation and to tie disease association loci to their target genes. But decades of study have shown that understanding the function of regulatory sequence requires experimental study at the individual locus. The genomics era has thus opened a gap between observational and perturbational approaches, as sequence analysis and biochemical feature profiling are more inherently scalable genome-wide. To address unique interpretive and experimental challenges that cannot be surmounted by the generation and analysis of DNA sequence data alone, I motivate the case for synthetic regulatory genomics representing a new perturbational approach to the analysis of regulatory architecture on a genomic scale.

First and foremost, synthetic regulatory genomics is inherently perturbational, distinguishing itself from the focus on sequence analysis and profiling that has characterized the genomic era so far. Conceptually, large-scale perturbational studies have a long history (86). But their application in mammalian systems has been constrained by the difficulty of engineering the relevant alleles at scale. Synthetic regulatory genomics specifically permits a focus on sufficiency, beyond the focus on necessity conferred by conventional genome editing. Further, as synthetic genomics renders large swaths of genomic sequence portable across species, the function of genomic sequence can be interrogated in the most appropriate model system.

This shift in focus to perturbational studies is enabled by a series of key technological developments, covered in more detail throughout this Review:

  1. Genome sequences and regulatory genomic maps

  2. DNA sources including synthetic oligonucleotides and BAC libraries

  3. Synthetic genomics cloning tools, high-capacity vectors, and strategies for assembly and manipulation in bacteria and yeast

  4. Genomic targeting schemes, specifically cassette exchange and programmable nuclease-mediated delivery

  5. Targeted and whole-genome sequencing for verification and readout

This Review introduces and motivates synthetic regulatory genomics, contextualizing it within key related technologies (Fig. 1) and identifying key motivating gaps in present approaches. I introduce key problems synthetic regulatory genomics is poised to address, its initial applications, practical technical details, and outline potential future developments.

Fig. 1.

Fig. 1.

Engineering genomic perturbations at scale. Schematic characterizing genome engineering approaches along two axes. Degree of control refers to the length of sequence which can be controlled, the control over the sequence (random vs. designed), and control over genomic location.

Relation to Deep Mutational Scanning

Synthetic regulatory genomics has many parallels to deep mutational scanning, which uses high-throughput approaches to characterize protein function (36). Accordingly, synthetic regulatory genomics can learn from the path blazed to incorporate functional classification of protein-coding variation into clinical practice (78). It is therefore worthwhile to outline key conceptual similarities and differences between the two fields. First, protein-coding effects are intrinsically local in that they first affect a linear sequence of amino acids. This dramatically constrains the experimental scope, even when accounting for the potential role of transcript structure or secondary effects of altered protein function. In contrast, promoter-distal noncoding sequence intrinsically affects the function of distant elements, and multiple regulatory elements combine in complex fashion to regulate a target gene. Identification of local effects on nearby regulatory element(s) is only a prerequisite to ultimately determining the effect on a target gene.

Thus protein function is intrinsically decoupled from genomic context, which can be abstracted away as a factor that ultimately affects protein levels. Consequently, deep mutational scanning platforms focus less on editing endogenous loci or avoiding genomic scars which might affect the regulatory landscape. Indeed, the simplest mutational scan can be performed using a cDNA construct, perhaps even in another species such as yeast. Understanding the function of the protein in context hews closer to the area of synthetic regulatory genomics, requiring consideration of organismal expression patterns and accounting for transcriptional splicing.

Despite these differences, both fields involve a high-throughput genomics approach to the sequence-to-function problem. Both fields are inherently experimental, and must grapple with the total inadequacy of current methods to saturate the relevant sequence space, whether it be all possible protein-coding mutations or all nucleotide changes in the human genome. This leads to a common pressure to impose structure on the vast search space by sharing information among related classes. For example, protein assays can target general phenotypes such as thermodynamic stability (4), or assays tuned to a broad class of proteins. Synthetic regulatory genomics will have to chart a similar path through sequence space, identifying classes of regulatory architecture and characteristic benchmark functions which enable sharing information genome-wide.

EXISTING TECHNOLOGIES FOR STUDYING GENOME REGULATION

Reporter Assays

Reporter assays have been the stalwart of high-throughput functional genomic analysis. The most basic approaches to assess transcriptional activity in vivo employ simplified transcriptional reporter constructs. These truncated reporter constructs typically fuse a single enhancer at a short distance to a heterologous minimal promoter. High-throughput sequencing and programmable oligonucleotide synthesis have enabled modern extensions of workhorse techniques – termed multiplexed reporter assays (MPRA) – to survey thousands of variants in enhancer and promoter sequences.

Reporter plasmids are most simply employed through transient transfection, without integrating into the genome of the host cell. However, the short time before the non-replicating plasmid is diluted out (24–72 hours) limits the ability to address many phenotypes. Transient transfection also results in uncontrolled copy number and variable activity (80). Furthermore, transiently transfected plasmids do not develop canonical chromatin structure (71, 111). As chromatin in eukaryotes is not merely a passive obstacle to transcription but instead a major regulator of individual transcription factor (TF) binding (134), regulatory function strongly depends on the modification state of surrounding chromatin (5, 79). Reporter assays have shown substantial expression differences between episomal and integrated reporters (53). Furthermore, the function of regulatory sequence depends critically on its native genomic context in vivo, and endogenous regulation relies upon specific interactions with multiple distal elements finely tuned for control of a specific promoter sequence. The requirement for cooperative TF occupancy imposes numerous local context dependencies within an individual regulatory site, including the specific combination of TFs present and their spacing and orientation (119). Thus, while transient reporter assays can be applied at great scale, they struggle to recapitulate many key aspects of endogenous regulation.

In contrast to transient assays, reporters can be integrated randomly into the genome whereby the reporter assumes a chromatin conformation dependent on its target locus. But with random integration, the reporter is not directed to any particular genomic location; the reporter often integrates in multiple copies (sometimes in tandem), and the result may include flanking viral or bacterial sequence that induces silencing. Increased efficiency and some control over random integration can be obtained with transposases like Sleeping Beauty (104) or PiggyBac (2, 59). Indicative of the complex orchestration of long-range communication underlying cellular patterns of DNA accessibility, the transgenesis of artificial constructs results in highly variable expression patterns depending on insertion site, termed position-effect variegation (91, 133).

Single-copy integration into a defined locus enables reproducible functional assessment of reporter activity. This approach has been significantly facilitated by integration into double-strand breaks created by site-specific zinc finger, TALE, and CRISPR–Cas nucleases (24, 43). However, these approaches still rely on a shortened reporter which may eliminate essential regulatory properties of the full endogenous locus (or create artifactual ones). Interactions between regulatory elements and target genes may be influenced by interposing genes (97) or regulatory elements (19). Indeed, there are numerous examples of regulatory elements interacting with cognate target gene(s) located at distances of tens to hundreds of kilobases away (61, 122). Thus functional investigation out-of-context risks investigating properties unique to the reporter model rather than of the endogenous locus, underscoring the difficulty of modeling these factors ectopically.

Genome Editing of Endogenous Loci

Genome editing offers the potential to perform functional assessment in a native context. A major bottleneck to the manipulation of endogenous loci in mammalian cells remains the inefficiency of transfection and of homologous recombination in many primary cells. Designer nucleases including TALENs and CRISPR–Cas have dramatically reduced the difficulty of studying regulatory elements in their endogenous context in mammalian cells. Cleavage with these nucleases is repaired efficiently by non-homologous end joining (NHEJ), resulting in short insertions or deletions at targeted loci. The exact product of repair is not precisely defined, other than by location, and thus this approach is best-suited to disruption of existing regulatory sequence. By coupling a nuclease targeting a given regulatory site with a synthetic library used as a template, homology-directed repair can be induced to introduce controlled changes in multiplex, albeit at much reduced efficiency (32). Similarly, base editing (67) and prime editing (20, 48, 74) provide an increased degree of control over the sequence edits generated and have been adapted to screening modalities.

Genome editing can be multiplexed for assessment at many loci simultaneously using a screening approach that employs pooled guide RNA libraries to engineer a genetically heterogeneous population of cells in culture (27, 110, 130). Screening has been used to couple targeted generation of regulatory mutation with functional selection approaches to dissect regulatory elements (16, 57, 101, 107, 127). A typical experimental design uses indirect readout and sequences a proxy barcode rather than the edits themselves, meaning that editing efficiency or break resolution cannot be directly assessed (6).

An alternate strategy uses epigenetic editing. Fusion of endonuclease-deficient CRISPR–Cas or TALE targeting domains to transcriptional activator or repressor domains enables precise alteration of locus-specific histone modifications and TF binding for gain-of-function studies (42, 79). However, this approach has a rather low resolution limited by the activity of the activator or repressor domains themselves and is subject to similar constraints on precisely verifying the location and potency of epigenetic modifications.

Thus, genome editing has dramatically reduced the difficulty of tailored alteration of endogenous loci and regulatory elements, enabling assessment of changes at single sites in cell culture or even whole model organisms. This approach is well-suited to analysis of individual sites, but does not scale to multiple functional elements at the same site. In addition to the inherent difficulty of simultaneously targeting multiple sites, genome editing operates in trans and is difficult to apply in an allelic fashion and thus unsuited for generating multiple simultaneous changes on a single haplotype.

SYNTHETIC REGULATORY GENOMICS

Synthetic regulatory genomics broadly encompasses the manipulation of genomic sequence at large length scales and with great flexibility. The economic limits on DNA synthesis and technical challenges of delivering these payloads into cells impose a series of tradeoffs that inform design of a successful synthetic regulatory genomics experiment. Here I give an overview of the technical aspects of genome rewriting (Fig. 2).

Fig. 2.

Fig. 2.

Schematic of DNA assembly, mESC engineering, verification, and analysis pipelines. Adapted from (11).

Large Construct Engineering

The ability to engineer large lengths of DNA underpins synthetic regulatory genomics. Large constructs spanning hundreds of kilobases such as bacterial or yeast artificial chromosomes (BAC or YAC) can encompass entire loci including all distal regulatory elements. These approaches have been used, for example, to enable precise recapitulation of phenotypes in mouse models such as the developmental switch from fetal to adult hemoglobin (95) and modeling disease-associated genetic variation (96). The efficiency of homologous recombination in yeast permits extensive payload engineering. However, the requisite technical sophistication has restricted their application when compared to the rapid spread of nuclease-based genome editing. In addition, targeted delivery to mammalian cells has historically been a challenge

The availability of megabase-scale DNA synthesis and assembly technologies (54) promises to enable large-scale generation and analysis of locus-scale constructs incorporating arbitrary sequence changes. But while the cost of DNA synthesis continues to decrease, the cost of fully synthetic payloads would substantially limit accessibility of synthetic regulatory genomics. A fully synthetic approach theoretically enables complete design freedom (15, 68), but most designs include substantial sequence which can be sourced from DNA available more efficiently through other means. Thus PCR can be used to generate kilobase-scale amplicons at the cost of oligonucleotide primers, reaction reagents, and thermocycler plates. Starting from a BAC representing the region of interest as the template substantially streamlines amplification and primer design compared to starting from genomic DNA. However, genomic BAC libraries are a legacy resource used less intensively than in the past, and fulfilment of individual orders may be at risk, potentially leaving a gap until the DNA synthesis cost further drops.

Both synthetic oligonucleotides and PCR amplicons are limited in length, prompting development of different strategies to assemble these kilobase-scale fragments into longer molecules (63, 140). In vitro assembly approaches have a key role, given their simplicity and scalability to generation of multiplexed libraries. Gibson assembly is a flexible method permitting assembly of long and complex sequences (41). Golden Gate assembly was employed in widely used protocols for cloning TALE nucleases that assemble dozens of 102-nt repeat modules in a single-step reaction (106). Golden Gate assembly can be a high-efficiency route to generation of large libraries, but it imposes design constraints: both the absence of internal Type IIS restriction sites and the requirement for short linker sequences between fragments. Both approaches exhibit reduced efficiency as the number of fragments to be assembled scales up. Bacterial recombineering uses the “Red” (recombination-deficient) system of the λ bacteriophage, to introduce simple modifications in bacterially hosted vectors (14, 47, 85). The yeast S. cerevisiae permits flexible assembly of many shorter segments based on short overlaps (40, 140). Once in yeast, payloads can be easily modified using homologous recombination or CRISPR–Cas nucleases (141). Additionally, delivery to mammalian cells generally requires a separate step recovering DNA from yeast into bacteria, which makes it cumbersome to generate complex pooled libraries.

In summary, there are a wide variety of approaches to assembling large DNA constructs which present different tradeoffs in technical complexity and scalability. Different assembly approaches may perform better for certain applications, for example in differing tolerance for repetitive sequences. But assembly in general offers significantly increased control compared to a genome editing approach. Furthermore, a major advantage of the genome rewriting approach is that, given the cheap and rapid availability of both short- and long-read sequencing, payloads can easily be sequence verified prior to delivery.

Targeted Genomic Delivery Schemes

Delivery of DNA to the cell nucleus and its integration into the genome is a major challenge (34). A key distinction is whether delivery is targeted to the endogenous locus, ectopically to a safe-harbor locus, or randomly (untargeted). Targeting the endogenous locus leaves intact the surrounding regulatory context, which may permit simplification of the assembly and delivery strategy by rewriting a smaller genomic region. Ectopic integration can enable testing sufficiency of a given payload, or as a comparison point to identify interactions with regulatory elements at the endogenous locus (90, 98).

Recombinase-mediated cassette exchange (RMCE) is the most broadly-used targeting approach to deliver large payloads to mammalian genomes. By flanking a payload with recombinase sites, the payload can be targeted to a genomic region marked by the same tandem recombinase sites (10, 109). The two recombinase sites must be orthogonal, in that their sequence diverges sufficiently to avoid undesired resolution as a deletion or inversion. While RMCE minimally requires tandem recombinase sites, usually a more complex landing pad is also installed to facilitate delivery and selection. A landing pad can be targeted using short homology arms in conjunction with CRISPR–Cas cut sites flanking the desired integration site. It is straightforward to select for integration of the landing pad through an included positive selection marker. An advantage of the landing pad strategy is that further deliveries to the landing pad line require no double-stranded breaks, and many distinct payloads can be compared in a controlled setting.

This landing pad strategy can employ a wide variety of selection schemes, based on the classic positive/negative selection paradigm for gene targeting (72). In the simplest case, the payload can include an intact expression cassette for a positive selection marker. However, this does not specifically select for correct integration and may favor off-target integration or transient presence. Inducible cassette exchange (ICE) uses a gene trap which confers resistance only upon successful integration (52) (Table 1). Both types of positive selection methods share the disadvantage of requiring an ectopic transcription unit, which may disrupt the function of the locus under study. A second class of methods uses negative selection to screen for loss of the landing pad. For example, a Big-IN landing pad expresses thymidine kinase to confer ganciclovir resistance or PIGA to confer proaerolysin resistance (12) (Fig. 3, Table 1). In this case, upon selection for landing pad loss (and thus payload delivery), the only scars are the two short recombinase sites. A disadvantage is that counterselection does not specifically select for correct integration, so there may be background from landing pad silencing or genomic deletions that include the landing pad. A variety of other approaches implement similar concepts, including RMGR (129) and DICE (144).

Table 1.

Features of selected big-DNA delivery methods.

Delivery Method Mechanism Scar Selection scheme Nuclease exposure Iterative? Delivery targeting Repeated deliveries
ICE Recombinase (RMCE) Transcription units Positive selection for correct integrant Once, during LP installation No Single (landing pad) Yes
Big-IN Recombinase (RMCE) 2 flanking recombinase sites (e.g., 34 bp for Cre-lox) Negative selection against LP retention Once, during LP installation No Single landing pad Yes
mSwAP-In Nuclease Landing pad (removeable with extra step) Positive/negative selection At each step No Multiple sites or alleles Possible

Fig. 3.

Fig. 3.

Big-IN delivery scheme using the example of the Sox2 LCR. The BL6 allele of the 41-kb Sox2 LCR (right) was replaced with a landing pad in mESCs. Landing pad integration was aided by CRISPR/Cas9 using a pair of gRNAs targeting both the replaced allele and the landing pad plasmid and by short homology arms that facilitate homology-directed repair. Landing pad mESCs were selected with puromycin, while ganciclovir (GCV) selects against integration of the landing pad vector backbone (BB). Cre recombinase-mediated cassette exchange (RMCE) enabled the replacement of each landing pad with a series of payloads. Transfected cells were transiently selected with blasticidin, followed by counterselection of landing pad mESCs with proaerolysin. Adapted from (11).

The requisite recombinase can be expressed by the landing pad, but generally this requires control by an inducible promoter (such as in the ICE method) to avoid constant recombinase expression which may be toxic to cells. More simply, the recombinase may be expressed through transient transfection which is still efficient and generally nontoxic. This creates risk that some cells could retain the recombinase plasmid, which can be attenuated through targeted genotyping or sequencing.

A second class of methods exemplified by mSwAP-IN (139) (Table 1) and UKiS (87) integrate large DNAs using nuclease-targeted double-strand breaks and rely on endogenous repair machinery for delivery. Similar to RMCE, a landing pad cassette is first installed to enable positive/negative selection. This approach additionally permits iterative delivery rounds. The key to iterative re-writing lies in using two marker cassettes expressing orthogonal positive and negative selective markers. A marker cassette remains at the end of the last rewriting step, but it can optionally be removed in a final step facilitated by counterselection. In contrast to recombinase-based schemes, each round requires expression of a CRISPR–Cas nuclease, which may lead to off-target edits or other undesired genomic events.

Recombinase-based methods must necessarily be used on a single allele in a diploid genome to avoid unwanted recombination between multiple sites. This requirement must be considered during landing pad installation, and can be addressed through an allele-specific installation (say with a single-nucleotide variant [SNV] overlapping a CRISPR–Cas Protospacer Adjacent Motif [PAM] site), or by screening to identify a heterozygous clone. mSwAP-IN is more flexible, as specificity is conferred exclusively through CRISPR–Cas targeting, and thus can generate biallelically engineered cells.

While current delivery approaches have supported a wide variety of productive studies, they impose several major limitations on scaling up synthetic regulatory genomics. First, per-cell delivery efficiency is low, necessitating selection approaches that increase the requisite scale of cell culture and require careful attention to the fidelity of the genomic outcome. Second, application to a new model system involves significant effort to implement landing pad installation, positive/negative selection, and cloning from single cells than is required for simpler and more broadly available genome engineering systems based on CRISPR–Cas. Third, the requirement for an intermediate landing pad to install exogenous recombinase sites and enable positive/negative selection adds a step and constrains programmable integration at high multiplicity of infection. Expanding beyond these constraints within existing systems is discussed in the “Scaling Cellular Readouts” section.

Programmable integration methods could help simplify this multistage delivery process and increase flexibility for multiplexing across loci. TwinPE and PASSIGE (3, 92) and PASTE (136) use CRISPR–Cas prime editing to insert recombinase sites for subsequent recombinase-mediated delivery. Engineered recombinases could offer increased efficiency and directly-programmable specificity (30). STITCHR uses a retrotransposase targeted through fusion with a CRISPR–Cas9 nickase (31). RNA-directed bridge recombinases have also been shown to support large integration (94). These programmable methods could offer increased flexibility for landing pad installation, and potentially for direct integration of much larger payloads.

More broadly, future experiments will have to address in more detail how the sequence (58, 66) and chromatin status of the transfected DNA payloads (108) influence the establishment of chromatin state once integrated chromosomally. In particular, it remains an open question whether temporary landing pad installation and selection cassette expression significantly affects the locus function following delivery.

Verification of Genome Rewriting Outcomes

Genome rewriting involves multiple DNA components including the landing pad, a large complex payload DNA, and plasmids required for delivery and the subsequent selection strategy. Each component adds to the risk of unexpected outcomes. Thus careful consideration must be paid to evaluate the efficiency of the delivery and selection approach, and to avoid misinterpreting results due to incomplete characterization. All verification approaches are targeted in some way, whether experimentally and/or informatically, so the aphorism, “You get what you look for” applies in that analysis requires careful consideration of possible events. In this context, perfect verification is generally unachievable and thus there is an economic tradeoff between detection of on- and off-target events and addressing other sources of error and irreproducibility.

There are three critical classes of events or signatures to evaluate with particular attention: (i) confirmation of integration junctions which confirm correct delivery; (ii) exclusion of undesired sequence, e.g. of landing pads or transiently transfected plasmids; (iii) confirmation of sequence presence, e.g. to rule out internal deletions in long payloads or escape from counterselection through genomic deletion.

PCR genotyping can serve as a flexible and efficient basis for screening for correct clones. The basic techniques and equipment are widely available. Sequence verification provides a richer data set enabling detailed verification, and important information for troubleshooting, but requires consideration of a broader set of implementation details.

In this context, whole-genome sequencing can be a useful option for initial work. Despite the cost, it requires little sample processing, the requisite experience and equipment are broadly available, and local and turnkey sequencing services are available to trade off turnaround time with cost. Furthermore, whole-genome sequencing can be used to comprehensively characterize the starting cells, such as parental and/or landing pad cell lines. Sequencing using Oxford Nanopore can employ targeted adaptive sampling to gain a modest (~4x) enrichment of the target region.

While PCR amplicon sequencing is used heavily for verifying the results of genome editing, it is less useful for large-scale genome rewriting. In conjunction with long-read sequencing, it can yield high-depth targeted sequencing of regions up to 5–10 kb. However, genomic PCR often requires optimization. Furthermore, it requires a more complex multiplex strategy to target regions that cannot be amplified as a single amplicon.

Short-read targeted sequencing using in-solution hybridization capture (Capture-seq) provides the highest throughput and lowest cost per-sample. Sequencing libraries are enriched through hybridization to biotinylated bait and retained through washes using magnetic beads. As hybridization capture is broadly deployed for exome resequencing, there is a broad base of experience with the technique. One specific consideration is the generation of bait required for targeting. One option is to order commercial custom synthetic oligonucleotides, like those used for exome capture. However, this approach can be expensive with significant turnaround time, which can interfere with rapid technological development and unseen circumstances. A second option is to biotinylate locally available DNA using nick translation (12, 137). This straightforward protocol enables perfectly matching the bait to the engineered cells and payloads, maximizing the value of the resequencing. It is advisable to use multiple baits, including from a genomic BAC to enrich for reads surrounding the engineering site, from the payload and vector backbone, and from any other plasmids transfected into cells, such as landing pad, recombinase, or CRISPR–Cas. Finally, it is valuable to include the parental line as a control.

Sequence verification must distinguish between the engineered and any other allele(s) present. Sometimes, two alleles in a diploid genome can be distinguished by the presence of SNVs or short indels. The rate of informative variation can be high in some circumstances, such as F1 hybrids of diverse mouse subspecies (55, 112). At a rate of ~1 SNV per 100 bp, BL6xCAST (M. Musculus Domesticus C57BL/6 × M. Musculus Castaneus) hybrid genomes permit mapping the majority of short reads in an allele-specific manner. In human populations, the rate of heterozygous sites is low enough that most short reads are not expected to overlap a SNV and will thus be uninformative as to their allelic origin. But the likelihood of a read being informative is proportional to its length, so long reads (5–10 kb) result in a high rate of allelic identification using natural human polymorphism. The presence of informative variation can be estimated in broad strokes ahead of time through population-level sequencing, but the ideal is to sequence the target cells before starting. In contrast, allele-specific mapping is more straightforward when delivering DNA from another species. For example, for the humanization of a locus in mouse cells, sequence divergence is high with the partial exception of exonic sequence, enabling confident assignment of reads originating from the payload. Alternatively, selected variation can be engineered into the payload to facilitate verification, but this requires careful judgement and some luck to ensure the marker variants are indeed neutral.

Long-read sequencing using the Pacific Biosciences or Oxford Nanopore technologies naturally complements genome rewriting approaches and can simplify some aspects of sequence verification. In particular, long-read assembly can reduce the need for alignment to a reference and provide contiguity through regions difficult to assemble from short reads. This can be worthwhile purely to simplify analysis. A major limitation is the limited availability of targeted enrichment options compatible with long-read platforms. CRISPR–Cas size-selection approaches provide a more modest enrichment relative to hybridization capture, are not easy to multiplex for multiple samples, and trade off significant wet lab involvement for a modest reduction in sequencing costs. Oxford Nanopore adaptive sampling is simple to implement, but yields even more limited enrichment. Thus long-read sequencing can play a key role for initial engineering or small numbers of deliveries, but its cost scales poorly with multiple deliveries.

While standard reference genome mapping and de novo assembly pipelines provide significant immediate value, there is significant added value in applying custom analyses to identify events poorly represented by the reference genome. We have found the most value in parallel mapping both to a mammalian reference genome (e.g. hg38 or mm10), and custom references representing the payload, landing pad, and plasmids. This enables both integration with other annotations anchored to the reference genome, as well as coverage analysis at expected and unexpected sequences which may not appear in the reference genome. Furthermore, standard variant callers can detect mismatches to the custom reference which indicate assembly or delivery problems. Analysis of reads mapped in parallel to two references can be used to identify paired-end reads that support junctions between those sequences. For example, the tool bamintersect can identify genome-payload junctions from reads mapped to the reference genome and payload reference, using (12).

An alternate approach involves creating a custom reference for the payload integrated at a genomic site. This simplifies mapping and early analysis but implies that each payload/delivery site combination has a unique coordinate system which is also distinct from the reference genome.

For either reference approach, SNVs and short indels can be identified using standard tools like bcftools (22) or GATK (126), and larger structural variation using tools like DELLY (102).

In the long run, development of graph-based representations for pangenome analysis will likely provide significant utility for more flexible analysis of sequencing data in the context of large-scale genome engineering.

Finally, a typical workflow will involve repeated use of a restricted sequence set in DNA cloning, assembly, and preparation, high-sensitivity detection approaches such as PCR and targeted next-generation sequencing, and genomic readouts. These circumstances create a high risk for contamination which can confound sequencing data analysis. Thus utmost stringency is required to compartmentalize lab operations to minimize contamination, while remaining aware of the possibility during data analysis.

Generation of Diversity In Vivo

While the ability to rewrite large segments of genomic DNA enables dramatically new experimental designs, discovery experiments are likely to focus rapidly on a small subset of functional elements within that interval. For follow-up experiments focusing on a small portion of the original locus, re-assembly and delivery of large (>100 kb) DNA segments can represent a significant burden. Thus subsequent generation of diversity in vivo has significant appeal as a scalable strategy. The most basic strategy is to use orthogonal recombinase sites to permit a second RMCE delivery at a functional element of interest. Or, installation of multiple landing pads which use orthogonal recombinases can permit modifying multiple sites on the same allele. A limit here is the availability of orthogonal recombinase sites, and the increased complexity from managing delivery of multiple recombinases and payload compatibility. Delivering payloads to a pool of cells with landing pads installed at different sites, each at one copy per cell, permits assessment of the effect of context on a library of functional elements (73).

More broadly, analysis of locus architecture benefits from systematic follow-up of DNA sequence features like distance, spacing, or position. Transposases like Sleeping Beauty (76) or piggyBac (13) can be used to isolate genomic context as a variable by integrating a reporter into tens of thousands of sites which are then mapped and expression profiled in multiplex, such as using thousands of reporters integrated in parallel (TRIP) (2, 50, 104). Transposases can also be used to generate diversity in position surrounding a launch point (29, 145).

Larger genomic perturbations can be generated with a Synthetic Chromosome Rearrangement and Modification by LoxP-mediated Evolution (SCRaMbLE) strategy that uses prime editing or PiggyBac to integrate cassettes containing loxPsym sites to catalyze subsequent rearrangement (56, 99).

OPEN CHALLENGES IN GENE REGULATION

The ENCODE data have shown that regulatory elements are distributed throughout the genome, with only 5% at promoters, and the remainder evenly distributed between intragenic and intergenic regions (82). Understanding the full significance of distal regulatory variation will require connecting local activities to molecular effects more proximal to phenotype. But connecting sequence variants that affect TF occupancy with a downstream molecular function, such as alterations in gene expression, remains a challenge. Thus a key question is how distal regulatory elements interact specifically with the promoters of certain genes.

Genomic Locus Architecture

A major unresolved question is how multiple enhancers and other elements combine to regulate gene expression at a genomic locus covering approximately 100 kb–1 Mb. Distal regulatory elements rely on diverse mechanisms to confer specificity to their interaction with nearby promoters (125). There is some evidence that enhancers interact preferentially with proximal genes (25, 97). Sequence features of particular enhancers may also specify preferential activation of promoters bearing compatible sequence features (64, 75, 81, 88, 138). Vertebrate enhancer activation appears to occur primarily in cis, and thus enhancer activity can also be modulated by the activity of nearby or intervening elements (18, 19, 33, 44). Yet none of these features is sufficiently precise to enable de novo prediction of target genes from genomic sequence.

On a larger length scale, the clustering of functionally related genes along a chromosome results in domains of similar regulatory features (115). Imaging studies have long depicted organization of chromatin into distinct territories, suggesting that enhancer-promoter communication is locally compartmentalized to some extent (60). Chromatin conformation studies using 5C and Hi-C methods have uncovered large-scale organization into megabase-scale topologically associating domains (TADs) (26). The interior of these domains corresponds to condensed regions of chromatin while the borders are delineated by transcriptionally active genes and accessible regulatory elements (28, 123). But these mapping studies do not address the effect of these domains on regulatory activity within. Depletion of the major architectural protein CTCF disrupts genomic architecture but shows little effect on gene expression. Evidence from high-throughput surveys of randomly integrated constructs suggests that the effect of enhancers is compartmentalized to some extent by distance (2, 118, 145), but it remains a challenge to predict whether other genomic features constrain enhancer-promoter communication. This represents a key gap for synthetic regulatory genomics to address.

Enhancer Interaction

A related question is how nearby enhancers affect the activities of their neighbors. Multiple variants within individual GWAS loci overlap DHSs and may each exert an independent effect on the expression of their target gene(s) (17, 21), but it is unclear how multiple enhancers at a given locus interact to affect gene expression. On one side of the spectrum, the effect of multiple enhancers on gene expression could be the sum of their individual spatiotemporal activities (69), such as even-skipped stripe 2 expression in the Drosophila melanogaster embryo (113). Alternatively, nearby regulatory elements controlling a given gene may provide redundancy rather than strict additivity (51). Yet this view is challenged by the failure of individual enhancer elements to fully recapitulate developmental expression in integrated constructs (91). Alternatively, regulatory elements could be tightly tuned to the function of their neighbors. Locus control regions (LCRs) were identified as DNA elements exerting dominant control on downstream gene expression regardless of chromosomal context (62). The modern term ‘super enhancer’ describes an extended region of regulatory elements frequently found surrounding key cell-identity and TF genes (49, 93, 100). Recent work has shed light on the extent to which these enhancer clusters are more than the sum of their parts.

Application of Synthetic Regulatory Genomics to Dissect Enhancer Function

Synthetic regulatory genomics has been used to investigate the function of the enhancer clusters at several key loci. The Sox2 locus control region (LCR) is located over 100 kb downstream from Sox2. The LCR is required for Sox2 expression in mouse embryonic stem cells (mESCs) (65, 142), and comprises 10 DNaseI hypersensitive sites (DHSs) and CTCF sites active in mESCs covering 41 kb. Brosh et al. investigated whether these individual DHSs function independently or are influenced by their surrounding DHSs. They delivered minimal payloads containing subsets of LCR DHSs and measured activity by allele-specific Sox2 qRT-PCR (11). They found that a minimal payload containing only the four core LCR DHSs recapitulated 88% of the activity of the full wild-type LCR. Activities of single DHSs varied substantially: while DHSs 24 and 26 restored 30% and 14% of Sox2 expression, respectively, DHS 23 was completely inactive on its own. But when DHS 23 was delivered with either DHSs 24 or 26, it augmented the activity of each individual enhancer nearly twofold. Investigation of DHSs 19 and 20 showed neither has activity alone but both can synergize to activate Sox2 when linked to DHS 26. Similar results have been identified at the Fgf5 (121), Hba (8), and NMU (143) loci.

Offering a complementary example of a locus with multiple genes, the Igf2/H19 locus is the paradigmatic example of imprinted expression. Allele-specific binding of CTCF at the imprinting control region (ICR) directs the activity of a distal enhancer cluster to Igf2 or H19 (7, 46). The locus includes: downstream CTCF sites; the H19 enhancer cluster comprising 11 DHSs; H19; the ICR, Igf2; and upstream CTCF sites. Thus this locus affords an opportunity to understand how locus architecture influences enhancer selectivity for these and other important developmentally regulated genes.

Ordoñez et al. installed a Big-IN landing pad in mESCs from an F1 hybrid BL6xCAST cross to replace the 157-kb Igf2/H19 locus sufficient to recapitulate imprinted expression ectopically (1, 90). Engineering of single-nucleotide variants in Igf2 enabled allele-specific quantification of expression from the engineered and wild-type alleles of both Igf2 and H19 using qRT-PCR. mESCs were differentiated into mesoendoderm and analyzed for expression of the engineered H19 and Igf2 alleles. Taking advantage of the synthetic system to recapitulate the standard boundary model of the Igf2/H19 locus (7, 46), they engineered genetic control of CTCF binding at the ICR, whose presence restricts enhancer activity to H19 and whose absence boosts Igf2 expression.

Surprisingly, this analysis at the Igf2/H19 locus showed that even payloads lacking the H19 enhancer cluster demonstrated residual expression of both genes (90), suggesting that the known H19 enhancers are not the only regulatory elements contributing to expression. This suggests that Igf2/H19 expression is also modulated by a previously unknown contribution of regulatory elements outside the 157-kb engineered region. These interactions spanning hundreds of kilobases highlight the power of our approach to identify context sensitivity by contrasting activity at the endogenous locus with that at defined ectopic sites.

To contrast the function of these two model enhancer clusters, Ordoñez et al. swapped the Sox2 LCR with the Igf2/H19 enhancers in different genomic contexts. Surprisingly, Sox2 expression remained negligible even after differentiation to mesendodermal cells. Even delivering the full Igf2/H19 locus, but with Sox2 and its promoter replacing Igf2 and its promoter failed to be expressed at the Sox2 locus. This showed the H19 enhancers were completely unable to function in the context at the Sox2 locus. In contrast to their behavior at the Sox2 locus, the H19 enhancers were able to activate Sox2 expression when in their native location. This suggests that the H19 enhancer cluster relies on surrounding sequence at its native locus, or that other elements surrounding the Sox2 locus suppress the activity of ectopic enhancer clusters. These results underscore the importance of experimental systems for systematic dissection of genomic context.

Ordoñez and colleagues expanded upon this showing that multiple pairs of enhancer DHSs from the Prdm14, Nanog, and Sall1 loci can synergize with partner DHSs in different patterns (89). They further used a synthetic approach to show that synergy depended on the spacing between DHSs. They used putatively neutral sequence generated by reversing but not complementing the human HPRT1 gene (15) as a substrate to analyze DHS spacing. By cloning a restriction digest of these payloads between pairs of DHSs, they found that Sox2 expression decreased monotonically as the spacer distance between DHSs increased to ~4 kb. This synthetic approach to studying enhancer-enhancer spacing fits with several similar studies focused on enhancer-promoter spacing at different length scales. Delivery of enhancers to landing pads installed at +75 kb, +25 kb, and +1.5 kb from the Nanog promoter showed synergy among DHSs (120). Wu engineered a dual landing pad system to enable independent delivery of promoter and enhancer pairs separated by variable spacing (ranging from 3–50 kb) at a housekeeping locus (135). Similarly, activation of the Sox2 promoter by its LCR is sensitive to the distance between them (45, 145). This work highlights the utility of systematic perturbations of genomic parameters like enhancer and promoter distance and spacing.

DISSECTION OF GWAS LOCI

The most pressing application of the noncoding sequence-to-function problem is in human genetics. GWAS have implicated a vast swath of noncoding loci in an impressive variety of human diseases and traits, promising to unlock the genetic underpinnings of common human diseases and traits. However, few of these loci have been followed up with detailed functional investigation (77, 84, 128). A key prerequisite for investigation of an association locus is a clear accounting of the responsible gene(s). But credible sets for most association loci still include dozens of variants even after fine-mapping (39). It remains unclear how many functional variants exist at a given disease locus, and often a single causal variant is assumed for simplicity. Further, there is evidence that multiple variants at a given locus may all exert an effect on gene expression (21). The potential for long-distance gene regulation further expands the range of implicated genes. Thus the success of GWAS has at the same time highlighted our lack of understanding of regulatory genome function, and a genomic approach for identification of target genes is a major unsolved problem (35).

Synthetic regulatory genomics naturally offers the potential to dissect intact GWAS loci, directly interrogating haplotypes that extend for tens or hundreds of kilobases from their target genes. But the low effect sizes of typical GWAS loci and uncertainty about the sensitivity of the cellular model and readout (see the “Translating Molecular Effect Sizes into Traits and Phenotypes” section) raise the risk of making major technological investments only to result in an underpowered analysis. Rather than studying human polymorphism, engineered alleles may offer a more promising initial strategy. Similar to the use of molecular phenotypes and quantitative traits as intermediates to the organismal traits under study or the use of mouse knockouts, studying larger engineered perturbations can reduce uncertainty about the basic locus architecture, model system, and readout. Once these basic parameters are outlined, the precise effect of the human haplotypes can then be worked out or validated in detail.

Understanding the regulatory landscape at a GWAS locus is likely more important than the exact causal variants, and a genomic understanding of the effect of variation on gene expression may additionally illuminate disease mechanism or potential therapeutic targets. Broad analysis of regulatory architecture at disease-implicated loci rather than precise recapitulation of associated variation can reveal novel regulatory function that may be of higher interest than common alleles. For example, functional dissection of the BCL11A locus identified through common variants in GWAS prompted a therapeutic strategy for disruption of an erythroid-specific enhancer (16, 127).

In addition to target gene(s), noncoding variation systematically implicates upstream regulatory factors. Understanding the regulatory network context is particularly important as buffering and homeostasis are universal features of physiologic regulation. A striking example is MyoD, the expression of which is sufficient in culture to convert fibroblasts into myoblasts, yet its mouse knockout has no phenotype as its function is buffered in vivo by Myf5 (132). Knockout organisms may demonstrate relevant phenotypes for unrelated reasons (114), and single knockout experiments can provide only part of the picture (83). Moreover, the thousands of low-effect size variants uncovered by GWAS present an acute challenge to the reductionist knockout paradigm. Thus systems regulatory genomics has a major role to play by permitting more sophisticated genetic analysis of the regulatory architecture of target genes, and their network context in terms of upstream and downstream regulators. Such gene network analysis offers the potential to infer more easily druggable members of the same pathway as well as the direction of effect for perturbation at alternate points in regulatory pathways.

Readout Approaches

The effects of perturbing locus architecture might manifest functionally in a variety of ways, including by impacting the magnitude of gene expression, tissue specificity, timing during development and differentiation, or response to environmental stimuli (including pharmacologic agents). Characterizing the exact consequences to gene expression in an appropriate model may provide insight into otherwise inaccessible aspects of disease etiology. Engineered alleles could potentially be analyzed for a wide variety of molecular effects, including at the chromatin, RNA, or protein levels. For example, enhancers are thought to function by increasing the frequency of cells in a population undergoing transcription (131). Assays performed on bulk cells detect the effects of regulatory variation as changes in population averages. It is now possible to analyze a variety of genomic traits on the single-cell level, which may help minimize the confounding effect of heterogeneous cell populations or incompletely penetrant variants. Gene expression itself is a broad term capturing both dynamic and steady-state transcript levels, which are each regulated at multiple levels, including initiation, elongation, and transcript stability, and there are a variety of post-transcriptional influences on the levels and localization of a protein product. In all cases, detection sensitivity is inherently tied to both the inherent properties of the locus as well as technical characteristics of the readout modality. Focusing on the most proximate molecular phenotypes like chromatin effects and nascent transcription increases detection sensitivity of large-scale approaches by reducing the influence of downstream regulation.

A major challenge to analyzing engineered alleles remains consideration of non-local effects. While it is straightforward to assess the effect of a variant on local TF binding and chromatin features, target genes of distal regulatory elements cannot yet be reliably predicted a priori. Thus, assessment across a full genomic locus can increase confidence, and genome-wide analysis may be necessary to fully contextualize the functional consequences of regulatory variation.

Translating Molecular Effect Sizes into Traits and Phenotypes

This Review has stressed the importance of studying regulatory elements in their appropriate context. While the focus thus far has been on genomic context, cellular and organismal context are just as important to consider. Echoing the ease of identifying conserved genes across model organisms following the widespread availability of genomic sequence, synthetic regulatory genomics makes DNA sequence portable across species and further unifies model organism and human genetics. This portability will facilitate experimental investigation of phenotypes in animal models where human cell culture models are insufficient or inappropriate (139). Fully realizing the potential of synthetic regulatory genomics will require a clear-headed assessment of the tradeoffs of working in different cellular or species contexts.

Despite genomic tools capable of profiling tens of thousands of molecular traits (for example, DNA accessibility or gene expression), how effect sizes at the molecular level relate to organismal phenotype remains poorly understood. It is not presently possible to predict whether alterations at the molecular level will affect organismal phenotype. Even small changes in transcript expression may in fact represent a significant biological effect size in a developmental context. Importantly, biological relevance must be distinguished from the technical sensitivity to detect minute alterations to TF activity or gene expression.

Despite these uncertainties, all phenotypes are ultimately underpinned by molecular effects. With this in mind, analysis of molecular effects promises increased power to investigate the underpinnings of traits of interest in an analogous fashion to that for association studies for quantitative traits intermediate to molecular traits and disease states (124). Thus, genome-scale approaches to study functional variation on a biochemical level underpin systems-level analysis of genome function and therefore are critical to further study of the genetics of important diseases and traits.

FUTURE CONSIDERATIONS

Scaling Cellular Readouts

Scale is a critical consideration in the genomic age. Used appropriately, high-throughput experiments are not merely technological tours de force or substitutes for thoughtful logic, but rather a strategic escape from the constraints of our limited datasets, which would otherwise prevent generalization. High-throughput perturbational experiments are thus an essential component of a predictive understanding of genome function. Perhaps the biggest near-term experimental limitation in scale comes from the clonal isolation required for arrayed genotyping and phenotyping. Thus multiplexing payloads within pooled cell culture is a natural path for increasing scale. RMCE schemes inherently support delivery of complex pooled payload libraries (70, 89, 143). Below I discuss the challenges in genotyping and phenotyping these pooled deliveries.

Fluorescence provides a straightforward pooled readout strategy. A selective marker can be engineered into target gene(s) using CRISPR–Cas. Alternatively, a two-stage big-DNA delivery approach using orthogonal recombinase sites can enable engineering fluorescence while minimizing the size of the payload pool to interrogate. Alternatively, FlowFISH can be used to provide fluorescent readout without requiring genome engineering (38, 74, 103). This is particularly advantageous for multiple target genes which may lie far from the delivery site. In these cases, the genotype of surviving cells can be read out by sequencing the pool to quantify prevalence of a barcode uniquely present in each individual payload. Flow cytometry of a fluorescent readout enables sorting into multiple bins, which can be used to obtain higher precision effect size estimates (23).

Long-read single-molecule footprinting technologies such as Fiber-seq (116) provide an attractive readout modality given that the same sequencing read includes both genotype and phenotype. The challenge is that methylation readout is inherently incompatible with PCR amplification, and targeted Fiber-seq using CRISPR–Cas enrichment (9) does not presently yield the high coverage depth required for phenotyping multiplex pools. Cytidine deaminase approaches like DAF-seq (117) are PCR-compatible and thus offer orders of magnitude more headroom for high-depth pooled readout.

The delivery method has a key role in facilitating multiplexing. Brosh et al. reported a 49% failure rate at the genotyping stage (11). This failure rate is notably lower for shorter payloads (<10%). In a system that relies on counterselection to avoid scars after delivery like Big-IN, landing pad loss through lox-lox site recombination or larger deletions comprises a significant proportion of these failed clones. These failures can often be recognized as such, but may waste significant experimental capacity in a pooled setting. In contrast, mis-delivered payloads including internal deletions, duplications, or ectopically delivered payloads can mislead the analysis and may be difficult to recognize without sequence verification. Sequence verification of the library before delivery is straightforward, but for pooled deliveries, it can only establish an upper limit on error rate. Thus, future optimizations in delivery, including alternate recombinases and landing pad configurations could prove as key enablers of large-scale readout.

Extrapolation, Prediction, and Modeling

The search space represented by mammalian genomes is too large to explore exhaustively with perturbational methods. Even scalable genomic methods can only scratch the surface of potential genomic sequence across relevant cell types and states. Addressing the genomic-wide scale of the sequence-to-function problem will require efficient exploration of the most informative sequence space, and modeling approaches which retain predictive value beyond their training sets. Active learning has been used in conjunction with MPRAs to traverse enhancer sequence space (37). Models such as Enformer which predict genomic regulatory tracks from sequence can be improved by inclusion of synthetic regulatory genomics data (105). Future studies could combine highly multiplexed reporter assays which may inform on basic TF-DNA interaction with in-context genome engineering that reflects on genome architecture. Technologies permitting multiplexed delivery and readout and generation of further diversity in vivo are naturally suited to exploring the sequence space in ways that will be informative for modeling.

CONCLUSIONS

Technology to survey human sequence variation has dramatically increased in scale throughout the genomics era, along with concomitant increases in the breadth of traditional biochemical assays. This rapid progress in observational methods has opened a gap with perturbational approaches, leaving us without complementary high-throughput methods to study regulatory variation at endogenous loci. This Review has focused on a new technological approach called synthetic regulatory genomics which aims to fill this gap. Realizing this potential will require coordinated efforts by both the human genetics and gene regulation communities. As cell and animal models are siloed by disease areas, filling this gap will require fostering collaboration between disease-focused researchers and genomic technology experts. Finally, the ability to engineer complex vectors with sophisticated expression profiles could improve efficacy and safety for gene therapy and cell therapies.

ACKNOWLEDGMENTS

M.T.M. was partly supported by US National Institutes of Health grants RM1HG009491 and R01MH136353 and the Irma T. Hirschl/Monique Weill-Caulier Research Award.

Footnotes

DISCLOSURE STATEMENT

M.T.M. is listed as an inventor on a patent application describing Big-IN.

LITERATURE CITED

  • 1.Ainscough JF, Koide T, Tada M, Barton S, Surani MA. 1997. Imprinting of Igf2 and H19 from a 130 kb YAC transgene. Development. 124(18):3621–32 [DOI] [PubMed] [Google Scholar]
  • 2.Akhtar W, de Jong J, Pindyurin AV, Pagie L, Meuleman W, et al. 2013. Chromatin position effects assayed by thousands of reporters integrated in parallel. Cell. 154(4):914–27 [DOI] [PubMed] [Google Scholar]
  • 3.Anzalone AV, Gao XD, Podracky CJ, Nelson AT, Koblan LW, et al. 2022. Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing. Nat Biotechnol. 40(5):731–40 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Araya CL, Fowler DM, Chen W, Muniez I, Kelly JW, Fields S. 2012. A fundamental protein property, thermodynamic stability, revealed solely from large-scale measurements of protein function. Proc Natl Acad Sci U S A. 109(42):16858–63 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Archer TK, Lefebvre P, Wolford RG, Hager GL. 1992. Transcription factor loading on the MMTV promoter: a bimodal mechanism for promoter activation. Science. 255(5051):1573–76 [DOI] [PubMed] [Google Scholar]
  • 6.Becerra B, Wittibschlager S, Patel ZM, Kutschat AP, Delano J, et al. 2024. CRISPR-CLEAR: Nucleotide-Resolution Mapping of Regulatory Elements via Allelic Readout of Tiled Base Editing. bioRxiv. 2024.09.09.612085 [Google Scholar]
  • 7.Bell AC, Felsenfeld G. 2000. Methylation of a CTCF-dependent boundary controls imprinted expression of the Igf2 gene. Nature. 405(6785):482–85 [DOI] [PubMed] [Google Scholar]
  • 8.Blayney JW, Francis H, Rampasekova A, Camellato B, Mitchell L, et al. 2023. Super-enhancers include classical enhancers and facilitators to fully activate gene expression. Cell. 186(26):5826–5839.e18 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Bohaczuk SC, Amador ZJ, Li C, Mallory BJ, Swanson EG, et al. 2024. Resolving the chromatin impact of mosaic variants with targeted Fiber-seq. Genome Res. 34(12):2269–78 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Bouhassira EE, Westerman K, Leboulch P. 1997. Transcriptional behavior of LCR enhancer elements integrated at the same chromosomal locus by recombinase-mediated cassette exchange. Blood. 90(9):3332–44 [PubMed] [Google Scholar]
  • 11.Brosh R, Coelho C, Ribeiro-Dos-Santos AM, Ellis G, Hogan MS, et al. 2023. Synthetic regulatory genomics uncovers enhancer context dependence at the Sox2 locus. Mol Cell. 83(7):1140–1152.e7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Brosh R, Laurent JM, Ordoñez R, Huang E, Hogan MS, et al. 2021. A versatile platform for locus-scale genome rewriting and verification. Proc Natl Acad Sci U S A. 118(10):e2023952118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Cadiñanos J, Bradley A. 2007. Generation of an inducible and optimized piggyBac transposon system. Nucleic Acids Res. 35(12):e87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Caldwell BJ, Bell CE. 2019. Structure and mechanism of the Red recombination system of bacteriophage λ. Prog Biophys Mol Biol. 147:33–46 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Camellato BR, Brosh R, Ashe HJ, Maurano MT, Boeke JD. 2024. Synthetic reversed sequences reveal default genomic states. Nature. 628(8007):373–80 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Canver MC, Smith EC, Sher F, Pinello L, Sanjana NE, et al. 2015. BCL11A enhancer dissection by Cas9-mediated in situ saturating mutagenesis. Nature. 527(7577):192–97 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Chatterjee S, Kapoor A, Akiyama JA, Auer DR, Lee D, et al. 2016. Enhancer Variants Synergistically Drive Dysfunction of a Gene Regulatory Network In Hirschsprung Disease. Cell. 167(2):355–368.e10 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Choi OR, Engel JD. 1988. Developmental regulation of beta-globin gene switching. Cell. 55(1):17–26 [DOI] [PubMed] [Google Scholar]
  • 19.Chung JH, Whiteley M, Felsenfeld G. 1993. A 5’ element of the chicken beta-globin domain serves as an insulator in human erythroid cells and protects against position effect in Drosophila. Cell. 74(3):505–14 [DOI] [PubMed] [Google Scholar]
  • 20.Cirincione A, Simpson D, Yan W, McNulty R, Ravisankar P, et al. 2025. A benchmarked, high-efficiency prime editing platform for multiplexed dropout screening. Nat Methods. 22(1):92–101 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Corradin O, Saiakhova A, Akhtar-Zaidi B, Myeroff L, Willis J, et al. 2014. Combinatorial effects of multiple enhancer variants in linkage disequilibrium dictate levels of gene expression to confer susceptibility to common traits. Genome Research. 24(1):1–13 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, et al. 2021. Twelve years of SAMtools and BCFtools. Gigascience. 10(2):giab008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.de Boer CG, Ray JP, Hacohen N, Regev A. 2020. MAUDE: inferring expression changes in sorting-based CRISPR screens. Genome Biol. 21(1):1–16 [Google Scholar]
  • 24.Dickel DE, Zhu Y, Nord AS, Wylie JN, Akiyama JA, et al. 2014. Function-based identification of mammalian enhancers using site-specific integration. Nature Methods. 11(5):566–71 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Dillon N, Trimborn T, Strouboulis J, Fraser P, Grosveld F. 1997. The effect of distance on long-range chromatin interactions. Molecular Cell. 1(1):131–39 [DOI] [PubMed] [Google Scholar]
  • 26.Dixon JR, Selvaraj S, Yue F, Kim A, Li Y, et al. 2012. Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature. 485(7398):376–80 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Doench JG. 2018. Am I ready for CRISPR? A user’s guide to genetic screens. Nat Rev Genet. 19(2):67–80 [DOI] [PubMed] [Google Scholar]
  • 28.Eagen KP, Hartl TA, Kornberg RD. 2015. Stable Chromosome Condensation Revealed by Chromosome Conformation Capture. Cell. 163(4):934–46 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Eder M, Moene CJI, Dauban L, Magnitov M, Drayton J, et al. 2025. Functional maps of a genomic locus reveal confinement of an enhancer by its target gene. Science. 389(6766):eads6552. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Fanton A, Bartie LJ, Martins JQ, Tran VQ, Goudy L, et al. 2025. Site-specific DNA insertion into the human genome with engineered recombinases. Nat Biotechnol [Google Scholar]
  • 31.Fell CW, Villiger L, Lim J, Hiraizumi M, Tagliaferri D, et al. 2025. Reprogramming site-specific retrotransposon activity to new DNA sites. Nature. 642(8069):1080–89 [DOI] [PubMed] [Google Scholar]
  • 32.Findlay GM, Boyle EA, Hause RJ, Klein JC, Shendure J. 2014. Saturation editing of genomic regions by multiplex homology-directed repair. Nature. 513(7516):120–23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Foley KP, Engel JD. 1992. Individual stage selector element mutations lead to reciprocal changes in beta- vs. epsilon-globin gene transcription: genetic confirmation of promoter competition during globin gene switching. Genes & Development. 6(5):730–44 [DOI] [PubMed] [Google Scholar]
  • 34.Fong JHC, Ceroni F. 2025. Transgene integration in mammalian cells: The tools, the challenges, and the future. Cell Syst. 16(12):101426. [DOI] [PubMed] [Google Scholar]
  • 35.Forgetta V, Jiang L, Vulpescu NA, Hogan MS, Chen S, et al. 2022. An effector index to predict target genes at GWAS loci. Hum Genet. 141(8):1431–47 [DOI] [PubMed] [Google Scholar]
  • 36.Fowler DM, Fields S. 2014. Deep mutational scanning: a new style of protein science. Nature Methods. 11(8):801–7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Friedman RZ, Ramu A, Lichtarge S, Wu Y, Tripp L, et al. 2025. Active learning of enhancers and silencers in the developing neural retina. Cell Syst. 16(1):101163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Fulco CP, Munschauer M, Anyoha R, Munson G, Grossman SR, et al. 2016. Systematic mapping of functional enhancer-promoter connections with CRISPR interference. Science. 354(6313):769–73 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Gaulton KJ, Ferreira T, Lee Y, Raimondo A, Mägi R, et al. 2015. Genetic fine mapping and genomic annotation defines causal mechanisms at type 2 diabetes susceptibility loci. Nat Genet. 47(12):1415–25 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Gibson DG, Benders GA, Axelrod KC, Zaveri J, Algire MA, et al. 2008. One-step assembly in yeast of 25 overlapping DNA fragments to form a complete synthetic Mycoplasma genitalium genome. Proc Natl Acad Sci U S A. 105(51):20404–9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Gibson DG, Young L, Chuang R-Y, Venter JC, Hutchison CA, Smith HO. 2009. Enzymatic assembly of DNA molecules up to several hundred kilobases. Nat Methods. 6(5):343–45 [DOI] [PubMed] [Google Scholar]
  • 42.Gilbert LA, Horlbeck MA, Adamson B, Villalta JE, Chen Y, et al. 2014. Genome-Scale CRISPR-Mediated Control of Gene Repression and Activation. Cell. 159(3):647–61 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Gisselbrecht SS, Barrera LA, Porsch M, Aboukhalil A, Estep PW, et al. 2013. Highly parallel assays of tissue-specific enhancers in whole Drosophila embryos. Nature Methods. 10(8):774–80 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Gray S, Szymanski P, Levine M. 1994. Short-range repression permits multiple enhancers to function autonomously within a complex promoter. Genes & Development. 8(15):1829–38 [DOI] [PubMed] [Google Scholar]
  • 45.Hansen KL, Adachi AS, Braccioli L, Kadvani S, Boileau RM, et al. 2025. Synergy between regulatory elements can render cohesin dispensable for distal enhancer function. Science, p. eadt4221 [Google Scholar]
  • 46.Hark AT, Schoenherr CJ, Katz DJ, Ingram RS, Levorse JM, Tilghman SM. 2000. CTCF mediates methylation-sensitive enhancer-blocking activity at the H19/Igf2 locus. Nature. 405(6785):486–89 [DOI] [PubMed] [Google Scholar]
  • 47.Heintz N 2001. BAC to the future: the use of bac transgenic mice for neuroscience research. Nat Rev Neurosci. 2(12):861–70 [DOI] [PubMed] [Google Scholar]
  • 48.Herger M, Kajba CM, Buckley M, Cunha A, Strom M, Findlay GM. 2025. High-throughput screening of human genetic variants by pooled prime editing. Cell Genom. 5(4):100814. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Hnisz D, Abraham BJ, Lee TI, Lau A, Saint-André V, et al. 2013. Super-enhancers in the control of cell identity and disease. Cell. 155(4):934–47 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Hong CKY, Wu Y, Erickson AA, Li J, Federico AJ, Cohen BA. 2024. Massively parallel characterization of insulator activity across the genome. Nat Commun. 15(1):8350. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Hong J-W, Hendrix DA, Levine MS. 2008. Shadow enhancers as a source of evolutionary novelty. Science. 321(5894):1314. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Iacovino M, Bosnakovski D, Fey H, Rux D, Bajwa G, et al. 2011. Inducible cassette exchange: a rapid and efficient system enabling conditional gene expression in embryonic stem and primary cells. Stem Cells. 29(10):1580–88 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Inoue F, Kircher M, Martin B, Cooper GM, Witten DM, et al. 2017. A systematic comparison reveals substantial differences in chromosomal versus episomal encoding of enhancer activity. Genome Research. 27(1):38–52 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.James JS, Dai J, Chew WL, Cai Y. 2025. The design and engineering of synthetic genomes. Nat Rev Genet. 26(5):298–319 [DOI] [PubMed] [Google Scholar]
  • 55.Keane TM, Goodstadt L, Danecek P, White MA, Wong K, et al. 2011. Mouse genomic variation and its effect on phenotypes and gene regulation. Nature. 477(7364):289–94 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Koeppel J, Ferreira R, Vanderstichele T, Riedmayr LM, Peets EM, et al. 2025. Randomizing the human genome by engineering recombination between repeat elements. Science. 387(6733):eado3979. [DOI] [PubMed] [Google Scholar]
  • 57.Korkmaz G, Lopes R, Ugalde AP, Nevedomskaya E, Han R, et al. 2016. Functional genetic screens for enhancer elements in the human genome using CRISPR-Cas9. Nature Biotechnology. 34(2):192–98 [Google Scholar]
  • 58.Krebs AR, Dessus-Babus S, Burger L, Schübeler D. 2014. High-throughput engineering of a mammalian genome reveals building principles of methylation states at CG rich regions. Elife. 3:e04094. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Lalanne J-B, Regalado SG, Domcke S, Calderon D, Martin BK, et al. 2024. Multiplex profiling of developmental cis-regulatory elements with quantitative single-cell expression reporters. Nat Methods. 21(6):983–93 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Lanctôt C, Cheutin T, Cremer M, Cavalli G, Cremer T. 2007. Dynamic genome architecture in the nuclear space: regulation of gene expression in three dimensions. Nature Reviews Genetics. 8(2):104–15 [Google Scholar]
  • 61.Lettice LA, Heaney SJH, Purdie LA, Li L, de Beer P, et al. 2003. A long-range Shh enhancer regulates expression in the developing limb and fin and is associated with preaxial polydactyly. Human Molecular Genetics. 12(14):1725–35 [DOI] [PubMed] [Google Scholar]
  • 62.Li Q, Peterson KR, Fang X, Stamatoyannopoulos G. 2002. Locus control regions. Blood. 100(9):3077–86 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Li W, Hung P-H, Matsui T, Levy SF, Sherlock G. 2025. Scaling DNA engineering. Trends Biotechnol. 43(10):2399–2409 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Li X, Noll M. 1994. Compatibility between enhancers and promoters determines the transcriptional specificity of gooseberry and gooseberry neuro in the Drosophila embryo. The EMBO journal. 13(2):400–406 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Li Y, Rivera CM, Ishii H, Jin F, Selvaraj S, et al. 2014. CRISPR reveals a distal super-enhancer required for Sox2 expression in mouse embryonic stem cells. PLoS One. 9(12):e114485. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Lienert F, Wirbelauer C, Som I, Dean A, Mohn F, Schübeler D. 2011. Identification of genetic elements that autonomously determine DNA methylation states. Nature Genetics. 43(11):1091–97 [DOI] [PubMed] [Google Scholar]
  • 67.Lue NZ, Liau BB. 2023. Base editor screens for in situ mutational scanning at scale. Mol Cell. 83(13):2167–87 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Luthra I, Jensen C, Chen XE, Salaudeen AL, Rafi AM, de Boer CG. 2024. Regulatory activity is the default DNA state in eukaryotes. Nat Struct Mol Biol. 31(3):559–67 [DOI] [PubMed] [Google Scholar]
  • 69.Maeda RK, Karch F. 2011. Gene expression in time and space: additive vs hierarchical organization of cis-regulatory regions. Current opinion in genetics & development. 21(2):187–93 [DOI] [PubMed] [Google Scholar]
  • 70.Majchrzycka B, Mundlos S, Krebs AR, Ibrahim DM. 2025. Enhancer-promoter compatibility is mediated by the promoter-proximal region. bioRxiv. 2025.10.14.682013 [Google Scholar]
  • 71.Mallory BJ, Tullius TW, Biar CG, Gustafson JA, Bohaczuk SC, et al. 2025. Principles and functional consequences of plasmid chromatinization in mammalian cells. bioRxiv. 2025.05.27.656122 [Google Scholar]
  • 72.Mansour SL, Thomas KR, Capecchi MR. 1988. Disruption of the proto-oncogene int-2 in mouse embryo-derived stem cells: a general strategy for targeting mutations to non-selectable genes. Nature. 336(6197):348–52 [DOI] [PubMed] [Google Scholar]
  • 73.Maricque BB, Chaudhari HG, Cohen BA. 2019. A massively parallel reporter assay dissects the influence of chromatin structure on cis-regulatory activity. Nat Biotechnol. 37(1):90–95 [Google Scholar]
  • 74.Martyn GE, Montgomery MT, Jones H, Guo K, Doughty BR, et al. 2025. Rewriting regulatory DNA to dissect and reprogram gene expression. Cell. 188(12):3349–3366.e23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Mastrangelo IA, Courey AJ, Wall JS, Jackson SP, Hough PV. 1991. DNA looping and Sp1 multimer links: a mechanism for transcriptional synergism and enhancement. Proceedings of the National Academy of Sciences of the United States of America. 88(13):5670–74 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Mátés L, Chuah MKL, Belay E, Jerchow B, Manoj N, et al. 2009. Molecular evolution of a novel hyperactive Sleeping Beauty transposase enables robust stable gene transfer in vertebrates. Nature Genetics. 41(6):753–61 [DOI] [PubMed] [Google Scholar]
  • 77.Maurano MT, Humbert R, Rynes E, Thurman RE, Haugen E, et al. 2012. Systematic localization of common disease-associated variation in regulatory DNA. Science. 337(6099):1190–95 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.McEwen AE, Tejura M, Fayer S, Starita LM, Fowler DM. 2025. Multiplexed assays of variant effect for clinical variant interpretation. Nat Rev Genet [Google Scholar]
  • 79.Mendenhall EM, Williamson KE, Reyon D, Zou JY, Ram O, et al. 2013. Locus-specific editing of histone modifications at endogenous enhancers. Nature Biotechnology. 31(12):1133–36 [Google Scholar]
  • 80.Mercola M, Goverman J, Mirell C, Calame K. 1985. Immunoglobulin heavy-chain enhancer requires one or more tissue-specific factors. Science. 227(4684):266–70 [DOI] [PubMed] [Google Scholar]
  • 81.Merli C, Bergstrom DE, Cygan JA, Blackman RK. 1996. Promoter specificity mediates the independent regulation of neighboring genes. Genes & Development. 10(10):1260–70 [DOI] [PubMed] [Google Scholar]
  • 82.Meuleman W, Muratov A, Rynes E, Halow J, Lee K, et al. 2020. Index and biological spectrum of human DNase I hypersensitive sites. Nature. 584(7820):244–51 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Miklos GL, Rubin GM. 1996. The role of the genome project in determining gene function: insights from model organisms. Cell. 86(4):521–29 [DOI] [PubMed] [Google Scholar]
  • 84.Minikel EV, Painter JL, Dong CC, Nelson MR. 2024. Refining the impact of genetic evidence on clinical success. Nature. 629(8012):624–29 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Murphy KC. 2016. λ Recombination and Recombineering. EcoSal Plus. 7(1): [Google Scholar]
  • 86.Myers RM, Tilly K, Maniatis T. 1986. Fine structure genetic analysis of a beta-globin promoter. Science. 232(4750):613–18 [DOI] [PubMed] [Google Scholar]
  • 87.Ohno T, Akase T, Kono S, Kurasawa H, Takashima T, et al. 2022. Biallelic and gene-wide genomic substitution for endogenous intron and retroelement mutagenesis in human cells. Nat Commun. 13(1):4219. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Ohtsuki S, Levine M, Cai HN. 1998. Different core promoters possess distinct regulatory activities in the Drosophila embryo. Genes & Development. 12(4):547–56 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Ordonez R, Ribeiro-dos-Santos AM, McLoughlin C, Ellis G, Ashe HJ, et al. 2025. Synthetic genomic dissection of enhancer context sensitivity and synergy. bioRxiv. 2025.08.13.669251 [Google Scholar]
  • 90.Ordoñez R, Zhang W, Ellis G, Zhu Y, Ashe HJ, et al. 2024. Genomic context sensitizes regulatory elements to genetic disruption. Mol Cell. 84(10):1842–1854.e7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Palmiter RD, Brinster RL. 1986. Germ-line transformation of mice. Annual review of genetics. 20:465–99 [Google Scholar]
  • 92.Pandey S, Gao XD, Krasnow NA, McElroy A, Tao YA, et al. 2025. Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing. Nat Biomed Eng. 9(1):22–39 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 93.Parker SCJ, Stitzel ML, Taylor DL, Orozco JM, Erdos MR, et al. 2013. Chromatin stretch enhancer states drive cell-specific gene regulation and harbor human disease risk variants. Proceedings of the National Academy of Sciences of the United States of America. 110(44):17921–26 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Perry NT, Bartie LJ, Katrekar D, Gonzalez GA, Durrant MG, et al. 2025. Megabase-scale human genome rearrangement with programmable bridge recombinases. Science, p. eadz0276 [Google Scholar]
  • 95.Peterson KR, Clegg CH, Huxley C, Josephson BM, Haugen HS, et al. 1993. Transgenic mice containing a 248-kb yeast artificial chromosome carrying the human beta-globin locus display proper developmental control of human globin genes. Proceedings of the National Academy of Sciences of the United States of America. 90(16):7593–97 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Peterson KR, Li QL, Clegg CH, Furukawa T, Navas PA, et al. 1995. Use of yeast artificial chromosomes (YACs) in studies of mammalian development: production of beta-globin locus YAC mice carrying human globin developmental mutants. Proceedings of the National Academy of Sciences of the United States of America. 92(12):5655–59 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Peterson KR, Stamatoyannopoulos G. 1993. Role of gene order in developmental control of human gamma- and beta-globin gene expression. Molecular and Cellular Biology. 13(8):4836–43 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Pinglay S, Bulajić M, Rahe DP, Huang E, Brosh R, et al. 2022. Synthetic regulatory reconstitution reveals principles of mammalian Hox cluster regulation. Science. 377(6601):eabk2820. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Pinglay S, Lalanne J-B, Daza RM, Kottapalli S, Quaisar F, et al. 2025. Multiplex generation and single-cell analysis of structural variants in mammalian genomes. Science. 387(6733):eado5978. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 100.Pott S, Lieb JD. 2014. What are super-enhancers? Nature Genetics. 47(1):8–12 [Google Scholar]
  • 101.Rajagopal N, Srinivasan S, Kooshesh K, Guo Y, Edwards MD, et al. 2016. High-throughput mapping of regulatory DNA. Nature Biotechnology. 34(2):167–74 [Google Scholar]
  • 102.Rausch T, Zichner T, Schlattl A, Stütz AM, Benes V, Korbel JO. 2012. DELLY: structural variant discovery by integrated paired-end and split-read analysis. Bioinformatics. 28(18):i333–39 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Reilly SK, Gosai SJ, Gutierrez A, Mackay-Smith A, Ulirsch JC, et al. 2021. Direct characterization of cis-regulatory elements and functional dissection of complex genetic associations using HCR-FlowFISH. Nat Genet. 53(8):1166–76 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Ribeiro-Dos-Santos AM, Hogan MS, Luther RD, Brosh R, Maurano MT. 2022. Genomic context sensitivity of insulator function. Genome Res. 32(3):425–36 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Ribeiro-Dos-Santos AM, Maurano MT. 2025. Iterative improvement of deep learning models using synthetic regulatory genomics. Genome Res. 35(11):2539–49 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Sanjana NE, Cong L, Zhou Y, Cunniff MM, Feng G, Zhang F. 2012. A transcription activator-like effector toolbox for genome engineering. Nat Protoc. 7(1):171–92 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Sanjana NE, Wright J, Zheng K, Shalem O, Fontanillas P, et al. 2016. High-resolution interrogation of functional elements in the noncoding genome. Science. 353(6307):1545–49 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.Schübeler D, Lorincz MC, Cimbora DM, Telling A, Feng YQ, et al. 2000. Genomic targeting of methylated DNA: influence of methylation on transcription, replication, chromatin structure, and histone acetylation. Mol Cell Biol. 20(24):9103–12 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Seibler J, Schübeler D, Fiering S, Groudine M, Bode J. 1998. DNA cassette exchange in ES cells mediated by Flp recombinase: an efficient strategy for repeated modification of tagged loci by marker-free constructs. Biochemistry. 37(18):6229–34 [DOI] [PubMed] [Google Scholar]
  • 110.Shalem O, Sanjana NE, Hartenian E, Shi X, Scott DA, et al. 2014. Genome-scale CRISPR-Cas9 knockout screening in human cells. Science. 343(6166):84–87 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 111.Sharon E, Kalma Y, Sharp A, Raveh-Sadka T, Levo M, et al. 2012. Inferring gene regulatory logic from high-throughput measurements of thousands of systematically designed promoters. Nature Biotechnology. 30(6):521–30 [Google Scholar]
  • 112.Silver LM. 1995. Mouse Genetics: Concepts and Applications. Oxford University Press [Google Scholar]
  • 113.Small S, Blair A, Levine M. 1992. Regulation of even-skipped stripe 2 in the Drosophila embryo. The EMBO journal. 11(11):4047–57 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.Smemo S, Tena JJ, Kim K-H, Gamazon ER, Sakabe NJ, et al. 2014. Obesity-associated variants within FTO form long-range functional connections with IRX3. Nature. 507(7492):371–75 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Sproul D, Gilbert N, Bickmore WA. 2005. The role of chromatin structure in regulating the expression of clustered genes. Nature Reviews Genetics. 6(10):775–81 [Google Scholar]
  • 116.Stergachis AB, Debo BM, Haugen E, Churchman LS, Stamatoyannopoulos JA. 2020. Single-molecule regulatory architectures captured by chromatin fiber sequencing. Science. 368(6498):1449–54 [DOI] [PubMed] [Google Scholar]
  • 117.Swanson EG, Mao Y, Mallory BJ, Vollger MR, Bohaczuk SC, et al. 2025. Mapping single-cell diploid chromatin fiber architectures using DAF-seq. Nat Biotechnol [Google Scholar]
  • 118.Symmons O, Uslu VV, Tsujimura T, Ruf S, Nassari S, et al. 2014. Functional and topological characteristics of mammalian regulatory domains. Genome Research. 24(3):390–400 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Thanos D, Maniatis T. 1995. Virus induction of human IFN beta gene expression requires the assembly of an enhanceosome. Cell. 83(7):1091–1100 [DOI] [PubMed] [Google Scholar]
  • 120.Thomas HF, Feng S, Haslhofer F, Huber M, García Gallardo M, et al. 2025. Enhancer cooperativity can compensate for loss of activity over large genomic distances. Mol Cell. 85(2):362–375.e9 [DOI] [PubMed] [Google Scholar]
  • 121.Thomas HF, Kotova E, Jayaram S, Pilz A, Romeike M, et al. 2021. Temporal dissection of an enhancer cluster reveals distinct temporal and functional contributions of individual elements. Mol Cell. 81(5):969–982.e13 [DOI] [PubMed] [Google Scholar]
  • 122.Thurman RE, Rynes E, Humbert R, Vierstra J, Maurano MT, et al. 2012. The accessible chromatin landscape of the human genome. Nature. 489(7414):75–82 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Ulianov SV, Khrameeva EE, Gavrilov AA, Flyamer IM, Kos P, et al. 2016. Active chromatin and transcription play a key role in chromosome partitioning into topologically associating domains. Genome Research. 26(1):70–84 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Ulirsch JC, Lareau CA, Bao EL, Ludwig LS, Guo MH, et al. 2019. Interrogation of human hematopoiesis at single-cell and single-variant resolution. Nat Genet. 51(4):683–93 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.van Arensbergen J, van Steensel B, Bussemaker HJ. 2014. In search of the determinants of enhancer-promoter interaction specificity. Trends in cell biology. 24(11):695–702 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, et al. 2013. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr Protoc Bioinformatics. 43(1110):11.10.1–11.10.33 [Google Scholar]
  • 127.Vierstra J, Reik A, Chang K-H, Stehling-Sun S, Zhou Y, et al. 2015. Functional footprinting of regulatory DNA. Nature Methods. 12(10):927–30 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Visscher PM, Wray NR, Zhang Q, Sklar P, McCarthy MI, et al. 2017. 10 Years of GWAS Discovery: Biology, Function, and Translation. Am J Hum Genet. 101(1):5–22 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Wallace HAC, Marques-Kranc F, Richardson M, Luna-Crespo F, Sharpe JA, et al. 2007. Manipulating the mouse genome to engineer precise functional syntenic replacements with human sequence. Cell. 128(1):197–209 [DOI] [PubMed] [Google Scholar]
  • 130.Wang T, Wei JJ, Sabatini DM, Lander ES. 2014. Genetic screens in human cells using the CRISPR-Cas9 system. Science. 343(6166):80–84 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Weintraub H 1988. Formation of stable transcription complexes as assayed by analysis of individual templates. Proceedings of the National Academy of Sciences of the United States of America. 85(16):5819–23 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.Weintraub H 1993. The MyoD family and myogenesis: redundancy, networks, and thresholds. Cell. 75(7):1241–44 [DOI] [PubMed] [Google Scholar]
  • 133.Wilson C, Bellen HJ, Gehring WJ. 1990. Position effects on eukaryotic gene expression. Annual review of cell biology. 6:679–714 [Google Scholar]
  • 134.Wolffe A 1999. Chromatin: Structure and Function. Academic Press. 3rd ed. [Google Scholar]
  • 135.Wu Y, Li J, Bartley-Dier EL, Pitts C, Cohen BA. 2025. Long-range massively parallel reporter assay reveals rules of distal enhancer-promoter interactions. bioRxiv. 2025.04.21.649048 [Google Scholar]
  • 136.Yarnall MTN, Ioannidi EI, Schmitt-Ulms C, Krajeski RN, Lim J, et al. 2023. Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat Biotechnol. 41(4):500–512 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Yigit E, Zhang Q, Xi L, Grilley D, Widom J, et al. 2013. High-resolution nucleosome mapping of targeted regions using BAC-based enrichment. Nucleic Acids Res. 41(7):e87. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Zabidi MA, Arnold CD, Schernhuber K, Pagani M, Rath M, et al. 2015. Enhancer-core-promoter specificity separates developmental and housekeeping gene regulation. Nature. 518(7540):556–59 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Zhang W, Golynker I, Brosh R, Fajardo A, Zhu Y, et al. 2023. Mouse genome rewriting and tailoring of three important disease loci. Nature. 623(7986):423–31 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.Zhang W, Mitchell LA, Bader JS, Boeke JD. 2020. Synthetic Genomes. Annu Rev Biochem. 89:77–101 [DOI] [PubMed] [Google Scholar]
  • 141.Zhao Y, Coelho C, Lauer S, Majewski M, Laurent JM, et al. 2023. CREEPY: CRISPR-mediated editing of synthetic episomes in yeast. Nucleic Acids Res. 51(13):e72. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Zhou HY, Katsman Y, Dhaliwal NK, Davidson S, Macpherson NN, et al. 2014. A Sox2 distal enhancer cluster regulates embryonic stem cell differentiation potential. Genes & Development. 28(24):2699–2711 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 143.Zhou Z, Li A, Zhang J, Yu H, Ozer A, Lis JT. 2025. Robust regulatory interplay of enhancers, facilitators, and promoters in a native chromatin context. bioRxiv. 2025.07.07.663560 [Google Scholar]
  • 144.Zhu F, Gamboa M, Farruggio AP, Hippenmeyer S, Tasic B, et al. 2014. DICE, an efficient system for iterative genomic editing in human pluripotent stem cells. Nucleic Acids Res. 42(5):e34. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Zuin J, Roth G, Zhan Y, Cramard J, Redolfi J, et al. 2022. Nonlinear control of transcription through enhancer-promoter interactions. Nature. 604(7906):571–77 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES