Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 May 20:2025.09.10.675455. [Version 4] doi: 10.1101/2025.09.10.675455

Genome-Wide Annotation of Promoter-Enhancer Interactions via Chromatin Loops in the Hybrid Rat Diversity Panel

Panjun Kim 1, Burt M Sharp 1, Robert W Williams 1, Hao Chen 2
PMCID: PMC13228554  PMID: 42239386

Abstract

Despite the prominence of laboratory rats in behavioral neuroscience and complex trait genetics, a significant gap exists in the functional rat genomics database that explains the regulation of gene expression at the genome level. To address this, we analyzed genome-wide Hi-C data from the frontal cortex of ten strains in the Hybrid Rat Diversity Panel. While originally generated to improve the rat genome assembly, these data provided a unique opportunity to characterize the regulatory landscape of the adult rat brain. We identified an average of 5,899 ± 1,997 (STD) loops per sample and integrated these with over 3 million curated CTCF binding sites, which serve as architectural anchors for chromatin loops. Our multi-stage filtering workflow identified 15,085 unique, high-confidence regulatory interactions throughout the genome. As a validation of our approach, we observed that genes with the highest loop counts, including Foxo1, Gja1, and Spry2, were significantly enriched for developmental processes, reflecting the role of 3D genome structure in maintaining adult neuronal identity. The resulting resource addresses a critical deficiency in rat functional genomics and provides a foundation for dissecting the genetic architecture of complex traits.

INTRODUCTION

The spatiotemporal control of gene transcription is fundamental to the development and maintenance of complex biological systems (1). This regulation is primarily mediated by distal cis-regulatory elements, such as enhancers, which integrate physiological and environmental signals to drive cell-type-specific transcriptional profiles (1, 2). In the central nervous system, these regulatory mechanisms underlie the expression of complex behavioral phenotypes (3–5). Genetic variants in these non-coding regions are frequently associated with phenotypic diversity and susceptibility to psychiatric conditions (6–8). Identifying the genes regulated by these elements is essential for characterizing the molecular mechanisms that link genotype to phenotype.

The laboratory rat (Rattus norvegicus) is a prominent model for biomedical research, offering a physiological complexity and behavioral repertoire, including reward processing and social behavior, that parallels human conditions more closely than other rodent models (9–11). Although the rat reference genome has recently reached a quality comparable to human and mouse standards (12–14), its functional annotation remains significantly underdeveloped (15). This gap is underscored by the exhaustive functional mapping available for other major mammalian models through initiatives such as ENCODE and Roadmap Epigenomics (15). No such comprehensive resource exists for the rat, leaving much of its regulatory landscape undefined. Furthermore, because regulatory relationships are highly tissue- and cell-type specific, the lack of data from relevant brain regions hinders the interpretation of how genetic variation influences physiological and behavioral outcomes (8).

Mapping the physical interactions between distal enhancers and their target promoters is a critical step toward overcoming this limitation. The three-dimensional (3D) organization of the genome facilitates these contacts, often through chromatin loops organized within topologically associating domains (TADs) (2, 16, 17). These architectural features are typically anchored by proteins such as the CCCTC-binding factor (CTCF) and cohesin, which bring distant enhancers, where transcription factors bind, into proximity with gene promoters to regulate transcription (1, 18). High-throughput chromosome conformation capture (Hi-C) provides a robust methodology for resolving these interactions at a genome-wide scale, enabling the identification of promoter-enhancer (P-E) loops that control gene expression (1, 18). However, while high-resolution mapping of enhancer-promoter contacts has been extensively characterized in various human and mouse cell types (1), equivalent comprehensive 3D genomic and epigenomic annotations for the rat lag significantly behind those of other major mammalian models (11, 15).

In this study, we use Hi-C to delineate the 3D regulatory architecture of the rat frontal cortex, a region central to executive function and behavioral regulation. We analyzed data from ten inbred strains within the Hybrid Rat Diversity Panel (HRDP), a reference population designed for high genetic and phenotypic diversity (19, 20). These data, originally generated for de novo genome assembly of individual HRDP strains, represent a high-resolution resource that we repurposed to characterize the regulatory landscape of the rat brain. By integrating these contact maps with over 3 million CTCF binding sites predicted using FIMO, we identified 15,085 high-confidence regulatory interactions across the rat genome, with each interaction assigned to a specific target gene.

The validity of these gene-specific loops is strongly supported by the finding that genes with the highest loop density, including Foxo1, Gja1, and Spry2, are significantly enriched for developmental processes essential for maintaining adult neuronal identity. This enrichment is consistent with the role of 3D genome architecture in preserving transcriptional states required for mature neuronal function. This resource addresses a critical deficiency in rat functional genomics, providing a foundational tool for the research community to investigate the molecular drivers of behavioral and physiological diversity.

RESULTS

Characterizing loops from 10 Hi-C rat samples

The production of these sequencing data has been described in detail (21), and all data have been deposited to the NIH Short Read Archive under accession number PRJNA1197090. Briefly, Arima Hi-C libraries were extracted from the frontal cortex of 10 rats, each a unique strain or F1 hybrid: SHR/OlaIpcv, BN-Lx/Cub, BXH6/Cub, HXB2/Ipcv, HXB10/Ipcv, HXB23/Ipcv, HXB31/Ipcv, LE/Stm, F344/Stm, and SHR/OlaIpcv x BN/NHsdMcwi F1, selected to maximize genetic distances based on a phylogenetic tree we generated before (14). Data were processed using the Juicer pipeline (v1.6) (22), aligned to the mRatBN7.2/rn7 reference genome. On average, 570.2 million paired-end reads were generated per sample, although LE/Stm and F344/Stm samples yielded fewer reads (21). Figure 1a summarizes the distribution of several categories of reads as proportions of the total read pairs. Unalignable reads, which include chimeric, ambiguous, and unmapped categories, accounted for an average of 11.4% of the total sequenced read pairs. Among these, unmapped reads represented only 1.3%. The overall proportion of unalignable reads was the lowest among all read categories and showed relatively low variability across samples (STD = 2.1%). PCR and optical duplicates accounted for 21.4% of the reads. Notably, samples such as LE/Stm (8.0%) and F344/Stm (8.6%) showed markedly lower duplication rates. The variation in duplication rates across samples (STD = 11.0%) most likely reflects technical differences. Unique reads constituted the largest proportion of total sequenced read pairs across all samples (67 ± 10.0%), but the extent varied considerably among samples. We retained uniquely mapped reads with a mapping quality threshold of 30 (i.e., < 0.1% chance of misalignment) to ensure reliable Hi-C contact inference.

Figure 1. Depth of unique reads determines the number of detected loops.

Figure 1.

a. Proportions of unique, duplicate, and chimeric/unmapped reads in Hi-C data across 10 samples. Hi-C sequencing reads from 10 samples were processed using the Juicer pipeline (v1.6) and categorized into unique reads, PCR/optical duplicates, and unalignable reads (including chimeric ambiguous, and unmapped). Unique reads comprised the majority across most samples. b. Positive correlation between loop counts and sequencing depth across 10 samples.

We further processed the Hi-C sequencing data using HiCCUPS (18) from the Juicer pipeline to annotate loop structures. In total, 58,992 loops were identified across the 10 samples, with persample counts ranging from 2,903 to 9,131 (Table S1). To evaluate the relationship between sequencing metrics and chromatin loop detection, we compared the number of loops identified per sample as a function of three sequencing depth metrics: total reads, unique reads, and alignable reads (Figure 1b). Loop counts were strongly correlated with sequencing depth for all three metrics examined (total reads, alignable reads, and uniquely mapped reads), with Pearson correlation coefficient ranging from 0.91 to 0.96 (all p < 0.001; Figure 1b), indicating that sequencing depth is a primary determinant in loop detection. In addition, several other QC metrics (Figure S1), such as the number of Hi-C contacts and the proportion of non-chimeric alignable reads, also showed strong positive associations with loop count. These results underscore the primary role of sequencing depth and library complexity in identifying chromatin structures.

We investigated how loop counts varied across three loop calling resolutions (5K, 10K, and 25K), considering all chromosomes (Figure 2). Overall, the number of annotated loops generally increased with broader resolution, particularly from 5K to 10K, with more modest gains observed between 10K and 25K. An exception to this trend was the Y chromosome, which showed a decrease in loop counts at 25K. Across the genome, the number of detected loops was strongly and positively correlated with chromosome length (r = 0.88, p = 5.47 × 10 ) as expected, reflecting more chromatin regulatory interactions in larger chromosomes.

Figure 2. Genome-wide distribution of loop counts detected at three resolutions.

Figure 2.

Loops were annotated on each chromosome at 5K, 10K, and 25K resolutions using HiCCUPS. In general, higher resolution (5K) yielded fewer detected loops. Larger chromosomes tend to have more loops, reflecting a greater number of regulatory interactions.

We then examined loops that were detected in multiple samples at each resolution. Figure 3a summarizes the mean number and variability of loops shared (i.e., have identical anchoring positions) between any two samples at each resolution. Lower resolution (25K) yielded a greater number of shared loops, accompanied by reduced variability. Specifically, the average number of shared loops increased from 774 ± 468 at 5K resolution to 1,379 ± 420 at 10K and 1,853 ± 225 at 25K. In addition, there was substantial variability in the total number of loops per sample (Figure 3b), with HXB31 showing the highest count (9131), and 54.6% classified as shared loops compared to 45.4% unique loops. Conversely, F344/Stm, which had the lowest total loop count (2903), demonstrates 74.8% shared loops and 25.2% unique. Notably, samples with fewer detected loops exhibited a higher proportion of shared loops, likely due to relatively low sequencing depth in these samples.

Figure 3. The majority of the loops are shared between samples.

Figure 3.

a. Number of shared chromatin loops between any pairs of samples at 5K, 10K, and 25K resolutions. The number of shared loops was greater, and variability was lower at lower (i.e., 25K) detection resolution. Error bars represent standard deviations. b. Distribution of unique and shared loops across rat samples. Chromatin loops are classified as either unique (sample-specific) or shared (observed in multiple samples) across each of the ten samples. Samples with greater sequencing depth (e.g. HXB31 and HXB23) have more loops detected and higher proportion of unique loops, compared to samples with lower number of sequencing depth (e.g., LE/Stm and F344/Stm).

After combining loops from all three resolutions (Figure S2), we observed a high degree of sharing, suggesting a conserved chromatin topology among these related strains. In contrast, fewer connections in other strains reflect greater biological divergence, combined with artifacts of technical variability. Figure S3 further quantifies pairwise loop sharing across three resolutions (5K, 10K, and 25K) using heatmaps. At 5K resolution, F344/Stm, HXB2, and LE/Stm exhibited a relatively higher proportion of shared loops compared to higher resolutions. This pattern may, in part, result from lower total loop counts in these samples. Together, these data reveal both conserved and strain-specific features of chromatin looping among a diverse collection of rat genomes.

Genome-wide annotation of CTCF binding sites

To further validate the functional relevance of these loops, we examined the genomic features that are critical for P-E interactions. CTCF plays a critical role in the formation of chromatin loops by acting as a boundary element that binds to specific DNA motifs. Accordingly, mapping the positional distribution of CTCF motifs provides a framework for identifying loops that may contribute to transcriptional regulation. Using 100 CTCF motifs curated from six databases (Table S2), we identified 6,551,641 potential CTCF binding sites across the rat genome using FIMO (23). The ideogram in Figure 4a illustrates the genomic distribution of CTCF elements across the rat genome. CTCF binding sites were not uniformly distributed across the chromosomes. We found chromosome length was not correlated with CTCF density per Mb (Pearson r = −0.202, p = 0.367). In contrast, CTCF density was strongly correlated with gene density (Figure 4b-4d). Using NCBI annotations, CTCF density was significantly correlated with the density of all genes (Pearson r = 0.922, p = 1.03e-09), protein-coding genes (r = 0.877, p = 8.51e-08), and lncRNA genes (r = 0.891, p = 2.66e-08). Similar correlations were observed using Ensembl annotations for all genes (r = 0.910, p = 4.56e-09), protein-coding genes (r = 0.872, p = 1.24e-07), and lncRNA genes (r = 0.796, p = 9.32e-06).

Figure 4. Chromosomal distribution of CTCF binding site density and its correlation with gene density across the rat genome.

Figure 4.

a. Genome-wide distribution of CTCF binding sites across all rat chromosomes (mRatBN7.2 assembly). Each vertical bar represents a chromosome, with color-coded segments reflecting the density of CTCF binding sites within non-overlapping 1 Mb genomic bins. The heatmap scale ranges from low (blue, minimum number of 0 CTCF in a bin) to high (red, maximum number of 5711 CTCF in a bin) density, allowing for visual comparison of CTCF abundance across different chromosomal regions. b–d. Chromosome-level correlation between gene density (genes per Mb) and CTCF binding site density (CTCF sites per Mb) using NCBI gene annotations. Each data point (blue dots) represents a single chromosome. b. All genes (n = 42,049), c. protein-coding genes (n = 21,944), and d. lncRNA genes (n = 7,815). Pearson correlation coefficients and p-values are shown in each panel. Linear regression fits (red lines) highlight the positive association between gene density and CTCF density.

Positional patterns of CTCF, TSS, and promoter near loop anchors

In addition to CTCF binding sites, we also used 42,925 Ensembl start-codon annotations and 12,601 experimentally validated promoter coordinates. After filtering and removing redundant entries based on identical coordinates (chromosome, start, and end positions), we retained 3,191,859 unique CTCF binding sites, 21,725 distinct TSS, and 12,529 unique promoter regions for downstream analyses. We next examined the distribution of CTCF motif binding sites, TSS, and promoter regions in the vicinity of chromatin loops. We used a final set of 31,019 loops shorter than 2Mb (18, 24). For the positional distribution analysis, we analyzed a slightly reduced set of 30,928 loops, excluding the 91 loops located at the end of a chromosome that exceeded the 3L flanking region definition.

Figure 5 and Figure S4 present density plots and histograms, respectively, illustrating the relative positions of CTCF (a), TSS (b), and promoters (c) across all loops. Each plot spans flanking regions equivalent to one loop length on either side of the anchor. The density plots in Figure 5 are stratified by three different resolutions (5K, 10K, and 25K). The bimodal pattern is evident across all resolutions, with distinct peaks showing enrichment of CTCF (Figure 5a) occurring symmetrically at positions flanking the loop region (approximately at 0 and +1, which represent the start and end of the loop, after normalization of loop length). The pronounced enrichment at loop boundaries underscores the architectural role of CTCF in anchoring chromatin loops, a feature that is consistently maintained across resolutions. While this pattern remains stable, peak intensity tends to increase with higher resolution. A bimodal enrichment of TSS (b) and promoter (c) is also evident at loop boundaries; however, the peaks in these figures appear broader, and resolution-dependent differences in peak height are less distinct.

Figure 5. Density plots for positional distribution of CTCF binding sites, TSS, and promoters relative to chromatin loops.

Figure 5.

Each plot spans ±1 loop length from the loop boundaries, with resolution-specific curves overlaid. a. CTCF binding sites exhibit distinct peaks at both loop anchors (x = 0 and 1) across all resolutions, indicating consistent enrichment at loop boundaries regardless of resolution. b. TSS density displays an enrichment near loop anchors with lower and broader amplitude compared to CTCF. c. The distribution of the promoter also displays a similar profile near both anchors resembling the pattern observed for TSS.

At the level of each chromosome, density plots and histograms for the relative location of CTCF, TSS, and promoter elements showed more diverse distributional patterns around chromatin loop anchors. These patterns ranged from distinct, symmetric peaks at both loop anchor locations to asymmetric or weakly defined profiles, indicating variability in their spatial positioning and enrichment (Figure S5, S6, and S7).

Informed by the positional distribution of CTCF motifs, TSS, and promoter elements, we examined genomic features near the anchors of chromatin loops to establish a threshold for selecting loops that mediate P-E interaction. We used the final set of 31,019 loops shorter than 2 Mb for this and all subsequent structural filtering analyses. For CTCF, we quantified the number of CTCF motif binding sites present in both anchors of the loops defined by HiCCUPS. The distribution of CTCF binding site counts in the anchors showed a distinct bimodal pattern across all detection resolutions (Figure 6). A prominent peak corresponded to the majority of loops, centered around approximately 32 CTCF motifs per anchor for loops detected at 5K or 10K resolution, and approximately 42 motifs for those detected at 25K resolution. In addition, a minor peak with 1 or 2 CTCF sites likely arose from artifacts of the loop detection method. Regardless of resolution, the trough between these two populations occurred at approximately 6 CTCF sites and served as a data-driven threshold to distinguish likely functional loops from those likely falsely detected by HiCCUPS.

Figure 6. Distribution of CTCF binding site counts at loop anchors by direction and resolution.

Figure 6.

Histograms showing the distribution of CTCF binding site counts (log -transformed) at loop anchors, stratified by loop resolution (5K, 10K, and 25K) and direction (upstream vs. downstream), where upstream and downstream are labeled as “UP” and “DOWN” respectively in the plot. Each panel represents a specific combination of direction and resolution, illustrating how CTCF occupancy varies across the genomic scale and loop orientation. A distinct bimodal distribution is evident across all panels. The first peak, consisting of loops with a low number of CTCF domains (log (ctcf_count) < 2.5), is likely noise resulting from the loop detection method. Of the 31,019 loops analyzed, only 111 lacked detectable CTCF binding sites at both the upstream and downstream anchors.

Identifying loops mediating P-E interaction

We conducted extensive annotation and filtering steps, as shown in Figure 7, to identify chromatin loops that are likely to mediate P-E interactions. These filtering steps are based on the assumption that one loop anchor is located near the start of a gene, that is, the promoter or TSS, while the other anchor corresponds to a regulatory element, such as an enhancer, and that the loop is flanked by strong CTCF signals.

Figure 7. Workflow for the identification of high-confidence gene regulatory chromatin loops.

Figure 7.

The filtering pipeline began with an initial aggregate of 58,992 loops identified across ten biological samples. Duplicate loops were merged to eliminate redundancy (27,219 loops removed), followed by the exclusion of long-range interactions exceeding 2 Mb to prioritize intradomain contacts (754 loops removed). The remaining loops were subjected to a dual-criteria assessment: (1) Architectural filtering, which required both loop anchors to contain a minimum of six unique CTCF binding sites; and (2) Functional annotation, which selected loops containing a valid TSS or promoter region within 200kb of the anchor-to-feature distance. Only interactions satisfying both the architectural and functional criteria were retained, yielding a final dataset of 15,085 high-confidence loops.

We first assigned the nearest promoter or TSS to each anchor. Only TSS or promoters located on the same side of the loop, defined by the loop midpoint, were considered. After excluding distance-tied cases, we obtained 61,969 anchors with a single TSS or promoter assignment.

Three additional anchors retained two valid assignments each, resulting in a total of 61,975 anchor–regulatory element assignments. Next, we examined the distance between the assigned promoter or TSS and the corresponding anchor. As shown in Figure S8, more than half of the promoters or TSS overlap with the anchors. However, the distance distribution exhibits a long tail. We reasoned that a minority of promoters or TSS located more than 200 kb away from the anchor are unlikely to be biologically relevant. Applying this distance threshold retained 11,919 loops associated with TSS and 5,729 loops associated with promoter regions, resulting in a total of 17,648 loops selected for downstream analysis.

We next filtered these loops for CTCF signals. Applying the threshold of at least six CTCF motifs in both anchors identified 25,268 loops, which represents 81.5% of the total 31,019 loops analyzed. The intersection of these two sets was then assessed to determine the extent to which CTCF anchored loops were also supported by TSS or promoter based assignments. This analysis identified 10,125 loops that satisfied both the TSS and CTCF based criteria, and 4,960 loops that satisfied both the promoter and CTCF based criteria, resulting in a total of 15,085 loops supported by both structural (CTCF) and functional (TSS or promoter) evidence.

The genomic location of the chromatin loops likely mediating P-E interactions, as illustrated in Figure 8a for the entire genome and Figure 8b for Chr 1, showed that they are widely distributed across chromosomes, though some genomic regions are more densely populated than others, while others appear depleted. Circos plots of the loops for other chromosomes are provided in Figure S9.

Figure 8. Circos plots of high-confidence chromatin loops across the genome (a) and chromosome 1 (b), stratified by resolution.

Figure 8.

a. genome-wide chromatin loops stratified by resolution (5K, 10K, and 25K). b. loops located on chromosome 1. Each arc represents a chromatin loop between two genomic loci, with arc heights scaled logarithmically to reflect their genomic span. These plots illustrate the broad yet non-uniform distribution of loops and enable visual comparison across resolution levels.

We next counted the number of loops each gene is associated with. We found a total of 54 genes associated with 11 or more loops (Table S3). The top three most connected genes, Foxo1, Gja1, and Spry2, were annotated with 19 loops each. To investigate potential functional convergence among the interaction-rich genes, we performed gene ontology enrichment analysis using g:Profiler (25) on the top 54 genes, corresponding to those with 11 or more interactions (Figure 9). Analysis using the top 25 genes, each with at least 13 interactions, showed similar results. We found a coherent enrichment in embryonic development, morphogenesis, cellular proliferation and other broader developmental growth and stimulus response functions. Many of these ontological categories are related and contain overlapping genes, such as Spry2, Foxc1, and Tbx3.

Figure 9. Functional enrichment of top interaction-rich genes reveals developmental and proliferative biological processes.

Figure 9.

Functional enrichment analysis of the top 54 genes with the highest numbers of loop interactions was performed using g:Profiler. The figure highlights the presence of key developmental and proliferative processes, including embryonic organ development, tissue morphogenesis, and cell population proliferation, across representative gene loci.

DISCUSSION

We generated a map of chromatin loops for the frontal cortex of ten rat strains from the HRDP. Integrating Hi-C data with genomic annotations identified 15,085 loops with strong likelihood of modulating gene expression. By calling loops across ten diverse genetic backgrounds and merging them into a consensus map, we established a robust multi-strain regulatory resource. This dataset accounts for the inherent genetic diversity of the HRDP, providing a foundation for investigating the 3D landscape of rat genomes and facilitating the dissection of genetic architecture underlying complex behavioral traits.

CTCF is an 11-zinc-finger DNA-binding protein involved in regulating the 3D structure and function of the eukaryotic genome (26, 27). Current models posit that the 3D genome is organized via loop extrusion, a process where cohesin extrudes chromatin until halted by inward-facing CTCF anchors (28, 29). In the absence of experimental CTCF data for the rat genome, we used the FIMO motif scanning tool to predict CTCF binding sites on the mRatBN7.2 reference genome. The predictive accuracy of this method is supported by comparative studies in human and mouse genomes, where FIMO-predicted CTCF sites showed an approximately 80% true positive rate against experimental data (30). To increase the recall of CTCF sites, we included 100 curated CTCF motifs from multiple databases (Table S2). While this approach may increase the number of false positive sites, our analysis of the location of CTCF sites near loop anchors revealed a distinct bimodal distribution as expected (Figure 5). We observed this pattern across all three resolutions, consistent with an architectural role for CTCF in loop formation (18, 31, 32). We also observed that the predicted CTCF binding sites are not uniformly distributed across chromosomes (Figure 4a). Instead, they are associated with gene density, a pattern observed across independent gene annotations and multiple gene classes, such as protein-coding gene and long non-coding gene (Figure 4b-d). This also provides additional evidence that the predicted CTCF binding sites are reliable despite the use of cross-species motifs.

To systematically identify functional loops facilitating P-E interactions, we implemented a multi-stage filtering workflow (Figure 7). We first enriched for structurally stable loops by requiring at least six CTCF binding sites at both anchors based on our data. Under this criterion, 81.5% of the loops were retained. It is possible that this high stringency might have excluded those dynamic, tissue-specific P-E interactions mediated by factors such as the Mediator complex in the absence of dense CTCF clusters (33). We next require that at least one anchor located in the vicinity (i.e. <200 kb) of a promoter or TSS. This distance threshold aligns with established mammalian models where the majority of functional interactions occur within a 200 kb radius of the anchors (34). This step resulted in the exclusion of 13,043 loops, which likely were structural in nature rather than regulatory. In cases where both anchors contained TSS/promoter annotations, we resolved the ambiguity by selecting the element closest to its respective anchor midpoint as the primary target.

The biological relevance of the loops in our dataset is supported by the presence of many genes with a large number of loops (Table S3), such as Foxo1, Gja1, Spry2, Tgfb2 and Tbx3. Most of these genes are involved in embryonic development as shown by GO enrichment analysis (Figure 9). Consistent with the findings that key developmental genes frequently possess ten or more redundant enhancers (35), we infer that the dense loop architectures around these loci correspond to complex multi enhancer regulatory landscapes. These enhancers work together to precisely control the timing and location of gene expression (36, 37). Furthermore, multiple loops facilitate the formation of localized nuclear microenvironments that concentrate transcription factors and coactivators (38). For instance, consistent with previous reports of enhancer-rich regulatory landscapes at the Tbx3-Tbx5-Lhx5 locus in the developing limbs of mouse embryos (39), we found 16 loops for Tbx3. Similarly, we found 17 loops for Tgfb2, which is in agreement with a report of 15 distinct promoter-interacting regions in human triple-negative breast cancer cells (40). This high degree of regulatory complexity acts as a “regulatory buffer,” whereby enhancer and suppressor loops converge in the regulation of a specific gene. Functionally, this buffering refers to the capacity of redundant regulatory elements to maintain consistent target gene expression levels despite fluctuations or reductions in upstream transcription factors, and environmental variations such as temperature shifts (35, 41, 42). This promotes phenotypic robustness by safeguarding proper developmental outcomes against inevitable environmental stressors or genetic variability (38, 39, 42).

Our finding that “developmental” multi loop architectures persist in the adult rat frontal cortex supports the idea that these regulatory structures are essential to maintain neuronal identity, connectivity and homeostasis throughout the lifespan. For instance, complex enhancer clusters (e.g., the Sox2 SRR2–18 cluster) are essential not only for initial development but also for the maintenance of the neural stem and progenitor cell phenotype (43). Given that the mammalian postnatal brain retains diverse neuronal and glial cell types defined by distinct gene expression profiles (44, 45), the high loop density around genes like Foxo1 and Spry2 likely implicates the ongoing transcriptional regulation of these genes in the preservation of cell-type specific identity. Furthermore, enhancers regulating proliferation and identity in human neural cells often act on genes associated with neuropsychiatric disorders (e.g., schizophrenia and bipolar disorder) (46), indicating that persistence and the strict control of these regulatory networks within the adult brain are critical for the prevention of cognitive deficits.

To further examine the neural functions related to loop-rich genes, we utilized the text-mining tool GeneCup (47). The results highlighted that several genes are collectively implicated in critical neural processes, including intracellular signaling, transcriptional regulation, neurotransmission and neural maintenance, and are involved in the pathology of neurodegenerative diseases. For example, Gja1 (Connexin43) is a key driver of astrocytespecific networks dysregulated in Alzheimer’s disease (48). Astrocyte-specific overexpression of Gja1 has been found to enhance neuronal viability and mitochondrial recovery in models of traumatic brain injury, highlighting its neuroprotective capacity (49). In the dopaminergic systems, Tgfb2 variants are associated with susceptibility to Parkinson’s disease, likely due to its role in supporting dopaminergic neuron survival (50). Additionally, Foxo1 acts as a critical mediator of psychological stress-induced neuroinflammation (51). Elevated Foxo1 expression was observed in the hippocampus following stress exposure, and its knockdown markedly attenuated stress-induced inflammatory factors in microglia (51, 52). Taken together, the fact that these critical neuroprotective and disease-associated genes are embedded within exceptionally dense chromatin loop networks in our cortical data is highly insightful. It suggests that, much like their functions in other brain regions, the mammalian brain relies on this massive 3D regulatory redundancy to securely buffer the expression level of these genes, thereby ensuring the lifelong functional integrity of the adult cortex against environmental stress and aging.

Although the laboratory rat is an important model for biomedical research, its genomic functional annotation lags behind that of human and mouse (14). This work provides functional data by mapping chromatin interactions across multiple strains in the rat brain. The resulting dataset enhances the utility of the rat as a model organism by helping connect genetic variation to gene regulation. Similar to humans, many disease-associated variants found in rat genetic studies are located in non-coding regions, which complicates the identification of their target genes and functional mechanisms (53, 54). Our data address this challenge by providing physical links between distal regulatory elements and their promoters. This enables a more precise interpretation of non-coding variants by connecting them to specific genes within interacting loop anchors. Furthermore, these data can be integrated with genome-wide association study results using tools like H-MAGMA (55) to prioritize candidate genes. Integrating this chromatin interaction map with transcriptomic and epigenomic data will further enable the prioritization of disease-relevant loci and the inference of their upstream transcriptional regulators (56).

While many of the loops we identified likely mediate P-E interactions, some could connect promoters to silencers, which also have roles in gene regulation (57). Future studies integrating this dataset with other functional genomics data, such as epigenomic maps of active and repressive chromatin marks (e.g., H3K27ac and H3K27me3) and transcriptomic data, will be necessary to further investigate the functional consequences of these 3D interactions.

There are also a few limitations of this study. Consistent with previous Hi-C studies (58), we found that sequencing depth was a primary factor in loop detection, with lower depth yielding fewer, more commonly shared loops between samples. The resolution of detection also influenced the number of loops we identified. Lower resolutions (e.g., 25K) captured a broader set of long-range interactions, whereas higher resolutions (5K) provided more precise boundary definitions but fewer total loops. We thus combined loops from all three resolutions to construct a comprehensive catalog. However, these findings underscore a technical consideration for comparative studies: technical variability can confound the identification of biological differences, such as strain or treatment specific loops. Consequently, our current catalog primarily reflects a consensus regulatory landscape of the rat frontal cortex, rather than strain-to-strain variability. In addition, we generated these data using DNA from the rat frontal cortex, a brain region involved in executive functions and complex behaviors (59). We selected this region for its relevance to the regulatory architecture of many behavioral traits. However, this choice of tissue limits the generalizability of our findings to other tissues, as specific enhancer-promoter loops vary substantially across different cell types (60). It is plausible that other cortical or non-cortical brain areas exhibit alternative chromatin looping patterns.

METHODS

Hi-C data generation.

Hi-C library construction and sequencing were conducted for the frontal cortex of 10 laboratory rat strains (SHR/OlaIpcv, BN-Lx, BXH6, HXB2, HXB10, HXB23, HXB31, LE/Stm, F344/Stm, and SHR/OlaIpcvxBN/NHsdMcwi) using Arima Hi-C kits, as described previously (21).

Hi-C reads mapping and loop annotation.

We processed the raw sequencing data using the Juicer pipeline (v1.6) (22) and annotated chromatin loops with HiCCUPS (18), implemented at Juicer Tools v1.22.01 in a CUDA 8 environment on Google Colab (https://colab.research.google.com). For two strains, SHR/OlaIpcv (SRR34514260) and HXB10/Ipcv (SRR34514259), SRA submission records under BioProject PRJNA1197090 indicate that each Hi-C sample is represented by a single sample accession but contains two sequencing sets/libraries. SHR/OlaIpcv includes the 592A (S1) and 592B (S2) libraries, whereas HXB10/Ipcv includes the 607B (S3) and 607C (S4) libraries. For HXB10/Ipcv, the two sets were merged and analyzed as a single sample-level Hi-C dataset. In contrast, loop calling and downstream analyses for SHR/OlaIpcv were based on the higher-quality 592B dataset only rather than the merged dataset, because the 592A library showed substantially poorer Hi-C library quality, including a markedly higher PCR duplicate rate (70.8% for 592A versus 11.9% for 592B, where PCR duplicate rate = PCR duplicate read pairs / total sequenced read pairs). HiCCUPS was run at its default resolutions of 5, 10, and 25 kb with Knight–Ruiz (KR) matrix balancing normalization and the --ignore-sparsity flag enabled to ensure consistent loop detection across all chromosomes regardless of local contact density.

Genome-wide CTCF site prediction.

CTCF motifs were curated from multiple publicly available databases (Table S2) including JASPAR 2022 (61), HOCOMOCO v11 (62), SwissRegulon (63), Jolma2013 (64), CTCFBSDB (65), and CIS-BP (ver.3) (66), based on datasets used in the R/Bioconductor data package for CTCF binding in human and mouse (30). Each set of motifs was converted to MEME format using utilities from the MEME Suite (67), including transfac2meme, uniprobe2meme, and jaspar2meme. The resulting MEME files were consolidated into a single MEME-format file. To comprehensively identify all potential CTCF binding sites across the genome, we retained all motifs (Table S2) collected from the databases during FIMO scanning, rather than filtering for redundancy prior to analysis. This approach allows for the capture of subtle variations in motif representations across different databases, thereby enhancing sensitivity. To generate the optional background file for FIMO, we employed the fasta-get-markov utility from the MEME Suite on the mRatBN7.2 genome assembly (command: fasta-get-markov rn7.fa output_background_file.txt). The input MEME file with the background file was scanned against the rat reference genome (mRatBN7.2/rn7) using FIMO (v5.5.4) (23) from the MEME Suite (v5.5.4), installed via Conda (v4.1.0).

Transcription start sites (TSS), promoter, and gene annotation.

TSS annotations were derived from the GTF file (release version of 113) downloaded from ENSEMBL (https://ftp.ensembl.org/pub/release-113/gtf/rattus_norvegicus/Rattus_norvegicus.mRatBN7.2.113.gtf.gz). The initial dataset was subsequently refined by selecting tag and gene_biotype annotated as “Ensembl_canonical” and “protein_coding”, respectively. After removing redundant TSS assigned to the same gene, a final set of 21,725 unique TSS was retained, each corresponding to a distinct gene. Promoter coordinates were obtained from a bed file of Rn_EPDnew_001_rn6.bed (https://epd.expasy.org/ftp/epdnew/R_norvegicus/001/) in the Eukaryotic Promoter Database (EPD) for Rattus norvegicus (68). An initial set of 12,601 promoters based on the Rnor6.0 genome assembly was obtained. These coordinates of the promoters were converted to the mRatBN7.2 (rn7) reference genome using UCSC LiftOver tool (69) with a UCSC Chain file, rn6ToRn7.over.chain.gz (https://hgdownload.soe.ucsc.edu/goldenPath/rn6/liftOver/). Of the 12,533 promoters mapped to rn7, 12,529 unique entries were retained after removing redundant coordinates per gene. We defined the gene's genomic coordinate based on the coordinates of the last exon (i.e., the exon with the highest rank per transcript). The exon datasets were retrieved from the same GTF file above and from UCSC (https://hgdownload.soe.ucsc.edu/goldenPath/rn7/bigZips/genes/refGene.gtf.gz). This gene coordinate information was subsequently integrated with the TSS and promoter datasets. To associate these promoters with their corresponding gene coordinates, we first mapped EPD promoter identifiers to Ensembl Gene IDs provided by the EPD database using its cross-reference table (https://epd.expasy.org/ftp/epdnew/R_norvegicus/001/db/promoter_ensembl.txt).

Defining the flanking region of chromatin loops for examining the distribution of CTCF, TSS and promoter across loops.

To examine the distribution of genomic features across chromatin loops, we restricted the analysis to loops shorter than 2 megabases (Mb) (18, 24), resulting in the set of 31,019 from an initial non-redundant set of 31,773. We then defined the loop length (L) as the genomic distance between the midpoints of the two anchors in a loop, where each midpoint was calculated as the center between the coordinate pairs (e.g., x1, x2 or y1, y2) of each anchor, with the length of the anchor corresponding to the loop resolution. This definition ensures that L reflects the physical span of the loop between its central anchor positions. Next, we symmetrically extended the region by adding flanking segments of length (L) outward from each of the anchor midpoints. This resulted in a 3L-wide window: the original loop (L) plus two flanking regions (L on each side). These extended windows, which exceeded chromosome boundaries, were excluded (91 loops) because they distort the relative positions of CTCF motifs, TSS, and promoter elements. For the remaining 30,928 loops, we annotated the genomic positions of CTCF motifs, TSS, and promoter regions within the corresponding 3L windows. Their spatial distribution along these normalized windows was then quantified by converting each feature’s coordinate into a relative position using the following transformation:

Relative Position=3*(Feature coordinate-Loop Start)/L-1

Sequential filtering of annotated chromatin loops to define functional interactions.

To establish a high-confidence set of functionally relevant chromatin interactions, we applied a multi-step filtering pipeline (Figure 7). We began by aggregating only unique loops identified across the ten samples (27,219 loops removed), followed by filtering out 754 chromatin loops with a genomic distance exceeding 2Mb to ensure the biological relevance of cis-regulatory contacts (70–73), retaining a total of 31,019 loops for downstream analysis.

We then prioritized the loops based on the presence of the architectural protein CTCF, a key regulator of loop formation (28, 74, 75). Using FIMO-predicted CTCF motifs, we quantified the number of unique binding sites within each loop anchor. We retained only loops in which both anchors contained at least 6 CTCF binding sites (Figure 6). Applying this criterion yielded a total of 25,268 loops.

In parallel, we assigned TSS and promoter regions to chromatin loop anchors based on a proximity based mapping approach. Specifically, we first restricted the search to TSS and promoters located on the anchor facing side of the loop midpoint, extending outward along the chromosome. After applying this directional constraint, we assigned to each anchor the TSS or promoters with the shortest distance to its midpoint using the distanceToNearest function with the select = “all” option in GenomicRanges R package (76). Then, we applied a multi stage refinement procedure to obtain a single assignment per anchor for anchors associated with multiple candidate elements due to the “all” option. We removed candidates lacking gene annotations (i.e., candidates with unavailable gene coordinates), prioritized promoter annotations when both promoter and TSS labels were present for the same gene, enforced gene identity consistency through trimmed gene identifiers, and resolved remaining multi gene cases using transcript level features (including exon structure) with additional manual curation to remove low confidence or uninformative gene symbols (e.g., LOC annotations and Gzmbl2). These steps produced a unique gene assignment for nearly all loop anchors, yielding 61,969 anchors with a single matched gene and three exceptional loops in which both anchors retained two valid assignments (six anchors in total). In sum, this resulted in 61,975 anchor–regulatory element distance measurements. The distance distribution of these finalized assignments was plotted (Figure S8), and 200 kb was selected as the distance threshold.

Next, under the assumption that each chromatin loop corresponds to a single P-E interaction, we removed 13,043 loops in which neither anchor carried a valid TSS or promoter assignment, leaving 17,976 loops with at least one assigned anchor. Among these, 12,707 loops had only one anchor with a valid assignment and were directly retained. For the remaining 5,269 loops in which both anchors carried valid assignments, we selected the anchor whose assigned element was closer to its midpoint. After applying the 200 kb distance threshold to the finalized TSS/promoter-anchor assignments, 17,648 loops were retained.

To obtain a final set of functional loops, we intersected the 17,648 loops from TSS/promoter assignments with the 25,268 loops retained after CTCF motif-based filtering (≥6 binding sites per anchor). This last step yielded 15,085 high-confidence loops jointly supported by the two independent processes (Figure 7).

Quantification of loop-based interactions in gene-level and functional enrichment analysis.

Functional enrichment analysis was performed using g:Profiler (Figure 9) on the top 54 genes (Table S3). These genes were ranked by the highest number of valid loop interactions, ranging from a minimum of 11 to a maximum of 19 interactions per gene.

ACKNOWLEDGMENTS

We thank Rachel Ward for generating the Hi-C libraries. We gratefully acknowledge the contributions to the Hybrid Rat Diversity Panel (HRDP): Dr. Melinda R. Dwinell (Medical College of Wisconsin) provided the breeders; the Center for Integrative and Translational Genomics at the University of Tennessee Health Science Center (UTHSC) supported colony maintenance; and Mr. Angel Garcia Martinez and Ms. Caroline Jones assisted with breeding. Sequencing data were generated by the UTHSC Molecular Resource Center and University of Tennessee Genomics Core. Computational analyses utilized the University of Tennessee Infrastructure for Scientific Applications and Advanced Computing (ISAAC). This work was supported by grants U01 DA-053672 from NIH/NIDA (to B.M.S., R.W.W., and H.C.).

Footnotes

COMPETING INTEREST STATEMENT

The authors declare no competing interests.

REFERENCES

  • 1.Schoenfelder S. and Fraser P. (2019) Long-range enhancer-promoter contacts in gene expression control. Nat. Rev. Genet., 20, 437–455. [DOI] [PubMed] [Google Scholar]
  • 2.Lupiáñez D.G., Kraft K., Heinrich V., Krawitz P., Brancati F., Klopocki E., Horn D., Kayserili H., Opitz J.M., Laxova R., et al. (2015) Disruptions of topological chromatin domains cause pathogenic rewiring of gene-enhancer interactions. Cell, 161, 1012–1025. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Kohl J., Babayan B.M., Rubinstein N.D., Autry A.E., Marin-Rodriguez B., Kapoor V., Miyamishi K., Zweifel L.S., Luo L., Uchida N., et al. (2018) Functional circuit architecture underlying parental behaviour. Nature, 556, 326–331. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Chiel H.J. and Beer R.D. (1997) The brain has a body: adaptive behavior emerges from interactions of nervous system, body and environment. Trends Neurosci., 20, 553–557. [DOI] [PubMed] [Google Scholar]
  • 5.Sinha S., Jones B.M., Traniello I.M., Bukhari S.A., Halfon M.S., Hofmann H.A., Huang S., Katz P.S., Keagy J., Lynch V.J., et al. (2020) Behavior-related gene regulatory networks: A new level of organization in the brain. Proc. Natl. Acad. Sci. U. S. A., 117, 23270–23279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Shao Y., Bajikar S.S., Tirumala H.P., Gutierrez M.C., Wythe J.D. and Zoghbi H.Y. (2021) Identification and characterization of conserved noncoding cis-regulatory elements that impact Mecp2 expression and neurological functions. Genes Dev., 35, 489–494. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Guo M.G., Reynolds D.L., Ang C.E., Liu Y., Zhao Y., Donohue L.K.H., Siprashvili Z., Yang X., Yoo Y., Mondal S., et al. (2023) Integrative analyses highlight functional regulatory variants associated with neuropsychiatric diseases. Nat. Genet., 55, 1876–1891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Maurano M.T., Humbert R., Rynes E., Thurman R.E., Haugen E., Wang H., Reynolds A.P., Sandstrom R., Qu H., Brody J., et al. (2012) Systematic localization of common disease-associated variation in regulatory DNA. Science, 337, 1190–1195. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Iannaccone P.M. and Jacob H.J. (2009) Rats! Dis. Model. Mech., 2, 206–210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Parker C.C., Chen H., Flagel S.B., Geurts A.M., Richards J.B., Robinson T.E., Solberg Woods L.C. and Palmer A.A. (2013) Rats are the smart choice: Rationale for a renewed focus on rats in behavioral genetics. Neuropharmacology, 76 Pt B, 250–258. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.de Jong T.V., Chen H., Brashear W.A., Kochan K.J., Hillhouse A.E., Zhu Y., Dhande I.S., Hudson E.A., Sumlut M.H., Smith M.L., et al. (2022) mRatBN7.2: familiar and unfamiliar features of a new rat genome reference assembly. Physiol. Genomics, 54, 251–260. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Howe K., Dwinell M., Shimoyama M., Corton C., Betteridge E., Dove A., Quail M.A., Smith M., Saba L., Williams R.W., et al. (2021) The genome sequence of the Norway rat, Rattus norvegicus Berkenhout 1769. Wellcome Open Res, 6, 118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Li K., Smith M.L., Blazier J.C., Kochan K.J., Wood J.M.D., Howe K., Kwitek A.E., Dwinell M.R., Chen H., Ciosek J.L., et al. (2024) Construction and evaluation of a new rat reference genome assembly, GRCr8, from long reads and long-range scaffolding. Genome Res., 34, 2081–2093. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.de Jong T.V., Pan Y., Rastas P., Munro D., Tutaj M., Akil H., Benner C., Chen D., Chitre A.S., Chow W., et al. (2024) A revamped rat reference genome improves the discovery of genetic diversity in laboratory rats. Cell Genom., 4, 100527. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Cox O.H., Seifuddin F., Guo J., Pirooznia M., Boersma G.J., Wang J., Tamashiro K.L.K. and Lee R.S. (2024) Implementation of the Methyl-Seq platform to identify tissue- and sexspecific DNA methylation differences in the rat epigenome. Epigenetics, 19, 2393945. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Chen Z., Snetkova V., Bower G., Jacinto S., Clock B., Dizehchi A., Barozzi I., Mannion B.J., Alcaina-Caro A., Lopez-Rios J., et al. (2024) Increased enhancer-promoter interactions during developmental enhancer activation in mammals. Nat. Genet., 56, 675–685. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Dixon J.R., Selvaraj S., Yue F., Kim A., Li Y., Shen Y., Hu M., Liu J.S. and Ren B. (2012) Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature, 485, 376–380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Rao S.S.P., Huntley M.H., Durand N.C., Stamenova E.K., Bochkov I.D., Robinson J.T., Sanborn A.L., Machol I., Omer A.D., Lander E.S., et al. (2014) A 3D Map of the Human Genome at Kilobase Resolution Reveals Principles of Chromatin Looping. Cell, 159, 1665– 1680. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Dwinell M.R., Takizawa A., Tutaj M., Malloy L., Schilling R., Endsley A., Demos W.M., Smith J.R., Wang S.J., De Pons J., et al. (2025) Establishing the hybrid rat diversity program: a resource for dissecting complex traits. Mamm. Genome, 36, 25–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Tabakoff B., Smith H., Vanderlinden L.A., Hoffman P.L. and Saba L.M. (2019) Networking in biology: The Hybrid Rat Diversity Panel. Methods Mol. Biol., 2018, 213–231. [DOI] [PubMed] [Google Scholar]
  • 21.Kim P., Ward R.R., Sharp B.M., Williams R.W. and Chen H. (2025) Hi-C sequencing data from frontal cortex of laboratory rats. Sci. Data, 12, 1910. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Durand N.C., Shamim M.S., Machol I., Rao S.S.P., Huntley M.H., Lander E.S. and Aiden E.L. (2016) Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. cels, 3, 95–98. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Grant C.E., Bailey T.L. and Noble W.S. (2011) FIMO: scanning for occurrences of a given motif. Bioinformatics, 27, 1017–1018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Reyna J., Fetter K., Ignacio R., Marandi C.C.A., Ma A., Rao N., Jiang Z., Figueroa D.S., Bhattacharyya S. and Ay F. (2025) Loop Catalog: a comprehensive HiChIP database of human and mouse samples. Genome Biol., 26, 175. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Kolberg L., Raudvere U., Kuzmin I., Adler P., Vilo J. and Peterson H. (2023) g:Profilerinteroperable web service for functional enrichment analysis and gene identifier mapping (2023 update). Nucleic Acids Res., 51, W207–W212. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Gosalia N. and Harris A. (2015) Chromatin dynamics in the regulation of CFTR expression. Genes (Basel), 6, 543–558. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cermakova K. and Hodges H.C. (2018) Next-generation drugs and probes for chromatin biology: From targeted protein degradation to phase separation. Molecules, 23, 1958. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Sanborn A.L., Rao S.S.P., Huang S.-C., Durand N.C., Huntley M.H., Jewett A.I., Bochkov I.D., Chinnappan D., Cutkosky A., Li J., et al. (2015) Chromatin extrusion explains key features of loop and domain formation in wild-type and engineered genomes. Proc. Natl. Acad. Sci. U. S. A., 112, E6456–65. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Rodríguez-Carballo E., Lopez-Delisle L., Willemin A., Beccari L., Gitto S., Mascrez B. and Duboule D. (2020) Chromatin topology and the timing of enhancer function at the HoxD locus. Proc. Natl. Acad. Sci. U. S. A., 117, 31231–31241. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Dozmorov M.G., Mu W., Davis E.S., Lee S., Triche T.J.,Jr, Phanstiel,D.H. and Love,M.I. (2022) CTCF: an R/bioconductor data package of human and mouse CTCF binding sites. Bioinform. Adv., 2, vbac097. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Vietri Rudan M., Barrington C., Henderson S., Ernst C., Odom D.T., Tanay A. and Hadjur S. (2015) Comparative Hi-C reveals that CTCF underlies evolution of chromosomal domain architecture. Cell Rep., 10, 1297–1309. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.de Wit E., Vos E.S.M., Holwerda S.J.B., Valdes-Quezada C., Verstegen M.J.A.M., Teunissen H., Splinter E., Wijchers P.J., Krijger P.H.L. and de Laat W. (2015) CTCF binding polarity determines chromatin looping. Mol. Cell, 60, 676–684. [DOI] [PubMed] [Google Scholar]
  • 33.Richter W.F., Nayak S., Iwasa J. and Taatjes D.J. (2022) The Mediator complex as a master regulator of transcription by RNA polymerase II. Nat. Rev. Mol. Cell Biol., 23, 732–749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Yang J., McGovern A., Martin P., Duffus K., Ge X., Zarrineh P., Morris A.P., Adamson A., Fraser P., Rattray M., et al. (2020) Analysis of chromatin organization and gene expression in T cells identifies functional genes for rheumatoid arthritis. Nat. Commun., 11, 4402. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Kvon E.Z., Waymack R., Gad M. and Wunderlich Z. (2021) Enhancer redundancy in development and disease. Nat. Rev. Genet., 22, 324–336. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Delpretti S., Montavon T., Leleu M., Joye E., Tzika A., Milinkovitch M. and Duboule D. (2013) Multiple enhancers regulate Hoxd genes and the Hotdog LncRNA during cecum budding. Cell Rep., 5, 137–150. [DOI] [PubMed] [Google Scholar]
  • 37.Rodríguez-Carballo E., Lopez-Delisle L., Zhan Y., Fabre P.J., Beccari L., El-Idrissi I., Huynh T.H.N., Ozadam H., Dekker J. and Duboule D. (2017) The HoxD cluster is a dynamic and resilient TAD boundary controlling the segregation of antagonistic regulatory landscapes. Genes Dev., 31, 2264–2281. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Tsai A., Alves M.R. and Crocker J. (2019) Multi-enhancer transcriptional hubs confer phenotypic robustness. Elife, 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Osterwalder M., Barozzi I., Tissières V., Fukuda-Yuzawa Y., Mannion B.J., Afzal S.Y., Lee E.A., Zhu Y., Plajzer-Frick I., Pickle C.S., et al. (2018) Enhancer redundancy provides phenotypic robustness in mammalian development. Nature, 554, 239–243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Bejjani F., Evanno E., Mahfoud S., Tolza C., Zibara K., Piechaczyk M. and Jariel-Encontre I. (2023) Multiple Fra-1-bound enhancers showing different molecular and functional features can cooperate to repress gene transcription. Cell Biosci., 13, 129. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Waymack R., Fletcher A., Enciso G. and Wunderlich Z. (2020) Shadow enhancers can suppress input transcription factor noise through distinct regulatory logic. Elife, 9, e59351. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Frankel N., Davis G.K., Vargas D., Wang S., Payre F. and Stern D.L. (2010) Phenotypic robustness conferred by apparently redundant transcriptional enhancers. Nature, 466, 490– 493. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Tobias I.C., Moorthy S.D., Shchuka V.M., Langroudi L., Cherednychenko M., Gillespie Z.E., Duncan A.G., Tian R., Gajewska N.A., Di Roberto R.B., et al. (2025) A Sox2 enhancer cluster regulates region-specific neural fates from mouse embryonic stem cells. G3 (Bethesda), 15. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Rosenberg A.B., Roco C.M., Muscat R.A., Kuchina A., Sample P., Yao Z., Graybuck L.T., Peeler D.J., Mukherjee S., Chen W., et al. (2018) Single-cell profiling of the developing mouse brain and spinal cord with split-pool barcoding. Science, 360, 176–182. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Zeisel A., Hochgerner H., Lönnerberg P., Johnsson A., Memic F., van der Zwan J., Häring M., Braun E., Borm L.E., La Manno G., et al. (2018) Molecular architecture of the mouse nervous system. Cell, 174, 999–1014.e22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Geller E., Noble M.A., Morales M., Gockley J., Emera D., Uebbing S., Cotney J.L. and Noonan J.P. (2024) Massively parallel disruption of enhancers active in human neural stem cells. Cell Rep., 43, 113693. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Gunturkun M.H., Flashner E., Wang T., Mulligan M.K., Williams R.W., Prins P. and Chen H. (2022) GeneCup: mining PubMed and GWAS catalog for gene-keyword relationships. G3 (Bethesda), 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Kajiwara Y., Wang E., Wang M., Sin W.C., Brennand K.J., Schadt E., Naus C.C., Buxbaum J. and Zhang B. (2018) GJA1 (connexin43) is a key regulator of Alzheimer’s disease pathogenesis. Acta Neuropathol. Commun., 6, 144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Ren D., Zheng P., Feng J., Gong Y., Wang Y., Duan J., Zhao L., Deng J., Chen H., Zou S., et al. (2020) Overexpression of astrocytes-specific GJA1–20k enhances the viability and recovery of the neurons in a rat model of traumatic brain injury. ACS Chem. Neurosci., 11, 1643–1650. [DOI] [PubMed] [Google Scholar]
  • 50.Goris A., Williams-Gray C.H., Foltynie T., Brown J., Maranian M., Walton A., Compston D.A.S., Barker R.A. and Sawcer S.J. (2007) Investigation of TGFB2 as a candidate gene in multiple sclerosis and Parkinson’s disease. J. Neurol., 254, 846–848. [DOI] [PubMed] [Google Scholar]
  • 51.Zhu Y., Geng X., Stone C., Guo S., Syed S. and Ding Y. (2022) Forkhead Box 1(FoxO1) mediates psychological stress-induced neuroinflammation. Neurol. Res., 44, 483–495. [DOI] [PubMed] [Google Scholar]
  • 52.Cattane N., Di Benedetto M.G., D’Aprile I., Riva M.A. and Cattaneo A. (2024) Dissecting the long-term effect of stress early in life on FKBP5: The role of miR-20b-5p and miR-29c-3p. Biomolecules, 14, 371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.de Guglielmo G., Carrette L., Kallupi M., Brennan M., Boomhower B., Maturin L., Conlisk D., Sedighim S., Tieu L., Fannon M.J., et al. (2024) Large-scale characterization of cocaine addiction-like behaviors reveals that escalation of intake, aversion-resistant responding, and breaking-points are highly correlated measures of the same construct. Elife, 12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Gunturkun M.H., Wang T., Chitre A.S., Garcia Martinez A., Holl K., St Pierre C., Bimschleger H., Gao J., Cheng R., Polesskaya O., et al. (2022) Genome-wide association study on three behaviors tested in an open field in heterogeneous stock rats identifies multiple loci implicated in psychiatric disorders. Front. Psychiatry, 13, 790566. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Sey N.Y.A., Hu B., Mah W., Fauni H., McAfee J.C., Rajarajan P., Brennand K.J., Akbarian S. and Won H. (2020) A computational tool (H-MAGMA) for improved prediction of brain-disorder risk genes by incorporating brain chromatin interaction profiles. Nat. Neurosci., 23, 583–593. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Duttke S.H., Montilla-Perez P., Chang M.W., Li H., Chen H., Carrette L.L.G., de Guglielmo G., George O., Palmer A.A., Benner C., et al. (2022) Glucocorticoid receptorregulated enhancers play a central role in the gene regulatory networks underlying drug addiction. Front. Neurosci., 16, 858427. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Segert J.A., Gisselbrecht S.S. and Bulyk M.L. (2021) Transcriptional silencers: Driving gene expression with the brakes on. Trends Genet., 37, 514–527. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Parker S.M., Davis E.S. and Phanstiel D.H. (2023) Guiding the design of well-powered Hi-C experiments to detect differential loops. Bioinform. Adv., 3, vbad152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Miller E.K. and Cohen J.D. (2001) An integrative theory of prefrontal cortex function. Annu. Rev. Neurosci., 24, 167–202. [DOI] [PubMed] [Google Scholar]
  • 60.Bonev B., Cohen N.M., Szabo Q., Fritsch L., Papadopoulos G.L., Lubling Y., Xu X., Lv X., Hugnot J.-P., Tanay A., et al. (2017) Multiscale 3D Genome Rewiring during Mouse Neural Development. Cell, 171, 557–572.e24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Castro-Mondragon J.A., Riudavets-Puig R., Rauluseviciute I., Lemma R.B., Turchi L., BlancMathieu R., Lucas J., Boddie P., Khan A., Manosalva Pérez N., et al. (2022) JASPAR 2022: the 9th release of the open-access database of transcription factor binding profiles. Nucleic Acids Res., 50, D165–D173. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Kulakovskiy I.V., Vorontsov I.E., Yevshin I.S., Sharipov R.N., Fedorova A.D., Rumynskiy E.I., Medvedeva Y.A., Magana-Mora A., Bajic V.B., Papatsenko D.A., et al. (2018) HOCOMOCO: towards a complete collection of transcription factor binding models for human and mouse via large-scale ChIP-Seq analysis. Nucleic Acids Res., 46, D252– D259. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Pachkov M., Balwierz P.J., Arnold P., Ozonov E. and van Nimwegen E. (2013) SwissRegulon, a database of genome-wide annotations of regulatory sites: recent updates. Nucleic Acids Res., 41, D214–20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Jolma A., Yan J., Whitington T., Toivonen J., Nitta K.R., Rastas P., Morgunova E., Enge M., Taipale M., Wei G., et al. (2013) DNA-binding specificities of human transcription factors. Cell, 152, 327–339. [DOI] [PubMed] [Google Scholar]
  • 65.Ziebarth J.D., Bhattacharya A. and Cui Y. (2013) CTCFBSDB 2.0: a database for CTCFbinding sites and genome organization. Nucleic Acids Res., 41, D188–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Weirauch M.T., Yang A., Albu M., Cote A.G., Montenegro-Montero A., Drewe P., Najafabadi H.S., Lambert S.A., Mann I., Cook K., et al. (2014) Determination and inference of eukaryotic transcription factor sequence specificity. Cell, 158, 1431–1443. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Bailey T.L., Boden M., Buske F.A., Frith M., Grant C.E., Clementi L., Ren J., Li W.W. and Noble W.S. (2009) MEME SUITE: tools for motif discovery and searching. Nucleic Acids Res., 37, W202–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Dreos R., Ambrosini G., Groux R., Cavin Périer R. and Bucher P. (2017) The eukaryotic promoter database in its 30th year: focus on non-vertebrate organisms. Nucleic Acids Res., 45, D51–D55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Hinrichs A.S., Karolchik D., Baertsch R., Barber G.P., Bejerano G., Clawson H., Diekhans M., Furey T.S., Harte R.A., Hsu F., et al. (2006) The UCSC Genome Browser Database: update 2006. Nucleic Acids Res., 34, D590–8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.He B., Chen C., Teng L. and Tan K. (2014) Global view of enhancer-promoter interactome in human cells. Proc. Natl. Acad. Sci. U. S. A., 111, E2191–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Laverré A., Tannier E. and Necsulea A. (2022) Long-range promoter-enhancer contacts are conserved during evolution and contribute to gene expression robustness. Genome Res., 32, 280–296. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Perlman B.S., Burget N., Zhou Y., Schwartz G.W., Petrovic J., Modrusan Z. and Faryabi R.B. (2024) Enhancer-promoter hubs organize transcriptional networks promoting oncogenesis and drug resistance. Nat. Commun., 15, 8070. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Bower G., Hollingsworth E.W., Jacinto S.H., Alcantara J.A., Clock B., Cao K., Liu M., Dziulko A., Alcaina-Caro A., Xu Q., et al. (2025) Range extender mediates long-distance enhancer activity. Nature, 643, 830–838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74.Bonev B. and Cavalli G. (2016) Organization and function of the 3D genome. Nat. Rev. Genet., 17, 661–678. [DOI] [PubMed] [Google Scholar]
  • 75.Hansen A.S., Pustova I., Cattoglio C., Tjian R. and Darzacq X. (2017) CTCF and cohesin regulate chromatin loop stability with distinct dynamics. Elife, 6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Lawrence M., Huber W., Pagès H., Aboyoun P., Carlson M., Gentleman R., Morgan M.T. and Carey V.J. (2013) Software for computing and annotating genomic ranges. PLoS Comput. Biol., 9, e1003118. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES