Abstract
The zinc finger–associated domain (ZAD)–containing C2H2 zinc finger proteins (ZAD-ZnFs) represent the most abundant class of transcription factors that emerged during insect evolution, yet their molecular diversity and biological functions remain largely unclear. Here, we established a systematic CRISPR-based protein-tagging approach that enables direct, unambiguous comparison of nuclear localization and genome-wide binding profiles of endogenous ZAD-ZnFs in developing Drosophila embryos. Evidence is provided that a subset of ZAD-ZnFs forms nuclear condensates through the stacking of the N-terminal ZAD dimerization surface. Disruption of condensation activity leads to misregulation of genome-wide binding profiles and lethality, underscoring its functional and physiological significance in development. Integrative chromatin immunoprecipitation sequencing and Micro-C analyses reveal that many ZAD-ZnFs colocalize with core insulator proteins such as CCCTC-binding factor and Centrosomal protein 190 kD to control the formation of topological boundaries. We suggest that the diverse molecular functions of ZAD-ZnFs have evolutionarily arisen from their ancestral role as insulator-binding proteins.
Identification of ZAD-ZnF genes as key regulators of genome organization during Drosophila embryogenesis.
INTRODUCTION
The zinc finger–associated domain (ZAD)–containing C2H2 zinc finger proteins (ZAD-ZnFs) represent the most abundant class of transcription factors that has emerged during insect evolution (1). ZAD-ZnFs consist of an N-terminal ZAD and arrays of C2H2 zinc finger domains at the C terminus, flanked by flexible linker regions. Similar to the POZ/BTB and SCAN domains in mammals and other species, the N-terminal ZAD functions as a protein dimerization scaffold through the stacking of the hydrophobic surface (2–6). More than 90 ZAD-ZnF genes are present in the Drosophila melanogaster genome, but the number of ZAD-ZnF genes greatly varies across related Drosophila species (7). Despite such rapid evolutionary dynamics (1, 7, 8), ZAD-ZnF genes often encode proteins with essential biological functions. For example, ZAD-ZnF genes Molting Defective, Ouija Board, and Séance encode transcriptional activators responsible for the expression of essential ecdysone biosynthesis genes during larval development (9, 10). On the other hand, evolutionarily young ZAD-ZnF genes Oddjob (Odj) and Nicknack (Nnk) as well as newly identified Identity crisis (Idc) gene encode DNA binding proteins that associate with heterochromatic regions enriched with the repressive H3K9me3 histone mark (7, 11, 12). The ZAD-ZnFs Kipferl (Kipf) and Trailblazer were recently reported to function as positive regulators of the Piwi–Piwi-interacting RNA (piRNA)–mediated transposon silencing pathway in reproductive tissues (13, 14). In addition, the ZAD-ZnFs Motif 1 Binding Protein (M1BP), Pita, Zinc-finger protein interacting with CP190 (ZIPIC), and Zeste-white 5 (Zw5) are thought to help organize higher-order genome topology through the association with insulator elements (15–17). As an exceptional case, a ZAD-ZnF gene weckle (wek) is suggested to encode a membrane-localized cytoplasmic protein that acts as an adaptor protein in the Toll pathway during the establishment of the dorsal-ventral axis in early embryos (18). Overall, these findings are consistent with the idea that individual ZAD-ZnFs have functionally diverged to control widespread biological processes during the evolution of Drosophila species under natural environments. However, the evolutionary origin and common ancestral features of ZAD-ZnFs remain uncertain.
As exemplified above, extensive efforts have been made to elucidate the biological roles and molecular properties of individual ZAD-ZnF genes over decades. However, we are still far from a comprehensive understanding of their in vivo functions because the systematic characterization of endogenous ZAD-ZnFs in a unified experimental setup is critically lacking. To overcome this hurdle, we established a unified protein-tagging system that enables comprehensive, direct comparison of nuclear localization, genome-wide binding profiles, and in vivo functionalities of endogenous ZAD-ZnFs in developing Drosophila embryos. CRISPR-Cas9–mediated insertion of a DNA cassette encoding a green fluorescent protein (GFP)–3xFLAG tag immediately upstream of the stop codon permits super-resolution live-imaging and chromatin immunoprecipitation sequencing (ChIP-seq) analyses of individual ZAD-ZnFs by minimizing potential experimental noise arising from the fixation protocol for immunostaining (19) and the variability of antibody quality for ChIP experiments (20, 21). Using our unique experimental framework, we identified a distinct subset of ZAD-ZnFs that forms nuclear condensates via the N-terminal ZAD dimerization surface in living embryos. Disruption of condensation activity leads to misregulation of genome-wide binding profiles and lethality, underscoring its functional and physiological significance in development. We also show that a class of ZAD-ZnFs exhibiting uniform nuclear localization does not rely on the N-terminal ZAD to exert their biological functions during embryogenesis, highlighting rapid functional and molecular divergence within this family. Furthermore, integrative Micro-C and ChIP-seq analyses reveal that many ZAD-ZnFs colocalize with well-characterized core insulator proteins such as CCCTC-binding factor (CTCF), Suppressor of Hairy-wing [Su(Hw)], Boundary element–associated factor of 32kD (BEAF-32), and Centrosomal protein 190kD (CP190) to control the formation of topological boundaries, suggesting that insulator-binding activity is a common ancestral feature of ZAD-ZnFs. We suggest that ZAD-ZnFs constitute a class of genome organizers with diverse molecular properties, likely shaped by rapid functional divergence during insect evolution.
RESULTS
Establishment of CRISPR-Cas9–mediated protein-tagging system
To systematically visualize the localization patterns of endogenous proteins in living Drosophila, we individually inserted a sequence cassette encoding GFP-3xFLAG or mCherry-3xFLAG tag immediately upstream of the stop codon of target genes (fig. S1A). Resulting genome-edited strains were then crossed with a Cre-expressing strain to remove the 3xP3-dsRed selection marker from the targeted locus. Insertion of the desired sequence cassette has been verified by polymerase chain reaction (PCR) analysis of purified genomic DNA for all the genome-edited strains produced in this study. As a test example, we initially targeted the well-characterized vasa gene, which is essential for the development of ovarian tissues through the formation of a perinuclear compartment called nuage in nurse cells (22, 23). A Western blotting signal against an anti-FLAG antibody was specifically seen upon genome editing (fig. S1B). In addition, nuage formation was reproducibly observed by live-imaging analysis of Vasa-mCherry-3xFLAG fusion protein in dissected ovaries (fig. S1C). The embryo hatching rate was essentially unchanged even after genome editing of the vasa gene (fig. S1D), supporting the idea that our CRISPR-Cas9–mediated tagging approach enables live visualization of protein products of endogenous key developmental genes in a noninterfering manner.
As another test example, we next focused on the well-characterized DNA binding protein GAGA factor (GAF), which is encoded by the Trithorax-like (Trl) gene in Drosophila (24, 25). A number of imaging studies reported that GAF protein forms a cluster or condensate within a nucleus in developing embryos [e.g., (26–29)]. A sequence cassette encoding either GFP-3xFLAG or mCherry-3xFLAG tag has been individually inserted immediately upstream of the stop codon at the endogenous Trl locus. Resulting genome-edited strains were both homozygous viable, suggesting that the inserted GFP-3xFLAG or mCherry-3xFLAG tag does not impede either transcription or translation. As in the case of vasa, a clear Western blotting signal against an anti-FLAG antibody was specifically detected upon genome editing (fig. S1E). Consistent with previous literature (27–29), both GAF-GFP-3xFLAG and GAF-mCherry-3xFLAG formed a cluster within a nucleus in developing embryos at nuclear cycle 14 (nc14) (fig. S1F), indicating that the unique localization patterns of endogenous protein products are not affected by the choice of fluorescent proteins. Condensate-like localization was not seen when GFP-3xFLAG tag alone was expressed in conjunction with an SV40 nuclear localization signal (fig. S1G), ruling out the possibility that the inserted GFP-3xFLAG tag itself drives condensate assembly in a nonspecific manner.
Visualization of endogenous insulator proteins in early embryos
Having established our CRISPR-Cas9–mediated protein-tagging strategy, we first sought to explore the localization patterns of so-called insulator proteins, a class of nuclear proteins that plays a central role in organizing the three-dimensional (3D) genome structure in Drosophila [reviewed in Bortle and Corces (30)]. In this study, we particularly focused on six well-characterized insulator proteins, namely Su(Hw), CTCF, BEAF-32, CP190, Lethal (3) malignant brain tumor [L(3)mbt], and Female sterile (1) homeotic long isoform [Fs(1)h-L] (31–36). As in the case of GAF, a GFP-3xFLAG tag has been inserted immediately upstream of the stop codon of individual target genes. All the resulting genome-edited strains were homozygous viable, suggesting that the GFP-3xFLAG tag does not impede the functionality of these targeted genes. Successful tagging was confirmed by Western blot analysis using anti-FLAG antibody (fig. S2A). Previous studies suggested that insulator proteins form condensates called “insulator bodies” within the nucleus of dissected imaginal discs or in cultured cells [e.g., (34, 37–42)]. However, it remains uncertain whether endogenous insulator proteins actually form “insulator bodies” in living Drosophila, especially given a recent finding that the traditional fixation protocol for immunostaining can significantly alter the appearance of protein condensates (19). Airyscan super-resolution live-imaging analysis of the genome-edited strains revealed that each insulator protein exhibits distinct patterns of nuclear localization in early embryos at the same developmental stage (~15 min after entry into nc14) (Fig. 1A). For example, Su(Hw) has long been postulated as a major constituent of “insulator bodies” (37–39, 42), but its localization pattern was more uniform than those of other insulator proteins such as CP190 and Fs(1)h-L (Fig. 1A). In addition, unlike these proteins, CTCF and L(3)mbt formed only a few bright foci per nucleus, while BEAF-32 was more uniformly distributed in the cytoplasm and nucleus without assembling clear nuclear foci (Fig. 1A). Together, these results suggest that it is not likely that all the tested insulator proteins coalesce to form common “insulator bodies” to regulate 3D genome organization in developing embryos. This view is consistent with a recent ChIP-seq study reporting distinct distribution profiles of insulator proteins in nc14 embryos (43). We will use our genome-edited insulator strains for subsequent ChIP-seq analysis, which will be discussed in the later part of this study.
Fig. 1. Live visualization of endogenous insulator proteins and ZAD-ZnFs in early embryos.
(A) Airyscan imaging of GFP-3xFLAG–tagged endogenous insulator proteins in living embryos. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous genome-edited strains were used for the analysis. The maximum-intensity–projected images are shown. (B) Airyscan imaging of GFP-3xFLAG–tagged endogenous ZAD-ZnFs in living embryos. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous genome-edited strains were used for the analysis. The maximum-intensity–projected images are shown.
Systematic visualization of endogenous ZAD-ZnFs in early embryos
To systematically characterize the localization patterns of ZAD-ZnFs in Drosophila, we targeted 14 ZAD-ZnF genes expressed in early embryos for CRISPR-Cas9–mediated insertion of the GFP-3xFLAG tag at the C-terminal end (fig. S1A). Successful tagging was confirmed by Western blot analysis using anti-FLAG antibody (fig. S2B). While all the ZAD-ZnFs share a common structural arrangement (i.e., an N-terminal ZAD followed by a flexible unstructured linker and an array of C-terminal C2H2 zinc finger domains), their nuclear localization was found to be very different (Fig. 1B). For example, Nnk, Odj, M1BP, Wek, CG17361, and CG30020 reproducibly formed a cluster or condensate within a nucleus, while others [Idc, Zinc-finger protein (Zif), D19B, Trade embargo (Trem), CG4282, Kipf, Zw5, and D19A] were more uniformly distributed in nc14 embryos. Among these, M1BP and Zw5 are well characterized as insulator proteins (16, 17, 44, 45). We noticed that the localization patterns of M1BP and Zw5 are distinct from those of non–ZAD-ZnF insulator proteins such as CP190, Fs(1)h-L, and others at the same developmental stage (Fig. 1A), further reinforcing the idea that hypothetical “insulator bodies” do not represent cocondensation of all insulator proteins at the same site to shape 3D genome topology. Another notable finding here is the nuclear localization of Wek in nc14 embryos (Fig. 1B) because a previous GAL4/UAS-driven transgene analysis reported that Wek localizes to the plasma membrane and acts as an adaptor of the Toll pathway to control dorsal-ventral patterning at this developmental stage (18). Single-molecule RNA fluorescence in situ hybridization (FISH) coupled with immunofluorescence confirmed that the expression profiles of the key dorsal-ventral patterning genes twist (twi) and zerknüllt (zen), as well as the nuclear localization of the Dorsal activator, are not compromised in our Wek-GFP knock-in strain (fig. S3A). Nuclear localization of endogenous Wek in nc14 embryos was also seen when only a minimal 2xHA tag was fused to the C terminus of the protein (fig. S3B), ruling out the possibility that the relatively large GFP tag prevented its membrane localization. Our subsequent ChIP-seq analysis further supports the idea that Wek acts as a nuclear DNA binding protein that associates with insulator elements through recognition of a specific sequence motif (Figs. 7 and 8 and fig. S9), challenging the traditional view that Wek exerts its function as a membrane-localized adaptor protein of the Toll pathway in early embryos.
Fig. 7. Topological boundaries are associated with ZAD-ZnF cluster.
(A and B) Micro-C map along with ChIP-seq profiles of ZAD-ZnFs and non–ZAD-ZnF insulator proteins at the Abd-B/abd-A Hox locus (A) and the Scr/ftz Hox locus (B). Publicly available Micro-C data in nc14 embryos [Batut et al. (53); SRR14136009] were used for the analysis.
Fig. 8. Intrinsic insulator-binding activity of ZAD-ZnFs.
(A) Heatmap visualization of BEAF-32, CTCF, Su(Hw), and CP190 ChIP-seq distribution sorted by the strength of BEAF-32 (top panels), CTCF (top middle panel), Su(Hw) (bottom middle panel), or CP190 (bottom panels) ChIP-seq peaks. (B) Heatmap visualization of ChIP-seq distribution of non–ZAD-ZnF insulator proteins and ZAD-ZnFs sorted by the strength of BEAF-32 ChIP-seq peaks. (C) Heatmap visualization of ChIP-seq distribution of non–ZAD-ZnF insulator proteins and ZAD-ZnFs sorted by the strength of CP190 ChIP-seq peaks. (D) Topological boundaries are divided into quantiles according to the boundary strength calculated from the insulation score (53). The numbers of proteins containing ChIP-seq peaks were determined for individual boundaries. List of proteins used for the analysis is shown in fig. S12B. In total, 336 (Q1), 336 (Q2), 337 (Q3), and 337 (Q4) boundaries were analyzed. The box indicates the lower (25%) and upper (75%) quantile, and the black line indicates the median. The P values of two-sided Wilcoxon rank sum test are shown.
ZAD is required for the in vivo function of Nnk and CG30020/Brother of Nnk
To elucidate the mechanism underlying the clustering of ZAD-ZnFs and its biological significance, we decided to focus on two ZAD-ZnFs, Nnk and CG30020, as they form clear clusters within a nucleus (Fig. 1B). Previous biochemical and x-ray crystallography studies reported that the N-terminal ZAD mediates homotypic protein-protein interactions through the stacking of a hydrophobic surface (fig. S4, A and B, left) (2, 5, 6). We initially tested whether this dimerization activity of ZAD could be reliably predicted using the AlphaFold3 algorithm (46). It turned out that the previously reported x-ray crystal structure of ZAD dimers was nicely recapitulated, including the positioning of the flexible regions and the Zn2+ ion coordination sites (fig. S4, A and B, right). Then, the AlphaFold3 algorithm was applied for the analysis of Nnk. It was predicted that its ZAD also forms a dimer with high confidence (Fig. 2, A and B), in a way very similar to the experimentally verified Grauzone and M1BP ZAD dimers (fig. S4, A and B). To examine the biological role of ZAD, we individually expressed Nnk wild-type (WT) and ΔZAD mutant as a GFP-3xFLAG fusion protein under the control of the native Nnk promoter (Fig. 2C). Cluster formation of Nnk was completely abolished upon the loss of the N-terminal ZAD (Fig. 2D), indicating that the dimerization activity of ZAD is required for the assembly of the Nnk cluster in early embryos. We then analyzed CG30020 following the same experimental strategy (Fig. 2E). As in the case of Nnk, AlphaFold3 predicted that the CG30020 ZAD also serves as a dimerization scaffold (Fig. 2F). When CG30020 WT and ΔZAD mutant were individually expressed as a GFP-3xFLAG fusion protein under the control of the native CG30020 promoter (Fig. 2G), only the WT construct was able to assemble a cluster within a nucleus (Fig. 2H). On the basis of molecular similarities between CG30020 and Nnk, we named CG30020 as Brother of Nnk or BroN (Fig. 2E).
Fig. 2. The N-terminal ZAD is essential for Nnk and BroN functions.
(A) Schematic representation of the full-length Nnk protein. (B) AlphaFold3 prediction of the Nnk ZAD dimer. Protein structure was visualized using UCSF ChimeraX (104). pTM and iPTM values are used as proxies for the accuracy of predicting the overall protein folding and the relative positioning of the monomers within the dimer, respectively (hereinafter the same). (C) Western blot analysis of GFP-3xFLAG–tagged endogenous Nnk and exogenously expressed Nnk rescue constructs. Two- to 4-hour embryos from corresponding homozygous strains were used for the analysis. As a control, 2- to 4-hour yw embryos were used. (D) Airyscan imaging of GFP-3xFLAG–tagged WT and ΔZAD Nnk expressed under the control of the native Nnk promoter. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous strains were used for the analysis. The maximum-intensity–projected images are shown. (E) Schematic of the full-length CG30020/BroN protein. (F) AlphaFold3 prediction of the BroN ZAD dimer. Protein structure was visualized using UCSF ChimeraX (104). (G) Western blot analysis of GFP-3xFLAG–tagged endogenous BroN and exogenously introduced BroN rescue constructs. Two- to 4-hour embryos from corresponding homozygous strains were used for the analysis. As a control, 2- to 4-hour yw embryos were used. (H) Airyscan imaging of GFP-3xFLAG–tagged WT and ΔZAD BroN expressed under the control of the native BroN promoter. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous strains were used for the analysis. The maximum-intensity–projected images are shown. (I) Full-locus deletion alleles (Nnk[Δ1] and BroN[Δ1]) were produced by CRISPR-Cas9–mediated genome editing. (J and K) ZAD-deficient transgenes failed to restore the lethal phenotype of homozygous Nnk[Δ1] (J) and BroN[Δ1] (K) mutants.
To further characterize the in vivo functionalities of Nnk and BroN genes, we used CRISPR-Cas9–mediated genome editing to remove the entire transcription unit from the endogenous loci (Fig. 2I; Nnk[Δ1] and BroN[Δ1] denote the resulting full-locus deletion alleles). Consistent with previous RNAi-mediated knockdown experiments (7), the homozygous Nnk[Δ1] mutant exhibited a complete lethal phenotype despite of its young evolutionary age (Fig. 2J). Transgene rescue assays demonstrated that only the ZAD-containing WT construct was able to restore viability (Fig. 2J), implicating that ZAD-mediated clustering activity is required for the in vivo function of Nnk. As in the case of Nnk[Δ1], the homozygous BroN[Δ1] mutant also exhibited a complete lethal phenotype (Fig. 2K), which was rescued only by the ZAD-containing WT construct (Fig. 2K). Overall, these data are consistent with the idea that the correct functionalities of Nnk and BroN rely on the N-terminal ZAD, which has the ability to mediate condensate assembly in developing embryos.
ZAD guides Nnk/BroN into H3K9me3-enriched heterochromatic regions
It has been previously reported that exogenously expressed Nnk and its paralog Odj preferentially accumulate at pericentromeric H3K9me3-enriched heterochromatic regions in S2 cultured cells (7), but their endogenous binding profiles in the context of Drosophila development are unknown. To address this point, we performed ChIP-seq analysis using a well-established anti–FLAG M2 antibody (47) for all the genome-edited strains produced in this study (see Materials and Methods), allowing an unambiguous comparison of the genome-wide distribution of ZAD-ZnFs and other insulator proteins by minimizing potential experimental noise arising from the variability of antibody quality for individual target proteins (20, 21). Initially, to test the validity of our experimental approach, we first performed ChIP-seq analysis against endogenous CTCF and Su(Hw) using our genome-edited strains (fig. S2A). Comparison of our data with publicly available CTCF and Su(Hw) ChIP-seq profiles (43) revealed essentially the same genome-wide distribution patterns of these proteins (fig. S5), suggesting that GFP-3xFLAG tagging does not alter the binding profiles of these key DNA binding proteins. We then performed ChIP-seq analysis of endogenous Nnk, BroN, and Odj in nc14 embryos by following the experimental procedure described in a previous study (43) with minor modifications (see Materials and Methods). We found that the newly identified BroN globally colocalizes with Nnk and Odj (Fig. 3, A to D). These three proteins were found to be highly accumulated at H3K9me3-enriched heterochromatic regions defined by Cut&Tag analysis in nc14 embryos (48), but unexpectedly, the majority of peaks (~70%) are present at H3K9me3-depleted regions (Fig. 3, A and C), suggesting that the H3K9me3 histone modification is not a prerequisite for the recruitment of these ZAD-ZnFs onto the genome. Analysis of ChIP-seq peak distribution revealed that Odj, Nnk, and BroN are mainly localized at transcription start sites (TSSs) and intragenic regions (fig. S4C). To elucidate the role of ZAD-mediated clustering activity in vivo (Fig. 2, D and H), we further analyzed the genome-wide distribution of ZAD-deficient Nnk and BroN mutants by ChIP-seq. Before proceeding with this experiment, we first confirmed that unmodified WT Nnk and BroN expressed as transgenes exhibit essentially the same ChIP-seq profiles as their endogenous counterparts in nc14 embryos (fig. S4, D and E). Peaks of Nnk and BroN ZAD-deficient mutants were found to be selectively lost from H3K9me3-enriched heterochromatic regions, while other peaks were largely maintained (Fig. 3, E and F, and fig. S6, A to C), suggesting that the N-terminal ZAD facilitates association of these proteins with H3K9me3-enriched heterochromatic regions in developing embryos. Given our preceding finding that ZAD-deficient mutants were unable to rescue the lethal phenotype of the full-locus deletion alleles produced in this study (Fig. 2, J and K), we suggest that ZAD-dependent recruitment of Nnk and BroN into H3K9me3-enriched heterochromatic regions is essential for the proper progression of Drosophila development. It is tempting to speculate that the N-terminal ZAD served as a key scaffold driving a coevolutionary arms race between rapid alternations of heterochromatic regions and heterochromatin-interacting ZAD-ZnFs during insect evolution (7).
Fig. 3. ZAD guides Nnk and BroN into H3K9me3-enriched repressive regions.
(A) Heatmap visualization of Odj, Nnk, and BroN ChIP-seq peaks along with H3K9me3 Cut&Tag signal (48) sorted by the strength of Odj (top panels), Nnk (middle panel), or BroN (bottom panels) signal. (B) Integrative Genome Viewer (IGV) genome browser tracks showing Odj, Nnk, and BroN ChIP-seq coverage at the H3K9me3-enriched region. (C) IGV genome browser tracks showing Odj, Nnk, and BroN ChIP-seq coverage at the H3K9me3-depleted region. (D) UpSet plot showing the cobinding profiles of Odj, Nnk, and BroN. (E) Heatmap visualization of WT and ΔZAD Nnk along with H3K9me3 Cut&Tag signal (48) sorted by the strength of WT Nnk. (F) Heatmap visualization of WT and ΔZAD BroN along with H3K9me3 Cut&Tag signal (48) sorted by the strength of WT BroN. See figs. S3 and S4.
Cross-species analysis of evolutionary dynamic Nnk orthologs
To experimentally explore the functionality of the evolutionary dynamic Nnk gene (48) in other Drosophila species, we isolated the protein coding sequences of Nnk orthologs from Drosophila simulans, Drosophila yakuba, Drosophila suzukii, Drosophila ficusphila, and Drosophila persimilis and expressed them as GFP-3xFLAG fusion proteins by placing under the control of the D. melanogaster native Nnk promoter (Fig. 4A). Airyscan super-resolution imaging revealed that three (D. simulans, D. yakuba, and D. suzukii) of the five Nnk orthologs form clear nuclear condensates, as seen for D. melanogaster Nnk in early embryos, whereas only a weak condensation or uniform nuclear localization was seen for the evolutionarily distant D. ficusphila and D. persimilis Nnk (Fig. 4B). We then performed ChIP-seq analysis to examine the genome-wide distributions of the Nnk orthologs. Our analysis revealed that heterochromatin-association activity is conserved among the condensation-prone Nnk orthologs (D. simulans, D. yakuba, and D. suzukii) but is less evident in the D. ficusphila and D. persimilis orthologs (Fig. 4C). Among the five tested orthologs, rescue assay further revealed that the lethal phenotype of the homozygous Nnk[Δ1] mutant can be restored by introducing the condensation-prone Nnk orthologs from three evolutionarily close species (D. simulans, D. yakuba, and D. suzukii) (Fig. 4D). Overall, our data suggest a rapid functional divergence of the Nnk gene, including its heterochromatin-association activity, during Drosophila evolution.
Fig. 4. Comparative analysis of evolutionary dynamic Nnk orthologs.
(A) Phylogeny of 20 Drosophila genomes. OrthoFinder (98) was used for the analysis. The tree was rooted using Drosophila mojavensis, Drosophila virilis, and Drosophila grimshawi as outgroups. (B) Airyscan imaging of GFP-3xFLAG–tagged Nnk orthologs expressed under the control of the native D. melanogaster Nnk promoter. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous strains were used for the analysis. The maximum-intensity–projected images are shown. (C) Heatmap visualization of ChIP-seq peaks of Nnk orthologs along with H3K9me3 Cut&Tag signal (48) sorted by the strength of D. melanogaster Nnk signal. ChIP-seq peaks were divided into two groups according to ZAD dependency of D. melanogaster Nnk (Fig. 3E). (D) Evolutionary distant Nnk orthologs failed to restore the lethal phenotype of homozygous Nnk[Δ1] mutant.
ZAD-independent function of D19B during early embryogenesis
We next focused on the analysis of two neighboring paralogous genes, D19A and D19B (Fig. 5A), as an example of nonclustering ZAD-ZnFs in developing embryos (Fig. 1B). Western blot analysis revealed that D19B is more abundantly expressed than its paralog D19A in early embryos (Fig. 5B). To experimentally test their in vivo functions, we used CRISPR-Cas9–mediated genome editing to remove the entire transcription unit of D19A and D19B individually from the endogenous tandem locus (Fig. 2I) (D19A[Δ1] and D19B[Δ1] denote the resulting full-locus deletion alleles). As a control, a recently reported ZAD-ZnF CG2678/Kipf locus (14) was similarly engineered to remove the entire transcription unit [referred to as kipf[Δ3] because kipf[Δ1] and kipf[Δ2] alleles have already been described by Baumgartner et al. (14)]. Contrary to Nnk[Δ1] and BroN[Δ1], the resulting D19A[Δ1], D19B[Δ1], and kipf[Δ3] mutants were homozygous viable. We then measured the hatching rates of embryos collected from homozygous mutant adults. Consistent with a previous genetic study (14), the null mutation of kipf resulted in a reduction of embryo hatching rate (Fig. 5C), presumably due to partial derepression of transposable elements. Similarly, there was a clear reduction in the hatching rate of embryos obtained from homozygous D19A[Δ1] mutant adults (Fig. 5C). Notably, there was more than 90% reduction in the hatching rate when embryos obtained from homozygous D19B[Δ1] mutant adults were analyzed (Fig. 5C), indicating that D19B plays an essential role in the correct progression of embryogenesis. Supporting this view, live visualization of His2Av-emiRFP670 in embryos obtained from homozygous D19B[Δ1] mutant adults revealed catastrophic nuclear defects, including asynchronous mitosis, disordered nuclei, and M phase arrest (Fig. 5D). These severe negative effects were not seen when control yw virgin females were crossed with homozygous D19B[Δ1] mutant males (Fig. 5E), indicating that the loss of maternal supply of D19B is responsible for the phenotype we observed. Given a recent finding that Kipf helps to protect genome integrity through the suppression of transposable elements by facilitating piRNA production in developing ovaries (14), we examined the possibility that D19A and D19B function in the same pathway. Consistent with the previous study, Kipf-GFP-3xFLAG formed bright foci within nurse cell nuclei that are thought to represent a subset of dual-strand piRNA clusters (fig. S7, A and B). In contrast, both D19A and D19B showed very uniform nuclear distribution in nurse cell and surrounding follicle cell nuclei (fig. S7 A and B). Derepression of Kipf-target transposable elements, Burdock and 3S18 (bel), was seen only in homozygous kipf[Δ3] ovaries in quantitative reverse-transcription (qRT)–PCR analysis (fig. S7C). This result was further supported by RNA sequencing (RNA-seq) analysis of dissected ovaries from homozygous D19B[Δ1] mutant females (fig. S7, D and E). We therefore concluded that D19A and D19B act differently from Piwi-piRNA pathway–related ZAD-ZnF Kipf (14).
Fig. 5. ZAD-independent function of D19B in early embryos.
(A) Two paralogous ZAD-ZnF genes D19B and D19A are tandemly placed at the locus. (B) Western blot analysis of GFP-3xFLAG–tagged endogenous D19A and D19B proteins. Two- to 4-hour embryos from corresponding homozygous strains were used for the analysis. As a control, 2- to 4-hour yw embryos were used. (C) Measurement of the embryo hatching rate. Embryos from corresponding homozygous full-locus deletion mutant adults were used for the analysis. (D) Confocal images of His2Av-emiRFP670 in embryos from WT or homozygous D19B[Δ1] mutant adults. The maximum-intensity–projected images of His2Av-emiRFP670 are shown. (E) Measurement of the embryo hatching rate. Maternal supply of D19B is essential for the correct progression of embryogenesis. (F) Schematic of the full-length D19B protein. (G) AlphaFold3 prediction of the D19B ZAD dimer. Protein structure was visualized using UCSF ChimeraX (104). (H) Western blot analysis of GFP-3xFLAG–tagged endogenous D19B and WT/ΔZAD D19B rescue constructs reintroduced into the endogenous locus via recombinase-mediated cassette exchange (RMCE). Two- to 4-hour embryos from corresponding homozygous mutant adults were used for the analysis. As a control, 2- to 4-hour yw embryos were used. (I) Airyscan imaging of GFP-3xFLAG–tagged WT and ΔZAD D19B expressed from the endogenous locus. Images were taken ~15 min after entry into nc14. Embryos from corresponding homozygous strains were used for the analysis. The maximum-intensity–projected images are shown. (J) Measurement of embryo hatching rate. Embryos from corresponding homozygous RMCE strains were used for the analysis.
Motivated by our finding that Nnk and BroN exert their functions through the N-terminal ZAD domain (Figs. 2 and 3), we then sought to elucidate the role of the D19B ZAD in detail (Fig. 5F). As in the case of Nnk and BroN, AlphaFold3 predicted that the D19B ZAD also forms a dimer through its hydrophobic surface (Fig. 5G). Using a recombinase-mediated cassette exchange (RMCE) approach (49), WT and ΔZAD rescue constructs were individually reintroduced into the endogenous D19B[Δ1] mutant allele. Western blot analysis showed that both rescue constructs are expressed at levels comparable to endogenous D19B (Fig. 5H). Unlike Nnk and BroN (Fig. 2, D and H), both WT and ΔZAD D19B exhibited a very similar pattern of nuclear localization in developing embryos (Fig. 5I). Moreover, measurement of embryo hatching rates revealed that the developmental defects seen in D19B-null embryos were almost fully restored by introducing either WT or ΔZAD D19B constructs (Fig. 5J), indicating that D19B does not rely on the N-terminal ZAD to exert its function during embryogenesis. Our data also showed that the larva-to-pupa and pupa-to-adult transition rates, as well as the embryo hatching rate under heat-stress conditions, are not affected by the loss of the N-terminal ZAD (fig. S8, A and B). Similarly, the homozygous D19A-null mutation did not cause defects in the larva-to-pupa or pupa-to-adult transition (fig. S8A). We then measured the adult survival of the WT and ΔZAD rescue strains and, intriguingly, found that ZAD-deficient rescue females exhibited an extended life span (fig. S8C). Although the molecular basis of this sex-specific longevity effect remains unclear, this phenotype is reminiscent of female-biased life span extension linked to reduced insulin signaling pathway in Drosophila [e.g., (50–52)]. Overall, our data are consistent with the idea that each ZAD-ZnF has differential dependencies on the N-terminal ZAD for exerting its biological functions.
Comparative ChIP-seq analysis of two paralogous D19A and D19B proteins
To determine the distribution patterns of the two paralogous proteins D19A and D19B genome wide, we next performed ChIP-seq analysis using the same standard M2 anti-FLAG antibody for both. As shown in the schematics (Fig. 6A), D19A and its paralog D19B share a very similar domain organization, including the number of zinc finger domains and their amino acid composition (Figs. 5F and 6B), as well as the N-terminal ZAD dimerization scaffold (Fig. 6C). Despite these extensive molecular similarities, we noticed that large fractions of D19A and D19B are recruited to distinct genomic locations without overlapping with H3K9me3, Nnk, and BroN peaks (Fig. 6, D to F, and fig. S6, D and E). Motif analysis suggested that D19B recognizes not only the common “GGGTAT” motif shared with D19A but also a distinct second motif for its chromatin association (fig. S9). These results imply that the two paralogous proteins D19A and D19B have functionally diverged after gene duplication during Drosophila evolution, with D19B playing a more dominant role than D19A in the early embryo (Fig. 5C). To characterize the molecular function of the N-terminal D19B ZAD, we then performed ChIP-seq analysis of ΔZAD D19B expressed in the D19B[Δ1] mutant background. Consistent with our preceding hatching rate analysis (Fig. 5J), D19B binding profiles were found to be mostly unchanged even in the absence of the N-terminal ZAD (Fig. 6, G and H, and fig. S6F). Thus, we concluded that condensation-free D19B does not require the N-terminal ZAD dimerization surface to exert its function during embryogenesis, which sharply contrasts with the strict ZAD dependency of the condensate-forming Nnk and BroN (Figs. 2 and 3).
Fig. 6. Functional divergence of two paralogous proteins D19A and D19B.
(A) Schematic representation of the full-length D19A protein. (B) Similarity score calculated for individual zinc finger domains of D19B and BroN relative to those of D19A. Zinc finger domains of BroN were used as a negative control. (C) AlphaFold3 prediction of the D19A ZAD dimer. Protein structure was visualized using UCSF ChimeraX (104). (D) Heatmap visualization of D19A, D19B, Nnk, and BroN ChIP-seq distribution along with H3K9me3 Cut&Tag signal (48) sorted by the strength of D19A (top panels) or D19B (bottom panels) ChIP-seq peaks. (E) UpSet plot showing the binding profiles of D19A and D19B. (F) IGV genome browser tracks showing D19A and D19B ChIP-seq coverage. (G) Heatmap visualization of WT and ΔZAD D19B ChIP-seq peaks. (H) IGV genome browser tracks showing WT and ΔZAD D19B ChIP-seq coverage.
Insulator-binding activity is a common feature of ZAD-ZnFs
To systematically characterize the genome-wide distribution of ZAD-ZnFs in developing embryos, we performed ChIP-seq analysis for all the GFP-3xFLAG–tagged endogenous ZAD-ZnFs along with the non–ZAD-ZnF insulator proteins produced in this study (Fig. 1 and fig. S2). Integrative analysis of the obtained ChIP-seq profiles and publicly available Micro-C data of nc14 embryos (53) revealed a preferential enrichment of ZAD-ZnFs at topological boundaries (Fig. 7 and fig. S10A). For example, the Abd-B/abd-A Hox locus at the Bithorax complex consists of a series of topologically associating domains (TADs) bordered by well-characterized insulator elements (Fab-8, Fab-7, Fab-6, Mcp, Fab-4, Fab-3, and Fub) [reviewed in Kyrchanova et al. (54)]. It was found that many of the ZAD-ZnFs, together with well-characterized core insulator proteins (e.g., CP190, CTCF, and BEAF-32), cooccupy these insulator elements (Fig. 7A). Among these, enrichment of ZAD-ZnFs was more evident at sharp topological boundaries such as Fub, Fab-6, Fab-8, and the 5′ Abd-B TAD border (Fig. 7A) (55). Similar results were seen for the well-characterized insulator elements responsible for shaping 3D genome topology at the ftz and eve pair-rule loci (Fig. 7B and fig. S10A) (56–60). These data together suggest that ZAD-ZnFs have intrinsic insulator-binding activity in the early Drosophila embryo. It also came to our attention that the tethering element that mediates long-range focal contact at the ftz/Scr locus (Fig. 7B, arrowhead) (53) is also enriched with many ZAD-ZnFs (Fig. 7B, orange region). It is tempting to speculate that not only GAF (61) but also ZAD-ZnFs play a role in organizing long-range focal interactions in developing embryos (see Discussion).
ZAD-ZnFs colocalize with known insulator proteins genome wide
To further address the degree of cooccupancy between ZAD-ZnFs and non–ZAD-ZnF core insulator proteins genome wide, we analyzed the distribution profiles of ZAD-ZnFs using the ChIP-seq peaks of BEAF-32, CTCF, Su(Hw), and CP190 as viewpoints. Consistent with a previous study in nc14 embryos (43), the differential distribution patterns of BEAF-32, CTCF, and Su(Hw) were nicely reproduced in our hands (Fig. 8A), verifying the quality of our ChIP-seq dataset. Having confirmed this, we next visualized the distribution of all tested ZAD-ZnFs at the four viewpoints. It turned out that ZAD-ZnFs globally colocalize with BEAF-32, CP190, and CTCF in developing Drosophila embryos (Fig. 8, B and C, and fig. S10B). Colocalization of ZAD-ZnFs and Su(Hw) was also seen genome wide, but the strong Su(Hw) peaks tended to be depleted of binding of ZAD-ZnFs (fig. S10C), as well as other insulator proteins (Fig. 8A). The two paralogous ZAD-ZnFs D19A and D19B did not show clear enrichment at the peaks of any of the four viewpoint insulator proteins genome wide (Fig. 8, B and C, and fig. S10, B and C), implicating that they are recruited to only a specific subset of topological boundaries, such as the 5′ Abd-B TAD border and the SF2 insulator (Fig. 7), to exert their locus-specific functions. Motif analysis suggested that ZAD-ZnFs are generally recruited to their binding sites through recognition of unique sequence motifs that are distinct from those of non–ZAD-ZnF DNA binding proteins (fig. S9). Clear sequence motifs were not reliably recovered for Nnk, Odj, and BroN, suggesting that these factors are indirectly recruited to chromatin, possibly through association with other DNA binding proteins. Overall, extensive degree of genome-wide colocalization with core insulator proteins suggests that the diverse molecular functions of ZAD-ZnFs have evolutionarily arisen from their ancestral role as insulator-binding proteins. Supporting this view, Nnk orthologs from five different Drosophila species similarly showed genome-wide colocalization profiles with other insulator proteins in early embryos (Fig. 9, A and B, and fig. S11). As a representative case, the boundary-association activities of Nnk orthologs can be seen at the Abd-B/abd-A Hox locus (Fig. 9C).
Fig. 9. Insulator-binding activity of Nnk is conserved across Drosophila species.
(A) Heatmap visualization of ChIP-seq distribution of insulator proteins and Nnk orthologs sorted by the strength of BEAF-32 ChIP-seq peaks. (B) Heatmap visualization of ChIP-seq distribution of insulator proteins and Nnk orthologs sorted by the strength of CP190 ChIP-seq peaks. (C) Micro-C map along with ChIP-seq profiles of Nnk orthologs at the Abd-B/abd-A Hox locus. Publicly available Micro-C data in nc14 embryos [Batut et al. (53); SRR14136009] were used for the analysis.
Next, to address the mechanism underlying genome-wide cooccupancy of ZAD-ZnFs with other insulator proteins, we analyzed the ChIP-seq profiles of the ZAD-deficient Nnk and BroN mutants (Fig. 2). It turned out that both Nnk and BroN retain colocalization with BEAF-32 and CP190 genome wide, even in the absence of the N-terminal ZAD and the accompanying condensation activity (Fig. 2, D and H, and fig. S12A), suggesting that the C-terminal zinc finger arrays but not the N-terminal ZAD are responsible for their association with insulator elements. To further explore the functional link between boundary strength and the ChIP-seq binding profiles of ZAD-ZnFs, we then calculated the number of proteins containing peaks at individual topological boundaries (fig. S12B). This analysis revealed that there is a positive correlation between boundary strength and the number of proteins associating with the corresponding DNA regions (Fig. 8D), implicating that the cooccupancy of ZAD-ZnFs and other non–ZAD-ZnF insulator proteins fosters the establishment and/or maintenance of topological boundaries during early embryogenesis (see Discussion).
Loss of ZAD-ZnFs leads to locus-specific misregulation of TADs
Last, to directly test the role of ZAD-ZnFs in TAD formation during early embryogenesis, we performed high-resolution Micro-C analysis using four homozygous knock-out strains produced in this study (kipf[Δ3], D19A[Δ1], CG17361[Δ1], and CG4282[Δ1]) and the control yw strain. Our data revealed that the loss of each ZAD-ZnF alters genome topology at a subset of chromosomal locations in a locus-specific manner. Notably, we observed emergence of small sub-TADs upon the loss of ZAD-ZnFs within TADs where the corresponding ZAD-ZnFs localize at their boundaries (Fig. 10A), suggesting that they help ensure the specificity of genome organization by preventing ectopic boundary formation. We also identified cases in which TAD boundaries were simply lost in D19A[Δ1] and CG17361[Δ1] mutants (Fig. 10B). Overall, these Micro-C results are consistent with the idea that ZAD-ZnFs play an active role in controlling genome topology in a context dependent manner. On the basis of these results, we named CG17361 and CG4282 Boundary element-associated factor of 25 kD (BEAF-25) and Boundary element-associated factor of 75 kD (BEAF-75), respectively. The observed changes occurred only at distinct subsets of genomic locations for each mutant, indicating that individual ZAD-ZnFs exert locus-specific functions in shaping genome topology. Our Micro-C data showed that the functionality of insulator elements at the Abd-B/abd-A Hox locus was essentially unaffected in all tested mutants (fig. S13A), suggesting that ZAD-ZnFs act redundantly at these highly cooccupied boundaries to confer robust topological insulation. This idea was further supported by analyses of the ftz and eve pair-rule loci (fig. S13, B and C).
Fig. 10. Misregulation of genome topology in ZAD-ZnF mutant embryos.
(A) Micro-C maps along with ChIP-seq profiles of corresponding ZAD-ZnFs. Representative genomic loci where new topological boundaries were gained upon the loss of ZAD-ZnFs are shown. (B) Micro-C maps along with ChIP-seq profiles of corresponding ZAD-ZnFs. Representative genomic loci where existing topological boundaries were lost upon the loss of ZAD-ZnFs are shown.
DISCUSSION
ZAD-ZnFs as insulator-binding proteins
A notable finding of this study is that not only the previously characterized Zw5 and M1BP (16, 17) but also other ZAD-ZnFs are highly accumulated at the topological boundaries of key developmental loci in developing Drosophila embryos (Fig. 7 and fig. S10A). Analysis of ChIP-seq datasets further revealed extensive cooccupancy of ZAD-ZnFs and other non–ZAD-ZnF insulator proteins genome wide (Fig. 8 and fig. S10, B and C), giving rise to the possibility that clustered binding of ZAD-ZnFs helps to organize 3D genome topology. Given recent findings that most topological boundaries are maintained even after the depletion of well-characterized core insulator proteins such as CTCF, BEAF-32, and CP190 (43, 62), it is tempting to speculate that insulator-associating ZAD-ZnFs help to ensure the robustness of topological boundary formation in concert with other insulator proteins during Drosophila development. This idea was supported by our Micro-C analysis of highly occupied topological boundaries at the Abd-B/abd-A Hox locus (fig. S13A) as well as the eve and ftz pair-rule loci (fig. S13, B and C). In mammals, ubiquitously expressed DNA binding proteins YY1 and Ldb1 are suggested to mediate long-range interactions through homotypic protein-protein interactions mediated by their dimerization domains (63, 64). Similarly in Drosophila, it was recently shown that the dimerization activity of the N-terminal POZ/BTB domain of GAF can mediate long-range interactions over large genomic distances (61). In analogy to the actions of YY1, Ldb1, and GAF, it can be possible that ZAD-ZnFs use their dimerization activity to help establish topological boundaries through homotypic protein-protein interactions. This mechanism might also be at work during the establishment of long-range focal interactions mediated by tethering elements (Fig. 7B, orange region) (53). We noticed that the structure of the CP190 POZ/BTB dimer predicted by AlphaFold3 appears similar to that of ZAD dimers, whereas POZ/BTB dimers from other architectural proteins (GAF, Mod(mdg4), and Lola) exhibit distinct structural configurations (fig. S14), suggesting that the modes of POZ/BTB dimerization are more diverse than those of ZAD dimers.
Our data suggest that the N-terminal ZAD is dispensable for the recruitment of ZAD-ZnFs to insulator elements (fig. S12A), but this does not necessarily rule out the possibility that the dimerization activity of the N-terminal ZAD helps organize higher-order genome topology after loading onto insulator elements. It is important to note that ZAD-ZnF functions, including their insulator-binding activity and the role of the N-terminal ZAD domain, may vary across developmental stages and tissue types. Future studies will be required to elucidate the molecular functions of individual insulator-associated ZAD-ZnFs in shaping 3D genome topology, and to examine additional developmental contexts (e.g., larval, imaginal disc, and adult stages) to determine whether ZAD-ZnFs act as universal genome organizers or acquire stage- or tissue-specific roles.
Heterochromatin association of ZAD-ZnFs
It has been previously reported that exogenously expressed Nnk and Odj are highly accumulated at H3K9me3-enriched heterochromatic regions in S2 cultured cells (7). Consistent with this observation, our ChIP-seq analysis of endogenous Nnk and Odj revealed their preferential association with H3K9me3-enriched heterochromatic regions in developing embryos (Fig. 3). We further demonstrated that a newly identified CG30020/BroN cooperates with Nnk and Odj to associate with their target sites across the genome (Fig. 3). While the detailed molecular functions of these H3K9me3-associating ZAD-ZnFs remain to be elucidated at this point, our data indicate that ZAD-dependent condensation activity is required for their correct in vivo functionalities (Fig. 2). It has been previously suggested that phase separation of HP1α mediates dynamic compaction of heterochromatic regions to silence gene expression in both mammals and Drosophila (65, 66). Similarly, Drosophila GAF subnuclear foci have been reported to contribute to the silencing of heterochromatic AAGAG satellite repeats (28). Given these findings, it is conceivable that ZAD-dependent cocondensation of Nnk, Odj, and BroN helps to ensure gene silencing at heterochromatic regions in developing embryos. This mechanism might involve the physical isolation of heterochromatic regions into repressive ZAD-ZnF condensates, thereby excluding active transcriptional machinery from silenced loci. A similar exclusion mechanism has been suggested to operate in PRC1-mediated gene silencing in mammals (67). Strict ZAD dependency of H3K9me3-association activities of Nnk and BroN (Fig. 3) gives rise to the possibility that the N-terminal ZAD served as a key scaffold driving a coevolutionary arms race between rapid alternations of heterochromatic regions and heterochromatin-interacting ZAD-ZnFs during insect evolution (7).
Dynamic modulation of ZAD-ZnF functions
While all the ZAD-ZnFs share a common structural arrangement (i.e., an N-terminal ZAD followed by an array of zinc finger domains at the C terminus), only a subset of ZAD-ZnFs exhibits condensation activity in living embryos (Fig. 1B). This implicates that the dimerization activity of the N-terminal ZAD is not sufficient to mediate the local accumulation of ZAD-ZnFs within a nucleus. Kipf forms very bright foci that are thought to represent a subset of dual-strand piRNA clusters in nurse cell nuclei (fig. S7, A and B), but its nuclear distribution is rather uniform in early embryos (Fig. 1B), suggesting that the protein localization of ZAD-ZnFs can be dynamically modulated during the progression of Drosophila development. Given a recent finding that the interaction between Kipf and Rhino, a germline-specific variant of HP1, is responsible for the formation of bright Kipf foci in nurse cell nuclei (14), we speculate that the presence or absence of binding partners of individual ZAD-ZnFs can flexibly tune their genome-wide distribution and in vivo functionalities. It is conceivable that the gain or loss of binding partners has also contributed to the molecular diversification and functional specification of evolutionarily dynamic ZAD-ZnFs in other insect species. As a future study, it would be intriguing to explore the roles of individual ZAD-ZnFs in developmental contexts other than early embryos. We believe that our current study represents a critical starting point toward a comprehensive understanding of the diverse, multifaceted functionalities of ZAD-ZnFs in Drosophila and other insect species.
MATERIALS AND METHODS
Site-specific transgenesis by phiC31 system
GFP-3xFLAG–fusion Nnk and BroN rescue constructs were integrated into a unique attP landing site on the second chromosome using VK00002 strain (68). SV40NLS-GFP-3xFLAG and His2Av-emiRFP670 transgenes were integrated into a unique attP landing site on the third chromosome using VK00033 strain and the second chromosome using VK00028, respectively (68). Microinjection was performed as previously described (69). Briefly, 0- to 1-hour embryos were collected and dechorionated with bleach. Aligned embryos were dried with silica gel for ~8 to 9 min and covered with FL-100 1000CS silicone oil (Shin-Etsu Silicone). Subsequently, microinjection was performed using FemtoJet (Eppendorf) and DM IL LED inverted microscope (Leica) equipped with M-152 Micromanipulator (Narishige). Injection mixture typically contains plasmid DNA (~900 ng/μl), 5 mM KCl, and 0.1 mM phosphate buffer (pH 6.8). mini-white marker was used for screening.
CRISPR-Cas9–mediated genome editing
Genome-edited fly lines were obtained by CRISPR-Cas9–mediated homology directed repair. For GFP-3xFLAG/mCherry-3xFLAG insertion, guide RNA (gRNA) was designed to target protospacer adjacent motif (PAM)–containing sequence adjacent to the stop codon. pCFD3 gRNA expression plasmid, pBS-Hsp70-Cas9 plasmid (Addgene, #46294), and pBS-GFP-3xFLAG/mCherry-3xFLAG-dsRed-loxP donor plasmid containing 5′ and 3′ homology arms were coinjected using the standard yw laboratory strain. Injection mixture typically contains pCFD3 gRNA expression plasmid (~300 ng/μl), pBS-Hsp70-Cas9 plasmid (~600 ng/μl), and donor plasmid (~300 ng/μl). Microinjection was performed as described above. 3xP3-dsRed was used for screening. Resulting genome-edited flies were then crossed with y1 w67c23 P{Crey}1b; D*/TM3, Sb1 (BDSC, #851) or y1 w67c23 P{Crey}1b; snaSco/CyO (BDSC, #766) to remove 3xP3-dsRed marker from the endogenous locus. For full-locus deletion, two gRNAs were designed to target PAM-containing sequences upstream of the TSS and downstream of the stop codon. A pair of pCFD3 gRNA expression plasmids, pBS-Hsp70-Cas9 plasmid, and pBS-3xFLAG-attP-dsRed-attP donor plasmid containing 5′and 3′ homology arms were coinjected using the standard yw laboratory strain. Injection mixture typically contains pCFD3 gRNA expression plasmids (~200 ng/μl), pBS-Hsp70-Cas9 plasmid (~400 ng/μl), and donor plasmid (~300 ng/μl). 3xP3-dsRed was used for screening. Deletion and insertion were confirmed by the PCR analysis of purified genomic DNA. DNA oligos used to build gRNA expression plasmids are listed in table S1.
Recombinase-mediated cassette exchange
D19B rescue constructs and p3xP3-EGFP.vas-int.NLS plasmid (Addgene, #60948) were coinjected into embryos obtained from the D19B[Δ1]/tm3, Kr-GFP[w+] strain. Injection mixture typically contains rescue plasmid (~500 ng/μl) and phiC31 expression plasmid (~800 ng/μl). Loss of 3xP3-dsRed marker was used for screening.
Western blotting
Twenty of 2- to 4-hour embryos were hand collected, and 50 μl of 2× SDS sample buffer was added. Subsequently, embryos were crushed with a pestle and boiled at 94°C for ~3 min. Insoluble pellet was removed by centrifugation; and lysates were separated by SDS–polyacrylamide gel electrophoresis, transferred to Immobilon-P PVDF membranes (Merck, IPVH0001), and immunoblotted with antibodies. The primary antibodies used were as follows: mouse monoclonal anti–FLAG M2 antibody (Sigma-Aldrich, F3165; 1:10,000) and mouse monoclonal anti–α-tubulin antibody (Sigma-Aldrich, T5168; 1:40,000). The membranes were then incubated with secondary antibodies conjugated with horseradish peroxidase. Chemiluminescence was detected by a FUSION Solo.7S.Edge (Vilber-Lourmat).
Confocal imaging analysis
In figs. S1C and S7 (A and B), virgin female flies were collected and aged for 2 days before ovary dissection. Ovaries were dissected into 1× phosphate-buffered saline (PBS) buffer and mounted between a polyethylene membrane (UBE Film) and a coverslip (18 mm by 18 mm) and embedded in FL-100-450CS (Shin-Etsu Silicone). Dissected ovaries were imaged by an FV4000 confocal microscope equipped with SilVIR detector (Evident). UPlanXApo 40x/1.4 numerical aperture (NA) oil immersion objective was used. Fluorescence of mCherry and GFP were excited using 561- and 488-nm lasers, respectively. Images were acquired with following settings: 2048 by 2048 pixels and 16-bit depth. In Fig. 5D, fluorescence of emiRFP670 was excited using a 640-nm laser. Embryos were imaged by an FV4000 confocal microscope equipped with SilVIR detector (Evident). Images were acquired with following settings: 2048 by 2048 pixels and 16-bit depth. In fig. S3A, data acquisition was performed using a Zeiss LSM 900 confocal microscope. Plan-Apochromat 20x/0.8 NA objective was used. Images were acquired with following settings: 512 by 512 pixels with a pixel size of 0.78 μm, 16-bit depth, and 21 z-slices separated by 0.5 μm. In fig. S3B, data acquisition was performed using a Zeiss LSM 900 confocal microscope. Plan-Apochromat 40x/1.4 NA oil immersion objective was used. Images were acquired with following settings: 512 by 512 pixels with a pixel size of 0.078 μm, 16-bit depth, and 21 z-slices separated by 0.5 μm. Fluorescence of 4′,6-diamidino-2-phenylindole (DAPI), Alexa Fluor 488, Cy3, and Cy5 was excited using 405-, 488-, 561-, and 640-nm lasers, respectively.
Super-resolution imaging with Airyscan2 system
In Figs. 1, 2 (D and H), 4B, and 5I and fig. S1 (F and G), data acquisition was performed using Airyscan2 imaging system. Embryos at nc14 were imaged using a Zeiss LSM 900 equipped with Airyscan2 detector. During imaging, temperature was kept in between 22.0° and 23.5°C. Plan-Apochromat 63x/1.4 NA oil immersion objective was used. Images were acquired with following settings: 1180 by 708 pixels (Fig. 1 and fig. S1G) or 1180 by 732 pixels (Figs. 2, D and H, 4B, and 5I) with a pixel size of 0.043 μm, 16-bit depth, and 67 z-slices separated by 0.15 μm. In fig. S1F, GAF-GFP and GAF-mCherry images were acquired with following settings to achieve optimal Airyscan imaging setup: 1180 by 708 pixels with a pixel size of 0.043 μm, 16-bit depth, and 67 z-slices separated by 0.15 μm for GAF-GFP; 1024 by 612 pixels with a pixel size of 0.049 μm, 16-bit depth, and 56 z-slices separated by 0.18 μm for GAF-mCherry. Fluorescence of mCherry and GFP were excited using 561- and 488-nm lasers, respectively. To enable high-speed scanning of embryo samples, Airyscan Multiplex SR-4Y acquisition mode was used. Raw Z-stack images were subjected to 3D Airyscan processing using ZEN 3.1 software (Zeiss) with the standard mode.
Single-molecule inexpensive FISH and immunofluorescence staining
Single-molecule inexpensive FISH (smiFISH) and immunostaining were performed with minor modifications to previously described methods (70–72). Embryos aged 2 to 4 hours were dechorionated and fixed in a fix solution (4 ml of 1× PBS, 1 ml of 37% formaldehyde, and 5 ml of heptane) for 45 min at room temperature. The aqueous phase was removed, and ~10 ml of 100% methanol was added. The mixture was shaken for 1 min to devitellinize embryos. Devitellinized embryos were washed with methanol and stored in methanol at –30°C until use. smiFISH probes targeting endogenous twi and zen genes were designed using the Biosearch Technologies Stellaris RNA FISH probe designer tool (https://oligos.biosearchtech.com/products/rna-fish/). The following sequence was added to the 5′ end of each 20-nucleotide probe: 5′-CCT CCT AAG TTT CGA GCT GGA CTC AGT G-3′, which is the reverse complement of the X FLAP sequence described by Tsanov et al. (73). The X FLAP sequence (5′-CAC TGA GTC CAG CTC GAA ACT TAG GAG G-3′), labeled with Cy3 or Cy5 at both the 5′ and 3′ ends, was synthesized by Eurofins Genomics. Probes and the X FLAP were annealed and stored at –30°C until use. A final concentration of 80 nM of annealed Cy3-labeled twi and Cy5-labeled zen probes was used for hybridization with embryos in the dark at 37°C for ~14 hours in hybridization buffer (10% dextran sodium sulfate 5000, 2× SSC, and 10% deionized formamide). After hybridization, embryos were washed with wash buffer (2× SSC and 10% deionized formamide) at 37°C. For immunofluorescence staining, embryos were incubated with primary antibody in blocking solution [1.5× western blocking reagent (Roche) in 1× PBS-T] for ~17 hours at 4°C. Anti-dorsal (Developmental Studies Hybridoma Bank, 7A4; 1:50) and anti-hemagglutinin (Roche, 11867423001; 1:500) were used as primary antibodies. Embryos were then washed with PBS-T, blocked for 30 min with blocking solution, and incubated with secondary antibody diluted 1:500 in blocking solution for 2 hours at room temperature. Alexa Fluor 488–conjugated anti-mouse immunoglobulin G (IgG; Thermo Fisher Scientific, A21202) and anti-rat IgG (Thermo Fisher Scientific, A11006) were used as secondary antibodies. After washing and DAPI staining, embryos were mounted in ProLong Gold Antifade Mountant (Thermo Fisher Scientific). The sequences of smiFISH probes are listed in table S2.
Scoring of embryo hatching rates
To quantify embryo viability, more than 70 virgin females and males were individually collected for 3 days, mated with each other, and aged for 2 additional days. Following two ~1-hour prelays, embryos were collected for 3 hours and incubated on apple juice plates for ~48 hours at 25°C. Subsequently, the number of unhatched embryos was counted to calculate the embryo hatching rate, expressed as the percentage of hatched embryos from total. For embryos laid by D19B[Δ1] females, due to the difficulty to visually distinguish unhatched embryos from hatched embryos, the number of hatched embryos was instead measured by counting the number of wandering larvae. To assess embryo hatching rates under heat-stress conditions, the same procedures were followed, except that embryos were incubated on apple juice plates at 30°C instead of 25°C.
Viability tests for larvae and pupae
More than 70 virgin females and males were individually collected for 3 days, mated with each other, and aged for 2 additional days. Following two ∼1-hour prelays, embryos were collected for 3 hours and incubated for ~24 hours at 25°C. Then, 10 sets of 25 larvae were picked and transferred into the food in individual vials, which had been supplemented with 3 ml of Milli-Q water and mixed to increase its moisture content. Pupae that developed before day 7 after egg laying were marked on the vial walls to score the larva-to-pupa transition rates. The eclosion status of these pupae was examined on day 11 after egg laying to calculate the pupa-to-adult transition rates. Those completely emerged from the pupal case were scored as eclosed.
Life span assay
Life span assay was performed basically as described in Linford et al. (74) with some modifications. Newly eclosed flies were collected over a 30-hour period and then allowed to mate for 2 days. Subsequently, males and females were sorted under mild CO2 anesthesia. We placed 30 flies of the same sex into individual vials. For each strain, 10 vial replicates were prepared. Every 2 to 3 days, flies were transferred to new vials without anesthesia, and the number of dead or censored flies was recorded. At each transfer, vial positions were randomized to minimize location-based bias. Flies were incubated in a room under controlled conditions of ~24° to 25°C and ~40 to 60% humidity. A custom Python code was used for data analysis. Survival curves were estimated using the Kaplan-Meier method. For merged survivorship data, the numbers of alive, dead, and censored flies from all replicates were summed at each time point, and the Kaplan-Meier method was applied. SciPy (75) was used for the statistical comparison.
qRT-PCR
Virgin female flies were collected for 3 days and subsequently fed yeast paste for 3 days before ovary dissection. Ovaries were dissected into 1× PBS. Total RNA was purified from five pairs of ovaries using TRIzol (Thermo Fisher Scientific). One microgram of total RNAs was reverse transcribed by PrimeScript RT Reagent Kit with gDNA Eraser (Takara). qRT-PCR was performed using TB Green Premix Ex Taq II (Takara) and the LightCycler 480 System II (Roche). Serial dilutions of cDNA were used for qRT-PCR calibrations. Melting-curve analysis was performed to confirm the amplification of a single product for each target. Following pairs of primers were used for the analysis: rp49: (5′-CCG CTT CAA GGG ACA GTA TCT G-3′) and (5′-ATC TCG CCG CAG TAA ACG C-3′); Burdock: (5′-GGA CAA ACT CCA TTT CCG ACT C-3′) and (5′-TCC CTG AGC CTG ACT TGT GT-3′); and 3S18 (bel): (5′-CAA TCG CTT CTT CTC GCT GG-3´) and (5´-CAG GTG TCT GGG GAC TCT TC-3′).
RNA-seq library preparation
Four hundred nanograms of total RNA isolated from ovaries was used for ribosomal RNA (rRNA) depletion using biotinylated antisense oligos against 18S and 28S rRNAs as previously described (76). The rRNA-depleted samples were purified with RNAClean XP beads (Beckman Coulter). Libraries were prepared using the TruSeq Stranded mRNA Library Prep Kit (Illumina, 20020594) according to the manufacturer’s instructions. The multiplexed libraries were sequenced on a NovaSeq X Plus sequencer (Illumina) using paired-end 150 base-pair (bp) reads.
Computational analysis of RNA-seq data
Adapter sequences were trimmed with Fastp (77). The trimmed reads were then aligned to the D. melanogaster reference genome (BDGP6.32) using STAR (78) with the following parameters: --outMultimapperOrder Random --outSAMtype SAM --outReadsUnmapped Fastx --outFilterMultimapNmax 1000 --winAnchorMultimapNmax 1000 --limitOutSAMoneReadBytes 2000000. SAM files were converted to bam files by samtools (79). The expression levels of genes and transposable elements were quantified with TEcount (80) with gene annotation from Ensembl (BDGP6.32) and transposable element annotations downloaded from the TEcount authors’ website (www.mghlab.org/software/tetranscripts). Differential expression analysis was performed with DESeq2 (81) for genes and transposable elements with >1 read counts. To quantify the expression of transposon transcripts in kipf-depleted versus control ovaries, RNA-seq data from previous publication (14) were processed following the same protocol.
Embryo collection and fixation for ChIP
Aged adult flies were discarded from vials, and newly eclosed flies were collected twice: in the evening of the following day and in the morning 2 days later. Flies were transferred to population cages with apple juice agar plates supplemented with yeast paste and allowed to mate for 1 day at 25°C. On the following morning, plates were exchanged twice, and embryos were collected over a 30-min window and aged for 2 hours and 10 min at 25°C. This resulted in embryo collections corresponding to 2 hours and 10 min to 2 hours and 40 min after egg laying, which corresponds to nc14. Collected embryos were dechorionated by immersion in bleach (Kao) for 2 min, followed by thorough rinsing with water. Embryos were fixed by rotating in a mixture of heptane and PBS-T containing 1.8% formaldehyde (final concentration in the aqueous phase) for 10 min. Fixation was quenched by adding glycine to a final concentration of 250 mM and rotating for 5 min. Embryos were pelleted by centrifugation at 1000g for 1 min at 25°C. Fixed embryos were flash frozen in liquid nitrogen and stored at –80°C until further use.
Preparation of antibody-conjugated beads for ChIP
Dynabeads Protein G magnetic beads (Invitrogen) were diluted by adding 10 μl of beads to 100 μl of blocking solution [PBS containing bovine serum albumin (BSA; 5 mg/ml)]. After a brief spin down, the beads were collected using a magnetic stand and washed twice with blocking solution. The beads were then resuspended in 500 μl of blocking solution, followed by the addition of 2 μl of either anti–FLAG M2 monoclonal antibody (Sigma-Aldrich, F1804) or anti–RNA Polymerase II monoclonal antibody CTD4H8 (Merck Millipore, 05-623). The mixture was rotated at 4°C for 6 hours. After incubation, the beads were washed three times with blocking solution and lastly resuspended in 12 μl of blocking solution.
ChIP
Approximately 50 μl of nc14 embryos were washed once with PBS-T and centrifuged at 1000g for 1 min at 25°C. The pellet was resuspended in lysis buffer [15 mM HEPES (pH 7.5), 15 mM NaCl, 60 mM KCl, 4 mM MgCl2, 0.5% Triton X-100, 0.5 mM dithiothreitol (DTT), and 1× EDTA-free protease inhibitor (Nacalai Tesque)] and homogenized using a Dounce homogenizer. The lysate was centrifuged at 3000g for 3 min at 4°C, and the pellet was further diluted with lysis buffer and pipetted 30 times to disperse nuclei. The nuclear suspension was sequentially washed once with lysis buffer and once with wash buffer [15 mM HEPES (pH 7.5), 200 mM NaCl, 1 mM EDTA, 0.5 mM EGTA, and 1× EDTA-free protease inhibitor], with 30 pipetting strokes at each step. Nuclei were then resuspended in 300 μl of sonication buffer [15 mM HEPES (pH 7.5), 140 mM NaCl, 1 mM EDTA, 0.5 mM EGTA, 0.1% sodium deoxycholate, 0.5% sodium N-lauroylsarcosinate, and 1× EDTA-free protease inhibitor] and subjected to sonication using a Bioruptor Pico (Diagenode) set to 30-s on/30-s off cycles, intensity: high, for 7 cycles. Four of the six available holders were used, with standard 1.5-ml tubes. The chromatin fraction was obtained by centrifugation at 20,000g for 15 min at 4°C. The supernatant was supplemented with 1% Triton X-100 in a total volume of 600 μl. One-twelfth of the sample (50 μl) was reserved as input and stored at –30°C. Immunoprecipitation was performed overnight at 4°C by adding the entire volume of antibody-coupled beads (prepared as described above) diluted in 12 μl of blocking solution. Beads were then washed four times with RIPA buffer [50 mM HEPES (pH 7.5), 500 mM LiCl, 1 mM EDTA, 1% NP-40, and 0.7% sodium deoxycholate] and once with TE buffer containing 50 mM NaCl. After centrifugation at 960g for 3 min at 4°C, the supernatant was completely removed. Beads were resuspended in 200 μl of elution buffer [50 mM Tris-HCl (pH 8.0), 10 mM EDTA, and 1% SDS] and incubated at 65°C for 30 min. The eluate was collected using a magnetic stand and transferred to a new tube. For the input, 150 μl of elution buffer was added to the 50-μl reserved fraction. De–cross-linking was performed by incubating both ChIP and input samples at 65°C for 6 hours. Afterward, 200 μl of TE buffer was added. Samples were treated with RNase A (final concentration of 10 μg/ml) at 37°C for 30 min and with proteinase K (final concentration of 0.5 mg/ml) at 55°C for 2 hours to remove RNA and proteins. DNA was purified by phenol-chloroform extraction, followed by ethanol precipitation. The final DNA pellet was resuspended in 50 μl of nuclease-free water.
ChIP-seq library preparation
ChIP-seq libraries were prepared using the NEBNext Ultra II DNA Library Prep Kit for Illumina (New England Biolabs), starting with 15 μl of ChIP DNA and 5 μl of input DNA. During adaptor ligation, the adaptor reagent was diluted 25-fold for ChIP DNA and 10-fold for input DNA. Adapter-ligated DNA was PCR amplified using Illumina index primers (NEBNext Multiplex Oligos for Illumina, New England Biolabs), with 12 cycles for ChIP DNA and 6 cycles for input DNA. PCR products ranging from 100 to 1000 bp were size selected and purified using AMPure XP beads (Beckman Coulter). Size-selected libraries were validated by TapeStation D5000HS kit (Agilent Technologies). The resulting libraries were sent to BGI Japan for circularization and sequenced on a DNBSEQ-G400 platform (MGI Tech) using paired-end 100-bp reads.
Computational analysis of ChIP-seq data
Before mapping, low-quality reads were removed, and adapter sequences were trimmed using Fastp (77) with the options -q 30 -n 5 -t 1 -l 20 -w 16. The filtered reads were mapped to the D. melanogaster reference genome (dm6) using Bowtie2 (82). Only uniquely mapped reads were retained for analysis using the samtools (79) view command with the -q 42 option. ChIP-seq experiments were performed with at least two biological replicates. Because the replicates showed high correlation, they were merged, and the combined libraries were used for downstream analyses. Peak calling was performed by MACS2 (83) using the input sample as control with -f BAMPE q = 0.01 option, and called peaks were considered as regions of protein binding. Because ChIP-seq of insulator proteins such as CTCF are known to tend to give rise to false-positive peaks, often referred to as “phantom peaks” (84), ChIP-seq using anti–FLAG M2 monoclonal antibody (Sigma-Aldrich, F1804) for the control yw strain (the parental strain for all GFP-3xFLAG tag knock-in strains) was performed to identify such peaks. In total, 269 genomic regions were identified as phantom peaks (summarized in table S3). Any peak in the experimental samples that overlapped by even 1 bp with these phantom peaks was excluded from further analysis. Overlaps were determined using the bedtools intersect command (85). For heatmaps and profile plots, ChIP scores were calculated using deepTools (86). BigWig files were generated from BAM files using bamCoverage or bamCompare, and matrix data were computed using computeMatrix, followed by plotting. In bamCompare, input samples were used as control for signal subtraction. All genome browser views and heatmap visualizations used ChIP tracks with input-subtracted signals. Correlation of mapping files was confirmed using the multiBamSummary program in deepTools. For upset plots, the narrowPeak files obtained from MACS2 callpeak were processed using the GRanges function in the GenomicRanges package (87) and visualized with the UpSetR package (88). Quantitative comparison of ChIP peaks was performed using MAnorm (89). For visualization of ChIP-seq tracks, IGV (90) was used. For visualization of ChIP-seq tracks alongside Micro-C data, CoolBox (91) was used. For Fig. 8D, genomic positions and strengths of topological boundaries were obtained from the published Micro-C dataset GSE171396 (the file GSE171396_Batut_MicroC_yw_CC14_Boundaries_min1.25.bed.gz) (53). In this dataset, boundaries were identified using genome-wide 400-bp resolution insulation scores calculated with a 10-kb window. The boundary strength corresponds to the boundary score defined in the original study, which quantifies the depth of the local minimum in the insulation score. ChIP-seq quality metrics, including alignment rates, were summarized in table S4. As a quality metrics for ChIP-seq, library complexity (the proportion of unique reads in the entire library, indicating low levels of PCR bias) was calculated using BAM files by the DROMPA3 (92) software. To evaluate the signal-to-noise ratio of the ChIP signal, the strand-shift profile and synthetic Jensen-Shannon divergence were calculated using BAM files by SSP (93) and plotFingerprint program in deepTools, respectively. DNA motifs enriched at each ChIP-seq binding site were identified using the MEME Suite (94). For motif discovery, the top 100 genomic regions from the narrowPeak file generated by MACS2 callpeak were used. Among the detected motifs, only those that showed partial overlap with ChIP-seq peaks upon manual inspection using IGV were considered as bona fide motifs.
Analysis of zinc finger domain similarity
Domain compositions were defined following the annotations provided in the UniProt database for each protein. To assess the similarity of C2H2 zinc finger domains among D19A, D19B, and BroN, C2H2 motifs were further identified using the regular expression “(.{2})C(.{2}|.{4})C(.{12})H(.{3,5})H” [reviewed in Pabo et al. (95)] with custom Python code. Each zinc finger domain in D19B or BroN was then aligned to the corresponding zinc finger domain in D19A using Biopython (96), ensuring that core Cys and His residues were always aligned with each other. The alignment results are presented in table S5. Similarity scores were defined as the percentage of perfectly aligned residues in the zinc finger domain of interest, relative to the total number of residues in the corresponding D19A zinc finger domain.
Identifying Nnk orthologs
For each species, we performed a blastp search (https://blast.ncbi.nlm.nih.gov/Blast.cgi) using the previously reported partial amino acid sequences of the Nnk orthologs (7) as a query. The top hit from the corresponding species, based on the lowest E value, was defined as its ortholog. Full-length amino acid sequences of the orthologs were then converted into a CDS, with codon usage optimized for D. melanogaster. DNA fragments containing the CDSs of Nnk orthologs were chemically synthesized (Thermo Fisher Scientific). All amino acid sequences of the identified orthologs are listed in table S6.
Construction of species phylogenetic tree
FASTA files containing the amino acid sequences of all gene products were downloaded from the National Center for Biotechnology Information (NCBI) website for each species. The genome assemblies and corresponding gene annotations used for the analysis are as follows: dm6 (FlyBase Release 6.54) for D. melanogaster, ASM438219v2 (NCBI Release 101) for Drosophila sechellia, Prin_Dsim_3.1 (NCBI Release 103) for D. simulans, DereRS2 (NCBI Release 101) for Drosophilia erecta, Prin_Dyak_Tai18E2_2.1 (NCBI Release 102) for D. yakuba, ASM1815383v1 (NCBI Release 102) for Drosophila eugracilis, RU_DBia_V1.1 (NCBI Release 103) for Drosophila biarmipes, CBGP_Dsuzu_IsoJpt1.0 (GCF_043229965.1-RS_2025_01) for D. suzukii, DtakHiC1v2 (GCF_030179915.1-RS_2024_12) for Drosophila takahashii, ASM1815250v1 (NCBI Release 102) for Drosophilia elegans, ASM1815211v1 (NCBI Release 102) for Drosophila rhopaloa, ASM1815226v1 (NCBI Release 102) for D. ficusphila, DkikHiC1v2 (GCF_030179895.1-RS_2024_12) for Drosophila kikkawai, ASM1763931v2 (NCBI Release 102) for Drosophila ananassae, DperRS2 (NCBI Release 101) for D. persimilis, UCI_Dpse_MV25 (GCF_009870125.1-RS_2025_01) for Drosophila pseudoobscura, UCI_dwil_1.1 (NCBI Release 102) for Drosophila willistoni, ASM1815329v1 (NCBI Release 103) for Drosophila grimshawi, ASM1815372v1 (NCBI Release 102) for D. mojavensis, Dvir_AGI_RSII-ME (GCF_030788295.1-RS_2024_12) for D. virilis. The rooted species phylogenetic tree was constructed by OrthoFinder (97–99). The tree was visualized in R using ape (100).
Embryo collection and fixation for Micro-C
Aged adult flies were discarded from vials, and newly eclosed flies were collected twice: in the evening of the following day and in the morning 2 days later. Flies were transferred to population cages with apple juice agar plates supplemented with yeast paste and allowed to mate for 1 day at 25°C. On the following morning, plates were exchanged twice, and embryos were collected over a 30-min window and aged for 2 hours and 10 min at 25°C. This resulted in embryo collections corresponding to 2 hours and 10 min to 2 hours and 40 min after egg laying, which corresponds to nc14. Collected embryos were dechorionated by immersion in bleach (Kao) for 2 min, followed by thorough rinsing with water. Embryos were fixed by rotating in a mixture of 6 ml of heptane, 2 ml of PBS-T, and 500 μl of 16% freshly prepared formaldehyde for 10 min. Fixation was quenched by adding 2 ml of 2 M Tris-HCl (pH 7.5) and rotating for 5 min. Embryos were pelleted by centrifugation at 1000g for 1 min at 25°C. Fixed embryos were washed twice with PBS-T and stored at 4°C, while an additional six to seven collection rounds were performed. Pooled embryos were subjected to secondary cross-linking by rotating in a mixture of 1 ml of PBS-T containing 3 mM DSG and 3 mM EGS for 45 min. Fixation was quenched by adding 330 μl of 2 M Tris-HCl (pH 7.5) and rotating for 5 min. Embryos were pelleted by centrifugation at 1000g for 1 min at 25°C, flash frozen in liquid nitrogen, and stored at –80°C until further use.
Micro-C
Approximately 50 μl of nc14 embryos was washed once with PBS-T and centrifuged at 1000g for 1 min at 25°C. The pellet was resuspended in 1 ml of MB1 buffer [10 mM Tris (pH 8.0), 50 mM NaCl, 5 mM MgCl2, 1 mM CaCl2, 0.2% NP-40, and 1× EDTA-free protease inhibitor (Nacalai Tesque)] and homogenized using a Dounce homogenizer. The lysate was transferred to low-bind 1.5-ml tubes and incubated on ice for 20 min. The lysate was centrifuged at 10,000g for 5 min at 4°C, and the pellet was resuspended in 1 ml of MB1 buffer and pipetted 60 times to disperse nuclei. The nuclear suspension was centrifuged again at 10,000g for 5 min at 4°C, and the pellet was resuspended in 100 μl of MB1 buffer and pipetted 60 times to further disperse nuclei. Chromatin was digested with a predetermined amount (~3 to 4 μl) of micrococcal nuclease (New England Biolabs) for 10 min at 37°C to yield a 90% mononucleosome/10% dinucleosome ratio. The digestion was stopped by adding 5 μl of 0.1 M EGTA and incubating for 10 min at 65°C. The digested chromatin was collected by centrifugation at 10,000g for 5 min at 4°C and washed twice with 1 ml of MB2 buffer [10 mM Tris (pH 8.0), 50 mM NaCl, and 5 mM MgCl2]. The pellet was resuspended in 45 μl of end-chewing mix [5 μl of NEBuffer r2.1 (New England Biolabs), 10 μl of 10 mM adenosine 5′-triphosphate (ATP), 2.5 μl of 100 mM DTT, 2.5 μl of T4 PNK (New England Biolabs), and 25 μl of Milli-Q water], pipetted 60 times, and incubated for 15 min at 37°C. After adding 5 μl of Klenow Fragment (New England Biolabs), the reaction was pipetted 60 times and incubated for another 15 min at 37°C. Next, 25 μl of end-labeling mix [5 μl of 1 mM biotin-14-dATP (Active Motif), 5 μl of 1 mM biotin-14-dCTP (Jena Bioscience), 0.5 μl of 10 mM dGTP (Takara), 0.5 μl of 10 mM dTTP (Roche), 2.5 μl of 10× T4 DNA Ligase Buffer (New England Biolabs), 0.25 μl of BSA (10 mg/ml), and 11.25 μl of Milli-Q water] was added. The reaction was pipetted 60 times and incubated for 45 min at 25°C. The reaction was stopped by adding 5 μl of 0.5 M EDTA and incubating for 20 min at 65°C. Chromatin was collected by centrifugation at 12,000g for 5 min at 4°C and washed once with 1 ml of MB3 buffer [10 mM Tris (pH 8.0) and 10 mM MgCl2]. The pellet was resuspended in 250 μl of end-ligation mix [25 μl of 10× T4 DNA Ligase Buffer (New England Biolabs), 2.5 μl of BSA (10 mg/ml), 12.5 μl of T4 DNA Ligase (New England Biolabs; 400 U/μl), and 210 μl of Milli-Q water], pipetted 60 times, and then incubated for 4 hours at room temperature with gentle rotation. The chromatin was collected by centrifugation at 16,000g for 5 min at 4°C and resuspended in 100 μl of unligated-end purification mix [10 μl of 10× NEBuffer #1 (New England Biolabs), 5 μl of Exonuclease III (New England Biolabs; 100 U/μl), and 85 μl of Milli-Q water]. The reaction was pipetted 60 times and incubated for 1 hour at 37°C. Then, 100 μl ChIP elution buffer was added. De–cross-linking and DNA precipitation were performed as described in the ChIP procedure, and the resulting pellet was resuspended in 20 μl of TE buffer.
Micro-C library preparation
Five microliters of Streptavidin C1 beads (Thermo Fisher Scientific) was washed with 300 μl of Tween wash buffer [5 mM Tris (pH 7.5), 0.5 mM EDTA, 1 M NaCl, and 0.05% Tween-20] and resuspended in 300 μl of 2× biotin binding buffer [10 mM Tris (pH 7.5), 1 mM EDTA, and 2 M NaCl]. A mixture of 290 μl of Milli-Q water, 300 μl of the bead suspension, and 10 μl of the Micro-C sample was incubated for 20 min at room temperature with gentle rotation. The bead-bound sample was washed twice with 600 μl of Tween wash buffer and resuspended in 50 μl of 0.1× TE buffer. Sequencing libraries were prepared using the NEBNext Ultra II DNA Library Prep Kit for Illumina (New England Biolabs). For adaptor ligation, the adaptor reagent was diluted 10-fold. After the adaptor ligation step, the beads were washed twice with 200 μl of Tween wash buffer, once with 150 μl of 0.1× TE buffer, and resuspended in 15 μl of 0.1× TE buffer. Adapter-ligated, bead-bound DNA was PCR amplified using Illumina index primers (NEBNext Multiplex Oligos for Illumina, New England Biolabs) for 8 cycles. PCR products ranging from 100 to 1000 bp were size selected and purified using AMPure XP beads (Beckman Coulter). The size-selected libraries were validated using the TapeStation D5000HS kit (Agilent Technologies). The resulting libraries were sent to BGI Japan for circularization and sequenced on a DNBSEQ-G400 platform (MGI Tech) using paired-end 100-bp reads.
Computational analysis of Micro-C data
Micro-C data generated in this study and public Micro-C data from a previous publication (53) were analyzed using the 4DN Hi-C analysis pipeline (https://data.4dnucleome.org/resources/data-analysis/hi_c-processing-pipeline). Briefly, adapter sequences were trimmed from paired-end reads with Fastp (77), and then reads were mapped to the D. melanogaster reference genome (BDGP6.32) using BWA (101) with the option -SP5M. Alignments were filtered to retain reads with alignment quality ≥ 3, sorted, and deduplicated using pairtools (102). Micro-C experiments were performed with at least two biological replicates. Because the replicates showed high correlation, they were merged, and the combined libraries were used for downstream analyses. The deduplicated pairs were then selected using pairtools with the query “(pair_type == “UU”) or (pair_type == “UR”) or (pair_type == “RU”).” Matrix aggregation was performed with Cooler (103). For visualization of contact map or insulation score of Micro-C data, CoolBox (91) was used. For Fig. 10, Micro-C contact matrices were balanced using Cooler to enable quantitative comparison between the two Micro-C samples. Micro-C sequence quality metrics, such as reads number, alignment reads number, and cis/trans ratio, were summarized in table S7.
Reagents and analysis pipelines
Reagents and analysis pipelines used in this study are summarized in table S8, and plasmids are listed in the Supplementary Materials.
Acknowledgments
We thank H. Takishita and M. Sato for fly husbandry and the Bloomington Drosophila Stock Center for fly strains. We are also grateful to Y. Kishi, Y. W. Iwasaki, C. Takeuchi, and R. Nakato for advice on ChIP-seq and Micro-C experiments and members of the Fukaya laboratory for critical comments on the manuscript.
Funding:
T.F. was supported by FOREST (JPMJFR214W) and CREST program (JPMJCR25T2) from the JST (Japan Science and Technology Agency), the Grant-in-Aid for Scientific Research (A) (25H00967), the Grant-in-Aid for Transformative Research Areas (A) (24H02327) from JSPS (Japan Society for the Promotion of Science), and the research grant from the Takeda Science Foundation. R.S. was supported by AMED ASPIRE (JP23jf0126003) and the Grant-in-Aid for Young Scientists (25K18438) from JSPS. Y.U. was supported by JSPS fellowship (24KJ0843). S.M. was supported by a Grant-in-Aid for Scientific Research (C) (23K05631) from the JSPS and the research grants from the Sumitomo Foundation, Nakajima Foundation, and Inamori Foundation.
Author contributions:
Conceptualization: T.F. Methodology: R.S., Y.U., S.M., and T.F. Validation: R.S., Y.U., S.M., and T.F. Formal analysis: R.S., Y.U., S.M., and T.F. Investigation: R.S., Y.U., S.M., and T.F. Software: R.S. Resources: T.F. Data curation: R.S., Y.U., and T.F. Writing—original draft: T.F. Writing—review and editing: R.S., Y.U., S.M., and T.F. Visualization: R.S., Y.U., S.M., and T.F. Supervision: T.F. Project administration: T.F. Funding acquisition: R.S., Y.U., S.M., and T.F.
Competing interests:
The authors declare that they have no competing interests.
Data, code, and materials availability:
All sequencing data have been deposited in the Gene Expression Omnibus under accession code GEO: GSE295515. All raw sequencing reads have been deposited in the NCBI Sequence Read Archive (SRA) under BioProject PRJNA1254419 (www.ncbi.nlm.nih.gov/bioproject/PRJNA1254419). For direct access to the run metadata and data selector, please see the following: www.ncbi.nlm.nih.gov/Traces/study/?acc=PRJNA1254419. All reagents and fly strains generated in this study are available from the corresponding author (T.F., tfukaya@iqb.u-tokyo.ac.jp) upon request. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.
Supplementary Materials
The PDF file includes:
Supplementary Materials and Methods
Figs. S1 to S14
Legends for tables S1 to S8
References
Other Supplementary Material for this manuscript includes the following:
Tables S1 to S8
REFERENCES
- 1.Chung H., Schäfer U., Jäckle H., Böhm S., Genomic expansion and clustering of ZAD-containing C2H2 zinc-finger genes in Drosophila. EMBO Rep. 3, 1158–1162 (2002). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Jauch R., Bourenkov G. P., Chung H.-R., Urlaub H., Reidt U., Jäckle H., Wahl M. C., The zinc finger-associated domain of the Drosophila transcription factor grauzone is a novel zinc-coordinating protein-protein interaction module. Structure 11, 1393–1402 (2003). [DOI] [PubMed] [Google Scholar]
- 3.Payre F., Buono P., Vanzo N., Vincent A., Two types of zinc fingers are required for dimerization of the serendipity δ transcriptional activator. Mol. Cell. Biol. 17, 3137–3145 (1997). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ruez C., Payre F., Vincent A., Transcriptional control of Drosophila bicoid by Serendipity δ: Cooperative binding sites, promoter context, and co-evolution. Mech. Dev. 78, 125–134 (1998). [DOI] [PubMed] [Google Scholar]
- 5.Bonchuk A., Boyko K., Fedotova A., Nikolaeva A., Lushchekina S., Khrustaleva A., Popov V., Georgiev P., Structural basis of diversity and homodimerization specificity of zinc-finger-associated domains in Drosophila. Nucleic Acids Res. 49, 2375–2389 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Bonchuk A. N., Boyko K. M., Nikolaeva A. Y., Burtseva A. D., Popov V. O., Georgiev P. G., Structural insights into highly similar spatial organization of zinc-finger associated domains with a very low sequence similarity. Structure 30, 1004–1015.e4 (2022). [DOI] [PubMed] [Google Scholar]
- 7.Kasinathan B., Colmenares S. U., McConnell H., Young J. M., Karpen G. H., Malik H. S., Innovation of heterochromatin functions drives rapid evolution of essential ZAD-ZNF genes in Drosophila. eLife 9, e63368 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Chung H.-R., Löhr U., Jäckle H., Lineage-specific expansion of the zinc finger associated domain ZAD. Mol. Biol. Evol. 24, 1934–1943 (2007). [DOI] [PubMed] [Google Scholar]
- 9.Komura-Kawa T., Hirota K., Shimada-Niwa Y., Yamauchi R., Shimell M., Shinoda T., Fukamizu A., O’Connor M. B., Niwa R., The Drosophila zinc finger transcription factor Ouija board controls ecdysteroid biosynthesis through specific regulation of spookier. PLOS Genet. 11, e1005712 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Uryu O., Ou Q., Komura-Kawa T., Kamiyama T., Iga M., Syrzycka M., Hirota K., Kataoka H., Honda B. M., King-Jones K., Niwa R., Cooperative control of ecdysone biosynthesis in Drosophila by transcription factors Séance, Ouija board, and molting defective. Genetics 208, 605–622 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Shapiro-Kulnane L., Selengut M., Salz H. K., Safeguarding Drosophila female germ cell identity depends on an H3K9me3 mini domain guided by a ZAD zinc finger protein. PLOS Genet. 18, e1010568 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Swenson J. M., Colmenares S. U., Strom A. R., Costes S. V., Karpen G. H., The composition and organization of Drosophila heterochromatin are heterogeneous and dynamic. eLife 5, e16096 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Chen P., Pan K. C., Park E. H., Luo Y., Lee Y. C. G., Aravin A. A., Escalation of genome defense capacity enables control of an expanding meiotic driver. Proc. Natl. Acad. Sci. U.S.A. 122, e2418541122 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Baumgartner L., Handler D., Platzer S. W., Yu C., Duchek P., Brennecke J., The Drosophila ZAD zinc finger protein Kipferl guides Rhino to piRNA clusters. eLife 11, e80067 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Maksimenko O., Bartkuhn M., Stakhov V., Herold M., Zolotarev N., Jox T., Buxa M. K., Kirsch R., Bonchuk A., Fedotova A., Kyrchanova O., Renkawitz R., Georgiev P., Two new insulator proteins, Pita and ZIPIC, target CP190 to chromatin. Genome Res. 25, 89–99 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Bag I., Chen S., Rosin L. F., Chen Y., Liu C.-Y., Yu G.-Y., Lei E. P., M1BP cooperates with CP190 to activate transcription at TAD borders and promote chromatin insulator activity. Nat. Commun. 12, 4170 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Gaszner M., Vazquez J., Schedl P., The Zw5 protein, a component of the scs chromatin domain boundary, is able to block enhancer-promoter interaction. Genes Dev. 13, 2098–2107 (1999). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Chen L.-Y., Wang J.-C., Hyvert Y., Lin H.-P., Perrimon N., Imler J.-L., Hsu J.-C., Weckle is a zinc finger adaptor of the toll pathway in dorsoventral patterning of the Drosophila embryo. Curr. Biol. 16, 1183–1193 (2006). [DOI] [PubMed] [Google Scholar]
- 19.Irgen-Gioro S., Yoshida S., Walling V., Chong S., Fixation can change the appearance of phase separation in living cells. eLife 11, e79903 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Magnitov M. D., Maresca M., Alonso Saiz N., Teunissen H., Dong J., Sathyan K. M., Braccioli L., Guertin M. J., de Wit E., ZNF143 is a transcriptional regulator of nuclear-encoded mitochondrial genes that acts independently of looping and CTCF. Mol. Cell 85, 24–41.e11 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Narducci D. N., Hansen A. S., Putative looping factor ZNF143/ZFP143 is an essential transcriptional regulator with no looping function. Mol. Cell 85, 9–23.e9 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Schüpbach T., Wieschaus E., Maternal-effect mutations altering the anterior-posterior pattern of the Drosophila embryo. Rouxs Arch. Dev. Biol. 195, 302–317 (1986). [DOI] [PubMed] [Google Scholar]
- 23.Liang L., Diehl-Jones W., Lasko P., Localization of vasa protein to the Drosophila pole plasm is independent of its RNA-binding and helicase activities. Development 120, 1201–1211 (1994). [DOI] [PubMed] [Google Scholar]
- 24.Soeller W. C., Oh C. E., Kornberg T. B., Isolation of cDNAs encoding the Drosophila GAGA transcription factor. Mol. Cell. Biol. 13, 7961–7970 (1993). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Farkas G., Gausz J., Galloni M., Reuter G., Gyurkovics H., Karch F., The Trithorax-like gene encodes the Drosophila GAGA factor. Nature 371, 806–808 (1994). [DOI] [PubMed] [Google Scholar]
- 26.Raff J. W., Kellum R., Alberts B., The Drosophila GAGA transcription factor is associated with specific regions of heterochromatin throughout the cell cycle. EMBO J. 13, 5977–5983 (1994). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Gaskill M. M., Gibson T. J., Larson E. D., Harrison M. M., GAF is essential for zygotic genome activation and chromatin accessibility in the early Drosophila embryo. eLife 10, e66668 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gaskill M. M., Soluri I. V., Branks A. E., Boka A. P., Stadler M. R., Vietor K., Huang H.-Y. S., Gibson T. J., Mukherjee A., Mir M., Blythe S. A., Harrison M. M., Localization of the Drosophila pioneer factor GAF to subnuclear foci is driven by DNA binding and required to silence satellite repeat expression. Dev. Cell 58, 1610–1624.e8 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Bellec M., Dufourt J., Hunt G., Lenden-Hasse H., Trullo A., Zine El Aabidine A., Lamarque M., Gaskill M. M., Faure-Gautron H., Mannervik M., Harrison M. M., Andrau J.-C., Favard C., Radulescu O., Lagha M., The control of transcriptional memory by stable mitotic bookmarking. Nat. Commun. 13, 1176 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Van Bortle K., Corces V. G., Nuclear organization and genome function. Annu. Rev. Cell Dev. Biol. 28, 163–187 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Parkhurst S. M., Harrison D. A., Remington M. P., Spana C., Kelley R. L., Coyne R. S., Corces V. G., The Drosophila su(Hw) gene, which controls the phenotypic effect of the gypsy transposable element, encodes a putative DNA-binding protein. Genes Dev. 2, 1205–1215 (1988). [DOI] [PubMed] [Google Scholar]
- 32.Moon H., Filippova G., Loukinov D., Pugacheva E., Chen Q., Smith S. T., Munhall A., Grewe B., Bartkuhn M., Arnold R., Burke L. J., Renkawitz-Pohl R., Ohlsson R., Zhou J., Renkawitz R., Lobanenkov V., CTCF is conserved from Drosophila to humans and confers enhancer blocking of the Fab-8 insulator. EMBO Rep. 6, 165–170 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Zhao K., Hart C. M., Laemmli U. K., Visualization of chromosomal domains with boundary element-associated factor BEAF-32. Cell 81, 879–889 (1995). [DOI] [PubMed] [Google Scholar]
- 34.Pai C.-Y., Lei E. P., Ghosh D., Corces V. G., The centrosomal protein CP190 is a component of the gypsy chromatin insulator. Mol. Cell 16, 737–748 (2004). [DOI] [PubMed] [Google Scholar]
- 35.Richter C., Oktaba K., Steinmann J., Müller J., Knoblich J. A., The tumour suppressor L(3)mbt inhibits neuroepithelial proliferation and acts on insulator elements. Nat. Cell Biol. 13, 1029–1039 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Kellner W. A., Van Bortle K., Li L., Ramos E., Takenaka N., Corces V. G., Distinct isoforms of the Drosophila Brd4 homologue are present at enhancers, promoters and insulator sites. Nucleic Acids Res. 41, 9274–9283 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Gerasimova T. I., Corces V. G., Polycomb and trithorax group proteins mediate the function of a chromatin insulator. Cell 92, 511–521 (1998). [DOI] [PubMed] [Google Scholar]
- 38.Gerasimova T. I., Byrd K., Corces V. G., A chromatin insulator determines the nuclear localization of DNA. Mol. Cell 6, 1025–1035 (2000). [DOI] [PubMed] [Google Scholar]
- 39.Ghosh D., Gerasimova T. I., Corces V. G., Interactions between the Su(Hw) and Mod(mdg4) proteins required for gypsy insulator function. EMBO J. 20, 2518–2527 (2001). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Capelson M., Corces V. G., The ubiquitin ligase dTopors directs the nuclear organization of a chromatin insulator. Mol. Cell 20, 105–116 (2005). [DOI] [PubMed] [Google Scholar]
- 41.Golovnin A., Melnikova L., Volkov I., Kostuchenko M., Galkin A. V., Georgiev P., ‘Insulator bodies’ are aggregates of proteins but not of insulators. EMBO Rep. 9, 440–445 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Schoborg T., Rickels R., Barrios J., Labrador M., Chromatin insulator bodies are nuclear structures that form in response to osmotic stress and cell death. J. Cell Biol. 202, 261–276 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Cavalheiro G. R., Girardot C., Viales R. R., Pollex T., Cao T. B. N., Lacour P., Feng S., Rabinowitz A., Furlong E. E. M., CTCF, BEAF-32, and CP190 are not required for the establishment of TADs in early Drosophila embryos but have locus-specific roles. Sci. Adv. 9, eade1085 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Li J., Gilmour D. S., Distinct mechanisms of transcriptional pausing orchestrated by GAGA factor and M1BP, a novel transcription factor. EMBO J. 32, 1829–1841 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Hug C. B., Grimaldi A. G., Kruse K., Vaquerizas J. M., Chromatin architecture emerges during zygotic genome activation independent of transcription. Cell 169, 216–228.e19 (2017). [DOI] [PubMed] [Google Scholar]
- 46.Abramson J., Adler J., Dunger J., Evans R., Green T., Pritzel A., Ronneberger O., Willmore L., Ballard A. J., Bambrick J., Bodenstein S. W., Evans D. A., Hung C.-C., O’Neill M., Reiman D., Tunyasuvunakool K., Wu Z., Žemgulytė A., Arvaniti E., Beattie C., Bertolli O., Bridgland A., Cherepanov A., Congreve M., Cowen-Rivers A. I., Cowie A., Figurnov M., Fuchs F. B., Gladman H., Jain R., Khan Y. A., Low C. M. R., Perlin K., Potapenko A., Savy P., Singh S., Stecula A., Thillaisundaram A., Tong C., Yakneen S., Zhong E. D., Zielinski M., Žídek A., Bapst V., Kohli P., Jaderberg M., Hassabis D., Jumper J. M., Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Beugelink J. W., Sweep E., Janssen B. J. C., Snijder J., Pronker M. F., Structural basis for recognition of the FLAG-tag by anti-FLAG M2. J. Mol. Biol. 436, 168649 (2024). [DOI] [PubMed] [Google Scholar]
- 48.Atinbayeva N., Valent I., Zenk F., Loeser E., Rauer M., Herur S., Quarato P., Pyrowolakis G., Gomez-Auli A., Mittler G., Cecere G., Erhardt S., Tiana G., Zhan Y., Iovino N., Inheritance of H3K9 methylation regulates genome architecture in Drosophila early embryos. EMBO J. 43, 2685–2714 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Zhang X., Koolhaas W. H., Schnorrer F., A versatile two-step CRISPR- and RMCE-based strategy for efficient genome engineering in Drosophila. G3 4, 2409–2418 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Tatar M., Kopelman A., Epstein D., Tu M.-P., Yin C.-M., Garofalo R. S., A mutant Drosophila insulin receptor homolog that extends life-span and impairs neuroendocrine function. Science 292, 107–110 (2001). [DOI] [PubMed] [Google Scholar]
- 51.Clancy D. J., Gems D., Harshman L. G., Oldham S., Stocker H., Hafen E., Leevers S. J., Partridge L., Extension of life-span by loss of CHICO, a Drosophila insulin receptor substrate protein. Science 292, 104–106 (2001). [DOI] [PubMed] [Google Scholar]
- 52.Giannakou M. E., Goss M., Jünger M. A., Hafen E., Leevers S. J., Partridge L., Long-lived Drosophila with overexpressed dFOXO in adult fat body. Science 305, 361–361 (2004). [DOI] [PubMed] [Google Scholar]
- 53.Batut P. J., Bing X. Y., Sisco Z., Raimundo J., Levo M., Levine M. S., Genome organization controls transcriptional dynamics during development. Science 375, 566–570 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Kyrchanova O., Mogila V., Wolle D., Magbanua J. P., White R., Georgiev P., Schedl P., The boundary paradox in the Bithorax complex. Mech. Dev. 138, 122–132 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Kyrchanova O., Ivlieva T., Toshchakov S., Parshikov A., Maksimenko O., Georgiev P., Selective interactions of boundaries with upstream region of Abd-B promoter in Drosophila bithorax complex and role of dCTCF in this process. Nucleic Acids Res. 39, 3042–3052 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56.Li M., Ma Z., Liu J. K., Roy S., Patel S. K., Lane D. C., Cai H. N., An organizational hub of developmentally regulated chromatin loops in the Drosophila antennapedia complex. Mol. Cell. Biol. 35, 4018–4029 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Belozerov V. E., Majumder P., Shen P., Cai H. N., A novel boundary element may facilitate independent gene regulation in the Antennapedia complex of Drosophila. EMBO J. 22, 3113–3121 (2003). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Fujioka M., Wu X., Jaynes J. B., A chromatin insulator mediates transgene homing and very long-range enhancer-promoter communication. Development 136, 3077–3087 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Fujioka M., Sun G., Jaynes J. B., The Drosophila eve insulator Homie promotes eve expression and protects the adjacent gene from repression by polycomb spreading. PLOS Genet. 9, e1003883 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Fujioka M., Mistry H., Schedl P., Jaynes J. B., Determinants of chromosome architecture: Insulator Pairing in cis and in trans. PLOS Genet. 12, e1005889 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Li X., Tang X., Bing X., Catalano C., Li T., Dolsten G., Wu C., Levine M., GAGA-associated factor fosters loop formation in the Drosophila genome. Mol. Cell 83, 1519–1526.e4 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Kaushal A., Mohana G., Dorier J., Özdemir I., Omer A., Cousin P., Semenova A., Taschner M., Dergai O., Marzetta F., Iseli C., Eliaz Y., Weisz D., Shamim M. S., Guex N., Lieberman Aiden E., Gambetta M. C., CTCF loss has limited effects on global genome architecture in Drosophila despite critical regulatory functions. Nat. Commun. 12, 1011 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Weintraub A. S., Li C. H., Zamudio A. V., Sigova A. A., Hannett N. M., Day D. S., Abraham B. J., Cohen M. A., Nabet B., Buckley D. L., Guo Y. E., Hnisz D., Jaenisch R., Bradner J. E., Gray N. S., Young R. A., YY1 Is a Structural Regulator of enhancer-promoter loops. Cell 171, 1573–1588.e28 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Deng W., Lee J., Wang H., Miller J., Reik A., Gregory P. D., Dean A., Blobel G. A., Controlling long-range genomic interactions at a native locus by targeted tethering of a looping factor. Cell 149, 1233–1244 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Larson A. G., Elnatan D., Keenen M. M., Trnka M. J., Johnston J. B., Burlingame A. L., Agard D. A., Redding S., Narlikar G. J., Liquid droplet formation by HP1α suggests a role for phase separation in heterochromatin. Nature 547, 236–240 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Strom A. R., Emelyanov A. V., Mir M., Fyodorov D. V., Darzacq X., Karpen G. H., Phase separation drives heterochromatin domain formation. Nature 547, 241–245 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Plys A. J., Davis C. P., Kim J., Rizki G., Keenen M. M., Marr S. K., Kingston R. E., Phase separation of Polycomb-repressive complex 1 is governed by a charged disordered region of CBX2. Genes Dev. 33, 799–813 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Venken K. J. T., He Y., Hoskins R. A., Bellen H. J., P[acman]: A BAC transgenic platform for targeted insertion of large DNA fragments in D. melanogaster. Science 314, 1747–1751 (2006). [DOI] [PubMed] [Google Scholar]
- 69.Ringrose L., Transgenesis in Drosophila melanogaster. Methods Mol. Biol. 561, 3–19 (2009). [DOI] [PubMed] [Google Scholar]
- 70.Hamamoto K., Umemura Y., Makino S., Fukaya T., Dynamic interplay between non-coding enhancer transcription and gene activity in development. Nat. Commun. 14, 826 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Kawasaki K., Fukaya T., Functional coordination between transcription factor clustering and gene activity. Mol. Cell 83, 1605–1622.e9 (2023). [DOI] [PubMed] [Google Scholar]
- 72.Calvo L., Ronshaugen M., Pettini T., smiFISH and embryo segmentation for single-cell multi-gene RNA quantification in arthropods. Commun. Biol. 4, 352 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Tsanov N., Samacoits A., Chouaib R., Traboulsi A.-M., Gostan T., Weber C., Zimmer C., Zibara K., Walter T., Peter M., Bertrand E., Mueller F., smiFISH and FISH-quant – A flexible single RNA detection approach with super-resolution capability. Nucleic Acids Res. 44, e165 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Linford N. J., Bilgir C., Ro J., Pletcher S. D., Measurement of lifespan in Drosophila melanogaster. J. Vis. Exp. 71, e50068 (2013). [Google Scholar]
- 75.Virtanen P., Gommers R., Oliphant T. E., Haberland M., Reddy T., Cournapeau D., Burovski E., Peterson P., Weckesser W., Bright J., van der Walt S. J., Brett M., Wilson J., Millman K. J., Mayorov N., Nelson A. R. J., Jones E., Kern R., Larson E., Carey C. J., Polat İ., Feng Y., Moore E. W., VanderPlas J., Laxalde D., Perktold J., Cimrman R., Henriksen I., Quintero E. A., Harris C. R., Archibald A. M., Ribeiro A. H., Pedregosa F., van Mulbregt P., Vijaykumar A., Bardelli A. P., Rothberg A., Hilboll A., Kloeckner A., Scopatz A., Lee A., Rokem A., Woods C. N., Fulton C., Masson C., Häggström C., Fitzgerald C., Nicholson D. A., Hagen D. R., Pasechnik D. V., Olivetti E., Martin E., Wieser E., Silva F., Lenders F., Wilhelm F., Young G., Price G. A., Ingold G.-L., Allen G. E., Lee G. R., Audren H., Probst I., Dietrich J. P., Silterra J., Webber J. T., Slavič J., Nothman J., Buchner J., Kulick J., Schönberger J. L., de Miranda Cardoso J. V., Reimer J., Harrington J., Rodríguez J. L. C., Nunez-Iglesias J., Kuczynski J., Tritz K., Thoma M., Newville M., Kümmerer M., Bolingbroke M., Tartre M., Pak M., Smith N. J., Nowaczyk N., Shebanov N., Pavlyk O., Brodtkorb P. A., Lee P., McGibbon R. T., Feldbauer R., Lewis S., Tygier S., Sievert S., Vigna S., Peterson S., More S., Pudlik T., Oshima T., Pingel T. J., Robitaille T. P., Spura T., Jones T. R., Cera T., Leslie T., Zito T., Krauss T., Upadhyay U., Halchenko Y. O., Vázquez-Baeza Y., SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat. Methods 17, 261–272 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Thompson M. K., Kiourlappou M., Davis I., Ribo-Pop: Simple, cost-effective, and widely applicable ribosomal RNA depletion. RNA 26, 1731–1742 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Chen S., Zhou Y., Chen Y., Gu J., fastp: An ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Dobin A., Davis C. A., Schlesinger F., Drenkow J., Zaleski C., Jha S., Batut P., Chaisson M., Gingeras T. R., STAR: Ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Li H., Handsaker B., Wysoker A., Fennell T., Ruan J., Homer N., Marth G., Abecasis G., Durbin R., Genome Project Data Processing Subgroup , The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078–2079 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Jin Y., Tam O. H., Paniagua E., Hammell M., TEtranscripts: A package for including transposable elements in differential expression analysis of RNA-seq datasets. Bioinformatics 31, 3593–3599 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Love M. I., Huber W., Anders S., Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15, 550 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Langmead B., Salzberg S. L., Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Zhang Y., Liu T., Meyer C. A., Eeckhoute J., Johnson D. S., Bernstein B. E., Nusbaum C., Myers R. M., Brown M., Li W., Liu X. S., Model-based analysis of ChIP-seq (MACS). Genome Biol. 9, R137 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Jain D., Baldi S., Zabel A., Straub T., Becker P. B., Active promoters give rise to false positive ‘Phantom Peaks’ in ChIP-seq experiments. Nucleic Acids Res. 43, 6959–6968 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Quinlan A. R., Hall I. M., BEDTools: A flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Ramírez F., Ryan D. P., Grüning B., Bhardwaj V., Kilpert F., Richter A. S., Heyne S., Dündar F., Manke T., deepTools2: A next generation web server for deep-sequencing data analysis. Nucleic Acids Res. 44, W160–W165 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87.Lawrence M., Huber W., Pagès H., Aboyoun P., Carlson M., Gentleman R., Morgan M. T., Carey V. J., Software for computing and annotating genomic ranges. PLOS Comput. Biol. 9, e1003118 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Conway J. R., Lex A., Gehlenborg N., UpSetR: An R package for the visualization of intersecting sets and their properties. Bioinformatics 33, 2938–2940 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 89.Shao Z., Zhang Y., Yuan G.-C., Orkin S. H., Waxman D. J., MAnorm: A robust model for quantitative comparison of ChIP-seq data sets. Genome Biol. 13, R16 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 90.Robinson J. T., Thorvaldsdóttir H., Winckler W., Guttman M., Lander E. S., Getz G., Mesirov J. P., Integrative genomics viewer. Nat. Biotechnol. 29, 24–26 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Xu W., Zhong Q., Lin D., Zuo Y., Dai J., Li G., Cao G., CoolBox: A flexible toolkit for visual analysis of genomics data. BMC Bioinformatics 22, 489 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 92.Nakato R., Itoh T., Shirahige K., DROMPA: Easy-to-handle peak calling and visualization software for the computational analysis and validation of ChIP-seq data. Genes Cells 18, 589–601 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 93.Nakato R., Shirahige K., Sensitive and robust assessment of ChIP-seq read distribution using a strand-shift profile. Bioinformatics 34, 2356–2363 (2018). [DOI] [PubMed] [Google Scholar]
- 94.Bailey T. L., Johnson J., Grant C. E., Noble W. S., The MEME Suite. Nucleic Acids Res. 43, W39–W49 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 95.Pabo C. O., Peisach E., Grant R. A., Design and selection of novel Cys2 His2 zinc finger proteins. Annu. Rev. Biochem. 70, 313–340 (2001). [DOI] [PubMed] [Google Scholar]
- 96.Cock P. J. A., Antao T., Chang J. T., Chapman B. A., Cox C. J., Dalke A., Friedberg I., Hamelryck T., Kauff F., Wilczynski B., de Hoon M. J. L., Biopython: Freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422–1423 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 97.D. M. Emms, S. Kelly, STAG: Species tree inference from all genes. bioRxiv 267914 [Preprint] (2018). 10.1101/267914. [DOI]
- 98.Emms D. M., Kelly S., OrthoFinder: Phylogenetic orthology inference for comparative genomics. Genome Biol. 20, 238 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 99.Emms D. M., Kelly S., STRIDE: Species tree root inference from gene duplication events. Mol. Biol. Evol. 34, 3267–3278 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 100.Paradis E., Schliep K., Ape 5.0: An environment for modern phylogenetics and evolutionary analyses in R. Bioinformatics 35, 526–528 (2019). [DOI] [PubMed] [Google Scholar]
- 101.Li H., Durbin R., Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25, 1754–1760 (2009). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 102.Open2C, Abdennur N., Fudenberg G., Flyamer I. M., Galitsyna A. A., Goloborodko A., Imakaev M., Venev S. V., Pairtools: From sequencing data to chromosome contacts. PLOS Comput. Biol. 20, e1012164 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 103.Abdennur N., Mirny L. A., Cooler: Scalable storage for Hi-C data and other genomically labeled arrays. Bioinformatics 36, 311–316 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 104.Pettersen E. F., Goddard T. D., Huang C. C., Meng E. C., Couch G. S., Croll T. I., Morris J. H., Ferrin T. E., UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. 30, 70–82 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 105.Fukaya T., Lim B., Levine M., Rapid rates of Pol II elongation in the Drosophila embryo. Curr. Biol. 27, 1387–1391 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Materials and Methods
Figs. S1 to S14
Legends for tables S1 to S8
References
Tables S1 to S8
Data Availability Statement
All sequencing data have been deposited in the Gene Expression Omnibus under accession code GEO: GSE295515. All raw sequencing reads have been deposited in the NCBI Sequence Read Archive (SRA) under BioProject PRJNA1254419 (www.ncbi.nlm.nih.gov/bioproject/PRJNA1254419). For direct access to the run metadata and data selector, please see the following: www.ncbi.nlm.nih.gov/Traces/study/?acc=PRJNA1254419. All reagents and fly strains generated in this study are available from the corresponding author (T.F., tfukaya@iqb.u-tokyo.ac.jp) upon request. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.










