Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

bioRxiv logoLink to bioRxiv
[Preprint]. 2026 May 21:2026.05.19.726322. [Version 1] doi: 10.64898/2026.05.19.726322

Combinatorial pioneer transcription factor binding reinforces bivalent epigenetic states to preserve lineage fidelity

Gerardo Mirizio 1, Morgan Buckley 1, Katie Ludwig 1, Satoshi Matsui 1, Samuel Sampson 1, Hee-Woong Lim 2, Makiko Iwafuchi 1,*
PMCID: PMC13228549  PMID: 42239318

Abstract

Combinatorial binding of transcription factors (TFs) is central to eukaryotic gene regulation, providing regulatory specificity and robustness to cell fate control. However, its impact on epigenetic regulation remains poorly understood. Here, we show that the pioneer TFs GATA6, EOMES, and SOX17 cooperate with the zinc finger TF PRDM1 to recruit Polycomb Repressive Complexes (PRCs) and establish enhancers marked by H3K4me1 and PRC-associated histone modifications during endoderm development. Increasing the number and diversity of pioneer TFs bound at enhancers drives synergistic nucleosome remodeling and promotes the formation of “hyper-bivalent” enhancers that reinforce repression of alternative-lineage programs. Together, our findings demonstrate that combinatorial pioneer TF binding creates locally accessible regions that facilitate recruitment of not only active but also PRC-associated epigenetic regulators to preserve lineage fidelity during development.

INTRODUCTION

Combinatorial binding of transcription factors (TFs) occurs across thousands of lineage-specific enhancers and represents a fundamental principle of eukaryotic gene regulation1,2. Functionally, the presence of multiple binding sites within enhancers can increase cell-type specificity and strengthen TF occupancy, thereby enabling robust lineage-specific gene expression3,4. Mechanistically, TFs cooperate directly through protein-protein interactions or indirectly through nucleosome-mediated or DNA-guided mechanisms5–8. Among TFs, a specialized subset known as pioneer TFs possesses the unique ability to bind nucleosomal DNA and locally open compacted chromatin, thereby making lineage-specific enhancers accessible to additional TFs, which in turn stabilize pioneer TF binding9. Thus, even pioneer TFs rely on combinatorial binding with other TFs to achieve stable, cell-type-specific chromatin occupancy. An important unanswered question is how combinatorial TF binding impacts transcriptional function beyond stabilizing enhancer occupancy.

Accumulating evidence indicates that pioneer TFs directly or indirectly recruit activating epigenetic regulators, including CBP/p300 for depositing H3K27ac (an active enhancer mark) and KMT2C/2D for H3K4me1 (an active/primed enhancer mark)10–12. We recently demonstrated that the pioneer TF FOXA cooperates with the PR domain zinc finger TF PRDM1 to establish a distinct “bivalent” state that is transcriptionally repressed yet poised for activation, characterized by H3K4me1 and Polycomb Repressive Complex (PRC)-associated modifications (H2AK119ub1 and H3K27me3)13. These findings raised the possibility that combinatorial TF codes are translated into distinct epigenetic states that ultimately impact transcriptional regulation.

To test this hypothesis, we built upon our prior discovery of pioneer TF-PRDM1 cooperation during endoderm differentiation as a tractable model13. Using a doxycycline-inducible CRISPRi system, we show that PRDM1 knockdown in definitive endoderm results in robust de-repression of genes associated with alternative lineage programs, establishing PRDM1 as a key regulator of lineage fidelity. Mechanistically, we found that multiple pioneer TFs (GATA6, SOX17, and EOMES) in definitive endoderm cooperate with PRDM1 to promote the establishment of bivalent chromatin states marked by PRC-associated modifications and H3K4me1 at target regulatory elements, extending our previous observations of FOXA-PRDM1 cooperation in foregut endoderm and demonstrating that this mechanism represents a general feature of pioneer TF function. Integrative cistromic, protein-protein interaction analyses, and massively parallel reporter assays revealed that combinatorial binding of pioneer TFs and PRDM1 modulates regulatory output in either antagonistic or synergistic manners, depending on the chromatin context established upon their binding. At active or weak enhancers, pioneer TF activity counteracts PRDM1-mediated repression, whereas at bivalent enhancers, they cooperate to reinforce gene silencing. Notably, the chromatin accessibility at PRDM1 binding sites and enrichment of bivalent marks at the flanking regions progressively increase with the number and diversity of pioneer TFs bound, culminating in the formation of “hyper-bivalent” enhancers. Together, our findings support a model in which combinatorial pioneer TF binding establishes locally accessible regions that reinforce recruitment of not only activating regulators, but also the PRC machinery, thereby stabilizing epigenetic states and preserving lineage fidelity.

RESULTS

PRDM1 safeguards endoderm identity by repressing mesodermal and ectodermal genes

Several studies have demonstrated PRDM1 expression in anterior embryonic structures in mouse, frog, and zebrafish, suggesting a role in endoderm specification during vertebrate development14–16. However, the temporal dynamics of PRDM1 expression and its function in human endoderm remain poorly understood. To define the expression dynamics and function of PRDM1 in human endoderm differentiation, we differentiated human pluripotent stem cells (hPSCs) along the endodermal lineage using a stepwise protocol to generate homogeneous populations of mesendoderm (ME), definitive endoderm (DE), posterior foregut (pFG), and liver bud progenitor (LBP) cells (Figure 1A)13,17. RNA sequencing (RNA-seq) analysis revealed that PRDM1 expression was induced upon exit from pluripotency, peaked at the DE and pFG stages, and declined during LBP differentiation (Figure 1B).

Figure 1. PRDM1 safeguards endodermal identity by repressing mesodermal and ectodermal genes.

Figure 1.

(a) Schematic representation of the differentiation protocol from Human Pluripotent Stem Cells (hPSCs) to Liver Bud Progenitors (LBP) and equivalent stages in early mouse development (Adapted from Ang et al., Cell Reports, 2018).

(b) Transcription factor expression matrix and line plot derived from RNA-seq experiments showing the expression profile of marker genes at each differentiation stage. Note that PRDM1 is highly expressed at the definitive endoderm (DE) and posterior foregut (pFG) stages.

(c) Schematic representation of the doxycycline-inducible CRISPRi system to knockdown PRDM1 expression. Note that the sgRNAs targeting PRDM1 are constitutively expressed whereas the dCas9 enzyme is induced upon doxycycline administration.

(d) Assessment of PRDM1 knockdown efficiency by Western Blot and Immunofluorescence during differentiation from hPSCs to pFG (n = 3 independent CRISPRi monoclonal cell lines).

(e) Schematic representation of RNA-seq experiments and Gene Ontology analysis in Control and PRDM1-KD at DE and pFG stages.

(f) Volcano plots and principal component analysis (PCA) of RNA-seq data from Control and PRDM1-KD at DE and pFG stages. Differential expression was assessed using the Wald test, with P-values corrected for multiple testing using the Benjamini–Hochberg method (FDR ≤ 0.05) and an absolute log2 fold change ≤ 1.5. PCA was performed on variance-stabilized counts.

(g) Gene Ontology (GO) enrichment analysis of differentially expressed genes at DE upon PRDM1 knockdown. GO terms were summarized using the ReviGO R package. Statistical significance was determined using the hypergeometric test with Benjamini–Hochberg correction (FDR ≤ 0.05).

To investigate the function of PRDM1 in endoderm differentiation, we used a doxycycline (Dox)-inducible CRISPR interference (CRISPRi) system targeting PRDM1 (PRDM1-KD) and assessed the effects of PRDM1 knockdown in DE and pFG cells (Figure 1C)13. Dox treatment initiated at the onset of differentiation resulted in greater than 75% reduction of PRDM1 expression at the DE and pFG stages, a developmental window in which PRDM1 is robustly expressed (Figure 1D). To identify PRDM1-regulated genes, we performed RNA-seq in control and PRDM1-KD cells at the DE and pFG stages (Figure 1E). PRDM1 knockdown resulted in 570 differentially expressed genes at the DE stage and 2,049 at the pFG stage, likely reflecting increased secondary effects at the later developmental stage (FC ≥ 1.5; FDR ≤ 0.05) (Figure 1F). To prioritize potential direct targets of PRDM1, we focused subsequent analyses on the DE stage. Gene Ontology analysis of up-regulated genes revealed enrichment for mesoderm-associated developmental programs (e.g., muscle cell fate, heart morphogenesis, cartilage development, and ossification) and ectoderm-associated programs (e.g., synapse assembly, synaptic transmission, and cell migration). These findings suggest that PRDM1 safeguards endoderm identity by repressing alternative lineage gene programs during endoderm development (Figure 1G).

PRDM1 promotes Polycomb-associated epigenetic states at mesodermal and ectodermal genes through physical interactions with PRC1 and PRC2

While PRDM1 acts as a DNA-binding scaffold protein and recruits co-repressors and epigenetic modifiers to enforce gene silencing18, its epigenetic partners appear to be highly context-dependent19–21. To investigate the epigenetic mechanisms underlying PRDM1-mediated gene regulation in definitive endoderm, we performed chromatin immunoprecipitation followed by sequencing (ChIP-seq) for PRDM1 and a comprehensive panel of histone modifications, including active (H3K27ac), active/primed enhancer-associated (H3K4me1), active/primed promoter-associated (H3K4me3), PRC2-mediated (H3K27me3), PRC1-mediated (H2AK119ub1) and heterochromatin-related (H3K9me2/3) marks in control and PRDM1-KD DE cells (Figure 2A). To assess the direct effects of PRDM1, we focused our analysis on histone marks surrounding PRDM1 binding sites. PRDM1 knockdown in DE cells resulted in a significant reduction of PRC-associated histone marks (H2AK119ub1, p ≤ 0.01; H3K27me3, p ≤ 0.01) and H3K9me2 (p ≤ 0.01), along with an increase in the active histone mark H3K27ac (p ≤ 0.01). In contrast, no significant changes were observed in H3K9me3, H3K4me1, and H3K4me3 marks (Figure 2B). These results suggest that PRDM1 contributes to the establishment and/or maintenance of both Polycomb-associated repressive chromatin marks during early endoderm development. The reduction in H3K9me2 is also consistent with previous studies showing that PRDM1 nucleates the formation of a complex containing the H3K9me2 methyltransferase G9a20.

Figure 2. PRDM1 promotes Polycomb-associated epigenetic states at mesodermal and ectodermal genes through physical interactions with PRC1 and PRC2.

Figure 2.

(a) Schematic representation of TF and Histone ChIP-seq experiments on Control and PRDM1-KD cell lines at the DE stage.

(b) Average plots of ChIP-seq signal centered around PRDM1 peaks over active (H3K27ac), enhancer-related (H3K4me1), promoter-related (H3K4me3), PRC2-mediated (H3K27me3), PRC1-mediated (H2AK119ub1) and heterochromatin-related (H3K9me2/3) post-translational histone modifications in Control and PRDM1-KD cell lines at DE. Statistical significance was assessed by a Wilcoxon rank-sum test comparing Ctrl and KD samples. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

(c) Co-immunoprecipitation of PRDM1 followed by immunoblotting for RING1B, RYBP, and SUZ12 in nuclear extracts from cells at the DE stage.

(d) Heatmap of Histone ChIP-seq signal centered around PRDM1 peaks and clustered by ChromHMM categories at the DE stage. Annotation plots on the right show the difference of signal between PRDM1-KD and Control for each individual peak. GO pathway enrichement of Biological Processes associated with each peak cluster was performed using GREAT (hg38) with default association rules. Statistical significance was determined using the binomial test, and multiple testing correction was applied using the Benjamini–Hochberg method (FDR ≤ 0.05).

To test whether PRDM1 physically interacts with PRCs in DE cells, we performed co-immunoprecipitation (co-IP) experiments of endogenous nuclear extracts treated with Benzonase nuclease to minimize the possibility that detected interactions arise from DNA-mediated proximity. Co-IP experiments showed that PRDM1 interacts with the PRC1 core and catalytic subunit RING1B/RNF2, the non-canonical PRC1 core subunit RYBP, and the PRC2 core subunit SUZ12 (Figure 2C), suggesting that PRDM1 recruits PRCs to establish repressive chromatin domains at the DE stage.

To further characterize the chromatin landscape associated with PRDM1 occupancy, we annotated all PRDM1 binding sites using a Chromatin Hidden Markov Model (ChromHMM) dataset generated from hPSC-derived endoderm cells by the ENCODE Project. This allowed us to integrate our individual epigenomic datasets into unified chromatin states, enabling interpretation of PRDM1 occupancy in relation to defined epigenomic annotations. Genome-wide clustering of histone modifications surrounding PRDM1 binding sites revealed eight distinct chromatin states that exhibit hallmarks of active promoters (TssA; n = 582), gene bodies (GeneB; n = 432), active enhancers (EnhA; n = 251), weak enhancers (EnhWk; n = 568), bivalent promoters (TssBiv; n = 68), bivalent enhancers (EnhBiv; n = 879), heterochromatin (Het; n = 107), and quiescent regions (Quies; n = 590) (Figure 2D). As expected, active promoters and enhancer chromatin states are enriched for H3K27ac, whereas all other bins lack H3K27ac enrichment and are therefore largely silent. Notably, a substantial fraction of PRDM1 binding sites localized to bivalent enhancer regions, characterized by high enrichment of H3K4me1 together with PRC1- and PRC2-associated marks. Genes associated with these regions were enriched for mesoderm- and ectoderm-related biological processes such as “muscle organ development”, “endocardium morphogenesis,” and “neuron development.” Importantly, bivalent enhancer regions exhibited the most pronounced reductions in H2AK119ub1 and H3K27me3 upon PRDM1 knockdown. Interestingly, the marked increase in H3K27ac upon PRDM1 knockdown was observed at a subset of active promoters and active/weak enhancer regions, suggesting that PRDM1 contributes not only to repression of alternative lineage programs but also to fine-tuning the expression levels of actively transcribed genes in DE cells. Because H3K9me2-associated heterochromatin represented only a minor fraction of PRDM1 binding sites and lacked clear gene ontology enrichment (Figure 2D), subsequent analyses focused on PRDM1-mediated regulation of Polycomb-associated chromatin states. Together, these findings indicate that PRDM1 mediates transcriptional repression during endoderm development through physical interactions with PRC1 and PRC2 components and through the establishment of Polycomb-associated chromatin domains.

PRDM1 cooperates with multiple pioneer TFs to establish bivalent enhancers

Emerging evidence suggests that PRDM family TFs cooperate with diverse partner TFs to control lineage specification22,23. While we previously showed that PRDM1 cooperates with FOXA family TFs at the pFG stage, FOXA TFs are not fully induced at the DE stage. To identify specific TF partners of PRDM1 at the DE stage, we analyzed all 3,484 PRDM1-bound regions stratified by ChromHMM category using differential motif analysis. PRDM1-bound active and weak enhancers were enriched for motifs associated with FGF and TGF-β signaling pathway TFs, such as FOS/JUN (AP1 complex) and SMAD4. In contrast, PRDM1-bound bivalent enhancers were enriched for motifs of the definitive endoderm pioneer TFs GATA6, SOX17, and EOMES (Figure 3A). Given their high expression levels at the DE stage and their potential role in enhancer bivalency, we investigated the relationship between PRDM1 and these endoderm-enriched pioneer TFs. To this end, we performed ChIP-seq for GATA6, SOX17, and EOMES in control and PRDM1-KD DE cells (Figure 3B). PRDM1 knockdown reduced, but did not abolish, GATA6, SOX17, and EOMES occupancy at co-bound sites with PRDM1 (Figure 3C), suggesting that PRDM1 contributes to the stabilization of pioneer TF binding. Importantly, expression levels of GATA6, SOX17, and EOMES remained unchanged in PRDM1-KD cells (Figure 3C), indicating that the reduced occupancy was not secondary to altered TF expression. Analysis of unique and overlapping binding sites revealed a marked redistribution of pioneer TF binding from weak enhancers and gene bodies at pioneer-TF-only sites toward bivalent enhancers at regions co-bound by pioneer TFs and PRDM1 within a 200-bp window (Figure 3D). Genes associated with these co-bound regions included key regulators of ectoderm, mesoderm, and primordial germ cell specification that became upregulated following PRDM1 knockdown (Figure 3E).

Figure 3. PRDM1 cooperates with multiple pioneer TFs to establish bivalent enhancers.

Figure 3.

(a) Heatmap of transcription factor (TF) motif enrichment across ChromHMM states. Z-scores indicate relative motif enrichment across states.

(b) Schematic representation of pioneer TF ChIP-seq experiments on Control and PRDM1-KD cell lines at the DE stage.

(c) Average plots of ChIP-seq signal for GATA6, SOX17 and EOMES TFs centered around PRDM1 peaks in Control and PRDM1-KD cell lines at DE. Statistical significance was assessed by a Wilcoxon rank-sum test comparing Ctrl and KD samples. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

(d) Bar plots showing the percentage of ChIP-seq peaks assigned to each ChromHMM state for regions bound individually or co-bound by pioneer TFs and PRDM1, relative to the total number of peaks in each group.

(e) Genome browser tracks showing PRDM1 and pTF ChIP-seq signal, RNA-seq coverage in Control and PRDM1-KD cells, and ChromHMM state annotations at representative loci.

(f) Average plots of ChIP-seq signal in CRISPRi-Ctrl cells, centered around PRDM1 peaks, for PRC1-mediated (H2AK119ub1) and PRC2-mediated (H3K27me3) histone modifications as well as PRC1 (RING1B, RYBP) and PRC2 (SUZ12) subunits, across sites bound by PRDM1 alone, pTFs alone or co-bound by PRDM1 and pTFs.

(g) Schematic representation of PRDM1 and pTFs co-immunoprecipitation experiments in Control and RING1A/B-KD cell lines at DE.

(h) Co-immunoprecipitation of pTFs followed by immunoblotting for PRDM1 and RING1B in Control and RING1A/B-KD cells at the DE stage.

Next, we investigated whether endoderm pioneer TFs and PRDM1 cooperate to establish Polycomb-associated chromatin domains. We performed ChIP-seq for the PRC1 subunits RING1B and RYBP, the PRC2 subunit SUZ12, and the PRC-associated histone marks, and aggregated their signals around pioneer TFs-only, PRDM1-only, and PRDM1-pioneer TF co-bound sites. Enrichment of PRC subunits and PRC-associated histone marks was at basal levels at pioneer TF-only sites, increased at PRDM1-only sites, and exhibited pronounced enrichment at PRDM1-pioneer TF co-bound sites (Figure 3F). Interestingly, enrichment of PRC1 subunits, but not the PRC2 subunit SUZ12, was further increased at PRDM1-pioneer TF co-bound sites relative to PRDM1-only sites, suggesting that pioneer TFs potentiate PRDM1-mediated recruitment of PRC1.

To further investigate this cooperation mechanism, we used a Dox-inducible CRISPRi cell line targeting RING1A/B (RING1A/B-KD)13, the core components of the PRC1 complex, and assessed the impact of PRC1 inhibition on PRDM1-pioneer TF interactions in DE cells (Figure 3G). Co-immunoprecipitation experiments revealed that PRDM1 interacts with GATA6, EOMES, and SOX17 in control DE cells, and that these interactions were reduced by 40–60% upon RING1A/B knockdown (Figure 3H). Together, these results indicate that PRDM1 and pioneer TFs establish functional interactions on chromatin that are at least partially mediated by PRC1-dependent mechanisms.

A high-throughput functional assay identifies PRDM1-mediated repression within combinatorial enhancer motif ensembles

To further characterize the functional interactions between PRDM1 and pioneer TFs on enhancers in DE cells, we performed a lentivirus-based massively parallel reporter assay (lentiMPRA). This approach allowed us to test the regulatory activity of thousands of sequences simultaneously and capture direct relationships between combinatorial TF activity and transcription. To do so, we used a TF motif perturbation approach, where we mutated one or several motif instances within a regulatory sequence to estimate their effects on transcription.

To select our candidate cis-regulatory sequences (CRSs) and perturbation sites, we developed a computational tool called MPRA-ChIP that leverages ChIP-seq peak intersections between two or more TFs followed by peak-associated motif identification, with the aim of identifying functional motifs in a given set of regulatory regions. We started with all 3,484 PRDM1 ChIP-seq peaks identified in DE cells as candidate cis-regulatory sequences (CRSs) and intersected them with GATA6, SOX17, and EOMES ChIP-seq peaks at DE within a 220-bp window (Figure 4A, step 1–2). We then used FIMO to identify PRDM1 and pioneer TF motif occurrences at overlapping peaks and retained those meeting moderate (p ≤ 1e-03) and high (p ≤ 1e-04) stringency thresholds. Of all PRDM1 binding sites at DE, 1,698 regions (48.7%) contained a single PRDM1 peak-associated motif, whereas the remaining regions harbored PRDM1 together with GATA6, SOX17, and/or EOMES peak-associated motifs. To characterize this combinatorial motif organization, we classified PRDM1-bound regions according to sequence complexity, defined by the number of peak-associated motifs present (Figure 4A, step 2). Thus, regions were grouped into sites containing PRDM1 alone or PRDM1 together with one, two, or up to three unique pioneer TF motifs, reflecting endogenous combinations of PRDM1 with GATA6, SOX17, and/or EOMES binding sites.

Figure 4. A high-throughput functional assay identifies PRDM1-mediated repression within combinatorial pioneer TF motif ensembles.

Figure 4.

(a) Schematic representation of lentiMPRA experimental design. Candidate regulatory sequences (CRSs) were based on all 3,484 PRDM1 ChIP-seq peaks at definitive endoderm (DE). For mutagenesis site selection, CRSs were intersected with GATA6, SOX17 and EOMES ChIP-seq peaks at DE within a 220-bp window and FIMO motif scanning was used to identify PRDM1 and pTF motif occurrences at overlapping peaks. The library included a set of wild-type (WT) and perturbed (PERT) sequences, corresponding to wild-type sequences where PRDM1 and/or pTF motifs were mutated in every possible combination. In addition, a set of scramble (SCRAM) sequences, were all nucleotides in each wild-type sequences were shuffled, were used as negative controls. The selected oligonucleotides along with 15-bp barcodes were synthesized and cloned upstream of an SV40 promoter and EGFP reporter into a lentiviral vector and infected into hPSCs. The infected cells were differentiated to DE upon which DNA and RNA were simultaneously extracted and submitted for sequencing. MPRAFlow was used for DNA and RNA barcode association and quantification. MPRAnalyze was used to estimate of transcriptional activity induced by each sequence.

(b) Box plots of CRS activity across perturbation groups (pTF-PERT and PRDM1-PERT) compared to WT sequences plotted as a function of mutagenesis group. Statistical significance was assessed using a likelihood ratio test (LRT). P-values were computed for each comparison and corrected within each analysis using the Benjamini-Hochberg FDR correction (FDR ≤ 0.05).

(c) Box plots of CRS activity across perturbation groups (pTF-PERT and PRDM1-PERT) compared to WT sequences plotted as a function of the total number of mutated motifs. Statistical significance was assessed using a likelihood ratio test (LRT). P-values were computed for each comparison and corrected within each analysis using the Benjamini-Hochberg FDR correction (FDR ≤ 0.05).

For the lentiMPRA library design, we first performed a random downsampling of the single PRDM1 peak-associated motif group to balance the number of sequences containing single versus multiple peak-associated motifs. Then, we combined this subset of PRDM1-only sites with all sequences containing combinatorial PRDM1 and pioneer TF motifs and used them as wild-type (WT) sequences. To test the effect of each motif instance on CRS activity, we designed a set of perturbed (PERT) sequences, corresponding to wild-type sequences where PRDM1 and/or pioneer TF motifs were mutated in every possible combination, including groups where only the PRDM1 motif was mutated (PRDM1-only-PERT), both the PRDM1 and pioneer TF motifs were mutated (PRDM1+pTF-PERT) and only the pioneer TF motifs were mutated while leaving the PRDM1 motif intact (pTF-only-PERT). In addition to this, we designed a set of scramble sequences (SCRAM) as negative controls, where we shuffled all nucleotides of each wild-type sequence. Overall, we assayed a total of 15,336 CRSs (2,012 WT, 2,012 SCRAM and 11,312 PERT [2,012 PRDM1-only-PERT, 4,650 PRDM1+pTF-PERT and 4,650 pTF-only-PERT]) (Figure 4A, step 3). The selected oligonucleotides were synthesized and cloned upstream of an SV40 promoter and EGFP reporter into a lentiviral vector, as previously described24. To allow for robust evaluation of transcriptional activity while normalizing for integration-site biases, we aimed for each individual CRS to be associated with 100 unique 15-bp random barcodes, leading to a library complexity of ~1.5 M sequences (15,336 CRSs and negative controls × 100 barcodes).

After the synthesis and cloning steps, we transduced hPSCs with the lentiMPRA library pool at a multiplicity of infection (MOI) of 8 and differentiated them towards DE, upon which DNA and RNA were simultaneously extracted and submitted for sequencing (Figure 4A, step 4). The library infections were carried out using four biological replicates. We then used MPRAFlow24 and MPRAnalyze25 to associate barcodes with the cloned sequences, quantify DNA and RNA barcodes, and quantify the transcription rate (α) induced by each sequence (Figure 4A, step 5).

To investigate the individual and combinatorial contributions of PRDM1 and pioneer TFs to the activity of CRSs, we compared the MPRA signal of PRDM1-PERT and pTF-PERT relative to their corresponding WT sequences (Figure 4B). Upon perturbation of a single TF motif (1TF), the median activity of the CRSs increased relative to WT only when the mutated motif was PRDM1; however, it decreased when the GATA6, EOMES, or SOX17 motifs were mutated. These results indicate that PRDM1 imposes a repressive constraint on CRS activity, whereas pioneer TF binding promotes transcriptional activation at these shared regulatory elements. Upon perturbation of more than a single TF motif (2TF, 3TF, 4TF), we observed an overall reduction in median activity compared to WT sequences, regardless of whether the PRDM1 motif was mutated or not. These results suggest that PRDM1-mediated repression may be mitigated in the presence of co-binding pioneer TFs. Alternatively, simultaneous mutation of multiple motifs may broadly disrupt CRS function, limiting the ability to resolve the specific contribution of individual TFs within highly combinatorial motif contexts.

To better characterize the effect of PRDM1 perturbation within combinatorial motif contexts, we aggregated the activity of individual TF mutations within the PRDM1-PERT and pTF-PERT groups and plotted them as a function of the number of mutated motifs (Figure 4C). This analysis revealed that mutation of PRDM1 motifs partially buffered the decline in CRS activity observed as the number of mutated motifs increased. We then quantified whether PRDM1 and pioneer TF perturbations exhibit different sensitivity to motif loss by modeling logFC as a function of the number of mutated motifs and perturbation group using a linear mixed model with sequence identity as a random effect. The model showed that CRS activity decreases significantly with increasing numbers of mutated motifs (p < 0.001), while PRDM1 perturbations induce higher CRS activity than pioneer TF perturbations (p < 0.001). To determine whether this overall group effect was maintained across motif classes, we performed pairwise comparisons within each motif count and observed higher activity for PRDM1-PERT than for pTF-PERT constructs at equivalent motif counts (all p < 0.01), indicating that CRS activity reflects a balance between PRDM1-mediated repression and pioneer TF-dependent activation.

PRDM1 repressive activity increases with pioneer TF motif number and depends on the chromatin environment

To minimize confounding effects arising from the simultaneous disruption of multiple motifs, we restricted subsequent analyses to sequences containing a single perturbed motif. The CRS sequences were classified according to the number of TF motifs present within each CRS into four complexity groups from low to high: PRDM1-only, PRDM1+1pTF, PRDM1+2pTF, and PRDM1+3pTF (Figure 5A). We first characterized these sequences by assessing motif scores and spacing for each TF across all complexity groups. Notably, the motif stringency for PRDM1—but not for the pioneer TF motifs—decreased as motif complexity increased (Figure 5B). This is consistent with previous observations in which reduced motif stringency can facilitate cooperative binding of TFs within enhancers26,27. Under such conditions, PRDM1 is recruited to low-affinity motifs only when adjacent sites are occupied by pioneer TFs, which may promote recruitment by opening chromatin, stabilizing PRDM1-DNA interactions, or facilitating protein-protein interactions. Analysis of motif spacing within the CRSs showed that TF binding sites are arranged within compact regulatory modules across all complexity groups (Figure 5C). The total span of the motif module increased with motif complexity, expanding from approximately ~50 bp in PRDM1+1pTF sequences to ~100 bp in PRDM1+3pTF sequences, within a single-nucleosomal size. In contrast, the average pairwise motif distances remained relatively stable, ranging from approximately 48–58 bp even as additional motifs were incorporated. Together, these results indicate that increasing motif complexity primarily expands the combinatorial composition of the regulatory module while maintaining a constrained spatial organization of TF binding sites.

Figure 5. PRDM1 repressive activity increases linearly with pioneer TF motif number and depends on the chromatin environment.

Figure 5.

(a) Schematic representation of sequence classification into low (PRDM1-only), intermediate (PRDM1+1pTF), high (PRDM1+2pTF) and very high (PRDM1+3pTF) complexity groups.

(b) Box plots of motif score across sequence complexity groups. Statistical significance was assessed using a one-way ANOVA followed by Tukey’s multiple comparisons test. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

(c) Box plots of mean pairwise motif distance and total motif span across sequence complexity groups. Statistical significance was assessed using a one-way ANOVA followed by Tukey’s multiple comparisons test. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

(d) Box plots of CRS activity compared to WT upon mutation of single PRDM1, GATA6, SOX17 or EOMES motifs or SCRAM across sequence complexity groups. Statistical significance was assessed using a one-way ANOVA followed by Tukey’s multiple comparisons test. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

(e) Box plots of CRS activity across perturbation groups (pTF-only-PERT and PRDM1-only-PERT) compared to WT sequences across sequence complexity groups plotted as a function of ChromHMM category in DE cells. Statistical significance was assessed using a one-way ANOVA followed by Tukey’s multiple comparisons test. Significance thresholds: P ≤ 0.05 (*), P ≤ 0.01 (**), P ≤ 0.001 (***); ns, not significant.

We next assessed how perturbation of individual PRDM1 or pioneer TF motifs affects CRS activity across motif complexity groups (Figure 5D). Strikingly, the increase in CRS activity upon mutation of PRDM1 motifs relative to their WT counterparts became progressively stronger as motif complexity increases. In contrast, mutation of pioneer TF motifs produced the opposite effect, leading to progressively larger reductions in CRS activity as motif complexity increases. Altogether, these results suggest that increasing motif complexity enhances the intrinsic activating contribution of pioneer TFs, while simultaneously creating a regulatory context in which PRDM1 can exert maximal repressive activity.

To further dissect the functional consequences of PRDM1 or pioneer TF motif perturbation across complexity groups, we mapped all CRSs to endogenous chromatin states using a ChromHMM dataset generated in definitive endoderm cells (Figure 5E). Overall, mutation of PRDM1 motifs across sequence complexity groups consistently increased CRS activity regardless of chromatin state. In contrast, pioneer TF motif perturbations produced a sequence complexity-dependent decrease in CRS activity across most chromatin states. A notable exception emerged for CRSs mapping to bivalent enhancers, where pioneer TF motif mutations instead resulted in increased CRS activity that mirrored the PRDM1 motif mutation effect. This suggests that pioneer TFs do not exert an intrinsically activating or repressive function; instead, their effect on CRS activity depends on the chromatin environment that emerges from their binding and the subsequent recruitment of additional TFs.

Combinatorial binding of pioneer TF and PRDM1 establishes hyper-bivalent enhancers to safeguard cell-type-specific gene expression programs

Having established that PRDM1 repressive activity increases with pioneer TF motif numbers, we next asked whether this relationship is reflected within the chromatin context (Figure 6A). We first classified ChIP-seq peaks identified in DE cells according to sequence complexity into PRDM1-only or PRDM1+pioneer TF (PRDM1+1pTF, PRDM1+2pTF, and PRDM1+3pTF) groups. Additionally, we included regions bound by all pioneer TFs without PRDM1 (3pTF-only). Then, we compared the enrichment of individual histone marks across these five peak groups in DE cells. This analysis revealed that the enrichment of bivalent enhancer-associated histone modifications (H3K4me1, H3K27me3, H2AK119ub1) progressively increased with the number of co-bound pioneer TFs, culminating in the formation of “hyper-bivalent” enhancers at PRDM1+3pTF sites (Figure 6B). Importantly, bivalency-associated histone marks were not enriched at 3pTF-only sites. Instead, these regions were characterized by a strong enrichment of active enhancer-associated histone modifications (H3K27ac, H3K4me1).

Figure 6. Combinatorial binding of pioneer TFs and PRDM1 establishes hyper-bivalent enhancers to safeguard cell-type-specific gene expression programs.

Figure 6.

(a) Schematic representation of endogenous genomic sequence classification by complexity group in DE cells: PRDM1-only, PRDM1+1pTF, PRDM1+2pTF, PRDM1+3pTF and 3pTF-only sequences.

(b) Heatmap of Histone (H2AK119ub1, H3K27me3, H3K4me1, H3K4me3 and H3K27ac) and TF (PRDM1) ChIP-seq signal centered on PRDM1 peaks across endogenous sequence complexity groups.

(c) Schematic representation of low and high MNase-seq experiments in DE cells. In high MNase digestions, where >70% of chromatin is digested to mononucleosomes, nucleosomes in open chromatin are destroyed and annotated as nucleosome-free regions. In contrast, low MNase digestions, where only about 10% of chromatin is digested to mononucleosomes, preserve mononucleosomal DNA in open chromatin, which represents “fragile” or “accessible” nucleosomes.

(d) Aggregate footprinting and nucleosome occupancy profiles across endogenous sequence complexity groups derived from low and high MNase digestion. Low MNase conditions reveal protected transcription factor footprints, whereas high MNase conditions capture nucleosome positioning and occupancy.

(e) Genome browser tracks showing Histone and TF ChIP-seq signal across endogenous sequence complexity groups. Note the presence of hyper-bivalent enhancers at highly occupied loci co-bound by PRDM1+3pTFs.

Increasing evidence suggests that nucleosomes can promote cooperative TF binding at enhancers, even in the absence of direct TF-TF physical interactions7. Consistent with this model, several studies have shown that clusters of TF binding sites at enhancers are associated with high nucleosome occupancy6,7. We also found that PRDM1 and pioneer TF motifs within PRDM1+3pTF regions were localized within a single-nucleosome size (Fig. 5C). To investigate nucleosome organization across PRDM1-bound regions, we performed low- and high-MNase-seq in DE cells, enabling simultaneous profiling of TF footprints, accessible/fragile nucleosomes (low-MNase), and stable nucleosomes (high-MNase) (Figure 6C–D)28,29. Footprinting analysis using subnucleosomal-sized fragments derived from low-MNase-seq revealed a progressive increase in PRDM1 motif protection as the number of co-bound pioneer TFs increased, suggesting that increased sequence complexity stabilizes PRDM1 binding. De novo motif analysis using HOMER further showed that while high-affinity binding to the common GAAGT core in the PRDM1 motif was maintained across all groups, higher complexity sequences tolerated increasingly degenerate 5’ and 3’ motif extensions. Furthermore, high-MNase-seq analysis showed progressively greater nucleosome depletion around PRDM1-bound sites with increasing pioneer TF number, indicating extensive nucleosome remodeling by combinatorial pioneer TF binding. Representative genes associated with the 3pTF-only, PRDM1-only, and PRDM1+3pTF groups include SMAD2, a mediator of TGFβ/NODAL signaling, CSF1, a regulator of macrophage development and survival, and SIX3, a transcription factor involved in forebrain development, respectively (Figure 6E). Altogether, these results show that “hyper-bivalent” enhancers emerge at regions with high combinatorial pioneer TF occupancy, where extensive nucleosome remodeling stabilizes PRDM1 binding and reinforces bivalent chromatin states.

DISCUSSION

In this study, we addressed a central question of how combinatorial TF binding impacts epigenetic regulation at lineage-specific enhancers. Our results show that the endoderm pioneer TFs GATA6, EOMES, and SOX17 act synergistically to remodel nucleosomes and cooperate with the repressive TF PRDM1 to establish “hyper-bivalent” enhancers that preserve lineage fidelity. These domains are enriched for H3K4me1 and PRC-associated histone modifications and regulate mesoderm and ectoderm genes that are highly sensitive to PRDM1 loss. Together, our findings reveal that combinatorial pioneer TF binding establishes locally accessible chromatin platforms that facilitate recruitment of both activating and repressive epigenetic regulators to reinforce bivalent epigenetic states during human development.

An expanding body of work has established PRDMs as key players in development by promoting cell type-specific gene programs and repressing alternative cell states30. PRDM1 has been extensively studied across species and operates by recruiting chromatin-modifying enzymes and co-repressor complexes to mediate gene silencing18. In humans, PRDM1 has been primarily studied in the context of primordial germ cell specification and plasma cell differentiation. However, its role in human endoderm development and its context-dependent mechanisms remain unexplored. Here, we comprehensively profiled the response of active and repressive epigenetic marks to PRDM1 knockdown at the definitive endoderm stage and identified PRC-associated chromatin states, marked by H2AK119ub1 and H3K27me3, as major drivers of PRDM1-mediated repression. Mechanistically, PRDM1 cooperates with endoderm-enriched pioneer TFs and physically interacts with PRC1 and PRC2 subunits to prevent the activation of alternative lineage programs associated with mesoderm and ectoderm genes. Together, these findings provide insight into the context-dependent mechanisms underlying PRDM1-mediated silencing and identify Polycomb complexes as key mediators of its repressive activity in human endoderm.

A central function of pioneer TFs is to locally open compacted chromatin, allowing secondary TFs and co-regulators to access and establish functional regulatory complexes that activate gene expression31. A remarkable example of pioneer TF cooperativity is the formation of super enhancers, where high densities of TF binding drive the formation of clusters of closely spaced enhancers highly enriched for H3K27ac and co-activators that boost transcription of key cell identity genes4. By integrating genomic and epigenomic profiling with lentiMPRA, we identified a distinct class of stand-alone “hyper-bivalent” enhancers established by pioneer TF cooperativity to reinforce transcriptional silencing. First, hyper-bivalent enhancers exhibit extensive local nucleosome remodeling around PRDM1-bound sites. Single-molecule footprinting studies have established that TF binding and nucleosome occupancy are tightly coupled, with increased TF binding directly corresponding to decreased nucleosome occupancy7. Here, we show that this relationship is preserved at bivalent enhancers, indicating that PRCs efficiently nucleate at locally open chromatin rather than nucleosome-dense domains. Second, hyper-bivalent enhancers exhibit a suboptimal motif configuration for PRDM1. Clusters of suboptimal motifs are now understood to increase the cell-type specificity of enhancers in embryonic development27,32. Our data show that PRDM1 motifs become progressively more degenerate at sites with higher pioneer TF co-occupancy, indicating that low-affinity binding sites not only provide specificity for transcriptional activation but also for silencing of developmental enhancers. Third, hyper-bivalent enhancers provide specificity to alternative-lineage gene silencing. Despite millions of potential recognition motifs in the genome, only a small fraction is occupied in vivo, indicating that combinatorial TF interactions – not motif affinity – govern regulatory specificity9. By assessing the distribution of binding sites across lineage-specific gene programs, we found that hyper-bivalent enhancers are enriched near mesoderm- and ectoderm-specific genes. Furthermore, alternative-lineage genes become progressively de-repressed upon PRDM1 knockdown as the number of co-bound pioneer TFs increased, indicating that combinatorial binding confers both specificity and sensitivity to the silencing of alternative lineage programs.

The concept of pioneer TFs emerged from efforts to understand how new genetic networks are established in naïve or silent chromatin31. Large-scale epigenomic profiling by the Roadmap Epigenomics Consortium across more than 100 human tissues and cell types revealed that inactive chromatin states comprise nearly 70% of the epigenome33. Therefore, a better understanding of how chromatin structure represses gene expression is needed to determine how TFs overcome these barriers during gene network and cell fate changes. Overall, this study demonstrates that epigenetic repression is not a default state of chromatin but is actively established through the combinatorial action of pioneer TFs and repressive regulators to maintain lineage fidelity during human development. Importantly, our findings also suggest that the human silencing machinery operates through a distinct logic compared to Drosophila, where repression relies on solitary TFs, inaccessible chromatin, and minimal motif complexity34.

EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

Culture of human induced pluripotent stem cells

Human induced pluripotent stem cells (hPSCs; line 72.3, male, RRID: CVCL_A1BW) were obtained from the CCHMC Pluripotent Stem Cell Facility (PSCF)35 and cultured under feeder-free conditions in mTeSR1 (85850, STEMCELL Technologies) on plates coated with hESC-qualified Matrigel (354277, Corning) or Cultrex Stem Cell Qualified Reduced Growth Factor Basement Membrane Extract (3434–010-02, Bio-Techne). All hPSCs were maintained in a humidified incubator at 37 °C with 5% CO2. Spontaneously differentiated cells were manually removed prior to passaging and differentiation. For passaging, hPSC colonies were dissociated into small cell clusters using Gentle Cell Dissociation Reagent (07174, STEMCELL Technologies) and seeded onto Matrigel- or Cultrex-coated plates in mTeSR1.

The 72.3 hPSC line was distributed by the PSCF from a quality-controlled, cryopreserved bank prepared at passage 30. Cells were confirmed to be mycoplasma-free, to harbor a normal G-banded karyotype, and to have verified identity by STR profiling.

METHOD DETAILS

Generation of doxycycline-inducible CRISPRi hPSC lines

Doxycycline (Dox)-inducible CRISPRi host hPSCs were generated as previously described36. The plasmid pAAVS1-NDi-CRISPRi (Gen 1), encoding dCas9-KRAB-P2A-mCherry for targeted knock-in at the AAVS1 locus, was obtained as a gift from Dr. Bruce Conklin (Addgene plasmid #73497). The guide RNA sequence targeting the AAVS1 locus (sgRNA-T2: GGGGCCACTAGGGACAGGAT)37 was subcloned into the pX459M2-HF vector, a modified version of the pX459 v2.0 vector (Addgene plasmid #62988) carrying an optimized gRNA scaffold, to generate the plasmid pX459M2-HF-AAVS138,39.

hPSCs were reverse-transfected using TransIT®-LT1 Transfection Reagent (2304, Mirus) with 2 μg each of pX459M2-HF-AAVS1 and pAAVS1-NDi-CRISPRi (Gen 1). Colonies obtained after 8 days of selection with 100 μg/mL G418 were manually excised, expanded, and genotyped. A clone containing bi-allelic knock-in of a cassette encoding constitutive expression of the third generation modified rtTA protein (Tet-On 3G) and doxycycline-inducible dCas9-KRAB-P2A-mCherry at the AAVS1 locus was identified by Sanger sequencing and used for subsequent experiments.

Dox-inducible CRISPRi host hPSCs were distributed by the PSCF from a quality-controlled, cryopreserved bank prepared at passage 50, transduced with gRNAs (see below), and used in experiments up to passage 80 (CRISPRi-PRDM1) and passage 73 (CRISPRi-RING1A/B). Cells were confirmed to be mycoplasma-free, to harbor a normal G-banded karyotype, and to have verified identity by STR profiling.

Design and cloning of CRISPRi gRNAs into expression vectors

gRNAs targeting the PRDM1 locus (for CRISPRi-PRDM1) and the RING1 and RING1B/RNF2 loci (for CRISPRi-RING1A/B) were designed using a machine-learning-based CRISPRi sgRNA prediction algorithm40 or CRISPOR41. Two to four gRNAs per CRISPRi target, each driven by independent RNA polymerase III promoters, were cloned into a single lentiviral expression vector using a two-step Golden Gate cloning strategy, as previously described42.

Briefly, in step 1, annealed and phosphorylated oligonucleotides corresponding to each target genomic sequence were cloned into the pmU6-gRNA (Addgene plasmid #53187), phU6-gRNA (Addgene plasmid #53188), phH1-gRNA (Addgene plasmid #53186), or ph7SK-gRNA (Addgene plasmid #53189) vectors via Golden Gate assembly. In step 2, the resulting plasmids from step 1 were cloned by Golden Gate assembly into the pLV-GG-hUbC-EGFP vector (Addgene plasmid #84034), in which dsRED had been replaced with EGFP from Addgene plasmid #1168.

Viral production, transduction, and establishment of CRISPRi-PRDM1- and CRISPRi-RING1A/B monoclonal lines

Lentivirus was produced either by the CCHMC Viral Vector Core or in-house. Briefly, the gRNA-expression lentiviral vector, packaging plasmid (psPAX2), and envelope plasmid (pMD2.G) were transfected into HEK293T cells using the polycation polyethyleneimine (PEI) transfection method. Viral supernatants were passed through a 0.45-μm filter and concentrated either 1:66 using Lenti-X Concentrator (631231, Takara Bio) in-house or 1:255 by ultracentrifugation at the core facility.

Each lentivirus was transduced into CRISPRi hPSCs by spin infection43. CRISPRi hPSCs were dissociated into single cells using Accutase and resuspended in transduction medium consisting of 1–15 μL of lentivirus, 4–8 μg/mL polybrene (STR-1003-G, Sigma-Aldrich), and 10 μM Y-27632 (72304, STEMCELL Technologies) in mTeSR1. Cells were centrifuged at 3,200 × g for 30–90 min. The resulting cell pellet was resuspended in mTeSR1 supplemented with 10 μM Y-27632 and plated onto Matrigel-coated plates. The following day, the medium was replaced with fresh mTeSR1, and cells were maintained with daily medium changes until reaching sub-confluence.

EGFP-positive lentivirus-transduced cells were sorted using a MoFlo XDP (Beckman Coulter) or BD FACSAria II (BD Biosciences) cell sorter and cloned from single cells by limiting dilution, as previously described44. Briefly, clonal cells were cultured in mTeSR1 supplemented with 1x CloneR (05888, STEMCELL Technologies), plasmocin (1:500; ant-mpp, InvivoGen), and penicillin-streptomycin (1:100; 15140122, Thermo Fisher Scientific) for 2 days. On day 3, clonal cells were cultured in mTeSR1 supplemented with 1x CloneR. On day 4, an additional 25% volume of 1x CloneR-supplemented mTeSR1 was added. From day 5 onward, cells were maintained in mTeSR1 with daily medium changes.

Following expansion, a minimum of six monoclonal lines were tested for CRISPRi knockdown efficiency by RT-qPCR. To induce CRISPRi-mediated knockdown of target genes, cells were cultured in the presence of 500 nM to 2 μM doxycycline (D9891–5G, Sigma-Aldrich).

hPSC to endoderm differentiation

For hPSC differentiation, hPSCs were dissociated into single cells using Accutase (07922, STEMCELL Technologies; A6964–500ML, Sigma-Aldrich). Dissociated cells were plated onto Matrigel- or Cultrex-coated culture plates and cultured in mTeSR1 supplemented with 10 μM Y-27632 (72304, STEMCELL Technologies) one day prior to differentiation. Chemically defined basal media 2 (CDM2) and 3 (CDM3) were prepared as previously reported17, filtered through a 0.22-μm filter unit, and stored at −80 °C until use. Once thawed, CDM2 was used within 2 weeks and CDM3 within 3 weeks.

Endodermal induction was performed as previously described17,45. For mesendoderm (ME) induction, hPSCs were cultured in CDM2 supplemented with 100 ng/mL Activin A (800–0, Shenandoah; or GFH6, Cell Guidance Systems), 2 μM CHIR99021 (SML1046–25MG, Sigma-Aldrich), and 50 nM PI103 (2930–1, Tocris Bioscience) for 24 h. For definitive endoderm (DE) induction, ME cells were briefly washed with DMEM/F12 and cultured in CDM2 supplemented with 100 ng/mL Activin A and 250 nM LDN193189 (SML0559–5MG, Sigma-Aldrich) for 24 h. For posterior foregut (pFG) induction, DE cells were briefly washed with DMEM/F12 and cultured in CDM3 supplemented with 1 μM A83–01 (SML0788–5MG, Sigma-Aldrich), 2 μM all-trans retinoic acid (ATRA; R2625–100MG, Sigma-Aldrich), 10 ng/mL bFGF (PHG0261, Thermo Fisher Scientific), and 30 ng/mL BMP4 (314-BP-050, R&D Systems) for 24 h. For liver bud progenitor (LBP) induction, pFG cells were cultured in CDM3 supplemented with 10 ng/mL Activin A, 30 ng/mL BMP4, and 1 mM forskolin (F3917–10MG, Sigma-Aldrich) for 72 h. Differentiation media was changed every 24 h.

Western blotting

Between 1 and 2.5 μg of nuclear protein in Laemmli Sample Buffer (J60015, Alfa Aesar) was denatured at 98 °C for 5 min, resolved on 10%, 12.5%, or 15% SDS-PAGE gels at 60 V for 30 min followed by 100 V for 2 h, and transferred onto either 0.2- or 0.45-μm pore-size PVDF membranes (for chemiluminescence; LC2002, Thermo Fisher Scientific; IPSN07852, Millipore) or 0.2-μm nitrocellulose membranes (for fluorescence; 1620112, Bio-Rad) using a wet tank transfer system with a Mini Blot Module (NW2000, Thermo Fisher Scientific) at 20 V (chemiluminescence) or 10 V (fluorescence) for 1 h.

Transferred membranes were briefly washed with TBST (0.1% Tween-20 in TBS, for chemiluminescence) or TTBS (0.05% Tween-20 in TBS, for fluorescence) and blocked for 1 h at room temperature in blocking buffer consisting of either 5% Blotting-Grade Blocker (1706404, Bio-Rad) or 5% BSA (A9647–50G, Sigma-Aldrich) prepared in TBST (chemiluminescence) or TBS (fluorescence). Membranes were then washed with TBST or TTBS and incubated with primary antibodies diluted in blocking buffer at 4 °C overnight.

The following day, membranes were washed three to five times with TBST or TTBS for 5 min at room temperature and incubated with secondary antibodies diluted in 1% BSA in TBST (chemiluminescence) or 5% Blotting-Grade Blocker in TBS (fluorescence) for 1 h at room temperature. Membranes were washed three to five additional times with TBST or TTBS for 5 min at room temperature. Chemiluminescent signals were developed using ECL Select Western Blotting Detection Reagent (RPN2235, GE Healthcare) or Clarity™ Western ECL Substrate (170–5061, Bio-Rad) for 5 min. Chemiluminescence and fluorescence images were acquired using a ChemiDoc™ MP Imaging System (12003153, Bio-Rad).

Co-immunoprecipitation

Co-immunoprecipitation (co-IP) was performed with modifications to previously published protocols11,46. Cultured cells were briefly washed with DMEM/F12 and crosslinked with 1.5 mM DSP (22585, Thermo Fisher Scientific) in 1x PBS at room temperature for 30 min with gentle shaking. The crosslinking solution was aspirated, and the reaction was quenched with 30 mM Tris-HCl (pH 7.4) in 1x PBS for 20 min with gentle shaking. Crosslinked cells were washed twice with ice-cold 1x PBS and scraped from the plate in ice-cold 0.01% PVA-supplemented 1x PBS. The cell suspension was transferred to conical tubes and centrifuged at 350 ×g for 3–5 min at 4 °C. The supernatant was discarded, and the cell pellet was resuspended in ice-cold 0.01% PVA-supplemented 1x PBS and centrifuged at 800 ×g for 3–5 min at 4 °C. The supernatant was discarded, and the cell pellet was snap-frozen on dry ice and stored at −80 °C until further processing.

For nuclear protein extraction, frozen cell pellets were resuspended in Cell Lysis Buffer (5 mM HEPES-NaOH [pH 7.9], 10 mM KCl, 1 mM DTT, 0.5% NP-40, 1x protease inhibitor) and incubated on ice for 10 min. The lysate was centrifuged at 1700 ×g for 10 min at 4 °C, and the supernatant was discarded. The nuclear pellet was resuspended in Nuclear Lysis Buffer (100 mM NaCl, 25 mM HEPES-NaOH [pH 7.9], 1 mM MgCl2, 0.2 mM EDTA, 0.5% NP-40, 600 U/mL Benzonase [70664, Millipore], 1x protease inhibitor) and incubated for 4 h at 4 °C with rotation. The NaCl concentration of the nuclear lysate was adjusted to 200 mM and incubated for an additional 30 min at 4 °C with rotation. The nuclear lysate was centrifuged at maximum speed (16000 ×g) for 30 min at 4 °C. A small volume of the precleared lysate was reserved as input, and the remaining lysate was incubated with the appropriate amount of primary antibody (listed below) overnight at 4 °C with rotation.

For each immunoprecipitation reaction, 50 μL Dynabeads Protein G (10004D, Thermo Fisher Scientific) were washed twice with PBST (0.1% Tween-20 in PBS) and incubated with the antibody-lysate mixture for 2 h at 4 °C with rotation. Dynabeads were washed five times with IP Wash Buffer (100 mM NaCl, 25 mM HEPES-NaOH [pH 7.9], 1 mM MgCl2, 0.2 mM EDTA, 0.5% NP-40, 1x protease inhibitor), with each wash performed for 2 min at room temperature with rotation.

Dynabead-bound protein complexes (IP samples) were eluted in 1x Laemmli Sample Buffer (1610737EDU, Bio-Rad) for 20 min at 65 °C with shaking at 1000 rpm in a ThermoMixer F1.5 (EP5384000020, Eppendorf). For denaturing IP and input samples, 1/20 volume of 2-mercaptoethanol was added, followed by boiling for 5 min at 98 °C. Denatured samples were stored at −20 °C until analysis by Western blotting.

Chromatin immunoprecipitation

Cells were crosslinked with 1% formaldehyde (F79–500, Fisher Scientific) in 1x PBS for 10 min at room temperature. The crosslinking reaction was then quenched by adding 0.125 M glycine (15527013, Thermo Fisher Scientific) for 5 min at room temperature. Crosslinked cells were centrifuged at 600 ×g for 5 min at 4 °C, and the supernatant was discarded. The cell pellet was washed twice with ice-cold 1x PBS, centrifuged at 600 ×g for 5 min at 4 °C, snap-frozen on dry ice, and stored at −80 °C.

Crosslinked human cells were resuspended in Lysis Buffer 1 (10 mM Tris-HCl [pH 8.0], 10 mM NaCl, 0.5% NP-40, and 1x Complete Protease Inhibitor, EDTA-free) and incubated on ice for 10 min. The cells were then centrifuged at 665 ×g for 5 min at 4 °C, and the supernatant was discarded. The cell pellet was resuspended in Lysis Buffer 2 (50 mM Tris-HCl [pH 8.0], 10 mM EDTA, 0.32% SDS, and 1x Complete Protease Inhibitor, EDTA-free) and incubated on ice for 10 min. The lysate was diluted with IP Dilution Buffer (20 mM Tris-HCl [pH 8.0], 150 mM NaCl, 2 mM EDTA, 1% Triton X-100, and 1x Complete Protease Inhibitor, EDTA-free) and transferred to a milliTUBE ATA Fiber (520135, Covaris). Chromatin was sonicated using a Covaris S220 (500217, Covaris) for 3.5 min (TF ChIP) or 5.5 min (histone ChIP).

Insoluble debris was cleared by centrifugation at 12000 ×g for 5 min at 4 °C. The sonicated chromatin (supernatant) was transferred to a new tube and stored at −80 °C. For immunoprecipitation, 25–100 μL Dynabeads Protein G beads (10003D or 10004D, Thermo Fisher Scientific) were washed twice with PBS-T (0.02% Tween-20 in 1x PBS). The desired amounts of antibodies were conjugated to the washed beads for 2–6 h at 4 °C with rotation.

In parallel, frozen chromatin was thawed on ice, and the appropriate amount of chromatin (determined by DNA content) was diluted with IP Dilution Buffer at a 1:4 ratio. Antibody-conjugated Dynabeads were washed twice with IP Dilution Buffer, resuspended in the diluted chromatin, and incubated overnight at 4 °C with rotation. The following day, the antibody-chromatin solution was cleared and the beads were washed four times with the following buffers: FA Lysis Buffer (50 mM HEPES-KOH [pH 7.5], 150 mM NaCl, 2 mM EDTA, 1% Triton X-100, 0.1% sodium deoxycholate, and 1x Complete Protease Inhibitor, EDTA-free); NaCl Buffer (50 mM HEPES-KOH [pH 7.5], 500 mM NaCl, 2 mM EDTA, 1% Triton X-100, and 0.1% sodium deoxycholate); LiCl Buffer (100 mM Tris-HCl [pH 8.0], 500 mM LiCl, 1% NP-40, and 1% sodium deoxycholate); and 10 mM Tris-HCl (pH 8.0).

Chromatin was eluted with TES buffer (50 mM Tris-HCl [pH 8.0], 10 mM EDTA, and 1% SDS) for 45 min (15 min × 3) at 65 °C with shaking at 1000 rpm in a ThermoMixer F1.5 (EP5384000020, Eppendorf). ChIP and input samples were reverse crosslinked by adding NaCl to a final concentration of 200 mM and incubating for 2–8 h at 65 °C. Samples were treated with RNase A (50 ng/μL; EN0531, Thermo Fisher Scientific) for 0.5–1 h at 37 °C, followed by proteinase K treatment (0.2 mg/mL; 03115828001, Roche) for 2–4 h at 37 °C. DNA was purified by phenol-chloroform extraction followed by ethanol precipitation, and DNA concentration was measured using a Quantus fluorometer (E6150, Promega).

MNase assay

MNase-treated DNA samples were prepared following previously published protocols28,29. Briefly, definitive endoderm stage cells were dissociated and washed twice with Wash Buffer (20 mM HEPES-NaOH [pH 7.5], 150 mM NaCl, 0.5 M spermidine, and 1x Complete Protease Inhibitor, EDTA-free), centrifuged at 600 ×g for 3 min. The cell pellet was resuspended in Digitonin Buffer (0.02% digitonin in Wash Buffer). The cell suspension was then divided into aliquots of 200,000 cells per reaction and incubated with 2 U/mL (low) or 80 U/mL (high) MNase (LS004798, Worthington) for exactly 5 min at 37 °C with shaking at 300 rpm in a ThermoMixer F1.5 (EP5384000020, Eppendorf).

The MNase reaction was quenched by adding equal volumes of 2x Stop Buffer (340 mM NaCl, 20 mM EDTA, 4 mM EGTA, 0.02% digitonin, 0.05 mg/mL RNase A, and 0.05 mg/mL glycogen). For RNase treatment, samples were incubated with RNase A (EN0531, Thermo Fisher Scientific) at 37 °C for 30 min. Subsequently, samples were treated with 0.5% SDS and 0.2 mg/mL proteinase K for 2 h at 37 °C.

Finally, DNA was purified by phenol-chloroform extraction followed by ethanol precipitation. The DNA concentration was measured using a NanoDrop spectrophotometer (13–400-518, Thermo Fisher Scientific). Mono- and sub-nucleosomal DNA fragments were selected and purified using MagBind Magnetic Beads (M1378–01, Omega). The purified DNA concentration was measured using a Quantus fluorometer (E6150, Promega).

Library preparation and sequencing for ChIP-seq and MNase-seq

We prepared multiplexed libraries of ChIP and MNase-treated samples from two replicates (corresponding to two independent CRISPRi clones) using the NEBNext Ultra II DNA Library Prep Kit for Illumina (E7645, NEB). We followed the manufacturer’s protocol with minor modifications: after adaptor ligation, we performed library cleanup using AMPure XP Beads (A63880, Beckman Coulter) or MagBind Magnetic Beads (M1378–01, Omega) without size selection. After PCR enrichment of the adaptor-ligated libraries, two rounds of size selection were performed using magnetic beads.

Bulk RNA sequencing

One million cultured cells were dissociated into single cells using Accutase and mixed with 100,000 Drosophila Schneider 2 (S2) cells as a 10% spike-in control. Cells were pelleted and lysed using Aurum Total RNA Lysis Solution (7326802, Bio-Rad). Total RNA was extracted using the Aurum Total RNA Mini Kit (7326820, Bio-Rad) and submitted to Novogene for library preparation and paired-end 150-bp sequencing on an Illumina NovaSeq 6000 platform.

Bulk RNA-seq data analysis

Sequencing reads were aligned to the hg38 and dm6 combined genome using the STAR aligner47. Only uniquely and concordantly aligned read pairs were used for downstream analysis. Gene-level read counts were quantified using featureCounts48 with the options ‘-O --fracOverlap 0.8’. For spike-in-controlled analysis, reads mapped to Drosophila genes were quantified using the same approach. Differential gene expression analysis was performed using the Wald test implemented in DESeq249. Spike-in normalization was incorporated by calculating size factors using the ‘estimateSizeFactors’ function in DESeq2 based on Drosophila gene read counts. Human genes with fold change (FC) > 1.5 and false discovery rate (FDR) < 0.05 were defined as differentially expressed genes for downstream analysis. Gene ontology (GO) analysis was performed using Enrichr50, and the top significantly enriched GO terms were selected based on Benjamini-Hochberg-adjusted P values (derived from Fisher’s exact test) for presentation.

ChromHMM data analysis

ChromHMM tracks from hPSC-derived endoderm cells were obtained from the Roadmap Epigenomics Mapping Consortium (https://egg2.wustl.edu/roadmap/web_portal, epigenome ID E011, hg38 lift-over) and simplified from an expanded 18-state into an 8-state model by collapsing related chromatin states with similar functional annotations as follows: TssA (TssA, TssAFlnk, TssFlnkU and TssFlnkD); GeneB (Tx and TxWk); EnhA (EnhA1 and EnhA2); EnhWk (EnhG1, EnhG2 and EnhWk); Het (ZNF/Rpts and Het); TssBiv (TssBiv); EnhBiv (EnhBiv, RepPC and RepPCWk) and Quies (Quies).

lentiMPRA library cloning and sequence-barcode association

The lentiMPRA plasmid library was constructed as previously described24, with minor modifications. Briefly, the array-synthesized oligonucleotide pool was first amplified by a low-cycle (5-cycle) PCR using primers (5BC-AG-f01 and 5BC-AG-r01_SV40) designed to append an SV40 promoter and spacer sequence downstream of each candidate regulatory sequence. PCR products were purified using 1.8x MagBind Magnetic Beads (M1378–01, Omega) and subjected to a second amplification step consisting of 12 PCR cycles with primers (5BC-AG-f02 and 5BC-AG-r02) that introduce a 15-nucleotide random barcode sequence. The barcoded amplicons were cloned into the SbfI/AgeI-digested pLS-SceI backbone (Addgene plasmid #137725) using NEBuilder HiFi DNA Assembly Mix (E2621S, NEB). The assembled library was electroporated into 10-beta electrocompetent E. coli (C3020, NEB) using a MicroPulser Electroporator (1652100, Bio-Rad), and transformants were grown overnight on carbenicillin-containing agar plates. Plasmid DNA was isolated by midiprep (D4201, Zymo Research) from an estimated ~1 million independent colonies, corresponding to an average representation of approximately 100 unique barcodes per CRS.

To map barcode-sequence associations, the sequence-SV40-barcode region was PCR-amplified from the pooled plasmid library using primers containing Illumina flow-cell adapters (P7-pLSmP-ass-gfp and P5-pLSmP-ass-i#). The resulting amplicons were sequenced on a NextSeq X Plus platform using 150-bp paired-end reads with custom sequencing primers (R1: pLSmP-ass-seq-R1; index read: pLSmP-ass-seq-ind1; R2: pLSmP-ass-seq-R2_SV40) to obtain ~150 million total reads.

Lentiviral infection and barcode sequencing

Lentivirus was produced as previously described24. Briefly, 10–12 million HEK293T cells were seeded in a single T225 flask. Two days later, plasmid DNA (10 μg plasmid library, 6.5 μg psPAX2, and 3.5 μg pMD2.G) was diluted in 800 μL Opti-MEM (31985062, Thermo Fisher Scientific) to generate premix A. In parallel, 60 μL EndoFectin (EF001, GeneCopoeia) was diluted in 800 μL Opti-MEM to generate premix B. Premixes A and B were combined, incubated at room temperature for 15 min, and then added dropwise to the cells. Cells were incubated for 8–14 h, after which the media was replaced with 30 mL per flask of DMEM supplemented with 5% heat-inactivated FBS and 1x ViralBoost reagent (VB100, Alstem; 60 μL per 30 mL medium). Cells were incubated for an additional 48 h, after which lentiviral supernatant was collected, filtered through a 0.45-μm PES filter system (165–0045, Thermo Scientific), and concentrated using Lenti-X Concentrator (631232, Takara Bio).

Titration of the lentiMPRA library was conducted on 72.3 hPSC lines as previously described24, with minor modifications. Briefly, hPSC cells were plated at 1 × 106 cells per well in 6-well plates and incubated for 24 h. The following day, cells were transduced with serial volumes of lentivirus (0, 4, 8, 16, 32, or 64 μL) in medium supplemented with 8 μg/mL polybrene and spin-infected by centrifugation at 930 × g for 90 min at 32 °C. Transduction media was replaced with fresh medium after 4–8 h and cells were cultured for 3 days prior to genomic DNA extraction. Genomic DNA was extracted using the Wizard SV genomic DNA purification kit (A2361, Promega). The multiplicity of infection (MOI) was measured as the relative abundance of viral DNA (WPRE region, WPRE.F and WPRE.R) over that of genomic DNA [intronic region of LIPC gene, LP34.F and LP34.R] by qPCR using the PowerUp™ SYBR™ Green Master Mix (A25778, Thermo Fisher Scientific), according to the manufacturer’s protocol.

Lentiviral infection, DNA/RNA extraction, and barcode sequencing were all performed as previously described24, with minor modifications. Briefly, ~14 million cells (two 6-well plates) per replicate were infected with the lentiviral library at a MOI of 8 in medium supplemented with 8 μg/mL polybrene (TR-1003-G, Sigma-Aldrich). Three independent biological replicates were generated. One day after transduction, cells were differentiated to mesendoderm (ME), followed by differentiation to definitive endoderm (DE) the next day. On day 3, cells were harvested, and DNA and RNA were purified using TRIzol™ Reagent (15596026, Thermo Fisher Scientific). RNA was treated with TURBO DNA-free™ Kit (AM1907, Thermo Fisher Scientific) to remove contaminating DNA, and reverse-transcribed with SuperScript II (18064022, Thermo Fisher Scientific) using a barcode-specific primer containing a unique molecular identifier (UMI) (P7-pLSmp-assUMI-gfp). Barcode DNA/cDNA from each replicate were amplified with a 3-cycle PCR using specific primers (P7-pLSmp-assUMI-gfp and P5-pLSmP-5bc-i#) to add sample index and UMI. PCR products were then subjected to a second amplification step consisting of 16 cycles using P5 and P7 primers. Amplified fragments were purified and sequenced on an Illumina NovaSeq X Plus platform using 15-bp paired-end reads and 10-cycle dual-index reads with custom sequencing primers (R1, pLSmP-ass-seq-ind1; R2 (index read1 for UMI), pLSmP-UMI-seq; R3, pLSmP-bc-seq; R4 (index read2 for sample index), pLSmP-5bc-seqR2).

QUANTIFICATION AND STATISTICAL ANALYSIS

Mixed-effects modeling of lentiMPRA activity

Analyses were conducted in R using the lme4, lmerTest and emmeans packages. Differences in MPRA activity were assessed using linear mixed-effects models with Anchor included as a random intercept to account for repeated measurements across constructs derived from the same CRS. A linear trend model was fit as logFC ~ motif_count * PRDM1_group + (1 | Anchor), and a categorical model was fit as logFC ~ factor(motif_count) * PRDM1_group + (1 | Anchor) to estimate activity at each motif class. Estimated marginal means and pairwise contrasts were computed using emmeans, with Kenward–Roger degrees of freedom for inference.

ACKNOWLEDGMENTS

We thank the Digestive Disease Research Core Center at CCHMC (Pluripotent Stem Cell Facility, DNA sequencing Core, Bio-Imaging and Analysis Facility). This work was supported by the National Institutes of Health [P30 DK078392 to M.I., R01GM143161 to M.I.]; Cincinnati Children’s Research Foundation [Trustee Award to M.I. and H-W.L., Center for Pediatric Genomics Pilot Award to M.I.]

Footnotes

DECLARATION OF INTEREST

The authors declare that they have no competing interests.

REFERENCES

  • 1.Carey M. (1998). The enhanceosome and transcriptional synergy. Cell 92, 5–8. 10.1016/s0092-8674(00)80893-4. [DOI] [PubMed] [Google Scholar]
  • 2.Yamamoto K.R. (1985). Steroid receptor regulated transcription of specific genes and gene networks. Annu Rev Genet 19, 209–252. 10.1146/annurev.ge.19.120185.001233. [DOI] [PubMed] [Google Scholar]
  • 3.Hnisz D., Abraham B.J., Lee T.I., Lau A., Saint-Andre V., Sigova A.A., Hoke H.A., and Young R.A. (2013). Super-enhancers in the control of cell identity and disease. Cell 155, 934–947. 10.1016/j.cell.2013.09.053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Whyte W.A., Orlando D.A., Hnisz D., Abraham B.J., Lin C.Y., Kagey M.H., Rahl P.B., Lee T.I., and Young R.A. (2013). Master transcription factors and mediator establish super-enhancers at key cell identity genes. Cell 153, 307–319. 10.1016/j.cell.2013.03.035. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jolma A., Yin Y., Nitta K.R., Dave K., Popov A., Taipale M., Enge M., Kivioja T., Morgunova E., and Taipale J. (2015). DNA-dependent formation of transcription factor pairs alters their binding specificity. Nature 527, 384–388. 10.1038/nature15518. [DOI] [PubMed] [Google Scholar]
  • 6.Rao S., Ahmad K., and Ramachandran S. (2021). Cooperative binding between distant transcription factors is a hallmark of active enhancers. Mol Cell 81, 1651–1665 e1654. 10.1016/j.molcel.2021.02.014. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sonmezer C., Kleinendorst R., Imanci D., Barzaghi G., Villacorta L., Schubeler D., Benes V., Molina N., and Krebs A.R. (2021). Molecular Co-occupancy Identifies Transcription Factor Binding Cooperativity In Vivo. Mol Cell 81, 255–267 e256. 10.1016/j.molcel.2020.11.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Xie Z., Sokolov I., Osmala M., Yue X., Bower G., Pett J.P., Chen Y., Wang K., Cavga A.D., Popov A., et al. (2025). DNA-guided transcription factor interactions extend human gene regulatory code. Nature 641, 1329–1338. 10.1038/s41586-025-08844-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Iwafuchi-Doi M., and Zaret K.S. (2014). Pioneer transcription factors in cell reprogramming. Genes Dev 28, 2679–2692. 10.1101/gad.253443.114. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Heslop J.A., Pournasr B., Liu J.T., and Duncan S.A. (2021). GATA6 defines endoderm fate by controlling chromatin accessibility during differentiation of human-induced pluripotent stem cells. Cell Rep 35, 109145. 10.1016/j.celrep.2021.109145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Mukherjee S., Luedeke D.M., McCoy L., Iwafuchi M., and Zorn A.M. (2022). SOX transcription factors direct TCF-independent WNT/beta-catenin responsive transcription to govern cell fate in human pluripotent stem cells. Cell Rep 40, 111247. 10.1016/j.celrep.2022.111247. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Schroder C.M., Zissel L., Mersiowsky S.L., Tekman M., Probst S., Schule K.M., Preissl S., Schilling O., Timmers H.T.M., and Arnold S.J. (2025). EOMES establishes mesoderm and endoderm differentiation potential through SWI/SNF-mediated global enhancer remodeling. Dev Cell 60, 735–748 e735. 10.1016/j.devcel.2024.11.014. [DOI] [PubMed] [Google Scholar]
  • 13.Matsui S., Granitto M., Buckley M., Ludwig K., Koigi S., Shiley J., Zacharias W.J., Mayhew C.N., Lim H.W., and Iwafuchi M. (2024). Pioneer and PRDM transcription factors coordinate bivalent epigenetic states to safeguard cell fate. Mol Cell 84, 476–489 e410. 10.1016/j.molcel.2023.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Baxendale S., Davison C., Muxworthy C., Wolff C., Ingham P.W., and Roy S. (2004). The B-cell maturation factor Blimp-1 specifies vertebrate slow-twitch muscle fiber identity in response to Hedgehog signaling. Nat Genet 36, 88–93. 10.1038/ng1280. [DOI] [PubMed] [Google Scholar]
  • 15.de Souza F.S., Gawantka V., Gomez A.P., Delius H., Ang S.L., and Niehrs C. (1999). The zinc finger gene Xblimp1 controls anterior endomesodermal cell fate in Spemann's organizer. EMBO J 18, 6062–6072. 10.1093/emboj/18.21.6062. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Vincent S.D., Dunn N.R., Sciammas R., Shapiro-Shalef M., Davis M.M., Calame K., Bikoff E.K., and Robertson E.J. (2005). The zinc finger transcriptional repressor Blimp1/Prdm1 is dispensable for early axis formation but is required for specification of primordial germ cells in the mouse. Development 132, 1315–1325. 10.1242/dev.01711. [DOI] [PubMed] [Google Scholar]
  • 17.Loh K.M., Ang L.T., Zhang J., Kumar V., Ang J., Auyeong J.Q., Lee K.L., Choo S.H., Lim C.Y., Nichane M., et al. (2014). Efficient endoderm induction from human pluripotent stem cells by logically directing signals controlling lineage bifurcations. Cell Stem Cell 14, 237–252. 10.1016/j.stem.2013.12.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Bikoff E.K., Morgan M.A., and Robertson E.J. (2009). An expanding job description for Blimp-1/PRDM1. Curr Opin Genet Dev 19, 379–385. 10.1016/j.gde.2009.05.005. [DOI] [PubMed] [Google Scholar]
  • 19.Ancelin K., Lange U.C., Hajkova P., Schneider R., Bannister A.J., Kouzarides T., and Surani M.A. (2006). Blimp1 associates with Prmt5 and directs histone arginine methylation in mouse germ cells. Nat Cell Biol 8, 623–630. 10.1038/ncb1413. [DOI] [PubMed] [Google Scholar]
  • 20.Gyory I., Wu J., Fejer G., Seto E., and Wright K.L. (2004). PRDI-BF1 recruits the histone H3 methyltransferase G9a in transcriptional silencing. Nat Immunol 5, 299–308. 10.1038/ni1046. [DOI] [PubMed] [Google Scholar]
  • 21.Su S.T., Ying H.Y., Chiu Y.K., Lin F.R., Chen M.Y., and Lin K.I. (2009). Involvement of histone demethylase LSD1 in Blimp-1-mediated gene repression during plasma cell differentiation. Mol Cell Biol 29, 1421–1431. 10.1128/MCB.01158-08. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.He H., Bell S.M., Davis A.K., Zhao S., Sridharan A., Na C.L., Guo M., Xu Y., Snowball J., Swarr D.T., et al. (2024). PRDM3/16 regulate chromatin accessibility required for NKX2–1 mediated alveolar epithelial differentiation and function. Nat Commun 15, 8112. 10.1038/s41467-024-52154-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Okashita N., Suwa Y., Nishimura O., Sakashita N., Kadota M., Nagamatsu G., Kawaguchi M., Kashida H., Nakajima A., Tachibana M., and Seki Y. (2016). PRDM14 Drives OCT3/4 Recruitment via Active Demethylation in the Transition from Primed to Naive Pluripotency. Stem Cell Reports 7, 1072–1086. 10.1016/j.stemcr.2016.10.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Gordon M.G., Inoue F., Martin B., Schubach M., Agarwal V., Whalen S., Feng S., Zhao J., Ashuach T., Ziffra R., et al. (2020). lentiMPRA and MPRAflow for high-throughput functional characterization of gene regulatory elements. Nat Protoc 15, 2387–2412. 10.1038/s41596-020-0333-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Ashuach T., Fischer D.S., Kreimer A., Ahituv N., Theis F.J., and Yosef N. (2019). MPRAnalyze: statistical framework for massively parallel reporter assays. Genome Biol 20, 183. 10.1186/s13059-019-1787-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Crocker J., Abe N., Rinaldi L., McGregor A.P., Frankel N., Wang S., Alsawadi A., Valenti P., Plaza S., Payre F., et al. (2015). Low affinity binding site clusters confer hox specificity and regulatory robustness. Cell 160, 191–203. 10.1016/j.cell.2014.11.041. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Farley E.K., Olson K.M., Zhang W., Brandt A.J., Rokhsar D.S., and Levine M.S. (2015). Suboptimization of developmental enhancers. Science 350, 325–328. 10.1126/science.aac6948. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Iwafuchi-Doi M., Donahue G., Kakumanu A., Watts J.A., Mahony S., Pugh B.F., Lee D., Kaestner K.H., and Zaret K.S. (2016). The Pioneer Transcription Factor FoxA Maintains an Accessible Nucleosome Configuration at Enhancers for Tissue-Specific Gene Activation. Mol Cell 62, 79–91. 10.1016/j.molcel.2016.03.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Lim H.W., and Iwafuchi M. (2023). Profiling Accessible Chromatin and Nucleosomes in the Mammalian Genome. Methods Mol Biol 2599, 59–68. 10.1007/978-1-0716-2847-8_6. [DOI] [PubMed] [Google Scholar]
  • 30.Hohenauer T., and Moore A.W. (2012). The Prdm family: expanding roles in stem cells and development. Development 139, 2267–2282. 10.1242/dev.070110. [DOI] [PubMed] [Google Scholar]
  • 31.Zaret K.S. (2020). Pioneer Transcription Factors Initiating Gene Network Changes. Annu Rev Genet 54, 367–385. 10.1146/annurev-genet-030220-015007. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Lim F., Solvason J.J., Ryan G.E., Le S.H., Jindal G.A., Steffen P., Jandu S.K., and Farley E.K. (2024). Affinity-optimizing enhancer variants disrupt development. Nature 626, 151–159. 10.1038/s41586-023-06922-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Roadmap Epigenomics C., Kundaje A., Meuleman W., Ernst J., Bilenky M., Yen A., Heravi-Moussavi A., Kheradpour P., Zhang Z., Wang J., et al. (2015). Integrative analysis of 111 reference human epigenomes. Nature 518, 317–330. 10.1038/nature14248. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Hofbauer L., Pleyer L.M., Reiter F., Schleiffer A., Vlasova A., Serebreni L., Huang A., and Stark A. (2024). A genome-wide screen identifies silencers with distinct chromatin properties and mechanisms of repression. Mol Cell 84, 4503–4521 e4514. 10.1016/j.molcel.2024.10.041. [DOI] [PubMed] [Google Scholar]
  • 35.McCracken K.W., Cata E.M., Crawford C.M., Sinagoga K.L., Schumacher M., Rockich B.E., Tsai Y.H., Mayhew C.N., Spence J.R., Zavros Y., and Wells J.M. (2014). Modelling human development and disease in pluripotent stem-cell-derived gastric organoids. Nature 516, 400–404. 10.1038/nature13863. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Mandegar M.A., Huebsch N., Frolov E.B., Shin E., Truong A., Olvera M.P., Chan A.H., Miyaoka Y., Holmes K., Spencer C.I., et al. (2016). CRISPR Interference Efficiently Induces Specific and Reversible Gene Silencing in Human iPSCs. Cell Stem Cell 18, 541–553. 10.1016/j.stem.2016.01.022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Oceguera-Yanez F., Kim S.I., Matsumoto T., Tan G.W., Xiang L., Hatani T., Kondo T., Ikeya M., Yoshida Y., Inoue H., and Woltjen K. (2016). Engineering the AAVS1 locus for consistent and scalable transgene expression in human iPSCs and their differentiated derivatives. Methods 101, 43–55. 10.1016/j.ymeth.2015.12.012. [DOI] [PubMed] [Google Scholar]
  • 38.Chen B., Gilbert L.A., Cimini B.A., Schnitzbauer J., Zhang W., Li G.W., Park J., Blackburn E.H., Weissman J.S., Qi L.S., and Huang B. (2013). Dynamic imaging of genomic loci in living human cells by an optimized CRISPR/Cas system. Cell 155, 1479–1491. 10.1016/j.cell.2013.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Ran F.A., Hsu P.D., Wright J., Agarwala V., Scott D.A., and Zhang F. (2013). Genome engineering using the CRISPR-Cas9 system. Nat Protoc 8, 2281–2308. 10.1038/nprot.2013.143. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Horlbeck M.A., Gilbert L.A., Villalta J.E., Adamson B., Pak R.A., Chen Y., Fields A.P., Park C.Y., Corn J.E., Kampmann M., and Weissman J.S. (2016). Compact and highly active next-generation libraries for CRISPR-mediated gene repression and activation. Elife 5. 10.7554/eLife.19760. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Concordet J.P., and Haeussler M. (2018). CRISPOR: intuitive guide selection for CRISPR/Cas9 genome editing experiments and screens. Nucleic Acids Res 46, W242–W245. 10.1093/nar/gky354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kabadi A.M., Ousterout D.G., Hilton I.B., and Gersbach C.A. (2014). Multiplex CRISPR/Cas9-based genome engineering from a single lentiviral vector. Nucleic Acids Res 42, e147. 10.1093/nar/gku749. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Kodaka Y., Asakura Y., and Asakura A. (2017). Spin infection enables efficient gene delivery to muscle stem cells. Biotechniques 63, 72–76. 10.2144/000114576. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Matsui S., Shiley J.R., Granitto M., Ludwig K., Buckley M., Koigi S., Mirizio G., Hu Y.C., Mayhew C.N., and Iwafuchi M. (2025). Protocol for establishing inducible CRISPR interference system for multiple-gene silencing in human pluripotent stem cells. STAR Protoc 6, 103736. 10.1016/j.xpro.2025.103736. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Ang L.T., Tan A.K.Y., Autio M.I., Goh S.H., Choo S.H., Lee K.L., Tan J., Pan B., Lee J.J.H., Lum J.J., et al. (2018). A Roadmap for Human Liver Differentiation from Pluripotent Stem Cells. Cell Rep 22, 2190–2205. 10.1016/j.celrep.2018.01.087. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Cattoglio C., Pustova I., Walther N., Ho J.J., Hantsche-Grininger M., Inouye C.J., Hossain M.J., Dailey G.M., Ellenberg J., Darzacq X., et al. (2019). Determining cellular CTCF and cohesin abundances to constrain 3D genome models. Elife 8. 10.7554/eLife.40164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Dobin A., Davis C.A., Schlesinger F., Drenkow J., Zaleski C., Jha S., Batut P., Chaisson M., and Gingeras T.R. (2013). STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Liao Y., Smyth G.K., and Shi W. (2014). featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923–930. 10.1093/bioinformatics/btt656. [DOI] [PubMed] [Google Scholar]
  • 49.Love M.I., Huber W., and Anders S. (2014). Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15, 550. 10.1186/s13059-014-0550-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Kuleshov M.V., Jones M.R., Rouillard A.D., Fernandez N.F., Duan Q., Wang Z., Koplev S., Jenkins S.L., Jagodnik K.M., Lachmann A., et al. (2016). Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res 44, W90–97. 10.1093/nar/gkw377. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from bioRxiv are provided here courtesy of Cold Spring Harbor Laboratory Preprints

RESOURCES