Summary
Quantitative comparison of ChIP-seq profiling between experimental conditions or samples remains technically challenging for the epigenetics field. Here, we report a strategy combining the use of well-defined cellular spike-in ratios of orthologous species’ chromatin and a bioinformatic analysis pipeline to facilitate highly quantitative comparisons of 2D chromatin sequencing across experimental conditions. We find that the PerCell methodology results in efficient and consistent levels of spike-in vs. experimental genomic reads. We demonstrate use of the method and pipeline to enable quantitative, internally normalized chromatin sequencing on zebrafish embryos and human cancer cells. Overall, we propose the PerCell method to enable cross-species comparative epigenomics and promote uniformity of data analyses and sharing across labs.
Keywords: quantitative epigenomics, sarcoma, transcription factor, genomics normalization, cross-species epigenomics, chromatin sequencing, Nextflow
Graphical abstract

Highlights
-
•
Presents a versatile, highly quantitative approach for analyzing protein-genome binding
-
•
Integrates cell-based chromatin spike-in with a flexible bioinformatic pipeline
-
•
Demonstrates quantitative epigenetic comparisons across cell states and models
Motivation
Chromatin immunoprecipitation sequencing (ChIP-seq) has become a foundational methodology for the field of epigenetics and studies of histone modifications and transcription factor localization across the genome. However, it has remained challenging to use this technology to quantitatively compare the amount of signal at a given locus between experimental conditions or samples. For the chromatin community, we introduce a low-cost and flexible method enabling quantitative comparisons across experimental conditions and cellular contexts via our freely released bioinformatic pipeline.
Tallan et al. report a versatile approach for highly quantitative protein-genome binding, demonstrating applications for human and zebrafish cells. The proposed PerCell methodology and analysis strategy are modular and accessible, enabling cross species comparative epigenomics.
Introduction
Chromatin provides an interrelated network for structural and functional regulation of the DNA content in a eukaryotic cell. A shared challenge throughout the study of chromatin is how to quantitatively compare epigenetic feature abundance and localization across the genome, from one experiment to another, or even one lab to another. This is especially important if the experimental question entails deleting epigenetic regulators or causing global alterations in genome structure. To better understand chromatin regulation across the genome, our community would benefit from a low-cost, universal, and accessible process to generate, compare, and share chromatin sequencing data so that we can, together, define and understand the integrated patterns of the epigenome. This has motivated our current project: to provide an outward-facing methodology encompassing both wet lab and dry lab in order to enhance universal data generation, analysis, and sharing across teams.
High-throughput DNA/chromatin sequencing technologies, including ChIP-seq (chromatin immunoprecipitation sequencing),1,2,3,4 have become foundational for the study of epigenetic regulators of chromatin and catalyzed a set of broadly popular methodologies. ChIP-seq can be applied to map histone modifications, transcription factor localization on chromatin, and indeed a myriad of other epigenetic regulatory events. The advantages of ChIP-seq include its reliance on largely non-enzymatic processes that avoid potential alterations of the regions of the epigenome that are sequenced. Enzymatic alternatives may cut preferentially at open chromatin,5,6 and we hypothesize that such a bias can skew results and our view of the epigenome in favor of accessible regions. In the context of methods dependent on endonucleases or transposases for chromatin fragmentation, ChIP-seq remains highly useful even as a significant number of newer methods have been developed, which are conceptually built upon the assay.
Despite widespread use over many years, ChIP-seq as it is commonly performed has key limitations affecting its generalizability, data comparisons across teams of investigators, and indeed even across experiments in the same lab. These limitations have yet to be addressed in a highly quantitative and low-cost manner including both wet lab procedure and semi-automated data analysis. Chief among the limitations is our inability to perform highly quantitative ChIP-seq comparisons across experimental conditions or cell lines. That is, while ChIP-seq can often be helpful in answering qualitative questions, such as where a specific event occurs on chromatin (e.g., histone modification, transcription factor [TF] binding), it is very difficult to draw confident conclusions in situations where a histone modification or TF binding event occurs at the same locus but at varying magnitudes. This information is often critical and varies depending on specific treatment conditions, timing of DNA replication within the cell cycle, or across tumor cell lines of varying ploidy or genome sizes.
Additionally, as we progress toward an era of highly interdisciplinary science, epigenetics is often integrated with preclinical drug treatments (e.g., entinostat,7,8 a clinical histone deacetylase [HDAC] oncology candidate, and A485,9 a P300 preclinical oncology candidate), many of which have potential to alter global histone modification states. When evaluating preclinical molecules that target histone modifications globally, there is a fundamental data processing challenge that arises and needs to be met. Paradoxically, if changes occur everywhere on chromatin, then the changes are challenging to measure through next-generation sequencing compared to more specific, local epigenetic changes.
In attempts to address these fundamental limitations, methods have been introduced that include the addition of internal “spike-in” DNA content, measured by weight. This exogenous DNA can be used to set a baseline across replicates and conditions, thereby enabling comparisons with greater clarity. However, some previous ChIP-seq protocols incorporating spike-in approaches were not designed for multicellular organisms.10 Others used only highly divergent species’ DNA for spike-in,11 fixed amounts of spike-in chromatin,12 or both,13 creating technical hurdles related to inconsistent representation of the spike-in sequencing reads within datasets. While useful for providing quantitative comparisons in certain contexts, these methods also have drawbacks in that they can result in significant limitations in sequencing efficiency and the choice of antibody,11 unnecessarily complicated procedures and the requirement for multiple antibodies,13 or an inability to compare results across distinct genetic backgrounds.12,13 This last point is especially significant given the high frequency of substantial changes in genomic content (aneuploidy) in cancers and the immortalized cell lines often studied using ChIP-seq. Hence, we sought to establish a universal and low-cost strategy for data comparisons with PerCell ChIP-seq, which we describe here.
Our approach takes advantage of highly related “orthologous” genomes that can be resolved through next-generation sequencing even when indexed together within the same experimental sample. We developed the orthologous cell spike-in to be within a framework of closely related genomes (e.g., human, mouse, and zebrafish) and the mixing of experimental and spike-in cells at fixed ratios early within the workflow. The use of cells and not purified DNA for spike-in content enables sensitive quantification and comparison of local and global differences in histone modification and transcription factor abundance/localization across diverse treatment conditions and genomic contexts (Figure 1).
Figure 1.
Rationale for PerCell-normalized semi-automated ChIP-seq analysis
Shown are conceptual representations and comparisons of ChIP results after analyzing cells of varying ploidy (1n, 2n, and 4n) via conventional spike-in or PerCell normalization. The PerCell methodology specifically enables the detection of changes in relative signal across cellular backgrounds with different chromatin content and the use of a single antibody.
PerCell ChIP-seq enables efficient, cost-effective, non-commercial, and reproducible strategies for epigenome sequencing. Our approaches do not rely on manufactured kits, can be performed with reagents commonly found in most molecular biology labs, and thus provide benefits across experience and resource levels for researchers. Our method will generate rigorous maps of protein-genome binding and cis-chromatin interactions across cell states and in vivo. Further, paired with our experimental approach, our work provides an intentionally designed analytical pipeline for spike-in ChIP-seq analyses. While key examples of this have been reported,14,15,16,17 several approaches lack version control or containerization, and others lack a streamlined wet lab-to-dry lab software component with the advantages that ours provides. Our pipeline (Figure 2) is amenable to a range of spike-in setups, although we emphasize mouse-to-human and human-to-zebrafish quantitative epigenomics. Written using the Nextflow18 workflow management software, including customized implementations from the nf-core project,19 our pipeline is highly scalable and reproducible across multiple execution platforms.
Figure 2.
PerCell workflow
(A) Workflow summary for the PerCell spike-in ChIP-seq wet lab procedure. Experimental cells (blue) and spike-in cells (red) of orthogonal species are separately crosslinked with chromatin-associated proteins via formaldehyde fixation before being mixed at known ratios that remain consistent for each sample. Experimental and spike-in cells proceed together throughout the remainder of the protocol. Cells are sonicated to solubilize DNA fragments of appropriate length, incubated with an antibody targeting a protein of interest, immunoprecipitated, and washed. Crosslinking is reversed to yield chromatin enriched for sequences interacting with the protein of interest, and the DNA is prepared for sequencing with identifying indexing as necessary and finally sequenced.
(B) Subway map for the semi-automated PerCell normalization analysis pipeline with cellular spike-in (red), experimental chromatin (blue), or the workflow with no spike-in option (yellow). With a streamlined user interface, our pipeline will automate the central processes for cellular spike-in chromatin normalization per cell, including downsampling for peak calling and generation of bigWig files. In our updated semi-automated pipeline, we include the option to either include or not include spike-in, which creates versatility and applicability for an increased range of researchers and projects.
Results
Overview of the PerCell workflow and bioinformatic pipeline
PerCell ChIP-seq is an optimized, flexible, and reliable methodology for the normalization of ChIP-seq data. The approach generates quantitative comparisons of genome-wide TF and histone modification abundance through the use of highly similar orthologous genomes for both experimental and spike-in cells. While it builds from the rationale behind past normalization strategies,10,11,12,13 we further optimized the method to increase signal-to-noise ratio and additionally enable comparisons of samples across distinct genetic backgrounds. Thus, key features of the PerCell workflow (Figure 2A) include the spike-in of cells rather than calculated amounts of chromatin from a closely related species, the mixing of experimental and spike-in cells at fixed ratios prior to cellular sonication, and the ability to use a single antibody directed against the protein or histone modification of interest. To accompany the optimized PerCell wet lab workflow and facilitate analyzing datasets generated by it, we also developed and freely released a bioinformatic ChIP-seq pipeline (Figure 2B) capable of normalizing based on spike-in content included within sequencing data (see STAR Methods for detailed information on pipeline outputs and software used).
Comparisons of PerCell spike-in efficiency and previous approaches
We took special consideration to benchmark PerCell ChIP-seq against previous approaches for chromatin or DNA spike-in. Specifically, in the context of highly quantitative comparisons between our methodological approach and previous approaches,11,12,13 we uncovered key differences in efficiency. When analyzing the number of reads aligned to the experimental (human) versus spike-in (mouse) genomes, we expect approximately 22.5% of mapped reads aligning to the spike-in genome, given a 3:1 human:mouse cellular ratio and accounting for the human genome being roughly 15% larger than the mouse genome. This is barring any effects on enrichment of one sample over another by the antibody used, and indeed we find generally consistent percentages of 21%–34% of input sample mapped reads aligning to the spike-in genome (Table S1). For immunoprecipitated samples, this range shifted slightly to 16%–25%. This shift is likely due to a higher affinity for the target of interest in experimental (as compared to spike-in) cells and may vary depending on the antibody used and the similarity between genomes. This is in marked contrast to previous approaches,11,12,13 where spike-in percentages ranged widely, and reads aligning to the spike-in genome occasionally even surpassed the number of reads aligning to the experimental genome. The relative percentages of spike-in versus experimental chromatin are important to consider due to overall read depth and sensitivity within individual experiments as well as the potential for limitations on signal-to-noise ratio for sequencing experiments.
While the specifics vary in their choice of spike-in and normalization strategies, the protocols employed by ChIP-Rx,11 SAP,12 and Active-Motif13 methodologies result in spike-in percentages ranging from 4%–65%, 1%–21%, and <1%–25% (Figure S1; Table S3). These considerable variations in spike-in efficiency might necessitate downsampling some samples to a fraction (in some cases <5%) of their total experimental read depth, but all three methods, in fact, use normalization approaches that involve downscaling some samples while upscaling others, in effect inflating read counts or signal beyond the original sequencing data. For comparison, our spike-in strategy involves less variation in experimental vs. spike-in content (Tables S1 and S2), enabling our pipeline to avoid upsampling, which we find to be a more rigorous approach when processing chromatin sequencing datasets.
While our methodology uses downsampling to remove reads proportionally to the number of spike-in reads for each sample, alternative approaches for normalizing ChIP-seq data include adjusting all samples to a predetermined number of reads, such as 1 million in RRPM (reference-adjusted reads per million).11 Normalizing in this way can result in some samples being disproportionately affected by scaling up or down, therefore substantially decreasing (downsampling) or increasing (upsampling) the apparent quality of those samples’ sequencing data. Another approach involves normalizing to the median number of reads within a dataset.17 This helps moderate the disproportionate effects of normalizing to a given value. However, as with any approach that involves upsampling, it runs the risk of introducing additional noise and even artificial data; i.e., reads and peaks that do not correspond to true biological/sequenced data. SpikChIP15,16 software, although difficult to install due to lack of versioning or containerization, introduces an additional computational method for normalizing experimental ChIP-seq data. This statistical method uses a locally estimated scatterplot smoothing (LOESS)-based local regression strategy to dynamically apply scaling factors with relatively higher effects on regions assigned as peaks. While the SpikChIP normalization strategy is a compelling quantitative methodology for some experimental setups, it necessitates the use of spike-in material from a relatively dissimilar genome vs. experimental cells and a secondary antibody confirmed to have no cross-reactivity with the experimental genome, which may provide some disadvantages or limitations in certain experimental setups.
The choice of species identity for spike-in cells (or, in the case of other methods, chromatin) is similarly important. Previous methods have relied on the use of spike-in species distantly related to the species investigated with experimental portions of samples, significantly limiting the choice of immunoprecipitation targets to conserved proteins11 or necessitating the use of a secondary antibody known to have no cross-reactivity across the two species.13 The use of fixed amounts of chromatin,11,12 rather than known ratios of experimental and spike-in cells, similarly limits any quantitative interpretation of binding across distinct cellular backgrounds or cell states (Figure 1, conventional spike-in ChIP). Proper alignment and filtering using our bioinformatic normalization pipeline reveals that even highly similar species (e.g., human and mouse) can be used for spike-in, with only negligible numbers of ambiguous reads (Figure S2). This made us confident that our PerCell approach could reliably use only a single antibody, reducing costs and workflow complexity. Further simplifying the PerCell protocol, the use of more highly related cell lines enables more even and consistent sonication across experimental and spike-in cells, which may have contributed to the high variations in spike-in content with previous methods.11
Detecting global alterations in histone modifications after drug treatments
Our approach allows for the precise quantification of comparisons of global changes in the deposition of histone H3 acetylation after drug treatments in patient-derived rhabdomyosarcoma cells, with clinically promising molecules that either globally increase (entinostat, 0.5 μM) or globally decrease (A485, 0.5 μM) chromatin acetylation (Figure 3). Without PerCell measurements of differential histone H3 Lys-27 acetylation, it becomes challenging to meaningfully interpret genome-wide data on the placement of this mark after either drug treatment (Figures 3A, 3C, and 3E), whereas PerCell normalization enables direct quantitative comparisons (Figures 3B, 3D, and 3F). Specifically, in the context of the TCF3 locus (Figures 3C and 3D) and RPL22 (Figures 3E and 3F), we observe dramatic improvements in our ability to interpret quantifications of H3 acetylation data after PerCell normalization, and there are many other examples as well. When we examine the mapping of experimental and cellular spike-in reads from orthologous chromatin, we observe that there is less than 0.4% overlap between mouse/human mapped reads (Table S2), meaning that the overwhelming majority of experimental or spike-in reads are mapped accurately to the correct genomes.
Figure 3.
Data from human cancer cells (rhabdomyosarcoma Rh4 with C2C12 mouse myoblast cellular spike-in) with anti-H3K27ac PerCell ChIP-seq after DMSO, 0.5 μM A485, or 0.5 μM entinostat treatment
(A and B) Non-normalized (A) and PerCell-normalized (B) H3K27ac ChIP-seq signal from Rh4 cells treated with DMSO, 0.5 μM A485 (a histone acetyltransferase inhibitor), or 0.5 μM entinostat (a histone deacetylase inhibitor). Average signal is shown in arbitrary units (a.u.) across the entire genome, divided into 5-kb bins.
(C–F) Non-normalized (C and E) and PerCell-normalized (D and F) signal of the same samples at representative loci decorated with histone H3 acetylation, including TCF3 (C and D) and RPL22 (E and F), respectively.
Quantifying in vivo changes in histone modifications after induction of an oncogenic TF-chimera
Of note, we report a specific application of our PerCell method to precisely map the genome-wide binding of a TF-chimera, which is an oncogenic driver in a rare childhood sarcoma (STAR Methods; Figure 4). We developed rapid and in vivo expression systems in zebrafish rhabdomyosarcoma models to understand initial fusion oncoprotein epigenetic activities (Figures 4A and 4B; cf., Kucinski et al.20), based on earlier contributions in this area.21 Notably, when we rapidly express the TF-chimera PAX3::FOXO1 in vivo, PerCell normalization of anti-FOXO1 ChIP-seq22,23 reveals quantitative changes in the induction of H3K27ac at key target genes and binding sites (Figures 4C and 4D). This is especially notable given that (1) these binding sites normally lack H3K27ac and (2) enhancer activation is typically driven by pluripotency transcription factors during these early embryonic zebrafish timepoints.24,25 However, as we note in our related manuscript,20 PAX3::FOXO1 regulatory activities also result in a quantifiable redistribution and, therefore, relative loss of H3K27ac at various loci, as detected via PerCell ChIP-seq. Excitingly, our approach enables in vivo genome-wide binding measurements for PAX3::FOXO1 within hours after its expression, revealing chromatin state changes including those shown at the her3 and tfap2b loci (Figures 4E and 4F). Here, the PerCell methodology is even more accurate and efficient in zebrafish/human systems than in human/mouse systems, and we find only approximately ∼0.001% of mapped reads overlapping between zebrafish/human genomes (Table S2; Figure S2). To our knowledge, this is the first reported in vivo genome-wide binding dataset for PAX3::FOXO1 and the initial report of highly quantitative differential comparisons of chromatin modifications states in vivo within hours of PAX3::FOXO1 induction. We anticipate further applications of the PerCell methodology to measure TFs and TF-chimeras in human cancers and their responsiveness to the activities of clinical or preclinical molecules that function by altering chromatin modification states.
Figure 4.
Zebrafish injected with control vs. PAX3::FOXO1 (P3F) mRNA in anti-H3K27ac and anti-FOXO1 (P3F) PerCell ChIP-seq experiments
(A and B) Experimental design workflow, including PAX3::FOXO1 (P3F) and control mRNA constructs (A) and our zebrafish embryo injection strategy (B).
(C and D) Profile plots of non-normalized (left) versus normalized (right) data for control and P3F construct injection at upregulated genes (C) or known P3F genomic binding sites (D).
(E and F) Genome browser views at the her3 (E) and the tfap2b gene loci (F), where P3F expression in zebrafish embryos increases adjacent histone H3K27ac levels, which can be quantified with PerCell ChIP-seq. In each case, the magnitude of P3F binding is also measured with our PerCell workflow.
Determining differences in protein-chromatin interactions due to variations in genome size
We also foresee utility in the context of quantitative examination of protein-genome binding events with altered genome sizes, which are common in human cancers.26 When we synthetically filter the apparent amount of cellular spike-in reads from experimental samples measuring CTCF binding, we are able to use PerCell normalization to obtain meaningful information for relative protein-genome binding with 25%, 50%, 75%, and 100% of spike-in (mouse) reads retained (Figures 5A and 5B). With specific examination of the MYOD1 (Figures 5C and 5D) and FGFR4 (Figures 5E and 5F) loci, our PerCell workflow enables direct comparisons of CTCF binding levels with high resolution. There are key implications here, for contextualizing the loop extrusion model,27 where examination not only of the convergence of CTCF sites bracketing loop domains but also quantifying magnitudes of CTCF-genome binding could illuminate a central logic of chromatin folding. Thus, we illustrate the utility of our approach in three distinct chromatin contexts: for architectural machinery (Figure 5), for the deposition of histone modifications and their drug sensitivities (Figure 3), and for cross-species genome-wide mapping and quantification of TF-chimeras that are penetrant drivers of childhood cancer (Figure 4). Further development of these workflows for defining altered protein-genome binding for TFs, other architectural regulators (e.g., cohesin), a diversity of histone marks, and other epigenetic contexts will be of high interest for researchers in the areas of pediatric oncology, epigenetics, and genomics.
Figure 5.
Quantitative PerCell differences in relative genome sizes
(A and B) Non-normalized (A) and PerCell-normalized (B) CTCF ChIP signal in Rh4 human rhabdomyosarcoma cancer cells with varying percentages of cellular spike-in reads removed (resulting in 25%, 50%, 75%, or 100% of reads being retained) before application of the PerCell pipeline. Signal is shown ±250 bp of consensus CTCF called peaks.
(C–F) Non-normalized (C and E) and PerCell-normalized (D and F) signal at the MYOD1 (C and D) or FGFR4 (E and F) locus.
Discussion
PerCell ChIP-seq assays and their analysis using our pipeline are expected to generate qualitative readouts of protein:DNA interactions or histone modification localization similar to conventional ChIP-seq workflows while additionally providing relative quantifications for these data across distinct replicates, lab groups, cellular treatment conditions, and genetic backgrounds (Figure 1). We expect our spike-in based normalization method and pipeline to enable the user to identify global changes in transcription factor or histone modification abundance throughout the genome while also accounting for differences in the chromatin states being investigated (e.g., due to aneuploidy). Additionally, the flexible choice of spike-in cells in a closely related species decreases barriers associated with requiring multiple antibodies or specifically prepared chromatin/cell lines.
There is a diversity of applications for our methodology. These include, but are not limited to, comparisons of drug treatments, genetic deletions, and developmental changes in vivo. With our workflows, comparisons are possible across treatment conditions and inhibition or functional disruption of global chromatin modifiers including targeted protein degradation, CRISPR-Cas9 knockout, and knockdowns. Additionally, our system has utility in human cancer cells with variable genomic content and in a cell cycle context where the amount of chromatin per cell is no longer fixed while the cell ratios remain constant (i.e., in S phase of the cell cycle).
Our analytical workflow is described in Figure 2, with a detailed overview of timing required per step in Table S4 (STAR Methods). The required files for running our bioinformatic pipeline are fastq files, either gzip compressed or uncompressed. These files should contain sequencing reads for ChIP-seq or similar chromatin-based assays. While designed for handling mixed experimental and spike-in sequencing reads, the pipeline can also be run to skip all portions involving spike-in reads (e.g., for comparing spike-in normalization effects on analyses) and is thus suitable for analyzing data without spike-in included.
Limitations of the study
While our pipeline was created with PerCell ChIP-seq experiments in mind, its functionality is not limited to this protocol. The same normalization approach and, hence, the same analytical processes, may be used for other assays involving spike-in chromatin, including the spike-in of fixed quantities of purified exogenous DNA as part of a standard ChIP-seq, Cut&Run,5 or Cut&Tag6 assay. We anticipate that our PerCell experimental and semi-automated analysis workflows will be widely used to increase reproducibility across laboratories and across experimental conditions in cellular models and in vivo.
Our methodology requires relative consistency in terms of calculating cell number within the same set of experiments, while it does not per se depend on absolute cell counts. Regardless, consistency across replicates is crucial. That is, using consistent relative numbers of counted cells in each compared sample is important. With this in mind, we highly recommend large-scale aliquoting of pre-counted cells for PerCell workflows and particularly those used for orthologous genome (spike-in) chromatin. This is also important to minimize experimental variation that cannot be accounted for through normalization. We agree with a previous report15 that, in situations where determining cell number is not possible, alternative spike-in methods involving fixed amounts of chromatin12,13 may be considered as alternatives.
While different affinities for experimental vs. spike-in targets are expected in any workflow using conserved orthologous targets where a primary antibody is reactive across species, or for dual-antibody approaches (i.e., setups with a Drosophila antibody paired with Drosophila-specific epitopes), this generally is not a concern for PerCell analysis. However, this issue becomes problematic for data interpretation if the difference is significant enough to drop the percentage of spike-in reads below ∼2% (where small variations in spike-in content would drastically impact normalization) or above the point where read depth dedicated to experimental ChIP data is adversely affected. Even in cases where an antibody has a uniquely specific affinity for a target in experimental chromatin, nonspecific (but consistent levels of) immunoprecipitation of orthologous chromatin may provide sufficient spike-in reads, especially if the ratio of experimental:spike-in cells is decreased. That is, spike-in reads for our methodology do not necessarily need to occur within defined peaks (Figure S3, scenario 1), which will inevitably differ from the peaks that occur across the experimental genome. This is also a key difference from other spike-in workflows, where chromatin regions are separated into peak and non-peak binned groups.15,16 As is the case for all input samples, a consistent number of nonspecific or background reads (Figure S3, scenario 2) can be used successfully for the normalization of experimental immunoprecipitated samples as long as they (1) remain consistent across replicates and (2) occur within the acceptable range of experimental:spike-in read ratios discussed above.
In cases where the above considerations remain a challenge, alternative methods where a second antibody is uniquely directed to a target found in spike-in chromatin11,13 may be preferable (see Table S5 for an overview of some key methods in the literature). The use of more distantly related species for spike-in cells is not without its own drawbacks, however. Sonication efficiency, or the rate at which experimental and spike-in chromatin shears into fragment lengths suitable for immunoprecipitation and sequencing, is a critical consideration in the choice of appropriate spike-in content and could be dependent on cell type or species of origin, among other considerations. Using more closely related species may be a significant aid in finding cells that have similar DNA fragment length distribution profiles after sonication. Notably, we found that zebrafish embryonic cells and human rhabdomyosarcoma cells sonicate with equal conditions,20 suggesting that testing is critical to support experimental goals. Additionally, using a second antibody intentionally directed to spike-in chromatin necessitates verifying that there is, in fact, no “reverse” cross-reactivity with this antibody toward experimental chromatin, which is especially relevant for highly conserved orthologous targets. These factors may be further complicated in setups where comparisons are being made across multiple experimental cell types or treatment conditions. We also note that our bioinformatic pipeline remains amenable to the analysis of these alternative experimental setups.
The PerCell analysis workflow is facile for rapid data processing of standard, non-spike-in ChIP-seq, which may be useful in cases where the ratio of spike-in versus test chromatin may be affected by PCR cycle number in library preparation or by test cell types that generate very different sonication patterns from control spike-in cells. For example, combinations of human immune cells and mouse myoblast cells may be more challenging because the differentiation states may result in vastly different sensitivities to the mechanical shearing forces of sonication, but mouse fibroblasts and human myoblasts may be a technically more facile combination of experimental/spike-in cell chromatin based on comparative differentiation states and more similar distributions of chromatin sonication fragment lengths. Thus, consideration must be made for sonication efficiency across test and orthologous cell chromatin within an experimental setup, but the pipeline easily allows the user universal access to non-spike-in comparisons of the same datasets.
Resource availability
Lead contact
Benjamin Stanton (benjamin.stanton@nationwidechildrens.org) will fulfill any requests for additional information, data, and other resources.
Materials availability
This study did not generate new unique reagents.
Data and code availability
-
•
All ChIP-seq datasets generated with this protocol and presented here were deposited in the NCBI Gene Expression Omnibus (http://www.ncbi.nlm.nih.gov/geo/; GEO: GSE271986 and GSE270319).
-
•
All code for the PerCell analysis pipeline is publicly available on GitHub (https://github.com/lextallan/PerCell and https://doi.org/10.5281/zenodo.12730196). Downloading the pipeline is not necessary before running it, as Nextflow will automatically download it during the first execution (https://www.nextflow.io/docs/latest/sharing.html#how-it-works).
-
•
Any additional information needed to re-analyze the results reported in this paper is available from the lead contact upon request.
Acknowledgments
We thank the Nationwide Children’s Hospital (NCH) Animal Resources Core for their exceptional zebrafish husbandry, especially the Zebrafish Facility team members Dr. Laurie Goodchild, Dr. Carmen Arsuaga, Dr. Lindsey Ferguson, Logan Fehrenbach, Alex Kramer, and Logan Bern. Additional support was provided by the NCH Institute for Genomic Medicine, the NCH Genomics Services Laboratory, and the NCH High-Performance Computing group, who assisted with maintaining and using the NCH cluster. We also thank Dr. Matthew Kent for assistance with cloning. B.Z.S. is grateful to the American Cancer Society (RSG-23-1021178-01-DMC), St. Baldrick’s Foundation (career development award), National Institutes of Health (R01GM144601 and 1R01HL166520 - 01A1), and intramural funding from Nationwide Children's Hospital for supporting this work. We are grateful to all members of the Stanton and Kendall groups for helpful discussions. G.C.K. is grateful for support from NIH/NCI R01 grant R01CA272872, an Alex's Lemonade Stand Foundation “A” Award, a V Foundation for Cancer Research V Scholar Award, a CancerFree Kids New Idea Award, and startup funds from The Abigail Wexner Research Institute at Nationwide Children’s Hospital. B.Z.S. and G.C.K. were supported by a Nationwide Children's Hospital Seed Fund from the Center for Childhood Cancer Research. J.K. is supported by a T32 CA269052 Training Program in Basic and Translational Pediatric Oncology Research predoctoral fellowship. The Institute for Genomic Medicine is funded by the Nationwide Foundation Pediatric Innovation Fund and Ohio State University Comprehensive Cancer Center grant P30 CA016058. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Further, the content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Author contributions
Conceptualization, A.T., B.S., S.L., and B.Z.S.; methodology, A.T., B.S., S.L., M.W., and B.Z.S.; formal analysis A.T., J.K., C.T., and M.W.; investigation, A.T., J.K., and B.S.; resources, Q.L. and J.Q.; writing – original draft, A.T., J.K., and B.Z.S.; writing – review & editing, A.T., J.K., G.C.K., and B.Z.S.; visualization, A.T., J.K., and B.Z.S.; supervision, G.C.K. and B.Z.S.; project administration, G.C.K. and B.Z.S.; funding acquisition, G.C.K. and B.Z.S.
Declaration of interests
B.Z.S., M.W., and B.S. declare a related patent application, US20230365637A1, titled “Identification of pax3-foxo1 binding genomic regions” with current assignee as Nationwide Children’s Hospital, November 16, 2023. S.L. is currently employed by Tempus AI, while her contributions to the work under submission have all been performed at Nationwide Children's Hospital, which is reflected by her affiliation on the title page of the manuscript.
STAR★Methods
Key resources table
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Antibodies | ||
| anti-FOXO1 Ab C29H4 | Cell Signaling Technologies | Cat#2880S; RRID:AB_2106495 |
| anti-H3K27ac Ab | Active Motif | Cat#39133; RRID:AB_2561016 |
| anti-CTCF Ab | Cell Signaling Technologies | Cat#2889; RRID:AB_2086794 |
| Chemicals, peptides, and recombinant proteins | ||
| DMSO, molecular biology grade | SIGMA | Cat#D8418 |
| Entinostat | Selleckchem | Cat#S1053 |
| A485, synthetic | Jun Qi lab, Dana-Farber Cancer Institute | N/A |
| Formaldehyde (methanol-free) | Fisher Scientific | Cat#PI28906 |
| Turbo DNAse | ThermoFisher | Cat#AM2238 |
| Glycine | SIGMA | Cat# G7126 |
| Pronase | SIGMA | Cat#11459643001 |
| Proteinase K, recombinant PCR grade | SIGMA | Cat#3115828001 |
| MinElute PCR Purification kit | QIAGEN | Cat#28006 |
| Fetal Bovine Serum (10 x 50 mL) | Fisher Scientific | Cat#A3160502 |
| DMEM | Quality Biological | Cat#112-013-101CS |
| 2% E-Gel EX Agarose Gels | Fisher Scientific | Cat#G401002 |
| Dynabeads Protein-A | Fisher Scientific | Cat#10002D |
| Dynabeads Protein G | Fisher Scientific | Cat#10-003-D |
| T4 DNA Ligase | NEB | Cat#M0202L |
| End-Repair Enzyme mix (T4 DNA Pol + T4 PNK) | NEB | Cat#E6051AAVIAL |
| Klenow Fragment (3–5 exo-) | NEB | Cat# M0212S |
| 100 mM dATP | NEB | Cat#R0141 |
| 10 mM ATP | Fisher Scientific | Cat#10-297-018 |
| 2.5 mM dNTP mix | Fisher Scientific | Cat#10-297-018 |
| Phusion Taq 2X Master Mix | NEB | Cat#M0531S |
| Negative control ChIP-qPCR primers for zebrafish | Active motif | Cat#71035 |
| 100X SYBR Green | Fisher Scientific | Cat#S7563 |
| AMPure XP beads | Beckman Coulter | Cat#A63880 |
| Phusion High-Fidelity PCR Master Mix with HF Buffer | NEB | Cat#M0531S |
| NP-40/IGEPAL CA-630 | SIGMA | Cat#I3021-100ML |
| Deoxycholate | SIGMA | Cat#D6750-100G |
| Triton X-100 | SIGMA | Cat#T8787-250mL |
| Trypsin (sterile) | Fisher Scientific | Cat#25200114 |
| Protease Inhibitor Cocktail | Active Motif | Cat#37491 |
| GlutaMax Supplement | Fisher Scientific | Cat#35050061 |
| Water (molecular biology grade) | Quality Biological | Cat#351-029-101CS |
| Critical commercial assays | ||
| MinElute PCR Purification kit | QIAGEN | Cat#28006 |
| Deposited data | ||
| Raw and analyzed ChIP-seq data (human, mouse spike-in) | This paper | GSE271986 |
| Raw and analyzed ChIP-seq data (zebrafish, human spike-in) | This paper; Kucinski et al.20 | GSE270319 |
| Experimental models: Cell lines | ||
| Rh4 cells | Laboratory of Dr. Peter Houghton (UTHSCSA) | N/A |
| Rh30 cells | Laboratory of Dr. Peter Houghton (UTHSCSA) | N/A |
| C2C12 cells | ATCC | Cat#CRL-1772 |
| Experimental models: Organisms/strains | ||
| WIK zebrafish | Zebrafish International Resource Center (ZIRC; https://zebrafish.org/) | Cat#ZL84 |
| Oligonucleotides | ||
| Positive control ChIP-qPCR primers for zebrafish, nrp2a locus | FWD: CACAGCACTCATAAGCGAAGC | Kucinski et al.20 |
| N/A | REV: TTGCACGGCGGTAAACAATC | Kucinski et al.20 |
| Negative control ChIP-qPCR primers for human Rh30 cells, SOX18 locus | FWD: GCTCTTGGTTCTCTGTCCCT | This paper |
| N/A | REV: AGACAGACTGTGATGTGGGG | This paper |
| Positive control ChIP-qPCR primers for human Rh30 cells, FGFR4 locus | FWD: AAATTTGACCTTCGTCGGCAC | This paper |
| N/A | REV: CAGCTGTTGGCGATTTCACG | This paper |
| Positive control ChIP-qPCR primers for human Rh30 cells, QKI locus | FWD: TGCATGCTGGTGACAGATCA | This paper |
| N/A | REV: ACAGCGTCCTCTTTCAGCTT | This paper |
| Positive control ChIP-qPCR primers for human Rh30 cells, MYCN locus: | FWD: TCTCCAATTCTCGCCTTCAC | This paper |
| N/A | REV: GCGCTAACAGGTTTCTGTCC | This paper |
| PE Adapter Oligo Mix | N/A | Gryder et al., 202028 |
| Illumina next generation sequencing index primers PE2.0 (i7) and PE1.0 (i5) | N/A | N/A |
| Software and algorithms | ||
| PerCell pipeline | This paper | https://github.com/lextallan/PerCell; https://doi.org/10.5281/zenodo.12730196 |
| Singularity | Kurtzer et al., 201729 | https://docs.sylabs.io/guides/latest/admin-guide/installation.html |
| Nextflow | Di Tommaso et al., 201718 | https://www.nextflow.io/docs/latest/install.html |
| nf-core | Ewels et al., 202019 | https://zenodo.org/records/7139814 |
| Cutadapt | Martin, 201130 | https://doi.org/10.14806/ej.17.1.200 |
| FastQC | Andrews, 201031 | http://www.bioinformatics.babraham.ac.uk/projects/fastqc/ |
| Bowtie2 | Langmead and Salzberg, 201232 | https://bowtie-bio.sourceforge.net/bowtie2/index.shtml |
| Chromap | Zhang et al., 202133 | https://haowenz.github.io/chromap/ |
| BEDTools | Quinlan and Hall, 201034 | https://bedtools.readthedocs.io/en/latest/ |
| SAMtools | Danecek et al., 202135 | https://www.htslib.org/ |
| Picard | 201936 | https://broadinstitute.github.io/picard/ |
| MACS2 | Zhang et al., 200837 | https://macs3-project.github.io/MACS/ |
| UCSC tools | Kent et al., 201038 | https://github.com/ucscGenomeBrowser/kent-core |
| HOMER | Heinz et al., 201039 | http://homer.ucsd.edu/homer/index.html |
| MultiQC | Ewels et al., 201640 | https://multiqc.info/ |
| IDR | Li et al., 201141 | http://projecteuclid.org/euclid.aoas/1318514284 |
| Other | ||
| Eppendorf Thermomixer | ThermoFisher | Cat#RS232C |
| Qubit Fluorometer | ThermoFisher | Cat#Q33239 |
| EpiShear Probe Sonicator (110 V) | Active Motif (interchangeable with other sonication platforms) | Cat# 53051 |
| EpiShear Cooled Sonication Platform (1.5 mL) | Active motif (interchangeable with other sonication platforms) | Cat# 53080 |
| E-Gel Power Snap Electrophoresis Device | ThermoFisher | Cat# G8100 |
| T100 Thermal Cycler | BioRad | Cat#1861096 |
| LUNA Automated Cell Counter | Logos Biosystem | Cat#L10001 |
| 0.5X Danieau’s buffer | (29 mM NaCl, 0.35 mM KCl, 0.2 mM MgSO4·4H2O, 0.3 mM Ca(NO3)2·4H2O, 2.5 mM HEPES) | N/A |
| Deyolking buffer | (55 mM NaCl, 1.8 mM KCl, 1.25 mM NaHCO3) | N/A |
| Embryo buffer 3 | (5 mM NaCl, 0.17 mM KCl, 0.33 mM CaCl2, 0.33 mM MgSO4) | N/A |
| ChIP wash Buffer 1 | (TE pH 8.0, with 0.1% SDS, 0.1% sodium deoxycholate, and 1% Triton X-100) | Kidder and Zhao, 201442 |
| ChIP wash Buffer 2 | TE pH 8.0, with 1% Triton X-100, 0.1% SDS, 0.1% sodium deoxycholate, and 200 mM NaCl) | Kidder and Zhao, 201442 |
| ChIP wash Buffer 3 | (TE pH 8.0 with 250 mM LiCl, 0.5% NP-40, 0.5% sodium deoxycholate) | Kidder and Zhao, 201442 |
Experimental model and study participant details
Cell culture
PAX3::FOXO1 fusion-positive Rh4 and Rh30 human rhabdomyosarcoma cells (both gifts from Dr. Peter Houghton at UT Health San Antonio) and C2C12 mouse myoblast cells were cultured in DMEM supplemented with Glutamax, FBS (10% v/v), and penicillin-streptomycin.
Zebrafish
Zebrafish (Danio rerio) were housed in an AAALAC-accredited, USDA-registered, OLAW-assured facility at Nationwide Children’s Hospital. Adult zebrafish were kept in mixed-sex groups of 5–12 fish per liter. System water was 28°C (conductivity, 510 to 600 μS; pH, 7.3 to 7.7; hardness, 80 ppm; alkalinity, 80 ppm; dissolved oxygen, greater than 6 mg/L; ammonia, 0 ppm; nitrate, 0 to 0.5 ppm; and nitrite, 0 ppm) and carbon-filtered (20-μm pleated particulate filter, 40W UV light exposure) from municipal tap water. From 5 to 30 days old, the fish were fed live rotifers three times daily and after 30 days were fed a commercial pelleted diet twice daily. In our facility, zebrafish sex is not determined until 30–60 days old, therefore zebrafish embryo micro-injections were sex-blinded.
Method details
PerCell ChIP-seq sample preparation
Timing: 1 h (sample generation and collection will vary based on experimental set-up)
Experimental and spike-in samples/cells were collected and dissociated into single-cell suspensions into Falcon or microcentrifuge tubes. Cells were centrifuged at 300g for 5 min at room temperature. Note that experimental and spike-in samples MUST be from orthogonal species, so samples can be de-convoluted during genomic sequence alignment. Spike-in sample should be collected at the same time for all experimental conditions to minimize technical variation.
Supernatant was carefully removed and cells suspended in PBS and fixed in 1% formaldehyde for 10 min at room temperature. Cells were inverted initially and again at every 2-3 min to ensure thorough and even fixation. Fixation was quenched by addition of glycine solution to a final concentration of 125 mM and incubated on ice for 5 min. Cells were again inverted as needed throughout to ensure thorough quenching.
Fixed cells were pelleted via centrifugation at 4°C, 1,200g for 5 min, with supernatant carefully removed and cells resuspended in PBS containing proteinase inhibitors. Cell counts were obtained and aliquoted as desired into microcentrifuge tubes. It is recommended to store cells in equal aliquots of 1–6 million cells, however aliquots can be later pooled to reach the desired cell number. Cell aliquots were pelleted via centrifugation at 4°C, 1,200g for 5 min. Cell pellets can be stored for future use at −80°C following snap freezing or immediately used.
PerCell chromatin immunoprecipitation
Timing: 2–3 days.
Fixed cells were thawed slowly on ice and experimental and spike-in cells combined in TE buffer (pH 8.0) supplemented with proteinase inhibitors. TE volume should be determined based on the recommended sonication volume according to the specific sonicator being used.
For the purposes of downstream PerCell ChIP-seq analysis, please note that the ratio of experimental to spike-in cells MUST be consistent across samples as the pipeline determines the scaling ratio based on the consistent spike-in ratio in each sample. It is recommended to have ∼25% of reads be from the spike-in cells. However, this pipeline has successfully analyzed datasets with as high as ∼75% spike-in reads and as low as 2.5% spike-in reads. A higher percentage of spike-in cells may help serve as carrier chromatin to decrease the loss of the experimental sample throughout the protocol but will result in lower sequencing depth for your experimental samples.
We have included a table (Table S6) with recommended spike-in ratios for species currently supported by our pipeline, both based on this work and previously reported quantitative ChIP-seq methodologies.11,13,15,16 While the minimum total number of recommended spike-in reads has previously been reported to be one million,15 we do not find such a strict requirement to be necessary for our normalization strategy or analyses. We note that the requisite number of cells for successful ChIP-seq assays can vary widely, depending on factors including the prevalence of the target of interest, its occupancy time and binding footprint within the genome, the sonication efficiency of cells, and treatment conditions being investigated. However, for our reported human:mouse and zebrafish:human experiments (see also Kucinski et al.20), we recommend 6M human cells mixed with 2M spike-in mouse cells and 2M zebrafish cells mixed with 250k spike-in human cells, respectively.
Samples were sonicated for the desired time and according to guidelines specific to the sonicator. We recommend sonicating the samples for the minimum amount of time required to achieve the desired fragmentation. For a histone modification ChIP-seq, DNA should be enriched around 200 bp. Meanwhile, a more transient protein-DNA interaction (i.e., transcription factor or chromatin remodeler) ChIP-seq sample should be sonicated until the DNA is evenly distributed across fragment length to prevent over-sonication of the isotope.
5 μL of sonicated DNA was removed and kept to serve as input DNA and to assess sonication efficiency. The remaining DNA can be stored at 4°C until sonication efficiency is confirmed. The choice of experimental and spike-in cells should ideally be made to ensure sonication at similar rates, which can be evaluated by sonicating both samples individually while determining the ideal sonication condition. However, experimental and spike-in cells MUST be sonicated together for this protocol to account for (and evaluate) variations across samples and conditions.
Input DNA was reverse cross-linked with 5 μL sonicated DNA, 20 μL TE buffer, 1 μL 10% SDS, and 1μL 18.5 mg/mL Proteinase K at 65°C overnight (less than 24 h). Alternatively, a shorter reaction can be achieved via using 5 μL sonicated DNA, 8 μL 18.5 mg/mL Proteinase K, 4 μL 5M NaCl, and 90 μL H2O at 65°C with 400 rpm in a ThermoMixer for 1 h.
De-crosslinked input DNA was purified with a Qiagen MinElute PCR Purification kit (although other similar kits/protocols can be easily substituted). In order to maximize yield of purified DNA, we recommend repeating the column binding step two times and incubating with 36 μL of pre-warmed EB buffer for 4 min. After elution, 1 μL of purified input DNA was run on an agarose E-gel (2%, EX) to confirm sonication efficiency.
If samples are under-sonicated (i.e., DNA fragments are enriched at a greater length or is minimally enriched at 200) the remaining DNA can be further sonicated and de-crosslinking/purification repeated. If samples are over-sonicated (i.e., fragment enrichment at less than 200 bp) chromatin immunoprecipitation should be repeated with new cells.
After successful sonication, the remaining sonicated DNA (ChIP DNA) was adjusted to a final buffer concentration of 1% Triton X-100, 0.1% SDS, 0.1% sodium deoxycholate, and 200 mM NaCl. Samples were then centrifuged at 4°C, <18,000g for 10 min to spin-down any remaining cellular debris. Supernatant was carefully pipetted into a clean microcentrifuge tube and incubated with 2–4 μg of your desired primary antibody (or according to the antibody manufacturer’s recommendation) for 2 h at 4°C with constant overhead rotation.
While the antibody of choice must be specific enough to experimental chromatin to enable a sufficient signal:noise ratio for quality ChIP-seq analysis, the antibody does not necessarily need to be equally directed to targets within the spike-in cells. The methodology is effective as long as the antibody pulls down a consistent amount of spike-in chromatin. In cases where the antibody of choice pulls down spike-in chromatin inefficiently (e.g., <5% of aligned reads to the spike-in genome), we would recommend increasing the number of spike-in cells added to each sample. Alternatively, as we have demonstrated that even closely related species can be reliably distinguished (see Figure S2), the selection of a more closely related spike-in species or cell type would be beneficial.
40 μL per sample of Dynabeads Protein-A or Dynabeads Protein-G, according to the primary antibody isotype, was placed on a magnetic rack. Once beads were isolated, they were resuspended in 200-400μL of 1% Triton X-100, 0.1% SDS, 0.1% sodium deoxycholate, and 200 mM NaCl with brief overhead rotation. After washing, beads were again isolated and finally re-suspended in a volume equal to the initial bead volume of 1% Triton X-100, 0.1% SDS, 0.1% sodium deoxycholate, and 200 mM NaCl. 40 μL of these equilibrated Dynabeads were added to each sample and mixed overnight at 4°C with overhead rotation.
The immunoprecipitated beads were washed with the following buffers. Washes were performed by i) isolating beads via placing the microcentrifuge tube on a magnetic rack, ii) removing the supernatant, iii) adding 250–500 μL of the respective wash buffer, iv) briefly rotating until beads are re-suspended, and v) repeating as indicated.
Buffer 1: 2 total washes of TE with 0.1% SDS, 0.1% sodium deoxycholate, and 1% Triton X-100.
Buffer 2: 2 total washes of TE with 1% Triton X-100, 0.1% SDS, 0.1% sodium deoxycholate, and 200 mM NaCl.
Buffer 3: 2 total washes of TE with 250 mM LiCl, 0.5% NP-40, 0.5% sodium deoxycholate.
Buffer 4: 1 total wash with TE.
Immunoprecipitated DNA was reverse cross-linked overnight by suspending beads in 100 μL TE, 2.5 μL 10% SDS, and 5 μL 18.5 mg/mL Proteinase K with overnight incubation at 65°C (less than 24 h). Beads were placed on a magnetic rack and the DNA-containing supernatant was removed to a clean microcentrifuge tube. Immunoprecipitated DNA was purified with a Qiagen MinElute PCR Purification kit. As above, we recommend repeating the column binding step two times and incubating with 36 μL warm EB buffer for 4 min to maximize DNA yield.
qPCR was used to confirm immunoprecipitation efficiency using 1 μL of ChIP and input DNA for regions in experimental and/or spike-in cells. Immunoprecipitated and input DNA can be stored at −20°C until ready for library preparation.
Library preparation
Timing: 1–2 days.
12-14 μL of ChIP DNA or 2–4 μL of input DNA was taken for library preparation. We recommended the preparation of no more than 8 libraries simultaneously. If multiple days of library preparation are required, stagger replicates to account for potential batch effects.
End-repair reaction was performed for each sample using Lucigen’s End-It DNA End-Repair Kit (although we have had success with similar kits as well) to combine the following:
34 μL DNA + H2O (e.g., 2μL input DNA +32μL H2O; 12uL ChIP DNA +22μL H2O)
5 μL 10X End Repair buffer (330 mM Tris-acetate pH 7, 660 mM potassium acetate, 5 mM DTT)
5 μL 2.5 mM dNTP.
5 μL 10 mM ATP.
1 μL End-Repair Enzyme mix (T4 DNA Pol + T4 PNK)
The entire mixture was incubated for 45 min at room temperature on a rocker. DNA was purified with a Qiagen MinElute PCR Purification kit and eluted in 32 μL EB. An A-tailing reaction was next performed on the purified result for 30 min on a ThermoMixer at 37°C, 300rpm, as below:
32 μL DNA (from End repair)
5 μL 10X Klenow Buffer (NEBuffer 2)
10 μL 1 mM dATP.
3 μL 5 units/μL Klenow Exo.
A-tailed DNA was purified with a Qiagen MinElute PCR Purification kit and eluted in 22 μL EB. Complete linker ligation reaction with 30–60 min of room temperature incubation on a rocker.
22 μL DNA (from A-tailing)
3 μL 10X T4 DNA Ligase Buffer.
2 μL PE Adapter Oligo Mix.
3 μL T4 DNA Ligase (400 units/μL)
(PE Adapter Oligo mix can be prepared, in advance, by diluting top and bottom adapter primers to a final concentration of 15 μM each. The adapter mix was heated to 95°C for 5 min and cool at room temperature for 2h)
Properly ligated products were selected with AMPure beads (beads were warmed to room temperature before selection). 30 μL (1.0X) AMPure beads were added to the mixture, mixed well by pipette, and incubated at room temperature for 5 min followed by isolation on a magnetic rack for another 5 min. Supernatant was removed and the beads washed with fresh 80% ethanol on the magnetic rack before being air dried for 10 min to remove excess ethanol. The beads should be dried but not cracked. The exact incubation time may vary. 23 μL EB was added to elute DNA off of beads. and incubated at room temperature for 5 min followed by isolation on a magnetic rack for another 5 min. The supernatant was collected, and can be stored at −20°C until ready for amplification.
DNA was amplified via the SYBR-based qRT-PCR reaction below. We recommend only using approximately half of the eluted DNA as the other half can be retained in case libraries are over-amplified or if more DNA is needed for sequencing following library preparation.
Amplification reaction:
11.4 μL DNA
12.5 μL Phusion Taq 2X Master Mix
0.5 μL PCR primer lnPE1.0 (2X diluted) (i5)
0.5 μL PCR primer lnPE2.0 (2X diluted) (i7)
0.15 μL 100X SYBR Green
PCR protocol:
98°C for 30 s.
Cycle - 98°C for 10 s, 65°C for 30 s, 72°C for 30 s.
Note that amplification reactions should be stopped before they plateau to minimize PCR duplicates. It is recommended to start with 10 cycles for input DNA or highly abundant histone marks/factors ChIPs and 15 cycles for less abundant histone marks/factors. Cycle number will vary by experiment. Amplified libraries were purified using AMPure beads (beads were warmed to room temperature before selection). 45 μL (1.8X) AMPure beads were added to each sample, pipetted well to mix, and incubated at room temperature for 5 min followed by isolation on a magnetic rack for another 5 min.
Supernatant was removed and the beads washed with fresh 80% ethanol on a magnetic rack before being air dried for 10 min to remove excess ethanol. The beads should be dried but not cracked. 20 μL EB was added to each sample’s beads to elute indexed DNA samples. Samples were incubated at room temperature for 5 min and then DNA isolated from beads on a magnetic rack over another 5 min. Supernatant was collected and stored at −20°C.
DNA libraries may be assessed by 1) running on E-Gel or bioanalyzer to confirm library fragments are ∼300 bp and/or 2) obtaining DNA concentration by fluorometer/spectrophotometer.
Specific notes for ChIP-seq with zebrafish PAX3::FOXO1 mRNA injection model
Adult wildtype zebrafish (Danio rerio) were maintained in an aquatics facility in compliance with the Guide for the Care and Use of Laboratory Animals. However, various mutant zebrafish strains can be used to assess potential cooperativity in the model system. The human PAX3::FOXO1 coding sequence was synthesized and cloned as previously reported,43 with the addition of Pac1 and Asc1 enzyme digestion sites. A PAX3FOXO1-2A-sfGFP plasmid was generated for protein visualization by cloning the PAX3::FOXO1 product into pCS2+MCS-P2A-sfGFP (Addgene plasmid #74668). mRNA was generated by SP6 reaction of linearized PAX3FOXO1-2A-sfGFP DNA with Not1 digestion. Single-cell zebrafish embryos were injected with 100 ng/μL PAX3FOXO1-2A-sfGFP a drop size diameter of 0.15 mm. A control mRNA construct was similarly injected with equal molarity. DNA constructs were generated by linearizing the pCS2+MCS-P2A-sfGFP plasmid with Not1 digestion or amplifying out sfGFP and polyA tail from pCS2+MCS-P2A-sfGFP and transcribing DNA into mRNA as mentioned. Following injection, zebrafish cells were incubated at 32°C for 5.25 h to reach the proper developmental stage. Embryos were then dechorionated with pronase, de-yolked in deyolking buffer (55 mM NaCl, 1.8 mM KCl, 1.25 mM NaHCO3), and washed with 0.5X Danieau’s buffer (29 mM NaCl, 0.35 mM KCl, 0.2 mM MgSO4·4H2O, 0.3 mM Ca(NO3)2·4H2O, 2.5 mM HEPES). Embryos were dissociated into single cells by thoroughly pipetting in dPBS. Samples were fixed and stored until enough cells were collected. ChIP-seq and library preparation was completed as described above. Generally, we could obtain approximately 1 million cells with 500 embryos, which took 1 h of injecting. PAX3::FOXO1 ChIP-seq (4 μL anti-FOXO1 - Cell Signaling, C29H4) used 2 million zebrafish cells and 250,000 Rh30 cells for spike-in. Immunoprecipitation was validated with qPCR using SOX18 (negative) and FGFR4 (positive) primers for Rh30 cells and negative control (Active motif, 71035) and nrp2a (positive) primers for zebrafish cells. H3K27ac ChIP-seq (4 μL anti-H3K27ac - Active Motif, 39133) used 1 million zebrafish cells and 1 million Rh30 cells for spike-in. Immunoprecipitation was validated with qPCR using SOX18 (negative) and QKI or MYCN (positive) primers for Rh30 cells.
Bioinformatic analysis using PerCell pipeline
Required hardware
-
(1)
Computer with Unix-based operating system
-
(2)
A computer cluster or cloud-based computing platform for executing the pipeline (currently the pipeline is tested for compatibility with SLURM-managed platforms)
Required software
Nextflow18 installation: https://www.nextflow.io/docs/latest/install.html.
Singularity29 installation: https://docs.sylabs.io/guides/latest/admin-guide/installation.html.
Data setup and automated pipeline execution steps
Timing: 30 min-1 h setup; 2–8 h execution.
Install both Nextflow and Singularity. Ensure both are properly running on the system.
If necessary, download files for chosen experimental and spike-in reference genomes in fasta format. Example commands for downloading and uncompressing human (hg38), mouse (mm10), zebrafish (danRer11), and fly (dm6) from UCSC can be found below:
$ wget http://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/hg38.fa.gz
$ gunzip hg38.fa.gz
$ wget http://hgdownload.soe.ucsc.edu/goldenPath/mm10/bigZips/mm10.fa.gz
$ gunzip mm10.fa.gz
$ wget http://hgdownload.soe.ucsc.edu/goldenPath/danRer11/bigZips/danRer11.fa.gz
$ gunzip danRer11.fa.gz
$ wget http://hgdownload.soe.ucsc.edu/goldenPath/dm6/bigZips/dm6.fa.gz
$ gunzip dm6.fa.gz
Create a csv file with names of samples, full locations of fastq files, antibodies/treatment used, and associated control sample names. Example input csv files are available at the PerCell GitHub repository (https://github.com/lextallan/PerCell; https://doi.org/10.5281/zenodo.12730196).
-
(1)
The first line serves as column headings and must consist of the following: “sample,fastq_1,fastq_2,antibody,control”
-
(2)
Each additional line follows that format for each sequenced sample
-
(3)
Sample names should be unique, unless the same sample has been re-sequenced. Additional sequencing of the same sample can use identical sample names and the pipeline will merge the associated fastq files before alignment.
-
(4)
Samples with multiple replicates should have sample names ending in a similar pattern followed by the number of the replicate, e.g., “_1”/“_2” or “_rep1”/“_rep2”. These will be treated as replicates by the pipeline for downstream analyses.
-
(5)
fastq_1 and fastq_2 columns should contain the absolute path to the files, in an accessible directory. Files may be gzipped.
-
(6)
Antibody column should contain info about the antibody used as well as any treatments done on the sample. Samples with identical antibody columns will be combined for peak calling. For example, samples created using antibody A but distinct cell lines/treatments B or C could be written as “antibodyA_B” and “antibodyA_C”, respectively. For input samples, this column should be left blank.
-
(7)
Control column should contain the exact name of the sample for the associated input, including replicate information. For example, “sampleA-Input_rep1” or “sampleA-Input_rep2” and NOT “sampleA-Input”.
Optionally, the file named nextflow.config may be edited before running the pipeline with parameter choices (see execution step below for specific explanations) and options specific to the computing environment used (many organizations have suggested nextflow setup and config options, as collected by the nf-core community: https://nf-co.re/configs). The PerCell config file can be manually adjusted after directly downloading from our GitHub repository, (https://github.com/lextallan/PerCell/blob/master/nextflow.config).
Alternatively, parameter options can be specified at the command line at the point of pipeline execution.
Finally, execute the main pipeline script, PC.nf, with desired parameters immediately following (explained below). For ease of use, the entire pipeline and all potential parameters can be launched and customized from this single command (i.e., in one command line).
$ nextflow run lextallan/PerCell/PC.nf
Pipeline parameter options
Make careful note of whether one or two dashes are used in the following parameters. Single dashes are for general nextflow options, double dashes are for pipeline-specific parameter choices. Available and default choices are given for relevant parameters.
-profile singularity (tells nextflow to download/use singularity images for each process)
-qs: optional nextflow parameter to limit queue size, or number of processes that will be submitted at a time (e.g., to prevent overwhelming a shared cluster)
--input: path to csv input file (see above for required format)
--outdir: path to directory where the pipeline’s output will be found
--aligner: choice of alignment tool with which samples will be aligned, <bowtie2,chromap> (default: bowtie2)
--experimental: species identify from which the experimental cells originate, <human, mouse,zebrafish,fly> (default: human)
--spikein: species identify from which the cells used for spike-in originate, <human, mouse,zebrafish,fly> (default: mouse)
--human_fa: path to human reference fasta file, if used as either experimental or spike-in genome
--mouse_fa: path to mouse reference fasta file, if used as either experimental or spike-in genome
--zebrafish_fa: path to zebrafish reference fasta file, if used as if used as either experimental or spike-in genome
--fly_fa: path to fly reference fasta file, if used as if used as either experimental or spike-in genome
--skip_fastqc: choice of whether or not to skip FastQC quality control assessments, <true, false> (default: false)
--skip_trimming: set to true if fastq files have already been trimmed of adapter sequences with an algorithm such as TrimGalore, <true, false> (default: false)
--save_trimmed: if set to true, trimmed fastq files will be saved as output, <true, false> (default: false)
--override_spikeinfail: the pipeline will not attempt to downsample experimental samples if spike-in content is detected to be below 0.5% of properly aligned total reads, setting this parameter to true will ignore this threshold and attempt normalization regardless of detected spike-in reads, <true, false> (default: false)
--skip_downsample: set to true to disable PerCell normalization and carry out a standard ChIP-seq pipeline analysis, <true, false> (default: false)
--seed: random seed used to downsample reads (default: 0)
--macs2_peak_method: choice of method with which to score macs2 bedgraphs for peak calling, <ppois, qpois,subtract, logFE,FE,logLR, slogLR,max> (default: ppois)
--macs2_cutoff: p/q-value cutoff used for calling peaks; unlike macs2’s default ‘callpeak’ command, cutoffs are taken in -log10(x) form. If using IDR, a particularly relaxed cutoff is highly recommended. For ChIPs targeting broader histone marks or when there are >2 replicates, IDR is not recommended and a more stringent cutoff of 1.301 (p/q-value of 0.05) or even 2 (p/q-value of 0.01) is suggested instead (default: 0.2218)
--macs2_bigwig_method: choice of method with which to generate bigWig tracks for visualizing ChIP data within a genome browser, <ppois, qpois,subtract, logFE,FE,logLR, slogLR,max> (default: ppois)
--skip_idr: if the IDR (Irreproducible Discovery Rate, https://github.com/nboley/idr) framework should be used to statistically identify significant peaks across two replicates. According to authors, the software is not designed for use with particularly broad peaks (e.g., those call in a ChIP for H3K9me3) and requires using a relaxed macs2 cutoff value in order to properly function, <true, false> (default: true)
--idr_cutoff: cutoff value over which peaks will not be included within the output (default: 0.05)
--skip_consensus: choice of whether consensus peaks should be identified across replicates, especially recommended along with a more stringent macs2 cutoff if there are >2 replicates, <true, false> (default: false)
--skip_annotation: whether or not to use HOMER’s ‘annotatePeaks’ tool for annotation of peaks with nearest genes and genomic features <true, false> (default: false)
--skip_motif: whether or not to use HOMER’s ‘findMotifsGenome’ tool to find enriched DNA sequence motifs in called peaks <true, false> (default: false)
Pipeline output
After initiation of the pipeline, the user specified input csv (“--input”) is first checked to ensure proper formatting and that all files both exist and are accessible to the pipeline. All pipeline outputs are stored in a directory named after the choice of alignment tool (“--aligner”) within the user-specified output directory (“--outdir”). When the applicable options are enabled (i.e., when “--skip_trimming” and “--skip_fastqc” set to false), adapter sequences are trimmed from reads and quality control data is collected and made available to the user via the MultiQC software40 in readable html format (MultiQC-Report-for-PerCell-pipeline.html). For user convenience, trimmed fastq files can be retained and organized for future use via the “--save_trimmed” parameter and stored in the trimgalore directory.
Each fastq sequencing file is separately aligned to both the chosen experimental (“--experimental”) and spike-in (“--spikein”) genomes in a parallel fashion, with these intermediate alignments available in the aligned directory. Supported aligners include Bowtie232 and Chromap.33 Misaligned and duplicated reads are removed using the Samtools35 and Picard36 software suites, with post-filtering alignments and metrics data stored in the deduplicated directory.
Next, the percentages of overlapping reads (independently aligning to both genomes) and spike-in reads (as a fraction of all aligned reads) are calculated for each sample, with the results collected in a single csv file (overlap_report.csv). Since the accuracy of normalization can be difficult when the percentage of spike-in reads is too low, the pipeline is designed to automatically exclude any samples with a calculated percentage of spike-in reads below 0.5%. This functionality can be overridden, forcing the pipeline to attempt normalization for all samples, by setting the “--override_spikeinfail” parameter to true. Conversely, PerCell normalization can be disabled (e.g., for analysis of ChIP-seq data lacking spike-in) by setting the parameter “--skip_downsample” to true.
If PerCell normalization is enabled, scaling factors are calculated for each sample (scaling_factors.csv) based on the relative number of properly aligned spike-in reads. These factors are used to randomly (via the “--seed” parameter) downsample reads, with the now normalized alignment files stored in the downsampled directory. Each immunoprecipitated sample is matched with its corresponding control, and peaks are called and bigWig files generated using a customized script invoking subcommands from MACS2.37 This was necessary to avoid MACS2’s default callpeak command and its built-in normalization (based on library size) from obscuring our own (based on spike-in). Each sample’s immunoprecipitated/control pairing is used to score potential peaks using the selected method (“--macs2_peak_method”) and stored in the macs2_subcommand directory along with narrowPeak files containing peaks passing the cutoff given in the “--macs2_cutoff” parameter. Bedgraphs for each pair of immunoprecipitated and control samples are next compared using MACS2’s bdgcmp command based on the “--macs2_bigwig_method” parameter, before conversion into the visualizable bigWig format using tools from the UCSC software suite.38
Optionally, setting “--skip_idr” to false will use the IDR framework41 to identify statistically consistent peaks across two replicates (output in the idr directory) based on the chosen “--idr_cutoff” value. As an alternative, e.g., when a more stringent MACS2 peak calling cutoff is used, consensus peaks can instead be identified using BEDTools34 by setting “--skip_consensus” to false. Finally, the HOMER software suite39 can be used to annotate peaksets via the “--skip_annotation” and identify enriched motifs via “--skip_motif” parameters, with the output found within their respective subdirectories inside the homer directory.
Quantification and statistical analysis
All ChIP-seq samples were analyzed using the PerCell pipeline (https://github.com/lextallan/PerCell; https://doi.org/10.5281/zenodo.12730196). Reference genomes for experimental and spike-in species include hg38 (human), mm10 (mouse), danRer11 (zebrafish), and dm6 (fly). All analyses were done using default pipeline parameters, which follow. Reads were aligned using Bowtie232 and filtered with parameters “-F 2308 -q 30” by SAMtools.35 Reads were further filtered using Picard36 using the parameter “--REMOVE_DUPLICATES true.” For PerCell normalized samples, scaling factors were calculated by the pipeline based on the relative number of reads properly aligning to the spike-in genome. Read counts for each sample to be normalized were divided by the minimum, yielding a normalization factor ≤1 which was used to downsample properly aligned experimental reads using SAMtools’ “—subsample” option. Peaks were called using the pipeline’s custom MACS237 implementation with “-narrow_peak” option enabled and a cutoff value of 1.301 (p-value of 0.05). Peaks were called and bigwig files generated using the built-in Poisson distribution based “ppois” method.
ChIP-seq signal shown in profile plots (Figures 3A and 3B, 4C, 4D, 5A, and 5B) was calculated by taking the mean number of reads at each set of genomic regions both with and without PerCell normalization, as indicated by respective figure legends. Profile plots and signal tracks were generated using Trackplot.44,45
Published: May 19, 2025
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.crmeth.2025.101052.
Contributor Information
Meng Wang, Email: meng.wang@nationwidechildrens.org.
Genevieve C. Kendall, Email: genevieve.kendall@nationwidechildrens.org.
Benjamin Z. Stanton, Email: benjamin.stanton@nationwidechildrens.org.
Supplemental information
References
- 1.Roh T.-Y., Ngau W.C., Cui K., Landsman D., Zhao K. High-resolution genome-wide mapping of histone modifications. Nat. Biotechnol. 2004;22:1013–1016. doi: 10.1038/nbt990. [DOI] [PubMed] [Google Scholar]
- 2.Barski A., Cuddapah S., Cui K., Roh T.-Y., Schones D.E., Wang Z., Wei G., Chepelev I., Zhao K. High-Resolution Profiling of Histone Methylations in the Human Genome. Cell. 2007;129:823–837. doi: 10.1016/j.cell.2007.05.009. [DOI] [PubMed] [Google Scholar]
- 3.Mikkelsen T.S., Ku M., Jaffe D.B., Issac B., Lieberman E., Giannoukos G., Alvarez P., Brockman W., Kim T.-K., Koche R.P., et al. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells. Nature. 2007;448:553–560. doi: 10.1038/nature06008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Johnson D.S., Mortazavi A., Myers R.M., Wold B. Genome-wide mapping of in vivo protein-DNA interactions. Science. 2007;316:1497–1502. doi: 10.1126/science.1141319. [DOI] [PubMed] [Google Scholar]
- 5.Skene P.J., Henikoff S. An efficient targeted nuclease strategy for high-resolution mapping of DNA binding sites. Elife. 2017;6 doi: 10.7554/eLife.21856. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Kaya-Okur H.S., Wu S.J., Codomo C.A., Pledger E.S., Bryson T.D., Henikoff J.G., Ahmad K., Henikoff S. CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat. Commun. 2019;10:1930–2010. doi: 10.1038/s41467-019-09982-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Saito A., Yamashita T., Mariko Y., Nosaka Y., Tsuchiya K., Ando T., Suzuki T., Tsuruo T., Nakanishi O. A synthetic inhibitor of histone deacetylase, MS-27-275, with marked in vivo antitumor activity against human tumors. Proc. Natl. Acad. Sci. USA. 1999;96:4592–4597. doi: 10.1073/pnas.96.8.4592. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Rosato R.R., Almenara J.A., Grant S. The histone deacetylase inhibitor MS-275 promotes differentiation or apoptosis in human leukemia cells through a process regulated by generation of reactive oxygen species and induction of p21CIP1/WAF1 1. Cancer Res. 2003;63:3637–3645. [PubMed] [Google Scholar]
- 9.Lasko L.M., Jakob C.G., Edalji R.P., Qiu W., Montgomery D., Digiammarino E.L., Hansen T.M., Risi R.M., Frey R., Manaves V., et al. Discovery of a selective catalytic p300/CBP inhibitor that targets lineage-specific tumours. Nature. 2017;550:128–132. doi: 10.1038/nature24028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Hu B., Petela N., Kurze A., Chan K.-L., Chapard C., Nasmyth K. Biological chromodynamics: a general method for measuring protein occupancy across the genome by calibrating ChIP-seq. Nucleic Acids Res. 2015;43:e132. doi: 10.1093/nar/gkv670. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Orlando D.A., Chen M.W., Brown V.E., Solanki S., Choi Y.J., Olson E.R., Fritz C.C., Bradner J.E., Guenther M.G. Quantitative ChIP-Seq normalization reveals global modulation of the epigenome. Cell Rep. 2014;9:1163–1170. doi: 10.1016/j.celrep.2014.10.018. [DOI] [PubMed] [Google Scholar]
- 12.Bonhoure N., Bounova G., Bernasconi D., Praz V., Lammers F., Canella D., Willis I.M., Herr W., Hernandez N., Delorenzi M., CycliX Consortium Quantifying ChIP-seq data: a spiking method providing an internal reference for sample-to-sample normalization. Genome Res. 2014;24:1157–1168. doi: 10.1101/gr.168260.113. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Egan B., Yuan C.-C., Craske M.L., Labhart P., Guler G.D., Arnott D., Maile T.M., Busby J., Henry C., Kelly T.K., et al. An Alternative Approach to ChIP-Seq Normalization Enables Detection of Genome-Wide Changes in Histone H3 Lysine 27 Trimethylation upon EZH2 Inhibition. PLoS One. 2016;11 doi: 10.1371/journal.pone.0166438. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Guertin M.J., Cullen A.E., Markowetz F., Holding A.N. Parallel factor ChIP provides essential internal control for quantitative differential ChIP-seq. Nucleic Acids Res. 2018;46:e75. doi: 10.1093/nar/gky252. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Blanco E., Di Croce L., Aranda S. SpikChIP: a novel computational methodology to compare multiple ChIP-seq using spike-in chromatin. NAR Genom. Bioinform. 2021;3 doi: 10.1093/nargab/lqab064. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Blanco E., Ballaré C., Di Croce L., Aranda S. Quantitative comparison of multiple chromatin immunoprecipitation-sequencing (ChIP-seq) experiments with spikChIP. Methods Mol. Biol. 2023;2624:55–72. doi: 10.1007/978-1-0716-2962-8_5. [DOI] [PubMed] [Google Scholar]
- 17.Bressan D., Fernández-Pérez D., Romanel A., Chiacchiera F. SpikeFlow: automated and flexible analysis of ChIP-Seq data with spike-in control. NAR Genom. Bioinform. 2024;6 doi: 10.1093/nargab/lqae118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Di Tommaso P., Chatzou M., Floden E.W., Barja P.P., Palumbo E., Notredame C. Nextflow enables reproducible computational workflows. Nat. Biotechnol. 2017;35:316–319. doi: 10.1038/nbt.3820. [DOI] [PubMed] [Google Scholar]
- 19.Ewels P.A., Peltzer A., Fillinger S., Patel H., Alneberg J., Wilm A., Garcia M.U., Di Tommaso P., Nahnsen S. The nf-core framework for community-curated bioinformatics pipelines. Nat. Biotechnol. 2020;38:276–278. doi: 10.1038/s41587-020-0439-x. [DOI] [PubMed] [Google Scholar]
- 20.Kucinski J., Tallan A., Taslim C., Wang M., Cannon M.V., Silvius K.M., Stanton B.Z., Kendall G.C. Rhabdomyosarcoma fusion oncoprotein initially pioneers a neural signature in vivo. bioRxiv. 2024 doi: 10.1101/2024.07.12.603270. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kendall G.C., Watson S., Xu L., LaVigne C.A., Murchison W., Rakheja D., Skapek S.X., Tirode F., Delattre O., Amatruda J.F. PAX3-FOXO1 transgenic zebrafish models identify HES3 as a mediator of rhabdomyosarcoma tumorigenesis. Elife. 2018;7 doi: 10.7554/eLife.33800. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Sunkel B.D., Wang M., LaHaye S., Kelly B.J., Fitch J.R., Barr F.G., White P., Stanton B.Z. Evidence of pioneer factor activity of an oncogenic fusion transcription factor. iScience. 2021;24 doi: 10.1016/j.isci.2021.102867. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Wang M., Sreenivas P., Sunkel B.D., Wang L., Ignatius M., Stanton B.Z. The 3D chromatin landscape of rhabdomyosarcoma. NAR Cancer. 2023;5 doi: 10.1093/narcan/zcad028. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Leichsenring M., Maes J., Mössner R., Driever W., Onichtchouk D. Pou5f1 Transcription Factor Controls Zygotic Gene Activation In Vertebrates. Science. 2013;341:1005–1009. doi: 10.1126/science.1242527. [DOI] [PubMed] [Google Scholar]
- 25.Pálfy M., Schulze G., Valen E., Vastenhouw N.L. Chromatin accessibility established by Pou5f3, Sox19b and Nanog primes genes for activity during zebrafish genome activation. PLoS Genet. 2020;16 doi: 10.1371/journal.pgen.1008546. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Girish V., Lakhani A.A., Thompson S.L., Scaduto C.M., Brown L.M., Hagenson R.A., Sausville E.L., Mendelson B.E., Kandikuppa P.K., Lukow D.A., et al. Oncogene-like addiction to aneuploidy in human cancers. Science. 2023;381 doi: 10.1126/science.adg4521. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Rao S.S.P., Huntley M.H., Durand N.C., Stamenova E.K., Bochkov I.D., Robinson J.T., Sanborn A.L., Machol I., Omer A.D., Lander E.S., Aiden E.L. A 3D Map of the Human Genome at Kilobase Resolution Reveals Principles of Chromatin Looping. Cell. 2014;159:1665–1680. doi: 10.1016/j.cell.2014.11.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Gryder B.E., Khan J., Stanton B.Z. Measurement of differential chromatin interactions with absolute quantification of architecture (AQuA-HiChIP) Nat. Protoc. 2020;15:1209–1236. doi: 10.1038/s41596-019-0285-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Kurtzer G.M., Sochat V., Bauer M.W. Singularity: Scientific containers for mobility of compute. PLoS One. 2017;12 doi: 10.1371/journal.pone.0177459. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet. j. 2011;17:10. [Google Scholar]
- 31.Andrews, S. (2010). FastQC A quality control tool for high throughput sequence data. http://www.bioinformatics.babraham.ac.uk/projects/fastqc/.
- 32.Langmead B., Salzberg S.L. Fast gapped-read alignment with Bowtie 2. Nat. Methods. 2012;9:357–359. doi: 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Zhang H., Song L., Wang X., Cheng H., Wang C., Meyer C.A., Liu T., Tang M., Aluru S., Yue F., et al. Fast alignment and preprocessing of chromatin profiles with Chromap. Nat. Commun. 2021;12:6566. doi: 10.1038/s41467-021-26865-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Quinlan A.R., Hall I.M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–842. doi: 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Danecek P., Bonfield J.K., Liddle J., Marshall J., Ohan V., Pollard M.O., Whitwham A., Keane T., McCarthy S.A., Davies R.M., Li H. Twelve years of SAMtools and BCFtools. GigaScience. 2021;10 doi: 10.1093/gigascience/giab008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Broad Institute. Picard toolkit (2019). https://broadinstitute.github.io/picard/.
- 37.Zhang Y., Liu T., Meyer C.A., Eeckhoute J., Johnson D.S., Bernstein B.E., Nusbaum C., Myers R.M., Brown M., Li W., Liu X.S. Model-based analysis of ChIP-Seq (MACS) Genome Biol. 2008;9 doi: 10.1186/gb-2008-9-9-r137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Kent W.J., Zweig A.S., Barber G., Hinrichs A.S., Karolchik D. BigWig and BigBed: enabling browsing of large distributed datasets. Bioinformatics. 2010;26:2204–2207. doi: 10.1093/bioinformatics/btq351. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Heinz S., Benner C., Spann N., Bertolino E., Lin Y.C., Laslo P., Cheng J.X., Murre C., Singh H., Glass C.K. Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol. Cell. 2010;38:576–589. doi: 10.1016/j.molcel.2010.05.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Ewels P., Magnusson M., Lundin S., Käller M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics. 2016;32:3047–3048. doi: 10.1093/bioinformatics/btw354. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Li Q., Brown J.B., Huang H., Bickel P.J. Measuring reproducibility of high-throughput experiments. aoas. 2011;5:1752–1779. [Google Scholar]
- 42.Kidder B.L., Zhao K. Efficient library preparation for next-generation sequencing analysis of genome-wide epigenetic and transcriptional landscapes in embryonic stem cells. Methods Mol. Biol. 2014;1150:3–20. doi: 10.1007/978-1-4939-0512-6_1. [DOI] [PubMed] [Google Scholar]
- 43.Shapiro D.N., Sublett J.E., Li B., Downing J.R., Naeve C.W. Fusion of PAX3 to a member of the forkhead family of transcription factors in human alveolar rhabdomyosarcoma. Cancer Res. 1993;53:5108–5112. [PubMed] [Google Scholar]
- 44.Mayakonda A., Westermann F. Trackplot: a fast and lightweight R script for epigenomic enrichment plots. Bioinform. Adv. 2024;4 doi: 10.1093/bioadv/vbae031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Pohl A., Beato M. bwtool: a tool for bigWig files. Bioinformatics. 2014;30:1618–1619. doi: 10.1093/bioinformatics/btu056. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
-
•
All ChIP-seq datasets generated with this protocol and presented here were deposited in the NCBI Gene Expression Omnibus (http://www.ncbi.nlm.nih.gov/geo/; GEO: GSE271986 and GSE270319).
-
•
All code for the PerCell analysis pipeline is publicly available on GitHub (https://github.com/lextallan/PerCell and https://doi.org/10.5281/zenodo.12730196). Downloading the pipeline is not necessary before running it, as Nextflow will automatically download it during the first execution (https://www.nextflow.io/docs/latest/sharing.html#how-it-works).
-
•
Any additional information needed to re-analyze the results reported in this paper is available from the lead contact upon request.





