Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2015 Mar 2.
Published in final edited form as: Quant Biol. 2013 Mar 1;1(1):54–70. doi: 10.1007/s40484-013-0006-2

Computational methodology for ChIP-seq analysis

Hyunjin Shin 1, Tao Liu 1, Xikun Duan 2, Yong Zhang 2, X Shirley Liu 1,*
PMCID: PMC4346130  NIHMSID: NIHMS619734  PMID: 25741452

Abstract

Chromatin immunoprecipitation coupled with massive parallel sequencing (ChIP-seq) is a powerful technology to identify the genome-wide locations of DNA binding proteins such as transcription factors or modified histones. As more and more experimental laboratories are adopting ChIP-seq to unravel the transcriptional and epigenetic regulatory mechanisms, computational analyses of ChIP-seq also become increasingly comprehensive and sophisticated. In this article, we review current computational methodology for ChIP-seq analysis, recommend useful algorithms and workflows, and introduce quality control measures at different analytical steps. We also discuss how ChIP-seq could be integrated with other types of genomic assays, such as gene expression profiling and genome-wide association studies, to provide a more comprehensive view of gene regulatory mechanisms in important physiological and pathological processes.

INTRODUCTION

Recent advances in next generation sequencing (NGS) technologies have enabled scientists to investigate a variety of molecular events occurring on the genome with high resolution and accuracy [14]. These NGS technologies have been applied to many scientific and clinical areas including detection of genetic variations (e.g., SNP calling) and quantification of RNA transcripts (e.g., RNA-seq). One of the most successful NGS applications is chromatin immunoprecipitation (ChIP) accompanied by NGS, or ChIP-seq, which can map the in vivo genome-wide binding sites of DNA-binding proteins such as transcription factors (TFs) or modified histones.

In chromatin immunoprecipitation, cells are lysed [5] and protein-DNA interactions are crosslinked to form covalent bonds by formaldehyde or other chemical reagents. Then the crosslinked DNA is sheared by sonication or DNA-cutting enzymes (e.g., micrococcal nuclease, often called MNase) into 150–500 bp-long fragments. Those DNA fragments crosslinked with the DNA-binding factor of interest are immunoprecipitated using an antibody specific to the factor. ChIP can be applied to a wide range of DNA binding factors, including TFs, transcription co-activators, co-repressors, chromatin regulators, and modified histones. After reverse crosslinking the protein-DNA complexes, the pulled-down DNA fragments are PCR amplified and then subjected to massively parallel sequencing (see Metzker’s review for details of various NGS technologies) [1]. Finally, when the resulting ChIP-seq reads are mapped back to the genome, the locations of the factor-DNA interactions can be identified.

ChIP-seq can provide important insights towards gene regulatory process particularly in combination with transcriptomic profiles from expression microarrays or RNA-seq, since ChIP-seq can help identify genes directly regulated by the factor (see the section Integrate with gene expression profiles for details). For instance, ChIP-seq and its predecessor ChIP-chip (i.e., ChIP coupled with tiling microarray technologies) have been used to study how nuclear hormone receptors (e.g., androgen receptor or estrogen receptor) and their cofactors (e.g., FoxA1) cooperate to regulate gene expression in prostate and breast cancers [6,7]. ChIP-seq has also been employed in many stem cell studies to associate the core regulatory circuitries of stem cell TFs such as Oct4, Nanog or cMyc with cell fate [812]. These transcriptional regulation studies have led to the recent development of techniques to capture the higher-order chromatin interactions [9,1322], which provide clues as to functional interactions between TF binding sites and the promoters of their target genes.

Likewise, since the pioneering studies of histone mark ChIP-seq in the human CD4+ T-cells [23,24], researchers have actively adopted ChIP-seq to investigate the biological functions of many histone marks. The ENCODE and modENCODE consortia, for example, conducted ChIP-seq of important histone marks in many cell states in human, worm, and fly [25,26], and used statistical modelling to annotate different chromatin states by the combinatorial patterns of different histone marks [2729]. These studies help identify previously unannotated functional elements in the genomes.

A number of algorithms for ChIP-seq analysis have been developed. In this review, we will provide an overview of the analytical workflow for ChIP-seq and summarize the key concepts and challenges for each step. In particular, we will introduce useful quality control strategies at different analytical steps, which are useful for data interpretation. We will also discuss methods to integrate ChIP-seq with other types of high-throughput data for more comprehensive understanding of important physiological or pathological processes.

CONSIDERATIONS OF ChIP-seq EXPERIMENTAL DESIGN

The quality of ChIP-seq critically depends on the sensitivity and specificity of the antibody for a DNA-binding factor. Specific antibodies give strong and clean binding enrichment information, while weak and non-specific antibodies have increased background noise. Another important issue related to experimental design is the use of control experiments to adjust the bias caused by chromatin accessibility [30]. In most cases, DNA from chromatin input (i.e., chromatin sample before IP), mock immunoprecipitation (IP) (i.e., IP without antibody) or non-specific IP (e.g., IP against immunoglobulin G) has often been chosen as a control sample. For more detailed topics on ChIP-seq experimental design, the advantage and disadvantage of different types of control, we refer the readers to a recent review on ChIP-seq [31].

Advances in NGS technologies enable ChIP-seq to be conducted at greater genome coverage at lower price [3], and recover weaker binding events. Saturation of ChIP-seq depends on the nature of DNA-binding factor, antibody sensitivity, and specific research focus. It can be evaluated by sub-sampling total sequencing reads, and computing the recovery rate of ChIP-seq peaks [31]. In general, a deeper sequencing is recommended for factors with diffuse binding patterns or repressive functions than those with sharp binding patterns or active functions. It is also important to sequence the IP and control at comparable depth to allow unbiased peak calling. Along with the improvement of NGS technologies, the development of methods for multiplexing samples by barcoding (i.e., multiple samples are processed at a single lane) can increase efficiency without excessive increase in time or cost [3235].

ChIP-seq DATA ANALYSIS WORKFLOW

A number of computational and statistical tools have been proposed and developed for addressing specific aspects of ChIP-seq analysis. We will describe them in the following subsections. In addition, we would like to highlight integrative analysis platforms such as Cistrome [36] and CisGenome [37,38], which provide comprehensive work environments for users to conduct most of the necessary analyses in one place. Independent algorithms and tools for specific purposes will be introduced in the context of the analytical workflow for ChIP-seq within each of the subsections. It is also important to develop quality control (QC) standards at each analytical step, so we have included suggestions on the potentially useful QC strategies.

Reads mapping

Raw data from NGS platform often appear in fastq format, containing short DNA sequence and quality scores. In general, the first step of ChIP-seq analysis starts with mapping these raw reads to the reference genome. Several algorithms have been developed to quickly map millions to hundreds of millions of short sequencing reads. The popular ones include ELAND (Illumina®), Bowtie [39,40], BWA [41,42], MAQ [43], Stampy [44], Novoalign [45], and SOAP2 [46]. Among them, Bowtie, BWA, and SOAP2 adopt the Burrows-Wheeler transformation, which was originally developed as a data compression technique in the 1990s [39,41]. Choosing appropriate mapping software depends on sequencing platform, speed requirement, and hardware resources. For example, if data come from the SOLiD platform, tools supporting colorspace are the ideal choices, such as BWA. If speed is the major consideration, Bowtie is preferred. For further reading, See Bao et al.’s review paper for more detailed comparisons of the above aligners [47].

Quality control (QC) can be conducted for reads mapping. In order to simplify analysis, usually only reads mapped to one unique location in the genome (called uniquely mapped reads) with minimum allowed mismatches (e.g., up to two mismatches) are kept for downstream analysis. In case a ChIP-seq read is mapped to multiple locations on the genome, a general solution is to randomly assign one of the locations to it. Often, the ratio of the number of uniquely mapped reads over the total number of ChIP-seq reads can be an assessment of library quality. From statistics of publicly available ChIP-seq datasets, ratios of over 50% suggest good library quality. Another useful QC measure is the number of redundant reads that are mapped to the same genomic coordinates, because high redundancy rate suggests PCR amplification bias from limited ChIP material. The ratio of redundant reads over all mapped reads should ideally be below 50%.

Peak modeling and identification

The sequence reads mapped to the genome are subject to peak calling to detect regions with significant enrichment of ChIP signals with respect to the background (e.g., control if available). For most of ChIP-seq experiments with single-end sequencing, DNA fragments are sequenced from the 5′ ends; as a result, bimodal distributions, surrounding the true binding site, are formed from reads mapped on the + and − strands respectively (red and blue curves in Figure A). Therefore, to precisely detect the correct binding site, some peak callers empirically model the distance between the + and − strand modes [48,49], and extend the tags towards their 3′ direction by the estimated distance (Figure B). Then, the pile of the extended tags forms a peak (Figure B) and its summit represents the most probable binding location. One QC measure at this step is the ability of peak callers to properly model the +/− mode distance, and failing to model the distance suggests potential biases in the sonication and library construction steps, or that the factor of interest binds to diffuse regions rather than point sources.

Next, peak callers calculate the statistical significance (e.g., p value) of the enrichment level of ChIP signals in selected regions comparing to a background model. Due to the discrete nature of NGS data, most of the peak callers adopt the binomial [50], Poisson [48,51,52], negative binomial (almost equivalent to the Poisson based on local average) [37,53] distributions or simulation-based modelling [49,54] to compute statistical significance of the ChIP enrichment over background. A few widely-used peak callers include MACS [48], Sissr [55], SPP [49], and USeq [50]. In our opinion, different distributions are same in essence. For example, negative binomial is a generalized Poisson, and dynamic Poisson based on empirical local lambda is a more generalized version of negative binomial, etc. So they provide similar sensitivity and specificity. But different peak callers have their own rationality, and are different in positional accuracy of predicted binding sites. More detailed description and comparison of different peak callers are available in separate studies [31,5658].

Detected peaks also need to be checked in terms of quality, where false discovery rate (FDR) or fold change against the background (e.g., control tag counts in the same region) is often used. FDR is defined as the expected proportion of false positive peaks in a list of detected peaks [5961]. Several peak callers provide empirical [48,50,54,6264] or model-estimated [37,38,51,52] FDR or q-value (minimum FDR at a given p value cut-off) for each peak, and 5% FDR is the most commonly accepted value for peaks of good quality. The empirical FDR can be calculated as the number of control peaks passing certain cut-off divided by the number of ChIP-seq peaks passing the same cut-off. Model-estimated FDR can be computed through permutation or random sampling. Fold change, the ratio of tag counts between IP and control in the peak region, is also an intuitive measure of peak quality. A fold change of 5 is generally recommended as a reasonable cut-off, and an enough number (e.g., >50%) of peaks with over 20-fold is an indicator of good ChIP-enrichment.

In addition to the above mentioned peak callers, there are also peak callers with more specialized functions, such as calling positioned nucleosomes from nucleosome-resolution histone mark ChIP-seq (i.e., MNase-digested fragments) [65], identifying diffuse regions enriched by a broad mark [51], combining multiple ChIP-chip or ChIP-seq sets for the same factor for consensus peak calling [66]. As more and more ChIP-seq data (or existing ChIP-chip data) become available, these peak callers with special functions will become more useful for data integration.

Visualization tools

Most people view their ChIP-seq data either as signal profiles or as called peaks on a genome browser. The most widely used visualization environment is the University of California Santa Cruz (UCSC) genome browser (http://genome.ucsc.edu) [67] (Figure ). In addition to the standard browser functions, this web-based application also provides other important genomic information, including tracks for gene annotation (e.g., refseq or UCSC known genes), evolutionary conservation, annotated SNPs [68,69], and data from NIH funded genomics consortia such as ENCODE [70].

As a web server, UCSC genome browser has limitations in response speed, which are mostly related to the process capability of the server and Internet connection. In contrast, stand alone genome browsers such as IGV (http://www.broadinstitute.org/software/igv/home) [71] and IGB (http://www.bioviz.org/igb/) [72] can efficiently display large data sets and enable the user to visualize genomic datasets from either a local computer or a remote date warehouse and navigate quickly at multiple scales. It also supports the simultaneous display of other types of genomics data tracks such as aligned NGS reads, mutations, copy numbers, gene expression, DNA methylation, and gene annotations as well [71]. Other NGS genome browsers are also available and we refer readers to their individual publications [37,38,7379].

Use of replicates

Biological replicates can help identify more confident peaks for ChIP-seq experiments. Intuitively, successful ChIP-enrichments should be consistent across biological replicates. Replicates should have similar signal profiles and peak regions. One measure of consistency is the percentage (e.g., > 50%) of overlapping peaks between two replicates, which can be easily visualized using a Venn diagram (Figure A). Another measure is the correlation coefficient (e.g., over 0.6 between replicates) of ChIP-seq signals over selected genomic intervals (e.g., every 1 kb or 2 kb) or over the union peak regions of the replicates (Figure B).

The ENCODE/modENCODE consortia proposed a third statistical measure of replicate consistency called Irreproducible Discovery Rate (IDR). IDR measures the proportion of inconsistent peaks over the consistent peaks across replicates at certain threshold [80], and also considers whether the peak ranks (based on p value or FDR) in the replicates are correlated or not. Thus, IDR informs not only the significance of individual peaks but also the consistency between two replicates. As the decreasing cost of NGS allows increasing number of ChIP-seq experiments to have replicates, IDR will be useful to provide more robust peak calls from the replicates. A typical threshold of IDR is 1% [80].

Use of other prior genomic information for quality assurance

Since the genomic era began in the late 1990s, increasing amount of knowledge, including evolutionary conservation, TF binding motifs, and gene annotations, has been accumulated in the public domain. In this section, we will discuss how to use such knowledge to assess the quality of ChIP-seq data.

Use of evolutionary conservation

Cis-regulatory elements that harbour TF binding sites are in general under more evolutionary constraint. Therefore, most identified peaks in a successful ChIP-seq experiment show more evolutionary conservation than background sequences in the genome. PhastCons scores [81] (e.g., that downloaded from UCSC genome browser) provide precomputed evolutionary conservation score at every base in a reference genome. When aligning good transcription factor ChIP-seq peak at the peak summit or center, the average PhastCons conservation scores near the summit or center are typically higher than the surrounding regions (Figure 4A).

Figure 4. Useful QC measures for ChIP-seq peak calling.

Figure 4

(A) The average PhastCons score at TF binding sites can be used to assess the quality of ChIP-seq peak calls. The plot indicates that the center of NF-κB binding sites is more evolutionarily conserved than the background.

(B) If ChIP-seq is successful and the factor of interest has specific DNA binding motifs, the motifs should be significantly enriched near the summits of detected binding sites. This example shows that a NF-κB binding motif registered in JASPAR database was found at NF-κB binding sites.

(C) The pie chart visualizes the distribution of H3K36me3 peaks over different categories of elements such as promoter, UTRs, coding exon, and intron. Since H3K36me3 is associated with transcriptional elongation, its peaks are primarily present in exons (i.e., 9.6% in the pie chart) and introns (i.e., 76.7%).

(D) This trend was also observed using the meta-gene plot of H3K36me3 ChIP enrichment. Every gene was normalized to have the same length of 3 kb and then the average ChIP signal was profiled on the meta-gene including 1 kb upstream and downstream of TSS and TTS. The red and purple lines represent the average ChIP enrichments of H3K36me3 on highly expressed (top 10%, red) and lowly expressed (bottom 10%, purple) genes, respectively, which shows that H3K36me3 is positively correlated with gene expression levels.

Use of TF DNA binding motifs

TFs bind to DNA in a sequence-specific manner, so peaks from successful TF ChIP-seq should exhibit significant enrichment of the TF binding motif. Binding motifs can be represented as a position weight matrix or virtualized as a sequence logo, both of which indicate the nucleotide preference of the factor at each motif position (Figure 4B). Currently, several databases of known TF binding motifs are publicly available [8284]. In addition, many experimental and computational groups continue to derive better or previously uncharacterized TF motifs from new genomic technologies or data [85,86]. Enriched sequence motifs can be identified from either de novo methods or known motif scanning [8790] at TF ChIP-seq peaks. In addition, a motif could be identified with better confidence if its occurrences are more frequently at the peak summits or centers than at the surrounding regions [91]. A common procedure for motif finding is to focus on top ranked ChIP-seq peaks (typically top 1000 ranked by p value) in order to avoid noises from weak binding sites.

Use of DNase I hypersensitive sites and HOT regions

Deoxyribonuclease (DNase) I hypersensitive sites are chromatin regions that are more susceptible to cleavage by this DNA cutting enzyme than other regions. DNase I hypersensitivity often indicates DNA accessibility associated with a local reduction in nucleosome occupancy [9294]. DNase I hypersensitive sites are broadly enriched in bodies of highly expressed genes, and sharply enriched in functional regulatory sequences such as promoters and enhancers, which are targeted by many TFs. Recently, the ENCODE and Roadmap Epigenomics consortia released the DNase-seq data of many human and mouse cell lines and tissues [25,95]. These data provide a comprehensive repertoire of the locations of TFs, chromatin factors, and histone marks [25,95]. Peaks of a successful ChIP-seq experiment overlap over 80% with the DNase-seq peaks, and this could be another evidence of good data quality.

In published ChIP-seq studies, some regions called High-Occupancy Targets (HOT) are constantly enriched in almost all ChIP experiments. They often lack the binding motif of the factor of interest, and might be caused by protein-protein interactions of unknown factors with the factor of interest or experimental artefacts during the ChIP process [9698]. It is advisable to filter the HOT regions before downstream analysis, and high percentage of ChIP peaks in the HOT regions raises a red flag on data quality.

Use of the statistics of peak distribution

TF binding sites and modified histone regions are generally enriched for cis-regulatory elements. For example, many TFs bind more frequently in promoters than the rest of the genome (Figures 4C and 4D) and histone marks such as H3K36me3 are enriched over the exons of actively transcribed genes. Although these properties vary widely based on the factor or modification being tested, to test the consistency of a priori knowledge and observed enrichment patterns can be used to evaluate ChIP-seq data quality. Several software packages are available to provide summary statistics on the distribution of ChIP-seq peaks or visualize average ChIP-seq enrichment signals on annotated cis-regulatory elements [38,99].

Differential peak identification

Transcriptional and epigenetic regulation studies often need to identify differential binding of a factor between two or more biological conditions. For example, a recent work by Wang et al. showed androgen receptor (AR) binding changes when its co-factor FoxA1 was knocked-down in the prostate cell line LNCaP [100]. One simple method is to run peak calling in separate conditions followed by intersection analysis to identify unique peaks for each condition, but it might miscall a region when it is barely above and below the peak calling cut-off in the respective conditions. Another method is to conduct peak calling between the two ChIP-seq conditions treating one as the control [101], but it might erroneously detect a region that is weak in both conditions but significantly enriched in one condition over the other. A more specialized algorithm uses a Hidden Markov Model to detect differential binding through probabilistic modelling of the ChIP-seq profiles in two conditions [102], but the method has limited resolution. Next significant version of MACS: MACS2 (https://github.com/taoliu/MACS/) that attempts to address all these issues based on four paired treatment and control samples is currently available under continuous testing and improvement. There are some existing algorithms initially designed for differential expression studies on RNA-seq data, such as edgeR [103], DESeq [104], and bayseq [105], which can be modified and applied to ChIP-seq as well. As more ChIP-seq data over multiple conditions become available, the increasing importance of differential peak calling will accelerate the effort to develop better algorithms.

INTEGRATIVE ANALYSIS WITH OTHER GENOMIC DATA

Integrate with gene expression profiles

ChIP-seq data of TFs, chromatin factors, and histone marks often need to be interpreted in the context of gene regulation, which requires predicting the target genes that are regulated by the factor and its binding sites. The methods for determining the potential target genes of a factor generally depend on the factor type and binding patterns. For example, most TFs and some histone marks are characterized as sharp binding patterns on either proximal or distal regions to transcription start sites (TSSs). In such case, the distance between each binding site to its nearest gene tends to show a strong association with the expression level or dynamics in expression of the gene. Please keep in mind that, although this strategy works in many cases, it neglects the fact that there would be long-range interactions between cis-regulatory elements and their target genes. With more understanding on high-level chromatin structure, we would revise target gene prediction method extensively. On the other hand, for some other histone marks with pervasive enrichment over a region, the coverage as well as the enrichment level of the mark on the promoter, exons, or body of a gene are important variables for determining whether the gene is a target.

One important subject of integrative analysis of ChIP-seq with expression profiles is to infer whether a given factor mainly functions as a transcriptional activator or a repressor. In order to perform such analysis, expression profiling should be conducted where the activity of the factor is perturbed (e.g., through hormone activation or siRNA knockdown). If a histone mark is studied, gene expression can be profiledin the same way by perturbing the activity of the modifying enzyme that is associated with the histone mark. Once the target genes of binding are predicted, another useful analysis is to examine their potential biological functions. Before experiments are conducted to validate the function of binding, computational analysis of gene annotation/ontology could provide important clues as to whether the target genes are enriched for specific biological processes or pathways.

Target gene prediction

  1. Target gene prediction for factors with sharp binding patterns

    The effect of TF binding on gene expression often attenuates with the distance to the gene, but could reach hundreds of kb from the target genes [7]. Therefore, distance-based target gene prediction is useful for most TFs and some histone marks with sharp peaks in promoters or enhancers (i.e., distal intergenic or intronic regions). This group of histone marks includes H3K4me1/2/3 (H3K4me3 more enriched at promoters, H3K4me1 more enriched at enhancers, and H3K4me2 enriched at both), H3R17me3, most acetylation marks, and the histone variant H2A.Z. Simple target prediction could be based on the nearest distance (e.g., within 1 kb for promoter binding and within 10 kb for enhancer binding) between factor binding sites and TSSs of genes. If expression profiling with factor activity perturbation is available, limiting genes with significant differential expression between the perturbation conditions could refine the targets. An alternative to using the nearest distance is to count the number of binding sites within a given distance range (e.g., 100 kb) since more binding sites in proximity are thought of as evidence of increased regulatory potential of the factor to the target gene [106]. This approach could be further refined by weighting the different binding sites by their distances to the target gene [107].

  2. Target gene prediction for factors with pervasive binding patterns

    Chromatin factor and histone marks related to transcriptional elongation or repression often show pervasive enrichment patterns over broad regions. For instance, H3K36me3 is broadly present in the exons of expressed genes while H3K27me3 or PRC2 complex tend to enrich at silenced genes [108]. It is informative to stratify genes according to their expression levels and see the relative enrichment of the mark in each of the strata. Figure 4D displays average ChIP signals of H3K36me3 on the meta-gene (by dividing every gene into the same number of bins) at different gene expression levels in human lymphoblastoid cell line GM12878 [23]. Since H3K36me3 is a transcriptional elongation mark, top 10% expressed genes have significantly higher H3K36me3 enrichment than bottom 10% expressed genes. By computing the ChIP-seq coverage over gene body or exons (normalized by the gene or exon length) and selecting genes with the top (e.g., 30%) coverage, one can predict the H3K36me3 associated genes.

Factor function as transcriptional activator or repressor

The role of a factor as a transcriptional activator or repressor can be inferred by the expression changes of target genes when the activity of the factor is perturbed by activation, overexpression, knockdown, or knockout. For example, if genes with reduced expression in the knockdown or knockout of the factor are more likely to harbor the binding sites of the factor in proximity than genes with increased expression, this will be considered as important evidence that the factor acts as a transcriptional activator.

This concept is illustrated in Figure [109,110]. Using distance of binding to all the genes as background, we could see that in the prostate cancer cell line LNCaP, androgen receptor binds much closer to the up-regulated genes than down-regulated and background genes, which implies that it primarily serves as a transcriptional activator. In contrast, in the breast cancer cell line MCF-7, estrogen receptor appears to be equally close to both down- and up-regulated genes. The statistical significance of these observations can be evaluated using a non-parametric test such as the one-side Kolmogorov-Smirnov test.

Target gene annotation based on gene ontology and biological pathways

Researchers are often interested in seeing whether the target genes of a factor binding are enriched for specific pathways or biological processes. Many gene ontology or annotation analysis tools are publicly available, which are summarized at the Gene Ontology website (http://www.geneontology.org/GO.tools.shtml). A few particularly user-friendly tools that are not in the list include DAVID [111,112] and Panther [113], which take gene list as input, and GREAT [114], which takes binding site coordinates as input. GREAT assigns target genes based on gene-binding distance which might have some disadvantages, but it can detect function or pathway enrichment with better sensitivity because it could have a richer meta-annotation gene list [114]. Gene set enrichment analysis (GSEA) also conducts ontological analysis on the target genes, and its unique gene sets provides more extensive annotation searches [115].

Integrate with other TF ChIP-seq data

TF ChIP-seq experiments often reveal novel associations between multiple TFs. Motif discovery has been previously suggested as an important ChIP-seq QC measure, and it can also be used to predict collaborating factors [116]. Crosslinking stabilizes not only the covalent link between TFs and DNA, but also that between interacting TFs, which allows ChIP of one factor to enrich the targets of the collaborating factors. If motif discovery on the peaks of one factor finds not only the known motif of its own, but also a significantly enriched motif of another factor, it suggests that the two factors may interact. Sometimes many factors with the same DNA-binding domain share similar motifs, so to pinpoint the exact factor binding to the collaborating motif requires analysis of expression data. If a factor is highly expressed in the cell and correlated in expression with another factor being ChIPed cross a panel of related tissue samples [117], they may be putative collaborating partners. Protein-protein interaction experiments and prior literature might provide additional insights.

The candidate co-factors selected by the aforementioned methods are often further analyzed by ChIP-seq. The Venn diagram (or Euler diagram in case of more than two factors) is also used to see the potential association or colocalization between different DNA binding factors when the ChIP-seq data of these factors are all available. In many cases, the regions co-occupied by collaborative factors (i.e., the intersection in the Venn diagram) are considered to be more important for understanding the detailed gene regulation event governed by these factors together. There are publicly available software to call such colocalized regions from the peak lists of multiple factors in UCSC BED format [36,118,119]. In addition, heatmaps are very useful to display ChIP signals of different factors across their union of binding sites (e.g., from −1 kb to 1 kb from the binding summit or center). Coupled with clustering methods (e.g., hierarchical or k-means clustering), heatmap analysis can provide better binding site classification than the Venn diagram by grouping binding sites with similar ChIP enrichment patterns across multiple factors (see Figure for more details [120]).

Integrate ChIP-seq data from multiple organisms

Many biomedical experiments are conducted on model organisms, so it is interesting to investigate whether mechanisms discovered for a factor in model organisms such as round worm or fruit fly still hold true in more complex organisms such as human. As for ChIP-seq, this question can be answered by comparing ChIP-seq profiles of orthologous factor in multiple species. Several studies have found that while the binding sites of the factors diverge extensively between species, the target genes of highly conserved TFs are mostly preserved across species [121,122]. However, some recent studies examining ChIP-seq reads mapping to repetitive regions of the genome also found significant rewiring between factor and target gene over evolution through transposable element turnover [123,124]. To accurately detect enrichment in repeat regions, more efforts are needed to develop algorithms that can better utilize multiply mapped sequence reads.

Integrate with epigenetic ChIP-seq data

Since ChIP-seq of histone methylations [23] and acetylations [24] were first profiled in human CD4+ T-cells, increasing efforts to study histone marks in more cell states and organisms have yielded better understanding of epigenetic regulation. These include the mod/ENCODE and Roadmap Epigenomics consortia and individual studies [23,24,27,125]. For example, it has been revealed that many histone modifications delineate important elements such as active or repressed promoters (e.g., H3K4me3 or H3K27me3, respectively [126]), actively transcribed exons (e.g., H3K36me3 [127]), and active enhancers (e.g., H3K27ac [128] and H3K4me1 [129]). Also, histone marks such as H4K20me3 and H3K9me3 indicate transcriptionally silenced state over very broad heterochromatin domains. In addition, although different histone marks have different antibody specificities, many, especially histone acetylations, have similar enrichment characteristics and probably redundant regulatory mechanisms [24]. Several groups have attempted to use machine-learning algorithms such as hidden Markov models to infer combinatorial enrichment patterns of histone marks and other chromatin factors [2729,130,131]. These combinatorial patterns also reduce the redundancy in histone mark profiles. They can be used to segment the whole genome into regions of distinct chromatin signatures, and predict previously unidentified functional elements in the genome.

An interesting follow up is to identify enriched transcription factor motifs on the putative enhancers that are inferred from histone mark profiles. If conducted on different cell types and integrated with gene expression analysis, this could yield important insights to cell-type specific transcription factor activities [29]. Putative transcription factor binding sites can carry enhancer histone marks before active binding of transcription factor. This histone mark pattern often changes upon binding, which is characterized by a displacement of the nucleosome at the binding site and stronger marking and better positioning of the nucleosomes flanking the site. Therefore, motif analysis conducted on the dynamics of histone mark profiles at single-nucleosome resolution can potentially improve the prediction accuracy of transcription factor binding [91,132]. This approach has drawn attention recently as an effective pre-screening measure to find key transcription factors responding to developmental or environmental stimulations [106]. Likewise, in another study, the chromatin state dynamics based on histone mark profiles were used to link enhancers with their target promoters based on correlation estimation. Then, predictions were made for cell-type specific active or repressive TFs through the integration of motif analysis results and gene expression profiles [29].

Integrate with genome variation and disease data

DNA sequence variations such as single nucleotide polymorphisms (SNP) on transcription factor binding sites may influence the binding affinity of the factors. Several studies have attempted to assess the potential impact of allele-specific SNPs or indels on chromatin structure or TF binding sites [133136]. Interestingly, the majority of disease-associated loci discovered from genome-wide association studies (GWAS) are located on introns or distal intergenetic regions. ChIP-seq data of transcription factors or histone marks can provide hints on the mechanisms of these disease SNPs. For example, H3K4me2 ChIP-seq helped identify a tissue-specific enhancer in the cancer risk loci on the human 8q24 region that might regulate Myc expression [135].

Copy number variation (CNV) studies can also be linked with ChIP-seq analysis [137]. For example, while studying the enrichment of TF binding in amplified genomic regions can reveal novel pathological pathways originating from the CNVs, they might also yield false positive peaks calls [137]. Therefore, caution should be taken to reduce the false positive ChIP-seq peak calls, by using CNV information from other sources or proper chromatin input controls.

CONCLUSION

The recent improvements in NGS throughput have dramatically enhanced the dynamic range of ChIP-seq. Illumina® HiSeq 2000, for example, can generate up to 200 Gb of sequence data per run. In addition, technological developments in automation, batch processing, and multiplexing also improve data production efficiency and lower cost. We expect increasing number of laboratories to adopt ChIP-seq for gene regulation studies, and each study to generate ChIP-seq data in multiple factors in more physiological or pathological conditions.

Most published ChIP-seq studies in vertebrate species are conducted on cell lines, since tissue or tumour samples often have heterogeneity and insufficient cell count for ChIP-seq (106 cells recommended). While some groups isolate relatively homogenous cell population from tissue for ChIP-seq [23,24,138,139], others try to improve the ChIP-seq protocol to work on smaller starting material [140,141]. A recent study proposed a method for single-tube linear DNA amplification that can avoid possible artefacts and bias during the amplification process, so ChIP-seq can be conducted on a few thousand cells [142]. Third generation single-molecule sequencing technologies might be an alternative solution to small sample experiments that we look forward to.

In addition to ChIP-seq, other NGS genomic assays that profile genome-wide chromatin states have also been useful to understand gene regulation. A comprehensive and unambiguous set of genomic binding sites for a transcription factor can also be revealed by ChIP-exo [143]. DNA methylation profiles from bisulfite sequencing (BS-Seq) can identify both 5-methylcytosines and 5-hydroxymethylcytosinesat base-pair resolution [144147]. DNase-Seq is a cost effective method for profiling cis-regulatory elements in open chromatin regions. In addition, Hi-C [19] or ChIA-PET [1618] can help infer complex long-range three-dimensional chromatin interactions, for example between promoters and enhancers. Each of the above technique helps provide additional salient view of the chromatin, and they could be integrated with ChIP-seq to more clearly understand gene regulation.

For integrative studies, it is essential to establish computational pipelines that perform meta-analysis of multiple datasets using a variety of algorithms and tools. It is also crucial to integrate researchers’ unpublished data with relevant publicly available genomic data to effectively refine biological hypotheses. There are already public resources such as the UCSC genome browser that provide processed and curated data for such integrative analysis [70].

As more high-throughput assays on gene expression and DNA binding activities become available, there will be increasing need for systems biology approaches to infer the associative or causal relationships between genes or proteins. Most early systems biology efforts focused on inferring network structures of gene regulation based on co-expression patterns from microarray data [148152]. From a systems biology point of view, TF binding and epigenomic data can provide additional information on the linkage and directionality between nodes of transcription factors, chromatin factors, and other genes [153155].

NGS technologies have improved our ability to detect genetic variations in non-coding regions. Also, increasingly improved statistical methods have been proposed to distinguish true variations from sequencing artifacts or errors [156]. As we previously discussed, DNA sequence variations could affect TF binding affinities and nearby gene expression, so ChIP-seq can help to explore the function of GWAS-identified disease loci in non-coding sequences. In addition, the coming years might also see epigenome-wide association studies conducted on tissues of populations to better understand diseases susceptibility or mechanisms.

In conclusion, since its invention, ChIP-seq has become a powerful tool for revealing transcriptional and epigenetic regulation in many cell systems. It still continues to evolve through the development of more advanced sequencing and sample production technologies. For ChIP-seq analysis, many computational and statistical applications have been developed and are now being organized into more comprehensive analytical pipelines. Finally, increasing efforts to integrate ChIP-seq with other types of high-throughput genomic assays will offer a more comprehensive perspective on complex regulatory mechanisms controlling a variety of physiological and pathological processes.

Figure 1. A schematic of peak modelling for ChIP-seq.

Figure 1

(A) The bimodal distributions of + and − sequence reads (red and blue arrows, respectively) surrounding a transcription factor binding center (marked by the yellow vertical arrow). The distance (d) between the summits of the + and − distributions is considered to be an estimate of the length of DNA fragments pulled down by the antibody.

(B) The ChIP enrichment signal can be obtained as the count of the + and − sequence reads extended by the estimated d at every base.

Figure 2. An UCSC genome browser snapshot of ChIP enrichments of NF-κB and H3K79me2 in human lymphoblastoid cell line GM12878 from ENCODE data.

Figure 2

The bottom gene track shows RefSeq genes near the binding sites. While NF-κB shows sharp binding patterns (the 1st green track from the top), H3K79me2 diffuses over a broad region (the 3rd green track). The black horizontal bars represent peaks called by MACS. The control tracks of DNA inputs (the 2nd and 4th green tracks) are used to model the background including the chromatin bias in ChIP-seq. The significance of binding site detection is estimated considering the background from the control; therefore, region A (near chr11 62610000) of NF-κB is not identified as a peak because its background level is relatively high compared to region B (near chr11 62625000) although the ChIP enrichments of these regions are similar.

Figure 3. A sample analysis for two replicates of NF-κB ChIP-seq from ENCODE data.

Figure 3

(A) The Venn diagram of the peak sets identified from the two individual replicates at the same cut-off for peak calling (7281 and 7525 peaks, respectively). A large intersection indicates high consistency between the replicates (e.g., > 50%).

(B) Another consistency measure is to see the pair-wise correlation of ChIP enrichments of the replicates. Each dot represents the average ChIP enrichments of the replicates every 2 kb. Consistent replicates show a tight distribution around the local regression line (red), resulting in a large correlation coefficient (e.g., 0.86 in this example).

Figure 5. A distance-based method for inferring the role of a sharply binding TF as a transcriptional activator or repressor.

Figure 5

(A) Each curve is the cumulative distribution of the distances from genes to the nearest AR binding sites in prostate cancer cell line LNCaP. The red and green colors correspond to the sets of genes that are up- and down-regulated by DHT treatment, respectively. The black dotted line indicates the distance distribution of all genes, which can be used as a background distribution. From this analysis, it is can be seen that AR more directly regulates the up-regulated genes than the down-regulated ones.

(B) A similar analysis was done with breast cancer cell line MCF-7. When it compared to AR (A), ER regulates the up- and down-regulated genes by E2 treatment almost equally.

Figure 6. An example analysis for the ChIP-seq data of two collaborating factors.

Figure 6

(A) The Venn diagram of Rev-erbα and HDAC3 (histone deacetylase 3) binding in mouse liver. This shows that Rev-erbα and HDAC3 are highly colocalized. However, this example is an extreme case for highly collaborating factors. In general, collaborating factor ChIP-Seq peaks tend to show a smaller intersection than that between two replicates of the same factor.

(B) The scatter plot of the ChIP-seq enrichments of the factors at their union binding sites (blue dots) and a regression line (red line). The correlation coefficient of these two factors was also calculated (0.89).

(C) Heatmap analysis for Rev-erbα and HDAC3 binding sites. The heatmaps also confirm the high association of these two factors (left and right). The binding sites with similar binding patterns are grouped using k-means clustering (k=4) and distinguished by yellow horizontal lines. The 1st and 2nd groups reflect strand bias in binding. The individual groups can be further analyzed by associating with differentially expressed genes.

References

  • 1.Metzker ML. Sequencing technologies — the next generation. Nat Rev Genet. 2010;11:31–46. doi: 10.1038/nrg2626. [DOI] [PubMed] [Google Scholar]
  • 2.Ansorge WJ. Next-generation DNA sequencing techniques. New Biotechnol. 2009;25:195–203. doi: 10.1016/j.nbt.2008.12.009. [DOI] [PubMed] [Google Scholar]
  • 3.Kircher M, Heyn P, Kelso J. Addressing challenges in the production and analysis of illumina sequencing data. BMC Genomics. 2011;12:382. doi: 10.1186/1471-2164-12-382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Schuster SC. Next-generation sequencing transforms today’s biology. Nat Methods. 2008;5:16–18. doi: 10.1038/nmeth1156. [DOI] [PubMed] [Google Scholar]
  • 5.Solomon MJ, Larsen PL, Varshavsky A. Mapping protein-DNA interactions in vivo with formaldehyde: evidence that histone H4 is retained on a highly transcribed gene. Cell. 1988;53:937–947. doi: 10.1016/S0092-8674(88)90469-2. [DOI] [PubMed] [Google Scholar]
  • 6.Hurtado A, Holmes KA, Ross-Innes CS, Schmidt D, Carroll JS. FOXA1 is a key determinant of estrogen receptor function and endocrine response. Nat Genet. 2011;43:27–33. doi: 10.1038/ng.730. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Lupien M, Eeckhoute J, Meyer CA, Wang Q, Zhang Y, Li W, Carroll JS, Liu XS, Brown M. FoxA1 translates epigenetic signatures into enhancer-driven lineage-specific transcription. Cell. 2008;132:958–970. doi: 10.1016/j.cell.2008.01.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Young RA. Control of the embryonic stem cell state. Cell. 2011;144:940–954. doi: 10.1016/j.cell.2011.01.032. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Kagey MH, Newman JJ, Bilodeau S, Zhan Y, Orlando DA, van Berkum NL, Ebmeier CC, Goossens J, Rahl PB, Levine SS, et al. Mediator and cohesin connect gene expression and chromatin architecture. Nature. 2010;467:430–435. doi: 10.1038/nature09380. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Chen X, Xu H, Yuan P, Fang F, Huss M, Vega VB, Wong E, Orlov YL, Zhang W, Jiang J, et al. Integration of external signaling pathways with the core transcriptional network in embryonic stem cells. Cell. 2008;133:1106–1117. doi: 10.1016/j.cell.2008.04.043. [DOI] [PubMed] [Google Scholar]
  • 11.Kim J, Chu J, Shen X, Wang J, Orkin SH. An extended transcriptional network for pluripotency of embryonic stem cells. Cell. 2008;132:1049–1061. doi: 10.1016/j.cell.2008.02.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Rahl PB, Lin CY, Seila AC, Flynn RA, McCuine S, Burge CB, Sharp PA, Young RA. c-Myc regulates transcriptional pause release. Cell. 2010;141:432–445. doi: 10.1016/j.cell.2010.03.030. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Handoko L, Xu H, Li G, Ngan CY, Chew E, Schnapp M, Lee CW, Ye C, Ping JL, Mulawadi F, et al. CTCF-mediated functional chromatin interactome in pluripotent cells. Nat Genet. 2011;43:630–638. doi: 10.1038/ng.857. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Dostie J, Richmond TA, Arnaout RA, Selzer RR, Lee WL, Honan TA, Rubio ED, Krumm A, Lamb J, Nusbaum C, et al. Chromosome Conformation Capture Carbon Copy (5C): a massively parallel solution for mapping interactions between genomic elements. Genome Res. 2006;16:1299–1309. doi: 10.1101/gr.5571506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Espinoza CA, Ren B. Mapping higher order structure of chromatin domains. Nat Genet. 2011;43:615–616. doi: 10.1038/ng.869. [DOI] [PubMed] [Google Scholar]
  • 16.Fullwood MJ, Han Y, Wei CL, Ruan X, Ruan Y. Chromatin interaction analysis using paired-end tag sequencing. Curr Protoc Mol Biol. 2010;Chapter 21(Unit 21.15):1–25. doi: 10.1002/0471142727.mb2115s89. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Fullwood MJ, Liu MH, Pan YF, Liu J, Xu H, Mohamed YB, Orlov YL, Velkov S, Ho A, Mei PH, et al. An oestrogen-receptor-alpha-bound human chromatin interactome. Nature. 2009;462:58–64. doi: 10.1038/nature08497. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Li G, Fullwood MJ, Xu H, Mulawadi FH, Velkov S, Vega V, Ariyaratne PN, Mohamed YB, Ooi HS, Tennakoon C, et al. ChIA-PET tool for comprehensive chromatin interaction analysis with paired-end tag sequencing. Genome Biol. 2010;11:R22. doi: 10.1186/gb-2010-11-2-r22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T, Telling A, Amit I, Lajoie BR, Sabo PJ, Dorschner MO, et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science. 2009;326:289–293. doi: 10.1126/science.1181369. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Rusk N. When ChIA PETs meet Hi-C. Nat Methods. 2009;6:863. doi: 10.1038/nmeth1209-863. [DOI] [Google Scholar]
  • 21.Schoenfelder S, Sexton T, Chakalova L, Cope NF, Horton A, Andrews S, Kurukuti S, Mitchell JA, Umlauf D, Dimitrova DS, et al. Preferential associations between co-regulated genes reveal a transcriptional interactome in erythroid cells. Nat Genet. 2010;42:53–61. doi: 10.1038/ng.496. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Theodorou V, Carroll JS. Estrogen receptor action in three dimensions — looping the loop. Breast Cancer Res. 2010;12:303. doi: 10.1186/bcr2470. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Barski A, Cuddapah S, Cui K, Roh TY, Schones DE, Wang Z, Wei G, Chepelev I, Zhao K. High-resolution profiling of histone methylations in the human genome. Cell. 2007;129:823–837. doi: 10.1016/j.cell.2007.05.009. [DOI] [PubMed] [Google Scholar]
  • 24.Wang Z, Zang C, Rosenfeld JA, Schones DE, Barski A, Cuddapah S, Cui K, Roh TY, Peng W, Zhang MQ, et al. Combinatorial patterns of histone acetylations and methylations in the human genome. Nat Genet. 2008;40:897–903. doi: 10.1038/ng.154. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.The ENCODE Project Consortium. A user’s guide to the encyclopedia of DNA elements (ENCODE) PLoS Biol. 2011;9:e1001046. doi: 10.1371/journal.pbio.1001046. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.The modENCODE Consortium. Roy S, Ernst J, Kharchenko PV, Kheradpour P, Negre N, Eaton ML, Landolin JM, Bristow CA, Ma L, Lin MF, et al. Identification of functional elements and regulatory circuits by Drosophila modENCODE. Science. 2010;330:1787–1797. doi: 10.1126/science.1198374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Liu T, Rechtsteiner A, Egelhofer TA, Vielle A, Latorre I, Cheung MS, Ercan S, Ikegami K, Jensen M, Kolasinska-Zwierz P, et al. Broad chromosomal domains of histone modification patterns in C. elegans. Genome Res. 2011;21:227–236. doi: 10.1101/gr.115519.110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Kharchenko PV, Alekseyenko AA, Schwartz YB, Minoda A, Riddle NC, Ernst J, Sabo PJ, Larschan E, Gorchakov AA, Gu T, et al. Comprehensive analysis of the chromatin landscape in Drosophila melanogaster. Nature. 2011;471:480–485. doi: 10.1038/nature09725. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Ernst J, Kheradpour P, Mikkelsen TS, Shoresh N, Ward LD, Epstein CB, Zhang X, Wang L, Issner R, Coyne M, et al. Mapping and analysis of chromatin state dynamics in nine human cell types. Nature. 2011;473:43–49. doi: 10.1038/nature09906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Auerbach RK, Euskirchen G, Rozowsky J, Lamarre-Vincent N, Moqtaderi Z, Lefrançois P, Struhl K, Gerstein M, Snyder M. Mapping accessible chromatin regions using Sono-Seq. Proc Natl Acad Sci USA. 2009;106:14926–14931. doi: 10.1073/pnas.0905443106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Park PJ. ChIP-seq: advantages and challenges of a maturing technology. Nat Rev Genet. 2009;10:669–680. doi: 10.1038/nrg2641. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.de Magalhães JP, Finch CE, Janssens G. Next-generation sequencing in aging research: emerging applications, problems, pitfalls and possible solutions. Ageing Res Rev. 2010;9:315–323. doi: 10.1016/j.arr.2009.10.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Hamady M, Walker JJ, Harris JK, Gold NJ, Knight R. Error-correcting barcoded primers for pyrosequencing hundreds of samples in multiplex. Nat Methods. 2008;5:235–237. doi: 10.1038/nmeth.1184. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Kim JB, Porreca GJ, Song L, Greenway SC, Gorham JM, Church GM, Seidman CE, Seidman JG. Polony multiplex analysis of gene expression (PMAGE) in mouse hypertrophic cardiomyopathy. Science. 2007;316:1481–1484. doi: 10.1126/science.1137325. [DOI] [PubMed] [Google Scholar]
  • 35.Meyer M, Kircher M. Illumina sequencing library preparation for highly multiplexed target capture and sequencing. Cold Spring Harb Protoc. 2010 doi: 10.1101/pdb.prot5448. pdb.prot5448. [DOI] [PubMed] [Google Scholar]
  • 36.Liu T, Ortiz JA, Taing L, Meyer CA, Lee B, Zhang Y, Shin H, Wong SS, Ma J, Lei Y, et al. Cistrome: an integrative platform for transcriptional regulation studies. Genome Biol. 2011;12:R83. doi: 10.1186/gb-2011-12-8-r83. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Ji H, Jiang H, Ma W, Johnson DS, Myers RM, Wong WH. An integrated software system for analyzing ChIP-chip and ChIP-seq data. Nat Biotechnol. 2008;26:1293–1300. doi: 10.1038/nbt.1505. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Ji H, Jiang H, Ma W, Wong WH. Using CisGenome to analyze ChIP-chip and ChIP-seq data. Curr Protoc Bioinformatics. 2011;Chapter 2(Unit2):13. doi: 10.1002/0471250953.bi0213s33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009;10:R25. doi: 10.1186/gb-2009-10-3-r25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012;9:357–359. doi: 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25:1754–1760. doi: 10.1093/bioinformatics/btp324. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Li H, Durbin R. Fast and accurate long-read alignment with Burrows-Wheeler transform. Bioinformatics. 2010;26:589–595. doi: 10.1093/bioinformatics/btp698. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Li H, Ruan J, Durbin R. Mapping short DNA sequencing reads and calling variants using mapping quality scores. Genome Res. 2008;18:1851–1858. doi: 10.1101/gr.078212.108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Lunter G, Goodson M. Stampy: a statistical algorithm for sensitive and fast mapping of Illumina sequence reads. Genome Res. 2011;21:936–939. doi: 10.1101/gr.111120.110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Krawitz P, Rödelsperger C, Jäger M, Jostins L, Bauer S, Robinson PN. Microindel detection in short-read sequence data. Bioinformatics. 2010;26:722–729. doi: 10.1093/bioinformatics/btq027. [DOI] [PubMed] [Google Scholar]
  • 46.Li R, Yu C, Li Y, Lam TW, Yiu SM, Kristiansen K, Wang J. SOAP2: an improved ultrafast tool for short read alignment. Bioinformatics. 2009;25:1966–1967. doi: 10.1093/bioinformatics/btp336. [DOI] [PubMed] [Google Scholar]
  • 47.Bao S, Jiang R, Kwan W, Wang B, Ma X, Song YQ. Evaluation of next-generation sequencing software in mapping and assembly. J Hum Genet. 2011;56:406–414. doi: 10.1038/jhg.2011.43. [DOI] [PubMed] [Google Scholar]
  • 48.Zhang Y, Liu T, Meyer CA, Eeckhoute J, Johnson DS, Bernstein BE, Nusbaum C, Myers RM, Brown M, Li W, et al. Model-based analysis of ChIP-Seq (MACS) Genome Biol. 2008;9:R137. doi: 10.1186/gb-2008-9-9-r137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Kharchenko PV, Tolstorukov MY, Park PJ. Design and analysis of ChIP-seq experiments for DNA-binding proteins. Nat Biotechnol. 2008;26:1351–1359. doi: 10.1038/nbt.1508. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Nix DA, Courdy SJ, Boucher KM. Empirical methods for controlling false positives and estimating confidence in ChIP-Seq peaks. BMC Bioinformatics. 2008;9:523. doi: 10.1186/1471-2105-9-523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Zang C, Schones DE, Zeng C, Cui K, Zhao K, Peng W. A clustering approach for identification of enriched domains from histone modification ChIP-Seq data. Bioinformatics. 2009;25:1952–1958. doi: 10.1093/bioinformatics/btp340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Rozowsky J, Euskirchen G, Auerbach RK, Zhang ZD, Gibson T, Bjornson R, Carriero N, Snyder M, Gerstein MB. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls. Nat Biotechnol. 2009;27:66–75. doi: 10.1038/nbt.1518. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Ji H. Computational analysis of ChIP-seq data. Methods Mol Biol. 2010;674:143–159. doi: 10.1007/978-1-60761-854-6_9. [DOI] [PubMed] [Google Scholar]
  • 54.Fejes AP, Robertson G, Bilenky M, Varhol R, Bainbridge M, Jones SJ. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology. Bioinformatics. 2008;24:1729–1730. doi: 10.1093/bioinformatics/btn305. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Jothi R, Cuddapah S, Barski A, Cui K, Zhao K. Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data. Nucleic Acids Res. 2008;36:5221–5231. doi: 10.1093/nar/gkn488. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56.Garber M, Grabherr MG, Guttman M, Trapnell C. Computational methods for transcriptome annotation and quantification using RNA-seq. Nat Methods. 2011;8:469–477. doi: 10.1038/nmeth.1613. [DOI] [PubMed] [Google Scholar]
  • 57.Pepke S, Wold B, Mortazavi A. Computation for ChIP-seq and RNA-seq studies. Nat Methods. 2009;6:S22–S32. doi: 10.1038/nmeth.1371. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Wilbanks EG, Facciotti MT. Evaluation of algorithm performance in ChIP-seq peak detection. PLoS ONE. 2010;5:e11471. doi: 10.1371/journal.pone.0011471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc B. 1995;57:289–300. [Google Scholar]
  • 60.Storey JD. A direct approach to false discovery rates. J R Stat Soc B. 2002;64:479–498. [Google Scholar]
  • 61.Storey JD, Tibshirani R. Statistical significance for genomewide studies. Proc Natl Acad Sci USA. 2003;100:9440–9445. doi: 10.1073/pnas.1530509100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Valouev A, Johnson DS, Sundquist A, Medina C, Anton E, Batzoglou S, Myers RM, Sidow A. Genome-wide analysis of transcription factor binding sites based on ChIP-Seq data. Nat Methods. 2008;5:829–834. doi: 10.1038/nmeth.1246. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Tuteja G, White P, Schug J, Kaestner KH. Extracting transcription factor targets from ChIP-Seq data. Nucleic Acids Res. 2009;37:e113. doi: 10.1093/nar/gkp536. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Johnson DS, Mortazavi A, Myers RM, Wold B. Genome-wide mapping of in vivo protein-DNA interactions. Science. 2007;316:1497–1502. doi: 10.1126/science.1141319. [DOI] [PubMed] [Google Scholar]
  • 65.Zhang Y, Shin H, Song JS, Lei Y, Liu XS. Identifying positioned nucleosomes with epigenetic marks in human from ChIP-Seq. BMC Genomics. 2008;9:537. doi: 10.1186/1471-2164-9-537. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Chen Y, Meyer CA, Liu T, Li W, Liu JS, Liu XS. MM-ChIP enables integrative analysis of cross-platform and between-laboratory ChIP-chip or ChIP-seq data. Genome Biol. 2011;12:R11. doi: 10.1186/gb-2011-12-2-r11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D. The human genome browser at UCSC. Genome Res. 2002;12:996–1006. doi: 10.1101/gr.229102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 68.Fujita PA, Rhead B, Zweig AS, Hinrichs AS, Karolchik D, Cline MS, Goldman M, Barber GP, Clawson H, Coelho A, et al. The UCSC Genome Browser database: update 2011. Nucleic Acids Res. 2011;39:D876–D882. doi: 10.1093/nar/gkq963. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Karolchik D, Hinrichs AS, Furey TS, Roskin KM, Sugnet CW, Haussler D, Kent WJ. The UCSC Table Browser data retrieval tool. Nucleic Acids Res. 2004;32:D493–D496. doi: 10.1093/nar/gkh103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Raney BJ, Cline MS, Rosenbloom KR, Dreszer TR, Learned K, Barber GP, Meyer LR, Sloan CA, Malladi VS, Roskin KM, et al. ENCODE whole-genome data in the UCSC genome browser (2011 update) Nucleic Acids Res. 2011;39:D871–D875. doi: 10.1093/nar/gkq1017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Robinson JT, Thorvaldsdóttir H, Winckler W, Guttman M, Lander ES, Getz G, Mesirov JP. Integrative genomics viewer. Nat Biotechnol. 2011;29:24–26. doi: 10.1038/nbt.1754. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Nicol JW, Helt GA, Blanchard SG, Jr, Raja A, Loraine AE. The Integrated Genome Browser: free software for distribution and exploration of genome-scale datasets. Bioinformatics. 2009;25:2730–2731. doi: 10.1093/bioinformatics/btp472. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73.Donlin MJ. Using the Generic Genome Browser (GBrowse) Curr Protoc Bioinformatics. 2009;Chapter 9(Unit 9.9) doi: 10.1002/0471250953.bi0909s28. [DOI] [PubMed] [Google Scholar]
  • 74.Podicheti R, Dong Q. Administering GBrowse sites with WebGBrowse. Curr Protoc Bioinformatics. 2011;Chapter 9(Unit 9.14) doi: 10.1002/0471250953.bi0914s33. [DOI] [PubMed] [Google Scholar]
  • 75.Huang W, Marth G. EagleView: a genome assembly viewer for next-generation sequencing technologies. Genome Res. 2008;18:1538–1543. doi: 10.1101/gr.076067.108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Milne I, Bayer M, Cardle L, Shaw P, Stephen G, Wright F, Marshall D. Tablet — next generation sequence assembly visualization. Bioinformatics. 2010;26:401–402. doi: 10.1093/bioinformatics/btp666. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Nicol JW, Helt GA, Blanchard SG, Jr, Raja A, Loraine AE. The Integrated Genome Browser: free software for distribution and exploration of genome-scale datasets. Bioinformatics. 2009;25:2730–2731. doi: 10.1093/bioinformatics/btp472. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 78.Bao H, Guo H, Wang J, Zhou R, Lu X, Shi S. MapView: visualization of short reads alignment on a desktop computer. Bioinformatics. 2009;25:1554–1555. doi: 10.1093/bioinformatics/btp255. [DOI] [PubMed] [Google Scholar]
  • 79.Lewis SE, Searle SM, Harris N, Gibson M, Lyer V, Richter J, Wiel C, Bayraktaroglir L, Birney E, Crosby MA, et al. Apollo: a sequence annotation editor. Genome Biol. 2002;3:RESEARCH0082. doi: 10.1186/gb-2002-3-12-research0082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 80.Li QH, Brown JB, Huang H, Bickel PJ. Measuring reproducibility of high-throughput experiments. Ann Appl Stat. 2011;5:1752–1779. doi: 10.1214/11-AOAS466. [DOI] [Google Scholar]
  • 81.Siepel A, Bejerano G, Pedersen JS, Hinrichs AS, Hou M, Rosenbloom K, Clawson H, Spieth J, Hillier LW, Richards S, et al. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes. Genome Res. 2005;15:1034–1050. doi: 10.1101/gr.3715005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Robasky K, Bulyk ML. UniPROBE, update 2011: expanded content and search tools in the online database of protein-binding microarray data on protein-DNA interactions. Nucleic Acids Res. 2011;39:D124–D128. doi: 10.1093/nar/gkq992. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 83.Xie Z, Hu S, Blackshaw S, Zhu H, Qian J. hPDI: a database of experimental human protein-DNA interactions. Bioinformatics. 2010;26:287–289. doi: 10.1093/bioinformatics/btp631. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 84.Bryne JC, Valen E, Tang MH, Marstrand T, Winther O, da Piedade I, Krogh A, Lenhard B, Sandelin A. JASPAR, the open access database of transcription factor-binding profiles: new content and tools in the 2008 update. Nucleic Acids Res. 2008;36:D102–D106. doi: 10.1093/nar/gkm955. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.AlQuraishi M, McAdams HH. Direct inference of protein-DNA interactions using compressed sensing methods. Proc Natl Acad Sci USA. 2011;108:14819–14824. doi: 10.1073/pnas.1106460108. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 86.Nutiu R, Friedman RC, Luo S, Khrebtukova I, Silva D, Li R, Zhang L, Schroth GP, Burge CB. Direct measurement of DNA affinity landscapes on a high-throughput sequencing instrument. Nat Biotechnol. 2011;29:659–664. doi: 10.1038/nbt.1882. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 87.Bailey TL. DREME: motif discovery in transcription factor ChIP-seq data. Bioinformatics. 2011;27:1653–1659. doi: 10.1093/bioinformatics/btr261. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Machanick P, Bailey TL. MEME-ChIP: motif analysis of large DNA datasets. Bioinformatics. 2011;27:1696–1697. doi: 10.1093/bioinformatics/btr189. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 89.Liu XS, Brutlag DL, Liu JS. An algorithm for finding protein-DNA binding sites with applications to chromatin-immunoprecipitation microarray experiments. Nat Biotechnol. 2002;20:835–839. doi: 10.1038/nbt717. [DOI] [PubMed] [Google Scholar]
  • 90.Ma X, Kulkarni A, Zhang Z, Xuan Z, Serfling R, Zhang MQ. A highly efficient and effective motif discovery method for ChIP-seq/ChIP-chip data using positional information. Nucleic Acids Res. 2012;40:e50. doi: 10.1093/nar/gkr1135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 91.Meyer CA, He HH, Brown M, Liu XS. BINOCh: binding inference from nucleosome occupancy changes. Bioinformatics. 2011;27:1867–1868. doi: 10.1093/bioinformatics/btr279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 92.Bell O, Tiwari VK, Thomä NH, Schübeler D. Determinants and dynamics of genome accessibility. Nat Rev Genet. 2011;12:554–564. doi: 10.1038/nrg3017. [DOI] [PubMed] [Google Scholar]
  • 93.Crawford GE, Holt IE, Mullikin JC, Tai D, Blakesley R, Bouffard G, Young A, Masiello C, Green ED, Wolfsberg TG, et al. Identifying gene regulatory elements by genome-wide recovery of DNase hypersensitive sites. Proc Natl Acad Sci USA. 2004;101:992–997. doi: 10.1073/pnas.0307540100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 94.Sabo PJ, Humbert R, Hawrylycz M, Wallace JC, Dorschner MO, McArthur M, Stamatoyannopoulos JA. Genome-wide identification of DNaseI hypersensitive sites using active chromatin sequence libraries. Proc Natl Acad Sci USA. 2004;101:4537–4542. doi: 10.1073/pnas.0400678101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 95.Bernstein BE, Stamatoyannopoulos JA, Costello JF, Ren B, Milosavljevic A, Meissner A, Kellis M, Marra MA, Beaudet AL, Ecker JR, et al. The NIH Roadmap Epigenomics Mapping Consortium. Nat Biotechnol. 2010;28:1045–1048. doi: 10.1038/nbt1010-1045. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 96.Gerstein MB, Lu ZJ, Van Nostrand EL, Cheng C, Arshinoff BI, Liu T, Yip KY, Robilotto R, Rechtsteiner A, Ikegami K, et al. Integrative analysis of the Caenorhabditis elegans genome by the modENCODE project. Science. 2010;330:1775–1787. doi: 10.1126/science.1196914. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 97.Moorman C, Sun LV, Wang J, de Wit E, Talhout W, Ward LD, Greil F, Lu XJ, White KP, Bussemaker HJ, et al. Hotspots of transcription factor colocalization in the genome of Drosophila melanogaster. Proc Natl Acad Sci USA. 2006;103:12027–12032. doi: 10.1073/pnas.0605003103. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 98.Nègre N, Brown CD, Ma L, Bristow CA, Miller SW, Wagner U, Kheradpour P, Eaton ML, Loriaux P, Sealfon R, et al. A cis-regulatory map of the Drosophila genome. Nature. 2011;471:527–531. doi: 10.1038/nature09990. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 99.Shin H, Liu T, Manrai AK, Liu XS. CEAS: cis-regulatory element annotation system. Bioinformatics. 2009;25:2605–2606. doi: 10.1093/bioinformatics/btp479. [DOI] [PubMed] [Google Scholar]
  • 100.Wang D, Garcia-Bassets I, Benner C, Li W, Su X, Zhou Y, Qiu J, Liu W, Kaikkonen MU, Ohgi KA, et al. Reprogramming transcription by distinct classes of enhancers functionally defined by eRNA. Nature. 2011;474:390–394. doi: 10.1038/nature10006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 101.Cheung I, Shulha HP, Jiang Y, Matevossian A, Wang J, Weng Z, Akbarian S. Developmental regulation and individual differences of neuronal H3K4me3 epigenomes in the prefrontal cortex. Proc Natl Acad Sci USA. 2010;107:8824–8829. doi: 10.1073/pnas.1001702107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 102.Xu H, Wei CL, Lin F, Sung WK. An HMM approach to genome-wide identification of differential histone modification sites from ChIP-seq data. Bioinformatics. 2008;24:2344–2349. doi: 10.1093/bioinformatics/btn402. [DOI] [PubMed] [Google Scholar]
  • 103.Robinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics. 2010;26:139–140. doi: 10.1093/bioinformatics/btp616. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 104.Anders S, Huber W. Differential expression analysis for sequence count data. Genome Biol. 2010;11:R106. doi: 10.1186/gb-2010-11-10-r106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 105.Hardcastle TJ, Kelly KA. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data. BMC Bioinformatics. 2010;11:422. doi: 10.1186/1471-2105-11-422. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 106.Verzi MP, Shin H, He HH, Sulahian R, Meyer CA, Montgomery RK, Fleet JC, Brown M, Liu XS, Shivdasani RA. Differentiation-specific histone modifications reveal dynamic chromatin interactions and partners for the intestinal transcription factor CDX2. Dev Cell. 2010;19:713–726. doi: 10.1016/j.devcel.2010.10.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 107.Tang Q, Chen Y, Meyer C, Geistlinger T, Lupien M, Wang Q, Liu T, Zhang Y, Brown M, Liu XS. A comprehensive view of nuclear receptor cancer cistromes. Cancer Res. 2011;71:6940–6947. doi: 10.1158/0008-5472.CAN-11-2091. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 108.The ENCODE Project Consortium. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project. Nature. 2007;447:799–816. doi: 10.1038/nature05874. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 109.Wang Q, Li W, Zhang Y, Yuan X, Xu K, Yu J, Chen Z, Beroukhim R, Wang H, Lupien M, et al. Androgen receptor regulates a distinct transcription program in androgen-independent prostate cancer. Cell. 2009;138:245–256. doi: 10.1016/j.cell.2009.04.056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 110.Carroll JS, Meyer CA, Song J, Li W, Geistlinger TR, Eeckhoute J, Brodsky AS, Keeton EK, Fertuck KC, Hall GF, et al. Genome-wide analysis of estrogen receptor binding sites. Nat Genet. 2006;38:1289–1297. doi: 10.1038/ng1901. [DOI] [PubMed] [Google Scholar]
  • 111.Huang DW, Sherman BT, Lempicki RA. Bioinformatics enrichment tools: paths toward the comprehensive functional analysis of large gene lists. Nucleic Acids Res. 2009;37:1–13. doi: 10.1093/nar/gkn923. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 112.Huang DW, Sherman BT, Lempicki RA. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nat Protoc. 2009;4:44–57. doi: 10.1038/nprot.2008.211. [DOI] [PubMed] [Google Scholar]
  • 113.Thomas PD, Campbell MJ, Kejariwal A, Mi H, Karlak B, Daverman R, Diemer K, Muruganujan A, Narechania A. PANTHER: a library of protein families and subfamilies indexed by function. Genome Res. 2003;13:2129–2141. doi: 10.1101/gr.772403. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 114.McLean CY, Bristor D, Hiller M, Clarke SL, Schaar BT, Lowe CB, Wenger AM, Bejerano G. GREAT improves functional interpretation of cis-regulatory regions. Nat Biotechnol. 2010;28:495–501. doi: 10.1038/nbt.1630. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 115.Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, Paulovich A, Pomeroy SL, Golub TR, Lander ES, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci USA. 2005;102:15545–15550. doi: 10.1073/pnas.0506580102. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 116.Zhang Z, Chang CW, Goh WL, Sung WK, Cheung E. CENTDIST: discovery of co-associated factors by motif distribution. Nucleic Acids Res. 2011;39:W391–W399. doi: 10.1093/nar/gkr387. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 117.Carroll JS, Liu XS, Brodsky AS, Li W, Meyer CA, Szary AJ, Eeckhoute J, Shao W, Hestermann EV, Geistlinger TR, et al. Chromosome-wide mapping of estrogen receptor binding reveals long-range regulation requiring the forkhead protein FoxA1. Cell. 2005;122:33–43. doi: 10.1016/j.cell.2005.05.008. [DOI] [PubMed] [Google Scholar]
  • 118.Giardine B, Riemer C, Hardison RC, Burhans R, Elnitski L, Shah P, Zhang Y, Blankenberg D, Albert I, Taylor J, et al. Galaxy: a platform for interactive large-scale genome analysis. Genome Res. 2005;15:1451–1455. doi: 10.1101/gr.4086505. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 119.Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–842. doi: 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 120.Feng D, Liu T, Sun Z, Bugge A, Mullican SE, Alenghat T, Liu XS, Lazar MA. A circadian rhythm orchestrated by histone deacetylase 3 controls hepatic lipid metabolism. Science. 2011;331:1315–1319. doi: 10.1126/science.1198125. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 121.Odom DT, Dowell RD, Jacobsen ES, Gordon W, Danford TW, MacIsaac KD, Rolfe PA, Conboy CM, Gifford DK, Fraenkel E. Tissue-specific transcriptional regulation has diverged significantly between human and mouse. Nat Genet. 2007;39:730–732. doi: 10.1038/ng2047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 122.Schmidt D, Wilson MD, Ballester B, Schwalie PC, Brown GD, Marshall A, Kutter C, Watt S, Martinez-Jimenez CP, Mackay S, et al. Five-vertebrate ChIP-seq reveals the evolutionary dynamics of transcription factor binding. Science. 2010;328:1036–1040. doi: 10.1126/science.1186176. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 123.Chung D, Kuan PF, Li B, Sanalkumar R, Liang K, Bresnick EH, Dewey C, Keleş S. Discovering transcription factor binding sites in highly repetitive regions of genomes with multi-read analysis of ChIP-Seq data. PLOS Comput Biol. 2011;7:e1002111. doi: 10.1371/journal.pcbi.1002111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 124.Wang T, Zeng J, Lowe CB, Sellers RG, Salama SR, Yang M, Burgess SM, Brachmann RK, Haussler D. Species-specific endogenous retroviruses shape the transcriptional network of the human tumor suppressor protein p53. Proc Natl Acad Sci USA. 2007;104:18613–18618. doi: 10.1073/pnas.0703637104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 125.Eaton ML, Prinz JA, MacAlpine HK, Tretyakov G, Kharchenko PV, MacAlpine DM. Chromatin signatures of the Drosophila replication program. Genome Res. 2011;21:164–174. doi: 10.1101/gr.116038.110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 126.Bernstein BE, Mikkelsen TS, Xie X, Kamal M, Huebert DJ, Cuff J, Fry B, Meissner A, Wernig M, Plath K, et al. A bivalent chromatin structure marks key developmental genes in embryonic stem cells. Cell. 2006;125:315–326. doi: 10.1016/j.cell.2006.02.041. [DOI] [PubMed] [Google Scholar]
  • 127.Kolasinska-Zwierz P, Down T, Latorre I, Liu T, Liu XS, Ahringer J. Differential chromatin marking of introns and expressed exons by H3K36me3. Nat Genet. 2009;41:376–381. doi: 10.1038/ng.322. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 128.Creyghton MP, Cheng AW, Welstead GG, Kooistra T, Carey BW, Steine EJ, Hanna J, Lodato MA, Frampton GM, Sharp PA, et al. Histone H3K27ac separates active from poised enhancers and predicts developmental state. Proc Natl Acad Sci USA. 2010;107:21931–21936. doi: 10.1073/pnas.1016071107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 129.Heintzman ND, Stuart RK, Hon G, Fu Y, Ching CW, Hawkins RD, Barrera LO, Van Calcar S, Qu C, Ching KA, et al. Distinct and predictive chromatin signatures of transcriptional promoters and enhancers in the human genome. Nat Genet. 2007;39:311–318. doi: 10.1038/ng1966. [DOI] [PubMed] [Google Scholar]
  • 130.Ernst J, Kellis M. Discovery and characterization of chromatin states for systematic annotation of the human genome. Nat Biotechnol. 2010;28:817–825. doi: 10.1038/nbt.1662. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 131.Hoffman MM, Buske OJ, Wang J, Weng Z, Bilmes JA, Noble WS. Unsupervised pattern discovery in human chromatin structure through genomic segmentation. Nat Methods. 2012;9:473–476. doi: 10.1038/nmeth.1937. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 132.He HH, Meyer CA, Shin H, Bailey ST, Wei G, Wang Q, Zhang Y, Xu K, Ni M, Lupien M, et al. Nucleosome dynamics define transcriptional enhancers. Nat Genet. 2010;42:343–347. doi: 10.1038/ng.545. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 133.Kasowski M, Grubert F, Heffelfinger C, Hariharan M, Asabere A, Waszak SM, Habegger L, Rozowsky J, Shi M, Urban AE, et al. Variation in transcription factor binding among humans. Science. 2010;328:232–235. doi: 10.1126/science.1183621. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 134.McDaniell R, Lee BK, Song L, Liu Z, Boyle AP, Erdos MR, Scott LJ, Morken MA, Kucera KS, Battenhouse A, et al. Heritable individual-specific and allele-specific chromatin signatures in humans. Science. 2010;328:235–239. doi: 10.1126/science.1184655. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 135.Ahmadiyeh N, Pomerantz MM, Grisanzio C, Herman P, Jia L, Almendro V, He HH, Brown M, Liu XS, Davis M, et al. 8q24 prostate, breast, and colon cancer risk loci show tissue-specific long-range interaction with MYC. Proc Natl Acad Sci USA. 2010;107:9742–9746. doi: 10.1073/pnas.0910668107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 136.Birney E, Lieb JD, Furey TS, Crawford GE, Iyer VR. Allele-specific and heritable chromatin signatures in humans. Hum Mol Genet. 2010;19:R204–R209. doi: 10.1093/hmg/ddq404. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 137.Pickrell JK, Gaffney DJ, Gilad Y, Pritchard JK. False positive peaks in ChIP-seq and other sequencing-based functional assays caused by unannotated high copy number regions. Bioinformatics. 2011;27:2144–2146. doi: 10.1093/bioinformatics/btr354. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 138.Verzi MP, Shin H, Ho LL, Liu XS, Shivdasani RA. Essential and redundant functions of caudal family proteins in activating adult intestinal genes. Mol Cell Biol. 2011;31:2026–2039. doi: 10.1128/MCB.01250-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 139.Iyengar S, Ivanov AV, Jin VX, Rauscher FJ, 3rd, Farnham PJ. Functional analysis of KAP1 genomic recruitment. Mol Cell Biol. 2011;31:1833–1847. doi: 10.1128/MCB.01331-10. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 140.O’Geen H, Echipare L, Farnham PJ. Using ChIP-seq technology to generate high-resolution profiles of histone modifications. Methods Mol Biol. 2011;791:265–286. doi: 10.1007/978-1-61779-316-5_20. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 141.Adli M, Zhu J, Bernstein BE. Genome-wide chromatin maps derived from limited numbers of hematopoietic progenitors. Nat Methods. 2010;7:615–618. doi: 10.1038/nmeth.1478. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 142.Shankaranarayanan P, Mendoza-Parra MA, Walia M, Wang L, Li N, Trindade LM, Gronemeyer H. Single-tube linear DNA amplification (LinDA) for robust ChIP-seq. Nat Methods. 2011;8:565–567. doi: 10.1038/nmeth.1626. [DOI] [PubMed] [Google Scholar]
  • 143.Rhee HS, Pugh BF. Comprehensive genome-wide protein-DNA interactions detected at single-nucleotide resolution. Cell. 2011;147:1408–1419. doi: 10.1016/j.cell.2011.11.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 144.Cokus SJ, Feng S, Zhang X, Chen Z, Merriman B, Haudenschild CD, Pradhan S, Nelson SF, Pellegrini M, Jacobsen SE. Shotgun bisulphite sequencing of the Arabidopsis genome reveals DNA methylation patterning. Nature. 2008;452:215–219. doi: 10.1038/nature06745. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 145.Lister R, O’Malley RC, Tonti-Filippini J, Gregory BD, Berry CC, Millar AH, Ecker JR. Highly integrated single-base resolution maps of the epigenome in Arabidopsis. Cell. 2008;133:523–536. doi: 10.1016/j.cell.2008.03.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 146.Lister R, Pelizzola M, Dowen RH, Hawkins RD, Hon G, Tonti-Filippini J, Nery JR, Lee L, Ye Z, Ngo QM, et al. Human DNA methylomes at base resolution show widespread epigenomic differences. Nature. 2009;462:315–322. doi: 10.1038/nature08514. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 147.Xiang H, Zhu J, Chen Q, Dai F, Li X, Li M, Zhang H, Zhang G, Li D, Dong Y, et al. Single base-resolution methylome of the silkworm reveals a sparse epigenomic map. Nat Biotechnol. 2010;28:516–520. doi: 10.1038/nbt.1626. [DOI] [PubMed] [Google Scholar]
  • 148.Bar-Joseph Z, Gerber GK, Lee TI, Rinaldi NJ, Yoo JY, Robert F, Gordon DB, Fraenkel E, Jaakkola TS, Young RA, et al. Computational discovery of gene modules and regulatory networks. Nat Biotechnol. 2003;21:1337–1342. doi: 10.1038/nbt890. [DOI] [PubMed] [Google Scholar]
  • 149.Basso K, Margolin AA, Stolovitzky G, Klein U, Dalla-Favera R, Califano A. Reverse engineering of regulatory networks in human B cells. Nat Genet. 2005;37:382–390. doi: 10.1038/ng1532. [DOI] [PubMed] [Google Scholar]
  • 150.Friedman N. Inferring cellular networks using probabilistic graphical models. Science. 2004;303:799–805. doi: 10.1126/science.1094068. [DOI] [PubMed] [Google Scholar]
  • 151.Lee I, Date SV, Adai AT, Marcotte EM. A probabilistic functional network of yeast genes. Science. 2004;306:1555–1558. doi: 10.1126/science.1099511. [DOI] [PubMed] [Google Scholar]
  • 152.Liao JC, Boscolo R, Yang YL, Tran LM, Sabatti C, Roychowdhury VP. Network component analysis: reconstruction of regulatory signals in biological systems. Proc Natl Acad Sci USA. 2003;100:15522–15527. doi: 10.1073/pnas.2136632100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 153.Lemmens K, Dhollander T, De Bie T, Monsieurs P, Engelen K, Smets B, Winderickx J, De Moor B, Marchal K. Inferring transcriptional modules from ChIP-chip, motif and microarray data. Genome Biol. 2006;7:R37. doi: 10.1186/gb-2006-7-5-r37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 154.Liu X, Jessen WJ, Sivaganesan S, Aronow BJ, Medvedovic M. Bayesian hierarchical model for transcriptional module discovery by jointly modeling gene expression and ChIP-chip data. BMC Bioinformatics. 2007;8:283. doi: 10.1186/1471-2105-8-283. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 155.Youn A, Reiss DJ, Stuetzle W. Learning transcriptional networks from the integration of ChIP-chip and expression data in a non-parametric model. Bioinformatics. 2010;26:1879–1886. doi: 10.1093/bioinformatics/btq289. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 156.Kinde I, Wu J, Papadopoulos N, Kinzler KW, Vogelstein B. Detection and quantification of rare mutations with massively parallel sequencing. Proc Natl Acad Sci USA. 2011;108:9530–9535. doi: 10.1073/pnas.1105422108. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES