Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Barbara Wold mailto:woldb@caltech.edu, Georgi K. Marinov mailto:georgi@caltech.edu, Diane Trout mailto:diane@caltech.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus, we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Genome-wide occupancy maps of transcription factors (TFs) are generated by ChIP-seq. A ChIP-Seq experiment combines a chromatin immunoprecipitation (ChIP) experiment that enriches genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibody) with high-throughput short read sequencing of the enriched DNA fragments (Wold & Myers, 2008). Proteins are crosslinked to DNA (usually with formaldehyde), chromatin is sheared and immunoprecipitated with the antibody of interest. The immunoprecipitated material is turned into a sequencing library and sequenced. The sequencing reads are then aligned to the genome. A control sample consisting of sonicated chromatin that has not been immunoprecipitated or immunoprecipitated with a non-specific immunoglobulin is also sequenced. The ChIP and the control datasets are analyzed with a variety of software packages to identify regions occupied by the target protein. The sequencing data, alignments and analysis files for these experiments are available for download. In specific, the Ren lab examined RNA polymerase II (PolII), co-activator protein p300, the insulator protein CTCF, and two chromatin modification marks, H3K4me3 and H3K4me1, due to their demonstrated utilities in identifying promoters, enhancers and insulator elements (Barski et al., 2007; Blow et al., 2010; Heintzman et al., 2009; Kim et al., 2007; Kim et al., 2005a; Visel et al., 2009). Enrichment of H3K4me3 or PolII signals is a strong indicator of an active promoter, while the presence of p300 or H3K4me1 outside of promoter regions has been used as a mark for enhancers. CTCF binding sites are considered as a mark for potential insulator elements. For each transcription factor or chromatin mark in each tissue, ChIP-seq was carried out with at least two biological replicates. Each experiment produced 20-30 million monoclonal, uniquely mapped tags. Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of histone modifications Levels of three histone modifications are being determined. H3K4me1 (monomethylation of lysine 4 of histone H3) is a mark for active chromatin and in the absence of H3K4me3, it is one indicator of an enhancer. H3K4me3 (trimethylation of lysine 4 of histone H3) is highly enriched at active promoters. One repressive (Polycomb) mark, H3K27me3, is associated with some silenced genes. Maps of genomic DNA in chromatin with these histone modifications are generated by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for histone modifications are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Maps of occupancy of genomic DNA by transcription factors (TFs) are determined by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for TF binding are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Cells were grown according to the approved ENCODE cell culture protocols (http://genome-test.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Chromatin immunoprecipitation followed published methods (Johnson & Mortazavi et al., 2007) with the exception of certain experiments for which glutaraldehyde was added to the crosslink reaction. Information on the antibodies used is available via the metadata for each subtrack. Libraries were constructed using the Illumina ChIP-seq Sample Preparation Kit or using a modified protocol that includes the addition of multiplexing tags to the fragments. DNA fragments were repaired to generate blunt ends and a single A nucleotide was added to each end. Double-stranded Illumina adaptors or Double-stranded Illumina adaptors with multiplexing tags were ligated to the fragments. Ligation products were amplified by 18 cycles of PCR, and the DNA between 150-250 bp was gel purified. Completed libraries were quantified with Quant-iT dsDNA HS Assay Kit. The DNA library was sequenced on the Illumina GAII and GAIIx sequencing systems, and more recently, for multiplexed libraries, several of them were pooled and sequenced on the HiSeq platform. Cluster generation, linearization, blocking and sequencing primer reagents were provided in the Illumina Cluster Amplification kits. Older libraries were generated using 2 rounds of PCR. Matched input samples were sequenced for each variation of fixation conditions and the number of PCR rounds. Reads of 32 bp, 36 bp or 50 bp length were generated. Sequencing reads (fastq files) were assigned to the corresponding libraries based on the multiplexing tag for pooled libraries (all tags have been removed from reads in the fastq files available for download) or directly processed. Bowtie (Langmead et al., 2009) was used to map reads to the male or female version of the mouse genome (excluding the _random chromosomes in the assembly) depending on the cell line sex. The following parameters were used: "-v 2 -k 11 -m 10 -t --best --strata". Aligned reads were converted into rds files using the ERANGE package (Johnson & Mortazavi et al., 2007) and the findall.py program in ERANGE was used to identify enriched regions against the matching input sample. The following settings were used for point-source transcription factors: "--shift learn --ratio 3 --minimum 2 --listPeak --revbackground". For histone modifications, the settings were changed to "--notrim --nodirectionality --spacing 100 --ratio 3 --minimum 2 --listPeak --revbackground". Cells were grown according to the approved ENCODE cell culture protocols (http://genome.ucsc.edu/ENCODE/protocols/cell/mouse). RNA-Seq RNA samples from tissues and primary cells were extracted from Trizol® according to protocol (Invitrogen). PolyA+ RNA was purified with the Dynabeads mRNA purification kit (Invitrogen). The mRNA libraries were prepared for strand-specific sequencing as described previously (Parkhomchuk et al., 2009). Sequencing and Analysis Samples were sequenced on Illumina Genome Analyzer II, Genome Analyzer IIx and HiSeq 2000 platforms for 36 cycles. Image analysis, base calling and alignment to the mouse genome version mm9 were performed using Illumina's RTA. Alignment to the mouse genome was performed using TopHat (Trapnell et al., 2009). Wig files were generated by TopHat and expression levels were calculated with Cufflinks (Trapnell et al., 2010). Cells were grown according to the approved ENCODE cell culture protocols (http://genome.ucsc.edu/ENCODE/protocols/cell/mouse). Enrichment and Library Preparation Chromatin immunoprecipitation was performed according to Ren Lab ChIP Protocol (http://bioinformatics-renlab.ucsd.edu/RenLabChipProtocolV1.pdf). Library construction was performed according to Ren Lab Library Protocol (http://bioinformatics-renlab.ucsd.edu/RenLabLibraryProtocolV1.pdf). Sequencing and Analysis Samples were sequenced on Illumina Genome Analyzer II, Genome Analyzer IIx and HiSeq 2000 platforms for 36 cycles. Image analysis, base calling and alignment to the mouse genome version mm9 were performed using Illumina's RTA and Genome Analyzer Pipeline software. Alignment to the mouse genome was performed using ELAND or Bowtie (Langmead et al., 2009) with a seed length of 25 and allowing up to two mismatches. Only the sequences that mapped to one location were used for further analysis. Of those sequences, clonal reads, defined as having the same start position on the same strand, were discarded. BED and wig files were created using custom perl scripts. Cells were grown according to the approved ENCODE cell culture protocols. The chromatin immunoprecipitation followed published methods (Welch et al. 2004). Information on antibodies used is available via the hyperlinks in the "Select subtracks" menu. Samples passing initial quality thresholds (enrichment and depletion for positive and negative controls - if available - by quantitative PCR of ChIP material) are processed for library construction for Illumina sequencing, using the ChIP-seq Sample Preparation Kit purchased from Illumina. Starting with a 10 ng sample of ChIP DNA, DNA fragments were repaired to generate blunt ends and a single A nucleotide was added to each end. Double-stranded Illumina adaptors were ligated to the fragments. Ligation products were amplified by 18 cycles of PCR, and the DNA between 250-350 bp was gel purified. Completed libraries were quantified with Quant-iT dsDNA HS Assay Kit. The DNA library was sequenced on the Illumina Genome Analyzer II sequencing system, and more recently on the HiSeq. Cluster generation, linearization, blocking and sequencing primer reagents were provided in the Illumina Cluster Amplification kits. All samples are being determined as biological replicates except time course samples. The data displayed are from the pooled reads for all replicates, but individual replicates are available by download. The resulting 36-nucleotide sequence reads (fastq files) were moved to a data library in Galaxy, and the tools implemented in Galaxy were used for further processing via workflows (Blankenberg et al. 2010). The reads were mapped to the mouse genome (mm9 assembly) using the program bowtie (Langmead et al. 2009), and the files of mapped reads for the ChIP sample and from the "input" control (no antibody) were processed by MACs (Zhang et al. 2008) to call peaks for occupancy by transcription factors, using the parameters mfold=15, bandwidth=125. Because the signal for some histone modifications is not expected to be tightly localized (compared to a transcription factor), peak calling programs may not be appropriate. Thus in addition, we provide wiggle tracks with tag counts for every 10 bp segment. Per-replicate aligments and sequences are available for download at downloads page. Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). The chromatin immunoprecipitation followed published methods (Welch et al. 2004). Information on antibodies used is available via the hyperlinks in the "Select subtracks" menu. Samples passing initial quality thresholds (enrichment and depletion for positive and negative controls - if available - by quantitative PCR of ChIP material) are processed for library construction for Illumina sequencing, using the ChIP-seq Sample Preparation Kit purchased from Illumina. Starting with a 10 ng sample of ChIP DNA, DNA fragments were repaired to generate blunt ends and a single A nucleotide was added to each end. Double-stranded Illumina adaptors were ligated to the fragments. Ligation products were amplified by 18 cycles of PCR, and the DNA between 250-350 bp was gel purified. Completed libraries were quantified with Quant-iT dsDNA HS Assay Kit. The DNA library was sequenced on the Illumina Genome Analyzer II sequencing system, and more recently on the HiSeq. Cluster generation, linearization, blocking and sequencing primer reagents were provided in the Illumina Cluster Amplification kits. All samples are being determined as biological replicates except time course samples. The data displayed are from the pooled reads for all replicates, but individual replicates are available by download. The resulting 36-nucleotide sequence reads (fastq files) were moved to a data library in Galaxy, and the tools implemented in Galaxy were used for further processing via workflows (Blankenberg et al. 2010). The reads were mapped to the mouse genome (mm9 assembly) using the program bowtie (Langmead et al. 2009), and the files of mapped reads for the ChIP sample and from the "input" control (no antibody) were processed by MACs (Zhang et al. 2008) to call peaks for occupancy by transcription factors, using the parameters mfold=15, bandwidth=125. Per-replicate aligments and sequences are available for download at downloads page (http://hgdownload.cse.ucsc.edu/goldenPath/mm9/encodeDCC/wgEncodePsuTfbs/). Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). For details on the chromatin immunoprecipitation protocol used, see Euskirchen et. al., (2007), Rozowsky et. al. (2009) and Auerbach et. al. (2009). DNA recovered from the precipitated chromatin was sequenced on the Illumina (Solexa) sequencing platform and mapped to the genome using the Eland alignment program. ChIP-seq data was scored based on sequence reads (length ~30 bps) that align uniquely to the human genome. From the mapped tags, a signal map of ChIP DNA fragments (average fragment length ~ 200 bp) was constructed where the signal height is the number of overlapping fragments at each nucleotide position in the genome. Reads were pooled from all submitted replicates to generate the Peak and Signal files. Per-replicate aligments and sequences are available for download at downloads page (http://hgdownload.cse.ucsc.edu/goldenPath/mm9/encodeDCC/wgEncodeSydhTfbs/). For each 1 Mb segment of each chromosome, a peak height threshold was determined by requiring a false discovery rate <= 0.01 when comparing the number of peaks above said threshold to the number of peaks obtained from multiple simulations of a random null background with the same number of mapped reads (also accounting for the fraction of mapable bases for sequence tags in that 1 Mb segment). The number of mapped tags in a putative binding region is compared to the normalized (normalized by correlating tag counts in genomic 10 kb windows) number of mapped tags in the same region from an input DNA control. Using a binomial test, only regions that have a p-value = 0.01 are considered to be significantly enriched compared to the input DNA control. Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Fresh tissues were harvested from mice and the nuclei prepared according to the tissue appropriate protocol (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Digital DNaseI was performed by DNaseI digestion of intact nuclei, isolating DNaseI 'double-hit' fragments as described in Sabo et al. (2006), and direct sequencing of fragment ends (which correspond to in vivo DNaseI cleavage sites) using the Illumina IIx (and Illumina HiSeq by early 2011) platform (36 bp reads). Uniquely mapping high-quality reads were mapped to the genome using the bowtie aligner. DNaseI sensitivity is directly reflected in raw tag density, which is shown in the track as density of tags mapping within a 150 bp sliding window (at a 20 bp step across the genome). DNaseI sensitive zones (HotSpots) were identified using the HotSpot algorithm described in Sabo et al. (2004). 1.0% false discovery rate thresholds (FDR 0.01) were computed for each cell type by applying the HotSpot algorithm to an equivalent number of random uniquely mapping 36mers. DNaseI hypersensitive sites (DHSs or Peaks) were identified as signal peaks within FDR 1.0% hypersensitive zones using a peak-finding algorithm (I-max). Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Fresh tissues were harvested from mice and the nuclei prepared according to the tissue appropriate protocol (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Reads were aligned to mm9 reference using ABI BioScope software version 1.2.1. Colorspace FASTQ format files were created using Heng Li's solid2fastq.pl script version 0.1.4, representing 0,1,2,3 color codes with the letters A,C,G,T respectively. Signal files were created from the BAM alignments using BEDTools.

Project description:Rationale for the Mouse ENCODE project Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Maps of occupancy of genomic DNA by transcription factors (TFs) are determined by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for TF binding are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). The chromatin immunoprecipitation followed published methods (Welch et al. 2004). Information on antibodies used is available via the hyperlinks in the "Select subtracks" menu. Samples passing initial quality thresholds (enrichment and depletion for positive and negative controls - if available - by quantitative PCR of ChIP material) are processed for library construction for Illumina sequencing, using the ChIP-seq Sample Preparation Kit purchased from Illumina. Starting with a 10 ng sample of ChIP DNA, DNA fragments were repaired to generate blunt ends and a single A nucleotide was added to each end. Double-stranded Illumina adaptors were ligated to the fragments. Ligation products were amplified by 18 cycles of PCR, and the DNA between 250-350 bp was gel purified. Completed libraries were quantified with Quant-iT dsDNA HS Assay Kit. The DNA library was sequenced on the Illumina Genome Analyzer II sequencing system, and more recently on the HiSeq. Cluster generation, linearization, blocking and sequencing primer reagents were provided in the Illumina Cluster Amplification kits. All samples are being determined as biological replicates except time course samples. The data displayed are from the pooled reads for all replicates, but individual replicates are available by download. The resulting 36-nucleotide sequence reads (fastq files) were moved to a data library in Galaxy, and the tools implemented in Galaxy were used for further processing via workflows (Blankenberg et al. 2010). The reads were mapped to the mouse genome (mm9 assembly) using the program bowtie (Langmead et al. 2009), and the files of mapped reads for the ChIP sample and from the "input" control (no antibody) were processed by MACs (Zhang et al. 2008) to call peaks for occupancy by transcription factors, using the parameters mfold=15, bandwidth=125. Per-replicate aligments and sequences are available for download at downloads page (http://hgdownload.cse.ucsc.edu/goldenPath/mm9/encodeDCC/wgEncodePsuTfbs/).

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Barbara Wold mailto:woldb@caltech.edu, Georgi K. Marinov mailto:georgi@caltech.edu, Diane Trout mailto:diane@caltech.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). Rationale for the Mouse ENCODE project Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus, we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Genome-wide occupancy maps of transcription factors (TFs) are generated by ChIP-seq. A ChIP-Seq experiment combines a chromatin immunoprecipitation (ChIP) experiment that enriches genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibody) with high-throughput short read sequencing of the enriched DNA fragments (Wold & Myers, 2008). Proteins are crosslinked to DNA (usually with formaldehyde), chromatin is sheared and immunoprecipitated with the antibody of interest. The immunoprecipitated material is turned into a sequencing library and sequenced. The sequencing reads are then aligned to the genome. A control sample consisting of sonicated chromatin that has not been immunoprecipitated or immunoprecipitated with a non-specific immunoglobulin is also sequenced. The ChIP and the control datasets are analyzed with a variety of software packages to identify regions occupied by the target protein. The sequencing data, alignments and analysis files for these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Cells were grown according to the approved ENCODE cell culture protocols (http://genome-test.cse.ucsc.edu/ENCODE/protocols/cell/mouse). Chromatin immunoprecipitation followed published methods (Johnson & Mortazavi et al., 2007) with the exception of certain experiments for which glutaraldehyde was added to the crosslink reaction. Information on the antibodies used is available via the metadata for each subtrack. Libraries were constructed using the Illumina ChIP-seq Sample Preparation Kit or using a modified protocol that includes the addition of multiplexing tags to the fragments. DNA fragments were repaired to generate blunt ends and a single A nucleotide was added to each end. Double-stranded Illumina adaptors or Double-stranded Illumina adaptors with multiplexing tags were ligated to the fragments. Ligation products were amplified by 18 cycles of PCR, and the DNA between 150-250 bp was gel purified. Completed libraries were quantified with Quant-iT dsDNA HS Assay Kit. The DNA library was sequenced on the Illumina GAII and GAIIx sequencing systems, and more recently, for multiplexed libraries, several of them were pooled and sequenced on the HiSeq platform. Cluster generation, linearization, blocking and sequencing primer reagents were provided in the Illumina Cluster Amplification kits. Older libraries were generated using 2 rounds of PCR. Matched input samples were sequenced for each variation of fixation conditions and the number of PCR rounds. Reads of 32 bp, 36 bp or 50 bp length were generated. Sequencing reads (fastq files) were assigned to the corresponding libraries based on the multiplexing tag for pooled libraries (all tags have been removed from reads in the fastq files available for download) or directly processed. Bowtie (Langmead et al., 2009) was used to map reads to the male or female version of the mouse genome (excluding the _random chromosomes in the assembly) depending on the cell line sex. The following parameters were used: "-v 2 -k 11 -m 10 -t --best --strata". Aligned reads were converted into rds files using the ERANGE package (Johnson & Mortazavi et al., 2007) and the findall.py program in ERANGE was used to identify enriched regions against the matching input sample. The following settings were used for point-source transcription factors: "--shift learn --ratio 3 --minimum 2 --listPeak --revbackground". For histone modifications, the settings were changed to "--notrim --nodirectionality --spacing 100 --ratio 3 --minimum 2 --listPeak --revbackground".

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Barbara Wold mailto:woldb@caltech.edu, Georgi K. Marinov mailto:georgi@caltech.edu, Diane Trout mailto:diane@caltech.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus, we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Genome-wide occupancy maps of transcription factors (TFs) are generated by ChIP-seq. A ChIP-Seq experiment combines a chromatin immunoprecipitation (ChIP) experiment that enriches genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibody) with high-throughput short read sequencing of the enriched DNA fragments (Wold & Myers, 2008). Proteins are crosslinked to DNA (usually with formaldehyde), chromatin is sheared and immunoprecipitated with the antibody of interest. The immunoprecipitated material is turned into a sequencing library and sequenced. The sequencing reads are then aligned to the genome. A control sample consisting of sonicated chromatin that has not been immunoprecipitated or immunoprecipitated with a non-specific immunoglobulin is also sequenced. The ChIP and the control datasets are analyzed with a variety of software packages to identify regions occupied by the target protein. The sequencing data, alignments and analysis files for these experiments are available for download. In specific, the Ren lab examined RNA polymerase II (PolII), co-activator protein p300, the insulator protein CTCF, and two chromatin modification marks, H3K4me3 and H3K4me1, due to their demonstrated utilities in identifying promoters, enhancers and insulator elements (Barski et al., 2007; Blow et al., 2010; Heintzman et al., 2009; Kim et al., 2007; Kim et al., 2005a; Visel et al., 2009). Enrichment of H3K4me3 or PolII signals is a strong indicator of an active promoter, while the presence of p300 or H3K4me1 outside of promoter regions has been used as a mark for enhancers. CTCF binding sites are considered as a mark for potential insulator elements. For each transcription factor or chromatin mark in each tissue, ChIP-seq was carried out with at least two biological replicates. Each experiment produced 20-30 million monoclonal, uniquely mapped tags. Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of histone modifications Levels of three histone modifications are being determined. H3K4me1 (monomethylation of lysine 4 of histone H3) is a mark for active chromatin and in the absence of H3K4me3, it is one indicator of an enhancer. H3K4me3 (trimethylation of lysine 4 of histone H3) is highly enriched at active promoters. One repressive (Polycomb) mark, H3K27me3, is associated with some silenced genes. Maps of genomic DNA in chromatin with these histone modifications are generated by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for histone modifications are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of Occupancy by Transcription Factors Maps of occupancy of genomic DNA by transcription factors (TFs) are determined by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for TF binding are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Ross Hardison mailto:rch8@psu.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). Rationale for the Mouse ENCODE project Our knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. These are inferred to be under negative or positive selection, respectively, and we interpret these as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. Thus we will be able to discover which epigenetic features are conserved between mouse and human, and we can examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of DNA that with a function preserved in mammals versus that with a function in only one species will be discovered. The results will have a significant impact on our understanding of the evolution of gene regulation. Maps of histone modifications Levels of three histone modifications are being determined. H3K4me1 (monomethylation of lysine 4 of histone H3) is a mark for active chromatin and in the absence of H3K4me3, it is one indicator of an enhancer. H3K4me3 (trimethylation of lysine 4 of histone H3) is highly enriched at active promoters. One repressive (Polycomb) mark, H3K27me3, is associated with some silenced genes. Maps of genomic DNA in chromatin with these histone modifications are generated by ChIP-seq. This consists of two basic steps: chromatin immunoprecipitation (ChIP) is used to highly enrich genomic DNA for the segments bound by specific proteins (the antigens recognized by the antibodies) followed by massively parallel short read sequencing to tag the enriched DNA segments. Sequencing is done on the Illumina GAIIx and HiSeq. The sequence tags are mapped back to the mouse genome (Langmead et al. 2009), and a graph of the enrichment for histone modifications are displayed as the "Signal" track (essentially the counts of mapped reads per interval) and the deduced probable binding sites from the MACS program (Zhang et al. 2008) are shown in the "Peaks" track. Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download.

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Ross Hardison mailto:rch8@psu.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). Rationale for the Mouse ENCODE project: Knowledge of the function of genomic DNA sequences comes from three basic approaches. Genetics uses changes in behavior or structure of a cell or organism in response to changes in DNA sequence to infer function of the altered sequence. Biochemical approaches monitor states of histone modification, binding of specific transcription factors, accessibility to DNases and other epigenetic features along genomic DNA. In general, these features are associated with gene activity, but the precise relationships remain to be established. The third approach is evolutionary, using comparisons among homologous DNA sequences to find segments that are evolving more slowly or more rapidly than expected given the local rate of neutral change. Such changes are inferred to be under negative or positive selection, respectively, and interpreted as DNA sequences needed for a preserved (negative selection) or adaptive (positive selection) function. The ENCODE project aims to discover all the DNA sequences associated with various epigenetic features, with the reasonable expectation that these will also be functional (best tested by genetic methods). However, it is not clear how to relate these results with those from evolutionary analyses. The mouse ENCODE project aims to make this connection explicitly and with a moderate breadth. Assays identical to those being used in the ENCODE project are performed in cell types in mouse that are similar or homologous to those studied in the human project. The comparison will be used to discover which epigenetic features are conserved between mouse and human, and examine the extent to which these overlap with the DNA sequences under negative selection. The contribution of functional DNA preserved in mammals versus function in only one species will be discovered. The results will have a significant impact on the understanding of the evolution of gene regulation. Maps of DNaseI Sensitivity: DNaseI has long been used to map general chromatin accessibility, and DNaseI hypersensitivity is a universal feature of active cis-regulatory sequences. Maps of DNaseI sensitivity measured genome-wide are generated through DNaseI digestion, addition of linkers at the sites of cleavage, and library prep followed by massively parallel short read sequencing on the Illumina GAIIx and HiSeq platforms. The sequence tags are mapped back to the mouse genome, and a graph of the smoothed kernel density of DNaseI cleavage sites is displayed as the "Signal" track. This provides a quantitative estimate of the frequency of cleavage by DNaseI in the initial digest, which in turn is related to the accessibility of the DNA in the chromatin. Segments of greatest cleavage site density represent DNase hypersensitive sites (DHSs) and are identified as peaks by the F-seq program (Boyle et al. 2008). DHSs are candidates for any cis-regulatory module, including promoters, enhancers, insulators, and novel elements. The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Cells were grown and harvested according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse) for G1E and G1E-ER4. DNaseI hypersensitive sites were isolated using methods called DNase-seq or DNase-chip (Song and Crawford, 2010). Briefly, cells were lysed with NP40, and intact nuclei were digested with optimal levels of DNaseI enzyme. DNaseI-digested ends were captured from three different DNase concentrations, and material was sequenced using Illumina sequencing. The read length for sequences from DNase-seq is 20 bases long due to a MmeI cutting step of the approximately 50 kb DNA fragments extracted after DNaseI digestion. Sequences from each experiment were mapped to the mouse genome (mm9 assembly) using the program Bowtie (http://bowtie-bio.sourceforge.net/index.shtml) (Langmead et al., 2009). Reads mapping to more than one location were not removed. For such reads, only the best mapping result was used ("--best" option). Sequences from multiple lanes were combined for a single replicate and converted to the sam/bam format using SAMtools (http://samtools.sourceforge.net/). Using F-seq, the resulting digital signal was converted to a continuous wiggle track that employs a Parzen kernel density estimation to create base pair scores (Boyle et al., 2008). Discrete DNaseI HS sites (peaks) were identified from the DNase-seq F-seq density signal. Significant regions were determined by fitting the data to a gamma distribution to calculate p-values.

Project description:Nucleosomes are part of the first level of chromatin packaging. They each consist of a histone heterooctamer around which DNA wraps 1.6 times. The histone heterooctomamer is made up of two copies of histones 2A, 2B, 3 and 5. The segment of DNA wrapped around the histones, the so-called "core" fragment, is 147 base pairs long. Neighboring nucleosomes are separated from one another by a stretch of DNA called the "linker," whose size varies depending on organism, cell type, and even chromatin activity. Certain chromatin remodeling factors govern accessibility of DNA to regulatory proteins by repositioning nucleosomes to reveal regulatory sites that would otherwise be occluded by a nucleosome. In contrast to histone modifications such as methylation or acetylation, which are investigated by ChIP-seq, nucleosome positioning data are generated without immunoprecipitation (see Methods below). Instead, micrococcal nuclease is used to digest chromatin to apparent completion, the (well-defined and clearly visible) mononucleosomal core fragment fraction is isolated by gel purification, and one end is then sequenced. Mapping the sequence tag back to the genome reveals the precise position of one end of the core fragment that was protected by the nucleosome; the position of the other end can then simply be inferred by extending the read to a virtual length of 147 bases. Statistical analyses such as occupancy and positioning stringency can then be employed to analyze the local nucleosome landscape anywhere in the (mappable) genome or to infer global parameters of nucleosome organization. In the context of the ENCODE project, nucleosome positioning data are particularly valuable for analysis of the relationship between transcription factor binding, histone modifications, and gene activity. For a general primer on these types of data and analyses, refer to Valouev et al. (2008). For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf To isolate mononucleosome core DNA fragments from the GM12878 and K562 ENCODE cell lines we followed the micrococcal nuclease (MNase) digestion and isolation protocol as described in Johnson et al. (2006), Valouev et al. (2008), and Valouev et al. (2011) with the following modifications. The precise concentrations of the two flash-frozen cell samples received from the Snyder Lab were not known so, per our standard procedure, we performed a series of digestions titrating the amount of MNase to determine the concentration of MNase for optimal digestion of each sample. Final concentrations of 25 U/µL and 50 U/µL of MNase were used to digest the GM12878 cells and K562 cells respectively at 20°C for 12 min. All other steps in the digestion and isolation protocol were as described. Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell). K562 and GM12878 were each grown to ~2.5×108 cells. The cells were harvested, frozen and the nucleosome core isolation followed (Valouev et al. 2008). The SOLiD reads were mapped in color-space with the probabilistic mapper, DNAnexus (https://dnanexus.com/). The DNAnexus mapper measures and propagates mapping uncertainty by including both quality values and mismatches in the alignment score calculation. The scores are then scaled across all possible mappings of the read to estimate the posterior probability for alignment to each genomic location. Reads corresponding to posterior probability of correct mapping > 0.9 were reported. Nucleosome density signal maps (bedgraph and bigwig files) were generated by first shifting reads by 74 bp in the 5´ to 3´ direction and counting the total number of reads starting at each genomic coordinate on both strands. These counts are then smoothed using un-normalized kernel density smoothing with a triweight kernel. A bandwidth of 30 bp is used which is equivalent to a smoothing window of 60 bp. The smoothed counts at each position are then divided by the expected number of reads from an equivalent uniform distribution of reads in a ± 30 bp window around that position. If less than 25% of the positions in a ± 30 bp window around a genomic location are uniquely mappable or if the location is part of an assembly gap, the signal value at that position is considered unreliable and not recorded in the signal files. Hence, genomic coordinates that do not have any associated signal value should be considered missing or unreliable data. Genomic coordinates associated with a signal value of 0 are reliably mapable but do not have any signal in the dataset.

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Philip Cayting mailto:pcayting@stanford.edu). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). This track shows probable binding sites of the specified transcription factors (TFs) in the given cell types as determined by chromatin immunoprecipitation followed by high throughput sequencing (ChIP-seq). Each experiment is associated with an input signal, which represents the control condition where immunoprecipitation with non-specific immunoglobulin was performed in the same cell type. For each experiment (cell type vs. antibody) this track shows a graph of enrichment for TF binding (Signal), along with sites that have the greatest evidence of transcription factor binding, as identified by the PeakSeq algorithm (Peaks). The sequence reads, quality scores, and alignment coordinates from these experiments are available for download. For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell/mouse). For details on the chromatin immunoprecipitation protocol used, see Euskirchen et. al., (2007), Rozowsky et. al. (2009) and Auerbach et. al. (2009). DNA recovered from the precipitated chromatin was sequenced on the Illumina (Solexa) sequencing platform and mapped to the genome using the Eland alignment program. ChIP-seq data was scored based on sequence reads (length ~30 bps) that align uniquely to the human genome. From the mapped tags, a signal map of ChIP DNA fragments (average fragment length ~ 200 bp) was constructed where the signal height is the number of overlapping fragments at each nucleotide position in the genome. Reads were pooled from all submitted replicates to generate the Peak and Signal files. Per-replicate aligments and sequences are available for download at downloads page (http://hgdownload.cse.ucsc.edu/goldenPath/mm9/encodeDCC/wgEncodeSydhTfbs/). For each 1 Mb segment of each chromosome, a peak height threshold was determined by requiring a false discovery rate <= 0.01 when comparing the number of peaks above said threshold to the number of peaks obtained from multiple simulations of a random null background with the same number of mapped reads (also accounting for the fraction of mapable bases for sequence tags in that 1 Mb segment). The number of mapped tags in a putative binding region is compared to the normalized (normalized by correlating tag counts in genomic 10 kb windows) number of mapped tags in the same region from an input DNA control. Using a binomial test, only regions that have a p-value = 0.01 are considered to be significantly enriched compared to the input DNA control.

Project description:This data was generated by ENCODE. If you have questions about the data, contact the submitting laboratory directly (Florencia Pauli mailto:fpauli@hudsonalpha.org). If you have questions about the Genome Browser track associated with this data, contact ENCODE (mailto:genome@soe.ucsc.edu). This track is produced as part of the ENCODE Project. RNA-seq is a method for mapping and quantifying the transcriptome of any organism that has a genomic DNA sequence assembly (Mortazavi et al., 2008). Biological replicates of ENCODE cell lines were grown on separate culture plates, total RNA was purified and polyA selected two times. mRNA was then fragmented by magnesium-catalyzed hydrolysis, reverse transcribed to cDNA by random priming and amplified. The cDNA was sequenced on an Illumina Genome Analyzer (GAI or GAIIx). The DNA sequences were aligned to the NCBI Build37 (hg19) version of the human genome using the sequence alignment programs ELAND (Illumina) or Bowtie (Langmead et al., 2009). The first 10 residues of sequencing have a weak characteristic nucleotide bias of unknown origin. This RNA-seq protocol does not specify the coding strand. As a result, there will be ambiguity at loci where both strands are transcribed. This is the first NCBI Build37 (hg19) release of this track (Jan 2012). This release includes the 3 datasets (Jurkat, A549/DEX100nm, and A549/EtOH2pct) previously released on NCBI Build36 (hg18) and adds data for several more cell types and growth conditions in replicate. Four types of download files are available for each replicate including the Raw Data (fastq), Transcripts GencodeV7 (gtf), Raw Signal (bigwig), and Alignments (bam). For data usage terms and conditions, please refer to http://www.genome.gov/27528022 and http://www.genome.gov/Pages/Research/ENCODE/ENCODEDataReleasePolicyFinal2008.pdf Experimental Procedures Cells were grown according to the approved ENCODE cell culture protocols (http://hgwdev.cse.ucsc.edu/ENCODE/protocols/cell) except for H1-hESC for which frozen cell pellets were purchased from Cellular Dynamics. Cells were lysed in RLT buffer (Qiagen RNEasy kit) and processed on RNEasy midi columns according to the manufacturer's protocol, with the inclusion of the "on-column" DNase digestion step to remove residual genomic DNA. mRNA was isolated from at least 10 ug of total RNA with oligo(dT) two times (Dynabeads mRNA PurificationgKit, Invitrogen). Alternatively, cells were lysed and mRNA was purified directly two times with oligo(dT) (Dynabeads mRNA DIRECT Kit, Invitrogen). 100 ng of mRNA was fragmented by magnesium-catalyzed hydrolysis and reverse transcribed to cDNA by random priming according to the protocol in Mortazavi et al. (2008). cDNA was prepared for sequencing on the Genome Analyzer flowcell according to the protocol for the ChIPSeq DNA genomic DNA kit (Illumina). The sequencing libraries were size-selected around 225 bp and amplified with 15 rounds of PCR. Libraries were sequenced with an Illumina Genome Analyzer I or an Illumina Genome Analyzer IIx according to the manufacturer's recommendations. Single end reads of 36 nt in length were obtained. Data Processing and Analysis Fastq files were made from qseq files generated by the Illumina pipeline (Casava 1.7). The Raw Signal files (bigWig) were generated from bedgraph files and the score was calculated as the number of reads at that position divided by the total number of reads divided by one million. Casava export files were aligned to the NCBI Build37 (hg19) version of the human genome with ELAND (Illumina), generating SAM files. Fastq files of experiments that were previously aligned to NCBI Build36 (hg18) were aligned to NCBI Build37 (hg19) using Bowtie (Langmead et al., 2009; parameters: -S -n 2 -k 11 -m 10 --best), also generating SAM files. SAM files were converted to BAM with SAMtools (Li et al., 2009). Gene expression within Gencode.v7 (Harrow et al., 2006) gene models was estimated using Cufflinks v0.9.3 (Roberts et al., 2011). Estimates of transcript abundance were reported in Fragments Per Kilobase of exon per Million fragments mapped (FPKM). FPKM is calculated by dividing the total number of fragments that align to the gene model by the size of the spliced transcript (exons) in kilobases. This number is then divided by the total number of reads in millions for the experiment. FPKM is reported in the last column of the gtf (TranscriptGencV7) files. Raw Data (fastq), Raw Signal (bigWig), Alignments (bam) and Transcript Gencode V7 (gtf) files are available from the Downloads (http://hgwdev.cse.ucsc.edu/cgi-bin/hgFileUi?g=wgEncodeHaibRnaSeq) page.

Dataset Information

Histone Modifications by ChIP-seq from ENCODE/PSU

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets