Publication: ChiLin: a comprehensive ChIP-seq and DNase-seq quality control and analysis pipeline
Open/View Files
Date
2016
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
BioMed Central
The Harvard community has made this article openly available. Please share how this access benefits you.
Citation
Qin, Q., S. Mei, Q. Wu, H. Sun, L. Li, L. Taing, S. Chen, et al. 2016. “ChiLin: a comprehensive ChIP-seq and DNase-seq quality control and analysis pipeline.” BMC Bioinformatics 17 (1): 404. doi:10.1186/s12859-016-1274-4. http://dx.doi.org/10.1186/s12859-016-1274-4.
Research Data
Abstract
Background: Transcription factor binding, histone modification, and chromatin accessibility studies are important approaches to understanding the biology of gene regulation. ChIP-seq and DNase-seq have become the standard techniques for studying protein-DNA interactions and chromatin accessibility respectively, and comprehensive quality control (QC) and analysis tools are critical to extracting the most value from these assay types. Although many analysis and QC tools have been reported, few combine ChIP-seq and DNase-seq data analysis and quality control in a unified framework with a comprehensive and unbiased reference of data quality metrics. Results: ChiLin is a computational pipeline that automates the quality control and data analyses of ChIP-seq and DNase-seq data. It is developed using a flexible and modular software framework that can be easily extended and modified. ChiLin is ideal for batch processing of many datasets and is well suited for large collaborative projects involving ChIP-seq and DNase-seq from different designs. ChiLin generates comprehensive quality control reports that include comparisons with historical data derived from over 23,677 public ChIP-seq and DNase-seq samples (11,265 datasets) from eight literature-based classified categories. To the best of our knowledge, this atlas represents the most comprehensive ChIP-seq and DNase-seq related quality metric resource currently available. These historical metrics provide useful heuristic quality references for experiment across all commonly used assay types. Using representative datasets, we demonstrate the versatility of the pipeline by applying it to different assay types of ChIP-seq data. The pipeline software is available open source at https://github.com/cfce/chilin. Conclusion: ChiLin is a scalable and powerful tool to process large batches of ChIP-seq and DNase-seq datasets. The analysis output and quality metrics have been structured into user-friendly directories and reports. We have successfully compiled 23,677 profiles into a comprehensive quality atlas with fine classification for users. Electronic supplementary material The online version of this article (doi:10.1186/s12859-016-1274-4) contains supplementary material, which is available to authorized users.
Description
Other Available Sources
Keywords
ChIP-seq, DNase-seq, Quality atlas, Analysis pipeline
Terms of Use
This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service