Person: Zang, Chongzhi
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication MethylPurify: tumor purity deconvolution and differential methylation detection from single tumor DNA methylomes
(BioMed Central, 2014) Zheng, Xiaoqi; Zhao, Qian; Wu, Hua-Jun; Li, Wei; Wang, Haiyun; Meyer, Clifford; Qin, Qian Alvin; Xu, Han; Zang, Chongzhi; Jiang, Peng; Li, Fuqiang; Hou, Yong; He, Jianxing; Wang, Jun; Zhang, Peng; Zhang, Yong; Liu, XiaoleWe propose a statistical algorithm MethylPurify that uses regions with bisulfite reads showing discordant methylation levels to infer tumor purity from tumor samples alone. MethylPurify can identify differentially methylated regions (DMRs) from individual tumor methylome samples, without genomic variation information or prior knowledge from other datasets. In simulations with mixed bisulfite reads from cancer and normal cell lines, MethylPurify correctly inferred tumor purity and identified over 96% of the DMRs. From patient data, MethylPurify gave satisfactory DMR calls from tumor methylome samples alone, and revealed potential missed DMRs by tumor to normal comparison due to tumor heterogeneity. Electronic supplementary material The online version of this article (doi:10.1186/s13059-014-0419-x) contains supplementary material, which is available to authorized users.
Publication Analysis of optimized DNase-seq reveals intrinsic bias in transcription factor footprint identification
(2014) He, Housheng Hansen; Meyer, Clifford; Hu, Sheng'en Shawn; Chen, Mei-Wei; Zang, Chongzhi; Liu, Yin; Rao, Prakash K.; Fei, Teng; Xu, Han; Long, Henry; Liu, X. Shirley; Brown, MylesDNase-seq is a powerful technique for identifying cis-regulatory elements across the genome. We studied the key experimental parameters to optimize the performance of DNase-seq. We found that sequencing short 50-100bp fragments that accumulate in long inter-nucleosome linker regions is more efficient for identifying transcription factor binding sites than using longer fragments. We also assessed the potential of DNase-seq to predict transcription factor occupancy through the generation of nucleotide-resolution transcription factor footprints. In modeling the sequence-specific DNaseI cutting bias we found a surprisingly strong effect that varied over more than two orders of magnitude. This confounds DNaseI footprint analysis to the extent that the nucleotide resolution cleavage patterns at most transcription factor binding sites are derived from intrinsic DNaseI cleavage bias rather than from specific protein-DNA interactions. In contrast, quantitative comparison of DNaseI hypersensitivity between states can predict transcription factor occupancy associated with particular biological perturbations.
Publication Cistrome Data Browser: a data portal for ChIP-Seq and chromatin accessibility data in human and mouse
(Oxford University Press, 2017) Mei, Shenglin; Qin, Qian; Wu, Qiu; Sun, Hanfei; Zheng, Rongbin; Zang, Chongzhi; Zhu, Muyuan; Wu, Jiaxin; Shi, Xiaohui; Taing, Len; Liu, Tao; Brown, Myles; Meyer, Clifford; Liu, X. ShirleyChromatin immunoprecipitation, DNase I hypersensitivity and transposase-accessibility assays combined with high-throughput sequencing enable the genome-wide study of chromatin dynamics, transcription factor binding and gene regulation. Although rapidly accumulating publicly available ChIP-seq, DNase-seq and ATAC-seq data are a valuable resource for the systematic investigation of gene regulation processes, a lack of standardized curation, quality control and analysis procedures have hindered extensive reuse of these data. To overcome this challenge, we built the Cistrome database, a collection of ChIP-seq and chromatin accessibility data (DNase-seq and ATAC-seq) published before January 1, 2016, including 13 366 human and 9953 mouse samples. All the data have been carefully curated and processed with a streamlined analysis pipeline and evaluated with comprehensive quality control metrics. We have also created a user-friendly web server for data query, exploration and visualization. The resulting Cistrome DB (Cistrome Data Browser), available online at http://cistrome.org/db, is expected to become a valuable resource for transcriptional and epigenetic regulation studies.
Publication ChiLin: a comprehensive ChIP-seq and DNase-seq quality control and analysis pipeline
(BioMed Central, 2016) Qin, Qian; Mei, Shenglin; Wu, Qiu; Sun, Hanfei; Li, Lewyn; Taing, Len; Chen, Sujun; Li, Fugen; Liu, Tao; Zang, Chongzhi; Xu, Han; Chen, Yiwen; Meyer, Clifford; Zhang, Yong; Brown, Myles; Long, Henry W.; Liu, X. ShirleyBackground: Transcription factor binding, histone modification, and chromatin accessibility studies are important approaches to understanding the biology of gene regulation. ChIP-seq and DNase-seq have become the standard techniques for studying protein-DNA interactions and chromatin accessibility respectively, and comprehensive quality control (QC) and analysis tools are critical to extracting the most value from these assay types. Although many analysis and QC tools have been reported, few combine ChIP-seq and DNase-seq data analysis and quality control in a unified framework with a comprehensive and unbiased reference of data quality metrics. Results: ChiLin is a computational pipeline that automates the quality control and data analyses of ChIP-seq and DNase-seq data. It is developed using a flexible and modular software framework that can be easily extended and modified. ChiLin is ideal for batch processing of many datasets and is well suited for large collaborative projects involving ChIP-seq and DNase-seq from different designs. ChiLin generates comprehensive quality control reports that include comparisons with historical data derived from over 23,677 public ChIP-seq and DNase-seq samples (11,265 datasets) from eight literature-based classified categories. To the best of our knowledge, this atlas represents the most comprehensive ChIP-seq and DNase-seq related quality metric resource currently available. These historical metrics provide useful heuristic quality references for experiment across all commonly used assay types. Using representative datasets, we demonstrate the versatility of the pipeline by applying it to different assay types of ChIP-seq data. The pipeline software is available open source at https://github.com/cfce/chilin. Conclusion: ChiLin is a scalable and powerful tool to process large batches of ChIP-seq and DNase-seq datasets. The analysis output and quality metrics have been structured into user-friendly directories and reports. We have successfully compiled 23,677 profiles into a comprehensive quality atlas with fine classification for users. Electronic supplementary material The online version of this article (doi:10.1186/s12859-016-1274-4) contains supplementary material, which is available to authorized users.
Publication Network analysis of gene essentiality in functional genomics experiments
(BioMed Central, 2015) Jiang, Peng; Wang, Hongfang; Li, Wei; Zang, Chongzhi; Li, Bo; Wong, Yinling J.; Meyer, Cliff; Liu, Jun; Aster, Jon; Liu, X. ShirleyMany genomic techniques have been developed to study gene essentiality genome-wide, such as CRISPR and shRNA screens. Our analyses of public CRISPR screens suggest protein interaction networks, when integrated with gene expression or histone marks, are highly predictive of gene essentiality. Meanwhile, the quality of CRISPR and shRNA screen results can be significantly enhanced through network neighbor information. We also found network neighbor information to be very informative on prioritizing ChIP-seq target genes and survival indicator genes from tumor profiling. Thus, our study provides a general method for gene essentiality analysis in functional genomic experiments (http://nest.dfci.harvard.edu). Electronic supplementary material The online version of this article (doi:10.1186/s13059-015-0808-9) contains supplementary material, which is available to authorized users.
Publication NF-E2, FLI1 and RUNX1 collaborate at areas of dynamic chromatin to activate transcription in mature mouse megakaryocytes
(Nature Publishing Group, 2016) Zang, Chongzhi; Luyten, Annouck; Chen, Justina; Liu, X. Shirley; Shivdasani, RameshMutations in mouse and human Nfe2, Fli1 and Runx1 cause thrombocytopenia. We applied genome-wide chromatin dynamics and ChIP-seq to determine these transcription factors’ (TFs) activities in terminal megakaryocyte (MK) maturation. Enhancers with H3K4me2-marked nucleosome pairs were most enriched for NF-E2, FLI and RUNX sequence motifs, suggesting that this TF triad controls much of the late MK program. ChIP-seq revealed NF-E2 occupancy near previously implicated target genes, whose expression is compromised in Nfe2-null cells, and many other genes that become active late in MK differentiation. FLI and RUNX were also the motifs most enriched near NF-E2 binding sites and ChIP-seq implicated FLI1 and RUNX1 in activation of late MK, including NF-E2-dependent, genes. Histones showed limited activation in regions of single TF binding, while enhancers that bind NF-E2 and either RUNX1, FLI1 or both TFs gave the highest signals for TF occupancy and H3K4me2; these enhancers associated best with genes activated late in MK maturation. Thus, three essential TFs co-occupy late-acting cis-elements and show evidence for additive activity at genes responsible for platelet assembly and release. These findings provide a rich dataset of TF and chromatin dynamics in primary MK and explain why individual TF losses cause thrombopocytopenia.
Publication Partitioning heritability by functional annotation using genome-wide association summary statistics
(2015) Finucane, Hilary; Bulik-Sullivan, Brendan; Gusev, Alexander; Trynka, Gosia; Reshef, Yakir; Loh, Po-Ru; Anttila, Verneri; Xu, Han; Zang, Chongzhi; Farh, Kyle; Ripke, Stephan; Day, Felix R.; Consortium, ReproGen; Purcell, Shaun M.; Stahl, Eli; Lindstrom, Sara; Perry, John R. B.; Okada, Yukinori; Raychaudhuri, Soumya; Daly, Mark; Patterson, Nick; Neale, Benjamin; Price, AlkesRecent work has demonstrated that some functional categories of the genome contribute disproportionately to the heritability of complex diseases. Here, we analyze a broad set of functional elements, including cell-type-specific elements, to estimate their polygenic contributions to heritability in genome-wide association studies (GWAS) of 17 complex diseases and traits with an average sample size of 73,599. To enable this analysis, we introduce a new method, stratified LD score regression, for partitioning heritability from GWAS summary statistics while accounting for linked markers. This new method is computationally tractable at very large sample sizes, and leverages genome-wide information. Our results include a large enrichment of heritability in conserved regions across many traits; a very large immunological disease-specific enrichment of heritability in FANTOM5 enhancers; and many cell-type-specific enrichments including significant enrichment of central nervous system cell types in body mass index, age at menarche, educational attainment, and smoking behavior.
Publication High-dimensional genomic data bias correction and data integration using MANCIE
(Nature Publishing Group, 2016) Zang, Chongzhi; Wang, Tao; Deng, Ke; Li, Bo; Hu, Sheng'en; Qin, Qian; Xiao, Tengfei; Zhang, Shihua; Meyer, Clifford; He, Housheng Hansen; Brown, Myles; Liu, Jun; Xie, Yang; Liu, X. ShirleyHigh-dimensional genomic data analysis is challenging due to noises and biases in high-throughput experiments. We present a computational method matrix analysis and normalization by concordant information enhancement (MANCIE) for bias correction and data integration of distinct genomic profiles on the same samples. MANCIE uses a Bayesian-supported principal component analysis-based approach to adjust the data so as to achieve better consistency between sample-wise distances in the different profiles. MANCIE can improve tissue-specific clustering in ENCODE data, prognostic prediction in Molecular Taxonomy of Breast Cancer International Consortium and The Cancer Genome Atlas data, copy number and expression agreement in Cancer Cell Line Encyclopedia data, and has broad applications in cross-platform, high-dimensional data integration.