Person: Liu, Xiaole
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Sequence determinants of improved CRISPR sgRNA design
(Cold Spring Harbor Laboratory Press, 2015) Xu, Han; Xiao, Tengfei; Chen, Chen-Hao; Li, Wei; Meyer, Clifford; Wu, Qiu; Wu, Di; Cong, L; Zhang, Feng; Liu, Jun; Brown, Myles; Liu, XiaoleThe CRISPR/Cas9 system has revolutionized mammalian somatic cell genetics. Genome-wide functional screens using CRISPR/Cas9-mediated knockout or dCas9 fusion-mediated inhibition/activation (CRISPRi/a) are powerful techniques for discovering phenotype-associated gene function. We systematically assessed the DNA sequence features that contribute to single guide RNA (sgRNA) efficiency in CRISPR-based screens. Leveraging the information from multiple designs, we derived a new sequence model for predicting sgRNA efficiency in CRISPR/Cas9 knockout experiments. Our model confirmed known features and suggested new features including a preference for cytosine at the cleavage site. The model was experimentally validated for sgRNA-mediated mutation rate and protein knockout efficiency. Tested on independent data sets, the model achieved significant results in both positive and negative selection conditions and outperformed existing models. We also found that the sequence preference for CRISPRi/a is substantially different from that for CRISPR/Cas9 knockout and propose a new model for predicting sgRNA efficiency in CRISPRi/a experiments. These results facilitate the genome-wide design of improved sgRNA for both knockout and CRISPRi/a studies.
Publication A suite of web-based programs to search for transcriptional regulatory motifs
(Oxford University Press (OUP), 2004) Liu, Y.; Wei, L.; Batzoglou, S.; Brutlag, D. L.; Liu, Jun; Liu, XiaoleThe identification of regulatory motifs is important for the study of gene expression. Here we present a suite of programs that we have developed to search for regulatory sequence motifs: (i) BioProspector, a Gibbs-sampling-based program for predicting regulatory motifs from co-regulated genes in prokaryotes or lower eukaryotes; (ii) CompareProspector, an extension to BioProspector which incorporates comparative genomics features to be used for higher eukaryotes; (iii) MDscan, a program for finding protein–DNA interaction sites from ChIP-on-chip targets. All three programs examine a group of sequences that may share common regulatory motifs and output a list of putative motifs as position-specific probability matrices, the individual sites used to construct the motifs and the location of each site on the input sequences. The web servers and executables can be accessed at http://seqmotifs.stanford.edu.
Publication Model-Based Analysis of Two-Color Arrays (MA2C)
(BioMed Central, 2007) Zhu, Xiaopeng; Zhang, Xinmin; Chen, Runsheng; Manrai, Arjun K; Song, Jun S; Johnson, W. Evan; Li, Wei; Liu, Xiaole; Liu, JunA novel normalization method based on the GC content of probes is developed for two-color tiling arrays. The proposed method, together with robust estimates of the model parameters, is shown to perform superbly on published data sets. A robust algorithm for detecting peak regions is also formulated and shown to perform well compared to other approaches. The tools have been implemented as a stand-alone Java program called MA2C, which can display various plots of statistical analysis for quality control.
Publication Integrating regulatory motif discovery and genome-wide expression analysis
(Proceedings of the National Academy of Sciences, 2003) Conlon, E. M.; Liu, Xiaole; Lieb, J. D.; Liu, JunWe propose motif regressor for discovering sequence motifs upstream of genes that undergo expression changes in a given condition. The method combines the advantages of matrix-based motif finding and oligomer motif-expression regression analysis, resulting in high sensitivity and specificity. motif regressor is particularly effective in discovering expression-mediating motifs of medium to long width with multiple degenerate positions. When applied to Saccharomyces cerevisiae, motif regressor identified the ROX1 and YAP1 motifs from Rox1p and Yap1p overexpression experiments, respectively; predicted that Gcn4p may have increased activity in YAP1 deletion mutants; reported a group of motifs (including GCN4, PHO4, MET4, STRE, USR1, RAP1, M3A, and M3B) that may mediate the transcriptional response to amino acid starvation; and found all of the known cell-cycle regulation motifs from 18 expression microarrays over two cell cycles.
Publication MM-ChIP enables integrative analysis of cross-platform and between-laboratory ChIP-chip or ChIP-seq data
(Springer Science + Business Media, 2011) Chen, Yiwen; Meyer, Clifford; Liu, Tao; Li, Wei; Liu, Jun; Liu, XiaoleThe ChIP-chip and ChIP-seq techniques enable genome-wide mapping of in vivo protein-DNA interactions and chromatin states. The cross-platform and between-laboratory variation poses a challenge to the comparison and integration of results from different ChIP experiments. We describe a novel method, MM-ChIP, which integrates information from cross-platform and between-laboratory ChIP-chip or ChIP-seq datasets. It improves both the sensitivity and the specificity of detecting ChIP-enriched regions, and is a useful meta-analysis tool for driving discoveries from multiple data sources.
Publication Gene expression profiling of human breast tissue samples using SAGE-Seq
(Cold Spring Harbor Laboratory Press, 2010) Wu, Z. J.; Meyer, Clifford; Choudhury, S.; Shipitsin, M.; Maruyama, R.; Bessarabova, M.; Nikolskaya, T.; Sukumar, S.; Schwartzman, A.; Liu, Jun; Polyak, Kornelia; Liu, XiaoleWe present a powerful application of ultra high-throughput sequencing, SAGE-Seq, for the accurate quantification of normal and neoplastic mammary epithelial cell transcriptomes. We develop data analysis pipelines that allow the mapping of sense and antisense strands of mitochondrial and RefSeq genes, the normalization between libraries, and the identification of differentially expressed genes. We find that the diversity of cancer transcriptomes is significantly higher than that of normal cells. Our analysis indicates that transcript discovery plateaus at 10 million reads/sample, and suggests a minimum desired sequencing depth around five million reads. Comparison of SAGE-Seq and traditional SAGE on normal and cancerous breast tissues reveals higher sensitivity of SAGE-Seq to detect less-abundant genes, including those encoding for known breast cancer-related transcription factors and G protein–coupled receptors (GPCRs). SAGE-Seq is able to identify genes and pathways abnormally activated in breast cancer that traditional SAGE failed to call. SAGE-Seq is a powerful method for the identification of biomarkers and therapeutic targets in human disease.
Publication Inference of transcriptional regulation in cancers
(Proceedings of the National Academy of Sciences, 2015) Jiang, Peng; Freedman, Matthew; Liu, Jun; Liu, XiaoleWe developed an efficient and accurate computational framework, RABIT (regression analysis with background integration), and comprehensively integrated public transcription factor (TF)-binding profiles with TCGA tumor-profiling datasets in 18 cancer types. To systematically search for cancer-associated TFs, RABIT controls the effect of tumor-confounding factors on transcriptional regulation, such as copy number alteration, DNA methylation, and TF somatic mutation. Our predicted TF regulatory activity in tumors is highly consistent with the knowledge from cancer gene databases and reveals many previously unidentified cancer-associated TFs. We also analyzed RNA-binding protein regulation in cancer and demonstrated that RABIT is a general platform for predicting oncogenic gene expression regulators.