Person:

Liu, Jun

Loading...
Profile Picture

Email Address

AA Acceptance Date

Birth Date

Research Projects

Organizational Units

Job Title

Last Name

Liu

First Name

Jun

Name

Liu, Jun

Search Results

Now showing 1 - 7 of 7
  • Publication

    Association pattern discovery via theme dictionary models

    (Wiley-Blackwell, 2013) Deng, Ke; Geng, Zhi; Liu, Jun

    Discovering patterns from a set of text or, more generally, categorical data is an important problem in many disciplines such as biomedical research, linguistics, artificial intelligence and sociology. We consider here the well‐known ‘market basket’ problem that is often discussed in the data mining community, and is also quite ubiquitous in biomedical research. The data under consideration are a set of ‘baskets’, where each basket contains a list of ‘items’. Our goal is to discover ‘themes’, which are defined as subsets of items that tend to co‐occur in a basket. We describe a generative model, i.e. the theme dictionary model, for such data structures and describe two likelihood‐based methods to infer themes that are hidden in a collection of baskets. We also propose a novel sequential Monte Carlo method to overcome computational challenges. Using both simulation studies and real applications, we demonstrate that the new approach proposed is significantly more powerful than existing methods, such as association rule mining and topic modelling, in detecting weak and subtle interactions in the data.

  • Publication

    Bayesian Inference of Spatial Organizations of Chromosomes

    (Public Library of Science, 2013) Hu, Ming; Deng, Ke; Qin, Zhaohui; Dixon, Jesse; Selvaraj, Siddarth; Fang, Jennifer; Ren, Bing; Liu, Jun

    Knowledge of spatial chromosomal organizations is critical for the study of transcriptional regulation and other nuclear processes in the cell. Recently, chromosome conformation capture (3C) based technologies, such as Hi-C and TCC, have been developed to provide a genome-wide, three-dimensional (3D) view of chromatin organization. Appropriate methods for analyzing these data and fully characterizing the 3D chromosomal structure and its structural variations are still under development. Here we describe a novel Bayesian probabilistic approach, denoted as “Bayesian 3D constructor for Hi-C data” (BACH), to infer the consensus 3D chromosomal structure. In addition, we describe a variant algorithm BACH-MIX to study the structural variations of chromatin in a cell population. Applying BACH and BACH-MIX to a high resolution Hi-C dataset generated from mouse embryonic stem cells, we found that most local genomic regions exhibit homogeneous 3D chromosomal structures. We further constructed a model for the spatial arrangement of chromatin, which reveals structural properties associated with euchromatic and heterochromatic regions in the genome. We observed strong associations between structural properties and several genomic and epigenetic features of the chromosome. Using BACH-MIX, we further found that the structural variations of chromatin are correlated with these genomic and epigenetic features. Our results demonstrate that BACH and BACH-MIX have the potential to provide new insights into the chromosomal architecture of mammalian cells.

  • Publication

    On Delay Tomography: Fast Algorithms and Spatially Dependent Models

    (Institute of Electrical and Electronics Engineers (IEEE), 2012) Deng, Ke; Li, Yang; Zhu, Weiping; Geng, Zhi; Liu, Jun

    As an active branch of network tomography, delay tomography has received considerable attentions in recent years. However, most methods in the literature assume that the delays of different links are independent of each other, and pursuit sub-optimal estimate instead of the maximum likelihood estimate (MLE) due to computational challenges. In this paper, we propose a novel method to implement the EM algorithm widely used in delay tomography analysis for multicast networks. The proposed method makes use of a “delay pattern database” to avoid all redundant computations in the E-step, and is much faster than the traditional implementation. With the help of this new implementation, finding MLE for large networks, which was considered impractical previously, becomes an easy task. Taking advantage of this computational breakthrough, we further consider models for potential spatial dependence of links, and propose a novel adaptive spatially dependent model (ASDM) for delay tomography. In ASDM, Markov dependence among nearby links is allowed, and spatially dependent links (SDLs) can be automatically recognized via model selection. The superiority of the new methods is confirmed by simulation studies.

  • Publication

    Assumptions behind Intercoder Reliability Indices

    (Routledge, 2012) Zhao, Xinshu; Liu, Jun; Deng, Ke

    Inter-coder reliability is the most often used quantitative indicator of measurement quality in content studies. Researchers in psychology, sociology, education, medicine, marketing and other disciplines also use reliability to evaluate the quality of diagnosis, tests and other assessments. Many indices of reliability have been recommended for general use. This article analyzes 22, which are organized into 18 chance-adjusted and four non-adjusted indices. The chance-adjusted indices are further organized into three groups, including nine category-based indices, eight distribution-based indices, and one that is double based, on category and distribution. The main purpose of this work is to examine the assumptions behind each index. Most of the assumptions are unexamined in the literature, and yet these assumptions have implications for assessments of reliability that need to be understood, and that result in paradoxes and abnormalities. This article discusses 13 paradoxes and nine abnormalities to illustrate the 24 assumptions. To facilitate understanding, the analysis focuses on categorical scales with two coders, and further focuses on binary scales where appropriate. The discussion is situated mostly in analysis of communication content. The assumptions and patterns that we will discover will also apply to studies, evaluations and diagnoses in other disciplines with more coders, raters, diagnosticians, or judges using binary or multi-category scales. We will argue that a new index is needed. Before the new index can be established, we need guidelines for using the existing indices. This article will recommend such guidelines.

  • Publication

    Understanding spatial organizations of chromosomes via statistical analysis of Hi-C data

    (Springer Science + Business Media, 2013) Hu, Ming; Deng, Ke; Qin, Zhaohui; Liu, Jun

    Understanding how chromosomes fold provides insights into the transcription regulation, hence, the functional state of the cell. Using the next generation sequencing technology, the recently developed Hi-C approach enables a global view of spatial chromatin organization in the nucleus, which substantially expands our knowledge about genome organization and function. However, due to multiple layers of biases, noises and uncertainties buried in the protocol of Hi-C experiments, analyzing and interpreting Hi-C data poses great challenges, and requires novel statistical methods to be developed. This article provides an overview of recent Hi-C studies and their impacts on biomedical research, describes major challenges in statistical analysis of Hi-C data, and discusses some perspectives for future research.

  • Publication

    High-dimensional genomic data bias correction and data integration using MANCIE

    (Nature Publishing Group, 2016) Zang, Chongzhi; Wang, Tao; Deng, Ke; Li, Bo; Hu, Sheng'en; Qin, Qian; Xiao, Tengfei; Zhang, Shihua; Meyer, Clifford; He, Housheng Hansen; Brown, Myles; Liu, Jun; Xie, Yang; Liu, X. Shirley

    High-dimensional genomic data analysis is challenging due to noises and biases in high-throughput experiments. We present a computational method matrix analysis and normalization by concordant information enhancement (MANCIE) for bias correction and data integration of distinct genomic profiles on the same samples. MANCIE uses a Bayesian-supported principal component analysis-based approach to adjust the data so as to achieve better consistency between sample-wise distances in the different profiles. MANCIE can improve tissue-specific clustering in ENCODE data, prognostic prediction in Molecular Taxonomy of Breast Cancer International Consortium and The Cancer Genome Atlas data, copy number and expression agreement in Cancer Cell Line Encyclopedia data, and has broad applications in cross-platform, high-dimensional data integration.

  • Publication

    Fast parameter estimation in loss tomography for networks of general topology

    (Institute of Mathematical Statistics, 2016) Deng, Ke; Li, Yang; Zhu, Weiping; Liu, Jun

    As a technique to investigate link-level loss rates of a computer network with low operational cost, loss tomography has received considerable attentions in recent years. A number of parameter estimation methods have been proposed for loss tomography of networks with a tree structure as well as a general topological structure. However, these methods suffer from either high computational cost or insufficient use of information in the data. In this paper, we provide both theoretical results and practical algorithms for parameter estimation in loss tomography. By introducing a group of novel statistics and alternative parameter systems, we find that the likelihood function of the observed data from loss tomography keeps exactly the same mathematical formulation for tree and general topologies, revealing that networks with different topologies share the same mathematical nature for loss tomography. More importantly, we discover that a reparametrization of the likelihood function belongs to the standard exponential family, which is convex and has a unique mode under regularity conditions. Based on these theoretical results, novel algorithms to find the MLE are developed. Compared to existing methods in the literature, the proposed methods enjoy great computational advantages.