Person: Meng, Xiao-li
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Decoding the H-likelihood
(Institute of Mathematical Statistics, 2009) Meng, Xiao-liPublication Desired and feared — What do we do now and over the next 50 years?
(Informa UK Limited, 2009) Meng, Xiao-liAn intense debate about Harvard University’s General Education Curriculum demonstrates that statistics, as a discipline, is now both desired and feared. With this new status comes a set of enormous challenges. We no longer simply enjoy the privilege of playing in or cleaning up everyone’s backyard. We are now being invited into everyone’s study or living room, and trusted with the task of being their offspring’s first quantitative nanny. Are we up to such a nerve-wracking task, given the insignificant size of our profession relative to the sheer number of our hosts and their progeny? Echoing Brown and Kass’s “What Is Statistics?” (2009), this article further suggests ways to prepare our profession to meet the ever-increasing demand, in terms of both quantity and quality. Discussed are (1) the need to supplement our graduate curricula with a professional development curriculum (PDC); (2) the need to develop more subject oriented statistics (SOS) courses and happy courses at the undergraduate level; (3) the need to have the most qualified statisticians—in terms of both teaching and research credentials—to teach introductory statistical courses, especially those for other disciplines; (4) the need to deepen our foundation while expanding our horizon in both teaching and research; and (5) the need to greatly increase the general awareness and avoidance of unprincipled data analysis methods, through our practice and teaching, as a way to combat “incentive bias,” a main culprit of false discoveries in science, misleading information in media, and misguided policies in society.
Publication Quantifying the Fraction of Missing Information for Hypothesis Testing in Statistical and Genetic Studies
(Institute of Mathematical Statistics, 2008) Nicolae, Dan L.; Meng, Xiao-li; Kong, AugustineMany practical studies rely on hypothesis testing procedures applied to data sets with missing information. An important part of the analysis is to determine the impact of the missing data on the performance of the test, and this can be done by properly quantifying the relative (to complete data) amount of available information. The problem is directly motivated by applications to studies, such as linkage analyses and haplotype-based association projects, designed to identify genetic contributions to complex diseases. In the genetic studies the relative information measures are needed for the experimental design, technology comparison, interpretation of the data, and for understanding the behavior of some of the inference tools. The central difficulties in constructing such information measures arise from the multiple, and sometimes conflicting, aims in practice. For large samples, we show that a satisfactory, likelihood-based general solution exists by using appropriate forms of the relative Kullback–Leibler information, and that the proposed measures are computationally inexpensive given the maximized likelihoods with the observed data. Two measures are introduced, under the null and alternative hypothesis respectively. We exemplify the measures on data coming from mapping studies on the inflammatory bowel disease and diabetes. For small-sample problems, which appear rather frequently in practice and sometimes in disguised forms (e.g., measuring individual contributions to a large study), the robust Bayesian approach holds great promise, though the choice of a general-purpose “default prior” is a very challenging problem. We also report several intriguing connections encountered in our investigation, such as the connection with the fundamental identity for the EM algorithm, the connection with the second CR (Chapman–Robbins) lower information bound, the connection with entropy, and connections between likelihood ratios and Bayes factors. We hope that these seemingly unrelated connections, as well as our specific proposals, will stimulate a general discussion and research in this theoretically fascinating and practically needed area.
Publication Disparities in Defining Disparities: Statistical Conceptual Frameworks
(Wiley-Blackwell, 2008) Duan, Naihua; Meng, Xiao-li; Lin, Julia Y.; Chen, Chih-nan; Alegria, MargaritaMotivated by the need to meaningfully implement the Institute of Medicine's (IOM's) definition of health care disparity, this paper proposes statistical frameworks that lay out explicitly the needed causal assumptions for defining disparity measures. Our key emphasis is that a scientifically defensible disparity measure must take into account the direction of the causal relationship between allowable covariates that are not considered to be contributors to disparity and non-allowable covariates that are considered to be contributors to disparity, to avoid flawed disparity measures based on implausible populations that are not relevant for clinical or policy decisions. However, these causal relationships are usually unknown and undetectable from observed data. Consequently, we must make strong causal assumptions in order to proceed. Two frameworks are proposed in this paper, one is the conditional disparity framework under the assumption that allowable covariates impact non-allowable covariates but not vice versa. The other is the marginal disparity framework under the assumption that non-allowable covariates impact allowable ones but not vice versa. We establish theoretical conditions under which the two disparity measures are the same and present a theoretical example showing that the difference between the two disparity measures can be arbitrarily large. Using data from the Collaborative Psychiatric Epidemiology Survey, we also provide an example where the conditional disparity is misled by Simpson's paradox, whereas the marginal disparity approach handles it correctly.
Publication Inference, Statistical
(Macmillan Reference USA, 2008) Meng, Xiao-liPublication Discussion: One-step Sparse Estimates in Nonconcave Penalized Likelihood Models: Who Cares if It Is a White cat or a Black cat?
(Institute of Mathematical Statistics, 2008) Meng, Xiao-liPublication Rejoinder: Quantifying the Fraction of Missing Information for Hypothesis Testing in Statistical and Genetic Studies
(Institute of Mathematical Statistics, 2008) Nicolae, Dan L.; Meng, Xiao-li; Kong, AugustineFew authors would not be pleased when discussants implement their methods or follow-up on their ideas. It is therefore a professional joy to see every discussant doing both! Our heartfelt thanks go to all discussants, and to the Executive Editor, Ed George, for bringing us such joy! Incidentally, the three discussions cover nicely the three main parts of our paper. Zheng and Lo’s discussion centers on our motivating application, namely, designing follow-up strategies in genetic studies, but with the additional consideration of the uncertainty in the measures themselves. Doss’s discussion focuses on the second part of our paper, namely, the likelihood-based relative measure, but with applications to survival analysis where the use of partial likelihood reveals very interesting (and inevitably confusing) complications. Chang, Chen, Chien and Hsing (hereafter C3H) comment on the third part of our paper, the Bayesian mea- sures for small samples, and implement variations that are applied to problems in infectious disease research and isotonic regression. Our responses are organized in the aforementioned order. We very much appreciate all the key messages conveyed by the discussants, though for a few of them we offer alternative explanations. Some questions posed by the discussants make nice Ph.D. or master thesis topics, so we summarize them at the end of this rejoinder.