Person: Xia, Zongqi
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Modeling Disease Severity in Multiple Sclerosis Using Electronic Health Records
(Public Library of Science, 2013) Xia, Zongqi; Secor, Elizabeth; Chibnik, Lori; Bove, Riley; Cheng, Suchun; Chitnis, Tanuja; Cagan, Andrew; Gainer, Vivian S.; Chen, Pei J.; Liao, Katherine; Shaw, Stanley; Ananthakrishnan, Ashwin; Szolovits, Peter; Weiner, Howard; Karlson, Elizabeth; Murphy, Shawn; Savova, Guergana; Cai, Tianxi; Churchill, Susanne E.; Plenge, Robert M.; Kohane, Isaac; De Jager, PhilipObjective: To optimally leverage the scalability and unique features of the electronic health records (EHR) for research that would ultimately improve patient care, we need to accurately identify patients and extract clinically meaningful measures. Using multiple sclerosis (MS) as a proof of principle, we showcased how to leverage routinely collected EHR data to identify patients with a complex neurological disorder and derive an important surrogate measure of disease severity heretofore only available in research settings. Methods: In a cross-sectional observational study, 5,495 MS patients were identified from the EHR systems of two major referral hospitals using an algorithm that includes codified and narrative information extracted using natural language processing. In the subset of patients who receive neurological care at a MS Center where disease measures have been collected, we used routinely collected EHR data to extract two aggregate indicators of MS severity of clinical relevance multiple sclerosis severity score (MSSS) and brain parenchymal fraction (BPF, a measure of whole brain volume). Results: The EHR algorithm that identifies MS patients has an area under the curve of 0.958, 83% sensitivity, 92% positive predictive value, and 89% negative predictive value when a 95% specificity threshold is used. The correlation between EHR-derived and true MSSS has a mean R2 = 0.38±0.05, and that between EHR-derived and true BPF has a mean R2 = 0.22±0.08. To illustrate its clinical relevance, derived MSSS captures the expected difference in disease severity between relapsing-remitting and progressive MS patients after adjusting for sex, age of symptom onset and disease duration (p = 1.56×10−12). Conclusion: Incorporation of sophisticated codified and narrative EHR data accurately identifies MS patients and provides estimation of a well-accepted indicator of MS severity that is widely used in research settings but not part of the routine medical records. Similar approaches could be applied to other complex neurological disorders.
Publication Complex relation of HLA-DRB1*1501, age at menarche, and age at multiple sclerosis onset
(Wolters Kluwer, 2016) Bove, Riley; Chua, Alicia S.; Xia, Zongqi; Chibnik, Lori; De Jager, Philip; Chitnis, TanujaObjective: To examine the relationship between 2 markers of early multiple sclerosis (MS) onset, 1 genetic (HLA-DRB11501) and 1 experiential (early menarche), in 2 cohorts. Methods: We included 540 white women with MS or clinically isolated syndrome (N = 156 with genetic data available) and 1,390 white women without MS but with a first-degree relative with MS (Genes and Environment in Multiple Sclerosis [GEMS]). Age at menarche, HLA-DRB11501 status, and age at MS onset were analyzed. Results: In both cohorts, participants with at least 1 HLA-DRB11501 allele had a later age at menarche than did participants with no risk alleles (MS: mean difference = 0.49, 95% confidence interval [CI] = [0.03–0.95], p = 0.036; GEMS: mean difference = 0.159, 95% CI = [0.012–0.305], p = 0.034). This association remained after we adjusted for body mass index at age 18 (available in GEMS) and for other MS risk alleles, as well as a single nucleotide polymorphism near the HLA-A region previously associated with age of menarche (available in MS cohort). Confirming previously reported associations, in our MS cohort, every year decrease in age at menarche was associated with a 0.65-year earlier MS onset (95% CI = [0.07–1.22], p = 0.027, N = 540). Earlier MS onset was also found in individuals with at least 1 HLA-DRB11501 risk allele (mean difference = −3.40 years, 95% CI = [−6.42 to −0.37], p = 0.028, N = 156). Conclusions: In 2 cohorts, a genetic marker for earlier MS onset (HLA-DRB1*1501) was inversely related to earlier menarche, an experiential marker for earlier symptom onset. This finding warrants broader investigations into the association between the HLA region and hormonal regulation in determining the onset of autoimmune disease.
Publication Development of phenotype algorithms using electronic medical records and incorporating natural language processing
(BMJ Publishing Group Ltd., 2015) Liao, Katherine; Cai, Tianxi; Savova, Guergana K; Murphy, Shawn; Karlson, Elizabeth; Ananthakrishnan, Ashwin; Gainer, Vivian S; Shaw, Stanley; Xia, Zongqi; Szolovits, Peter; Churchill, Susanne; Kohane, IsaacElectronic medical records are emerging as a major source of data for clinical and translational research studies, although phenotypes of interest need to be accurately defined first. This article provides an overview of how to develop a phenotype algorithm from electronic medical records, incorporating modern informatics and biostatistics methods.
Publication Methods to Develop an Electronic Medical Record Phenotype Algorithm to Compare the Risk of Coronary Artery Disease across 3 Chronic Disease Cohorts
(Public Library of Science, 2015) Liao, Katherine; Ananthakrishnan, Ashwin; Kumar, Vishesh; Xia, Zongqi; Cagan, Andrew; Gainer, Vivian S.; Goryachev, Sergey; Chen, Pei; Savova, Guergana; Agniel, Denis; Churchill, Susanne; Lee, Jaeyoung; Murphy, Shawn; Plenge, Robert M.; Szolovits, Peter; Kohane, Isaac; Shaw, Stanley; Karlson, Elizabeth; Cai, TianxiBackground: Typically, algorithms to classify phenotypes using electronic medical record (EMR) data were developed to perform well in a specific patient population. There is increasing interest in analyses which can allow study of a specific outcome across different diseases. Such a study in the EMR would require an algorithm that can be applied across different patient populations. Our objectives were: (1) to develop an algorithm that would enable the study of coronary artery disease (CAD) across diverse patient populations; (2) to study the impact of adding narrative data extracted using natural language processing (NLP) in the algorithm. Additionally, we demonstrate how to implement CAD algorithm to compare risk across 3 chronic diseases in a preliminary study. Methods and Results: We studied 3 established EMR based patient cohorts: diabetes mellitus (DM, n = 65,099), inflammatory bowel disease (IBD, n = 10,974), and rheumatoid arthritis (RA, n = 4,453) from two large academic centers. We developed a CAD algorithm using NLP in addition to structured data (e.g. ICD9 codes) in the RA cohort and validated it in the DM and IBD cohorts. The CAD algorithm using NLP in addition to structured data achieved specificity >95% with a positive predictive value (PPV) 90% in the training (RA) and validation sets (IBD and DM). The addition of NLP data improved the sensitivity for all cohorts, classifying an additional 17% of CAD subjects in IBD and 10% in DM while maintaining PPV of 90%. The algorithm classified 16,488 DM (26.1%), 457 IBD (4.2%), and 245 RA (5.0%) with CAD. In a cross-sectional analysis, CAD risk was 63% lower in RA and 68% lower in IBD compared to DM (p<0.0001) after adjusting for traditional cardiovascular risk factors. Conclusions: We developed and validated a CAD algorithm that performed well across diverse patient populations. The addition of NLP into the CAD algorithm improved the sensitivity of the algorithm, particularly in cohorts where the prevalence of CAD was low. Preliminary data suggest that CAD risk was significantly lower in RA and IBD compared to DM.
Publication High-Throughput Phenotyping With Electronic Medical Record Data Using a Common Semi-Supervised Approach (PheCAP)
(Springer Science and Business Media LLC, 2019-11-20) Zhang, Yichi; Cai, Tianrun; Yu, Sheng; Cho, Kelly; Hong, Chuan; Sun, Jiehuan; Huang, Jie; Xia, Zongqi; Castro, Victor; Gagnon, David; Savova, Guergana; Churchill, Susanne; Gaziano, John; Kohane, Isaac; Cai, Tianxi; Ho, Yuk-Lam; Ananthakrishnan, Ashwin; Shaw, Stanley; Gainer, Vivian; Link, Nicholas; Honerlaw, Jacqueline; Huong, Sicong; Karlson, Elizabeth; Plenge, Robert; Szolovits, Peter; O'Donnell, Christopher; Murphy, Shawn; Liao, KatherinePhenotypes are the foundation for clinical and genetic studies of disease risk and outcomes. The growth of biobanks linked to electronic medical record (EMR) data has both facilitated and increased the demand for efficient, accurate, and robust approaches for phenotyping millions of patients. Challenges to phenotyping using EMR data include variation in the accuracy of codes, as well as the high level of manual input required to identify features for the algorithm and to obtain gold standard labels. To address these challenges, we developed PheCAP, a high-throughput semi-supervised phenotyping pipeline. PheCAP begins with data from the EMR, including structured data and information extracted from the narrative notes using natural language processing (NLP). The standardized steps integrate automated procedures reducing the level of manual input, and machine learning approaches for algorithm training. PheCAP itself can be executed in 1-2 days if all data are available; however, the timing is largely dependent on the chart review step which typically requires at least 2 weeks. The final products of PheCAP include a phenotype algorithm, the probability of the phenotype for all patients, and a phenotype classification (yes/no).
Publication Genes and Environment in Multiple Sclerosis Project: A Platform to Investigate Multiple Sclerosis Risk
(Wiley, 2016-02) Xia, Zongqi; White, Charles C.; Owen, Emily K.; Von Korff, Alina; Clarkson, Sarah R.; McCabe, Cristin; Cimpean, Maria; Winn, Phoebe A.; Hoesing, Ashley; Steele, Sonya U.; Cortese, Irene C. M.; Chitnis, Tanuja; Weiner, Howard; Reich, Daniel S.; Chibnik, Lori; De Jager, PhilipThe Genes and Environment in Multiple Sclerosis (GEMS) project establishes a platform to investigate the events leading to MS in at-risk individuals. It has recruited 2,632 first-degree relatives from across the USA. Using an integrated genetic and environmental risk score, we identified subjects with twice the MS risk when compared to the average family member, and we report an initial incidence rate in these subjects that is 30 times greater than that of sporadic MS. We discuss the feasibility of large-scale studies of asymptomatic at-risk subjects that leverage modern tools of subject recruitment to execute collaborative projects.
Publication Improving Case Definition of Crohnʼs Disease and Ulcerative Colitis in Electronic Medical Records Using Natural Language Processing
(Oxford University Press (OUP), 2013-06) Ananthakrishnan, Ashwin; Cai, Tianxi; Savova, Guergana; Cheng, Su-Chun; Chen, Pei; Guzman, Raul; Gainer, Vivian S.; Murphy, Shawn; Szolovits, Peter; Xia, Zongqi; Shaw, Stanley; Churchill, Susanne; Karlson, Elizabeth; Kohane, Isaac; Plenge, Robert M.; Liao, KatherineIntroduction Prior studies identifying patients with inflammatory bowel disease (IBD) utilizing administrative codes have yielded inconsistent results. Our objective was to develop a robust electronic medical record (EMR) based model for classification of IBD leveraging the combination of codified data and information from clinical text notes using natural language processing (NLP).
Methods Using the EMR of 2 large academic centers, we created data marts for Crohn’s disease (CD) and ulcerative colitis (UC) comprising patients with ≥ 1 ICD-9 code for each disease. We utilized codified (i.e. ICD9 codes, electronic prescriptions) and narrative data from clinical notes to develop our classification model. Model development and validation was performed in a training set of 600 randomly selected patients for each disease with medical record review as the gold standard. Logistic regression with the adaptive LASSO penalty was used to select informative variables.
Results We confirmed 399 (67%) CD cases in the CD training set and 378 (63%) UC cases in the UC training set. For both, a combined model including narrative and codified data had better accuracy (area under the curve (AUC) for CD 0.95; UC 0.94) than models utilizing only disease ICD-9 codes (AUC 0.89 for CD; 0.86 for UC). Addition of NLP narrative terms to our final model resulted in classification of 6–12% more subjects with the same accuracy.
Conclusion Inclusion of narrative concepts identified using NLP improves the accuracy of EMR case-definition for CD and UC while simultaneously identifying more subjects compared to models using codified data alone.
Publication Normalization of Plasma 25-Hydroxy Vitamin D Is Associated with Reduced Risk of Surgery in Crohn’s Disease
(Oxford University Press (OUP), 2013-08-01) Ananthakrishnan, Ashwin; Cagan, Andrew; Gainer, Vivian S.; Cai, Tianxi; Cheng, Su-Chun; Savova, Guergana; Chen, Pei; Szolovits, Peter; Xia, Zongqi; De Jager, Philip; Shaw, Stanley; Churchill, Susanne; Karlson, Elizabeth; Kohane, Isaac; Plenge, Robert; Murphy, Shawn; Liao, KatherineIntroduction Vitamin D may have an immunological role in Crohn’s disease (CD) and ulcerative colitis (UC). Retrospective studies suggested a weak association between vitamin D status and disease activity but have significant limitations.
Methods Using a multi-institution inflammatory bowel disease (IBD) cohort, we identified all CD and UC patients who had at least one measured plasma 25-hydroxy vitamin D [25(OH)D]. Plasma 25(OH)D was considered sufficient at levels ≥ 30ng/mL. Logistic regression models adjusting for potential confounders were used to identify impact of measured plasma 25(OH)D on subsequent risk of IBD-related surgery or hospitalization. In a subset of patients where multiple measures of 25(OH)D were available, we examined impact of normalization of vitamin D status on study outcomes.
Results Our study included 3,217 patients (55% CD, mean age 49 yrs). The median lowest plasma 25(OH)D was 26ng/ml (IQR 17–35ng/ml). In CD, on multivariable analysis, plasma 25(OH)D < 20ng/ml was associated with an increased risk of surgery (OR 1.76 (1.24 – 2.51) and IBD-related hospitalization (OR 2.07, 95% CI 1.59 – 2.68) compared to those with 25(OH)D ≥ 30ng/ml. Similar estimates were also seen for UC. Furthermore, CD patients who had initial levels < 30ng/ml but subsequently normalized their 25(OH)D had a reduced likelihood of surgery (OR 0.56, 95% CI 0.32 – 0.98) compared to those who remained deficient.
Conclusion Low plasma 25(OH)D is associated with increased risk of surgery and hospitalizations in both CD and UC and normalization of 25(OH)D status is associated with a reduction in the risk of CD-related surgery.
Publication Paraneoplastic limbic encephalitis presenting as a neurological emergency: a case report
(Springer Science and Business Media LLC, 2010-03-24) Xia, Zongqi; Mehta, Brijesh P; Ropper, Allan; Kesari, SantoshIntroduction
Paraneoplastic limbic encephalitis remains a challenging clinical diagnosis with poor outcome if it is not recognized and treated early in the course of the disease.
Case Presentation
A 65-year-old Caucasian woman presented with generalized tonic-clonic seizures and increasing confusion shortly after a lung biopsy that led to the diagnosis of small-cell lung cancer. She had a complicated hospital course, and had recurrent respiratory distress due to aspiration pneumonia, and fluctuating mental status and seizures that were refractory to anti-epileptic drug treatment. Routine laboratory testing, magnetic resonance imaging of the brain, electroencephalogram, lumbar puncture, serum and cerebrospinal fluid tests for paraneoplastic antibodies, and chest computed tomography were performed on our patient. The diagnosis was paraneoplastic limbic encephalitis in the setting of small-cell lung cancer with positive N-type voltage-gated calcium channel antibody titer. Anti-epileptic drugs for seizures, chemotherapy for small-cell lung cancer, and intravenous immunoglobulin and steroids for paraneoplastic limbic encephalitis led to a resolution of her seizures and improved her mental status.
Conclusion
Early recognition of paraneoplastic limbic encephalitis and prompt intervention with immune therapies at the onset of presentation will probably translate into more favorable neurological outcomes.
Publication A Putative Alzheimer's Disease Risk Allele in PCK1 Influences Brain Atrophy in Multiple Sclerosis
(Public Library of Science (PLoS), 2010-11-30) Xia, Zongqi; Chibnik, Lori; Glanz, Bonnie; Liguori, Maria; Shulman, Joshua M.; Tran, Dong; Khoury, Samia; Chitnis, Tanuja; Holyoak, Todd; Weiner, Howard; Guttmann, Charles; De Jager, Philip L.Background Brain atrophy and cognitive dysfunction are neurodegenerative features of Multiple Sclerosis (MS). We used a candidate gene approach to address whether genetic variants implicated in susceptibility to late onset Alzheimer's Disease (AD) influence brain volume and cognition in MS patients.
Methods/Principal Findings MS subjects were genotyped for five single nucleotide polymorphisms (SNPs) associated with susceptibility to AD: PICALM, CR1, CLU, PCK1, and ZNF224. We assessed brain volume using Brain Parenchymal Fraction (BPF) measurements obtained from Magnetic Resonance Imaging (MRI) data and cognitive function using the Symbol Digit Modalities Test (SDMT). Genotypes were correlated with cross-sectional BPF and SDMT scores using linear regression after adjusting for sex, age at symptom onset, and disease duration. 722 MS patients with a mean (±SD) age at enrollment of 41 (±10) years were followed for 44 (±28) months. The AD risk-associated allele of a non-synonymous SNP in the PCK1 locus (rs8192708G) is associated with a smaller average brain volume (P = 0.0047) at the baseline MRI, but it does not impact our baseline estimate of cognition. PCK1 is additionally associated with higher baseline T2-hyperintense lesion volume (P = 0.0088). Finally, we provide technical validation of our observation in a subset of 641 subjects that have more than one MRI study, demonstrating the same association between PCK1 and smaller average brain volume (P = 0.0089) at the last MRI visit.
Conclusion/Significance Our study provides suggestive evidence for greater brain atrophy in MS patients bearing the PCK1 allele associated with AD-susceptibility, yielding new insights into potentially shared neurodegenerative process between MS and late onset AD.