Person: Cai, Tianxi
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Modeling Disease Severity in Multiple Sclerosis Using Electronic Health Records
(Public Library of Science, 2013) Xia, Zongqi; Secor, Elizabeth; Chibnik, Lori; Bove, Riley; Cheng, Suchun; Chitnis, Tanuja; Cagan, Andrew; Gainer, Vivian S.; Chen, Pei J.; Liao, Katherine; Shaw, Stanley; Ananthakrishnan, Ashwin; Szolovits, Peter; Weiner, Howard; Karlson, Elizabeth; Murphy, Shawn; Savova, Guergana; Cai, Tianxi; Churchill, Susanne E.; Plenge, Robert M.; Kohane, Isaac; De Jager, PhilipObjective: To optimally leverage the scalability and unique features of the electronic health records (EHR) for research that would ultimately improve patient care, we need to accurately identify patients and extract clinically meaningful measures. Using multiple sclerosis (MS) as a proof of principle, we showcased how to leverage routinely collected EHR data to identify patients with a complex neurological disorder and derive an important surrogate measure of disease severity heretofore only available in research settings. Methods: In a cross-sectional observational study, 5,495 MS patients were identified from the EHR systems of two major referral hospitals using an algorithm that includes codified and narrative information extracted using natural language processing. In the subset of patients who receive neurological care at a MS Center where disease measures have been collected, we used routinely collected EHR data to extract two aggregate indicators of MS severity of clinical relevance multiple sclerosis severity score (MSSS) and brain parenchymal fraction (BPF, a measure of whole brain volume). Results: The EHR algorithm that identifies MS patients has an area under the curve of 0.958, 83% sensitivity, 92% positive predictive value, and 89% negative predictive value when a 95% specificity threshold is used. The correlation between EHR-derived and true MSSS has a mean R2 = 0.38±0.05, and that between EHR-derived and true BPF has a mean R2 = 0.22±0.08. To illustrate its clinical relevance, derived MSSS captures the expected difference in disease severity between relapsing-remitting and progressive MS patients after adjusting for sex, age of symptom onset and disease duration (p = 1.56×10−12). Conclusion: Incorporation of sophisticated codified and narrative EHR data accurately identifies MS patients and provides estimation of a well-accepted indicator of MS severity that is widely used in research settings but not part of the routine medical records. Similar approaches could be applied to other complex neurological disorders.
Publication Testing a machine-learning algorithm to predict the persistence and severity of major depressive disorder from baseline self-reports
(2015) Kessler, Ronald; van Loo, Hanna M.; Wardenaar, Klaas J.; Bossarte, Robert M.; Brenner, Lisa A.; Cai, Tianxi; Ebert, David Daniel; Hwang, Irving; Li, Junlong; de Jonge, Peter; Nierenberg, Andrew; Petukhova, Maria; Rosellini, Anthony; Sampson, Nancy; Schoevers, Robert A.; Wilcox, Marsha A.; Zaslavsky, AlanHeterogeneity of major depressive disorder (MDD) illness course complicates clinical decision-making. While efforts to use symptom profiles or biomarkers to develop clinically useful prognostic subtypes have had limited success, a recent report showed that machine learning (ML) models developed from self-reports about incident episode characteristics and comorbidities among respondents with lifetime MDD in the World Health Organization World Mental Health (WMH) Surveys predicted MDD persistence, chronicity, and severity with good accuracy. We report results of model validation in an independent prospective national household sample of 1,056 respondents with lifetime MDD at baseline. The WMH ML models were applied to these baseline data to generate predicted outcome scores that were compared to observed scores assessed 10–12 years after baseline. ML model prediction accuracy was also compared to that of conventional logistic regression models. Area under the receiver operating characteristic curve (AUC) based on ML (.63 for high chronicity and .71–.76 for the other prospective outcomes) was consistently higher than for the logistic models (.62–.70) despite the latter models including more predictors. 34.6–38.1% of respondents with subsequent high persistence-chronicity and 40.8–55.8% with the severity indicators were in the top 20% of the baseline ML predicted risk distribution, while only 0.9% of respondents with subsequent hospitalizations and 1.5% with suicide attempts were in the lowest 20% of the ML predicted risk distribution. These results confirm that clinically useful MDD risk stratification models can be generated from baseline patient self-reports and that ML methods improve on conventional methods in developing such models.
Publication Pathprinting: An integrative approach to understand the functional basis of disease
(Springer Science + Business Media, 2013) Altschuler, Gabriel M; Hofmann, Oliver; Kalatskaya, Irina; Payne, Rebecca; Ho Sui, Shannan; Saxena, Uma; Krivtsov, Andrei V; Armstrong, Scott A; Cai, Tianxi; Stein, Lincoln; Hide, WinstonNew strategies to combat complex human disease require systems approaches to biology that integrate experiments from cell lines, primary tissues and model organisms. We have developed Pathprint, a functional approach that compares gene expression profiles in a set of pathways, networks and transcriptionally regulated targets. It can be applied universally to gene expression profiles across species. Integration of large-scale profiling methods and curation of the public repository overcomes platform, species and batch effects to yield a standard measure of functional distance between experiments. We show that pathprints combine mouse and human blood developmental lineage, and can be used to identify new prognostic indicators in acute myeloid leukemia. The code and resources are available at http://compbio.sph.harvard.edu/hidelab/pathprint
Publication Iodinated Contrast Opacification Gradients in Normal Coronary Arteries Imaged With Prospectively ECG-Gated Single Heart Beat 320-Detector Row Computed Tomography
(Ovid Technologies (Wolters Kluwer Health), 2009) Steigner, Michael; Mitsouras, Dimitrios; Whitmore, A. G.; Otero, H. J.; Wang, C.; Buckley, O.; Levit, N. A.; Hussain, A. Z.; Cai, Tianxi; Mather, R. T.; Smedby, O.; DiCarli, M. F.; Rybicki, Frank JohnBackground
To define and evaluate coronary contrast opacification gradients using prospectively ECG-gated single heart beat 320-detector row coronary angiography (CTA).
Methods and Results
Thirty-six patients with normal coronary arteries determined by 320 × 0.5 mm detector row coronary CTA were retrospectively evaluated with customized image post-processing software to measure Hounsfield Units (HU) at 1 mm intervals orthogonal to the artery center line. Linear regression determined correlation between mean HU and distance from the coronary ostium (regression slope defined as the distance gradient Gd), lumen cross-sectional area (Ga), and lumen short axis diameter (Gs). For each gradient, differences between the three coronary arteries were analyzed with ANOVA. Linear regression determined correlations between measured gradients, heart rate, body-mass index (BMI), and cardiac phase. To determine feasibility in lesions, all three gradients were evaluated in 22 consecutive patients with left anterior descending artery lesions greater than or equal to 50% stenosis. For all 3 coronary arteries in all patients, the gradients Ga and Gs were significantly different from zero (p<0.0001), highly linear (Pearson r values 0.77-0.84), and had no significant difference between the LAD, LCx, and RCA (p>0.503). The distance gradient Gd demonstrated nonlinearities in a small number of vessels and was significantly smaller in the RCA when compared to the left coronary system (p<0.001). Gradient variations between cardiac phases, heart rates, BMI, and readers were low. Gradients in patients with lesions were significantly different (p<0.021) than in patients considered normal by CTA.
Conclusions
Measurement of contrast opacification gradients from temporally uniform coronary CTA demonstrates feasibility and reproducibility in patients with normal coronary arteries. For all patients the gradients defined with respect to the coronary lumen cross-sectional area and short axis diameters are highly linear, not significantly influenced by the coronary artery (LAD vs LCx vs RCA), and have only small variation with respect to patient parameters. Preliminary evaluation of gradients across coronary artery lesions is promising but requires additional study.
Publication A Predictive Phosphorylation Signature of Lung Cancer
(Public Library of Science, 2009) Rikova, Klarisa; Merberg, David; Kasif, Simon; Steffen, Martin; Creighton, Chad; Wu, Chang-Jiun; Cai, TianxiBackground: Aberrant activation of signaling pathways drives many of the fundamental biological processes that accompany tumor initiation and progression. Inappropriate phosphorylation of intermediates in these signaling pathways are a frequently observed molecular lesion that accompanies the undesirable activation or repression of pro- and anti-oncogenic pathways. Therefore, methods which directly query signaling pathway activation via phosphorylation assays in individual cancer biopsies are expected to provide important insights into the molecular “logic” that distinguishes cancer and normal tissue on one hand, and enables personalized intervention strategies on the other. Results: We first document the largest available set of tyrosine phosphorylation sites that are, individually, differentially phosphorylated in lung cancer, thus providing an immediate set of drug targets. Next, we develop a novel computational methodology to identify pathways whose phosphorylation activity is strongly correlated with the lung cancer phenotype. Finally, we demonstrate the feasibility of classifying lung cancers based on multi-variate phosphorylation signatures. Conclusions: Highly predictive and biologically transparent phosphorylation signatures of lung cancer provide evidence for the existence of a robust set of phosphorylation mechanisms (captured by the signatures) present in the majority of lung cancers, and that reliably distinguish each lung cancer from normal. This approach should improve our understanding of cancer and help guide its treatment, since the phosphorylation signatures highlight proteins and pathways whose phosphorylation should be inhibited in order to prevent unregulated proliferation.
Publication Origins of lymphatic and distant metastases in human colorectal cancer
(American Association for the Advancement of Science (AAAS), 2017) Nahrendorf, Kamila; Reiter, Johannes; Brachtel, Elena; Lennerz, Jochen; van de Wetering, Marc; Rowan, Andrew; Cai, Tianxi; Clevers, Hans; Swanton, Charles; Nowak, Martin; Elledge, Stephen; Jain, RakeshThe spread of cancer cells from primary tumors to regional lymph nodes is often associated with reduced survival. One prevailing model to explain this association posits that fatal, distant metastases are seeded by lymph node metastases. This view provides a mechanistic basis for the TNM staging system and is the rationale for surgical resection of tumor-draining lymph nodes. Here we examine the evolutionary relationship between primary tumor, lymph node, and distant metastases in human colorectal cancer. Studying 213 archival biopsy samples from 17 patients, we used somatic variants in hypermutable DNA regions to reconstruct high-confidence phylogenetic trees. We found that in 65% of cases, lymphatic and distant metastases arose from independent subclones in the primary tumor, whereas in 35% of cases they shared common subclonal origin. Therefore, two different lineage relationships between lymphatic and distant metastases exist in colorectal cancer.
Publication Predicting Suicides After Psychiatric Hospitalization in US Army Soldiers
(American Medical Association (AMA), 2015) Kessler, Ronald; Warner, Christopher H.; Ivany, Christopher; Petukhova, Maria; Rose, Sherri; Bromet, Evelyn J.; Brown, Millard; Cai, Tianxi; Colpe, Lisa J.; Cox, Kenneth L.; Fullerton, Carol S.; Gilman, Stephen Edward; Gruber, M; Heeringa, Steven G.; Lewandowski-Romps, Lisa; Li, Junlong; Millikan-Bell, Amy M.; Naifeh, James A.; Nock, Matthew K.; Rosellini, Anthony; Sampson, Nancy; Schoenbaum, Michael; Stein, Murray B.; Wessely, Simon; Zaslavsky, Alan; Ursano, Robert J.IMPORTANCE: The US Army experienced a sharp increase in soldier suicides beginning in 2004. Administrative data reveal that among those at highest risk are soldiers in the 12 months after inpatient treatment of a psychiatric disorder. OBJECTIVE: To develop an actuarial risk algorithm predicting suicide in the 12 months after US Army soldier inpatient treatment of a psychiatric disorder to target expanded posthospitalization care. DESIGN, SETTING, AND PARTICIPANTS: There were 53,769 hospitalizations of active duty soldiers from January 1, 2004, through December 31, 2009, with International Classification of Diseases, Ninth Revision, Clinical Modification psychiatric admission diagnoses. Administrative data available before hospital discharge abstracted from a wide range of data systems (sociodemographic, US Army career, criminal justice, and medical or pharmacy) were used to predict suicides in the subsequent 12 months using machine learning methods (regression trees and penalized regressions) designed to evaluate cross-validated linear, nonlinear, and interactive predictive associations. MAIN OUTCOMES AND MEASURES: Suicides of soldiers hospitalized with psychiatric disorders in the 12 months after hospital discharge. RESULTS: Sixty-eight soldiers died by suicide within 12 months of hospital discharge (12.0% of all US Army suicides), equivalent to 263.9 suicides per 100,000 person-years compared with 18.5 suicides per 100,000 person-years in the total US Army. The strongest predictors included sociodemographics (male sex [odds ratio (OR), 7.9; 95% CI, 1.9-32.6] and late age of enlistment [OR, 1.9; 95% CI, 1.0-3.5]), criminal offenses (verbal violence [OR, 2.2; 95% CI, 1.2-4.0] and weapons possession [OR, 5.6; 95% CI, 1.7-18.3]), prior suicidality [OR, 2.9; 95% CI, 1.7-4.9], aspects of prior psychiatric inpatient and outpatient treatment (eg, number of antidepressant prescriptions filled in the past 12 months [OR, 1.3; 95% CI, 1.1-1.7]), and disorders diagnosed during the focal hospitalizations (eg, nonaffective psychosis [OR, 2.9; 95% CI, 1.2-7.0]). A total of 52.9% of posthospitalization suicides occurred after the 5% of hospitalizations with highest predicted suicide risk (3824.1 suicides per 100,000 person-years). These highest-risk hospitalizations also accounted for significantly elevated proportions of several other adverse posthospitalization outcomes (unintentional injury deaths, suicide attempts, and subsequent hospitalizations). CONCLUSIONS AND RELEVANCE: The high concentration of risk of suicide and other adverse outcomes might justify targeting expanded posthospitalization interventions to soldiers classified as having highest posthospitalization suicide risk, although final determination requires careful consideration of intervention costs, comparative effectiveness, and possible adverse effects.
Publication Identification of subjects with polycystic ovary syndrome using electronic health records
(BioMed Central, 2015) Castro, Victor; Shen, Yuanyuan; Yu, Sheng; Finan, Sean; Pau, Cindy Ta; Gainer, Vivian; Keefe, Candace C.; Savova, Guergana; Murphy, Shawn; Cai, Tianxi; Welt, Corrine K.Background: Polycystic ovary syndrome (PCOS) is a heterogeneous disorder because of the variable criteria used for diagnosis. Therefore, International Classification of Diseases 9 (ICD-9) codes may not accurately capture the diagnostic criteria necessary for large scale PCOS identification. We hypothesized that use of electronic medical records text and data would more specifically capture PCOS subjects. Methods: Subjects with PCOS were identified in the Partners Healthcare Research Patients Data Registry by searching for the term “polycystic ovary syndrome” using natural language processing (n = 24,930). A training subset of 199 identified charts was reviewed and categorized based on likelihood of a true Rotterdam PCOS diagnosis, i.e. two out of three of the following: irregular menstrual cycles, hyperandrogenism and/or polycystic ovary morphology. Data from the history, physical exam, laboratory and radiology results were codified and extracted from notes of definite PCOS subjects. Thirty-two terms were used to build an algorithm for identifying definite PCOS cases and applied to the rest of the dataset. The positive predictive value cutoff was set at 76.8 % to maximize the number of subjects available for study. A true positive predictive value for the algorithm was calculated after review of 100 charts from subjects identified as definite PCOS cases with at least two documented Rotterdam criteria. The positive predictive value was compared to that calculated using 200 charts identified using the ICD-9 code for PCOS (256.4; n = 13,670). In addition, a cohort of previously recruited PCOS subjects was submitted for algorithm validation. Results: Chart review demonstrated that 64 % were confirmed as definitely PCOS using the algorithm, with a 9 % false positive rate. 66 % of subjects identified by ICD-9 code for PCOS could be confirmed as definitely PCOS, with an 8.5 % false positive rate. There was no significant difference in the positive predictive values using the two methods (p = 0.2). However, the number of charts that had insufficient confirmatory data was lower using the algorithm (5 % vs 11 %; p < 0.04). Of 477 subjects with PCOS recruited and examined individually and present in the database as patients, 451 were found within the algorithm dataset. Conclusions: Extraction of text parameters along with codified data improves the confidence in PCOS patient cohorts identified using the electronic medical record. However, the positive predictive value was not significantly different when using ICD-9 codes or the specific algorithm. Further studies are needed to determine the positive predictive value of the two methods in additional electronic medical record datasets. Electronic supplementary material The online version of this article (doi:10.1186/s12958-015-0115-z) contains supplementary material, which is available to authorized users.
Publication Genome-wide Association Studies of Posttraumatic Stress Disorder in 2 Cohorts of US Army Soldiers
(American Medical Association (AMA), 2016) Stein, Murray B.; Chen, Chia-Yen; Ursano, Robert J.; Cai, Tianxi; Gelernter, Joel; Heeringa, Steven G.; Jain, Sonia; Jensen, Kevin P.; Maihofer, Adam X.; Mitchell, Colter; Nievergelt, Caroline M.; Nock, Matthew; Neale, Benjamin; Polimanti, Renato; Ripke, Stephan; Sun, Xiaoying; Thomas, Michael P.; Wang, Qian; Ware, Erin B.; Borja, Susan; Kessler, Ronald; Smoller, Jordan; undefined, undefinedImportance Posttraumatic stress disorder (PTSD) is a prevalent, serious public health concern, particularly in the military. The identification of genetic risk factors for PTSD may provide important insights into the biological foundation of vulnerability and comorbidity.
Objective To discover genetic loci associated with the lifetime risk for PTSD in 2 cohorts from the Army Study to Assess Risk and Resilience in Servicemembers (Army STARRS).
Design, Setting, and Participants Two coordinated genome-wide association studies of mental health in the US military contributed participants. The New Soldier Study (NSS) included 3167 unique participants with PTSD and 4607 trauma-exposed control individuals; the Pre/Post Deployment Study (PPDS) included 947 unique participants with PTSD and 4969 trauma-exposed controls. The NSS data were collected from February 1, 2011, to November 30, 2012; the PDDS data, from January 9 to April 30, 2012. The primary analysis compared lifetime DSM-IV PTSD cases with trauma-exposed controls without lifetime PTSD. Data were analyzed from March 18 to December 27, 2015.
Main Outcomes and Measures Association analyses for PTSD used logistic regression models within each of 3 ancestral groups (European, African, and Latino American) by study, followed by meta-analysis. Heritability and genetic correlation and pleiotropy with other psychiatric and immune-related disorders were estimated.
Results The NSS population was 80.7% male (6277 of 7774 participants; mean [SD] age, 20.9 [3.3] years); the PPDS population, 94.4% male (5583 of 5916 participants; mean [SD] age, 26.5 [6.0] years). A genome-wide significant locus was found in ANKRD55 on chromosome 5 (rs159572; odds ratio [OR], 1.62; 95% CI, 1.37-1.92; P = 2.34 × 10−8) and persisted after adjustment for cumulative trauma exposure (adjusted OR, 1.64; 95% CI, 1.39-1.95; P = 1.18 × 10−8) in the African American samples from the NSS. A genome-wide significant locus was also found in or near ZNF626 on chromosome 19 (rs11085374; OR, 0.77; 95% CI, 0.70-0.85; P = 4.59 × 10−8) in the European American samples from the NSS. Similar results were not found for either single-nucleotide polymorphism in the corresponding ancestry group from the PPDS sample, in other ancestral groups, or in transancestral meta-analyses. Single-nucleotide polymorphism–based heritability was nonsignificant, and no significant genetic correlations were observed between PTSD and 6 mental disorders or 9 immune-related disorders. Significant evidence of pleiotropy was observed between PTSD and rheumatoid arthritis and, to a lesser extent, psoriasis.
Conclusions and Relevance In the largest genome-wide association study of PTSD to date, involving a US military sample, limited evidence of association for specific loci was found. Further efforts are needed to replicate the genome-wide significant association with ANKRD55—associated in prior research with several autoimmune and inflammatory disorders—and to clarify the nature of the genetic overlap observed between PTSD and rheumatoid arthritis and psoriasis.
Publication Development of phenotype algorithms using electronic medical records and incorporating natural language processing
(BMJ Publishing Group Ltd., 2015) Liao, Katherine; Cai, Tianxi; Savova, Guergana K; Murphy, Shawn; Karlson, Elizabeth; Ananthakrishnan, Ashwin; Gainer, Vivian S; Shaw, Stanley; Xia, Zongqi; Szolovits, Peter; Churchill, Susanne; Kohane, IsaacElectronic medical records are emerging as a major source of data for clinical and translational research studies, although phenotypes of interest need to be accurately defined first. This article provides an overview of how to develop a phenotype algorithm from electronic medical records, incorporating modern informatics and biostatistics methods.
- «
- 1 (current)
- 2
- 3
- »