Person: Tonellato, Peter
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication The Translational Medicine Ontology and Knowledge Base: Driving Personalized Medicine By Bridging The Gap Between Bench And Bedside
(BioMed Central, 2011) Luciano, Joanne S; Andersson, Bosse; Batchelor, Colin; Bodenreider, Olivier; Denney, Christine K; Domarew, Christopher; Gambet, Thomas; Harland, Lee; Jentzsch, Anja; Kashyap, Vipul; Kos, Peter; Kozlovsky, Julia; Lebo, Timothy; Marshall, Scott M; McCusker, James P; McGuinness, Deborah L; Ogbuji, Chimezie; Pichler, Elgar; Samwald, Matthias; Schriml, Lynn; Whetzel, Patricia L; Stephens, Susie; Dumontier, Michel; Clark, Tim; Prud'hommeaux, Eric; Tonellato, Peter; Zhao, JunBackground: Translational medicine requires the integration of knowledge using heterogeneous data from health care to the life sciences. Here, we describe a collaborative effort to produce a prototype Translational Medicine Knowledge Base (TMKB) capable of answering questions relating to clinical practice and pharmaceutical drug discovery. Results: We developed the Translational Medicine Ontology (TMO) as a unifying ontology to integrate chemical, genomic and proteomic data with disease, treatment, and electronic health records. We demonstrate the use of Semantic Web technologies in the integration of patient and biomedical data, and reveal how such a knowledge base can aid physicians in providing tailored patient care and facilitate the recruitment of patients into active clinical trials. Thus, patients, physicians and researchers may explore the knowledge base to better understand therapeutic options, efficacy, and mechanisms of action. Conclusions: This work takes an important step in using Semantic Web technologies to facilitate integration of relevant, distributed, external sources and progress towards a computational platform to support personalized medicine. Availability: TMO can be downloaded from http://code.google.com/p/translationalmedicineontology and TMKB can be accessed at http://tm.semanticscience.org/sparql.
Publication Cloud Computing for Comparative Genomics
(BioMed Central, 2010) Wall, Dennis Paul; Kudtarkar, Parul; Fusaro, Vincent Alfred; Pivovarov, Rimma; Patil, Prasad; Tonellato, PeterBackground: Large comparative genomics studies and tools are becoming increasingly more compute-expensive as the number of available genome sequences continues to rise. The capacity and cost of local computing infrastructures are likely to become prohibitive with the increase, especially as the breadth of questions continues to rise. Alternative computing architectures, in particular cloud computing environments, may help alleviate this increasing pressure and enable fast, large-scale, and cost-effective comparative genomics strategies going forward. To test this, we redesigned a typical comparative genomics algorithm, the reciprocal smallest distance algorithm (RSD), to run within Amazon's Elastic Computing Cloud (EC2). We then employed the RSD-cloud for ortholog calculations across a wide selection of fully sequenced genomes. Results: We ran more than 300,000 RSD-cloud processes within the EC2. These jobs were farmed simultaneously to 100 high capacity compute nodes using the Amazon Web Service Elastic Map Reduce and included a wide mix of large and small genomes. The total computation time took just under 70 hours and cost a total of $6,302 USD. Conclusions: The effort to transform existing comparative genomics algorithms from local compute infrastructures is not trivial. However, the speed and flexibility of cloud computing environments provides a substantial boost with manageable cost. The procedure designed to transform the RSD algorithm into a cloud-ready application is readily adaptable to similar comparative genomics problems.
Publication The future of genomics in pathology
(Faculty of 1000 Ltd, 2012) Wall, Dennis Paul; Tonellato, PeterThe recent advances in technology and the promise of cheap and fast whole genomic data offer the possibility to revolutionise the discipline of pathology. This should allow pathologists in the near future to diagnose disease rapidly and early to change its course, and to tailor treatment programs to the individual. This review outlines some of these technical advances and the changes needed to make this revolution a reality.
Publication A Simulation Platform to Examine Heterogeneity Influence on Treatment
(American Medical Informatics Association, 2012) Chi, Chih-Lin; Fusaro, Vincent Alfred; Patil, Prasad; Crawford, Matthew A.; Content, Charles F.; Tonellato, PeterAlthough a protocol aims to guide treatment management and optimize overall outcomes, the benefits and harms for each individual vary due to heterogeneity. Some protocols integrate clinical and genetic variation to provide treatment recommendation; it is not clear whether such integration is sufficient. If not, treatment outcomes may be sub-optimal for certain patient sub-populations. Unfortunately, running a clinical trial to examine such outcome responses is cost prohibitive and requires a significant amount of time to conduct the study. We propose a simulation approach to discover this knowledge from electronic medical records; a rapid method to reach this goal. We use the well-known drug warfarin as an example to examine whether patient characteristics, including race and the genes CYP2C9 and VKORC1, have been fully integrated into dosing protocols. The two genes mentioned above have been shown to be important in patient response to warfarin.
Publication Personalized cloud-based bioinformatics services for research and education: Use cases and the elasticHPC package
(BioMed Central, 2012) El-Kalioby, Mohamed; Abouelhoda, Mohamed; Krüger, Jan; Giegerich, Robert; Sczyrba, Alexander; Wall, Dennis Paul; Tonellato, PeterBackground: Bioinformatics services have been traditionally provided in the form of a web-server that is hosted at institutional infrastructure and serves multiple users. This model, however, is not flexible enough to cope with the increasing number of users, increasing data size, and new requirements in terms of speed and availability of service. The advent of cloud computing suggests a new service model that provides an efficient solution to these problems, based on the concepts of "resources-on-demand" and "pay-as-you-go". However, cloud computing has not yet been introduced within bioinformatics servers due to the lack of usage scenarios and software layers that address the requirements of the bioinformatics domain. Results: In this paper, we provide different use case scenarios for providing cloud computing based services, considering both the technical and financial aspects of the cloud computing service model. These scenarios are for individual users seeking computational power as well as bioinformatics service providers aiming at provision of personalized bioinformatics services to their users. We also present elasticHPC, a software package and a library that facilitates the use of high performance cloud computing resources in general and the implementation of the suggested bioinformatics scenarios in particular. Concrete examples that demonstrate the suggested use case scenarios with whole bioinformatics servers and major sequence analysis tools like BLAST are presented. Experimental results with large datasets are also included to show the advantages of the cloud model. Conclusions: Our use case scenarios and the elasticHPC package are steps towards the provision of cloud based bioinformatics services, which would help in overcoming the data challenge of recent biological research. All resources related to elasticHPC and its web-interface are available at http://www.elasticHPC.org.
Publication Genotator: A Disease-Agnostic Tool for Genetic Annotation of Disease
(BioMed Central, 2010) Wall, Dennis Paul; Pivovarov, Rimma; Tong, Mark; Jung, Jae-Yoon; Fusaro, Vincent Alfred; DeLuca, Todd; Tonellato, PeterBackground: Disease-specific genetic information has been increasing at rapid rates as a consequence of recent improvements and massive cost reductions in sequencing technologies. Numerous systems designed to capture and organize this mounting sea of genetic data have emerged, but these resources differ dramatically in their disease coverage and genetic depth. With few exceptions, researchers must manually search a variety of sites to assemble a complete set of genetic evidence for a particular disease of interest, a process that is both time-consuming and error-prone. Methods: We designed a real-time aggregation tool that provides both comprehensive coverage and reliable gene-to-disease rankings for any disease. Our tool, called Genotator, automatically integrates data from 11 externally accessible clinical genetics resources and uses these data in a straightforward formula to rank genes in order of disease relevance. We tested the accuracy of coverage of Genotator in three separate diseases for which there exist specialty curated databases, Autism Spectrum Disorder, Parkinson's Disease, and Alzheimer Disease. Genotator is freely available at http://genotator.hms.harvard.edu. Results: Genotator demonstrated that most of the 11 selected databases contain unique information about the genetic composition of disease, with 2514 genes found in only one of the 11 databases. These findings confirm that the integration of these databases provides a more complete picture than would be possible from any one database alone. Genotator successfully identified at least 75% of the top ranked genes for all three of our use cases, including a 90% concordance with the top 40 ranked candidates for Alzheimer Disease. Conclusions: As a meta-query engine, Genotator provides high coverage of both historical genetic research as well as recent advances in the genetic understanding of specific diseases. As such, Genotator provides a real-time aggregation of ranked data that remains current with the pace of research in the disease fields. Genotator's algorithm appropriately transforms query terms to match the input requirements of each targeted databases and accurately resolves named synonyms to ensure full coverage of the genetic results with official nomenclature. Genotator generates an excel-style output that is consistent across disease queries and readily importable to other applications.
Publication Systems Analysis of Inflammatory Bowel Disease Based on Comprehensive Gene Information
(BioMed Central, 2012) Suzuki, Satoru; Takai-Igarashi, Takako; Fukuoka, Yutaka; Wall, Dennis Paul; Tanaka, Hiroshi; Tonellato, PeterBackground: The rise of systems biology and availability of highly curated gene and molecular information resources has promoted a comprehensive approach to study disease as the cumulative deleterious function of a collection of individual genes and networks of molecules acting in concert. These "human disease networks" (HDN) have revealed novel candidate genes and pharmaceutical targets for many diseases and identified fundamental HDN features conserved across diseases. A network-based analysis is particularly vital for a study on polygenic diseases where many interactions between molecules should be simultaneously examined and elucidated. We employ a new knowledge driven HDN gene and molecular database systems approach to analyze Inflammatory Bowel Disease (IBD), whose pathogenesis remains largely unknown. Methods and Results: Based on drug indications for IBD, we determined sibling diseases of mild and severe states of IBD. Approximately 1,000 genes associated with the sibling diseases were retrieved from four databases. After ranking the genes by the frequency of records in the databases, we obtained 250 and 253 genes highly associated with the mild and severe IBD states, respectively. We then calculated functional similarities of these genes with known drug targets and examined and presented their interactions as PPI networks. Conclusions: The results demonstrate that this knowledge-based systems approach, predicated on functionally similar genes important to sibling diseases is an effective method to identify important components of the IBD human disease network. Our approach elucidates a previously unknown biological distinction between mild and severe IBD states.
Publication COSMOS: Python library for massively parallel workflows
(Oxford University Press, 2014) Gafni, Erik; Luquette, Joe; Lancaster, Alex K.; Hawkins, Jared; Jung, Jae-Yoon; Souilmi, Yassine; Wall, Dennis P.; Tonellato, PeterSummary: Efficient workflows to shepherd clinically generated genomic data through the multiple stages of a next-generation sequencing pipeline are of critical importance in translational biomedical science. Here we present COSMOS, a Python library for workflow management that allows formal description of pipelines and partitioning of jobs. In addition, it includes a user interface for tracking the progress of jobs, abstraction of the queuing system and fine-grained control over the workflow. Workflows can be created on traditional computing clusters as well as cloud-based services. Availability and implementation: Source code is available for academic non-commercial research purposes. Links to code and documentation are provided at http://lpm.hms.harvard.edu and http://wall-lab.stanford.edu. Contact: dpwall@stanford.edu or peter_tonellato@hms.harvard.edu. Supplementary information: Supplementary data are available at Bioinformatics online.
Publication Analysis of sequence-based copy number variation detection tools for cancer studies
(American Medical Informatics Association, 2013) Nabavi, Sheida; Cai, Zhengqiu; Tonellato, PeterPublication Histopathologic Alterations Associated with Global Gene Expression Due to Chronic Dietary TCDD Exposure in Juvenile Zebrafish
(Public Library of Science, 2014) Liu, Qing; Spitsbergen, Jan M.; Cariou, Ronan; Huang, Chun-Yuan; Jiang, Nan; Goetz, Giles; Hutz, Reinhold J.; Tonellato, Peter; Carvan, Michael J.The goal of this project was to investigate the effects and possible developmental disease implication of chronic dietary TCDD exposure on global gene expression anchored to histopathologic analysis in juvenile zebrafish by functional genomic, histopathologic and analytic chemistry methods. Specifically, juvenile zebrafish were fed Biodiet starter with TCDD added at 0, 0.1, 1, 10 and 100 ppb, and fish were sampled following 0, 7, 14, 28 and 42 d after initiation of the exposure. TCDD accumulated in a dose- and time-dependent manner and 100 ppb TCDD caused TCDD accumulation in female (15.49 ppb) and male (18.04 ppb) fish at 28 d post exposure. Dietary TCDD caused multiple lesions in liver, kidney, intestine and ovary of zebrafish and functional dysregulation such as depletion of glycogen in liver, retrobulbar edema, degeneration of nasal neurosensory epithelium, underdevelopment of intestine, and diminution in the fraction of ovarian follicles containing vitellogenic oocytes. Importantly, lesions in nasal epithelium and evidence of endocrine disruption based on alternatively spliced vasa transcripts are two novel and significant results of this study. Microarray gene expression analysis comparing vehicle control to dietary TCDD revealed dysregulated genes involved in pathways associated with cardiac necrosis/cell death, cardiac fibrosis, renal necrosis/cell death and liver necrosis/cell death. These baseline toxicological effects provide evidence for the potential mechanisms of developmental dysfunctions induced by TCDD and vasa as a biomarker for ovarian developmental disruption.