Person:

Avillach, Paul

Loading...
Profile Picture

Email Address

AA Acceptance Date

Birth Date

Research Projects

Organizational Units

Job Title

Last Name

Avillach

First Name

Paul

Name

Avillach, Paul

Search Results

Now showing 1 - 10 of 12
  • Publication

    Data Extraction and Management in Networks of Observational Health Care Databases for Scientific Research: A Comparison of EU-ADR, OMOP, Mini-Sentinel and MATRICE Strategies

    (AcademyHealth, 2016) Gini, Rosa; Schuemie, Martijn; Brown, Jeffrey; Ryan, Patrick; Vacchi, Edoardo; Coppola, Massimo; Cazzola, Walter; Coloma, Preciosa; Berni, Roberto; Diallo, Gayo; Oliveira, José Luis; Avillach, Paul; Trifirò, Gianluca; Rijnbeek, Peter; Bellentani, Mariadonata; van Der Lei, Johan; Klazinga, Niek; Sturkenboom, Miriam

    Introduction: We see increased use of existing observational data in order to achieve fast and transparent production of empirical evidence in health care research. Multiple databases are often used to increase power, to assess rare exposures or outcomes, or to study diverse populations. For privacy and sociological reasons, original data on individual subjects can’t be shared, requiring a distributed network approach where data processing is performed prior to data sharing. Case Descriptions and Variation Among Sites: We created a conceptual framework distinguishing three steps in local data processing: (1) data reorganization into a data structure common across the network; (2) derivation of study variables not present in original data; and (3) application of study design to transform longitudinal data into aggregated data sets for statistical analysis. We applied this framework to four case studies to identify similarities and differences in the United States and Europe: Exploring and Understanding Adverse Drug Reactions by Integrative Mining of Clinical Records and Biomedical Knowledge (EU-ADR), Observational Medical Outcomes Partnership (OMOP), the Food and Drug Administration’s (FDA’s) Mini-Sentinel, and the Italian network—the Integration of Content Management Information on the Territory of Patients with Complex Diseases or with Chronic Conditions (MATRICE). Findings: National networks (OMOP, Mini-Sentinel, MATRICE) all adopted shared procedures for local data reorganization. The multinational EU-ADR network needed locally defined procedures to reorganize its heterogeneous data into a common structure. Derivation of new data elements was centrally defined in all networks but the procedure was not shared in EU-ADR. Application of study design was a common and shared procedure in all the case studies. Computer procedures were embodied in different programming languages, including SAS, R, SQL, Java, and C++. Conclusion: Using our conceptual framework we found several areas that would benefit from research to identify optimal standards for production of empirical knowledge from existing databases.an opportunity to advance evidence-based care management. In addition, formalized CM outcomes assessment methodologies will enable us to compare CM effectiveness across health delivery settings.

  • Publication

    Combining clinical and genomics queries using i2b2 – Three methods

    (Public Library of Science, 2017) Murphy, Shawn; Avillach, Paul; Bellazzi, Riccardo; Phillips, Lori; Gabetta, Matteo; Eran, Ally; McDuffie, Michael T.; Kohane, Isaac

    We are fortunate to be living in an era of twin biomedical data surges: a burgeoning representation of human phenotypes in the medical records of our healthcare systems, and high-throughput sequencing making rapid technological advances. The difficulty representing genomic data and its annotations has almost by itself led to the recognition of a biomedical “Big Data” challenge, and the complexity of healthcare data only compounds the problem to the point that coherent representation of both systems on the same platform seems insuperably difficult. We investigated the capability for complex, integrative genomic and clinical queries to be supported in the Informatics for Integrating Biology and the Bedside (i2b2) translational software package. Three different data integration approaches were developed: The first is based on Sequence Ontology, the second is based on the tranSMART engine, and the third on CouchDB. These novel methods for representing and querying complex genomic and clinical data on the i2b2 platform are available today for advancing precision medicine.

  • Publication

    A database of human exposomes and phenomes from the US National Health and Nutrition Examination Survey

    (Nature Publishing Group, 2016) Patel, Chirag; Pho, Nam; McDuffie, Michael T.; Easton-Marks, Jeremy; Kothari, Cartik; Kohane, Isaac; Avillach, Paul

    The National Health and Nutrition Examination Survey (NHANES) is a population survey implemented by the Centers for Disease Control and Prevention (CDC) to monitor the health of the United States whose data is publicly available in hundreds of files. This Data Descriptor describes a single unified and universally accessible data file, merging across 255 separate files and stitching data across 4 surveys, encompassing 41,474 individuals and 1,191 variables. The variables consist of phenotype and environmental exposure information on each individual, specifically (1) demographic information, physical exam results (e.g., height, body mass index), laboratory results (e.g., cholesterol, glucose, and environmental exposures), and (4) questionnaire items. Second, the data descriptor describes a dictionary to enable analysts find variables by category and human-readable description. The datasets are available on DataDryad and a hands-on analytics tutorial is available on GitHub. Through a new big data platform, BD2K Patient Centered Information Commons (http://pic-sure.org), we provide a new way to browse the dataset via a web browser (https://nhanes.hms.harvard.edu) and provide application programming interface for programmatic access.

  • Publication

    Identifying Cases of Type 2 Diabetes in Heterogeneous Data Sources: Strategy from the EMIF Project

    (Public Library of Science, 2016) Roberto, Giuseppe; Leal, Ingrid; Sattar, Naveed; Loomis, A. Katrina; Avillach, Paul; Egger, Peter; van Wijngaarden, Rients; Ansell, David; Reisberg, Sulev; Tammesoo, Mari-Liis; Alavere, Helene; Pasqua, Alessandro; Pedersen, Lars; Cunningham, James; Tramontan, Lara; Mayer, Miguel A.; Herings, Ron; Coloma, Preciosa; Lapi, Francesco; Sturkenboom, Miriam; van der Lei, Johan; Schuemie, Martijn J.; Rijnbeek, Peter; Gini, Rosa

    Due to the heterogeneity of existing European sources of observational healthcare data, data source-tailored choices are needed to execute multi-data source, multi-national epidemiological studies. This makes transparent documentation paramount. In this proof-of-concept study, a novel standard data derivation procedure was tested in a set of heterogeneous data sources. Identification of subjects with type 2 diabetes (T2DM) was the test case. We included three primary care data sources (PCDs), three record linkage of administrative and/or registry data sources (RLDs), one hospital and one biobank. Overall, data from 12 million subjects from six European countries were extracted. Based on a shared event definition, sixteeen standard algorithms (components) useful to identify T2DM cases were generated through a top-down/bottom-up iterative approach. Each component was based on one single data domain among diagnoses, drugs, diagnostic test utilization and laboratory results. Diagnoses-based components were subclassified considering the healthcare setting (primary, secondary, inpatient care). The Unified Medical Language System was used for semantic harmonization within data domains. Individual components were extracted and proportion of population identified was compared across data sources. Drug-based components performed similarly in RLDs and PCDs, unlike diagnoses-based components. Using components as building blocks, logical combinations with AND, OR, AND NOT were tested and local experts recommended their preferred data source-tailored combination. The population identified per data sources by resulting algorithms varied from 3.5% to 15.7%, however, age-specific results were fairly comparable. The impact of individual components was assessed: diagnoses-based components identified the majority of cases in PCDs (93–100%), while drug-based components were the main contributors in RLDs (81–100%). The proposed data derivation procedure allowed the generation of data source-tailored case-finding algorithms in a standardized fashion, facilitated transparent documentation of the process and benchmarking of data sources, and provided bases for interpretation of possible inter-data source inconsistency of findings in future studies.

  • Publication

    Evaluating the Impact of Computerized Provider Order Entry on Medical Students Training at Bedside: A Randomized Controlled Trial

    (Public Library of Science, 2015) Wack, Maxime; Puymirat, Etienne; Ranque, Brigitte; Georgin-Lavialle, Sophie; Pierre, Isabelle; Tanguy, Aurelia; Ackermann, Felix; Mallet, Celine; Pavie, Juliette; Boultache, Hakima; Durieux, Pierre; Avillach, Paul

    Objective: To evaluate the impact of computerized provider order entry (CPOE) at the bedside on medical students training. Materials and Methods We conducted a randomized cross-controlled educational trial on medical students during two clerkship rotations in three departments, assessing the impact of the use of CPOE on their ability to place adequate monitoring and therapeutic orders using a written test before and after each rotation. Students’ satisfaction with their practice and the order placement system was surveyed. A multivariate mixed model was used to take individual students and chief resident (CR) effects into account. Factorial analysis was applied on the satisfaction questionnaire to identify dimensions, and scores were compared on these dimensions. Results: Thirty-six students show no better progress (beginning and final test means = 69.87 and 80.98 points out of 176 for the control group, 64.60 and 78.11 for the CPOE group, p = 0.556) during their rotation in either group, even after adjusting for each student and CR, but show a better satisfaction with patient care and greater involvement in the medical team in the CPOE group (p = 0.035*). Both groups have a favorable opinion regarding CPOE as an educational tool, especially because of the order reviewing by the supervisor. Conclusion: This is the first randomized controlled trial assessing the performance of CPOE in both the progress in prescriptions ability and satisfaction of the students. The absence of effect on the medical skills must be weighted by the small time scale and low sample size. However, students are more satisfied when using CPOE rather than usual training.

  • Publication

    Development of the Precision Link Biobank at Boston Children’s Hospital: Challenges and Opportunities

    (MDPI, 2017) Bourgeois, Florence; Avillach, Paul; Kong, Sek Won; Heinz, Michelle M.; Tran, Tram A.; Chakrabarty, Ramkrishna; Bickel, Jonathan; Sliz, Piotr; Borglund, Erin M.; Kornetsky, Susan; Mandl, Kenneth

    Increasingly, biobanks are being developed to support organized collections of biological specimens and associated clinical information on broadly consented, diverse patient populations. We describe the implementation of a pediatric biobank, comprised of a fully-informed patient cohort linking specimens to phenotypic data derived from electronic health records (EHR). The Biobank was launched after multiple stakeholders’ input and implemented initially in a pilot phase before hospital-wide expansion in 2016. In-person informed consent is obtained from all participants enrolling in the Biobank and provides permission to: (1) access EHR data for research; (2) collect and use residual specimens produced as by-products of routine care; and (3) share de-identified data and specimens outside of the institution. Participants are recruited throughout the hospital, across diverse clinical settings. We have enrolled 4900 patients to date, and 41% of these have an associated blood sample for DNA processing. Current efforts are focused on aligning the Biobank with other ongoing research efforts at our institution and extending our electronic consenting system to support remote enrollment. A number of pediatric-specific challenges and opportunities is reviewed, including the need to re-consent patients when they reach 18 years of age, the ability to enroll family members accompanying patients and alignment with disease-specific research efforts at our institution and other pediatric centers to increase cohort sizes, particularly for rare diseases.

  • Publication

    CodeMapper: semiautomatic coding of case definitions. A contribution from the ADVANCE project

    (John Wiley and Sons Inc., 2017) Becker, Benedikt F.H.; Avillach, Paul; Romio, Silvana; van Mulligen, Erik M.; Weibel, Daniel; Sturkenboom, Miriam C.J.M.; Kors, Jan A.

    Abstract Background: Assessment of drug and vaccine effects by combining information from different healthcare databases in the European Union requires extensive efforts in the harmonization of codes as different vocabularies are being used across countries. In this paper, we present a web application called CodeMapper, which assists in the mapping of case definitions to codes from different vocabularies, while keeping a transparent record of the complete mapping process. Methods: CodeMapper builds upon coding vocabularies contained in the Metathesaurus of the Unified Medical Language System. The mapping approach consists of three phases. First, medical concepts are automatically identified in a free‐text case definition. Second, the user revises the set of medical concepts by adding or removing concepts, or expanding them to related concepts that are more general or more specific. Finally, the selected concepts are projected to codes from the targeted coding vocabularies. We evaluated the application by comparing codes that were automatically generated from case definitions by applying CodeMapper's concept identification and successive concept expansion, with reference codes that were manually created in a previous epidemiological study. Results: Automated concept identification alone had a sensitivity of 0.246 and positive predictive value (PPV) of 0.420 for reproducing the reference codes. Three successive steps of concept expansion increased sensitivity to 0.953 and PPV to 0.616. Conclusions: Automatic concept identification in the case definition alone was insufficient to reproduce the reference codes, but CodeMapper's operations for concept expansion provide an effective, efficient, and transparent way for reproducing the reference codes.

  • Publication

    Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful API

    (Oxford University Press, 2017) Gutiérrez-Sacristán, Alba; Guedj, Romain; Korodi, Gabor; Stedman, Jason; Furlong, Laura I; Patel, Chirag; Kohane, Isaac; Avillach, Paul

    Abstract Motivation In the era of big data and precision medicine, the number of databases containing clinical, environmental, self-reported and biochemical variables is increasing exponentially. Enabling the experts to focus on their research questions rather than on computational data management, access and analysis is one of the most significant challenges nowadays. Results: We present Rcupcake, an R package that contains a variety of functions for leveraging different databases through the BD2K PIC-SURE RESTful API and facilitating its query, analysis and interpretation. The package offers a variety of analysis and visualization tools, including the study of the phenotype co-occurrence and prevalence, according to multiple layers of data, such as phenome, exposome or genome. Availability and implementation The package is implemented in R and is available under Mozilla v2 license from GitHub (https://github.com/hms-dbmi/Rcupcake). Two reproducible case studies are also available (https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/SSCcaseStudy_v01.ipynb, https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/NHANEScaseStudy_v01.ipynb). Contact paul_avillach@hms.harvard.edu Supplementary information Supplementary data are available at Bioinformatics online.

  • Publication

    Screening pregnant women for suicidal behavior in electronic medical records: diagnostic codes vs. clinical notes processed by natural language processing

    (BioMed Central, 2018) Zhong, Qiu-Yue; Karlson, Elizabeth; Gelaye, Bizu; Finan, Sean; Avillach, Paul; Smoller, Jordan; Cai, Tianxi; Williams, Michelle

    Background: We examined the comparative performance of structured, diagnostic codes vs. natural language processing (NLP) of unstructured text for screening suicidal behavior among pregnant women in electronic medical records (EMRs). Methods: Women aged 10–64 years with at least one diagnostic code related to pregnancy or delivery (N = 275,843) from Partners HealthCare were included as our “datamart.” Diagnostic codes related to suicidal behavior were applied to the datamart to screen women for suicidal behavior. Among women without any diagnostic codes related to suicidal behavior (n = 273,410), 5880 women were randomly sampled, of whom 1120 had at least one mention of terms related to suicidal behavior in clinical notes. NLP was then used to process clinical notes for the 1120 women. Chart reviews were performed for subsamples of women. Results: Using diagnostic codes, 196 pregnant women were screened positive for suicidal behavior, among whom 149 (76%) had confirmed suicidal behavior by chart review. Using NLP among those without diagnostic codes, 486 pregnant women were screened positive for suicidal behavior, among whom 146 (30%) had confirmed suicidal behavior by chart review. Conclusions: The use of NLP substantially improves the sensitivity of screening suicidal behavior in EMRs. However, the prevalence of confirmed suicidal behavior was lower among women who did not have diagnostic codes for suicidal behavior but screened positive by NLP. NLP should be used together with diagnostic codes for future EMR-based phenotyping studies for suicidal behavior. Electronic supplementary material The online version of this article (10.1186/s12911-018-0617-7) contains supplementary material, which is available to authorized users.

  • Publication

    Adverse obstetric and neonatal outcomes complicated by psychosis among pregnant women in the United States

    (BioMed Central, 2018) Zhong, Qiu-Yue; Gelaye, Bizu; Fricchione, Gregory; Avillach, Paul; Karlson, Elizabeth; Williams, Michelle

    Background: Adverse obstetric and neonatal outcomes among women with psychosis, particularly affective psychosis, has rarely been studied at the population level. We aimed to assess the risk of adverse obstetric and neonatal outcomes among women with psychosis (schizophrenia, affective psychosis, and other psychoses). Methods: From the 2007 – 2012 National (Nationwide) Inpatient Sample, 23,507,597 delivery hospitalizations were identified. From the same hospitalization, International Classification of Diseases diagnosis codes were used to identify maternal psychosis and outcomes. Adjusted odds ratios (aOR) and 95% confidence intervals (CI) were obtained using logistic regression. Results: The prevalence of psychosis at delivery was 698.76 per 100,000 hospitalizations. After adjusting for sociodemographic characteristics, smoking, alcohol/substance abuse, and pregnancy-related hypertension, women with psychosis were at a heightened risk for cesarean delivery (aOR = 1.26; 95% CI: 1.23 - 1.29), induced labor (aOR = 1.05; 95% CI: 1.02 - 1.09), antepartum hemorrhage (aOR = 1.22; 95% CI: 1.14 - 1.31), placental abruption (aOR = 1.22; 95% CI: 1.13 - 1.32), postpartum hemorrhage (aOR = 1.18; 95% CI: 1.10 - 1.27), premature delivery (aOR = 1.40; 95% CI: 1.36 - 1.46), stillbirth (aOR = 1.37; 95% CI: 1.23 - 1.53), premature rupture of membranes (aOR = 1.22; 95% CI: 1.15 - 1.29), fetal abnormalities (aOR = 1.49; 95% CI: 1.38 - 1.61), poor fetal growth (aOR = 1.26; 95% CI: 1.19 - 1.34), and fetal distress (aOR = 1.14; 95% CI: 1.10 - 1.18). Maternal death during hospitalizations (aOR = 1.00; 95% CI: 0.30 - 3.31) and excessive fetal growth (aOR = 1.06; 95% CI: 0.98 - 1.14) were not statistically significantly associated with psychosis. Conclusions: Pregnant women with psychosis have elevated risk of several adverse obstetric and neonatal outcomes. Efforts to identify and manage pregnancies complicated by psychosis may contribute to improved outcomes. Electronic supplementary material The online version of this article (10.1186/s12884-018-1750-0) contains supplementary material, which is available to authorized users.