Person:

Irizarry, Rafael

Loading...
Profile Picture

Email Address

AA Acceptance Date

Birth Date

Research Projects

Organizational Units

Job Title

Last Name

Irizarry

First Name

Rafael

Name

Irizarry, Rafael

Search Results

Now showing 1 - 10 of 17
  • Publication

    quantro: a data-driven approach to guide the choice of an appropriate normalization method

    (BioMed Central, 2015) Hicks, Stephanie C.; Irizarry, Rafael

    Normalization is an essential step in the analysis of high-throughput data. Multi-sample global normalization methods, such as quantile normalization, have been successfully used to remove technical variation. However, these methods rely on the assumption that observed global changes across samples are due to unwanted technical variability. Applying global normalization methods has the potential to remove biologically driven variation. Currently, it is up to the subject matter experts to determine if the stated assumptions are appropriate. Here, we propose a data-driven alternative. We demonstrate the utility of our method (quantro) through examples and simulations. A software implementation is available from http://www.bioconductor.org/packages/release/bioc/html/quantro.html. Electronic supplementary material The online version of this article (doi:10.1186/s13059-015-0679-0) contains supplementary material, which is available to authorized users.

  • Publication

    Large hypomethylated blocks as a universal defining epigenetic alteration in human solid tumors

    (BioMed Central, 2014) Timp, Winston; Bravo, Hector Corrada; McDonald, Oliver G; Goggins, Michael; Umbricht, Chris; Zeiger, Martha; Feinberg, Andrew P; Irizarry, Rafael

    Background: One of the most provocative recent observations in cancer epigenetics is the discovery of large hypomethylated blocks, including single copy genes, in colorectal cancer, that correspond in location to heterochromatic LOCKs (large organized chromatin lysine-modifications) and LADs (lamin-associated domains). Methods: Here we performed a comprehensive genome-scale analysis of 10 breast, 28 colon, nine lung, 38 thyroid, 18 pancreas cancers, and five pancreas neuroendocrine tumors as well as matched normal tissue from most of these cases, as well as 51 premalignant lesions. We used a new statistical approach that allows the identification of large hypomethylated blocks on the Illumina HumanMethylation450 BeadChip platform. Results: We find that hypomethylated blocks are a universal feature of common solid human cancer, and that they occur at the earliest stage of premalignant tumors and progress through clinical stages of thyroid and colon cancer development. We also find that the disrupted CpG islands widely reported previously, including hypermethylated island bodies and hypomethylated shores, are enriched in hypomethylated blocks, with flattening of the methylation signal within and flanking the islands. Finally, we found that genes showing higher between individual gene expression variability are enriched within these hypomethylated blocks. Conclusion: Thus hypomethylated blocks appear to be a universal defining epigenetic alteration in human cancer, at least for common solid tumors. Electronic supplementary material The online version of this article (doi:10.1186/s13073-014-0061-y) contains supplementary material, which is available to authorized users.

  • Publication

    MAGeCK enables robust identification of essential genes from genome-scale CRISPR/Cas9 knockout screens

    (BioMed Central, 2014) Li, Wei; Xu, Han; Xiao, Tengfei; Cong, Le; Love, Michael I.; Zhang, Feng; Irizarry, Rafael; Liu, Jun; Brown, Myles; Liu, X Shirley

    We propose the Model-based Analysis of Genome-wide CRISPR/Cas9 Knockout (MAGeCK) method for prioritizing single-guide RNAs, genes and pathways in genome-scale CRISPR/Cas9 knockout screens. MAGeCK demonstrates better performance compared with existing methods, identifies both positively and negatively selected genes simultaneously, and reports robust results across different experimental conditions. Using public datasets, MAGeCK identified novel essential genes and pathways, including EGFR in vemurafenib-treated A375 cells harboring a BRAF mutation. MAGeCK also detected cell type-specific essential genes, including BCR and ABL1, in KBM7 cells bearing a BCR-ABL fusion, and IGF1R in HL-60 cells, which depends on the insulin signaling pathway for proliferation. Electronic supplementary material The online version of this article (doi:10.1186/s13059-014-0554-4) contains supplementary material, which is available to authorized users.

  • Publication

    Accounting for cellular heterogeneity is critical in epigenome-wide association studies

    (BioMed Central, 2014) Jaffe, Andrew E; Irizarry, Rafael

    Background: Epigenome-wide association studies of human disease and other quantitative traits are becoming increasingly common. A series of papers reporting age-related changes in DNA methylation profiles in peripheral blood have already been published. However, blood is a heterogeneous collection of different cell types, each with a very different DNA methylation profile. Results: Using a statistical method that permits estimating the relative proportion of cell types from DNA methylation profiles, we examine data from five previously published studies, and find strong evidence of cell composition change across age in blood. We also demonstrate that, in these studies, cellular composition explains much of the observed variability in DNA methylation. Furthermore, we find high levels of confounding between age-related variability and cellular composition at the CpG level. Conclusions: Our findings underscore the importance of considering cell composition variability in epigenetic studies based on whole blood and other heterogeneous tissue sources. We also provide software for estimating and exploring this composition confounding for the Illumina 450k microarray.

  • Publication

    Flexible expressed region analysis for RNA-seq with derfinder

    (Oxford University Press, 2017) Collado-Torres, Leonardo; Nellore, Abhinav; Frazee, Alyssa C.; Wilks, Christopher; Love, Michael I.; Langmead, Ben; Irizarry, Rafael; Leek, Jeffrey T.; Jaffe, Andrew E.

    Differential expression analysis of RNA sequencing (RNA-seq) data typically relies on reconstructing transcripts or counting reads that overlap known gene structures. We previously introduced an intermediate statistical approach called differentially expressed region (DER) finder that seeks to identify contiguous regions of the genome showing differential expression signal at single base resolution without relying on existing annotation or potentially inaccurate transcript assembly. We present the derfinder software that improves our annotation-agnostic approach to RNA-seq analysis by: (i) implementing a computationally efficient bump-hunting approach to identify DERs that permits genome-scale analyses in a large number of samples, (ii) introducing a flexible statistical modeling framework, including multi-group and time-course analyses and (iii) introducing a new set of data visualizations for expressed region analysis. We apply this approach to public RNA-seq data from the Genotype-Tissue Expression (GTEx) project and BrainSpan project to show that derfinder permits the analysis of hundreds of samples at base resolution in R, identifies expression outside of known gene boundaries and can be used to visualize expressed regions at base-resolution. In simulations, our base resolution approaches enable discovery in the presence of incomplete annotation and is nearly as powerful as feature-level methods when the annotation is complete. derfinder analysis using expressed region-level and single base-level approaches provides a compromise between full transcript reconstruction and feature-level analysis. The package is available from Bioconductor at www.bioconductor.org/packages/derfinder.

  • Publication

    Modeling of RNA-seq fragment sequence bias reduces systematic errors in transcript abundance estimation

    (2016) Love, Michael I.; Hogenesch, John B.; Irizarry, Rafael
  • Publication

    A benchmark for RNA-seq quantification pipelines

    (BioMed Central, 2016) Teng, Mingxiang; Love, Michael I.; Davis, Carrie A.; Djebali, Sarah; Dobin, Alexander; Graveley, Brenton R.; Li, Sheng; Mason, Christopher E.; Olson, Sara; Pervouchine, Dmitri; Sloan, Cricket A.; Wei, Xintao; Zhan, Lijun; Irizarry, Rafael

    Obtaining RNA-seq measurements involves a complex data analytical process with a large number of competing algorithms as options. There is much debate about which of these methods provides the best approach. Unfortunately, it is currently difficult to evaluate their performance due in part to a lack of sensitive assessment metrics. We present a series of statistical summaries and plots to evaluate the performance in terms of specificity and sensitivity, available as a R/Bioconductor package (http://bioconductor.org/packages/rnaseqcomp). Using two independent datasets, we assessed seven competing pipelines. Performance was generally poor, with two methods clearly underperforming and RSEM slightly outperforming the rest. Electronic supplementary material The online version of this article (doi:10.1186/s13059-016-0940-1) contains supplementary material, which is available to authorized users.

  • Publication

    Every Body Counts: Measuring Mortality From the COVID-19 Pandemic

    (American College of Physicians, 2020-09-11) Kiang, Mathew; Irizarry, Rafael; Buckee, Caroline; Balsari, Satchit

    As of mid-August 2020, more than 170 000 U.S. residents have died of coronavirus disease 2019 (COVID-19); however, the true number of deaths resulting from COVID-19, both directly and indirectly, is likely to be much higher. The proper attribution of deaths to this pandemic has a range of societal, legal, mortuary, and public health consequences. This article discusses the current difficulties of disaster death attribution and describes the strengths and limitations of relying on death counts from death certificates, estimations of indirect deaths, and estimations of excess mortality. Improving the tabulation of direct and indirect deaths on death certificates will require concerted efforts and consensus across medical institutions and public health agencies. In addition, actionable estimates of excess mortality will require timely access to standardized and structured vital registry data, which should be shared directly at the state level to ensure rapid response for local governments. Correct attribution of direct and indirect deaths and estimation of excess mortality are complementary goals that are critical to our understanding of the pandemic and its effect on human life.

  • Publication

    Erratum to: A benchmark for RNA-seq quantification pipelines

    (BioMed Central, 2016) Teng, Mingxiang; Love, Michael I.; Davis, Carrie A.; Djebali, Sarah; Dobin, Alexander; Graveley, Brenton R.; Li, Sheng; Mason, Christopher E.; Olson, Sara; Pervouchine, Dmitri; Sloan, Cricket A.; Wei, Xintao; Zhan, Lijun; Irizarry, Rafael
  • Publication

    Challenges and emerging directions in single-cell analysis

    (BioMed Central, 2017) Yuan, Guo-Cheng; Cai, Long; Elowitz, Michael; Enver, Tariq; Fan, Guoping; Guo, Guoji; Irizarry, Rafael; Kharchenko, Peter; Kim, Junhyong; Orkin, Stuart; Quackenbush, John; Saadatpour, Assieh; Schroeder, Timm; Shivdasani, Ramesh; Tirosh, Itay

    Single-cell analysis is a rapidly evolving approach to characterize genome-scale molecular information at the individual cell level. Development of single-cell technologies and computational methods has enabled systematic investigation of cellular heterogeneity in a wide range of tissues and cell populations, yielding fresh insights into the composition, dynamics, and regulatory mechanisms of cell states in development and disease. Despite substantial advances, significant challenges remain in the analysis, integration, and interpretation of single-cell omics data. Here, we discuss the state of the field and recent advances and look to future opportunities.