Person: McCarroll, Steven
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Improved detection of global copy number variation using high density, non-polymorphic oligonucleotide probes
(BioMed Central, 2008) Shen, Fan; Huang, Jing; Fitch, Karen R.; Truong, Vivi B.; Kirby, Andrew; Chen, Wenwei; Zhang, Jane; Liu, Guoying; McCarroll, Steven; Jones, Keith W.; Shapero, Michael H.Background: DNA sequence diversity within the human genome may be more greatly affected by copy number variations (CNVs) than single nucleotide polymorphisms (SNPs). Although the importance of CNVs in genome wide association studies (GWAS) is becoming widely accepted, the optimal methods for identifying these variants are still under evaluation. We have previously reported a comprehensive view of CNVs in the HapMap DNA collection using high density 500 K EA (Early Access) SNP genotyping arrays which revealed greater than 1,000 CNVs ranging in size from 1 kb to over 3 Mb. Although the arrays used most commonly for GWAS predominantly interrogate SNPs, CNV identification and detection does not necessarily require the use of DNA probes centered on polymorphic nucleotides and may even be hindered by the dependence on a successful SNP genotyping assay. Results: In this study, we have designed and evaluated a high density array predicated on the use of non-polymorphic oligonucleotide probes for CNV detection. This approach effectively uncouples copy number detection from SNP genotyping and thus has the potential to significantly improve probe coverage for genome-wide CNV identification. This array, in conjunction with PCR-based, complexity-reduced DNA target, queries over 1.3 M independent NspI restriction enzyme fragments in the 200 bp to 1100 bp size range, which is a several fold increase in marker density as compared to the 500 K EA array. In addition, a novel algorithm was developed and validated to extract CNV regions and boundaries. Conclusion: Using a well-characterized pair of DNA samples, close to 200 CNVs were identified, of which nearly 50% appear novel yet were independently validated using quantitative PCR. The results indicate that non-polymorphic probes provide a robust approach for CNV identification, and the increasing precision of CNV boundary delineation should allow a more complete analysis of their genomic organization.
Publication Mapping Copy Number Variation by Population Scale Genome Sequencing
(Nature Publishing Group, 2011) Mills, Ryan Edward; Handsaker, Robert; Korn, Joshua; Nemesh, James; Shi, Xinghua; Lee, Charles; McCarroll, Steven; Altshuler, David; Gabriel, Stacey B.; Lander, Eric; Ambrogio, Lauren; Bloom, Toby; Cibulskis, Kristian; Fennell, Tim J.; Jaffe, David B.; Shefler, Erica; Sougnez, Carrie L.; Daly, Mark; DePristo, Mark A.; Ball, Aaron D.; Banks, Eric; Browning, Brian L.; Garimella, Kiran V.; Grossman, Sharon; Hanna, Matt; Hartl, Chris; Kernytsky, Andrew M.; Li, Heng; Maguire, Jared R.; McKenna, Aaron; Philippakis, Anthony Andrew; Poplin, Ryan E.; Price, Alkes; Rivas, Manuel A.; Sabeti, Pardis; Schaffner, Stephen; Shlyakhter, Ilya; Wilkinson, JaneGenomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.
Publication A Map of Human Genome Variation from Population Scale Sequencing
(Nature Publishing Group, 2010) Altshuler, David; Lander, Eric; Ambrogio, Lauren; Bloom, Toby; Cibulskis, Kristian; Fennell, Tim J.; Gabriel, Stacey B.; Jaffe, David B.; Shefler, Erica; Sougnez, Carrie L.; Lee, Charles; Mills, Ryan Edward; Shi, Xinghua; Daly, Mark; DePristo, Mark A.; Ball, Aaron D.; Banks, Eric; Browning, Brian L.; Garimella, Kiran V.; Grossman, Sharon; Handsaker, Robert; Hanna, Matt; Hartl, Chris; Kernytsky, Andrew M.; Korn, Joshua M.; Li, Heng; Maguire, Jared R.; McCarroll, Steven; Nemesh, James C.; McKenna, Aaron; Philippakis, Anthony Andrew; Poplin, Ryan E.; Price, Alkes; Rivas, Manuel A.; Sabeti, Pardis; Schaffner, Stephen; Shlyakhter, IlyaThe 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother–father–child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately (10^{−8}) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.
Publication Schizophrenia risk from complex variation of complement component 4
(2016) Sekar, Aswin; Rosen, Allison; de Rivera, Heather; Bell, Avery; Hammond, Timothy; Kamitaki, Nolan; Tooley, Katherine; Presumey, Jessy; Baum, Matt; Van Doren, Vanessa; Genovese, Giulio; Rose, Samuel A.; Handsaker, Robert; Daly, Mark; Carroll, Michael C.; Stevens, Beth; McCarroll, StevenSchizophrenia is a heritable brain illness with unknown pathogenic mechanisms. Schizophrenia’s strongest genetic association at a population level involves variation in the Major Histocompatibility Complex (MHC) locus, but the genes and molecular mechanisms accounting for this have been challenging to recognize. We show here that schizophrenia’s association with the MHC locus arises in substantial part from many structurally diverse alleles of the complement component 4 (C4) genes. We found that these alleles promoted widely varying levels of C4A and C4B expression and associated with schizophrenia in proportion to their tendency to promote greater expression of C4A in the brain. Human C4 protein localized at neuronal synapses, dendrites, axons, and cell bodies. In mice, C4 mediated synapse elimination during postnatal development. These results implicate excessive complement activity in the development of schizophrenia and may help explain the reduced numbers of synapses in the brains of individuals affected with schizophrenia.
Publication Using population admixture to help complete maps of the human genome
(2013) Genovese, Giulio; Handsaker, Robert; Li, Heng; Altemose, Nicolas; Lindgren, Amelia M.; Chambert, Kimberly; Pasaniuc, Bogdan; Price, Alkes; Reich, David; Morton, Cynthia; Pollak, Martin; Wilson, James G.; McCarroll, StevenTens of millions of base pairs of euchromatic human genome sequence, including many protein-coding genes, have no known location in the human genome. We describe an approach for localizing the human genome's missing pieces by utilizing the patterns of genome sequence variation created by population admixture. We mapped the locations of 70 scaffolds spanning four million base pairs of the human genome's unplaced euchromatic sequence, including more than a dozen protein-coding genes, and identified eight large novel inter-chromosomal segmental duplications. We find that most of these sequences are hidden in the genome's heterochromatin, particularly its pericentromeric regions. Many cryptic, pericentromeric genes are expressed in RNA and have been maintained intact for millions of years while their expression patterns diverged from those of paralogous genes elsewhere in the genome. We describe how knowledge of the locations of these sequences can inform disease association and genome biology studies.
Publication Biological Insights From 108 Schizophrenia-Associated Genetic Loci
(2014) Ripke, Stephan; Neale, Benjamin; Corvin, Aiden; Walters, James TR; Farh, Kai-How; Holmans, Peter A; Lee, Phil; Bulik-Sullivan, Brendan; Collier, David A; Huang, Hailiang; Pers, Tune H; Agartz, Ingrid; Agerbo, Esben; Albus, Margot; Alexander, Madeline; Amin, Farooq; Bacanu, Silviu A; Begemann, Martin; Belliveau, Richard A; Bene, Judit; Bergen, Sarah E; Bevilacqua, Elizabeth; Bigdeli, Tim B; Black, Donald W; Bruggeman, Richard; Buccola, Nancy G; Buckner, Randy; Byerley, William; Cahn, Wiepke; Cai, Guiqing; Campion, Dominique; Cantor, Rita M; Carr, Vaughan J; Carrera, Noa; Catts, Stanley V; Chambert, Kimberley D; Chan, Raymond CK; Chan, Ronald YL; Chen, Eric YH; Cheng, Wei; Cheung, Eric FC; Chong, Siow Ann; Cloninger, C Robert; Cohen, David; Cohen, Nadine; Cormican, Paul; Craddock, Nick; Crowley, James J; Curtis, David; Davidson, Michael; Davis, Kenneth L; Degenhardt, Franziska; Del Favero, Jurgen; Demontis, Ditte; Dikeos, Dimitris; Dinan, Timothy; Djurovic, Srdjan; Donohoe, Gary; Drapeau, Elodie; Duan, Jubao; Dudbridge, Frank; Durmishi, Naser; Eichhammer, Peter; Eriksson, Johan; Escott-Price, Valentina; Essioux, Laurent; Fanous, Ayman H; Farrell, Martilias S; Frank, Josef; Franke, Lude; Freedman, Robert; Freimer, Nelson B; Friedl, Marion; Friedman, Joseph I; Fromer, Menachem; Genovese, Giulio; Georgieva, Lyudmila; Giegling, Ina; Giusti-Rodríguez, Paola; Godard, Stephanie; Goldstein, Jacqueline I; Golimbet, Vera; Gopal, Srihari; Gratten, Jacob; de Haan, Lieuwe; Hammer, Christian; Hamshere, Marian L; Hansen, Mark; Hansen, Thomas; Haroutunian, Vahram; Hartmann, Annette M; Henskens, Frans A; Herms, Stefan; Hirschhorn, Joel; Hoffmann, Per; Hofman, Andrea; Hollegaard, Mads V; Hougaard, David M; Ikeda, Masashi; Joa, Inge; Julià, Antonio; Kahn, René S; Kalaydjieva, Luba; Karachanak-Yankova, Sena; Karjalainen, Juha; Kavanagh, David; Keller, Matthew C; Kennedy, James L; Khrunin, Andrey; Kim, Yunjung; Klovins, Janis; Knowles, James A; Konte, Bettina; Kucinskas, Vaidutis; Kucinskiene, Zita Ausrele; Kuzelova-Ptackova, Hana; Kähler, Anna K; Laurent, Claudine; Lee, Jimmy; Lee, S Hong; Legge, Sophie E; Lerer, Bernard; Li, Miaoxin; Li, Tao; Liang, Kung-Yee; Lieberman, Jeffrey; Limborska, Svetlana; Loughland, Carmel M; Lubinski, Jan; Lönnqvist, Jouko; Macek, Milan; Magnusson, Patrik KE; Maher, Brion S; Maier, Wolfgang; Mallet, Jacques; Marsal, Sara; Mattheisen, Manuel; Mattingsdal, Morten; McCarley, Robert William; McDonald, Colm; McIntosh, Andrew M; Meier, Sandra; Meijer, Carin J; Melegh, Bela; Melle, Ingrid; Mesholam-Gately, Raquelle; Metspalu, Andres; Michie, Patricia T; Milani, Lili; Milanova, Vihra; Mokrab, Younes; Morris, Derek W; Mors, Ole; Murphy, Kieran C; Murray, Robin M; Myin-Germeys, Inez; Müller-Myhsok, Bertram; Nelis, Mari; Nenadic, Igor; Nertney, Deborah A; Nestadt, Gerald; Nicodemus, Kristin K; Nikitina-Zake, Liene; Nisenbaum, Laura; Nordin, Annelie; O’Callaghan, Eadbhard; O’Dushlaine, Colm; O’Neill, F Anthony; Oh, Sang-Yun; Olincy, Ann; Olsen, Line; Van Os, Jim; Pantelis, Christos; Papadimitriou, George N; Papiol, Sergi; Parkhomenko, Elena; Pato, Michele T; Paunio, Tiina; Pejovic-Milovancevic, Milica; Perkins, Diana O; Pietiläinen, Olli; Pimm, Jonathan; Pocklington, Andrew J; Powell, John; Price, Alkes; Pulver, Ann E; Purcell, Shaun M; Quested, Digby; Rasmussen, Henrik B; Reichenberg, Abraham; Reimers, Mark A; Richards, Alexander L; Roffman, Joshua; Roussos, Panos; Ruderfer, Douglas M; Salomaa, Veikko; Sanders, Alan R; Schall, Ulrich; Schubert, Christian R; Schulze, Thomas G; Schwab, Sibylle G; Scolnick, Edward; Scott, Rodney J; Seidman, Larry Joel; Shi, Jianxin; Sigurdsson, Engilbert; Silagadze, Teimuraz; Silverman, Jeremy M; Sim, Kang; Slominsky, Petr; Smoller, Jordan; So, Hon-Cheong; Spencer, Chris C A; Stahl, Eli A; Stefansson, Hreinn; Steinberg, Stacy; Stogmann, Elisabeth; Straub, Richard E; Strengman, Eric; Strohmaier, Jana; Stroup, T Scott; Subramaniam, Mythily; Suvisaari, Jaana; Svrakic, Dragan M; Szatkiewicz, Jin P; Söderman, Erik; Thirumalai, Srinivas; Toncheva, Draga; Tosato, Sarah; Veijola, Juha; Waddington, John; Walsh, Dermot; Wang, Dai; Wang, Qiang; Webb, Bradley T; Weiser, Mark; Wildenauer, Dieter B; Williams, Nigel M; Williams, Stephanie; Witt, Stephanie H; Wolen, Aaron R; Wong, Emily HM; Wormley, Brandon K; Xi, Hualin Simon; Zai, Clement C; Zheng, Xuebin; Zimprich, Fritz; Wray, Naomi R; Stefansson, Kari; Visscher, Peter M; Adolfsson, Rolf; Andreassen, Ole A; Blackwood, Douglas HR; Bramon, Elvira; Buxbaum, Joseph D; Børglum, Anders D; Cichon, Sven; Darvasi, Ariel; Domenici, Enrico; Ehrenreich, Hannelore; Esko, Tõnu; Gejman, Pablo V; Gill, Michael; Gurling, Hugh; Hultman, Christina M; Iwata, Nakao; Jablensky, Assen V; Jönsson, Erik G; Kendler, Kenneth S; Kirov, George; Knight, Jo; Lencz, Todd; Levinson, Douglas F; Li, Qingqin S; Liu, Jianjun; Malhotra, Anil K; McCarroll, Steven; McQuillin, Andrew; Moran, Jennifer L; Mortensen, Preben B; Mowry, Bryan J; Nöthen, Markus M; Ophoff, Roel A; Owen, Michael J; Palotie, Aarno; Pato, Carlos N; Petryshen, Tracey L.; Posthuma, Danielle; Rietschel, Marcella; Riley, Brien P; Rujescu, Dan; Sham, Pak C; Sklar, Pamela; St Clair, David; Weinberger, Daniel R; Wendland, Jens R; Werge, Thomas; Daly, Mark; Sullivan, Patrick F; O’Donovan, Michael CSummary Schizophrenia is a highly heritable disorder. Genetic risk is conferred by a large number of alleles, including common alleles of small effect that might be detected by genome-wide association studies. Here, we report a multi-stage schizophrenia genome-wide association study of up to 36,989 cases and 113,075 controls. We identify 128 independent associations spanning 108 conservatively defined loci that meet genome-wide significance, 83 of which have not been previously reported. Associations were enriched among genes expressed in brain providing biological plausibility for the findings. Many findings have the potential to provide entirely novel insights into aetiology, but associations at DRD2 and multiple genes involved in glutamatergic neurotransmission highlight molecules of known and potential therapeutic relevance to schizophrenia, and are consistent with leading pathophysiological hypotheses. Independent of genes expressed in brain, associations were enriched among genes expressed in tissues that play important roles in immunity, providing support for the hypothesized link between the immune system and schizophrenia.
Publication Mutational heterogeneity in cancer and the search for new cancer genes
(2014) Lawrence, Michael S.; Stojanov, Petar; Polak, Paz; Kryukov, Gregory V.; Cibulskis, Kristian; Sivachenko, Andrey; Carter, Scott L.; Stewart, Chip; Mermel, Craig; Roberts, Steven A.; Kiezun, Adam; Hammerman, Peter S.; McKenna, Aaron; Drier, Yotam; Zou, Lihua; Ramos, Alex H.; Pugh, Trevor J.; Stransky, Nicolas; Helman, Elena; Kim, Jaegil; Sougnez, Carrie; Ambrogio, Lauren; Nickerson, Elizabeth; Shefler, Erica; Cortés, Maria L.; Auclair, Daniel; Saksena, Gordon; Voet, Douglas; Noble, Michael; DiCara, Daniel; Lin, Pei; Lichtenstein, Lee; Heiman, David I.; Fennell, Timothy; Imielinski, Marcin; Hernandez, Bryan; Hodis, Eran; Baca, Sylvan; Dulak, Austin M.; Lohr, Jens; Landau, Dan-Avi; Wu, Catherine; Melendez-Zajgla, Jorge; Hidalgo-Miranda, Alfredo; Koren, Amnon; McCarroll, Steven; Mora, Jaume; Crompton, Brian; Onofrio, Robert; Parkin, Melissa; Winckler, Wendy; Ardlie, Kristin; Gabriel, Stacey B.; Roberts, Charles W. M.; Biegel, Jaclyn A.; Stegmaier, Kimberly; Bass, Adam; Garraway, Levi; Meyerson, Matthew; Golub, Todd; Gordenin, Dmitry A.; Sunyaev, Shamil; Lander, Eric; Getz, GadMajor international projects are now underway aimed at creating a comprehensive catalog of all genes responsible for the initiation and progression of cancer. These studies involve sequencing of matched tumor–normal samples followed by mathematical analysis to identify those genes in which mutations occur more frequently than expected by random chance. Here, we describe a fundamental problem with cancer genome studies: as the sample size increases, the list of putatively significant genes produced by current analytical methods burgeons into the hundreds. The list includes many implausible genes (such as those encoding olfactory receptors and the muscle protein titin), suggesting extensive false positive findings that overshadow true driver events. Here, we show that this problem stems largely from mutational heterogeneity and provide a novel analytical methodology, MutSigCV, for resolving the problem. We apply MutSigCV to exome sequences from 3,083 tumor-normal pairs and discover extraordinary variation in (i) mutation frequency and spectrum within cancer types, which shed light on mutational processes and disease etiology, and (ii) mutation frequency across the genome, which is strongly correlated with DNA replication timing and also with transcriptional activity. By incorporating mutational heterogeneity into the analyses, MutSigCV is able to eliminate most of the apparent artefactual findings and allow true cancer genes to rise to attention.
Publication CNV analysis in a large schizophrenia sample implicates deletions at 16p12.1 and SLC1A1 and duplications at 1p36.33 and CGNL1
(Oxford University Press, 2013) Rees, Elliott; Walters, James T.R.; Chambert, Kimberly D.; O'Dushlaine, Colm; Szatkiewicz, Jin; Richards, Alexander L.; Georgieva, Lyudmila; Mahoney-Davies, Gerwyn; Legge, Sophie E.; Moran, Jennifer L.; Genovese, Giulio; Levinson, Douglas; Morris, Derek W.; Cormican, Paul; Kendler, Kenneth S.; O'Neill, Francis A.; Riley, Brien; Gill, Michael; Corvin, Aiden; Sklar, Pamela; Hultman, Christina; Pato, Carlos; Pato, Michele; Sullivan, Patrick F.; Gejman, Pablo V.; McCarroll, Steven; O'Donovan, Michael C.; Owen, Michael J.; Kirov, GeorgeLarge and rare copy number variants (CNVs) at several loci have been shown to increase risk for schizophrenia. Aiming to discover novel susceptibility CNV loci, we analyzed 6882 cases and 11 255 controls genotyped on Illumina arrays, most of which have not been used for this purpose before. We identified genes enriched for rare exonic CNVs among cases, and then attempted to replicate the findings in additional 14 568 cases and 15 274 controls. In a combined analysis of all samples, 12 distinct loci were enriched among cases with nominal levels of significance (P < 0.05); however, none would survive correction for multiple testing. These loci include recurrent deletions at 16p12.1, a locus previously associated with neurodevelopmental disorders (P = 0.0084 in the discovery sample and P = 0.023 in the replication sample). Other plausible candidates include non-recurrent deletions at the glutamate transporter gene SLC1A1, a CNV locus recently suggested to be involved in schizophrenia through linkage analysis, and duplications at 1p36.33 and CGNL1. A burden analysis of large (>500 kb), rare CNVs showed a 1.2% excess in cases after excluding known schizophrenia-associated loci, suggesting that additional susceptibility loci exist. However, even larger samples are required for their discovery.
Publication SnapShot-Seq: A Method for Extracting Genome-Wide, In Vivo mRNA Dynamics from a Single Total RNA Sample
(Public Library of Science, 2014) Gray, Jesse; Harmin, David; Boswell, Sarah; Cloonan, Nicole; Mullen, Thomas E.; Ling, Joseph J.; Miller, Nimrod; Kuersten, Scott; Ma, Yong-Chao; McCarroll, Steven; Grimmond, Sean M.; Springer, MichaelmRNA synthesis, processing, and destruction involve a complex series of molecular steps that are incompletely understood. Because the RNA intermediates in each of these steps have finite lifetimes, extensive mechanistic and dynamical information is encoded in total cellular RNA. Here we report the development of SnapShot-Seq, a set of computational methods that allow the determination of in vivo rates of pre-mRNA synthesis, splicing, intron degradation, and mRNA decay from a single RNA-Seq snapshot of total cellular RNA. SnapShot-Seq can detect in vivo changes in the rates of specific steps of splicing, and it provides genome-wide estimates of pre-mRNA synthesis rates comparable to those obtained via labeling of newly synthesized RNA. We used SnapShot-Seq to investigate the origins of the intrinsic bimodality of metazoan gene expression levels, and our results suggest that this bimodality is partly due to spillover of transcriptional activation from highly expressed genes to their poorly expressed neighbors. SnapShot-Seq dramatically expands the information obtainable from a standard RNA-Seq experiment.
Publication Discovery and genotyping of genome structural polymorphism by sequencing on a population scale
(2016) Handsaker, Robert; Korn, Joshua M.; Nemesh, James; McCarroll, StevenAccurate and complete analysis of genome variation in large populations will be required to understand the role of genome variation in complex disease. We present an analytical framework for characterizing genome deletion polymorphism in populations, using sequence data that are distributed across hundreds or thousands of genomes. Our approach uses population-level relationships to re-interpret the technical features of sequence data that often reflect structural variation. In the 1000 Genomes Project pilot, this approach identified deletion polymorphism across 168 genomes (sequenced at 4x average coverage) with sensitivity and specificity unmatched by other algorithms. We also describe a way to determine the allelic state or genotype of each deletion polymorphism in each genome; the 1000 Genomes Project used this approach to type 13,826 deletion polymorphisms (48 bp – 960 kbp) at high accuracy in populations. These methods offer a way to relate genome structural polymorphism to complex disease in populations.