Person: Samocha, Kaitlin E.
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Patterns and rates of exonic de novo mutations in autism spectrum disorders
(2013) Neale, Benjamin; Kou, Yan; Liu, Li; Ma'ayan, Avi; Samocha, Kaitlin E.; Sabo, Aniko; Lin, Chiao-Feng; Stevens, Christine; Wang, Li-San; Makarov, Vladimir; Polak, Paz; Yoon, Seungtai; Maguire, Jared; Crawford, Emily L.; Campbell, Nicholas G.; Geller, Evan T.; Valladares, Otto; Shafer, Chad; Liu, Han; Zhao, Tuo; Cai, Guiqing; Lihm, Jayon; Dannenfelser, Ruth; Jabado, Omar; Peralta, Zuleyma; Nagaswamy, Uma; Muzny, Donna; Reid, Jeffrey G.; Newsham, Irene; Wu, Yuanqing; Lewis, Lora; Han, Yi; Voight, Benjamin F.; Lim, Elaine; Rossin, Elizabeth; Kirby, Andrew; Flannick, Jason; Fromer, Menachem; Shakir, Khalid; Fennell, Tim; Garimella, Kiran; Banks, Eric; Poplin, Ryan; Gabriel, Stacey; DePristo, Mark; Wimbish, Jack R.; Boone, Braden E.; Levy, Shawn E.; Betancur, Catalina; Sunyaev, Shamil; Boerwinkle, Eric; Buxbaum, Joseph D.; Cook, Edwin H.; Devlin, Bernie; Gibbs, Richard A.; Roeder, Kathryn; Schellenberg, Gerard D.; Sutcliffe, James S.; Daly, MarkAutism spectrum disorders (ASD) are believed to have genetic and environmental origins, yet in only a modest fraction of individuals can specific causes be identified1,2. To identify further genetic risk factors, we assess the role of de novo mutations in ASD by sequencing the exomes of ASD cases and their parents (n= 175 trios). Fewer than half of the cases (46.3%) carry a missense or nonsense de novo variant and the overall rate of mutation is only modestly higher than the expected rate. In contrast, there is significantly enriched connectivity among the proteins encoded by genes harboring de novo missense or nonsense mutations, and excess connectivity to prior ASD genes of major effect, suggesting a subset of observed events are relevant to ASD risk. The small increase in rate of de novo events, when taken together with the connections among the proteins themselves and to ASD, are consistent with an important but limited role for de novo point mutations, similar to that documented for de novo copy number variants. Genetic models incorporating these data suggest that the majority of observed de novo events are unconnected to ASD, those that do confer risk are distributed across many genes and are incompletely penetrant (i.e., not necessarily causal). Our results support polygenic models in which spontaneous coding mutations in any of a large number of genes increases risk by 5 to 20-fold. Despite the challenge posed by such models, results from de novo events and a large parallel case-control study provide strong evidence in favor of CHD8 and KATNAL2 as genuine autism risk factors.
Publication Analysis of Rare, Exonic Variation amongst Subjects with Autism Spectrum Disorders and Population Controls
(Public Library of Science, 2013) Liu, Li; Sabo, Aniko; Neale, Benjamin; Nagaswamy, Uma; Stevens, Christine; Lim, Elaine; Bodea, Corneliu A.; Muzny, Donna; Reid, Jeffrey G.; Banks, Eric; Coon, Hillary; DePristo, Mark; Dinh, Huyen; Fennel, Tim; Flannick, Jason; Gabriel, Stacey; Garimella, Kiran; Gross, Shannon; Hawes, Alicia; Lewis, Lora; Makarov, Vladimir; Maguire, Jared; Newsham, Irene; Poplin, Ryan; Ripke, Stephan; Shakir, Khalid; Samocha, Kaitlin E.; Wu, Yuanqing; Boerwinkle, Eric; Buxbaum, Joseph D.; Cook, Edwin H., Jr.; Devlin, Bernie; Schellenberg, Gerard D.; Sutcliffe, James S.; Daly, Mark; Gibbs, Richard A.; Roeder, KathrynWe report on results from whole-exome sequencing (WES) of 1,039 subjects diagnosed with autism spectrum disorders (ASD) and 870 controls selected from the NIMH repository to be of similar ancestry to cases. The WES data came from two centers using different methods to produce sequence and to call variants from it. Therefore, an initial goal was to ensure the distribution of rare variation was similar for data from different centers. This proved straightforward by filtering called variants by fraction of missing data, read depth, and balance of alternative to reference reads. Results were evaluated using seven samples sequenced at both centers and by results from the association study. Next we addressed how the data and/or results from the centers should be combined. Gene-based analyses of association was an obvious choice, but should statistics for association be combined across centers (meta-analysis) or should data be combined and then analyzed (mega-analysis)? Because of the nature of many gene-based tests, we showed by theory and simulations that mega-analysis has better power than meta-analysis. Finally, before analyzing the data for association, we explored the impact of population structure on rare variant analysis in these data. Like other recent studies, we found evidence that population structure can confound case-control studies by the clustering of rare variants in ancestry space; yet, unlike some recent studies, for these data we found that principal component-based analyses were sufficient to control for ancestry and produce test statistics with appropriate distributions. After using a variety of gene-based tests and both meta- and mega-analysis, we found no new risk genes for ASD in this sample. Our results suggest that standard gene-based tests will require much larger samples of cases and controls before being effective for gene discovery, even for a disorder like ASD.
Publication A framework for the detection of de novo mutations in family-based sequencing data
(Nature Publishing Group, 2016) Francioli, Laurent; Cretu-Stancu, Mircea; Garimella, Kiran V; Fromer, Menachem; Kloosterman, Wigard P; Wijmenga, Cisca; Investigator, Principal; Swertz, Morris A; van Duijn, Cornelia M; Boomsma, Dorret I; Slagboom, PEline; van Ommen, Gertjan B; de Bakker, Paul IW; van Dijk, Freerk; Menelaou, Androniki; Neerincx, Pieter BT; Pulit, Sara L; Deelen, Patrick; Elbers, Clara C; Francesco Palamara, Pier; Pe'er, Itsik; Abdellaoui, Abdel; van Oven, Mannis; Vermaat, Martijn; Li, Mingkun; Laros, Jeroen FJ; Stoneking, Mark; de Knijff, Peter; Kayser, Manfred; Veldink, Jan H; van den Berg, Leonard H; Byelas, Heorhiy; den Dunnen, Johan T; Dijkstra, Martijn; Amin, Najaf; van der Velde, K Joeri; Hottenga, Jouke Jan; van Setten, Jessica; van Leeuwen, Elisabeth M; Kanterakis, Alexandros; Kattenberg, Mathijs; Karssen, Lennart C; van Schaik, Barbera DC; Bot, Jan; Nijman, Isaäc J; Renkens, Ivo; van Enckevort, David; Mei, Hailiang; Koval, Vyacheslav; Estrada, Karol; Medina-Gomez, Carolina; Ye, Kai; Lameijer, Eric-Wubbo; Moed, Matthijs H; Hehir-Kwa, Jayne Y; Handsaker, Robert E; McCarroll, Steven A; Sunyaev, Shamil R; Polak, Paz; Vuzman, Dana; Sohail, Mashaal; Hormozdiari, Fereydoun; Marschall, Tobias; Schönhuth, Alexander; Guryev, Victor; Slagboom, P Eline; Beekman, Marian B; de Craen, Anton JM; Suchiman, H Eka D; Hofman, Albert; Oostra, Ben; Isaacs, Aaron; Rivadeneira, Fernando; Uitterlinden, André G; Willemsen, Gonneke; Platteel, Mathieu; Pitts, Steven J; Potluri, Shobha; Sundar, Purnima; Cox, David R; Li, Qibin; Li, Yingrui; Du, Yuanping; Chen, Ruoyan; Cao, Hongzhi; Li, Ning; Cao, Sujie; Wang, Jun; Bovenberg, Jasper A; Brandsma, Margreet; Samocha, Kaitlin E.; Neale, Benjamin; Daly, Mark; Banks, Eric; DePristo, Mark AGermline mutation detection from human DNA sequence data is challenging due to the rarity of such events relative to the intrinsic error rates of sequencing technologies and the uneven coverage across the genome. We developed PhaseByTransmission (PBT) to identify de novo single nucleotide variants and short insertions and deletions (indels) from sequence data collected in parent-offspring trios. We compute the joint probability of the data given the genotype likelihoods in the individual family members, the known familial relationships and a prior probability for the mutation rate. Candidate de novo mutations (DNMs) are reported along with their posterior probability, providing a systematic way to prioritize them for validation. Our tool is integrated in the Genome Analysis Toolkit and can be used together with the ReadBackedPhasing module to infer the parental origin of DNMs based on phase-informative reads. Using simulated data, we show that PBT outperforms existing tools, especially in low coverage data and on the X chromosome. We further show that PBT displays high validation rates on empirical parent-offspring sequencing data for whole-exome data from 104 trios and X-chromosome data from 249 parent-offspring families. Finally, we demonstrate an association between father's age at conception and the number of DNMs in female offspring's X chromosome, consistent with previous literature reports.
Publication Analysis of protein-coding genetic variation in 60,706 humans
(2016) Lek, Monkol; Karczewski, Konrad; Minikel, Eric; Samocha, Kaitlin E.; Banks, Eric; Fennell, Timothy; O'Donnell-Luria, Anne H; Ware, James S; Hill, Andrew J; Cummings, Beryl; Tukiainen, Taru; Birnbaum, Daniel P; Kosmicki, Jack; Duncan, Laramie E; Estrada, Karol; Zhao, Fengmei; Zou, James; Pierce-Hoffman, Emma; Berghout, Joanne; Cooper, David N; Deflaux, Nicole; DePristo, Mark; Do, Ron; Flannick, Jason; Fromer, Menachem; Gauthier, Laura; Goldstein, Jackie; Gupta, Namrata; Howrigan, Daniel; Kiezun, Adam; Kurki, Mitja; Moonshine, Ami Levy; Natarajan, Pradeep; Orozco, Lorena; Peloso, Gina M; Poplin, Ryan; Rivas, Manuel A; Ruano-Rubio, Valentin; Rose, Samuel A; Ruderfer, Douglas M; Shakir, Khalid; Stenson, Peter D; Stevens, Christine; Thomas, Brett P; Tiao, Grace; Tusie-Luna, Maria T; Weisburd, Ben; Won, Hong-Hee; Yu, Dongmei; Altshuler, David; Ardissino, Diego; Boehnke, Michael; Danesh, John; Donnelly, Stacey; Elosua, Roberto; Florez, Jose; Gabriel, Stacey B; Getz, Gad; Glatt, Stephen J; Hultman, Christina M; Kathiresan, Sekar; Laakso, Markku; McCarroll, Steven; McCarthy, Mark I; McGovern, Dermot; McPherson, Ruth; Neale, Benjamin; Palotie, Aarno; Purcell, Shaun M; Saleheen, Danish; Scharf, Jeremiah; Sklar, Pamela; Sullivan, Patrick F; Tuomilehto, Jaakko; Tsuang, Ming T; Watkins, Hugh C; Wilson, James G; Daly, Mark; MacArthur, DanielSummary Large-scale reference data sets of human genetic variation are critical for the medical and functional interpretation of DNA sequence changes. We describe the aggregation and analysis of high-quality exome (protein-coding region) sequence data for 60,706 individuals of diverse ethnicities generated as part of the Exome Aggregation Consortium (ExAC). This catalogue of human genetic diversity contains an average of one variant every eight bases of the exome, and provides direct evidence for the presence of widespread mutational recurrence. We have used this catalogue to calculate objective metrics of pathogenicity for sequence variants, and to identify genes subject to strong selection against various classes of mutation; identifying 3,230 genes with near-complete depletion of truncating variants with 72% having no currently established human disease phenotype. Finally, we demonstrate that these data can be used for the efficient filtering of candidate disease-causing variants, and for the discovery of human “knockout” variants in protein-coding genes.
Publication A framework for the interpretation of de novo mutation in human disease
(2014) Samocha, Kaitlin E.; Robinson, Elise; Sanders, Stephan J.; Stevens, Christine; Sabo, Aniko; McGrath, Lauren M.; Kosmicki, Jack; Rehnström, Karola; Mallick, Swapan; Kirby, Andrew; Wall, Dennis P.; MacArthur, Daniel; Gabriel, Stacey B.; dePristo, Mark; Purcell, Shaun M.; Palotie, Aarno; Boerwinkle, Eric; Buxbaum, Joseph D.; Cook, Edwin H.; Gibbs, Richard A.; Schellenberg, Gerard D.; Sutcliffe, James S.; Devlin, Bernie; Roeder, Kathryn; Neale, Benjamin; Daly, MarkSpontaneously arising (‘de novo’) mutations play an important role in medical genetics. For diseases with extensive locus heterogeneity – such as autism spectrum disorders (ASDs) – the signal from de novo mutations (DNMs) is distributed across many genes, making it difficult to distinguish disease-relevant mutations from background variation. We provide a statistical framework for the analysis of DNM excesses per gene and gene set by calibrating a model of de novo mutation. We applied this framework to DNMs collected from 1,078 ASD trios and – while affirming a significant role for loss-of-function (LoF) mutations – found no excess of de novo LoF mutations in cases with IQ above 100, suggesting that the role of DNMs in ASD may reside in fundamental neurodevelopmental processes. We also used our model to identify ~1,000 genes that are significantly lacking functional coding variation in non-ASD samples and are enriched for de novo LoF mutations identified in ASD cases.
Publication Genetic risk for autism spectrum disorders and neuropsychiatric variation in the general population
(2016) Robinson, Elise; St. Pourcain, Beate; Anttila, Verneri; Kosmicki, Jack; Bulik-Sullivan, Brendan; Grove, Jakob; Maller, Julian; Samocha, Kaitlin E.; Sanders, Stephan J.; Ripke, Stephan; Martin, Joanna; Hollegaard, Mads V.; Werge, Thomas; Hougaard, David M.; Neale, Benjamin; Evans, David M.; Skuse, David; Mortensen, Preben Bo; Børglum, Anders D.; Ronald, Angelica; Smith, George Davey; Daly, MarkAlmost all genetic risk factors for autism spectrum disorders (ASDs) can be found in the general population, but the effects of that risk are unclear in people not ascertained for neuropsychiatric symptoms. Using several large ASD consortia and population based resources (total n>38,000), we find genomewide genetic links between ASDs and typical variation in social behavior and adaptive functioning. This finding is evidenced through both LD score correlation and de novo variant analysis, indicating that multiple types of genetic risk for ASDs influence a continuum of behavioral and developmental traits, the severe tail of which can result in an ASD or other neuropsychiatric disorder diagnosis. A continuum model should inform the design and interpretation of studies of neuropsychiatric disease biology.
Publication Synaptic, transcriptional, and chromatin genes disrupted in autism
(2014) De Rubeis, Silvia; He, Xin; Goldberg, Arthur P.; Poultney, Christopher S.; Samocha, Kaitlin E.; Cicek, A Ercument; Kou, Yan; Liu, Li; Fromer, Menachem; Walker, Susan; Singh, Tarjinder; Klei, Lambertus; Kosmicki, Jack; Fu, Shih-Chen; Aleksic, Branko; Biscaldi, Monica; Bolton, Patrick F.; Brownfeld, Jessica M.; Cai, Jinlu; Campbell, Nicholas J.; Carracedo, Angel; Chahrour, Maria H.; Chiocchetti, Andreas G.; Coon, Hilary; Crawford, Emily L.; Crooks, Lucy; Curran, Sarah R.; Dawson, Geraldine; Duketis, Eftichia; Fernandez, Bridget A.; Gallagher, Louise; Geller, Evan; Guter, Stephen J.; Hill, R. Sean; Ionita-Laza, Iuliana; Gonzalez, Patricia Jimenez; Kilpinen, Helena; Klauck, Sabine M.; Kolevzon, Alexander; Lee, Irene; Lei, Jing; Lehtimäki, Terho; Lin, Chiao-Feng; Ma'ayan, Avi; Marshall, Christian R.; McInnes, Alison L.; Neale, Benjamin; Owen, Michael J.; Ozaki, Norio; Parellada, Mara; Parr, Jeremy R.; Purcell, Shaun; Puura, Kaija; Rajagopalan, Deepthi; Rehnström, Karola; Reichenberg, Abraham; Sabo, Aniko; Sachse, Michael; Sanders, Stephan J.; Schafer, Chad; Schulte-Rüther, Martin; Skuse, David; Stevens, Christine; Szatmari, Peter; Tammimies, Kristiina; Valladares, Otto; Voran, Annette; Wang, Li-San; Weiss, Lauren A.; Willsey, A. Jeremy; Yu, Timothy W.; Yuen, Ryan K.C.; Cook, Edwin H.; Freitag, Christine M.; Gill, Michael; Hultman, Christina M.; Lehner, Thomas; Palotie, Aarno; Schellenberg, Gerard D.; Sklar, Pamela; State, Matthew W.; Sutcliffe, James S.; Walsh, Christopher; Scherer, Stephen W.; Zwick, Michael E.; Barrett, Jeffrey C.; Cutler, David J.; Roeder, Kathryn; Devlin, Bernie; Daly, Mark; Buxbaum, Joseph D.Summary The genetic architecture of autism spectrum disorder involves the interplay of common and rare variation and their impact on hundreds of genes. Using exome sequencing, analysis of rare coding variation in 3,871 autism cases and 9,937 ancestry-matched or parental controls implicates 22 autosomal genes at a false discovery rate (FDR) < 0.05, and a set of 107 autosomal genes strongly enriched for those likely to affect risk (FDR < 0.30). These 107 genes, which show unusual evolutionary constraint against mutations, incur de novo loss-of-function mutations in over 5% of autistic subjects. Many of the genes implicated encode proteins for synaptic, transcriptional, and chromatin remodeling pathways. These include voltage-gated ion channels regulating propagation of action potentials, pacemaking, and excitability-transcription coupling, as well as histone-modifying enzymes and chromatin remodelers, prominently histone post-translational modifications involving lysine methylation/demethylation.
Publication Refining the role of de novo protein truncating variants in neurodevelopmental disorders using population reference samples
(2017) Kosmicki, Jack; Samocha, Kaitlin E.; Howrigan, Daniel; Sanders, Stephan J.; Slowikowski, Kamil; Lek, Monkol; Karczewski, Konrad; Cutler, David J.; Devlin, Bernie; Roeder, Kathryn; Buxbaum, Joseph D.; Neale, Benjamin; MacArthur, Daniel; Wall, Dennis P.; Robinson, Elise; Daly, MarkRecent research has uncovered a significant role for de novo variation in neurodevelopmental disorders. Using aggregated data from 9246 families with autism spectrum disorder, intellectual disability, or developmental delay, we show ~1/3 of de novo variants are independently observed as standing variation in the Exome Aggregation Consortium’s cohort of 60,706 adults, and these de novo variants do not contribute to neurodevelopmental risk. We further use a loss-of-function (LoF)-intolerance metric, pLI, to identify a subset of LoF-intolerant genes that contain the observed signal of associated de novo protein truncating variants (PTVs) in neurodevelopmental disorders. LoF-intolerant genes also carry a modest excess of inherited PTVs; though the strongest de novo impacted genes contribute little to this, suggesting the excess of inherited risk resides lower-penetrant genes. These findings illustrate the importance of population-based reference cohorts for the interpretation of candidate pathogenic variants, even for analyses of complex diseases and de novo variation.
Publication Exome Sequencing in Schizophrenia-Affected Parent–offspring Trios Reveals Risk Conferred by Protein-Coding De Novo Mutations
(Springer Science and Business Media LLC, 2020-01-13) Howrigan, Daniel; Rose, Samuel A.; Samocha, Kaitlin E.; Fromer, Menachem; Cerrato, Felecia; Chen, Wei J.; Churchhouse, Claire; Chambert, Kimberly; Chandler, Sharon D.; Daly, Mark; Dumont, Ashley; Genovese, Giulio; Hwu, Hai-Gwo; Laird, Nan; Kosmicki, Jack; Moran, Jennifer L.; Singh, Tarjinder; McCarroll, Steven; Faraone, Stephen V.; Glatt, Stephen J.; Tsuang, Ming; Neale, BenjaminProtein-coding de novo mutations (DNMs) are significant risk factors in many neurodevelopmental disorders, whereas association with schizophrenia (SCZ) risk thus far has been modest. We analyze whole-exome sequence from 1,695 SCZ affected trios along with DNMs from 1,077 published SCZ trios to better understand their contribution to SCZ risk. Among 2,772 SCZ probands, exome-wide DNM burden remains modest. Gene set analyses reveal that SCZ DNMs are significantly concentrated in genes either highly brain expressed, under strong evolutionary constraint, and/or overlap with genes identified in other neurodevelopmental disorders. No single gene surpasses exome-wide significance, however sixteen genes are recurrently hit by protein-truncating DNMs, a 3.15-fold higher rate than the mutation model expectation (permuted 95% CI=1-10 genes, permuted p=3e-5). Overall, DNMs explain only a small fraction of SCZ risk, and larger samples are needed to identify individual risk genes, as coding variation across many genes confer risk for SCZ in the population.