Person: Reich, David
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication The Impact of Divergence Time on the Nature of Population Structure: An Example from Iceland
(Public Library of Science, 2009) Helgason, Agnar; Palsson, Snaebjorn; Stefansson, Hreinn; St. Clair, David; Andreassen, Ole A.; Kong, Augustine; Stefansson, Kari; Price, Alkes; Reich, DavidThe Icelandic population has been sampled in many disease association studies, providing a strong motivation to understand the structure of this population and its ramifications for disease gene mapping. Previous work using 40 microsatellites showed that the Icelandic population is relatively homogeneous, but exhibits subtle population structure that can bias disease association statistics. Here, we show that regional geographic ancestries of individuals from Iceland can be distinguished using 292,289 autosomal single-nucleotide polymorphisms (SNPs). We further show that subpopulation differences are due to genetic drift since the settlement of Iceland 1100 years ago, and not to varying contributions from different ancestral populations. A consequence of the recent origin of Icelandic population structure is that allele frequency differences follow a null distribution devoid of outliers, so that the risk of false positive associations due to stratification is minimal. Our results highlight an important distinction between population differences attributable to recent drift and those arising from more ancient divergence, which has implications both for association studies and for efforts to detect natural selection using population differentiation.
Publication Sensitive Detection of Chromosomal Segments of Distinct Ancestry in Admixed Populations
(Public Library of Science, 2009) Price, Alkes; Tandon, Arti; Patterson, Nick; Barnes, Kathleen C.; Rafaels, Nicholas; Ruczinski, Ingo; Beaty, Terri H.; Mathias, Rasika; Reich, David; Myers, SimonIdentifying the ancestry of chromosomal segments of distinct ancestry has a wide range of applications from disease mapping to learning about history. Most methods require the use of unlinked markers; but, using all markers from genome-wide scanning arrays, it should in principle be possible to infer the ancestry of even very small segments with exquisite accuracy. We describe a method, HAPMIX, which employs an explicit population genetic model to perform such local ancestry inference based on fine-scale variation data. We show that HAPMIX outperforms other methods, and we explore its utility for inferring ancestry, learning about ancestral populations, and inferring dates of admixture. We validate the method empirically by applying it to populations that have experienced recent and ancient admixture: 935 African Americans from the United States and 29 Mozabites from North Africa. HAPMIX will be of particular utility for mapping disease genes in recently admixed populations, as its accurate estimates of local ancestry permit admixture and case-control association signals to be combined, enabling more powerful tests of association than with either signal alone.
Publication Concept, Design and Implementation of a Cardiovascular Gene-centric 50 K SNP Array for Large-scale Genomic Association Studies
(Public Library of Science, 2008) Keating, Brendan J.; Tischfield, Sam; Murray, Sarah S.; Bhangale, Tushar; Price, Thomas S.; Glessner, Joseph T.; Galver, Luana; Barrett, Jeffrey C.; Grant, Struan F. A.; Farlow, Deborah N.; Chandrupatla, Hareesh R.; Ajmal, Saad; Papanicolaou, George J.; Guo, Yiran; Li, Mingyao; DerOhannessian, Stephanie; Bailey, Swneke D.; Montpetit, Alexandre; Edmondson, Andrew C.; Taylor, Kent; Gai, Xiaowu; Wang, Susanna S.; Fornage, Myriam; Shaikh, Tamim; Groop, Leif; Boehnke, Michael; Hall, Alistair S.; Hattersley, Andrew T.; Frackelton, Edward; Patterson, Nick; Chiang, Charleston W. K.; Kim, Cecelia E.; Fabsitz, Richard R.; Ouwehand, Willem; Munroe, Patricia; Caulfield, Mark; Drake, Thomas; Boerwinkle, Eric; Whitehead, A. Stephen; Cappola, Thomas P.; Samani, Nilesh J.; Lusis, A. Jake; Schadt, Eric; Wilson, James G.; Koenig, Wolfgang; McCarthy, Mark I.; Kathiresan, Sekar; Gabriel, Stacey B.; Hakonarson, Hakon; Anand, Sonia S.; Reilly, Muredach; Engert, James C.; Nickerson, Deborah A.; Rader, Daniel J.; FitzGerald, Garret A.; Reitsma, Pieter H.; Hansen, Mark; de Bakker, Paul; Price, Alkes; Reich, David; Hirschhorn, JoelA wealth of genetic associations for cardiovascular and metabolic phenotypes in humans has been accumulating over the last decade, in particular a large number of loci derived from recent genome wide association studies (GWAS). True complex disease-associated loci often exert modest effects, so their delineation currently requires integration of diverse phenotypic data from large studies to ensure robust meta-analyses. We have designed a gene-centric 50 K single nucleotide polymorphism (SNP) array to assess potentially relevant loci across a range of cardiovascular, metabolic and inflammatory syndromes. The array utilizes a “cosmopolitan” tagging approach to capture the genetic diversity across ∼2,000 loci in populations represented in the HapMap and SeattleSNPs projects. The array content is informed by GWAS of vascular and inflammatory disease, expression quantitative trait loci implicated in atherosclerosis, pathway based approaches and comprehensive literature searching. The custom flexibility of the array platform facilitated interrogation of loci at differing stringencies, according to a gene prioritization strategy that allows saturation of high priority loci with a greater density of markers than the existing GWAS tools, particularly in African HapMap samples. We also demonstrate that the IBC array can be used to complement GWAS, increasing coverage in high priority CVD-related loci across all major HapMap populations. DNA from over 200,000 extensively phenotyped individuals will be genotyped with this array with a significant portion of the generated data being released into the academic domain facilitating in silico replication attempts, analyses of rare variants and cross-cohort meta-analyses in diverse populations. These datasets will also facilitate more robust secondary analyses, such as explorations with alternative genetic models, epistasis and gene-environment interactions.
Publication Population Structure and Eigenanalysis
(Public Library of Science, 2006) Patterson, Nick; Price, Alkes; Reich, DavidCurrent methods for inferring population structure from genetic data do not provide formal significance tests for population differentiation. We discuss an approach to studying population structure (principal components analysis) that was first applied to genetic data by Cavalli-Sforza and colleagues. We place the method on a solid statistical footing, using results from modern statistics to develop formal significance tests. We also uncover a general “phase change” phenomenon about the ability to detect structure in genetic data, which emerges from the statistical theory we use, and has an important implication for the ability to discover structure in genetic data: for a fixed but large dataset size, divergence between two populations (as measured, for example, by a statistic like (F_{ST})) below a threshold is essentially undetectable, but a little above threshold, detection will be easy. This means that we can predict the dataset size needed to detect structure.
Publication Amerind Ancestry, Socioeconomic Status and the Genetics of Type 2 Diabetes in a Colombian Population
(Public Library of Science, 2012) Campbell, Desmond D.; Parra, Maria V.; Duque, Constanza; Gallego, Natalia; Franco, Liliana; Hünemeier, Tábita; Bortolini, Cátira; Villegas, Alberto; Bedoya, Gabriel; McCarthy, Mark I.; Ruiz-Linares, Andrés; Tandon, Arti; Price, Alkes; Reich, DavidThe “thrifty genotype” hypothesis proposes that the high prevalence of type 2 diabetes (T2D) in Native Americans and admixed Latin Americans has a genetic basis and reflects an evolutionary adaptation to a past low calorie/high exercise lifestyle. However, identification of the gene variants underpinning this hypothesis remains elusive. Here we assessed the role of Native American ancestry, socioeconomic status (SES) and 21 candidate gene loci in susceptibility to T2D in a sample of 876 T2D cases and 399 controls from Antioquia (Colombia). Although mean Native American ancestry is significantly higher in T2D cases than in controls (32% v 29%), this difference is confounded by the correlation of ancestry with SES, which is a stronger predictor of disease status. Nominally significant association (P<0.05) was observed for markers in: TCF7L2, RBMS1, CDKAL1, ZNF239, KCNQ1 and TCF1 and a significant bias (P<0.05) towards OR>1 was observed for markers selected from previous T2D genome-wide association studies, consistent with a role for Old World variants in susceptibility to T2D in Latin Americans. No association was found to the only known Native American-specific gene variant previously associated with T2D in a Mexican sample (rs9282541 in ABCA1). An admixture mapping scan with 1,536 ancestry informative markers (AIMs) did not identify genome regions with significant deviation of ancestry in Antioquia. Exclusion analysis indicates that this scan rules out ∼95% of the genome as harboring loci with ancestry risk ratios >1.22 (at P < 0.05).
Publication Inferring Admixture Histories of Human Populations Using Linkage Disequilibrium
(Genetics Society of America, 2013) Loh, Po-Ru; Lipson, Mark; Patterson, Nick; Moorjani, Priya; Pickrell, Joseph; Reich, David; Berger, BonnieLong-range migrations and the resulting admixtures between populations have been important forces shaping human genetic diversity. Most existing methods for detecting and reconstructing historical admixture events are based on allele frequency divergences or patterns of ancestry segments in chromosomes of admixed individuals. An emerging new approach harnesses the exponential decay of admixture-induced linkage disequilibrium (LD) as a function of genetic distance. Here, we comprehensively develop LD-based inference into a versatile tool for investigating admixture. We present a new weighted LD statistic that can be used to infer mixture proportions as well as dates with fewer constraints on reference populations than previous methods. We define an LD-based three-population test for admixture and identify scenarios in which it can detect admixture events that previous formal tests cannot. We further show that we can uncover phylogenetic relationships among populations by comparing weighted LD curves obtained using a suite of references. Finally, we describe several improvements to the computation and fitting of weighted LD curves that greatly increase the robustness and speed of the calculations. We implement all of these advances in a software package, ALDER, which we validate in simulations and apply to test for admixture among all populations from the Human Genome Diversity Project (HGDP), highlighting insights into the admixture history of Central African Pygmies, Sardinians, and Japanese.
Publication A direct characterization of human mutation based on microsatellites
(2012) Sun, James Xin; Helgason, Agnar; Masson, Gisli; Ebenesersdóttir, Sigríđur Sunna; Li, Heng; Mallick, Swapan; Gnerre, Sante; Patterson, Nick; Kong, Augustine; Reich, David; Stefansson, KariMutations are the raw material of evolution, but have been difficult to study directly. We report the largest study of new mutations to date: 2,058 germline changes discovered by analyzing 85,289 Icelanders at 2,477 microsatellites. The paternal-to-maternal mutation rate ratio is 3.3, and the rate in fathers doubles from age 20 to 58 whereas there is no association with age in mothers. Longer microsatellite alleles are more mutagenic and tend to decrease in length, whereas the opposite is seen for shorter alleles. We use these empirical observations to build a model that we apply to individuals for whom we have both genome sequence and microsatellite data, allowing us to estimate key parameters of evolution without calibration to the fossil record. We infer that the sequence mutation rate is 1.4–2.3×10−8 per base pair per generation (90% credible interval), and that human-chimpanzee speciation occurred 3.7–6.6 million years ago.
Publication The genetic prehistory of southern Africa
(Nature Pub. Group, 2012) Pickrell, Joseph; Patterson, Nick; Barbieri, Chiara; Berthold, Falko; Gerlach, Linda; Güldemann, Tom; Kure, Blesswell; Mpoloka, Sununguko Wata; Nakagawa, Hirosi; Naumann, Christfried; Lipson, Mark; Loh, Po-Ru; Lachance, Joseph; Mountain, Joanna; Bustamante, Carlos D.; Berger, Bonnie; Tishkoff, Sarah A.; Henn, Brenna M.; Stoneking, Mark; Reich, David; Pakendorf, BrigitteSouthern and eastern African populations that speak non-Bantu languages with click consonants are known to harbour some of the most ancient genetic lineages in humans, but their relationships are poorly understood. Here, we report data from 23 populations analysed at over half a million single-nucleotide polymorphisms, using a genome-wide array designed for studying human history. The southern African Khoisan fall into two genetic groups, loosely corresponding to the northwestern and southeastern Kalahari, which we show separated within the last 30,000 years. We find that all individuals derive at least a few percent of their genomes from admixture with non-Khoisan populations that began ∼1,200 years ago. In addition, the East African Hadza and Sandawe derive a fraction of their ancestry from admixture with a population related to the Khoisan, supporting the hypothesis of an ancient link between southern and eastern Africa.
Publication Reconstructing Native American Population History
(2013) Reich, David; Patterson, Nick; Campbell, Desmond; Tandon, Arti; Mazieres, Stéphane; Ray, Nicolas; Parra, Maria V.; Rojas, Winston; Duque, Constanza; Mesa, Natalia; García, Luis F.; Triana, Omar; Blair, Silvia; Maestre, Amanda; Dib, Juan C.; Bravi, Claudio M.; Bailliet, Graciela; Corach, Daniel; Hünemeier, Tábita; Bortolini, Maria-Cátira; Salzano, Francisco M.; Petzl-Erler, María Luiza; Acuña-Alonzo, Victor; Aguilar-Salinas, Carlos; Canizales-Quinteros, Samuel; Tusié-Luna, Teresa; Riba, Laura; Rodríguez-Cruz, Maricela; Lopez-Alarcón, Mardia; Coral-Vazquez, Ramón; Canto-Cetina, Thelma; Silva-Zolezzi, Irma; Fernandez-Lopez, Juan Carlos; Contreras, Alejandra V.; Jimenez-Sanchez, Gerardo; Gómez-Vázquez, María José; Molina, Julio; Carracedo, Ángel; Salas, Antonio; Gallo, Carla; Poletti, Giovanni; Witonsky, David B.; Alkorta-Aranburu, Gorka; Sukernik, Rem I.; Osipova, Ludmila; Fedorova, Sardana; Vasquez, René; Villena, Mercedes; Moreau, Claudia; Barrantes, Ramiro; Pauls, David; Excoffier, Laurent; Bedoya, Gabriel; Rothhammer, Francisco; Dugoujon, Jean Michel; Larrouy, Georges; Klitz, William; Labuda, Damian; Kidd, Judith; Kidd, Kenneth; Rienzo, Anna Di; Freimer, Nelson B.; Price, Alkes; Ruiz-Linares, AndrésThe peopling of the Americas has been the subject of extensive genetic, archaeological and linguistic research; however, central questions remain unresolved1–5. One contentious issue is whether the settlement occurred via a single6–8 or multiple streams of migration from Siberia9–15. The pattern of dispersals within the Americas is also poorly understood. To address these questions at higher resolution than was previously possible, we assembled data from 52 Native American and 17 Siberian groups genotyped at 364,470 single nucleotide polymorphisms. We show that Native Americans descend from at least three streams of Asian gene flow. Most descend entirely from a single ancestral population that we call “First American”. However, speakers of Eskimo-Aleut languages from the Arctic inherit almost half their ancestry from a second stream of Asian gene flow, and the Na-Dene-speaking Chipewyan from Canada inherit roughly one-tenth of their ancestry from a third stream. We show that the initial peopling followed a southward expansion facilitated by the coast, with sequential population splits and little gene flow after divergence, especially in South America. A major exception is in Chibchan-speakers on both sides of the Panama Isthmus, who have ancestry from both North and South America.
Publication Effects of cis and trans Genetic Ancestry on Gene Expression in African Americans
(Public Library of Science, 2008) Price, Alkes; Patterson, Nick; Hancks, Dustin C.; Myers, Simon; Reich, David; Cheung, Vivian G.; Spielman, Richard S.Variation in gene expression is a fundamental aspect of human phenotypic variation. Several recent studies have analyzed gene expression levels in populations of different continental ancestry and reported population differences at a large number of genes. However, these differences could largely be due to non-genetic (e.g., environmental) effects. Here, we analyze gene expression levels in African American cell lines, which differ from previously analyzed cell lines in that individuals from this population inherit variable proportions of two continental ancestries. We first relate gene expression levels in individual African Americans to their genome-wide proportion of European ancestry. The results provide strong evidence of a genetic contribution to expression differences between European and African populations, validating previous findings. Second, we infer local ancestry (0, 1, or 2 European chromosomes) at each location in the genome and investigate the effects of ancestry proximal to the expressed gene (cis) versus ancestry elsewhere in the genome (trans). Both effects are highly significant, and we estimate that 12±3% of all heritable variation in human gene expression is due to cis variants.