Person:

King, Gary

Loading...
Profile Picture

Email Address

AA Acceptance Date

Birth Date

Research Projects

Organizational Units

Job Title

Last Name

King

First Name

Gary

Name

King, Gary

Search Results

Now showing 1 - 10 of 32
  • Publication

    A Method of Automated Nonparametric Content Analysis for Social Science

    (Wiley-Blackwell, 2010) Hopkins, Daniel J.; King, Gary

    The increasing availability of digitized text presents enormous opportunities for social scientists. Yet hand coding many blogs, speeches, government records, newspapers, or other sources of unstructured text is infeasible. Although computer scientists have methods for automated content analysis, most are optimized to classify individual documents, whereas social scientists instead want generalizations about the population of documents, such as the proportion in a given category. Unfortunately, even a method with a high percent of individual documents correctly classified can be hugely biased when estimating category proportions. By directly optimizing for this social science goal, we develop a method that gives approximately unbiased estimates of category proportions even when the optimal classifier performs poorly. We illustrate with diverse data sets, including the daily expressed opinions of thousands of people about the U.S. presidency. We also make available software that implements our methods and large corpora of text for further analysis.

  • Publication

    Avoiding Randomization Failure in Program Evaluation, with Application to the Medicare Health Support Program

    (Mary Ann Liebert, 2011) King, Gary; Nielsen, Richard; Coberly, Carter; Pope, James E.; Wells, Aaron

    We highlight common problems in the application of random treatment assignment in large-scale program evaluation. Random assignment is the defining feature of modern experimental design, yet errors in design, implementation, and analysis often result in real-world applications not benefiting from its advantages. The errors discussed here cover the control of variability, levels of randomization, size of treatment arms, and power to detect causal effects, as well as the many problems that commonly lead to post-treatment bias. We illustrate these issues by identifying numerous serious errors in the Medicare Health Support evaluation and offering recommendations to improve the design and analysis of this and other large-scale randomized experiments.

  • Publication

    Estimating Partisan Bias of the Electoral College Under Proposed Changes in Elector Apportionment

    (Walter de Gruyter, 2013) Thomas, Ayende; Gelman, Andrew; King, Gary; Katz, Jonathan N.

    In the election for President of the United States, the Electoral College is the body whose members vote to elect the President directly. Each state sends a number of delegates equal to its total number of representatives and senators in Congress; all but two states (Nebraska and Maine) assign electors pledged to the candidate that wins the state's plurality vote. We investigate the effect on presidential elections if states were to assign their electoral votes according to results in each congressional district,and conclude that the direct popular vote and the current electoral college are both substantially fairer compared to those alternatives where states would have divided their electoral votes by congressional district.

  • Publication

    Google Flu Trends Still Appears Sick: An Evaluation of the 2013-2014 Flu Season

    (Social Science Electronic Publishing, 2014) Lazer, David; Kennedy, Ryan; King, Gary; Vespignani, Alessandro

    In response to its poor performance during the 2012-2013 flu season, Google Flu Trends (GFT) engineers announced a redesign of the GFT algorithm. Two changes were made: (1) dampening anomalous media spikes and (2) using ElasticNet, rather than regression, for estimation. This paper identifies several problems that persist in the new algorithm. First, the transparency problems identified in our earlier Science paper appear to have, if anything, become worse. Second, there are reasons to doubt whether a spike in media attention was the only, or primary, cause of GFT's errors. Finally, there is strong evidence that GFT is still not using all the information at its disposal to make accurate measurements of flu prevalence. While it is too early to give a complete evaluation of the new algorithm, these results are discouraging.

  • Publication

    Statistical Security for Social Security

    (Springer Nature, 2012) Soneji, Samir; King, Gary

    The financial viability of Social Security, the single largest U.S. government program, depends on accurate forecasts of the solvency of its intergenerational trust fund. We begin by detailing information necessary for replicating the Social Security Administration’s (SSA’s) forecasting procedures, which until now has been unavailable in the public domain. We then offer a way to improve the quality of these procedures via age- and sex-specific mortality forecasts. The most recent SSA mortality forecasts were based on the best available technology at the time, which was a combination of linear extrapolation and qualitative judgments. Unfortunately, linear extrapolation excludes known risk factors and is inconsistent with long-standing demographic patterns, such as the smoothness of age profiles. Modern statistical methods typically outperform even the best qualitative judgments in these contexts. We show how to use such methods, enabling researchers to forecast using far more information, such as the known risk factors of smoking and obesity and known demographic patterns. Including this extra information makes a substantial difference. For example, by improving only mortality forecasting methods, we predict three fewer years of net surplus, $730 billion less in Social Security Trust Funds, and program costs that are 0.66% greater for projected taxable payroll by 2031 compared with SSA projections. More important than specific numerical estimates are the advantages of transparency, replicability, reduction of uncertainty, and what may be the resulting lower vulnerability to the politicization of program forecasts. In addition, by offering with this article software and detailed replication information, we hope to marshal the efforts of the research community to include ever more informative inputs and to continue to reduce uncertainties in Social Security forecasts.

  • Publication

    Reverse-engineering censorship in China: Randomized experimentation and participant observation

    (American Association for the Advancement of Science (AAAS), 2014) King, Gary; Pan, Jennifer; Roberts, Margaret E.

    Existing research on the extensive Chinese censorship organization uses observational methods with well-known limitations. We conducted the first large-scale experimental study of censorship by creating accounts on numerous social media sites, randomly submitting different texts, and observing from a worldwide network of computers which texts were censored and which were not. We also supplemented interviews with confidential sources by creating our own social media site, contracting with Chinese firms to install the same censoring technologies as existing sites, and—with their software, documentation, and even customer support—reverse-engineering how it all works. Our results offer rigorous support for the recent hypothesis that criticisms of the state, its leaders, and their policies are published, whereas posts about real-world events with collective action potential are censored.

  • Publication

    Estimating Incidence Curves of Several Infections Using Symptom Surveillance Data

    (Public Library of Science, 2011) Goldstein, Edward; Cowling, Benjamin J.; Aiello, Allison E.; Takahashi, Saki; King, Gary; Lu, Ying; Lipsitch, Marc

    We introduce a method for estimating incidence curves of several co-circulating infectious pathogens, where each infection has its own probabilities of particular symptom profiles. Our deconvolution method utilizes weekly surveillance data on symptoms from a defined population as well as additional data on symptoms from a sample of virologically confirmed infectious episodes. We illustrate this method by numerical simulations and by using data from a survey conducted on the University of Michigan campus. Last, we describe the data needs to make such estimates accurate.

  • Publication

    Ensuring the Data-Rich Future of the Social Sciences

    (American Association for the Advancement of Science (AAAS), 2011) King, Gary

    Massive increases in the availability of informative social science data are making dramatic progress possible in analyzing, understanding, and addressing many major societal problems. Yet the same forces pose severe challenges to the scientific infrastructure supporting data sharing, data management, informatics, statistical methodology, and research ethics and policy, and these are collectively holding back progress. I address these changes and challenges and suggest what can be done.

  • Publication

    MatchIt: Nonparametric Preprocessing for Parametric Causal Inference

    (University of California, Los Angeles, 2011) Stuart, Elizabeth A.; King, Gary; Imai, Kosuke; Ho, Daniel

    MatchIt implements the suggestions of Ho, Imai, King, and Stuart (2007) for improving parametric statistical models by preprocessing data with nonparametric matching methods. MatchIt implements a wide range of sophisticated matching methods, making it possible to greatly reduce the dependence of causal inferences on hard-to-justify, but commonly made, statistical modeling assumptions. The software also easily fits into existing research practices since, after preprocessing data with MatchIt, researchers can use whatever parametric model they would have used without MatchIt, but produce inferences with substantially more robustness and less sensitivity to modeling assumptions. MatchIt is an R program, and also works seamlessly with Zelig.

  • Publication

    The Troubled Future of Colleges and Universities

    (Cambridge University Press, 2013) King, Gary; Sen, Maya

    The American system of higher education appears poised for disruptive change of potentially historic proportions due to massive new political, economic, and educational forces that threaten to undermine its business model, governmental support, and operating mission. These forces include dramatic new types of economic competition, difficulties in growing revenue streams as we had in the past, relative declines in philanthropic and government support, actual and likely future political attacks on universities, and some outdated methods of teaching and learning that have been unchanged for hundreds of years.