Person: Ho, Andrew
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Final report on the evaluation of the Growth Model Pilot Project
(2011) Hoffer, Thomas B.; Hedberg, E. C.; Brown, Kevin L.; Halverson, Marie L.; Reid-Brossard, Paki; Ho, Andrew; Furgol, KatherinePublication Estimating Achievement Gaps from Test Scores Reported in Ordinal "Proficiency" Categories
(2012) Ho, Andrew; Reardon, Sean F.Test scores are commonly reported in a small number of ordered categories. Examples of such reporting include state accountability testing, Advanced Placement tests, and English proficiency tests. This paper introduces and evaluates methods for estimating achievement gaps on a familiar standard-deviation-unit metric using data from these ordered categories alone. These methods hold two practical advantages over alternative achievement gap metrics. First, they require only categorical proficiency data, which are often available where means and standard deviations are not. Second, they result in gap estimates that are invariant to score scale transformations, providing a stronger basis for achievement gap comparisons over time and across jurisdictions. We find three candidate estimation methods that recover full-distribution gap estimates well when only censored data are available.
Publication The Epidemiology of Modern Test Score Use: Anticipating Aggregation, Adjustment, and Equating
(Informa UK Limited, 2013) Ho, AndrewPublication Publication The Dependence of Growth-Model Results on Proficiency Cut Scores
(Wiley-Blackwell, 2009) Ho, Andrew; Lewis, Daniel M.; Farris, Jason L. MacGregorStates participating in the Growth Model Pilot Program reference individual student growth against “proficiency” cut scores that conform with the original No Child Left Behind Act (NCLB). Although achievement results from conventional NCLB models are also cut-score dependent, the functional relationships between cut-score location and growth results are more complex and are not currently well described. We apply cut-score scenarios to longitudinal data to demonstrate the dependence of state- and school-level growth results on cut-score choice. This dependence is examined along three dimensions: 1) rigor, as states set cut scores largely at their discretion, 2) across-grade articulation, as the rigor of proficiency standards may vary across grades, and 3) the time horizon chosen for growth to proficiency. Results show that the selection of plausible alternative cut scores within a growth model can change the percentage of students “on track to proficiency” by more than 20 percentage points and reverse accountability decisions for more than 40% of schools. We contribute a framework for predicting these dependencies, and we argue that the cut-score dependence of large-scale growth statistics must be made transparent, particularly for comparisons of growth results across states.
Publication Response switching and self-efficacy in Peer Instruction classrooms
(American Physical Society (APS), 2015) Miller, Kelly; Schell, Julie; Ho, Andrew; Lukoff, Brian; Mazur, EricPeer Instruction, a well-known student-centered teaching method, engages students during class through structured, frequent questioning and is often facilitated by classroom response systems. The central feature of any Peer Instruction class is a conceptual question designed to help resolve student misconceptions about subject matter. We provide students two opportunities to answer each question—once after a round of individual reflection and then again after a discussion round with a peer. The second round provides students the choice to “switch” their original response to a different answer. The percentage of right answers typically increases after peer discussion: most students who answer incorrectly in the individual round switch to the correct answer after the peer discussion. However, for any given question there are also students who switch their initially right answer to a wrong answer and students who switch their initially wrong answer to a different wrong answer. In this study, we analyze response switching over one semester of an introductory electricity and magnetism course taught using Peer Instruction at Harvard University. Two key features emerge from our analysis: First, response switching correlates with academic selfefficacy. Students with low self-efficacy switch their responses more than students with high self-efficacy. Second, switching also correlates with the difficulty of the question; students switch to incorrect responses more often when the question is difficult. These findings indicate that instructors may need to provide greater support for difficult questions, such as supplying cues during lectures, increasing times for discussions, or ensuring effective pairing (such as having a student with one right answer in the pair). Additionally, the connection between response switching and self-efficacy motivates interventions to increase student self-efficacy at the beginning of the semester by helping students develop early mastery or to reduce stressful experiences (i.e., high-stakes testing) early in the semester, in the hope that this will improve student learning in Peer Instruction classrooms.
Publication Publication Discreteness Causes Bias in Percentage-Based Comparisons: A Case Study From Educational Testing
(Informa UK Limited, 2015) Yee, Darrick Shen-Wei; Ho, AndrewDiscretizing continuous distributions can lead to bias in parameter estimates. We present a case study from educational testing that illustrates dramatic consequences of discreteness when discretizing partitions differ across distributions. The percentage of test-takers who score above a certain cutoff score (percent above cutoff, or “PAC”) often describes overall performance on a test. Year-over-year changes in PAC, or ΔPAC, have gained prominence under recent U.S. education policies, with public schools facing sanctions if they fail to meet PAC targets. In this paper, we describe how test score distributions act as continuous distributions that are discretized inconsistently over time. We show that this can propagate considerable bias to PAC trends, where positive ΔPACs appear negative, and vice versa, for a substantial number of actual tests. A simple model shows that this bias applies to any comparison of PAC statistics in which values for one distribution are discretized differently from values for the other.
Publication Validation Methods for Aggregate-Level Test Scale Linking: A Rejoinder
(American Educational Research Association (AERA), 2021-03-15) Ho, Andrew; Reardon, Sean F.; Kalogrides, DemetraIn Reardon, Kalogrides, and Ho (2021), we developed precision-adjusted random effects models to estimate aggregate-level linking error, for populations and subpopulations, for averages and progress over time. We are grateful to past editor Dan McCaffrey for selecting our paper as the focal article for a set of commentaries from our colleagues, Daniel Bolt, Mark Davison, Alina von Davier, Tim Moses, and Neil Dorans. These commentaries reinforce important cautions and identify promising directions for future research. In this rejoinder, we clarify aspects of our originally proposed method. 1) Validation methods provide evidence of benefits and risks that different experts may weigh differently for different purposes. 2) Our proposed method differs from “standard mapping” procedures using the National Assessment of Educational Progress not only by using a linear (vs. equipercentile) link but also by targeting direct validity evidence about counterfactual aggregate scores. 3) Multilevel approaches that assume common score scales across states are indeed a promising next step for validation, and we hope that states enable researchers to use more of their common-core-era consortium test data for this purpose. Finally, we apply our linking method to an extended panel of data from 2009 to 2017 to show that linking recovery has remained stable.
Publication Specifying the Three Ws in Educational Measurement: Who Uses Which Scores for What Purpose?
(Wiley, 2022-12) Ho, AndrewI argue that understanding and improving educational measurement requires specificity about actors, scores, and purpose: Who uses which scores for what purpose? I show how this specificity complements Briggs’ frameworks for educational measurement that he presented in his 2022 address as president of the National Council on Measurement in Education.