Publication:

The new (educational) statistics: Properties of scales that matter

Loading...
Thumbnail Image

Date

2016

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Ho, A. D. (2016). The new (educational) statistics: Properties of scales that matter. Journal of Educational and Behavioral Statistics, 41, 94-99. https://doi.org/10.3102/1076998615621302

Abstract

David Thissen’s essay, Bad Questions: An Essay Involving Item Response Theory (2016), is an excellent contribution to the genre of commentaries on the field. It joins the likes of the piece by Thissen’s frequent collaborator, Howard Wainer (2010), who published 14 conversations about three things in this journal 6 years ago. Thissen asks and answers, dismissively, five of his titular “bad questions.” He concludes that what makes them bad “is the framing of the question that demands a yes-or-no, black and white, cut and dried response” (p. 10). He argues for a statistical education that values continua over dichotomies and categories, and I agree. However, I think Thissen, as well as scholars and students of statistics and measurement generally, underappreciate the utility and necessity of dichotomies by decision makers. Ultimately, I believe we can and should inform their dichotomies with meaningful scales and defensible procedures more often than we do. Conveniently, Item Response Theory (IRT) can be quite useful for this purpose. In this brief response, I distinguish between Thissen’s (2016) first three bad questions and his last two. His first three questions concern statistical and psychometric criteria, for IRT model fit, unidimensionality, and cardinality (interval scale properties), respectively. These are judgments about models and scores. His last two questions concern policy criteria, for student proficiency and teacher effectiveness, respectively. These are judgments about people. Thissen’s (2016) first three questions are bad in the way that statistical p values are bad. These questions yield incomplete information about practical significance when samples are small, and they are a needless distraction when samples are large. I agree with Thissen that we should shift our attention away from these questions and toward questions of practical significance. In contrast, I argue that the last two questions are bad in part because statisticians and psychometricians have done too little to help answer them. I suggest how we might help, by employing scale anchoring methods and investigating properties of “Frankenstein” score composites that are in common use.

Description

Other Available Sources

Research Data

Keywords

Endorsement

Review

Supplemented By

Related Stories