Publication: The new (educational) statistics: Properties of scales that matter
Open/View Files
Date
Authors
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
David Thissen’s essay, Bad Questions: An Essay Involving Item Response Theory (2016), is an excellent contribution to the genre of commentaries on the field. It joins the likes of the piece by Thissen’s frequent collaborator, Howard Wainer (2010), who published 14 conversations about three things in this journal 6 years ago. Thissen asks and answers, dismissively, five of his titular “bad questions.” He concludes that what makes them bad “is the framing of the question that demands a yes-or-no, black and white, cut and dried response” (p. 10). He argues for a statistical education that values continua over dichotomies and categories, and I agree. However, I think Thissen, as well as scholars and students of statistics and measurement generally, underappreciate the utility and necessity of dichotomies by decision makers. Ultimately, I believe we can and should inform their dichotomies with meaningful scales and defensible procedures more often than we do. Conveniently, Item Response Theory (IRT) can be quite useful for this purpose. In this brief response, I distinguish between Thissen’s (2016) first three bad questions and his last two. His first three questions concern statistical and psychometric criteria, for IRT model fit, unidimensionality, and cardinality (interval scale properties), respectively. These are judgments about models and scores. His last two questions concern policy criteria, for student proficiency and teacher effectiveness, respectively. These are judgments about people. Thissen’s (2016) first three questions are bad in the way that statistical p values are bad. These questions yield incomplete information about practical significance when samples are small, and they are a needless distraction when samples are large. I agree with Thissen that we should shift our attention away from these questions and toward questions of practical significance. In contrast, I argue that the last two questions are bad in part because statisticians and psychometricians have done too little to help answer them. I suggest how we might help, by employing scale anchoring methods and investigating properties of “Frankenstein” score composites that are in common use.