GSE Theses and Dissertations
Permanent URI for this collectionhttps://dash.harvard.edu/handle/1/13056148
Browse
Publication Essays on Educational Testing in an Era of Higher ("College-Ready") Standards
(2019-05-08) Thng, Yi Xe; Ho, Andrew D.; Miratrix, Luke W.; Taylor, Eric S.I present three essays on educational testing in an era of "college-ready" standards. My first essay evaluates evidence-based standard setting methods that select passing or “college-ready” cut scores using regression-based predictive relationships between test scores and college outcomes. I investigate the forms of evidence that can be derived using predictive methods, whether such evidence can enhance score interpretations, and if so, how such evidence can be used to inform standard setting. I find that to compensate for the poor predictive utility, cut scores derived from predictive methods may be overly stringent or lenient than the stringency required of the standard. This may result in "college-ready" cut scores that are higher than warranted. My second essay uses the case of Minnesota which set a high passing standard for its math high school exit exam, but later waived the passing requirement for obtaining a high school diploma and required failing students to take remediation. I investigate the impact of barely failing on students’ high school and college outcomes. I find some evidence that within this context, there may be a cohort-dependent impact on on-time high school graduation and enrollment in 4-year colleges for students who score barely below versus barely above the math passing score. If the second essay shows that the passing score has consequences, the first study may help to advance wiser selection of cut scores. It is well-documented that gains on high-stakes state tests for low-income children and racial minority children are not matched on state-level audit tests such as the NAEP (Ho, 2007) or other low-stakes tests in districts (Jacob, 2005). This raises concerns about the generalizability of findings from one test used to another. Using the case of Texas where the curriculum standards stayed the same but the newly introduced high-stakes test focused more on "college-readiness" standards, my third essay investigates whether measured district-SES score gaps change when the test changes. I find that after the assessment focused more on "college-readiness" standards, district-SES score gaps between the 90th and 10th district-SES percentiles widen slightly, but are of smaller magnitude than previously found using low-stakes audit tests for students.
Publication Estimating the Impact of Receiving a Higher Evaluation Rating on Early-Career Teacher Turnover and Its Relationship to School Context
(2015-05-08) Rosner, Jessica Lori; Koretz, Daniel; West, Martin R.; Fullerton, JonTeacher evaluation systems have gone through widespread changes in recent years. These new systems have the potential to influence early-career teachers’ career trajectories because of their increase in rigor and ties to formal and informal consequences. In this paper, I investigate the impact of receiving a higher evaluation rating on early-career attrition and movement among districts and schools in a medium-sized state in the United States. Using a regression discontinuity design that takes advantage of the cut scores between each pair of evaluation ratings, I find that receiving a higher rating slightly increases the probability that early-career teachers will move to different districts. I do not find any effect of receiving a higher rating on the probability of a teacher leaving teaching in the state or moving among schools within a district. I find similar effects of ratings on teacher attrition and movement among schools for early-career and experienced teachers. When I apply the regression discontinuity to three cut scores simultaneously, I find that early-career teachers who received a higher evaluation rating were more likely to move among districts than experienced teachers who received a higher rating. Additionally, I do not detect any differences in the impact of teacher evaluation ratings on early-career teacher turnover based on school context. Finally, while on average early-career teachers moved to schools that were similar to those they left, I find that teachers who received higher ratings were more likely to move to schools with slightly higher percentages of low-income and non-white students.
Publication Multiple Imputation Methods for Latent Profile Analysis in Education and the Behavioral Sciences
(2020-05-18) Waldman, Marcus R.; Ho, Andrew; McCoy, Dana; Masyn, Katherine E.Researchers in education and the behavioral sciences are increasingly conducting latent profile analysis by focusing investigations on subpopulations of individuals to describe individual differences with greater nuance. At the same time, missing data is practically inevitable, and even modest missingness rates can threaten the validity of inferences. Multiple imputation is a powerful strategy to treat missing data, but it suffers from key limitations in both the imputation and pooling phases when conducting latent profile analysis. In this dissertation, I conduct three studies to address these gaps. In the first study (Chapter 2), I evaluate whether recursive partitioning imputation algorithms better mitigate nonresponse bias than alternative missing data approaches that are common in practice. I find that recursive partitioning imputation algorithms perform well when sample sizes are large (N = 1,200), but not when sample sizes are small (N = 300) or when class separation is weak (i.e., entropy ≈ .74). In response, I propose a hybrid imputation procedure in the second study (Chapter 3); the proposed method embeds a finite mixture model to generate imputations using a joint modeling framework within a larger chained equations procedure. I demonstrate the hybrid imputation procedure using real-world data. In the final study (Chapter 4), I scrutinize current practices for conducting finite mixture model selection with missing data. I am not aware of any studies evaluating whether model selection decisions are sensitive to real-world missing data problems. I fill this gap in research by studying whether selection decisions are sensitive to missing data if either a full information maximum likelihood (FIML) or a multiple imputation strategy is employed. Two findings emerge. First, with regards to FIML, the BIC under extracts the true number of classes relative to the complete data condition in the presence of small sample sizes and small classes. Second, with regards to multiple imputation, current practices for pooling information criteria result in model selection decisions that poorly replicate the decisions that would have been made had the data been complete. I propose two remedial procedures for future practice, and, using simulations, I show that these two procedures outperform current practices.
Publication The Predictive Validity of Information From Clinical Practice Lessons: Experimental Evidence From Argentina
(2015-05-17) Ganimian, Alejandro; Murnane, Richard J.; Deming, David J.; Barrera-Osorio, FelipeA growing number of teacher preparation programs require trainees to practice teaching. Yet, there is almost no evidence on whether the performance of individuals during clinical practice lessons predicts how they fare once they enter the school system.
We address this question by taking advantage of the fact that an alternative pathway into teaching in Argentina requires admitted applicants to complete two weeks of clinical practice. We collect information both during clinical practice and the school year. During clinical practice, we measure the performance of teaching trainees using classroom observations and student surveys. During the school year, we measure their performance using classroom observations, student surveys, and principal surveys. We find that the overall performance of trainees during clinical practice predicts their overall performance during the school year, but this prediction is only statistically under certain model specifications. The performance of these individuals during clinical practice predicts their ratings on classroom observations during the school year. This relationship remains statistically significant even when we account for how trainees fare on the application and selection processes of the alternative pathway. We also find that the performance of trainees on a brief demonstration lesson, delivered during the selection process, predicts their performance on classroom observations during the school year. The predictive effect is smaller than that of clinical practice lessons, but it raises the question of whether the additional effort required to collect information during clinical practice is worth the improved predictive validity.Publication School Choice and Educational Opportunities: The Upper-Secondary Student-Assignment Process in Mexico City
(2015-05-15) Ortega Hesles, Maria Elena; Murnane, Richard; Reimers, Fernando; Barrera-Osorio, FelipeMany education systems around the world use a centralized admission process to assign students to schools. By definition, some applicants to oversubscribed schools are not offered admission to their most-preferred school. Thus, one naturally asks whether it makes a difference to applicants’ educational opportunities and outcomes which schools they apply to, are offered admission to, and eventually enroll in. Each year in Mexico City, about 300,000 teenagers apply for a seat at one of the nearly 650 public upper-secondary schools. In this centralized, merit-based admission process, applicants are assigned to a school based on entrance examination score and their ranked list of school choices, subject to school capacity constraints.
In this dissertation, I include two papers assessing data from the upper-secondary application cohorts in Mexico City from 2005 to 2009. In the first paper, I find evidence of socio-economic stratification across schools. I also find dissimilarities in the application behavior of individuals according to their socio-economic background, even for those with high achievement levels. Based on qualitative and quantitative data from a small sample of applicants, I suggest that in addition to differences in economic resources, asymmetries in access to information might help to explain disparities in the application behavior of individuals from different socio-economic backgrounds.
In the second paper, I capitalize on the natural experiment created at each oversubscribed public upper-secondary school in Mexico City by the imposition of exogenous admission cut-off scores. Using a regression-discontinuity design, I estimate that, on average, upper-secondary applicants who score just above the admission threshold for a more competitive school (i.e. a school with higher cut-off score and higher average examination scores) have lower probability of graduating on time and within 5 years than do applicants who scored just below the admission threshold. Given the high take-up rates of the offers of admission, I find that the effects for enrollment in a more competitive school are only slightly larger than they are in their analogous reduced-form estimates. In addition, I show that effects differ across the distribution of admission cut-off scores and for applicants with selected socio-demographic characteristics who scored just above the admission threshold.
Publication Separating the Signal From the Noise: An Examination of Student and Teacher Scores Based on Student Learning Objectives (SLOs) in One State
(2015-05-15) Buckley, Katie; Hill, Heather; Ho, Andrew; Murnane, RichardDespite the prevalence of student learning objectives (SLOs) in teacher evaluation systems throughout the United States, research on the validity of student and teacher SLO scores used for high-stakes decisions is lacking. For this reason, this dissertation is comprised of two chapters that examine student and teacher-level SLO performance data from select districts in one Race to the Top state. In Chapter 1, I describe the quality of student assessment data and the comparability of student scores across alternative growth targets. I find that in the first year of implementation, assessments from half of the courses in the sample contained indicators of poor data quality, including anomalous score distributions and small to negative correlations between student prescores and postscores. However, in the second year of implementation, when student SLO performance is incorporated into final teacher evaluation scores, far fewer assessments contained anomalous score distributions, and there is no evidence to suggest manipulation of student scores. In addition to the assessments, the choice of student growth target does have an impact on the comparability of student and teacher scores across districts and years.
Chapter 2 describes the validity and reliability of teacher SLO scores. I find that while teacher SLO scores are moderately stable across courses, they are not stable over time, likely due to changes made to the assessments and targets used to determine student SLO scores. Further, for teachers with both SLO scores and an alternative metric of performance based on student growth, the two metrics do not converge. Finally, teachers in courses with higher average student prescores and lower proportions of students with disabilities have slightly higher SLO scores. In general, results on teacher SLO scores were similar to those found with value-added based metrics of teacher performance. Findings from both chapters suggest that improvement in the quality of the assessments administered as well as greater consistency in the growth targets assigned to students, both within districts over time and across districts, will improve the validity of student and teacher SLO scores in this state.