GSE Scholarly Articles
Permanent URI for this collectionhttps://dash.harvard.edu/handle/1/3345928
Browse
Publication Abstract Representations of Attributed Emotion: Evidence From Neuroscience and Development
(2015-05-12) Skerry, Amy; Spelke, Elizabeth; Carey, Susan; Rebecca, Saxe; Nancy, KanwisherHumans can recognize others’ emotions based on overt cues such as facial expressions, affective vocalizations, or body posture, or by recruiting an abstract, causal theory of the conditions that tend to elicit different emotions. Whereas previous research has investigated the recognition of emotion in specific perceptual modalities (e.g. facial expressions), this dissertation focuses on the abstract representations that relate observable reactions to their antecedent causes. A combination of neuroimaging, behavioral, and developmental methods are used to shed light on the mechanisms that support various forms of emotion attribution, and to elucidate the core features or dimensions that structure the space of emotions we represent. Chapter 1 identifies brain regions that contain information about emotional valence conveyed either via facial expressions and or via animations depicting abstract situational information. These data reveal regions with modality-specific representations of emotional valence (i.e. patterns of activity that discriminate only positive versus negative facial expressions), as well as modality-independent representations: in medial prefrontal cortex (MPFC), the valence representation generalizes across stimuli, indicating a common neural code that abstracts away from specific perceptual features and is invariant to different forms of evidence. Building on evidence that young infants discriminate and respond to the emotional expressions of others, Chapter 2 investigates whether infants also represent these expressions in relation to the situations that elicit them. The results of several experiments demonstrate that infants within their first year of life have expectations about how facial and vocal displays of emotion relate to the valence of events that precede them. Whereas Chapters 1 and 2 focus on a simple binary distinction between positive and negative affect, Chapter 3 investigates a space of more fine-grained discriminations (e.g. someone feeling proud vs. grateful). A combination of multi- voxel pattern analyses and representational similarity analyses reveal brain regions containing abstract and high-dimensional representations of attributed emotion. Moreover, a set of causal features (encoding properties of eliciting events that vary between different emotions) outperforms more primitive dimensions in capturing neural similarities within these regions. Together, these studies provide a newly detailed characterization of the representations that structure emotion attribution, including their development and neural basis.
Publication Academic Discussions: An Analysis of Instructional Discourse and an Argument for an Integrative Assessment Framework
(American Educational Research Association (AERA), 2012) Elizabeth, Tracy; Ross Anderson, T. L.; Snow, E. H.; Selman, RobertThis article describes the structure of academic discussions during the implementation of a literacy curriculum in the upper elementary grades. The authors examine the quality of academic discussion, using existing discourse analysis frameworks designed to evaluate varying attributes of classroom discourse. To integrate the overlapping qualities of these models with researchers’ descriptions of effective discussion into a single instrument, the authors propose a matrix that (1) moves from a present/absent analytic tendency to a continuum-based model and (2) captures both social and cognitive facets of quality academic discourse. The authors conclude with a discussion of how this matrix could serve to align teachers’ and researchers’ identification of quality academic discussion and the process by which users could measure improvement in students’ discourse skills over time.
Publication Accelerating Antimicrobial Discovery With Controllable Deep Generative Models and Molecular Dynamics
(2021-03-11) Das, Payel; Sercu, Tom; Wadhawan, Kahini; Padhi, Inkit; Gehrmann, Sebastian; Cipcigan, Flaviu; Chenthamarakshan, Vijil; Strobelt, Hendrik; dos Santos, Cicero; Chen, Pin-Yu; Yang, Yi Yan; Tan, Jeremy P.K.; Hedrick, James; Crain, Jason; Mojsilovic, Aleksandra; MojsilovicDe novo therapeutic design is challenged by a vast chemical repertoire and multiple constraints such as high broad-spectrum potency and low toxicity. We propose CLaSS (Controlled Latent attribute Space Sampling) — an efficient computational method for attribute-controlled generation of molecules, which leverages guidance from classifiers trained on an informative latent space of molecules modeled using a deep generative autoencoder. We further screen the generated molecules for additional key attributes by using a set of deep learning classifiers in conjunction with novel physicochemical features derived from high-throughput molecular simulations. The proposed approach is employed for designing non-toxic antimicrobial peptides (AMPs) with strong broad-spectrum potency, which are emerging drug candidates for tackling antibiotic resistance. Synthesis and wet lab testing of only twenty designed sequences identified two novel and minimalist AMPs with high potency against diverse Gram-positive and Gram-negative pathogens, including the multidrug-resistant K. pneumoniae, as well as low in vitro and in vivo toxicity.Live-cell confocal imaging revealed that the bactericidal mode of action of the peptides occurs through membrane pore formation. Both antimicrobials mitigate the onset of drug resistance and are effective against antibiotic-resistant strains. The proposed approach thus presents a viable path for faster and efficient discovery of potent and selective broad-spectrum antimicrobials.
Publication Publication Adapting Educational Measurement to the Demands of Test-Based Accountability
(Informa UK Limited, 2015) Koretz, DanielAccountability has become a primary function of large-scale testing in the U.S. The pressure on educators to raise scores is vastly greater than it was several decades ago. Research has shown that high-stakes testing can generate behavioral responses that inflate scores, often severely. I argue that because of these responses, using tests for accountability necessitates major changes in the practices of educational measurement. The needed changes span the entire testing endeavor. This paper addresses implications for design, linking, and validation. It offers suggestions about possible new approaches and calls for research evaluating them.
Publication Adjusting treatment effect estimates by post-stratification in randomized experiments
(Wiley-Blackwell, 2012) Miratrix, Luke; Sekhon, Jasjeet S.; Yu, BinExperimenters often use post-stratification to adjust estimates. Post-stratification is akin to blocking, except that the number of treated units in each stratum is a random variable because stratification occurs after treatment assignment. We analyse both post-stratification and blocking under the Neyman–Rubin model and compare the efficiency of these designs. We derive the variances for a post-stratified estimator and a simple difference-in-means estimator under different randomization schemes. Post-stratification is nearly as efficient as blocking: the difference in their variances is of the order of 1/n2, with a constant depending on treatment proportion. Post-stratification is therefore a reasonable alternative to blocking when blocking is not feasible. However, in finite samples, post-stratification can increase variance if the number of strata is large and the strata are poorly chosen. To examine why the estimators’ variances are different, we extend our results by conditioning on the observed number of treated units in each stratum. Conditioning also provides more accurate variance estimates because it takes into account how close (or far) a realized random sample is from a comparable blocked experiment. We then show that the practical substance of our results remains under an infinite population sampling model. Finally, we provide an analysis of an actual experiment to illustrate our analytical results.
Publication Anchoring and Adjusting in Questionnaire Responses
(Informa UK (Taylor & Francis), 2012) Gehlbach, Hunter; Barge, ScottWhen ordering items on attitude/opinion questionnaires, do survey designers bias respondents’ answers by the mere act of choosing to organize their survey in a particular way? We hypothesize that, under specific frequently-occurring conditions, respondents employ an anchoring and adjusting strategy in which their response to an initial survey item provides a cognitive anchor from which they (insufficiently) adjust in answering the subsequent item. Three experiments indicate that respondents anchor and insufficiently adjust in certain situations, anchoring and adjusting leads to higher inter-item correlations between adjacent items, and these inflated correlations can (spuriously) increase the reliability estimate of the scale that they comprise and affect the resultant correlations with other measures. These effects are not consistently accounted for by a “superior memory search” explanation. In organizing their surveys, researchers may wish to combat this bias by intermixing items designed for different, but related constructs.
Publication An Applied Researcher’s Guide to Estimating Effects From Multisite Individually Randomized Trials: Estimands, Estimators, and Estimates
(Taylor & Francis, 2020) Miratrix, Luke; Weiss, Michael J.; Henderson, BritResearchers face many choices when conducting large-scale multisite individually randomized control trials. One of the most common quantities of interest in multisite RCTs is the overall average effect. Even this quantity is non-trivial to define and estimate. The researcher can target the average effect across individuals or sites. Furthermore, the researcher can target the effect for the experimental sample or a larger population. If treatment effects vary across sites, these estimands can differ. Once an estimand is selected, an estimator must be chosen. Standard estimators, such as fixed-effects regression, can be biased. We describe 15 estimators, consider which estimands they are appropriate for, and discuss their properties in the face of cross-site effect heterogeneity. Using data from 12 large multisite RCTs, we estimate the effect (and standard error) using each estimator and compare the results. We assess the extent that these decisions matter in practice and provide guidance for applied researchers.
Publication Archiving Blackness: Reimagining and Recreating the Archive(s) as Literary and Information Wake Work
(Journal of Contemporary Archival Studies, 2023-01-27) Gabriel, Jamillah R.“…we, Black people everywhere and anywhere we are, still produce in, into, and through the wake an insistence on existing: we insist Black being into the wake.”
– Christina Sharpe, In the Wake (2016)
In this paper, I introduce Christina Sharpe’s conceptualizations of wake and wake work, as they pertain to archiving the experiences of Blackness to better understand how the archive and archives are vital for those living and working in the wake of slavery. I am particularly interested in the wake work conducted both in literary works (speculative fiction) and at information sites (archives). To that end, I closely examine archives as they are presented in literature so as to explicate how these archival narratives created by Black authors perform wake work. Moreover, I make the connection between literary wake work, that which is performed by Black speculative fiction writers, and information wake work, that which is performed by Black community archivists, before delving into an analysis of the physical act of creating archives as the wake work of Black archivists. This investigation of wake work and archive(s) is meant to articulate Black life through a multidisciplinary lens, one that merges scholarship in Black studies, archives, information, and literature. My interrogation of archiving Blackness centers on the concepts of “wake” and “wake work,” and how they can be used to characterize the act of archiving the histories and the futures of Black people as an intervention towards coloring and diversifying the archival record.
Publication Art in the Advancement of Understanding
(2002) Elgin, CatherineCognitive progress often involves reconfiguring a domain, bringing previously unrecognized likenesses, differences, patterns and discrepancies to light. I argue that the arts effect such reconfigurations, enabling us to discern and appreciate the importance of aspects of the domain that we had previously overlooked or underemphasized. I argue that so-called ‘aesthetic devices’ like metaphor, fiction, and exemplification figure in our understanding of science as well as art. We cannot do justice to our scientific understanding while denying that art and its devices function cognitively.
Publication Artificial Intelligence and Educational Measurement: Opportunities and Threats
(American Educational Research Association (AERA), 2024-05-09) Ho, AndrewI review opportunities and threats that widely accessible Artificial Intelligence (AI)-powered services present for educational statistics and measurement. Algorithmic and computational advances continue to improve approaches to item generation, scale maintenance, test security, test scoring, and score reporting. Predictable misuses of AI for these purposes will result in biased scores, construct underrepresentation, and differential impact over time. Recent efforts to develop standards for AI use in testing like those of Burstein are promising. I argue that similar efforts to develop AI standards for educational measurement will benefit from increased attention to the context of test use and explicit commitment to ongoing monitoring of bias and scale drift over time.
Publication Assessing Instructional Explanations for Mathematical Procedures at Scale Using Animated Teaching Simulations
(University of Chicago Press, 2026-03) Hill, Heather; Garcia Coopersmith, Jeanette; Kleen, HannahStudent mastery of elementary mathematical procedures is foundational to learning in the discipline and to success in more advanced mathematics. Prior studies suggest that classroom instruction often focuses on the steps of procedures without also providing support for their meaning, for instance by emphasizing place value or justifying steps. However, studies on this topic are either outdated or limited in scale. In response, we analyze 324 teachers’ spoken instructional explanations in reaction to 6 animated teaching simulations covering 3 teaching tasks—explaining a procedure, addressing student confusion, and summarizing a nonstandard student method. An analysis of these data reveals that teachers primarily focus on the steps of procedures except when summarizing nonstandard student methods. Results provide clues about the nature of US classroom instruction and offer a new tool for evaluating the impact of efforts to change that instruction.
Publication Assessing Reading Comprehension in Bilinguals
(University of Chicago Press, 2006) August, Diane; Francis, David J.; Hsu, Han‐Ya Annie; Snow, CatherineA new measure of reading comprehension, the Diagnostic Assessment of Reading Comprehension (DARC), designed to reflect central comprehension processes while minimizing decoding and language demands, was pilot tested. We conducted three pilot studies to assess the DARC’s feasibility, reliability, comparability across Spanish and English, developmental sensitivity, and relation to standardized measures. The first study, carried out with 16 second‐through sixth‐grade English language learners, showed that the DARC items were at the appropriate reading level. The second pilot study, with 28 native Spanish‐speaking fourth graders who had scored poorly on the Woodcock‐Johnson Language Proficiency Reading Passages subtest, revealed a range of scores on the DARC, that yes‐no answers were valid indicators of respondents’ thinking, and that the Spanish and English versions of the DARC were comparable. The third study, carried out with 521 Spanish‐speaking students in kindergarten through grade 3, confirmed that different comprehension processes assessed by the DARC (text memory, text inferencing, background knowledge, and knowledge integration) could be measured independently, and that DARC scores were less strongly related to word reading than Woodcock‐Johnson comprehension scores. By minimizing the need for high levels of English oral proficiency or decoding ability, the DARC has the potential to reflect the central comprehension processes of second‐language readers of English more effectively than other measures.
Publication Assessment in Early Literacy Research
(Guilford Press, 2011) Snow, Catherine; Oh, Soojin S.Much of what we know about children’s language and literacy development derives from efforts to assess those skills. In fact, language and literacy development might be taken as a case study in the history of assessment—a local domain which displays the full range of tensions, challenges, and approaches that have characterized the field of behavioral assessment, and in particular, the assessment of young children. In this chapter, we discuss language and literacy assessment in young children as an illustrative special case of issues that extend far beyond the language/literacy domain. In that larger domain, as in this specific one, three key questions organize the information: For what purposes should we assess young children? What aspects of their functioning should be assessed? And how do we carry out assessments so as to get good, reliable information with only modest burden on the adult assessor or the child?
Publication Assessment of cognitive abilities in multiethnic countries: The case of the Wolof and Mandinka in the Gambia
(2010) Jukes, Matthew; Grigorenko, Elena L.Background: The use of cognitive tests is increasing in Africa but little is known about how such tests are affected by the great ethnic and linguistic diversity on the continent.
Aim: To assess ethnic and linguistic group differences in cognitive test performance in the West African country of the Gambia and to investigate the sources of these differences.
Samples: Study 1 included 579 participants aged 14–19 years from the Wolof and Mandinka ethnic groups of the Gambia. Study 2 included 41 participants aged 12–18 years from the two ethnic groups.
Methods: Study 1 assessed performance on six cognitive tests. Participants were also asked about their history of education, residence in the city, parental education, and family socio-economic status. Study 2 assessed performance on two versions of the digit span test. Recall of the numbers 1–5 were compared with recall of numbers 1–9 for both the Wolof (who count in base 5) and the Mandinka (who count in base 10).
Results: Study 1 established that Wolof performance was lower than that of the Mandinka on five out of six cognitive tests. In four of these tests, group differences were partially mediated by participation in primary school and migration to the city. Group differences were substantial for the digit span test and were not attenuated by mediating variables. Study 2 found that digit span among the Wolof was shorter than that of the Mandinka for numbers 1–9 but not for numbers 1–5.
Conclusions: Several suggestions are made on how to consider the ethnicity, language, education, and residence (urban vs. rural) of groups when conducting comparative cognitive assessments or collecting normative data.
Publication Assessment, technology and change
(2010) Clarke, Jody; Dede, ChristopherPublication Assisting students struggling with mathematics: Response to intervention (RtI) for elementary and middle schools.
(2009) Gersten, Russell; Beckmann, Sybilla; Clarke, Benjamin; Foegen, Anne; Marsh, Laurel; Star, Jon; Witzel, BradleyTaking early action may be key to helping students struggling with mathematics. The eight recommendations in this guide are designed to help teachers, principals, and administrators use Response to Intervention for the early detection, prevention, and support of students struggling with mathematics.
Publication Atom tracker: Designing a mobile augmented reality experience to support instruction about cycles and conservation of matter in outdoor learning environments
(2016) Kamarainen, Amy; Metcalf, Shari; Grotzer, Tina; Brimhall, C; Dede, ChristopherWe describe a mobile augmented reality (AR) experience called Atom Tracker designed to help middle school students better understand the cycling of matter in ecosystems with a focus on the concept of conservation of matter and the processes of photosynthesis and respiration. Location-based AR allows students to locate virtual "hotspots," where they interact with multiple representations including vision-based AR animations of virtual atoms during ecological processes such as photosynthesis and physical LEGO® -based representations of molecules. This design case describes the design rationale, the iterative design process, the context for implementation, and reflections on the success and limitations of the Atom Tracker AR experience. An augmented reality interface was chosen due to theoretical support for its utility in supporting interaction with multiple representations (both physical and virtual) of atoms and molecules, the ability to condense and expand temporal and spatial scales associated with ecological processes, and its ability to explicitly situate these representations in real-world contexts that could support learning. Two significant design challenges that we recognized were (a) appropriately leveraging narrative, student engagement and agency when designing around the topic of atoms and molecules, which are inanimate and invisible; and (b) designing for engagement with both virtual and physical resources available during the experience.
Publication Auditing for Score Inflation Using Self-Monitoring Assessments: Findings from Three Pilot Studies
(Taylor and Francis, 2016) Koretz, Daniel; Jennings, Jennifer L.; Hui, Leng Ng; Yu, Carol; Braslow, David; Langi, MeredithResearch has shown that test-based accountability programs often produce score inflation. Most studies have evaluated inflation by comparing trends on a high-stakes test and a lower-stakes audit test. However, Koretz and Benguin (2010) noted the weaknesses of using external audit tests and suggested instead using self-monitoring assessments (SMAs), which incorporate into high-stakes tests audit items that are not susceptible to test preparation aimed at more predictable items. This paper reports the results of the first three trials of the SMA approach, evaluating whether SMAs can detect inflation in a context in which it has been demonstrated to exist. The studies were conducted with the New York State mathematics tests in grades 4, 7, and 8 in 2011 and 2012. Despite a severe conservative bias created by numerous aspects of the study designs, we found that the audit component functioned as expected in many of the trials. The difference in performance between nonaudit and audit items was associated with factors that earlier research showed to be related to test preparation and score inflation, such as "bubble-student" status (scoring just below the Proficient cut in the previous year) and school poverty. However, a number of trials yielded null findings. These findings underscore the need for additional research investigating the optimal characteristics of audit items.