Publication: Towards a Cognitive Science of Large Language Models
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Cognitive science studies the mind and the latent processes that let us process inputs and respond accordingly. Traditionally, many forms of cognition could only be studied in humans and other animals, but recent advances in AI have begun to change this. This raises an important question: what latent processes and representations give rise to the complex behaviors that AI models produce? In this dissertation, I propose an approach to understanding AI that is analogous to cognitive science, which I call cognitive interpretability.
In Chapter 1, I develop a theory for how Large Language Models (LLMs) learn in-context (i.e., from prompts) by using tools from cognitive science. I test LLMs on judging and generating random sequences of coin flips, a domain that has been studied in humans, and show that Bayesian model selection can explain LLM behavior. Chapter 2 builds on this and uses cognitive models to test whether LLMs reason the same way as humans. Here, I use another domain adopted from human psychology, and I show that LLMs use a different mechanism from humans when assigning responsibility to collaborators. Chapter 3 presents human experiments showing that when people evaluate agents, they value not only behavior but also the algorithms that drive the behavior. This suggests that understanding AI is important not only for practical reasons, but also because humans naturally evaluate minds and agents in terms of underlying processes rather than behavior alone. Together, these studies represent early steps towards developing a cognitive science of LLMs.