Publication: Stabilize or Explore: A Circuit Switch for Odor Plume Tracking
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Intelligent systems do not simply register inputs; they transform them into internal signals that can guide behavior or prediction. Across both biological and artificial systems, successful performance depends not only on whether task-relevant information is represented, but on whether it is organized and converted into a form that downstream circuits or computational modules can use. This dissertation examines that transformation in two domains: olfactory navigation in animals and spatial reasoning in generative neural networks. In both cases, the central question is how structured input is converted into actionable internal representations that support robust task performance.
The first part of the dissertation investigates how Drosophila use recent sensory history to regulate navigation in turbulent odor environments. To follow a path, animals can accumulate patchy sensory evidence over time, and then use that accumulation to adjust their heading. Using a closed-loop virtual plume paradigm, together with circuit mapping, calcium imaging, and optogenetic perturbations, this work identifies a neural mechanism underlying that arbitration. Sparse odor encounters (such as those encountered outside a plume) evoke slow exploratory turning, whereas frequent odor encounters (such as those inside a plume) promote faster, straighter runs; this causes the fly's heading to align with a direction associated with frequent odor encounters. Activity in two mushroom body output neurons (MBONs) forms an opponent representation of recent odor encounter statistics, with an increase in odor encounter frequency producing higher MBON21 activity and lower MBON09 activity. These MBONs then modulate the brain's internal representation of travel direction in h$\Delta$B cells of the central complex. Activating MBON21 suppresses turning and increases path straightness, whereas activating h$\Delta$B biases turning away from the fly's pre-stimulus goal heading. Together, these results show how coordination between the mushroom body and central complex can convert odor encounter history into navigation control, enabling robust navigation in intermittent sensory environments.
The second part examines how diffusion transformers convert linguistic structure into spatially organized visual outputs. Although Diffusion Transformers (DiTs) have substantially advanced text-to-image generation, they often fail to realize the spatial relations specified in a prompt. Using a mechanistic interpretability approach, this work investigates how DiTs generate correct spatial relations between objects. Models of different sizes and with different text encoders are trained from scratch on a controlled image generation task involving two objects with specified attributes and spatial relations. Although all models achieve near-perfect task accuracy, the underlying mechanisms differ substantially as a function of text encoder choice. With random text embeddings, spatial-relation information is transmitted to image tokens through a two-stage circuit involving cross-attention heads that separately read spatial relations and single-object attributes. With a pretrained text encoder (T5), the model instead relies on a different circuit in which spatial-relation and object information are jointly read from a single fused text token. Although in-domain performance is similar across these settings, robustness to out-of-domain perturbations differs, indicating that comparable task performance can arise from distinct internal mechanisms with different generalization properties.
Taken together, these studies support a common claim: representation alone is not sufficient for successful computation. In both biological and artificial systems, information must be transformed, structured, and routed in ways that make it usable for downstream processes. By characterizing these transformations in both neural circuits and deep generative models, this dissertation advances a unified view of how internal representations support adaptive behavior and reliable task performance.