Publication:

Latent Structure of Affective Representations in Large Language Models

Loading...
Thumbnail Image

Date

2026-06-24

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Choi, Benjamin J.. 2026. Latent Structure of Affective Representations in Large Language Models. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing work has focused mainly on general geometric and topological properties of learned representations, but validating such findings is challenging due to a lack of ground-truth latent geometry. Emotion processing provides a compelling testbed for probing representational geometry, as emotions exhibit both categorical organization and continuous affective dimensions that are well-established in the psychology literature.

In this thesis, we investigate the latent structure of affective representations in LLMs using geometric data analysis tools. We present five main findings. First, we show that Gemma-2-9B, Mistral-7B, and LLaMA-3-70B-Instruct learn coherent internal representations of emotion that align with the valence--arousal model from psychology, as confirmed by statistically significant Procrustes alignment with human normative ratings. Second, we find that these representations exhibit nonlinear geometric structure consistent with the parabolic curvature of valence--arousal space, yet remain well-approximated by linear methods---providing nuanced evidence for the linear representation hypothesis commonly assumed in model interpretability. Third, we demonstrate that the geometry of the learned representation space can be leveraged to build well-calibrated uncertainty estimates for emotion classification, with expected calibration errors below 0.011 across all models tested. Fourth, we show that probe-derived directions in activation space can causally steer the emotional content of generated text, as confirmed by independent human raters. Fifth, a parallel analysis on human EEG brainwave data reveals that similar parabolic valence--arousal geometry emerges in human neural representations of emotion, suggesting structural parallels between artificial and biological affective processing. Taken together, these findings demonstrate that LLMs learn affective representations with remarkably meaningful geometric structure that can be leveraged for both interpretability and practical applications.

Description

Other Available Sources

Research Data

Keywords

interpretability, large language models, manifold learning, neural representations, representation geometry, Artificial intelligence, Applied mathematics, Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories