Publication:

Using Large Language Models to Predict Linguistic Neurological Intracranial Signals for Multimodal Sentences

Loading...
Thumbnail Image

Date

2026-06-02

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Steele, Alliyah Nicole. 2026. Using Large Language Models to Predict Linguistic Neurological Intracranial Signals for Multimodal Sentences. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

Prior work using fMRI and EEG has demonstrated correlations between large language model (LLM) representations and neural language signals, however these approaches are limited by low temporal resolution and small sample sizes. To address these gaps, we analyzed stereoelectroencephalography (SEEG) recordings from 16 multilingual participants completing an auditory and visual sentence task with semantic and syntactic violations. Per-electrode ridge regressions mapped hidden layer embeddings from three transformer models (BERT, GPT-2-XL, and mT5-XL) onto gamma-band neural signals across cortical and subcortical regions, without anatomical priors. Permutation testing and False Discovery Rate (FDR) correction were applied to control for false positives. Significant neural-LLM alignment was sparse however anatomically specific, clustering in right-hemisphere homologous language regions. The right inferior frontal sulcus (IFS) showed the highest significant explained variance (R2 = 0.27, p.001) driven primarily by responses to syntactic and semantic violations. This was consistent with previous research indicating right-hemisphere linguistic regional recruitment under increased linguistic processing demand. The insular cortex, despite not being a classical language area, showed selective alignment with BERT embeddings under grammatical conditions, supporting its proposed role in auditory-motor integration and grammar learning. These findings suggest transformer representations may capture aspects of both classical and non-classical language processing, particularly under conditions of linguistic surprisal. This work demonstrates the potential of transformer-SEEG alignment as a tool for mapping language-relevant cortical regions and has potential downstream implications for modeling learning disabilities, brain-computer interface decoding, and cognitively-inspired LLM architectures.

Description

Other Available Sources

Research Data

Keywords

AI, Linguistics, LLM, Neuro-AI Alignment, Neurolingustics, SEEG, Neurosciences

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories