Publication:

EudAImonia: Epistemic Governance in Large Language Models to Reduce Hallucination and Sycophancy

Loading...
Thumbnail Image

Date

2026-06-02

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Ravichandran, Katharina A. 2026. EudAImonia: Epistemic Governance in Large Language Models to Reduce Hallucination and Sycophancy. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

Large language models hallucinate false content and produce sycophantic outputs, yet the dominant approaches to these failures, such as Reinforcement Learning from Human Feedback, Constitutional AI, and fine-tuning, share a common omission: they target symptoms without first asking what function an LLM serves. This thesis begins with that prior question. I argue that a central class of LLM outputs functions as assertions, propositions put forward as true, and should therefore be governed by assertoric norms. Operationalizing Timothy Williamson’s Knowledge Norm of Assertion for mechanistic systems, I define the Epistemic Governance Norm of Assertion (EGNA), which requires (i) that a system possess the capacity to partition epistemically available from unavailable content, and (ii) that this partition govern what is asserted. Drawing on recent mechanistic interpretability research, I show that modern LLMs already approximate the first condition. Current training objectives, however, structurally decouple epistemic registration from assertoric output, producing hallucination and sycophancy as predictable consequences. I implement a two-stage EGNA scaffold enforcing epistemic classification as a precondition of assertion and evaluate it across five frontier models. The scaffold reduces hallucination to zero or near-zero rates across all models and substantially suppresses sycophancy, while maintaining competitive response rates on legitimate queries. A control scaffold preserving the two-stage structure without the epistemic partition leaves hallucination largely unchecked, isolating the partition as the operative mechanism. The results also reveal that social inference enters epistemic classification itself, a failure mode that purely behavioral interventions would not have surfaced. Ultimately, getting the prior question right is not a philosophical exercise but a precondition for these systems to perform their function well.

Description

Other Available Sources

Research Data

Keywords

Artificial Intelligence, Epistemic Governance, Hallucination, Knowledge, Socrates, Sycophancy, Computer science, Philosophy, Artificial intelligence

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories