Publication: Disambiguating Large Language Model Performance on the Ambiguities of Law, Reasoning, and the Future
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Over the past year, two separate cases of legal actors using LLM-based products to submit fictitious materials to American courtrooms have raised significant concern. In this thesis, I tackle one facet of this legal LLM usage problem, in arguing normatively for the need to teach LLMs legal reasoning. Drawing from legal theory but developing new definitions for the LLM context, I establish that legal reasoning requires reasoning over a central ambiguity in the application of general statutes to specific fact patterns — an inescapable ambiguity that I find traditional LLM reasoning research paradigms unequipped to handle. Thus, I contribute back to LLM reasoning research an experimental methodology that directly targets this central ambiguity and intuition in reasoning, as well as an evaluative framework that holistically captures true reasoning ability. I rework two prior legal ML datasets into benchmarks for a novel legal ambiguity identification task (available on HuggingFace). Conducting a preliminary experimental exploration of GPT-3.5 and GPT-4 models on this task, I discover some true ability at identifying legal ambiguity. Though it is currently weak and somewhat misaligned with legal practice, this ability shows promise for improvement, especially through model milestoning and fine-tuning. Overall, this thesis analyzes one component of ideal legal LLM usage from an interdisciplinary legal theory and computational perspective, toward progress in principled legal LLM implementations and more productive law and CS collaboration.