Publication: AI Systems for Understanding and Grounding Radiology Reports
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Radiology reports are the primary medium through which radiologists communicate, conveying critical clinical information in natural language. AI holds the potential to both generate and analyze these reports, yet two key challenges persist. First, as AI systems increasingly generate radiology reports, evaluating their accuracy remains an open problem. Second, a fundamental disconnect exists between the findings described in radiology reports and their corresponding locations in the imaging studies, limiting referring physicians, patients, and trainees who must interpret findings without explicit visual guidance. This thesis addresses these challenges through three interconnected contributions spanning report evaluation, report visualization, and clinical application.
First, we introduce CRIMSON, a clinically grounded evaluation metric for radiology report generation, along with two new benchmarks: RadJudge, a 30-case clinical judgment test suite, and RadPref, a 100-case radiologist preference benchmark. CRIMSON incorporates patient context and weights errors by clinical significance when comparing generated against reference reports. Validated against radiologist error counts from the ReXVal dataset, RadJudge, and RadPref, CRIMSON achieves stronger alignment with expert judgment than prior metrics.
Second, we introduce ReXGroundingCT, the first publicly available dataset linking free-text radiology findings to manually annotated 3D segmentation masks in chest CT scans. Designed through a multi-stage annotation pipeline, the dataset comprises 3,142 scans and 8,028 segmented findings. We benchmark state-of-the-art text-prompted segmentation models, demonstrating that current approaches fall substantially short of clinical utility, even after fine-tuning. Subsequently, we release the dataset, along with a public leaderboard to drive continued progress on this task.
Third, we present RadGame, an AI-powered platform for radiology education that brings together both report evaluation and grounding as a concrete downstream application. The platform teaches two core skills: localizing findings through interactive bounding-box annotation, and writing radiology reports with automated structured feedback. In a prospective multi-institutional study with 18 medical students, participants using RadGame achieved a 68% improvement in localization accuracy and a 31% improvement in report-writing scores, outperforming traditional passive learning methods.
Together, these contributions address the radiology AI pipeline from multiple angles: evaluating generated reports against clinical standards, spatial grounding of reports in imaging, and applying both to advance radiology education.