Harvard Medical School

Permanent URI for this communityhttps://dash.harvard.edu/handle/1/4454685

This community provides open access to material created by faculty, staff, and students of the Harvard Medical School. All material in the repository is also harvested by search engines (such as Google Scholar) and Open Archives Initiative data harvesters.

Browse

Search Results

Now showing 1 - 1 of 1
  • Publication

    Generating Actionable Clinical Summaries from EHR Data of Pediatric Patients Using Large Language Models

    (2026-05-15) Goel, Rishabh; Kohane, Isaac S; Fischer, Shira H; Mandl, Kenneth D

    Electronic health record (EHR) information overload is a well-documented driver of physician burnout, and pediatric settings face unique challenges including longitudinal developmental tracking, weight-based dosing, and age-specific reference ranges. This thesis presents a multi-agent large language model (LLM) framework for generating longitudinal clinical summaries from structured pediatric EHR data and evaluates these summaries against physician-authored summaries using the Provider Documentation Summarization Quality Instrument (PDSQI-9). The framework uses GPT-4.1 for summary generation and iterative improvement and Claude Opus 4.6 as an independent LLM-as-a-judge evaluator.

    Using the Pediatric Physicians Organization at Children's (PPOC) dataset, we developed a three-component pipeline consisting of a generation agent, an LLM-as-a-judge evaluation agent, and an improvement agent that operate in an iterative refinement loop. We selected a cohort of 10 pediatric patients stratified by clinical complexity (five with growth disorders, three with cancer-related histories, and two clinically healthy) and collected 30 physician-authored summaries from three practicing pediatricians through a custom web-based labeling tool.

    AI summaries refined through the feedback loop (AI+Feedback) achieved the highest mean PDSQI-9 score of 3.7 on a 5-point scale, significantly outperforming physician summaries. AI+Feedback scored higher than base AI summaries without feedback, though this difference was not statistically significant. AI+Feedback summaries prevailed on 8 of 10 patients and showed the largest quality gains for patients with cancer-related histories. However, AI+Feedback summaries were approximately 1.9 times longer than physician summaries and scored lower on the comprehensibility dimension. Exploratory analyses of temperature parameters revealed that the iterative refinement process is robust across a wide range of generation and improvement temperatures, with the feedback loop providing the greatest benefit when initial summary quality is low. These findings suggest that LLMs can produce pediatric clinical summaries that meet or exceed physician quality standards as measured by the PDSQI-9. However, the non-significant difference between AI+Feedback and base AI summaries, combined with the comprehensiveness-conciseness tradeoff, indicates that the iterative feedback loop’s primary value may lie in regularizing quality for complex cases rather than uniformly improving all summaries.