Publication: Latent Morphologic Representations in Pathology Foundation Models and Their Implications for Algorithmic Disparities and Clinical Prediction
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Histopathology has long been the gold standard for cancer diagnosis and remains central to clinical decision-making in oncology1,2. Through examination of hematoxylin-and-eosin (H&E)–stained whole-slide images, pathologists assess tumor architecture, cellular morphology, and microenvironmental context to determine tumor type, grade, and stage and to estimate disease aggressiveness. These morphologic features provide critical information for prognosis and treatment planning2. However, tumors that appear similar under routine histopathologic evaluation can nevertheless exhibit markedly different clinical trajectories, including recurrence and metastatic progression3-5. This discrepancy suggests that additional biological signals influencing tumor behavior may be present within tissue architecture but remain difficult to systematically quantify through conventional visual interpretation alone4,5. Recent advances in machine learning have begun to reveal that routine histologic images encode far richer biological information than previously appreciated6. Large-scale pathology foundation models trained through self-supervised learning can extract high-dimensional visual representations that capture subtle patterns of tissue composition7,8, nuclear morphology9, and spatial organization6,10,11. These models have demonstrated the ability to recover molecular alterations6,12, transcriptional programs13-15, and clinical outcomes6,15,16 directly from H&E slides, even when such signals lack clear morphologic correlates recognizable to pathologists. Together, these findings indicate that pathology images contain latent morphologic signals reflecting underlying biological processes that can be systematically captured through representation learning. Despite these advances, the biological meaning of the morphologic representations learned by pathology foundation models remains incompletely understood. While predictive performance has been widely demonstrated, it is often unclear which morphologic features drive model predictions and whether these signals reflect tumor-intrinsic biological programs or systematic variation related to patient populations17. Clarifying the nature of these signals is therefore essential both for improving the interpretability of pathology AI systems and for ensuring that such models produce reliable and equitable predictions when applied in clinical settings. The overarching goal of this thesis is to examine how pathology foundation models encode latent morphologic signals in histologic images and determine how these signals reflect both tumor biology and systematic variation that influences downstream clinical predictions. To address this objective, we focus on two complementary questions. First, we examine whether systematic variation in tissue morphology related to patient characteristics is captured within the representations learned by pathology foundation models and how such signals influence model performance. Second, we investigate whether routine pathology images contain conserved morphologic programs associated with tumor trajectories, including recurrence and distant metastatic progression. Together, these studies provide complementary perspectives on how pathology foundation models encode biologically structured variation in histologic images and demonstrate that linking model representations to quantifiable morphologic features can advance both the biological interpretation and the reliable clinical application of computational pathology systems.