Publication: From Hidden Activations to Human-Interpretable Concepts: Explaining and Interpreting Computational Pathology Classifiers Using Concept Learning
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Modern computational pathology models, particularly those based on foundation models and multiple-instance learning (MIL), achieve strong performance across diagnostic tasks but remain difficult to interpret due to their black-box nature. This lack of transparency limits their adoption in clinical workflows, where understanding the basis of a model’s prediction is essential for trust and validation. This thesis investigates whether human-interpretable nuclear morphology concepts can be recovered from the hidden-layer activations of MIL-based pathology classifiers and used to improve explainability without constraining model performance. In Aim 1, a proof-of-concept auxiliary concept model (CM) was developed to predict discrete nuclear morphology concepts from the hidden activations of a CHIEF-based LUAD/LUSC classifier trained on TCGA whole-slide images. Results demonstrate that concept signals are recoverable from intermediate representations, with higher recoverability in earlier attention layers and selective preservation of interpretable features such as nuclear size and shape. However, discretization introduced calibration limitations, restricting the CM’s ability to capture the continuous structure of nuclear morphologies. Aim 2 extends this framework by modeling concepts as continuous variables and benchmarking concept recoverability across MIL classifiers trained using four pathology foundation models: CHIEF, UNI, Virchow2, and GigaPath. A series of experiments evaluated the impact of CM architecture, loss functions, and concept-space designs using mean absolute error (MAE) and its variability across concepts. Nonlinear CMs improved performance over linear baselines, while concept clustering reduced bias toward highly correlated features. Across all settings, UNI consistently achieved the lowest MAE and variability, indicating stronger preservation of morphology-related information. Qualitative analyses further showed that predicted concepts are biologically grounded and spatially aligned with model attention, providing insight into both where and what features drive predictions. Overall, this work demonstrates the promise of learning concepts from hidden activations to enhance the interpretability and explainability of computational pathology models.