Publication:

Building a Generalized Decoder for Immune Profiling in Pathology

Loading...
Thumbnail Image

Date

2026-05-15

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Lin, Sunni Chinjo. 2026. Building a Generalized Decoder for Immune Profiling in Pathology. Masters Thesis, Harvard Medical School.

Abstract

Accurate immune profiling of the tumor microenvironment (TME) requires integrating nuclei, tissue, and spatial-level information across sequential reasoning steps. Recent advances in vision language models (VLMs) have enabled natural language interaction with histopathology images, yet existing pathology VLMs are predominantly limited to single-turn question-answering and lack the structured, multi-step reasoning required for clinical immune profiling. We present a dataset construction, fine-tuning, and evaluation framework for multi-turn visual question answering (VQA), applied to computational immune profiling of the TME. Our multi-turn conversational pipeline decomposes immune profiling into sequential, clinically-grounded subtasks: nucleus detection and classification, cell-type quantification, tissue region analysis, and TME summarization. We introduce a structured approach to generate multi-turn VQA training data from annotated histopathology images. Our dataset comprises 238,747 patches spanning 22 organ types from 7 datasets, including hematoxylin and eosin (H&E) and immunohistochemistry (IHC) slides, from which we construct 515,382 multi-turn conversation samples across 13 turn types that span pixel-level and patch-level spatial scales. We further introduce a modular, turn-aware evaluation framework that assesses turn-level accuracy, cross-turn consistency, and conversational coherence. To our knowledge, this is the first such framework for multi-turn histopathology VQA. On an external test set, our fine-tuned model achieves a median nuclei detection F1 of 0.667, outperforming HoverNet on nuclei detection and segmentation. On a held-out test set, it surpasses all four VLM baselines across multi-turn evaluation metrics. This work establishes the feasibility of fine-tuning a general-purpose VLM for structured immune profiling. We envision this framework as a generalizable basis for developing and benchmarking conversational VLMs across immune profiling tasks and broader digital pathology applications.

Description

Other Available Sources

Research Data

Keywords

Computational Pathology, Histopathology, Immune Profiling, Multi-Turn Visual Question Answering, Tumor Microenvironment, Vision Language Model, Bioinformatics, Artificial intelligence, Medical imaging

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories