A cross-border community for researchers with openness, equality and inclusion
Anatomy-Aware Vision-Language Learning for Medical Image Interpretation
ID:98 View protection:Participant Only Updated time:2026-07-22 16:09:56 Views:23 Online

Start Time:2026-07-30 16:10

Duration:15min

Session:[S4] Computer Vision and Pattern Recognition [S4-2] Computer Vision and Pattern Recognition

Abstract

Vision-language modeling has significantly advanced radiology by enabling models that jointly learn from medical images and radiology reports for tasks such as disease classification, report generation, and visual question answering. However, most existing approaches treat an entire medical image as a single entity during image-text alignment, overlooking the fine-grained anatomical reasoning process employed by radiologists. In clinical practice, radiologists systematically examine individual anatomical regions, associate findings with specific structures, and integrate these region-specific observations before arriving at a conclusion about the image as a whole. To address this limitation, we propose an anatomy-aware vision-language framework that learns anatomy-specific representations using dedicated anatomy tokens and anatomical segmentation masks. The framework further incorporates context-aware anatomical representations and jointly learns anatomical localization, aligns anatomical regions with their corresponding findings, and aligns global image representations with image-level disease categories within a unified vision-language framework. Extensive experiments on out-of-distribution datasets demonstrate the effectiveness of the proposed framework. The model achieves strong performance in zero-shot disease classification and anatomical segmentation, demonstrating robust generalization to unseen data and accurate localization of anatomical structures. Comprehensive ablation studies further validate the contribution of each component and the effectiveness of the proposed design choices.

Keywords
Vision-Language Models,Anatomy-Aware Learning,Medical Image Interpretation,Zero-Shot Classification,Medical Image Segmentation
Speaker
Shahab Ahmad Khan
University of Wisconsin–Madison

Post comments
Verification Code Change Another
All comments
Important Dates
  • Conference date

    07-30

    2026

    -

    08-01

    2026

  • 07-28 2026

    Draft paper submission deadline

  • 07-28 2026

    Registration deadline

Sponsored By

The United Societies of Science

Organized By

Kongunadu College of Engineering and Technology

Contact info
×

USS WeChat Official Account

USSsociety

Please scan the QR code to follow
the wechat official account.