Skip to content
Preprint

Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

This work evaluates three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs, and finds that image-text alignment achieves the most superior performance.

Abstract

Vision Transformers (ViTs) have shown immense potential in medical image analysis. However, standard pre-training via global image classification suffers from spatial collapse, where models rely heavily on background shortcuts rather than localising critical foreground lesions. To overcome this limitation and align visual evidence with precise medical semantics, we systematically investigate alternative pre-training paradigms.Specifically, we evaluate three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs. Notably, our empirical analysis reveals that while all three forms of structured supervision successfully alleviate the global pooling bottleneck and steer visual attention towards foreground regions, image-text alignment achieves the most superior performance. By embedding high-dimensional diagnostic logic, the cross-modal approach not only anchors attention on precise visual evidence but also enables profound abstract reasoning. Extensive experiments demonstrate that this semantically enriched pre-training fundamentally enhances the model's feature representation. Consequently, when fine-tuned for downstream clinical classification tasks, our models achieve superior accuracy and yield highly interpretable attention maps focused on true pathological features, vastly outperforming vanilla classification baselines.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal med...

Zihao Liu, Zhe Zhu, Xuzi Shi · 0 citations
Conference Open access Sep 2026

Hierarchical Conditional Energy Modeling for Medical Vision–Language Pretraining

Contrastive vision–language pretraining models such as CLIP align images and text in a shared embedding space but do not explicitly model or evaluate the hierarchical semantics common in medical image interpretation. We propose HCE-CLIP (Hierarchical Conditional Energy CLIP), a vision–language pretraining framework tha...

Cheng-Sheng Mao, Yuan Luo · 0 citations
Open access Aug 2026

TSPFusion: Tri-stream and prototype network for learning detail-semantic fusion in medical image segmentation

A tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage, and a Global Prototype Bank that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing pe...

Mohammed A. M. Elhassan, Qian-Fa Yuan, Zhizhong Xu et al. · 0 citations
Preprint Aug 2026

DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation

DistMedVL is proposed, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.

Jiaxuan Li, Qing Xu, Xiang-Jian He et al. · 0 citations
Sep 2026

Concept-Enhanced Multi-Scale Cross-Modal Alignment for Medical Visual Representation Learning.

Medical Vision-Language Pre-training (Med VLP) on paired medical images and reports has emerged as a promising direction for learning visual representations. However, current alignment approaches remain insufficient for learning fine-grained pathological details, largely due to the inherent difficulty of tokenizing rep...

Xiang-Min Kong, Xi-Bin Jia, Da-Wei Yang et al. · 0 citations
Preprint Aug 2026

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

This work presents MedPixel, a unified medical pixel-language model built around a shared language--mask interface that achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prom...

Haoyu Yang, Meixing Shi, Zeng-Jie Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.