This work proposes UnDA, an anchor-guided framework for unpaired cross-modal distillation that introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism, and dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision.
Abstract
Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions. To overcome these challenges, we propose UnDA, an anchor-guided framework for unpaired cross-modal distillation. Our approach introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism. To ensure robust knowledge transfer, we propose Uncertainty-Weighted Optimal Transport (UCT-OT), which dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision. Furthermore, a per-class ProtoNCE objective maintains stable prototype memories to enforce global discriminability across unpaired batches. Evaluations on representative segmentation tasks under strictly unpaired settings show consistent improvements in accuracy and boundary precision in the target modality, demonstrating that meaningful structural knowledge can be transferred across heterogeneous data sources without paired datasets.
Multimodal medical imaging benefits from the global context modeling of transformers, yet most existing models fuse modalities implicitly by channel concatenation, leaving cross-modal interaction unstructured or relying on costly multi-stream cross-attention. We propose Modality-Aware Token Interaction (MATI), an archi...
Selene Tomassini, Hafiza Ayesha Hoor Chaudhry, Alessandro Galdelli et al.· Proceedings of the Thirty-Fi...· 0 citations
Multimodal medical prediction often faces incomplete pairing: auxiliary modalities with complementary signal are available for only a subset of subjects (or none) and cannot be assumed at deployment. We introduce PANDA (Prototype Anchored Data Alignment), a two-stage framework that transfers auxiliary information to a...
Sheethal Bhat, Mahfuzur Rahman Chowdhury, P. A. Pérez-Toro et al.· 0 citations
Synthesizing missing modality medical images is a critical and challenging task. To achieve multimodality translation in a single model, existing methods usually construct an implicit modality-shared space with auxiliary loss and modulate the translation process with one-hot modality code. In this article, we propose a...
Le Hu, Qian-Di Yu, Fa-Ming Fang et al.· IEEE Transactions on Neural...· 0 citations
Automatic medical report generation (MRG) holds promise for alleviating radiologists’ workload, which has spurred growing interest in MRG for stroke diagnosis. However, existing approaches often fail to effectively model its long-range spatial dependencies and suppress phase-level noise in multi-slice sequences, lead...
Shaowei Shen· Poster Volume 0007 The 2026...· 0 citations
Multimodal medical images provide abundant and complementary diagnostic cues that are often indispensable for accurate clinical decision-making. However, various constraints, such as limited equipment availability, physical limitations of patients, etc. often lead to incomplete modality of medical data acquisition in t...
Jing Li, Qin-Kai Yu, Fei-Xiang Zhou et al.· Medical Image Analysis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.