Skip to content

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

Jul 2026 · arXiv.org · Vol abs/2607.21546 · 0 citations · 23 references
Computer Science

TL;DR

This work proposes UnDA, an anchor-guided framework for unpaired cross-modal distillation that introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism, and dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision.

Abstract

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions. To overcome these challenges, we propose UnDA, an anchor-guided framework for unpaired cross-modal distillation. Our approach introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism. To ensure robust knowledge transfer, we propose Uncertainty-Weighted Optimal Transport (UCT-OT), which dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision. Furthermore, a per-class ProtoNCE objective maintains stable prototype memories to enforce global discriminability across unpaired batches. Evaluations on representative segmentation tasks under strictly unpaired settings show consistent improvements in accuracy and boundary precision in the target modality, demonstrating that meaningful structural knowledge can be transferred across heterogeneous data sources without paired datasets.

View source

Similar papers

Conference Open access Sep 2026

Structured Modality-Aware Token Interaction for Multimodal Medical Imaging

Multimodal medical imaging benefits from the global context modeling of transformers, yet most existing models fuse modalities implicitly by channel concatenation, leaving cross-modal interaction unstructured or relying on costly multi-stream cross-attention. We propose Modality-Aware Token Interaction (MATI), an archi...

Selene Tomassini, Hafiza Ayesha Hoor Chaudhry, Alessandro Galdelli et al. · 0 citations
Preprint Aug 2026

PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

Multimodal medical prediction often faces incomplete pairing: auxiliary modalities with complementary signal are available for only a subset of subjects (or none) and cannot be assumed at deployment. We introduce PANDA (Prototype Anchored Data Alignment), a two-stage framework that transfers auxiliary information to a...

Sheethal Bhat, Mahfuzur Rahman Chowdhury, P. A. Pérez-Toro et al. · 0 citations
Aug 2026

PUNF: A Prompt-Driven Unified Unpaired Medical Image Translation With Normalizing Flow.

Synthesizing missing modality medical images is a critical and challenging task. To achieve multimodality translation in a single model, existing methods usually construct an implicit modality-shared space with auxiliary loss and modulate the translation process with one-hot modality code. In this article, we propose a...

Le Hu, Qian-Di Yu, Fa-Ming Fang et al. · 0 citations
Conference 2026

ProFuseGPT: Progressive Fusion with Contrastive Refinement for Long-Sequence Medical Report Generation

Automatic medical report generation (MRG) holds promise for alleviating radiologists’ workload, which has spurred growing interest in MRG for stroke diagnosis. However, existing approaches often fail to effectively model its long-range spatial dependencies and suppress phase-level noise in multi-slice sequences, lead...

Shaowei Shen · 0 citations
Open access Sep 2026

T2ID: Towards trustworthy incomplete modality diagnosis via dual-axis adaptive completion and confidence-aware integration.

Multimodal medical images provide abundant and complementary diagnostic cues that are often indispensable for accurate clinical decision-making. However, various constraints, such as limited equipment availability, physical limitations of patients, etc. often lead to incomplete modality of medical data acquisition in t...

Jing Li, Qin-Kai Yu, Fei-Xiang Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.