A novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance is proposed, highlighting the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.
Abstract
Cross-domain variability in medical imaging, arising from differences in scanners, acquisition protocols, and patient populations, remains a major challenge for reliable semantic segmentation. Existing unsupervised domain adaptation (UDA) methods predominantly rely on image-level transformations or feature alignment, which often fail to preserve anatomical consistency under large domain shifts. In this work, we propose a novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance. Specifically, we introduce a
Frequency-Conditioned Graph Diffusion
paradigm, where convolutional features are transformed into anatomical graphs to explicitly capture structural relationships. A latent diffusion process then progressively refines these graph embeddings, guided by frequency-aware contextual cues, enabling robust cross-domain alignment. To further enhance generalization, we integrate structural consistency regularization with adversarial latent alignment, eliminating the need for labeled target data. A dedicated decoder reconstructs dense segmentation maps, while stochastic diffusion sampling provides uncertainty estimates for improved potential clinical reliability. Extensive experiments on multiple public medical imaging benchmarks demonstrate that our method consistently outperforms state-of-the-art UDA approaches, achieving superior segmentation accuracy and robustness under significant domain shifts. These results highlight the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.
This work introduces a novel SFUDA framework built on Symmetrical Flow Matching, a unified generative model that segments an input image and synthesizes a source-like image from a mask within the same learned flow that outperforms SFUDA baselines and is competitive with conventional UDA methods.
Medical image segmentation is a key component of computer-aided diagnosis and treatment planning. Despite substantial progress in deep learning–based models, most existing approaches depend heavily on large annotated datasets and often fail to generalize across heterogeneous clinical environments, limiting their deploy...
V. Nguyen, Hoang Quan Luong, Phuc Ngoc Pham· IEEE International Conferenc...· 0 citations
Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-to-mask translation must preserve target structure while remaining computationally efficient. We propose Generative Embedding Translation (GET), a structured embedding-tran...
Md. Maklachur Rahman, M. H. Al Banna, Saraf Anjum et al.· 0 citations
An efficient diffusion framework that jointly diffuses a baseline scan and its follow-up residual, summed to synthesize the follow-up scan, while concurrently predicting a spatial uncertainty map, in a single reverse diffusion process is proposed.
A. Oliveras, Roger Marí, Rafael Redondo et al.· 1 citation· ⚡1
MDCL-UNet is proposed, a supervised multi-domain collaborative learning framework based on domain feature disentanglement that achieves consistently higher Dice scores and lower HD95 distances than single-dataset baselines and existing cross-dataset collaborative learning methods.
This study introduces a novel 3D registration framework centered on a dynamic wavelet transform module that achieves superior registration fidelity, highlighting its potential for practical clinical implementation.
Bo-Hua Chu, Bao-Ju Zhang, Bo Zhang et al.· Interdisciplinary Sciences C...· 0 citations