Generative Embedding Translation (GET), a structured embedding-translation framework that progressively transforms image embeddings into mask embeddings within the frozen latent space of a Stable Diffusion VAE, is proposed.
Abstract
Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-to-mask translation must preserve target structure while remaining computationally efficient. We propose Generative Embedding Translation (GET), a structured embedding-translation framework that progressively transforms image embeddings into mask embeddings within the frozen latent space of a Stable Diffusion VAE. GET uses a U-Net-style Embedding Translation Network with 1.07M trainable parameters, combining Mobile Bottleneck Convolutions, Subsampled Self-Attention, and Multi-scale Feature Enrichment for local modeling, global context, and multi-scale refinement. Across five medical segmentation datasets, GET outperforms generative, CNN, and Transformer baselines. Compared with the strongest generative baseline, GMS, GET improves average Dice and IoU by 0.93% and 1.26%, reduces HD95 by 0.81 pixels, and uses 31.41% fewer trainable parameters. Under bidirectional BUS-BUSI domain shift, GET further improves Dice and IoU by 3.51% and 3.39%, while reducing HD95 by 27.37 pixels. Our code is available at: https://github.com/maklachur/GET.
A novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance is proposed, highlighting the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.
GAN-based image translation has been widely used for cross-domain medical image synthesis. However, most existing methods follow a one-to-one mapping paradigm, requiring a separate model for each target domain and increasing the training cost in multi-target tasks such as multi-sequence MRI synthesis. Although recent o...
Jinhao Li, Kai Hu, Runze Wang et al.· Computerized Medical Imaging...· 0 citations
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
A tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage, and a Global Prototype Bank that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing pe...
Mohammed A. M. Elhassan, Qian-Fa Yuan, Zhizhong Xu et al.· Journal of King Saud Univers...· 0 citations
Medical image segmentation is still challenging, especially in semi-supervised scenarios where limited annotations are expected to support both accurate boundary delineation and coherent anatomical structures. We propose LR-GCF, a Local-Rotation-Driven Global Consistency Framework that couples strong local geometric pe...
Zhen Yang, Dong-Shuai Zhang, Yun-Liang Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.