Skip to content
Preprint

GET: Generative Embedding Translation for Medical Image Segmentation

Aug 2026 · 0 citations · 44 references
Engineering Computer Science

TL;DR

Generative Embedding Translation (GET), a structured embedding-translation framework that progressively transforms image embeddings into mask embeddings within the frozen latent space of a Stable Diffusion VAE, is proposed.

Abstract

Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-to-mask translation must preserve target structure while remaining computationally efficient. We propose Generative Embedding Translation (GET), a structured embedding-translation framework that progressively transforms image embeddings into mask embeddings within the frozen latent space of a Stable Diffusion VAE. GET uses a U-Net-style Embedding Translation Network with 1.07M trainable parameters, combining Mobile Bottleneck Convolutions, Subsampled Self-Attention, and Multi-scale Feature Enrichment for local modeling, global context, and multi-scale refinement. Across five medical segmentation datasets, GET outperforms generative, CNN, and Transformer baselines. Compared with the strongest generative baseline, GMS, GET improves average Dice and IoU by 0.93% and 1.26%, reduces HD95 by 0.81 pixels, and uses 31.41% fewer trainable parameters. Under bidirectional BUS-BUSI domain shift, GET further improves Dice and IoU by 3.51% and 3.39%, while reducing HD95 by 27.37 pixels. Our code is available at: https://github.com/maklachur/GET.

View source

Similar papers

#diffusion models Open access Aug 2026

Robust unsupervised domain adaptation for medical image segmentation via frequency-conditioned graph diffusion

A novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance is proposed, highlighting the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.

Usman Ahmad Usmani, Arunava Roy, Junzo Watada · 0 citations
Aug 2026

SOM-GAN: A structure-preserving one-to-multiple generative adversarial network for unpaired medical image synthesis.

GAN-based image translation has been widely used for cross-domain medical image synthesis. However, most existing methods follow a one-to-one mapping paradigm, requiring a separate model for each target domain and increasing the training cost in multi-target tasks such as multi-sequence MRI synthesis. Although recent o...

Jinhao Li, Kai Hu, Runze Wang et al. · 0 citations
Preprint Aug 2026

CiUNet: A Hybrid Swin-CNN UNet for Medical Image Segmentation

Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...

Bin Dong, Jing-Hong Chen · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Open access Aug 2026

TSPFusion: Tri-stream and prototype network for learning detail-semantic fusion in medical image segmentation

A tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage, and a Global Prototype Bank that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing pe...

Mohammed A. M. Elhassan, Qian-Fa Yuan, Zhizhong Xu et al. · 0 citations
Conference Open access Sep 2026

A Local-Rotation-Driven Global Consistency Framework with Dual-View Decoding for Semi-Supervised Medical Image Segmentation

Medical image segmentation is still challenging, especially in semi-supervised scenarios where limited annotations are expected to support both accurate boundary delineation and coherent anatomical structures. We propose LR-GCF, a Local-Rotation-Driven Global Consistency Framework that couples strong local geometric pe...

Zhen Yang, Dong-Shuai Zhang, Yun-Liang Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.