Sep 2026· International journal of imaging systems and technology (Print)· Vol 36· 0 citations· 43 references
TL;DR
A unified framework is presented that simultaneously models segmentation variability and expert‐specific behavior within a single architecture, enabling personalized predictions while preserving diversity and demonstrating consistent improvements over existing methods in both diversity and personalization metrics.
Abstract
Accurate medical image segmentation is often hindered by ambiguity arising from both image quality limitations and variability in expert interpretation. Although datasets with multiple annotators provide richer information about such uncertainty, most existing methods either compress these annotations into a single target, thereby ignoring individual differences, or generate multiple predictions without maintaining alignment with specific experts. As a result, current approaches fail to fully exploit the structure of multi‐rater data. We present a unified framework that simultaneously models segmentation variability and expert‐specific behavior within a single architecture, enabling personalized predictions while preserving diversity. The proposed method learns a shared latent representation that captures the range of plausible annotations and leverages it to generate both diverse and expert‐aligned segmentations. To achieve this, we introduce a Gaussian context‐guided attention mechanism that adaptively extracts relevant features from the latent space in a structured and parameter‐efficient manner. This design allows the model to reflect distinct annotation patterns across experts without relying on heavily parameterized attention modules. By jointly modeling shared anatomical knowledge and expert‐specific variations, the framework maintains consistency across predictions while adapting to individual annotator preferences. We evaluate the proposed approach on the LIDC‐IDRI dataset and a nasopharyngeal carcinoma (NPC‐170) dataset. The results demonstrate consistent improvements over existing methods in both diversity and personalization metrics. Furthermore, the model produces expert‐aligned predictions with enhanced interpretability and reduced computational overhead, suggesting potential applicability in clinical workflows where understanding annotation variability is essential.
Semi-supervised medical image segmentation methods have drawn wide attention as they reduce reliance on heavily annotated data. However, existing models suffer from confirmation bias with limited annotations, and structural or parameter coupling hinders self-correction, especially for medical images with ambiguous boun...
Dong-Sheng Wang, Xiao-Han Lang· Biomedical engineering and p...· 0 citations
Visual in-context learning (ICL) enables medical image segmentation models to adapt to new tasks using only a small set of annotated support images, without additional model fine-tuning. This makes visual ICL promising for clinical deployment, yet a fundamental limitation remains. Current ICL models lack the intrinsic...
Tiana-Tian-Ying Chen, Liang-Li Zhen, Yan-Yu Xu et al.· IEEE Transactions on Medical...· 0 citations
Precise medical image segmentation is essential to modern clinical workflows and biomedical research. However, current automated models often lack the flexibility, generalizability, and clinician control required to adapt to out-of-distribution data or novel classes without computationally expensive retraining. Further...
Paul Machauer, M. Reisert, Janis Keuper· IEEE Access· 0 citations
A novel structured latent UDA framework that performs domain alignment in a topology-aware representation space rather than directly modifying image appearance is proposed, highlighting the effectiveness of structured latent modeling and diffusion-based learning for robust domain-adaptive segmentation.
Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as ControlNet, offer an alternative by generating images conditioned on semantic masks and text. However,...
Aayushi Tyagi, Prathosh A. P., Mausam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.