Jul 2026· Journal of Electronic Imaging (JEI)· Vol 35, pp. 043009 - 043009· 0 citations· 54 references
Engineering
TL;DR
A layer-wise uncertainty-guided feature perturbation and refinement framework that operates directly in representation space and can be seamlessly integrated into pretrained transformer encoders is proposed, validating the effectiveness of uncertainty-guided representation diffusion for medical visual understanding.
Abstract
Abstract. Vision Transformers pretrained via self-supervised learning have demonstrated strong representation capability in natural image analysis and are increasingly adopted in medical imaging tasks. However, when transferred to medical domains, pretrained encoders often exhibit limited adaptability to ambiguous anatomical boundaries and low-contrast structures as contextual dependencies learned from natural images may not adequately capture uncertainty characteristics inherent in medical data. In this work, we propose a layer-wise uncertainty-guided feature perturbation and refinement framework that operates directly in representation space and can be seamlessly integrated into pretrained transformer encoders. The proposed method explicitly estimates spatial uncertainty from encoder features and performs controlled semantic diffusion in feature space, enabling selective refinement of ambiguous regions while maintaining stable representations elsewhere. The enhanced representations are integrated through residual modulation, enabling progressive adaptation without disrupting pretrained dynamics. The proposed framework is fully plug-and-play and can be inserted into each transformer layer without modifying the original architecture. Extensive experiments on multiple medical image analysis tasks demonstrate consistent performance improvements over strong transformer baselines, validating the effectiveness of uncertainty-guided representation diffusion for medical visual understanding.
An efficient diffusion framework that jointly diffuses a baseline scan and its follow-up residual, summed to synthesize the follow-up scan, while concurrently predicting a spatial uncertainty map, in a single reverse diffusion process is proposed.
A. Oliveras, Roger Marí, Rafael Redondo et al.· 1 citation· ⚡1
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...
Position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors, is proposed and implemented in EmbedVision, an interactive 3D Slicer-based workflow, and evaluated across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors.
A. Jamzad, Dilakshan Srikanthan, F. Akbarifar et al.· 0 citations
A unified framework is presented that simultaneously models segmentation variability and expert‐specific behavior within a single architecture, enabling personalized predictions while preserving diversity and demonstrating consistent improvements over existing methods in both diversity and personalization metrics.
A. Gharawi, M. Alahmadi· International journal of ima...· 0 citations
This work proposes Adaptive Feature Fusion U-Net (AFFUNet), an adaptive feature fusion (AFF) Transformer U-Net with a joint loss function, aiming to improve segmentation accuracy through dynamic feature fusion, hard sample reweighting, and explicit boundary optimization.