A novel IIS framework based on Multi-Layer Perceptron (MLP), which fuses clicks from users with feature representations extracted from the Transformer Encoder, which closes the gap between interactive user instructions and deep learning and provides a practical and flexible solution for accurate segmentation in many applications.
Abstract
Interactive Image Segmentation (IIS) is important for applications that require precise segmentation with minimal user intervention, for example, in medical imaging, 3D object segmentation and video editing. However, the existing IIS methods are often hampered by ambiguity of interaction, leading to suboptimal performance of segmentation. We propose a novel IIS framework based on Multi-Layer Perceptron (MLP), which fuses clicks from users with feature representations extracted from the Transformer Encoder. We enhance the segmentation accuracy by using the adaptive click-based refinement and optimise with focal loss to address class imbalance. Extensive experiments on benchmark datasets including GrabCut, Berkeley, SBD, DAVIS and Pascal VOC show that our method outperforms state-of-the-art techniques in terms of segmentation accuracy and robustness. The proposed method achieves lower NoC@85 and NoC@90 values across several benchmark datasets, with particularly strong performance on GrabCut and BDD100K. We use a Vision Transformer-based architecture with more advanced training methods to improve feature extraction and boundary delineation. This guarantees reliable convergence of the model and good generalisation. Qualitative and quantitative evaluations validate that the framework is effective in real-world segmentation tasks. Our method closes the gap between interactive user instructions and deep learning and provides a practical and flexible solution for accurate segmentation in many applications. The proposed framework differs from existing transformer-based IIS methods by incorporating a Vision Transformer encoder and a lightweight MLP segmentation head and by using normalised focal-loss optimisation to enhance the interaction efficiency while maintaining high segmentation accuracy. The source code of the proposed approach is available at https://github.com/abhigoku10/IIS.git for reproducibility and further research.
Accurate biomedical image segmentation is crucial for clinical diagnosis. Convolutional neural networks and Transformer-based models have been widely used for biomedical image segmentation and have improved segmentation accuracy across multiple imaging modalities. However, most existing methods rely heavily on supervis...
Rongjia Lin, Zhidong Yang, Zi-Heng Xu et al.· BMC Medical Imaging· 0 citations
. Medical image segmentation is a key technology for achieving precise medical care. Traditional methods rely on a large amount of labeled data and have limited generalization capabilities. Visual foundation models represented by CLIP and SAM have obtained general capabilities through large-scale pre-training, providin...
Jing-Yi Yu· Proceedings of the 4th Inter...· 0 citations
This survey provides a comprehensive evaluation of various deep learning-based segmentation architectures, covering a wide range of models, from traditional ones like FCN and PSPNet to more modern approaches like SegFormer and FAN, and proposes to evaluate the methods in terms of temporal consistency and corruption vul...
Ronny Velastegui, Maxim Tatarchenko, Sezer Karaoglu et al.· 0 citations
Accurate medical image segmentation plays an important role in ultrasound triage, colonoscopy screening, and histopathology analysis. However, segmentation performance often degrades substantially when input images are noisy or low-contrast and the model must operate on resource-constrained devices. This challenge main...
Peiyuan Wang, Yong-Jie Liang, Bizhong Wei et al.· Journal of King Saud Univers...· 0 citations
. Abstract In this research, a computer vision model was developed that enables object detection, shape identification, and orientation prediction using image segmentation. The GrabCut algorithm was used to segment images. GrabCut was initialized in two ways: first, by initializing the input image with a bounding box (...
M. K. Hussein, N. M. Almoosawi, Haider M. Al-Mashhadi· EEPES 2026· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.