Skip to content
Open access

Efficient Interactive Image Segmentation Using Multilayer Perceptron

2026 · IEEE Access · Vol 14, pp. 127017-127028 · 0 citations · 59 references
Computer Science

TL;DR

A novel IIS framework based on Multi-Layer Perceptron (MLP), which fuses clicks from users with feature representations extracted from the Transformer Encoder, which closes the gap between interactive user instructions and deep learning and provides a practical and flexible solution for accurate segmentation in many applications.

Abstract

Interactive Image Segmentation (IIS) is important for applications that require precise segmentation with minimal user intervention, for example, in medical imaging, 3D object segmentation and video editing. However, the existing IIS methods are often hampered by ambiguity of interaction, leading to suboptimal performance of segmentation. We propose a novel IIS framework based on Multi-Layer Perceptron (MLP), which fuses clicks from users with feature representations extracted from the Transformer Encoder. We enhance the segmentation accuracy by using the adaptive click-based refinement and optimise with focal loss to address class imbalance. Extensive experiments on benchmark datasets including GrabCut, Berkeley, SBD, DAVIS and Pascal VOC show that our method outperforms state-of-the-art techniques in terms of segmentation accuracy and robustness. The proposed method achieves lower NoC@85 and NoC@90 values across several benchmark datasets, with particularly strong performance on GrabCut and BDD100K. We use a Vision Transformer-based architecture with more advanced training methods to improve feature extraction and boundary delineation. This guarantees reliable convergence of the model and good generalisation. Qualitative and quantitative evaluations validate that the framework is effective in real-world segmentation tasks. Our method closes the gap between interactive user instructions and deep learning and provides a practical and flexible solution for accurate segmentation in many applications. The proposed framework differs from existing transformer-based IIS methods by incorporating a Vision Transformer encoder and a lightweight MLP segmentation head and by using normalised focal-loss optimisation to enhance the interaction efficiency while maintaining high segmentation accuracy. The source code of the proposed approach is available at https://github.com/abhigoku10/IIS.git for reproducibility and further research.

Read PDF

Similar papers

Open access Aug 2026

A similarity-aware network with contrastive optimization for biomedical image segmentation

Accurate biomedical image segmentation is crucial for clinical diagnosis. Convolutional neural networks and Transformer-based models have been widely used for biomedical image segmentation and have improved segmentation accuracy across multiple imaging modalities. However, most existing methods rely heavily on supervis...

Rongjia Lin, Zhidong Yang, Zi-Heng Xu et al. · 0 citations
Conference Open access 2026

Medical Image Segmentation Based on the CLIP and SAM Base Models

. Medical image segmentation is a key technology for achieving precise medical care. Traditional methods rely on a large amount of labeled data and have limited generalization capabilities. Visual foundation models represented by CLIP and SAM have obtained general capabilities through large-scale pre-training, providin...

Jing-Yi Yu · 0 citations
Review

Computer Vision and Image Understanding

This survey provides a comprehensive evaluation of various deep learning-based segmentation architectures, covering a wide range of models, from traditional ones like FCN and PSPNet to more modern approaches like SegFormer and FAN, and proposes to evaluate the methods in terms of temporal consistency and corruption vul...

Ronny Velastegui, Maxim Tatarchenko, Sezer Karaoglu et al. · 0 citations
Open access Aug 2026

Adaptive wavelet enhancement and stratified feature fusion for robust medical image segmentation

Accurate medical image segmentation plays an important role in ultrasound triage, colonoscopy screening, and histopathology analysis. However, segmentation performance often degrades substantially when input images are noisy or low-contrast and the model must operate on resource-constrained devices. This challenge main...

Peiyuan Wang, Yong-Jie Liang, Bizhong Wei et al. · 0 citations
Open access Sep 2026

Improving Image Segmentation Accuracy Using Dynamic Performance Analysis of the GrabCut Algorithm

. Abstract In this research, a computer vision model was developed that enables object detection, shape identification, and orientation prediction using image segmentation. The GrabCut algorithm was used to segment images. GrabCut was initialized in two ways: first, by initializing the input image with a bounding box (...

M. K. Hussein, N. M. Almoosawi, Haider M. Al-Mashhadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.