Skip to content
Conference

ResUNet++ -ViT: A Fusion Framework for Accurate High-Resolution Medical Image Segmentation

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1-8 · 0 citations · 15 references

Abstract

Segmentation of medical images is a crucial process for diagnosis and treatment planning. Yet, traditional CNN based models are often not able to convey high-level spatial relationships and intricate tissue boundaries in high-resolution medical images. To tackle these issues, this paper presents a novel ResUNet++–Vision Transformer (ResUNet++–ViT) framework incorporating both multi-scale local feature extraction and global contextual learning. The ResUNet++ backbone consists of residual blocks and nested skip connections, which are used to extract hierarchical features, and the Vision Transformer uses self-attention mechanisms to capture long-range dependencies. A fusion module allows for the integration of local and global features, resulting in better segmentation accuracy and preserving the boundary. The proposed model was tested on the ISBI 2012 Electron Microscopy Segmentation Challenge (EMSC) dataset, and obtained a Dice score of 0.960, IoU of 0.920, precision of 0.967, recall of 0.958, accuracy of 0.984 and Hausdorff distance of 2.76. The proposed framework is compared with FCN, UNet, Attention UNet, ResUNet, UNet++, Vision Transformer and TransUNet, and the results show its superiority. The results show that the combination of ResUNet++ and Vision Transformers greatly enhances the performance of segmentation, boundary delineation, and generalization in the field of advanced medical image analysis applications.

View source

Similar papers

Aug 2026

GLNet: global-to-local aware hybrid framework for medical image segmentation

GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.

Dengdi Sun, Longlong Liu, Xiao-Wei Zhao et al. · 0 citations
Open access 2026

Hybrid U-Net++–Vision Transformer Fusion for Accurate Brain MRI Image Segmentation and classifications

Automatic brain Magnetic Resonance Imaging (MRI) analysis plays a crucial role in computer-aided diagnosis, treatment planning, and disease monitoring by enabling accurate delineation and identification of brain tumors. However, manual segmentation is labor-intensive, time-consuming, and susceptible to inter-observer v...

Malathi Janapati, Shaheda Akthar · 0 citations
Open access Aug 2026

WVM-UNet: A Wavelet–Vision Mamba Framework for Enhanced Medical Image Segmentation

Results indicate the WVM-UNet architecture effectively captures discriminative features for precise medical image segmentation, and demonstrates the competitive performance of the method on multiple public datasets.

Yulong Yang, Wenchao Gao, Zheng-Guo Wu et al. · 0 citations
Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Sep 2026

LMDAU-Net: An Effective Lightweight Multi-scale Deformation Aggregation U-Net for Skin Lesion Segmentation.

Automatic skin lesion segmentation is a pivotal problem in the medical domain and an indispensable component in the computer-aided diagnosis program. Most convolutional neural network-based segmentation algorithms have demonstrated promising performance due to their ability to encode detail and semantic features effici...

Jun-Han Hu, Ming Liu, Jing Yang et al. · 0 citations
Open access Aug 2026

Integrating state space models and attention mechanisms for brain tumor segmentation in MRI

Brain tumor segmentation from MRI is clinically critical yet challenging due to heterogeneous appearance and irregular boundaries. Conventional CNN based methods lack effective global context modeling, while transformer-based approaches are computationally expensive and unstable on limited datasets. To address these,...

Saritha Saladi, Riyaz Hussain Shaik, Abhi Chevuri et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.