DIG-MambaNet achieves consistent and competitive performance across diverse target structures and imaging conditions, with improved boundary delineation and favorable overlap-based accuracy compared with representative CNN-, Transformer-, and Mamba-based methods.
Abstract
Reliable medical image segmentation remains challenging because models must preserve fine boundary details while maintaining global semantic consistency. CNNs capture local structures effectively but have limited long-range modeling ability, whereas Transformer-based methods improve global context at high computational cost. Mamba-based state space models offer efficient long-range modeling, but may weaken high-frequency textures and boundary cues. To address these limitations, we propose DIG-MambaNet, a Dual-path Interactive Guided Mamba Network for medical image segmentation. The network introduces a dual-path complementary modeling block (DCM Block), where a cross-feature spatial interaction module (CSIM) adaptively integrates CNN-based local features and Mamba-based global features. A source image-guided module (SIGM) injects high-frequency information from the original image to compensate for downsampling-induced detail loss, while an inter-layer detail refinement fusion module (IDRFM) improves encoder–decoder feature alignment during reconstruction. Experiments on 2018DSB, ISIC2018, JSUAH-Cerebellum, and CVC-ClinicDB, covering nuclei segmentation in microscopy images, skin lesion segmentation in dermoscopic images, fetal cerebellum segmentation in ultrasound images, and polyp segmentation in colonoscopy images, demonstrate that DIG-MambaNet achieves consistent and competitive performance across diverse target structures and imaging conditions, with improved boundary delineation and favorable overlap-based accuracy compared with representative CNN-, Transformer-, and Mamba-based methods.
MGA-UNet, a frequency-aware multi-scale encoder–decoder segmentation framework that integrates wavelet-based frequency decomposition with Mamba-based long-range dependency modelling, demonstrates that frequency-domain decomposition and state-space modelling can complement each other for accurate medical image segmentation.
Shuai-Kang Qiu, Xuan Wang, Kai-Le Su et al.· Italian National Conference...· 0 citations
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
TvaraNet is pro-posed, an extremely lightweight segmentation network designed to preserve boundary fidelity under strict efficiency constraints and achieves competitive or superior boundary-aware performance compared to heavier architectures.
Sridhatta Jayaram Aithal, Vandana Bharti· Proceedings of the Thirty-Fi...· 0 citations
Results indicate the WVM-UNet architecture effectively captures discriminative features for precise medical image segmentation, and demonstrates the competitive performance of the method on multiple public datasets.
Yulong Yang, Wenchao Gao, Zheng-Guo Wu et al.· Journal of Imaging· 0 citations
Medical image segmentation requires computational methods that accurately capture global context, local boundaries, and multiscale anatomical structures while remaining reproducible across different applications. This article presents a protocol for constructing, training, and evaluating a Multi-view Vision Mamba U-Shaped Network framework for two-dimensional medical image segmentation. The protocol provides a reproducible workflow that includes public dataset acquisition, image and mask preprocessing, network construction, model training, checkpoint selection, and quantitative and qualitative performance evaluation. The framework incorporates multiview feature scanning to capture complementary spatial, contour, scale, and boundary information and applies multistage feature fusion within a U-shaped encoder-decoder architecture to improve feature integration during segmentation. The protocol is demonstrated using publicly available skin lesion and abdominal organ segmentation datasets. Under the described implementation workflow, the framework achieves competitive segmentation performance using standard evaluation metrics. By following the procedures presented in this protocol, researchers can reproduce the model implementation, train the network using defined experimental settings, evaluate segmentation performance, and adapt the workflow for related medical image segmentation tasks requiring reproducible deep learning-based analysis.
Sufen Guo, Xueguang Li· Journal of Visualized Experi...· 0 citations
Existing medical image segmentation (MedISeg) models predominantly rely on convolutional neural networks (CNNs) and Transformer architectures. However, the limited receptive fields of CNNs and the quadratic computational cost of Transformers hinder their scalability and efficiency. Recently, receptance-weighted key-value (RWKV) has emerged as a promising linear-complexity alternative for global context modeling. In this article, we propose SRWKV, a shape-guided RWKV (SGR) model for parameter-efficient MedISeg. SRWKV introduces an SGR block that uses a shape prior predicted from the deepest encoder feature to guide token traversal during decoding, reducing foreground-background interleaving and improving structural coherence during sequence formation. In addition, we develop a deformable adaptive shift (DA-Shift) module that dynamically adjusts token interactions according to local context, enabling flexible receptive field adaptation for diverse anatomical structures. Extensive experiments across six MedISeg tasks on 11 datasets demonstrate that SRWKV achieves strong segmentation performance with a compact parameter footprint. Our code is available at https://github.com/ukeLin/SRWKV.
Chun-Li Yu, Yin-Hao Li, Zheng Zhao et al.· IEEE Transactions on Neural...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.