Aug 2026· Multimedia Systems· Vol 32· 0 citations· 57 references
TL;DR
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
DIG-MambaNet achieves consistent and competitive performance across diverse target structures and imaging conditions, with improved boundary delineation and favorable overlap-based accuracy compared with representative CNN-, Transformer-, and Mamba-based methods.
Yongkang Zhu, Tianyu Yu, Hongmei Li et al.· Journal of Imaging· 0 citations
MGA-UNet, a frequency-aware multi-scale encoder–decoder segmentation framework that integrates wavelet-based frequency decomposition with Mamba-based long-range dependency modelling, demonstrates that frequency-domain decomposition and state-space modelling can complement each other for accurate medical image segmentation.
Shuai-Kang Qiu, Xuan Wang, Kai-Le Su et al.· Italian National Conference...· 0 citations
Polyp segmentation in colonoscopy images plays a pivotal role in computer-aided medical diagnosis and the early prevention of colorectal cancer. However, existing methods often suffer from performance degradation when confronted with extreme polyp scale variation and polyp boundary ambiguity. To address these challenges, we propose the Staged Global-to-Local Cross-Scale Fusion Network (SGLF-Net), which adopts a novel staged global-to-local learning paradigm to progressively refine segmentation from coarse global semantics to fine-grained local details. Specifically, the Global Semantic Perception Stage integrates a Swin Transformer Encoder and a Dynamic Attentive Decoder (DAD) to construct comprehensive multi-scale contextual representations. The Local Detail Refinement Stage employs an Edge-aware Dynamic Attentive Decoder (E-DAD) to enhance structural fidelity and boundary precision through explicit edge-guided supervision. Furthermore, we introduce the Cross Spatial-Scale Feature Aggregation and Reconstitution (CSSAR) module, equipped with hybrid attention mechanisms, to facilitate efficient semantic structural interaction between the two cascaded stages. Extensive experiments on five public benchmark datasets demonstrate that SGLF-Net consistently outperforms state-of-the-art methods in both segmentation accuracy and boundary preservation.
Tan Guo, Wen-Han Zhang, Fu-Lin Luo et al.· IEEE journal of biomedical a...· 0 citations
Medical image segmentation is a key component of computer-aided diagnosis and treatment planning. Despite substantial progress in deep learning–based models, most existing approaches depend heavily on large annotated datasets and often fail to generalize across heterogeneous clinical environments, limiting their deployment in real-world settings characterized by domain shifts and scarce expert annotations. This paper presents a zero-shot learning framework named GroundMed-SAM for medical image segmentation. The framework integrates GroundingDINO for prompt-based region localization and MedSAM for mask generation. To address the weak alignment between visual features and medical semantics in GroundingDINO, which is pretrained on general domain image-text pairs, we introduce learnable medical text embeddings that explicitly parameterize domain-specific terminology in a continuous semantic space. These embeddings are optimized during training to better align medical concepts with visual representations, thereby strengthening text-image correspondence and improving detection-guided segmentation. The proposed framework preserves true zero-shot capability, enabling segmentation of previously unseen anatomical structures without task-specific labels. Extensive experiments on multiple public datasets across diverse modalities and clinical contexts demonstrate that our method achieves competitive segmentation performance in-domain while exhibiting superior robustness under cross-domain evaluation. Although supervised baselines outperform the proposed framework by only 3–5% on in-domain datasets, they experience substantial performance degradation when evaluated on unseen domains. Additionally, the framework achieves an AUC of 98.9 in endoscopic polyp detection, highlighting the effectiveness of the proposed medical-aware textual embeddings in guiding region localization. These results demonstrate the effectiveness of the proposed framework in improving cross-domain generalization for medical image segmentation with limited annotations.
V. Nguyen, Hoang Quan Luong, Phuc Ngoc Pham· IEEE International Conferenc...· 0 citations
A Parallel Mamba Dual-U Network that adopts a cascaded dual U-Net encoder-decoder for two-stage “coarse-to-fine” segmentation refinement and outperforms state-of-the-art methods on most core metrics, verifying its effectiveness and robustness for complex medical image segmentation.
Shao-Qiang Wang, Linhao Zhang, Guiling Shi et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.