A novel segmentation network, MDSF-Net, consisting of two key components: the Spatial-Mamba-Frequency Feature Block (SMFB) and the Hierarchical Semantic Link (HSL) module, which outperforms current state-of-the-art segmentation methods across key quantitative metrics.
In medical image segmentation, accurately identifying anatomical structure boundaries is a core task for clinical computer-aided diagnosis. Although vision foundation models, represented by the Segment Anything Model (SAM), have demonstrated strong generalization potential, their deployment in fully automated clinical scenarios remains constrained by severe domain shift, reliance on manual prompts, and significant boundary degradation in ambiguous anatomical transition zones. To address these challenges, this paper proposes GeoFuse-SAM, a multimodal geometric data fusion adaptation framework.
The core innovation of this framework lies in breaking the limitations of single semantic features. Through a Geometry-Guided Cross-Domain Attention Fusion (GCAF) module, it achieves deep data fusion between the raw image data and high-frequency geometric priors (Sobel gradient fields) derived from computer graphics. This cross-modal interaction mechanism provides explicit spatial guidance to the model, significantly enhancing the robustness of boundary recognition. Furthermore, we introduce a lightweight Parallel Fusion Adapter (PFA) to achieve medical semantic alignment, and propose a Parameter-Free Morphological Boundary Weighting (PMBW) strategy. This strategy utilizes morphological operators to pinpoint ambiguous boundary regions during the training phase and impose dynamic geometric constraints.
Experiments on two challenging medical datasets, BUSI and ISIC 2018, demonstrate that GeoFuse-SAM, operating in a fully automatic prompt-free mode, not only maintains leading region segmentation accuracy (achieving a Dice score of 89.85\% on ISIC 2018), but also effectively suppresses the boundary degradation phenomenon of foundation models in grayscale modalities on the core boundary metric HD95 (optimized to 15.97 on BUSI). Without introducing extra inference parameters, it exhibits superior edge fidelity compared to existing medical fine-tuned foundation models (such as SAM-Med2D). This study provides a high-fidelity, low-cost technical paradigm for robust medical image segmentation.
Qi-Yuan Wang· Poster Volume 0007 The 2026...· 0 citations
DIG-MambaNet achieves consistent and competitive performance across diverse target structures and imaging conditions, with improved boundary delineation and favorable overlap-based accuracy compared with representative CNN-, Transformer-, and Mamba-based methods.
Yongkang Zhu, Tianyu Yu, Hongmei Li et al.· Journal of Imaging· 0 citations
MGA-UNet, a frequency-aware multi-scale encoder–decoder segmentation framework that integrates wavelet-based frequency decomposition with Mamba-based long-range dependency modelling, demonstrates that frequency-domain decomposition and state-space modelling can complement each other for accurate medical image segmentation.
Shuai-Kang Qiu, Xuan Wang, Kai-Le Su et al.· Italian National Conference...· 0 citations
Multi-scale feature fusion is a cornerstone of encoder-decoder architectures in medical image segmentation, yet effectively integrating representations across stages remains a significant challenge due to the inherent semantic–spatial gap. Deep features encode abstract semantic context but lack spatial precision, whereas early-stage features preserve fine-grained details but suffer from limited semantic discriminability. Existing fusion mechanisms, which often rely on symmetric aggregation or simple skip connections, fail to explicitly model the semantic-to-spatial guidance necessary for precise alignment. To address this, we propose a Tri-stream Prototype Fusion Network (TSPFusion) that introduces three key innovations: (i) a tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage; (ii) a Global Prototype Bank (GPB) that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing persistent semantic priors across images; and (iii) a Detail–Semantic Feature Aligner (DSFA) that performs semantic-guided refinement of spatial features prior to fusion, preventing feature interference from direct concatenation. Additionally, an Adaptive Pyramid Context Decoder module aggregates multi-scale information with resolution-aware dynamic pooling, and a Gradient-Gated Spatial Attention head enforces boundary-sensitive structural consistency. Extensive experiments on four medical imaging benchmarks (CT and Ultrasound) demonstrate that TSPFusion achieves state-of-the-art performance 97.82±0.85% DSC on COVID19 lung CT, 81.76% DSC on COVID19-Seg, and 88.16% mDice on cross-dataset BUSI→\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\rightarrow $$\end{document}STU, while maintaining a compact 5.92M parameter footprint.
Mohammed A. M. Elhassan, Qian-Fa Yuan, Zhizhong Xu et al.· Journal of King Saud Univers...· 0 citations
Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration independently. This may cause semantic mismatch between global context and local boundaries, insufficient channel relationship modeling, weak spatial-channel interaction, and redundant representations. We propose CDGC-Net, a 3D medical image segmentation network that combines cooperative dual-scale spatial attention with grouped hierarchical channel modeling. With-in each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads to parallel local-window and global-sparse branches. The two branches capture fine spatial details and long-range anatomical context at the same feature level. Their outputs are concatenated into an $N\times C$ spatial representation and directly passed to Grouped Hierarchical Channel Attention (GHCA). GHCA organizes the channels into $r$ groups and models both within-group and cross-group dependencies. CDSA and GHCA reuse a shared key projection to maintain a consistent feature reference. Residual feature alignment subsequently integrates the refined features with the original representation. On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values of 86.96\%, 92.91\%, 82.56\%, and 93.52\%, respectively, exceeding the next-highest reported values by 0.39, 0.47, 0.17, and 0.32 percentage points. CDGC-Net contains 25.83M parameters and 28.62G FLOPs for an input size of $64\times128\times128$, reducing these quantities by 39.87\% and 40.30\%, respectively, relative to UNETR++. These results indicate a favorable trade-off between segmentation accuracy and computational complexity.
Zhe-Yang Jing, Qin Lu, Jianwang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.