HBTMNet: A Tri-Hybrid Boundary-Aware Transformer Mamba Network for Precise Skin Lesion Segmentation
Abstract
Accurate automated segmentation of melanoma from dermoscopic images remains a critical challenge in computer-aided diagnosis pipelines due to high intra-class diversity, fuzzy lesion boundaries, and structural artifacts. Conventional deep learning frameworks, including convolutional neural networks, vision transformers, and standard hybrid models, struggle to reconcile localized textural feature extraction, long-range global context modelling, and sharp edge delineation. We propose HBTMNet, a Tri-Hybrid Boundary-Aware Transformer Mamba Network unified within an end-to-end encoderdecoder paradigm. The architecture integrates a linear-complexity Swin-UMamba backbone using 2D selective-scan visual state-space blocks to capture global context without the quadratic memory overhead of self-attention. This is coupled with an encoder-side Global-Local Refinement block that implicitly models feature-level boundary uncertainty, alongside an explicit edge-guided pipeline using Difference Attention and Edge Generation Modules driven by offline Canny-derived geometric priors. Evaluated across four benchmark datasets – ISIC 2016, ISIC 2017, ISIC 2018, and ISIC 2019 – HBTMNet outperforms existing baselines. On ISIC 2017, it achieves a Dice similarity coefficient of 94.71% and an Intersection-over-Union of 93.48%, while substantially reducing boundary outlier errors as measured by the 95thpercentile Hausdorff Distance. These results support combining implicit, distribution-based boundary detection with explicit geometric edge guidance as an effective strategy for clinically reliable lesion segmentation