MSHA-Net: learning with multiscale hierarchical discriminative attention for few-shot medical segmentation
Abstract
Few-shot medical image segmentation, which aims to achieve accurate segmentation with limited labeled data, faces challenges including 1) the model’s need for strong generalization ability and 2) simultaneous extraction of global semantics and local boundaries. To address these issues, we propose MSHA-NET, a framework based on Multi-Scale Hierarchical Discriminative Attention for few-shot medical segmentation. MSHA-Net consists of three core components: a Triple-Branch Feature Extraction Module (TBFEM), a Two-Stream Attention Module (TSAM), and a Threshold-based Multi-scale Fusion Module (TMFM). First, TBFEM effectively extracts global semantics and local boundary details across multiple scales from support and query images. Then, TSAM harnesses self-attention and cross-attention mechanisms to adequately enhance and fuse multi-scale support and query features. Finally, via the features containing both global semantics and boundary details, TMFM adaptively generates and then integrates prediction masks from different scales to get the final segmentation map. To train MSHA-NET, super voxel-based self-supervised training process based on generating auxiliary data is developed to enhance the model’s generalization ability under data-scarce conditions. Unlike existing single-scale attention-based few-shot segmentation methods that struggle to balance global semantics and local boundary details, MSHA-Net introduces a hierarchical discriminative attention framework that explicitly decouples and fuses multi-scale features through three specialized modules. Experiments on the public ABD-CT and ABD-MR datasets show that MSHA-Net surpasses state-of-the-art methods by up to 10.62% in Dice scores, demonstrating its superior ability to handle diverse organ sizes and cross-modality generalization.