MASF-Net: efficient linear attention guided few-shot fine-grained image recognition
Abstract
Few-shot fine-grained image classification (FS-FGIC) aims to distinguish visually similar subcategories with only a handful of labeled examples, posing significant challenges due to subtle inter-class differences and large intra-class variations. Existing methods often fail to fully leverage complementary information from different network hierarchies or lack efficient mechanisms to focus on discriminative regions. To address these issues, we propose a novel Multi-scale Attention-based Similarity Fusion Network (MASF-Net). Our approach introduces three key components: (1) a Feature Enhancement Transformer (FET) module that performs multi-scale feature extraction and cross-scale interaction via linear attention; (2) a Cross-attention Relation Learner (CRL) module that models fine-grained semantic relationships between support and query sets with an efficient linear attention mechanism; and (3) a Multi-Scale Fusion Classifier (MSFC) module that adaptively integrates multi-level features for final classification. Extensive experiments on three challenging fine-grained benchmarks—CUB-200-2011, Stanford-Dogs, and Stanford-Cars—demonstrate that MASFNet consistently outperforms state-of-the-art methods in both 5-way 1-shot and 5-way 5-shot settings. Visualizations illustrate the model's ability to focus on discriminative regions and establish accurate semantic correspondences.