Aug 2026· Applied Sciences· Vol 16, pp. 8229· 0 citations· 10 references
TL;DR
The Equipment-Primed Network (EP-Net) is proposed, a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination and a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration.
Abstract
Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance or global context, making them vulnerable to equipment-detail loss and background interference. To address this problem, we propose the Equipment-Primed Network (EP-Net), a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination. EP-Net employs an EfficientNetV2-S-based Equipment Stream to capture localized equipment shapes and textures and a Swin-Tiny-based Behavior Stream to model the athlete’s spatial configuration and global scene context. We further introduce a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration. Unlike simple feature concatenation, CMCA uses equipment information to enhance action-relevant channels while reducing the relative influence of background-dominated responses. Experiments on the Sports-100 dataset show that EP-Net achieves a Top-1 accuracy of 98.80%, outperforming EfficientNetV2-S and Swin-Tiny by 3.00 and 2.35 percentage points, respectively. It also improves on naive dual-stream concatenation by 0.88 percentage points. Grad-CAM visualizations further indicate that EP-Net attends more consistently to discriminative equipment and human–equipment interaction regions. These results suggest that equipment-guided local–global feature interaction provides an effective solution to pose ambiguity and background interference in static fine-grained sports recognition.
CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al.· IEEE Internet of Things Jour...· 0 citations
Few-shot fine-grained image classification (FS-FGIC) aims to distinguish visually similar subcategories with only a handful of labeled examples, posing significant challenges due to subtle inter-class differences and large intra-class variations. Existing methods often fail to fully leverage complementary information f...
Jinyu Wang, Bing-Xin Xu, Weiguo Pan et al.· International Conference on...· 0 citations
FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.
Imtiaz Ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al.· 0 citations
ConCA is proposed, which pairs the mean with a shift-invariant negative-input entropy (NegEnt), computed via a softmax over the negated activations, forming a dual descriptor that jointly encodes magnitude and concentration.
The results indicate that the proposed GAESA-iFormer model achieves improved prediction accuracy on the current compressor dataset under limited experimental data.