Skip to content
Open access

EP-Net: An Equipment-Guided Dual-Stream CNN–Transformer Network for Fine-Grained Sports Image Classification

Aug 2026 · Applied Sciences · Vol 16, pp. 8229 · 0 citations · 10 references

TL;DR

The Equipment-Primed Network (EP-Net) is proposed, a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination and a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration.

Abstract

Fine-grained sports recognition from static images is challenging because visually similar sports often exhibit nearly identical human poses, whereas their decisive differences are encoded by small-scale equipment and subtle human–equipment interactions. Existing single-stream convolutional or Transformer-based models tend to emphasize either local appearance or global context, making them vulnerable to equipment-detail loss and background interference. To address this problem, we propose the Equipment-Primed Network (EP-Net), a heterogeneous dual-stream architecture that treats sports equipment as a primary semantic cue for action discrimination. EP-Net employs an EfficientNetV2-S-based Equipment Stream to capture localized equipment shapes and textures and a Swin-Tiny-based Behavior Stream to model the athlete’s spatial configuration and global scene context. We further introduce a Cross-Modal Channel Attention (CMCA) module that projects equipment features into the behavior-feature space and performs directional channel recalibration. Unlike simple feature concatenation, CMCA uses equipment information to enhance action-relevant channels while reducing the relative influence of background-dominated responses. Experiments on the Sports-100 dataset show that EP-Net achieves a Top-1 accuracy of 98.80%, outperforming EfficientNetV2-S and Swin-Tiny by 3.00 and 2.35 percentage points, respectively. It also improves on naive dual-stream concatenation by 0.88 percentage points. Grad-CAM visualizations further indicate that EP-Net attends more consistently to discriminative equipment and human–equipment interaction regions. These results suggest that equipment-guided local–global feature interaction provides an effective solution to pose ambiguity and background interference in static fine-grained sports recognition.

Read PDF

Similar papers

Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Conference Aug 2026

MASF-Net: efficient linear attention guided few-shot fine-grained image recognition

Few-shot fine-grained image classification (FS-FGIC) aims to distinguish visually similar subcategories with only a handful of labeled examples, posing significant challenges due to subtle inter-class differences and large intra-class variations. Existing methods often fail to fully leverage complementary information f...

Jinyu Wang, Bing-Xin Xu, Weiguo Pan et al. · 0 citations
Preprint Aug 2026

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.

Imtiaz Ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al. · 0 citations
Preprint Aug 2026

ConCA: Concentration-Aware Channel Attention for Fine-Grained Visual Recognition

ConCA is proposed, which pairs the mean with a shift-invariant negative-input entropy (NegEnt), computed via a softmax over the negated activations, forming a dual descriptor that jointly encodes magnitude and concentration.

Yushi Liu, Yu-Chen Tung · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.