Skip to content
Open access

MMViT: Bridging Mamba and Attention for efficient video action recognition in sports.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

This paper proposes MMViT (Multi-scale Mamba Visual Transformer) with a hierarchical design for improved recognition and objective efficiency and introduces a computation-downsampling decoupling (CDD) mechanism to preserve feature coverage during Mamba spatial scaling change.

Abstract

In sports AI, human action recognition (HAR) faces a challenge between the expensive Transformer and the one-dimensional state space models (SSMs). Although Transformer has proven success on video tasks, its high computational cost scales quadratically. In contrast, conventional SSMs like Mamba possess linear complexity, but also underperform in the recognition. In this paper, we propose MMViT (Multi-scale Mamba Visual Transformer) with a hierarchical design for improved recognition and objective efficiency. We employ a heterogeneous "Attention-Mamba-Attention" (A-M-A) strategy. It first uses Multi-scale Pooling Attention (MPA) for efficient capture of local spatial feature. As computation-heavy stages come, it transitions to Mamba module with linear complexity to efficiently model long-range temporal context. Finally, attention is re-introduced at latter stages for semantic feature fusion. Also, we introduce a computation-downsampling decoupling (CDD) mechanism to preserve feature coverage during Mamba spatial scaling change. We have validated our approach on SpaceJam and Basketball-51 datasets. Experiments show that MMViT achieves superior performance over strong baselines with substantial margins. Ablation studies show the significance of A-M-A, MPA and CDD. MMViT achieves competitive accuracy among evaluated models and provides a favorable accuracy-efficiency trade-off for video action recognition task.

Read PDF

Similar papers

Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Jul 2026

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

It is proved that the computationally cheaper split space-time attention is equivalent to full space-time attention and is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.

N. Tran, Fanghui Xue, Shuai Zhang et al. · 0 citations
Open access Jul 2026

Deep Features Evaluation Method of Human Action Recognition Based on Convolutional Neural Network

A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline a...

H. Khan, Altaf Hussain · 0 citations
Conference Jul 2026

AVT-PAC: A Pipeline for Multimodal Action Prediction and Captioning

Action prediction from frames and videos is a well-studied problem. Models trained with a single modality, mostly vision, will fail in low-light conditions. Recent works have attempted to predict action categories using vision-language and audio-visual models. A challenge, however, is that some dataset annotations lack...

A. R, Ambarish Parthasarathy, Sucharitha Devarakonda et al. · 0 citations
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.