This paper proposes MMViT (Multi-scale Mamba Visual Transformer) with a hierarchical design for improved recognition and objective efficiency and introduces a computation-downsampling decoupling (CDD) mechanism to preserve feature coverage during Mamba spatial scaling change.
Abstract
In sports AI, human action recognition (HAR) faces a challenge between the expensive Transformer and the one-dimensional state space models (SSMs). Although Transformer has proven success on video tasks, its high computational cost scales quadratically. In contrast, conventional SSMs like Mamba possess linear complexity, but also underperform in the recognition. In this paper, we propose MMViT (Multi-scale Mamba Visual Transformer) with a hierarchical design for improved recognition and objective efficiency. We employ a heterogeneous "Attention-Mamba-Attention" (A-M-A) strategy. It first uses Multi-scale Pooling Attention (MPA) for efficient capture of local spatial feature. As computation-heavy stages come, it transitions to Mamba module with linear complexity to efficiently model long-range temporal context. Finally, attention is re-introduced at latter stages for semantic feature fusion. Also, we introduce a computation-downsampling decoupling (CDD) mechanism to preserve feature coverage during Mamba spatial scaling change. We have validated our approach on SpaceJam and Basketball-51 datasets. Experiments show that MMViT achieves superior performance over strong baselines with substantial margins. Ablation studies show the significance of A-M-A, MPA and CDD. MMViT achieves competitive accuracy among evaluated models and provides a favorable accuracy-efficiency trade-off for video action recognition task.
CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al.· IEEE Internet of Things Jour...· 0 citations
It is proved that the computationally cheaper split space-time attention is equivalent to full space-time attention and is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
N. Tran, Fanghui Xue, Shuai Zhang et al.· arXiv.org· 0 citations
A compact and deployable Convolutional Neural Network-Long Short Term Memory (CNN-LSTM) framework that combines a 2D convolutional backbone (AlexNet) for frame-level descriptors with an LSTM head for sequence modeling is proposed, indicating a robust, real-time-capable solution for video understanding in both offline a...
H. Khan, Altaf Hussain· ICCK Transactions on Advance...· 0 citations
Action prediction from frames and videos is a well-studied problem. Models trained with a single modality, mostly vision, will fail in low-light conditions. Recent works have attempted to predict action categories using vision-language and audio-visual models. A challenge, however, is that some dataset annotations lack...
A. R, Ambarish Parthasarathy, Sucharitha Devarakonda et al.· International Conference on...· 0 citations
This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.
Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.