Skip to content

Efficient compressed video action recognition via frame selection and token merging in vision transformers

Aug 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 45 references

TL;DR

This work proposes a compressed-video-oriented framework, the Frame Selection and Token Merging for Efficient Compressed Video Transformer (FSTM-ECVT), which follows a dual-stream Transformer architecture equipped with a Global Multi-Modal Fusion module to effectively leverage the distinctive and complementary characteristics of RGB and compressed motion modalities.

View source

Similar papers

Jul 2026

ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression

ENCORE, an Event-Assisted Complementary Motion Refinement framework for learned video compression, employs Complementary Motion Representation to decompose aligned RGB-event features into common and modality-specific motion representations and identifies event-specific responses that are active and novel relative to RG...

Shuhan Ye, Hong Yu, Chenqi Kong et al. · 0 citations
Conference Aug 2026

Transformer-based heterogeneous attention and multiscale feature extraction network for video saliency prediction

Video saliency prediction aims to estimate the regions in dynamic scenes that are most likely to attract human attention. Although Transformer-based backbones have demonstrated strong performance in spatiotemporal representation learning, existing decoders still rely primarily on convolutional operations, which are lim...

Yunxiang Liao, Zeyu Zhao · 0 citations
Preprint Aug 2026

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens that achieves strong reconstruction and generation quality at a state-of-the-art compression ratio.

Yeonkyeong Lee, Hyun-Young Go, Jongmin Kim et al. · 0 citations
Conference Aug 2026

Sparsity-Aware Vision Transformer Network for Efficient Event-Based Action Recognition

Event cameras can capture human motion with low latency and high dynamic range, providing an emerging sensing paradigm for action recognition. A large body of recent work converts event data into a sequence of dense frames and feeds them into a video classification model for prediction. Although the frame-based methods...

Bo-Cheng Xie, Jian Liu, You-Xuan Fang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.