This work proposes a compressed-video-oriented framework, the Frame Selection and Token Merging for Efficient Compressed Video Transformer (FSTM-ECVT), which follows a dual-stream Transformer architecture equipped with a Global Multi-Modal Fusion module to effectively leverage the distinctive and complementary characteristics of RGB and compressed motion modalities.
ENCORE, an Event-Assisted Complementary Motion Refinement framework for learned video compression, employs Complementary Motion Representation to decompose aligned RGB-event features into common and modality-specific motion representations and identifies event-specific responses that are active and novel relative to RG...
Shuhan Ye, Hong Yu, Chenqi Kong et al.· arXiv.org· 0 citations
Video saliency prediction aims to estimate the regions in dynamic scenes that are most likely to attract human attention. Although Transformer-based backbones have demonstrated strong performance in spatiotemporal representation learning, existing decoders still rely primarily on convolutional operations, which are lim...
Yunxiang Liao, Zeyu Zhao· International Conference on...· 0 citations
KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens that achieves strong reconstruction and generation quality at a state-of-the-art compression ratio.
Yeonkyeong Lee, Hyun-Young Go, Jongmin Kim et al.· 0 citations
Event cameras can capture human motion with low latency and high dynamic range, providing an emerging sensing paradigm for action recognition. A large body of recent work converts event data into a sequence of dense frames and feeds them into a video classification model for prediction. Although the frame-based methods...
Bo-Cheng Xie, Jian Liu, You-Xuan Fang et al.· 2026 IEEE International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.