Skip to content
Conference

Transformer-based heterogeneous attention and multiscale feature extraction network for video saliency prediction

Aug 2026 · International Conference on Computer Vision and Pattern Analysis · Vol 14296, pp. 142960Y - 142960Y-8 · 0 citations · 15 references
Engineering

Abstract

Video saliency prediction aims to estimate the regions in dynamic scenes that are most likely to attract human attention. Although Transformer-based backbones have demonstrated strong performance in spatiotemporal representation learning, existing decoders still rely primarily on convolutional operations, which are limited in modeling long-range dependencies during saliency map reconstruction. To address this issue, we propose HAMFE-Net, a Transformer-based encoder-decoder network for video saliency prediction. Specifically, a heterogeneous attention module is introduced into the decoder to employ different attention mechanisms for different feature hierarchies. Global self-attention is applied to deep features to capture long-range spatiotemporal dependencies, while a lightweight local-global attention mechanism is adopted for middle- and low-level features to preserve structural details at low computational cost. In addition, a multiscale feature extraction module is designed to enhance cross-level feature fusion and improve saliency localization. Experimental results on DHF1K, Hollywood-2, and UCF-Sports demonstrate the effectiveness of the proposed method.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.