Transformer-based heterogeneous attention and multiscale feature extraction network for video saliency prediction
Abstract
Video saliency prediction aims to estimate the regions in dynamic scenes that are most likely to attract human attention. Although Transformer-based backbones have demonstrated strong performance in spatiotemporal representation learning, existing decoders still rely primarily on convolutional operations, which are limited in modeling long-range dependencies during saliency map reconstruction. To address this issue, we propose HAMFE-Net, a Transformer-based encoder-decoder network for video saliency prediction. Specifically, a heterogeneous attention module is introduced into the decoder to employ different attention mechanisms for different feature hierarchies. Global self-attention is applied to deep features to capture long-range spatiotemporal dependencies, while a lightweight local-global attention mechanism is adopted for middle- and low-level features to preserve structural details at low computational cost. In addition, a multiscale feature extraction module is designed to enhance cross-level feature fusion and improve saliency localization. Experimental results on DHF1K, Hollywood-2, and UCF-Sports demonstrate the effectiveness of the proposed method.