Transformer-based heterogeneous attention and multiscale feature extraction network for video saliency prediction
Video saliency prediction aims to estimate the regions in dynamic scenes that are most likely to attract human attention. Although Transformer-based backbones have demonstrated strong performance in spatiotemporal representation learning, existing decoders still rely primarily on convolutional operations, which are lim...