Skip to content
Preprint

Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

V-JEPA4A is introduced, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy that preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction.

Abstract

Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.

View source

Similar papers

Preprint Aug 2026

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

This work introduces TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization and improves motion-centric retrieval across model scales.

Yi-Chung Chen, P. Jacobson, Tom Lampo et al. · 0 citations
May 2025

Object Concepts Emerge from Motion

Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...

Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al. · 0 citations
Open access Aug 2026

Spatio-Temporal Masked Autoencoders with Contrastive Cross-Attention for Robust Video Semantic Segmentation in Adverse Autonomous Driving Conditions

ST-MaskContrast incorporates a decoupled space-time token masking strategy coupled with a multi-scale cross-attention temporal decoder that enforces invariant representations between clean and corrupted sequential frames via contrastive patch-level loss, providing unprecedented perceptual resilience for safety-critical...

Rashed Karim Bipul · 0 citations
2025

EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning

This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model's predictions.

Yuhan Liu, Linghui Fu, Zheng Yang et al. · 1 citation
Preprint Sep 2026

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

SV-WAM is proposed, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference and introduces a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning saf...

Jin-Yang Wang, Shi-Wei Li, Jun-Jian Wang et al. · 0 citations
Open access Sep 2026

Unified vision-centric pedestrian crossing action prediction via adaptive patch projection and proactive spatial rectification

ViCross is proposed, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization.

Yao Tian, Le Yang, Bing-Lu Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.