V-JEPA4A is introduced, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy that preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction.
Abstract
Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.
This work introduces TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization and improves motion-centric retrieval across model scales.
Yi-Chung Chen, P. Jacobson, Tom Lampo et al.· 0 citations
Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...
Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al.· Neural Information Processin...· 0 citations
ST-MaskContrast incorporates a decoupled space-time token masking strategy coupled with a multi-scale cross-attention temporal decoder that enforces invariant representations between clean and corrupted sequential frames via contrastive patch-level loss, providing unprecedented perceptual resilience for safety-critical...
Rashed Karim Bipul· International Journal of App...· 0 citations
This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model's predictions.
Yuhan Liu, Linghui Fu, Zheng Yang et al.· Neural Information Processin...· 1 citation
SV-WAM is proposed, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference and introduces a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning saf...
Jin-Yang Wang, Shi-Wei Li, Jun-Jian Wang et al.· 0 citations
ViCross is proposed, a vision-centric pedestrian crossing action prediction framework powered by multimodal large language models, which maintains target-centric reasoning from video frames without additional perception modules beyond first-frame target initialization.
Yao Tian, Le Yang, Bing-Lu Wang· Pattern Recognition· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.