Skip to content
Preprint

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

Combining DINOv2 with video SSL objectives such as VideoMAE substantially improves image classification and segmentation performance, but degrades video tracking and camera-pose estimation performance, revealing an important tradeoff between semantic and geometric representation learning.

Abstract

Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant benefits arise when effective models can be obtained with fewer resources. To better understand how self-supervised learning (SSL) objectives behave under resource constraints, we conduct a controlled study of image and video SSL objectives under matched data, architecture, and compute budgets. We compare contrastive, reconstruction, feature-prediction, and diffusion objectives and evaluate both standalone and jointly trained image-video SSL formulations across a diverse set of image and video understanding tasks. Our results show that DINOv2-style pretraining consistently provides the strongest overall performance under limited resources. Furthermore, combining DINOv2 with video SSL objectives such as VideoMAE substantially improves image classification and segmentation performance, but degrades video tracking and camera-pose estimation performance, revealing an important tradeoff between semantic and geometric representation learning. These findings suggest that combining image and video SSL objectives can be beneficial in resource-limited settings, while highlighting the need for improved methods that better balance semantic, temporal, and geometric supervision.

View source

Similar papers

Preprint Aug 2026

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate in...

Bonan Zhang, Shiyu Dong, Quan Hung Tran et al. · 1 citation

UvA-DARE (Digital Academic Repository) Elastic ViTs from Pretrained Models without Retraining

SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...

Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al. · 0 citations
Preprint Sep 2026

High-Fidelity Video Quality Assessment with VQA-Specific Saliency

No-reference video quality assessment (NR VQA) has recently seen promising progress with deep learning. However, video data is inherently large, and processing them with deep models incurs high computational cost. This challenge is particularly acute in VQA, where preserving original-resolution cues and dense temporal...

H. Gedik, Shashank Gupta, Alan C. Bovik · 0 citations
Preprint Aug 2026

LeVJEPA: Efficient&Scalable Video Pretraining without the Heuristics

LeVJEPA is introduced, the first video encoder trained under LeJEPA's collapse-free objective, which indicates that video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.

Lukas Kuhn, Lucas Maes, Giuseppe Serra et al. · 0 citations
Aug 2026

A visual-neural network for specific objects-of-interest inpainting

Experiments indicate that the proposed front end of Specific Object-of-Interest Imaging improves object-region reconstruction and boundary stability for several inpainting backbones, although the gains are not uniform across all metrics or categories.

Yong-Hao Wu, Rui-Rong Wang, Kang-Wen Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.