Preprint
Aug 2026
A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources
Combining DINOv2 with video SSL objectives such as VideoMAE substantially improves image classification and segmentation performance, but degrades video tracking and camera-pose estimation performance, revealing an important tradeoff between semantic and geometric representation learning.
B. B. Englert, G. Dubbelman
· 0 citations