Skip to content
Review Open access

A Comprehensive Survey on Self-Supervised Learning on Computer Vision: Moving Beyond the Supervised Paradigm

2026 · IEEE Access · Vol 14, pp. 127111-127148 · 0 citations · 90 references
Computer Science

TL;DR

This survey presents a seven-category taxonomy of self-supervised learning methods covering: 1) input reconstruction or restoration, 2) context prediction, 3) contrastive learning, 4) feature clustering, 5) self-distillation-based feature reconstruction, 6) redundancy reduction, and 7) masked image modeling.

Abstract

The availability of large-scale labeled datasets has driven advances in AI-based computer vision, yet supervised learning remains costly and impractical in domains where annotation is scarce. Self-supervised learning (SSL) addresses this by harnessing unlabeled data to learn rich, transferable representations without explicit supervision. This survey presents a seven-category taxonomy of self-supervised learning methods covering: 1) input reconstruction or restoration, 2) context prediction, 3) contrastive learning, 4) feature clustering, 5) self-distillation-based feature reconstruction, 6) redundancy reduction, and 7) masked image modeling, with coverage extended to recent methods that include DINOv2, I-JEPA, SparK, data2vec 2.0, V-JEPA, DINOv3, V-JEPA 2, V-JEPA 2.1, C-JEPA, PhiNet v2. We situate this work within the existing survey landscape by explicitly comparing our contributions with prior SSL reviews. Beyond method descriptions, we provide: a chronological timeline of SSL evolution from 2008 to 2026; a cross-paradigm comparative analysis evaluating all seven families along collapse risk, scalability, computational cost, and downstream transferability; a dedicated comparative analysis of anti-collapse mechanisms; critical limitations and trade-off analyses per method family; and systematic benchmarking evidence on ImageNet-1K, PASCAL VOC, COCO, and five public medical imaging datasets. We also contribute a practical method selection decision matrix, extended challenge discussions, and actionable open problems for future research.

Read PDF

Similar papers

Preprint Aug 2026

Far from the Crowd: Scalable Self-Supervised Learning via Geographic Isolation

Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure. We propose a curriculum learning strategy for self-supervised Earth observation that ranks samples by geographic isolation, a label-free proxy derived...

Daniele Rege Cambrin, Francesco Rossi, Mattia Varile · 0 citations
Open access Aug 2026

Semisupervised Adaptation of Vision-Language Models for Image Classification

Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.

M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al. · 0 citations
Open access Aug 2026

Enhancing self-supervised representation learning through lightweight learnable data augmentation

A lightweight learnable augmentation framework based on Extreme Learning Machines for self-supervised visual representation learning that improves linear evaluation performance relative to reproduced baselines across most settings and demonstrates that lightweight learnable augmentation can effectively enhance self-sup...

Mubarakah Alotaibi · 0 citations
Jul 2026

Object image retrieval based on self-supervised learning and zero-shot learning

A novel three-stage framework that integrates self-supervised contrastive pre-training with a dual-branch attention-driven architecture and employs an asymmetric deep hashing layer that promotes both high-level semantic consistency and computational efficiency through binary Hamming distance matching is proposed.

A. Nazari, Kambiz Rahbar, Hamidreza Moghassemi et al. · 0 citations
Preprint Aug 2026

Vernata: Self-Supervised Learning of LiDAR Point Representations

Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...

Oliver Lemke, Alexander Liniger, Abel Gawel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.