Skip to content

Three Necessary Principles for Self-Supervised Visual Representation Learning

Aug 2026 · 0 citations · 32 references
Computer Science

Abstract

We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.

View source

Similar papers

#machine learning Preprint Sep 2026

Sharp Rates and a One-Line Correction for Spectral Representation Learning

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum o...

Di-Er Tang, Jing-Yee Tan, Guang-Yue Han · 0 citations
Open access 2026

MISA: Mutual Information-Driven Separator With Spectral Alignment

This research introduces a domain generalization framework, MISA (Mutual Information-driven Separator with Spectral Alignment), which disentangles and learns semantic features via mutual information optimization and spectral alignment.

Sihwa Lee, Taehun Lee, Yoon Kim · 0 citations
Jul 2026

Unbiased Open World Regularization for Fair Self-Supervised Learning

The Unbiased Open World Regularization (UOWReg) framework is proposed, an encoder-only framework that effectively prevents the subpopulation collapse observed in standard SSL, successfully isolating micro-signatures even when heavily entangled with the global structure.

Léo Nicollier, M. Pic, Pablo Mus'e et al. · 0 citations
Preprint Aug 2026

LeVJEPA: Efficient&Scalable Video Pretraining without the Heuristics

LeVJEPA is introduced, the first video encoder trained under LeJEPA's collapse-free objective, which indicates that video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.

Lukas Kuhn, Lucas Maes, Giuseppe Serra et al. · 3 citations · ⚡1
Preprint Sep 2026

Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison

Gekko is a network that jointly performs cross-view completion, masked autoencoding, and per-pixel prediction of this relative improvement of the cross-view reconstruction error over a masked-autoencoder error, providing an additional binocular signal for all masked regions without any ground-truth 3D annotation.

T. Loiseau, Guillaume Bourmaud, Vincent Lepetit · 1 citation
Preprint Aug 2026

PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization

Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image...

Zhaorui Tan, Weimiao Yu, Xi Yang · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.