A lightweight learnable augmentation framework based on Extreme Learning Machines for self-supervised visual representation learning that improves linear evaluation performance relative to reproduced baselines across most settings and demonstrates that lightweight learnable augmentation can effectively enhance self-supervised representation learning across different frameworks and datasets.
Abstract
Data augmentation plays a central role in self-supervised learning, as the quality and diversity of augmented views strongly influence the learned representations. However, most existing self-supervised methods rely on fixed stochastic augmentation pipelines, while more adaptive alternatives often require expensive policy search, adversarial training, or additional optimization procedures. In this paper, we propose a lightweight learnable augmentation framework based on Extreme Learning Machines (ELM) for self-supervised visual representation learning. The proposed module predicts image-dependent transformation parameters and applies them through a differentiable augmentation operator, enabling joint optimization with the representation model while introducing minimal additional computational overhead. The framework is integrated into three representative self-supervised learning methods: SimCLR, BYOL, and SimSiam. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny ImageNet show that the proposed method consistently improves linear evaluation performance relative to reproduced baselines across most settings. In particular, the method yields notable gains on CIFAR datasets and remains effective on the more challenging Tiny ImageNet benchmark. A per-class difficulty analysis further shows that the proposed augmentation strategy substantially improves performance on hard classes, indicating stronger robustness to challenging categories while maintaining competitive overall performance. In general, the results demonstrate that lightweight learnable augmentation can effectively enhance self-supervised representation learning across different frameworks and datasets.
This work introduces ActiveAugment, a unified framework that treats augmentation selection as an online active learning problem, and reveals that the augmentation selection policy evolves meaningfully during training and that strategy choice has a direct impact on generalisation.
GenView is presented, a controllable framework that augments the diversity of positive views leveraging the power of pretrained generative models while preserving semantics, and an adaptive view generation method that dynamically adjusts the noise level in sampling to ensure the preservation of essential semantic meaning while introducing variability.
Xiao-Jie Li, Yibo Yang, Xiangtai Li et al.· 0 citations
Learning-state-aware dynamic generative data augmentation (LSADA) is proposed, which introduces a decoupled data augmentation and diffusion fusion strategy that applies strength-controlled transformations to class-relevant regions and generates diverse class-irrelevant regions, progressively fusing them to improve image diversity while preserving class semantics.
Ting Xiang, Chen-Xi Deng, Jinhui Zhao et al.· 0 citations
Self-supervised representation-guided generative dataset distillation (SRG) is proposed, a framework that translates the SSL geometry into diffusion guidance and consistently outperforms the evaluated generative baselines across multiple datasets and IPC settings.
Mingzhuo Li, Guang Li, Linfeng Ye et al.· 0 citations
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations