Jun 2026
BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations
BEST-RQ-2 is presented, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a two-step contextualize-then-predict pretraining scheme, and consistently outperforms one-stage baselines in overall transfer while keeping inference compute unchanged.
Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini
· arXiv.org · 0 citations