Skip to content

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

Jul 2026 · arXiv.org · Vol abs/2607.06392 · 1 citation · 36 references
Computer Science

TL;DR

This work establishes a task-agnostic analysis from three intrinsic per-layer perspectives: compression, geometry, and robustness to perturbations, and introduces the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability.

Abstract

Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called InsideSSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.

View source

Similar papers

Preprint Aug 2026

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle path, constraining representation capacity. We introduce Manifold-Constrained Hyper-Connections (mHC), reformulat- ing residual paths as a mu...

Zezhong Jin, Xiaoyu Wang, Zhe Li et al. · 0 citations
Preprint Jul 2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

A multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model is proposed and shows that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.

Xiaoyu Yang, Xuenan Xu, Wenyi Yu et al. · 0 citations
Preprint Aug 2026

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or cont...

Yihui Fu, Zhengyang Li, Tim Fingscheidt · 0 citations
Preprint Aug 2026

Trajectory Dynamics in Self-Supervised Learning Latent Space for Audio Deepfake Detection

On near-domain benchmarks, static and dynamic approaches perform comparably, and on harder cross-corpus benchmarks with diverse synthesis methods, trajectory dynamics provide substantial gains, confirming that temporal physiological constraints carry detection signal beyond utterance-level statistics.

T. Weber · 0 citations
Preprint Aug 2026

DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning

DINO-A is presented, an adaptation of self-distillation from vision to general audio representation learning, and it is traced to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2...

Tomasz Radzikowski, M. Modrzejewski, Przemysław Rokita · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.