Jul 2026
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual of linguistic content, and shows a 98.7% relative reduction in speaker probe accuracy.
Laurin Wagner, Bernhard Thallinger, Miroslav Stankovič et al.
· arXiv.org · 2 citations