What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes....