Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations
This work investigates ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input and shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recordin...