Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
This work presents the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels, and assesses fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder models on transcript-based versus direct audio processing.