Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 26 references
Abstract
Evaluating speech-to-speech translation (S2ST) systems remains challenging due to the multi-dimensional nature of speech, encompassing semantic accuracy, acoustic quality, and speaker characteristics. Existing approaches rely on either textbased or speech-based metrics, each capturing only partial aspects of translation quality. In this work, we conduct a systematic empirical analysis of S2ST evaluation metrics across multiple language pairs (fr-en, es-en, de-en, hi-en) using SeamlessM4T-v2 translations generated from the FLEURS dataset. We analyze ngram, neural text, and speech-embedding metrics at both corpus and sentence levels. Our analysis shows that text-based metrics are sensitive to linguistic variation but depend on automatic speech recognition (ASR), while speech-based metrics provide more consistent scores yet exhibit limited discriminative ability. Based on these observations, we propose a multi-dimensional evaluation framework that jointly assesses semantic adequacy and acoustic naturalness. The proposed framework improves semantic alignment and achieves strong correlation with speech naturalness across language pairs, enabling more reliable evaluation of S2ST systems.
A reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis, which reveals substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge.
Ali B. Jafar, Amal Sarmad, Shifa Yousaf et al.· 0 citations
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based evaluation suite that compares speech-only, text-only, and speech-text language models trained on matched data distributions and evaluated in matched generation settings. We evaluate generated continuations along multiple dimensions: semantic coherence, measured by transcribing generated speech and scoring it with a reference language model; local phonetic structure, measured by phone n-gram distributional statistics; speaker consistency and acoustic quality; and emotion-based distributional metrics. Across datasets, we find that joint speech-text modeling substantially improves semantic coherence. However, the improvement is not uniform across metrics: phone-level metrics change only modestly, speaker similarity and predicted quality are lower for speech-text continuations, while emotion-based distributional metrics improve. Compared with larger-scale speech-only models, our speech-text model closes much of the scaling gap in transcript-based semantic coherence, suggesting that text provides an efficient semantic training signal for spoken language modeling.
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.· 0 citations
Advancement of zero-shot text-to-speech synthesis is currently hindered for low-resource languages by the scarcity of large-scale, high-fidelity speech datasets. Traditional alignment-based methods require rare verbatim transcripts, while standard in-the-wild pipelines often rely on single-model automatic speech recognition and silence-based segmentation, leading to transcription errors and truncated prosody. To address these challenges for the Persian language, this paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data. Our approach integrates a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling. Furthermore, we employ a dual-model agreement mechanism, leveraging two distinct model architectures to filter unreliable transcriptions without ground truth. This pipeline yields a 2,400-hour multi-speaker dataset, the largest open-source speech resource available for Persian to date. Additionally, we provide the first comparative benchmark of speech quality assessment methods for Persian, releasing a human-annotated subset to facilitate future research.
Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al.· 0 citations
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Oluwanifemi Bamgbose, Simon Rosen, J. Shah et al.· 0 citations
Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.
J. Hebert, Amalia Zahra· Bulletin of Electrical Engin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.