Jul 2026· Data Technologies and Applications· pp. 1-21· 0 citations· 20 references
TL;DR
The results show that cascaded pipelines can produce semantically meaningful Ma’anyan–Indonesian translations even under high transcription error conditions, and indicate that ASR quality is the dominant determinant of overall speech translation performance, while larger LoRA-adapted MT models provide stronger robustness against noisy ASR outputs.
Abstract
This study aims to address the technological marginalization of endangered regional languages by evaluating speech-to-text translation for Dayak Ma’anyan, an extremely low-resource Austronesian language. In particular, it seeks to examine whether cascaded multilingual automatic speech recognition and machine translation models can provide effective Ma’anyan–Indonesian translation despite severe data scarcity.
This study employs a cascaded speech-to-text translation framework that combines two multilingual automatic speech recognition models, Whisper Large-v3 and SeamlessM4T v2, with two LoRA-adapted multilingual machine translation models, NLLB-200 3.3B and distilled 600M. Experiments are conducted in an extremely low-resource setting using limited parallel speech and text data. The proposed pipelines are evaluated at three levels: ASR transcription quality, machine translation performance and end-to-end semantic preservation.
The results show that cascaded pipelines can produce semantically meaningful Ma’anyan–Indonesian translations even under high transcription error conditions. Whisper substantially outperforms SeamlessM4T at the ASR stage, achieving a lower WER (0.464 vs 0.812) and yielding better downstream translation quality. Among the machine translation models, LoRA-adapted NLLB-200 3.3B achieves the best performance, with BLEU 31.00, chrF 58.91 and the highest end-to-end semantic similarity (SBERT 0.722). The findings further indicate that ASR quality is the dominant determinant of overall speech translation performance, while larger LoRA-adapted MT models provide stronger robustness against noisy ASR outputs.
This study provides, to the best of the authors’ knowledge, the first empirical benchmark for Ma’anyan–Indonesian speech-to-text translation. It contributes a systematic evaluation of multilingual ASR and LoRA-adapted MT combinations for endangered-language technology and offers empirical insight into the relative impact of ASR quality and MT model capacity in extremely low-resource cascaded speech translation.
Comparing empirical cross-lingual transfer with typology-based similarity, it is found that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training.
Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson et al.· 0 citations
The rapid expansion of online IT education has created a pressing need to bridge the English-Vietnamese language barrier for non-native learners, as subtitle-based solutions impose cognitive load and fail to preserve the lecturer’s vocal identity. This paper presents a cascade Speech-to-Speech Translation (S2ST) pipeli...
Trang Thi Thuy Pham, Nhut Minh Nguyen, T. Nguyen· International Conference on...· 0 citations
Speech-to-text translation for low-resource language pairs such as Vietnamese–English remains underexplored, despite growing demand in real-world applications. In this study, we present a systematic zero-shot benchmark on the Vietnamese test split of FLEURS, evaluating 30 configurations across three system families: co...
L. Nguyen· International Conference on...· 0 citations
The translationese content of MLLM generations is assessed and the key features that distinguish MLLM-generated text from typical translation-related interference are examined.
Maria R. Valentini, Téa Wright, Julisa Granados et al.· 0 citations
We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 language families. The corpus comprises approximat...
L. Gris, A. I. Ferreira, F. S. de Oliveira et al.· 0 citations
This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of par...
Jayanand A. Kamble, Shivajirao M. Jadhav, V. J. Kadam· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.