Skip to content

Mixed approach speech-to-text translation for endangered language

Jul 2026 · Data Technologies and Applications · pp. 1-21 · 0 citations · 20 references

TL;DR

The results show that cascaded pipelines can produce semantically meaningful Ma’anyan–Indonesian translations even under high transcription error conditions, and indicate that ASR quality is the dominant determinant of overall speech translation performance, while larger LoRA-adapted MT models provide stronger robustness against noisy ASR outputs.

Abstract

This study aims to address the technological marginalization of endangered regional languages by evaluating speech-to-text translation for Dayak Ma’anyan, an extremely low-resource Austronesian language. In particular, it seeks to examine whether cascaded multilingual automatic speech recognition and machine translation models can provide effective Ma’anyan–Indonesian translation despite severe data scarcity. This study employs a cascaded speech-to-text translation framework that combines two multilingual automatic speech recognition models, Whisper Large-v3 and SeamlessM4T v2, with two LoRA-adapted multilingual machine translation models, NLLB-200 3.3B and distilled 600M. Experiments are conducted in an extremely low-resource setting using limited parallel speech and text data. The proposed pipelines are evaluated at three levels: ASR transcription quality, machine translation performance and end-to-end semantic preservation. The results show that cascaded pipelines can produce semantically meaningful Ma’anyan–Indonesian translations even under high transcription error conditions. Whisper substantially outperforms SeamlessM4T at the ASR stage, achieving a lower WER (0.464 vs 0.812) and yielding better downstream translation quality. Among the machine translation models, LoRA-adapted NLLB-200 3.3B achieves the best performance, with BLEU 31.00, chrF 58.91 and the highest end-to-end semantic similarity (SBERT 0.722). The findings further indicate that ASR quality is the dominant determinant of overall speech translation performance, while larger LoRA-adapted MT models provide stronger robustness against noisy ASR outputs. This study provides, to the best of the authors’ knowledge, the first empirical benchmark for Ma’anyan–Indonesian speech-to-text translation. It contributes a systematic evaluation of multilingual ASR and LoRA-adapted MT combinations for endangered-language technology and offers empirical insight into the relative impact of ASR quality and MT model capacity in extremely low-resource cascaded speech translation.

View source

Similar papers

#natural language process... Preprint Sep 2026

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

Comparing empirical cross-lingual transfer with typology-based similarity, it is found that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training.

Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson et al. · 0 citations
Conference Aug 2026

A Cascade English–Vietnamese Speech-to-Speech Translation Framework for Online IT Education

The rapid expansion of online IT education has created a pressing need to bridge the English-Vietnamese language barrier for non-native learners, as subtitle-based solutions impose cognitive load and fail to preserve the lecturer’s vocal identity. This paper presents a cascade Speech-to-Speech Translation (S2ST) pipeli...

Trang Thi Thuy Pham, Nhut Minh Nguyen, T. Nguyen · 0 citations
Conference Aug 2026

Benchmarking Open-Source Vietnamese-English Speech-to-Text Translation Systems

Speech-to-text translation for low-resource language pairs such as Vietnamese–English remains underexplored, despite growing demand in real-world applications. In this study, we present a systematic zero-shot benchmark on the Vietnamese test split of FLEURS, evaluating 30 configurations across three system families: co...

L. Nguyen · 0 citations
#natural language process... Preprint Sep 2026

CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 language families. The corpus comprises approximat...

L. Gris, A. I. Ferreira, F. S. de Oliveira et al. · 0 citations
Open access Sep 2026

Bridging the linguistic divide: recent developments in machine translation for Indian languages

This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of par...

Jayanand A. Kamble, Shivajirao M. Jadhav, V. J. Kadam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.