It is found that transfer is strongest between closely related Turkic pairs, especially Turkish-Azerbaijani and Kazakh-Kyrgyz and that the same transfer source-transfer target pair can behave differently when the translation target changes.
Abstract
Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices. In this setting, each model is fine-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish-Azerbaijani and Kazakh-Kyrgyz. We also show that transfer direction matters, and that the same transfer source-transfer target pair can behave differently when the translation target changes. Latinization improves BLEU and chrF in several script-mismatched settings, but its effect is not uniform across metrics. Additional analyses show that transfer sources are mostly stable across different datasets and model settings.
Overall, the effectiveness of script unification depends on the language, the induced subword overlap, and the available supervision, while within-language coverage becomes more important when target-language supervision is available.
A PMI-based translation metric is proposed, which is less dependent on the target language and correlates strongly with chrF, and finds that CLA with English predicts translation quality comparably to or better than source-target CLA.
Adnan Al Ali, Kathy Hämmerl, Jindrich Libovický et al.· 0 citations
This work proposes an approach that integrates code-switching directly into masked language model pretraining, and introduces a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens.
Ruan Visser, Trienko L. Grobler, Marcel Dunaiski· 0 citations
The results indicate that a moderately sized, shared self-attention architecture can deliver production-quality multilin-gual translation within the resource constraints of an academic de-ployment, while surfacing clear directions – low-resource language coverage, domain adaptation, and speech-based extension – for con...
Darshan Gowda D H and Dr. Kruti R· International Journal of Adv...· 0 citations
Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents, rather than corpus contents.
This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of par...
Jayanand A. Kamble, Shivajirao M. Jadhav, V. J. Kadam· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.