Aug 2026· Frontiers in Artificial Intelligence· Vol 9· 0 citations· 35 references
Medicine
TL;DR
It is demonstrated that pretrained multilingual models effectively address TTG transformation for low-resource agglutinative languages through transfer learning, with practical applications for Kazakh Sign Language assistive systems and education.
Abstract
This study investigates the automatic transformation of Kazakh text into sign language glosses (Text-to-Gloss) using multilingual transformer-based models with emphasis on preserving morphological structure in a low-resource agglutinative language framework. Given the scarcity of high-quality intermediate representations for Kazakh Sign Language, a methodology for corpus formation was developed, resulting in a specialized dataset of 11 190 unique text−gloss pairs sourced from educational materials. To assess the effectiveness of transfer learning under low-resource constraints, we fine-tuned mT5-small and mT5-base models on this dataset. Experimental results indicate that the mT5-base architecture consistently outperforms the smaller variant across all metrics, achieving a 2.61-point improvement in BLEU (87.51 → 90.12) and a 2.47-point increase in ROUGE-1 F1 (89.61% → 92.08%), alongside gains in exact match (74.96% → 80.68%) and chrF (91.6 → 94.1). Statistical significance testing (p < 0.001) confirms the stability of this improvement. Error analysis reveals that larger model capacity particularly improves morphological handling and token coverage, which is critical for agglutinative languages, while maintaining semantic accuracy. These findings demonstrate that pretrained multilingual models effectively address TTG transformation for low-resource agglutinative languages through transfer learning, with practical applications for Kazakh Sign Language assistive systems and education.
A Contrastive Learning-based Chinese-English Scientific Translation Quality Evaluation model (C-TQE), which provides an effective solution for large-scale scientific translation quality assessment and facilitates the accurate international communication of multidisciplinary engineering research, including electromagnet...
This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation that aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4.
Oline Ranum, Edward Fish, Simon Hadfield et al.· 0 citations
Large Language Models struggle with dialectal and code-switched text like Kelantanese Malay and Manglish due to data-centric training that ignores non-standard morphology and phonology. This paper proposes a hybrid architecture that addresses these weaknesses by fusing three complementary linguistic representations — p...
Noor Ali Ikhwan Bin Noor Azmy, Nurzeatul Hamimah Abdul Hamid, Azliza Mohd Ali· 2026 7th International Confe...· 0 citations
Malayalam, a classical Dravidian language spoken by approximately 38 million people in the Indian state of Kerala, remains severely underrepresented in natural language processing research. This paper presents the systematic study of abstractive text summarization for Malayalam educational documents using mT5-base, a m...
Ajmal E. B., M. Rajesh· International Conference Inn...· 0 citations