Skip to content
Conference

Hybrid Model for Code-Switch and Dialect Translation in Low-Resource Malay Language

Sep 2026 · 2026 7th International Conference on Artificial Intelligence and Data Sciences (AiDAS) · pp. 251-255 · 0 citations · 24 references

Abstract

Large Language Models struggle with dialectal and code-switched text like Kelantanese Malay and Manglish due to data-centric training that ignores non-standard morphology and phonology. This paper proposes a hybrid architecture that addresses these weaknesses by fusing three complementary linguistic representations — phoneme-aware embeddings, morphology-aware embeddings, and an unsupervised syntax-aware attention mechanism — integrated via a Multi-Channel Heterogeneous Interaction Module and gated fusion strategy into a T5-based encoder-decoder backbone. Experiments on 1,885 sentence pairs across Manglish and Kelantanese dialect datasets show the full fusion model achieves BLEU scores of 64.64 and 41.01 respectively, substantially outperforming the T5-Mesolitica baseline. Among ablation configurations, morphology and phoneme fusion yields the strongest performance with BLEU scores of 67.45 and 84.25, confirming that linguistically grounded multi-level representation significantly improves translation quality for low-resource Malay varieties.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.