Hybrid Model for Code-Switch and Dialect Translation in Low-Resource Malay Language
Abstract
Large Language Models struggle with dialectal and code-switched text like Kelantanese Malay and Manglish due to data-centric training that ignores non-standard morphology and phonology. This paper proposes a hybrid architecture that addresses these weaknesses by fusing three complementary linguistic representations — phoneme-aware embeddings, morphology-aware embeddings, and an unsupervised syntax-aware attention mechanism — integrated via a Multi-Channel Heterogeneous Interaction Module and gated fusion strategy into a T5-based encoder-decoder backbone. Experiments on 1,885 sentence pairs across Manglish and Kelantanese dialect datasets show the full fusion model achieves BLEU scores of 64.64 and 41.01 respectively, substantially outperforming the T5-Mesolitica baseline. Among ablation configurations, morphology and phoneme fusion yields the strongest performance with BLEU scores of 67.45 and 84.25, confirming that linguistically grounded multi-level representation significantly improves translation quality for low-resource Malay varieties.