Skip to content
Open access

SwitchEmbed: Representation Learning for Arabic-English Code-Switched Text

2026 · International Conference on Data Technologies and Applications · pp. 712-719 · 0 citations · 28 references
Computer Science

Abstract

: Multilingual speakers often alternate between languages within a conversation, a phenomenon known as code-switching. This is common in Arabic-speaking communities, where Arabic and English are frequently mixed in everyday communication. Although recent advances in natural language processing have been driven by pretrained multilingual language models, these models are largely trained on monolingual data and often struggle to capture the abrupt language transitions and cross-lingual semantic interactions that characterize code-switched text. This work investigates representation learning for Arabic-English code-switched text at multiple levels. At the word level, we employ a code-switch-aware masked language modeling objective that captures token-level language variation and switch points. At the sentence level, we adopt a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants. To support this objective, we introduce CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset. The resulting embeddings are evaluated on sentiment analysis and named entity recognition. On sentiment analysis, the word-level pipeline improves F1-score by 2.17 percentage points over mBERT and 2.27 percentage points over XLM-R, while the sentence-level pipeline improves F1-score by 1.78 and 1.28 percentage points, respectively. In contrast, named entity recognition shows only marginal, statistically insignificant gains.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.