Aug 2026· Frontiers in Big Data· Vol 9· 0 citations· 18 references
Medicine
TL;DR
The first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging is provided, providing the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.
Abstract
Introduction Part Of Speech (POS) tagging is a fundamental task in Natural Language Processing (NLP) that assigns grammatical labels to words in a sentence. Code mixed text, which entails switching between two or more languages within a single conversation or a sentence, presents challenges for POS tagging. This investigation entailed a comprehensive study of deep learning approaches for cross linguistic POS tagging, focused on English Malayalam code mixed data prevalent on social media platforms. The study was carried out on linguistically complex and varied English Malayalam code mixed text from social media platforms with informal spellings, language switching, transliteration, slang, and unclear grammatical boundaries, reflecting the characteristics of informal online communication. We provide the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging. Methods We evaluated 14 state of the art model configurations that span traditional sequence labeling approaches and multilingual transformer architectures. Models were compared using standard performance metrics prevalent in the domain of data science, supplemented by normalized confusion matrices, error prone tag identification and micro-macro F1 gap analysis. Results Our results showed that CRF (No Lang) emerged as the most balanced model overall on macro F1 (all classes) of 0.8170, while (BiLSTM + CRF) achieved the highest macro F1 (seen classes) of 0.8831, precision of 0.9167, and recall of 0.875, though this reflects strong performance concentrated on frequent tag classes rather than balanced coverage across the full tag set. Notably, the pretrained multilingual transformers (mBERT, MuRIL), despite prior exposure to Malayalam during pretraining, were outperformed on several key metrics by CRF and BiLSTM models trained directly on the code mixed dataset. Discussion This finding was contrary to our expectation that existing multilingual knowledge would translate into a clear advantage on this task and merits further investigation.
This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.
Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al.· 0 citations
Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.
Chirag D. Shah, Shailesh A. Chaudhari· International journal of com...· 0 citations
Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring social media posts (486,974 tokens) rich in Turkish-English code-mixing. From this corpus, we construct a new Turkish-English benchmark for code-mixed language identification (LID) and named entity recognition (NER), comprising 15K expert-annotated tokens. Evaluating both decoder LLM and fine-tuned encoder baselines, we find that monolingual Turkish and English tokens are labeled reliably, but all models have high error rates on mixed-language tokens for both LID and NER. For morphologically integrated tokens, NER error rates were 5.2x and 6.3x higher for GPT-4o and Qwen, respectively. This highlights how morphological integration remains a challenge. We release the corpus, annotations, and code to support future computational and sociolinguistic research on Turkish-English code-mixing.
Ilayda Dogan, Phuong-Anh Nguyen-Le, Julia Mendelsohn· 0 citations
The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.
Maureen Otieno, L. Wanzare, Calvins Otieno· International Journal of Com...· 0 citations
Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate these systems on three different data configurations (Hindi, Gujarati, and Bengali) to predict language labels for individual tokens. We release a benchmark for language identification in code-mixed tokens with manually annotated test sets. We propose two approaches of code-mixed generation using parallel sentences of three languages. The trained models demonstrate the effectiveness of contextual embeddings for token-level language identification in multilingual social media text. For reproducibility and to facilitate future research, we publicly release our fine-tuned models.
Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al.· 0 citations
: Multilingual speakers often alternate between languages within a conversation, a phenomenon known as code-switching. This is common in Arabic-speaking communities, where Arabic and English are frequently mixed in everyday communication. Although recent advances in natural language processing have been driven by pretrained multilingual language models, these models are largely trained on monolingual data and often struggle to capture the abrupt language transitions and cross-lingual semantic interactions that characterize code-switched text. This work investigates representation learning for Arabic-English code-switched text at multiple levels. At the word level, we employ a code-switch-aware masked language modeling objective that captures token-level language variation and switch points. At the sentence level, we adopt a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants. To support this objective, we introduce CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset. The resulting embeddings are evaluated on sentiment analysis and named entity recognition. On sentiment analysis, the word-level pipeline improves F1-score by 2.17 percentage points over mBERT and 2.27 percentage points over XLM-R, while the sentence-level pipeline improves F1-score by 1.78 and 1.28 percentage points, respectively. In contrast, named entity recognition shows only marginal, statistically insignificant gains.
Mariam Rizkallah, Amani Ghonim, A. Sherif et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.