Skip to content
Open access

Language Identification in Transliteration-Based Code-Mixed Text: A Study on Telugu–English Data

Jun 2026 · International Journal of Computer Science and Engineering · Vol 14, pp. 34-41 · 0 citations

TL;DR

This work focuses on word-level Language Identification (LID) for transliterated text in informal Roman transliteration, and relies on character-based TF–IDF features and a set of traditional machine-learning models.

Abstract

In recent years, social media users in multilingual regions have begun mixing languages more freely, and Telugu–English combinations are among the most common examples in India. Much of this content appears in informal Roman transliteration, and the lack of uniform spelling makes automatic processing difficult. In this work, we focus on word-level Language Identification (LID) for such transliterated text. Our approach relies on character-based TF–IDF features and a set of traditional machine-learning models. In this study, we worked with four different models—Multinomial Naive Bayes, Logistic Regression, Random Forest, and Support Vector Machine—and evaluated them on an annotated set that included Telugu, English, Named Entity, and Universal tokens. Among the four, the SVM turned out to be the strongest, reaching an accuracy of 86% along with an F1-score of 0.85. The study also brings out some practical issues with real-world transliterated text, particularly class imbalance and the wide range of spelling variations. We conclude with possible directions for improvement, including the use of neural and transformer-based models that might capture more contextual cues in future versions of this system.

Read PDF

Similar papers

Open access Aug 2026

Linguistically Informed Machine Learning for Gujarati–English Code-Mixed Sentiment Classification: A Comparative Study of Feature Fusion Strategies

Code-mixed text, in which words from multiple languages are used within the same sentence, is common on social media platforms and poses significant challenges for conventional natural language processing techniques. This study focuses on five-class sentiment classification of of Gujarati–English code-mixed text. The dataset consists of 5 sentiment classes ranging from extremely negative to extremely positive distributed over 44,672 sentences. A linguistically informed framework is proposed that incorporates word-level annotations, including language identity, sentiment polarity, and intensifier presence, generated using a multi-task fine-tuned DistilBERT tagger. These annotations are aggregated into sentence-level handcrafted features and combined with conventional text representations, namely Bag of Words (BoW), Term Frequency–Inverse Document Frequency (TF-IDF), Word2Vec, and FastText. The proposed framework is evaluated using a late fusion approach based on out-of-fold (OOF) stacking and is compared with early fusion, where features are directly concatenated, and with embedding-only baselines. The statistical significance of performance differences is assessed using McNemar's test. Seven machine learning classifiers—Logistic Regression (LR), Multinomial Naïve Bayes (MNB), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Decision Tree (DT), Random Forest (RF), and Extreme Gradient Boost(XGB)—are evaluated. The hyperparameters of all classifiers are optimized using GridSearchCV and then kept fixed throughout the experiments to ensure a fair comparison. Class imbalance is addressed using cost-sensitive learning through class-weight adjustment. The experimental results show that early fusion of linguistic and text features consistently outperforms the embedding-only and late fusion approaches for most classifier–representation combinations. The best-performing model is XGB with FastText under the early fusion framework, achieving an accuracy of 0.78 and a macro F1-score of 0.74 on the unseen test set. Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text. In addition, it provides a comprehensive and statistically validated comparison of feature integration strategies for sentiment analysis in low-resource, code-mixed language settings.

Chirag D. Shah, Shailesh A. Chaudhari · 0 citations
Preprint Aug 2026

Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study

Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the English and Pnar language pair. Using articles collected from the Wyrta newspaper, we built a parallel corpus comprising of 10,234 sentences and trained phrase-based statistical machine translation (SMT) systems the models using 9,563 parallel corpora under three configurations for each direction using Moses, GIZA++ , KenLM, varying lexicalized reordering and minimum error rate training (MERT) tuning. The models are evaluated on a held out test set of 371 sentences, the best performing system achieves a BLEU score of 14.97 (chrF2: 33.42, TER: 77.60) for Pnar to English and 11.16 (chrF2: 31.38, TER: 93.51) for English to Pnar, establishing the first quantitative benchmark for this language pair. Lexicalized reordering improves translation quality by 3.73 BLEU points for Pnar to English, reflecting the structural shift from the source language's SOV word order to the target language's SVO order, whereas MERT tuning degrades BLEU performance under low resource conditions. Finally, we analyze the remaining translation errors, including morphological out of vocabulary (OOV) words, long-distance reordering and Khasi code mixing and discuss future directions toward neural and multilingual machine translation for Pnar.

Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh · 0 citations
Open access Jul 2026

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has become a serious concern. This problem is even more complex in online spaces where people often switch between English and Swahili within the same sentence or conversation. Such code-switching makes it difficult for existing systems to accurately detect harmful content, especially because there is limited labeled data and much of the language used is informal and context-dependent. This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models. In this work, transformer-based models such as BERT and AfriBERTa are adapted to better understand mixed-language communication. The models are trained on a carefully prepared dataset made up of real social media posts that reflect how people actually write and speak online. These posts are manually labeled to capture both direct and subtle forms of hate speech, including expressions that are influenced by local culture and everyday slang. The findings show that fine-tuned models perform better than traditional machine learning approaches, especially in terms of accuracy and overall detection quality. They are also more effective at handling informal language, abbreviations, and mixed grammar structures. Beyond performance, the study also looks at fairness and bias, emphasizing the need for systems that are sensitive to cultural and linguistic diversity. Overall, this work shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.

Kipkebut Andrew, Jepkemei Betty · 0 citations
Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Mohammed Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Open access 2026

Using English-Based NLP Tools for Domain-Specific Text in Foreign Languages

Social scientists often machine-translate foreign-language texts into English and apply English-based natural language processing tools without systematically evaluating translation quality or annotation efficiency. To address this problem, this study provides evidence-based guidance for researchers applying English-centric natural language processing to domain-specific foreign-language corpora. We provide and empirically validate a structured framework combining multi-system machine translation evaluation and active learning for domain-specific text classification. Using 11,493 parallel Spanish and Arabic sentences aligned to English, we compare four machine translation systems (Google Translate, Deep, DeepL, OPUS) using SacreBLEU, METEOR, COMET, and BERTScore quality scores. Across languages and metrics, machine translation systems yield statistically comparable performance. We then evaluate eight active learning strategies using ConfliBERT for political conflict classification under a 20% annotation budget, corresponding to 1,155 samples from the training split. Binary classification exceeds F1 = 0.90, while QuadClass multi-class performance peaks around F $1~\approx ~0.75$ . The Ensemble Intersection strategy achieves the highest performance in 53% of tasks and often matches or surpasses full-dataset results using only a fraction of labeled data. These results provide a practical workflow for researchers using English-based natural language processing tools on foreign-language, domain-specific corpora.

Naif Alatrush, Luay Abdeljaber, Javier Osorio et al. · 0 citations
Preprint Jul 2026

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages. In this work, we quantify this tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages. Under cl100k_base (used by GPT-3.5 and GPT-4), Indian languages experience an average 8.0x tokenization tax relative to English, reaching 13.0x for Malayalam, reducing the effective context window to as little as 12% of that available to English users for equivalent semantic content. We identify the primary mechanism behind this disparity: failed byte-pair merges that leave text fragmented into single-byte tokens, with merge failure strongly correlating with tokenizer tax (Pearson r = 0.89). We further show that this phenomenon is not an inherent property of Indic scripts but a consequence of tokenizer design. Multilingual tokenizers such as XLM-R and OpenAI's o200k_base reduce the average Indic tokenizer tax by 73%, demonstrating that the disparity is largely remediable. Beyond token statistics, we quantify a practical consequence by showing that, under fixed context budgets, Indian-language documents preserve substantially less original content than equivalent English documents. Finally, we examine the relationship between tokenizer fertility and reading comprehension performance on the Belebele benchmark, finding that the apparent correlation is largely explained by language resource availability rather than tokenizer behavior alone.

Priyanshu Srivastava · 0 citations