Sep 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 153-178· 0 citations
TL;DR
The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.
Abstract
Fake review detection in low-resource and code-mixed languages remains challenging due to informal writing styles, transliterated regional expressions, linguistic variability, and the limited availability of annotated datasets. This paper presents a hybrid LaBSE semantic and handcrafted feature fusion approach with machine learning for fake review detection in Roman Marathi code-mixed text. A real-time dataset comprising 2,287 Roman Marathi reviews collected from multiple online platforms is utilized to evaluate the proposed approach. The framework integrates opinion-mining features with multilingual semantic representations generated using Language-agnostic BERT Sentence Embedding (LaBSE) and handcrafted linguistic, behavioural, contextual, temporal, and metadata features to construct a comprehensive hybrid feature representation. The dataset is balanced using random oversampling and subsequently partitioned into training and testing subsets using an 80:20 ratio. Four machine learning classifiers, namely Random Forest, XGBoost, Support Vector Machine, and K-Nearest Neighbour, are evaluated using accuracy, precision, recall, and F1-score. Experimental results demonstrate that XGBoost achieves the best performance with 94.60% accuracy, 95.63% precision, 93.47% recall, and 94.54% F1-score, outperforming the other evaluated classifiers. The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.
The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.
Nikita Garg, Pritam Singh Negi· International Journal of Eng...· 0 citations
The development of internet-based platforms has had a major effect on how consumers make choices when shopping and how they view businesses. The increased availability of information about products via the same digital channels has resulted in an increase in fake user reviews, deceptive content intended to sway popular...
Ravi Pal, S. Prakash· International journal of com...· 0 citations
Automated Urdu hate-speech detection remains challenging because annotated resources are limited, orthography varies, and harmful meaning depends on both lexical and contextual cues. This study evaluates a leakage-safe lexical-semantic framework combining 5,000-dimensional word/bigram TF-IDF features with 768-dimension...
Haseeb Ullah, Husnain Saleem, Asia Kanwal et al.· Brain: Broad Research in Art...· 0 citations
Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.
Chirag D. Shah, Shailesh A. Chaudhari· International journal of com...· 0 citations
The rapid growth of e-commerce platforms has significantly increased the dependence on customer reviews, making the detection of fake reviews essential for preserving consumer trust and ensuring the credibility of online marketplaces.While substantial progress has been made for high-resource languages, limited progress...
F. S., A. M, Sobhana N. V.· International Conference Inn...· 0 citations
This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages by proposing a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual dat...
Satyapal Singh, Jarnail Singh, D. S· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.