Skip to content
Review Open access

Hybrid LaBSE Semantic and Handcrafted Feature Fusion with Machine Learning for Fake Review Detection in Roman Marathi Code-Mixed Text

Sep 2026 · International journal of computer information systems and industrial management applications · Vol 18, pp. 153-178 · 0 citations

TL;DR

The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.

Abstract

Fake review detection in low-resource and code-mixed languages remains challenging due to informal writing styles, transliterated regional expressions, linguistic variability, and the limited availability of annotated datasets. This paper presents a hybrid LaBSE semantic and handcrafted feature fusion approach with machine learning for fake review detection in Roman Marathi code-mixed text. A real-time dataset comprising 2,287 Roman Marathi reviews collected from multiple online platforms is utilized to evaluate the proposed approach. The framework integrates opinion-mining features with multilingual semantic representations generated using Language-agnostic BERT Sentence Embedding (LaBSE) and handcrafted linguistic, behavioural, contextual, temporal, and metadata features to construct a comprehensive hybrid feature representation. The dataset is balanced using random oversampling and subsequently partitioned into training and testing subsets using an 80:20 ratio. Four machine learning classifiers, namely Random Forest, XGBoost, Support Vector Machine, and K-Nearest Neighbour, are evaluated using accuracy, precision, recall, and F1-score. Experimental results demonstrate that XGBoost achieves the best performance with 94.60% accuracy, 95.63% precision, 93.47% recall, and 94.54% F1-score, outperforming the other evaluated classifiers. The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.

Read PDF

Similar papers

Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Review Open access Aug 2026

DESIGN AND ANALYSIS OF FAKE REVIEW DETECTION IN URDU LANGUAGE USING TRANSFORMER-AUGMENTED RoBERTa + LSTM HYBRID MACHINE LEARNING MODEL

The development of internet-based platforms has had a major effect on how consumers make choices when shopping and how they view businesses. The increased availability of information about products via the same digital channels has resulted in an increase in fake user reviews, deceptive content intended to sway popular...

Ravi Pal, S. Prakash · 0 citations
Open access Sep 2026

PSO-Based Hybrid Lexical-Semantic Feature Selection and Ensemble Learning Framework for Urdu Hate Speech Detection

Automated Urdu hate-speech detection remains challenging because annotated resources are limited, orthography varies, and harmful meaning depends on both lexical and contextual cues. This study evaluates a leakage-safe lexical-semantic framework combining 5,000-dimensional word/bigram TF-IDF features with 768-dimension...

Haseeb Ullah, Husnain Saleem, Asia Kanwal et al. · 0 citations
Open access Aug 2026

Linguistically Informed Machine Learning for Gujarati–English Code-Mixed Sentiment Classification: A Comparative Study of Feature Fusion Strategies

Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.

Chirag D. Shah, Shailesh A. Chaudhari · 0 citations
Conference Aug 2026

Malayalam Fake Review Detection Using Hybrid Feature Engineering and Stacked Ensemble Learning

The rapid growth of e-commerce platforms has significantly increased the dependence on customer reviews, making the detection of fake reviews essential for preserving consumer trust and ensuring the credibility of online marketplaces.While substantial progress has been made for high-resource languages, limited progress...

F. S., A. M, Sobhana N. V. · 0 citations
Review Open access Sep 2026

Automated Sentiment Analysis of Hindi Text using Machine Learning Techniques: A Lightweight and Scalable Framework for Regional Language NLP

This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages by proposing a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual dat...

Satyapal Singh, Jarnail Singh, D. S · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.