Hybrid LaBSE Semantic and Handcrafted Feature Fusion with Machine Learning for Fake Review Detection in Roman Marathi Code-Mixed Text
DOI:
https://doi.org/10.70917/ijcisim-2026-5423Keywords:
Fake Review Detection, Roman Marathi, Code-Mixed Text, LaBSE, Semantic EmbeddingsAbstract
Fake review detection in low-resource and code-mixed languages remains challenging due to informal writing styles, transliterated regional expressions, linguistic variability, and the limited availability of annotated datasets. This paper presents a hybrid LaBSE semantic and handcrafted feature fusion approach with machine learning for fake review detection in Roman Marathi code-mixed text. A real-time dataset comprising 2,287 Roman Marathi reviews collected from multiple online platforms is utilized to evaluate the proposed approach. The framework integrates opinion-mining features with multilingual semantic representations generated using Language-agnostic BERT Sentence Embedding (LaBSE) and handcrafted linguistic, behavioural, contextual, temporal, and metadata features to construct a comprehensive hybrid feature representation. The dataset is balanced using random oversampling and subsequently partitioned into training and testing subsets using an 80:20 ratio. Four machine learning classifiers, namely Random Forest, XGBoost, Support Vector Machine, and K-Nearest Neighbour, are evaluated using accuracy, precision, recall, and F1-score. Experimental results demonstrate that XGBoost achieves the best performance with 94.60% accuracy, 95.63% precision, 93.47% recall, and 94.54% F1-score, outperforming the other evaluated classifiers. The findings demonstrate the effectiveness of combining multilingual semantic information with explicit linguistic and contextual characteristics for identifying deceptive reviews in Roman Marathi code-mixed environments and establish an initial benchmark for fake review detection in this low-resource setting.