Automated Sentiment Analysis of Hindi Text using Machine Learning Techniques: A Lightweight and Scalable Framework for Regional Language NLP

Authors

  • Satyapal Singh Department of Artificial Intelligence & Data Science, S.E.A. College of Engineering & Technology, Bengaluru, Karnataka, India.
  • Jarnail Singh University School of Computer Applications and Technology, Rayat Bahra Professional University, Hoshiarpur, Punjab, India.
  • Dhivya Karunya S. Department of Artificial Intelligence & Data Science, S.E.A. College of Engineering & Technology, Bengaluru, Karnataka, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-5312

Keywords:

Hindi Sentiment Analysis, Regional Languages, TF-IDF Feature Extraction, Natural Language Processing, Machine Learning, Multinomial Naive Bayes, Logistic Regression, Low-Resource NLP, Opinion Mining

Abstract

The paper presents a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual data. This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages. Here methodology encompasses a systematic pipeline comprising data preprocessing, Term Frequency–Inverse Document Frequency (TF-IDF) feature extraction with unigram and bigram representations, and supervised classification using Multinomial Naive Bayes (MNB) and Logistic Regression (LR) algorithms. Experiments on the IIT Patna Movie Reviews Hindi Sentiment Analysis dataset (2,480 training, 310 validation, and 310 test samples spanning the negative, neutral, and positive classes) demonstrate that Logistic Regression substantially outperforms the Naive Bayes baseline, achieving 88.19% training accuracy and 54.84% test accuracy against 67.82% and 43.23% for MNB, together with a higher micro-averaged Receiver Operating Characteristic – Area Under the Curve (ROC–AUC) (0.742 vs. 0.653). Class-wise analysis shows that positive sentiment is the easiest to detect (LR F1 = 0.63), while the neutral class remains the most challenging (LR F1 = 0.39). Proposed framework offers a computationally efficient, interpretable, and scalable solution for Hindi sentiment analysis, with direct applicability to social media monitoring, customer feedback analysis, and opinion mining in regional language ecosystems. Experimental codes is made available as a fully executable Google Colab notebook to ensure reproducibility and facilitate future research extensions.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-01

How to Cite

Satyapal Singh, Jarnail Singh, & Dhivya Karunya S. (2026). Automated Sentiment Analysis of Hindi Text using Machine Learning Techniques: A Lightweight and Scalable Framework for Regional Language NLP. International Journal of Computer Information Systems and Industrial Management Applications, 18(21s), 888–898. https://doi.org/10.70917/ijcisim-2026-5312

Issue

Section

Original Articles