IMPROVED MULTILINGUAL HATE SPEECH DETECTION THROUGH INDEXED COPHENETIC MULTILAYEREDEXTREME MACHINE RELYING ON N-GRAMS AND HIDDEN MARKOV MODEL
DOI:
https://doi.org/10.70917/ijcisim-2026-5455Keywords:
Multilingual Hate Speech Detection, Machine Learning, Natural Language Processing, Multilayered Extreme Learnt Machine, Character level N-grams, Baum-Welch Hidden Markov, Cophenetic CorrelationAbstract
This study consider hate speech detection in multiple languages such as English, Tamil, and Tanglish by achieving higher accuracy, precision, recall, and minimal training time utilizing machine learning based Natural Language Processing (NLP) techniques. Proposed Indexed Cophenetic Multilayered Extreme Learning based on N-Grams and Hidden Markov (ICMEL-NGHM) technique introduced for achieving multilingual hate speech detection. Proposed technique and two existing methods are compared for effectiveness in hate speech identification with diverse metrics. Accuracy of proposed technique for detecting hate speech in English, Tamil, and Tanglish classes is 97.89%, 97.32%, and 96.34%,. Precision of proposed technique in detecting hate speech for English, Tamil, and Tanglish classes is 97.34%, 95.34%, and 94.32%. For English, Tamil, and Tanglish groups, Recall in identifying hate speech of proposed ICMEL-NGHM is 99.23%, 96.34%, and 98.23%. Proposed technique detects hate speech in 12.35 ms for English class, 10.55 ms for Tamil class, and 9.45 ms for Tanglish class. Results of proposed technique is better than state-of-the-art approaches.