A Hybrid SVM–XGBoost Framework for Semantic Classification of Ancient Indian Cosmological Texts

Authors

  • Ranjana Neb Department of Languages, Swami Vivekanand Subharti University, Meerut-250005, Uttar Pradesh, India.
  • Seema Sharma Department of Languages, Swami Vivekanand Subharti University, Meerut-250005, Uttar Pradesh, India.
  • Ashish Kumar Department of Languages, Swami Vivekanand Subharti University, Meerut-250005, Uttar Pradesh, India.
  • Nishi Raghav Department of Languages, Swami Vivekanand Subharti University, Meerut-250005, Uttar Pradesh, India.
  • Geeta Kapil Department of Hindi and Modern Indian Languages, Banasthali Vidyapith, P.O. Banasthali Vidyapith-304022 (Rajasthan)
  • Sanjita Verma Central Hindi Department, Tribhuvan University Kathmandu, Nepal.

DOI:

https://doi.org/10.70917/ijcisim-2026-4214

Keywords:

Natural Language Processing, Ancient Text Classification, Support Vector Machine (SVM), XGBoost, Computational Humanities, Artificial Intelligence, Digital Heritage Preservation, Knowledge Preservation

Abstract

The integration of artificial intelligence, natural language processing, and machine learning has created new opportunities for the analysis of traditional knowledge systems and ancient scientific texts, supporting digital learning, knowledge preservation, and cultural heritage research. This study presents a hybrid ensemble learning framework for the semantic classification of ancient Indian cosmological texts derived from the classical astronomical treatise Surya Siddhanta. The proposed framework transforms unstructured textual data into structured machine-readable representations using text preprocessing and Term Frequency–Inverse Document Frequency (TF-IDF) feature extraction techniques. Two machine learning models, namely Support Vector Machine (SVM) and Extreme Gradient Boosting (XGBoost), are combined using a weighted ensemble strategy to improve classification performance and model robustness. The dataset is categorized into multiple semantic classes including astronomy, cosmology, mathematics, time, history, and general knowledge. The proposed model is evaluated using standard performance metrics such as accuracy, precision, recall, and F1-score. Experimental results demonstrate that the hybrid ensemble model achieves high classification performance with an overall accuracy of 98%, macro F1-score of 0.92, and weighted F1-score of 0.98. The framework provides an efficient and interpretable solution for domain-specific text classification and contributes to computational humanities, digital heritage preservation, knowledge preservation, education, and sustainable access to ancient knowledge systems.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-04

How to Cite

Ranjana Neb, Seema Sharma, Ashish Kumar, Nishi Raghav, Geeta Kapil, & Sanjita Verma. (2026). A Hybrid SVM–XGBoost Framework for Semantic Classification of Ancient Indian Cosmological Texts. International Journal of Computer Information Systems and Industrial Management Applications, 18(14s), 301–313. https://doi.org/10.70917/ijcisim-2026-4214

Issue

Section

Original Articles