A Hybrid SVM–XGBoost Framework for Semantic Classification of Ancient Indian Cosmological Texts
DOI:
https://doi.org/10.70917/ijcisim-2026-4214Keywords:
Natural Language Processing, Ancient Text Classification, Support Vector Machine (SVM), XGBoost, Computational Humanities, Artificial Intelligence, Digital Heritage Preservation, Knowledge PreservationAbstract
The integration of artificial intelligence, natural language processing, and machine learning has created new opportunities for the analysis of traditional knowledge systems and ancient scientific texts, supporting digital learning, knowledge preservation, and cultural heritage research. This study presents a hybrid ensemble learning framework for the semantic classification of ancient Indian cosmological texts derived from the classical astronomical treatise Surya Siddhanta. The proposed framework transforms unstructured textual data into structured machine-readable representations using text preprocessing and Term Frequency–Inverse Document Frequency (TF-IDF) feature extraction techniques. Two machine learning models, namely Support Vector Machine (SVM) and Extreme Gradient Boosting (XGBoost), are combined using a weighted ensemble strategy to improve classification performance and model robustness. The dataset is categorized into multiple semantic classes including astronomy, cosmology, mathematics, time, history, and general knowledge. The proposed model is evaluated using standard performance metrics such as accuracy, precision, recall, and F1-score. Experimental results demonstrate that the hybrid ensemble model achieves high classification performance with an overall accuracy of 98%, macro F1-score of 0.92, and weighted F1-score of 0.98. The framework provides an efficient and interpretable solution for domain-specific text classification and contributes to computational humanities, digital heritage preservation, knowledge preservation, education, and sustainable access to ancient knowledge systems.