Implementation and Performance Evaluation of a Hybrid Machine Learning Framework for Web Content Mining

Authors

  • S. Zafar Mehdi Kazmi Department of Computer Application, Integral University, Lucknow, Uttar Pradesh, India.
  • Md. Faizan Farooqui Department of Computer Application, Integral University, Lucknow, Uttar Pradesh, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-3306

Keywords:

Hybrid Machine Learning, Web Content Mining, Knowledge Discovery in Databases, Named Entity Recognition, Transformer NLP, Computer Vision, Ensemble Learning, Transfer Learning Dynamic Content Extraction & Knowledge Graph.

Abstract

The rapid growth of web data has opened new doors for knowledge discovery in various domains, however web data extraction traditionally performed on web contents are increasingly facing cumulative challenges such as dynamic JavaScript rendering, sophisticated anti-bot counter measures, heterogeneous content structures, and the brittleness of rules-based templates to routine website evolution. The Hybrid Framework for Knowledge Discovery and Web Content Extraction (HFKDE) was designed, implemented, and extensively tested in a production environment, as presented in this paper. The framework includes a computer vision model for structural page analysis, a transformer-based NLP for semantic analysis, an ensemble model trained over a massive corpus for content classification, and RL for crawler policy optimisation. The framework is evaluated with four baselines which are traditional rule-based scraping, standard KDD (Knowledge Discovery in Databases) process, pure ML based frameworks, ontology-based frameworks across six domains namely news, e-commerce, academic repositories, social media, government portals and healthcare. An extensive testing carried on a 32,000-sample benchmark for over a duration of 30 days of continuous deployment shows that the HFKDE achieves 94-96 percent precision and 93-95 percent recall, with an average cross-domain generalization score of 92.8 percent. Thus, it outperforms all baselines by at least 7.9 percentage points. Using transfer learning for activity recognition decreases labeling effort by 90%, while the hybrid architecture reduces operational maintenance effort by 84%. Basing on this, principled multi-paradigm integration can resolve the gap between theoretical ML progress and the practical challenges of mining large-scale web data.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-18

How to Cite

S. Zafar Mehdi Kazmi, & Md. Faizan Farooqui. (2026). Implementation and Performance Evaluation of a Hybrid Machine Learning Framework for Web Content Mining. International Journal of Computer Information Systems and Industrial Management Applications, 18(8s), 718–738. https://doi.org/10.70917/ijcisim-2026-3306

Issue

Section

Original Articles