An Optimised ML Framework for CVD Detection, Classification and Early Prediction Using Big Data

Authors

  • Amar Paul Singh Dept of Computer Science Himachal Pradesh University-Shimla 171005 INDIA
  • Yogesh Mohan Dept of Computer Science Himachal Pradesh University-1Shimla 171005 INDIA

DOI:

https://doi.org/10.70917/ijcisim-2026-5019

Keywords:

Optimised, ML framework, CVD detection, classification, early prediction, big data

Abstract

Cardiovascular disease (CVD) remains one of the major causes of morbidity and mortality worldwide, creating a strong need for accurate, scalable, and interpretable computational methods for early detection and risk prediction. This study proposes an Optimized Ensemble Machine Learning Framework (OEMLF) for CVD detection, classification, and early prediction using integrated cardiovascular big data. The framework combines heterogeneous records from the Framingham Heart Study, UCI Heart Disease, Kaggle Heart Disease, and NHANES datasets into a unified repository. The proposed dataflow incorporates data harmonization, duplicate removal, missing-value imputation, outlier handling, normalization, and SMOTE-based class balancing. Clinically meaningful features are subsequently generated, while Mutual Information, SHAP-based feature importance, and Recursive Feature Elimination are jointly employed for hybrid feature selection. Bayesian hyperparameter optimization is then used to tune Logistic Regression, Random Forest, XGBoost, and LightGBM models. Their complementary predictions are integrated using a weighted soft-voting ensemble to improve robustness and generalization across heterogeneous patient records. Experimental evaluation demonstrates that the proposed OEMLF achieves 98.84% accuracy, 98.8% precision, 98.7% recall, 98.7% F1-score, and 99.8% ROC-AUC, outperforming the individual baseline models. The framework also reduces the feature representation from 45 standardized attributes to 40 informative features while eliminating missing values and improving class balance from 1:3.4 to approximately 1:1.1. Computational analysis indicates that although ensemble training requires additional processing time, the framework maintains a low patient-level prediction time of approximately 4.8 ms. SHAP-based explainability further identifies influential cardiovascular factors and provides patient-specific explanations for model decisions. The results indicate that the proposed framework can provide an effective combination of predictive accuracy, scalability, computational efficiency, and interpretability for large-scale cardiovascular screening and clinical decision-support applications.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-23

How to Cite

Amar Paul Singh, & Yogesh Mohan. (2026). An Optimised ML Framework for CVD Detection, Classification and Early Prediction Using Big Data. International Journal of Computer Information Systems and Industrial Management Applications, 18(19s), 260–282. https://doi.org/10.70917/ijcisim-2026-5019

Issue

Section

Original Articles