A Feature-Optimized Hybrid Supervised Ensemble Learning Framework for Early Heart Attack Prediction Using Artificial Intelligence and Machine Learning
DOI:
https://doi.org/10.70917/ijcisim-2026-3650Keywords:
heart-attack prediction, coronary heart disease, cardiovascular risk, artificial intelligence, machine learning, ensemble learning, stacking classifier, explainable AI, clinical decision supportAbstract
Background: While machine-learning models capture the complex interactions of clinical variables, small cardiovascular datasets may exhibit optimistic validation, data leakage, and mislabeling of outcomes.
Objective: This study sought to implement and assess a leakage-controlled, hybrid, supervised ensemble to classify the presence of coronary heart disease in the UCI Cleveland cohort. Even though the proposed title uses "heart attack prediction," the endpoint of this study is the determination of prevalent coronary disease at evaluation, and not the prediction of a future myocardial infarction.
Methods: The Cleveland dataset housed in the public domain consists of 303 samples and 13 features. The original 0–4 categorical variable of diagnostic severity was transformed into a binary variable with 0 for no disease and 1 for any diagnosis (1–4). The median and mode were imputed. Standardization and one-hot encoding were fitted only to training folds. Logistic regression, support vector machines, random forests, XGBoost, and LightGBM were optimized through stratified five-fold cross-validation. Optimized, weighted soft voting and logistic-regression stacking utilized out-of-fold probabilities. Evaluation of performance was conducted with an untouched 20% test set, and through two iterations of five-fold cross-validation. Discrimination, calibration, and uncertainty bootstrap were assessed, along with confusion-matrix metrics, paired tests, SHAP analysis, and permutation importance.
Results: The cohort consisted of 164 observations without disease, 139 with disease, and had six missing observations. From the 61-observation test set, stacking attained an accuracy of 91.8% with sensitivity of 92.9% and specificity of 90.9%, with an F1 score of 0.912, an MCC score of 0.836, and ROC-AUC of 0.960. This resulted in 30 true negatives, 3 false positives, 2 false negatives, and 26 true positives. For weighted voting, Brier score and ROC-AUC showed the best results with 0.0849 and 0.961, respectively. When comparing the rest of the models, the greatest mean ROC-AUC (0.898±0.042) was attained with weighted voting, but results of the seven models were comparable (Friedman p=0.628). The most important variables for this dataset were the number of major vessels, the chest-pain type, and the thallium-test type, and the induced angina and the ST depression and the maximum heart rate.
Conclusion: The use of heterogeneous ensembles on this small dataset proved discrimination and clinically relevant sensitivities. However, tuned single models have not shown any statistically relevant results. Before