Optimized malware Detection Model Employing Data Balancing and Ensemble Machine Learning Approaches
DOI:
https://doi.org/10.70917/ijcisim-2026-5504Keywords:
Malware detection, random forest classifier, feature selection, extra tree classifier, resource constrained device, execution timeAbstract
The rapid proliferation of malware presents serious challenges to digital security, especially for resource constrained devices where efficiency and memory usage are critical. This study proposes an efficient machine learning based malware detection framework using the Kaggle Malware Detection dataset. Preprocessing includes missing value handling, categorical label encoding, numeric feature standardization, and five fold cross validation. Feature selection is performed using the Extra Trees Classifier based on Gini impurity, followed by data balancing through under sampling and over sampling techniques. A wide range of algorithms is evaluated, including Logistic Regression, Support Vector Machine, K Nearest Neighbors, Gaussian and Bernoulli Naive Bayes, Decision Tree, Random Forest, XGBoost, Gradient Boosting, LightGBM, AdaBoost, CatBoost, Histogram based Gradient Boosting, and Extra Trees. Ensemble strategies such as Voting and Stacking classifiers are also explored. Experimental results show Random Forest achieves 98.97% accuracy without feature selection, while the balanced Stacking Classifier attains 99.90% accuracy. Explainable AI techniques, LIME and SHAP, provide feature level insights. A Flask based web application enables secure user interaction, real time preprocessing, prediction visualization, and classification of inputs as legitimate or malware.