A Generalized Spatio-Temporal Stacked Ensemble model for Urban Crime Type Classification Across Multi-City Crime Datasets
DOI:
https://doi.org/10.70917/ijcisim-2026-2913Keywords:
Crime Type Classification, Ensemble Learning, Stacking, Machine Learning, Crime Analytics, Urban Crime Data, Spatio-Temporal Features, Random Forest, Support Vector Machine, Gradient BoostingAbstract
Urban crime datasets contain large volumes of spatial, temporal, and categorical information that can be used to support crime analysis and decision-making. Accurate crime type classification remains challenging due to class imbalance, nonlinear relationships among features, and variations in crime patterns across different cities. Many existing studies rely on individual machine learning models and evaluate their performance on a single dataset, which limits the assessment of model generalization. This study presents a spatio-temporal stacked ensemble model for crime type classification using two large-scale urban crime datasets: the Chicago Crime Dataset and the Los Angeles Crime Dataset. The proposed model combines Random Forest, Logistic Regression, and Support Vector Machine as base classifiers, while Gradient Boosting is used as a meta-classifier to generate the final prediction. A preprocessing pipeline consisting of data cleaning, feature encoding, feature selection, feature scaling, and stratified sampling is applied to improve data quality and maintain consistency across datasets. The proposed model is evaluated using accuracy, precision, recall, and F1-score. On the Chicago Crime Dataset, the model achieved an accuracy of 99.9%, precision of 1.00, recall of 0.99, and F1-score of 0.99. On the Los Angeles Crime Dataset, it achieved an accuracy of 99.2%, precision of 0.99, recall of 0.99, and F1-score of 0.99. The results were compared with Logistic Regression, Random Forest, Support Vector Machine, and Gradient Boosting classifiers. The stacked ensemble produced the highest classification performance on both datasets and maintained consistent results across different crime distributions. The experimental findings indicate that combining heterogeneous classifiers through stacking can improve prediction accuracy and reduce the limitations of individual models. The proposed approach can support crime analysis systems that require reliable classification of crime incidents across different urban environments.