MACHINE LEARNING CLASSIFICATION OF COMMUTING ACCIDENT CLAIM OUTCOMES

Authors

  • Mohd Zaki Awang Chek Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA (UiTM) Perak Branch, Tapah Campus, 35400 Tapah, Perak, Malaysia
  • Hilmi Aziz Bukhori Faculty of Mathematics and Natural Sciences, Universitas Brawijaya, Malang 65145, Indonesia

DOI:

https://doi.org/10.70917/ijcisim-2026-3646

Keywords:

Commuting accidents, claim outcome prediction, machine learning, gradient boosting, decision tree, occupational safety, social security, insurance analytics, Malaysia

Abstract

Background: Commuting accident claims impose substantial occupational, social, and administrative burdens on social security systems. Despite the increasing availability of administrative claims databases, evidence on machine-learning approaches for predicting commuting accident claim outcomes remains limited, particularly in the context of national social insurance programmes.
Methods: This study developed and compared five supervised machine-learning algorithms for classifying commuting accident claim outcomes using 37,828 commuting accident claim records obtained from the Social Security Organisation (SOCSO), Malaysia. The analytical dataset comprised 57 demographic, accident, injury, and administrative variables, with approved claims representing 61.5% of all observations. The dataset was randomly partitioned into 70% training (n = 26,480) and 30% validation (n = 11,348) subsets using SAS Enterprise Miner. Five classification models were evaluated: Decision Tree, Interactive Decision Tree, Gradient Boosting, Logistic Regression, and Neural Network. Model performance was assessed using recall, receiver operating characteristic area under the curve (ROC-AUC), misclassification rate, and average squared error.
Results: Gradient Boosting achieved the highest predictive performance, producing a validation ROC-AUC of 0.998 and the lowest validation misclassification rate (0.0121). The Interactive Decision Tree demonstrated the greatest model stability, with minimal differences between the training and validation datasets in misclassification rate (0.000009) and average squared error (0.000022), while providing transparent and interpretable decision rules. Variable importance analysis consistently identified Reason_Code as the strongest predictor across several models. However, because this variable may be generated during the claim adjudication process, its inclusion requires careful verification to exclude potential target leakage and ensure unbiased model evaluation.
Conclusions: The findings demonstrate the potential of machine-learning techniques for accurately classifying commuting accident claim outcomes using large-scale social security administrative data. Among the evaluated models, Gradient Boosting provided the highest predictive accuracy, whereas the Interactive Decision Tree offered a favourable balance between predictive performance and interpretability for operational decision support. Nevertheless, future studies should employ leakage-controlled feature selection, independent or temporal validation datasets, calibration assessment, and fairness analyses before these models are considered for implementation in operational claim management systems.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-24

How to Cite

Mohd Zaki Awang Chek, & Hilmi Aziz Bukhori. (2026). MACHINE LEARNING CLASSIFICATION OF COMMUTING ACCIDENT CLAIM OUTCOMES. International Journal of Computer Information Systems and Industrial Management Applications, 18(10s), 738–749. https://doi.org/10.70917/ijcisim-2026-3646

Issue

Section

Original Articles