An Explainable Machine Learning Framework for Credit Risk Assessment and Optimization in the Indian Banking Sector

Authors

  • Amit Kumar Agrawal Department of Computer Science, Faculty of Information Technology and Computer Science, Parul University, Vadodara, Gujrat, India.
  • Vaibhav Gandhi Department of Computer Science, Faculty of Information Technology and Computer Science, Parul University, Vadodra, Gujrat, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-3503

Keywords:

Explainable Machine Learning, Credit Risk Assessment, Indian Banking Sector, Non-Performing Assets, XGBoost, SHAP, Ensemble Learning, Financial Risk Management, Priority Sector Lending, Credit Scoring

Abstract

Credit risk assessment remains a foundational function of commercial banking, yet the Indian banking sector's persistent non-performing asset (NPA) burden, heterogeneous borrower base, and expanding priority-sector and microfinance lending create distinct challenges that generic, globally trained credit scoring models do not adequately address. This paper proposes an Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer, and evaluates the framework using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India. The framework incorporates SMOTE-ENN based class-imbalance handling to address the low base default rate typical of retail lending portfolios, and an explainability-guided feature refinement step that uses SHAP attributions to iteratively prune low-value features and retain regulator-interpretable risk factors. The proposed model was evaluated on a dataset of 42,860 loan accounts and benchmarked against five baseline models. Results show that the proposed explainable ensemble achieves 93.2% accuracy, 91.0% precision, 89.4% recall, and a 90.2% F1-score, exceeding the strongest baseline (LightGBM) by 5.6 percentage points in F1-score, with an area under the ROC curve (AUC) of 0.967. Segment-wise analysis reveals materially higher default risk concentration in regional rural banks and small finance institutions relative to public and private sector banks, with debt-to-income ratio, credit bureau (CIBIL) score, and repayment delinquency history emerging as the most influential predictors across segments.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-21

How to Cite

Amit Kumar Agrawal, & Vaibhav Gandhi. (2026). An Explainable Machine Learning Framework for Credit Risk Assessment and Optimization in the Indian Banking Sector. International Journal of Computer Information Systems and Industrial Management Applications, 18(9s), 802–813. https://doi.org/10.70917/ijcisim-2026-3503

Issue

Section

Original Articles