Segment-Aware Credit Risk Modelling: A Multi-Algorithm Benchmark

Authors

  • Deep Shikha Faculty of Sciences (Computer Science), Suresh Gyan Vihar University, Jaipur, India
  • Himanshu Vasnani Department of Mechanical Engineering, Suresh Gyan Vihar University, Jaipur, India
  • Amit Kumar Verma Arrivo Edu Private Limited, Ahmedabad, Gujarat, India

DOI:

https://doi.org/10.70917/ijcisim-2026-3833

Keywords:

Credit risk modelling, Borrower segmentation, Machine learning, Recall optimization, Algorithm–segment affinity, Leakage-free validation

Abstract

Today, most lenders use the same predictive model across their whole loan portfolio, whether the applicant is a salaried employee with a regular income or a gig worker whose earnings are unpredictable. This homogeneous approach does reasonably well at the aggregate level but systematically underpredicts default risk for subgroups that differ from the typical borrower profile. In this paper we pose a simple question: does training a separate machine learning model for each distinct group of borrowers improve the ability to detect defaults within those groups? We classify the 733,451 leakage-free LendingClub loans (2013–2015) into four groups, based on two binary dimensions: financial capacity (annual income and debt-to-income ratio) and behavioral stability (employment length above or below three years). The resulting groups are Prime, Transitional, Constrained, and High-Risk. For each segment, we trained five algorithms (Logistic Regression, Random Forest, XGBoost, CatBoost, LightGBM) independently, and benchmarked against a non-segmented baseline. Recall improves in 13 out of 20 algorithm-segment combinations with the largest gains of +0.0851 (Prime, Random Forest), +0.0564 (Constrained, Logistic Regression) and +0.0555 (High-Risk, LightGBM). All four best-per-segment gains are statistically significant in both DeLong and McNemar’s tests (p < 0.001). Besides the headline numbers, we document a reproducible pattern, which we call algorithm-segment affinity, in which the structure of a borrower segment reliably predicts which algorithm will be best within that segment. We also see that gradient boosting models require about 100,000-150,000 training records per segment to consistently outperform simpler alternatives. Implications for Indian FinTech and the NBFCs functioning under the RBI Model Risk Management Guideines also discussed.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-28

How to Cite

Deep Shikha, Himanshu Vasnani, & Amit Kumar Verma. (2026). Segment-Aware Credit Risk Modelling: A Multi-Algorithm Benchmark. International Journal of Computer Information Systems and Industrial Management Applications, 18(11s), 1089–1100. https://doi.org/10.70917/ijcisim-2026-3833

Issue

Section

Original Articles