Threshold-Calibrated XGBoost with TreeSHAP Explainability for Fraud Detection in Peer-to-Peer Digital Payments

Authors

  • Dhaval Chandarana Gyanmanjari Institute of Technology, Gyanmanjari Innovative University, Bhavnagar, Gujarat, India
  • Hitesh Nimbark Gyanmanjari Institute of Technology, Gyanmanjari Innovative University, Bhavnagar, Gujarat, IndiaHitesh Nimbark

DOI:

https://doi.org/10.70917/ijcisim-2026-4041

Keywords:

Fraud detection, explainable artificial intelligence, SHAP, XGBoost, decision-threshold calibration, class imbalance, peer-to-peer payments, deployment readiness

Abstract

The rapid diffusion of peer-to-peer (P2P) digital payment infrastructures has intensified the demand for fraud detection systems that are simultaneously accurate, interpretable and operationally deployable. A substantial body of recent work optimises headline accuracy while leaving decision-threshold behaviour, minority-class recall, attribution transparency and deployment readiness largely unaddressed. This paper presents an end-to-end, explainable and threshold-aware fraud detection framework and evaluates it on a banking transaction corpus of 10,000 records exhibiting a moderate class imbalance of 2.52:1 (28.38 % fraudulent). Three structurally distinct learners — K-Nearest Neighbours (KNN), a one-dimensional Convolutional Neural Network (CNN) and Extreme Gradient Boosting (XGBoost) — are trained on an identical preprocessing pipeline and an identical stratified 80/20 partition, so that observed differences are attributable to inductive bias rather than to experimental artefacts. At the conventional operating point (τ = 0.50) the CNN attains the highest accuracy (0.7170) yet a recall of only 0.0035, flagging just 2 of 568 fraudulent transactions and rendering it operationally inert despite recording the highest area under the precision–recall curve (AUPRC = 0.4813). This dissociation between ranking quality and decision quality is the central empirical finding of the study: it demonstrates that accuracy, and even threshold-free ranking metrics, can conceal complete failure at the deployed operating point. Recalibrating the XGBoost decision threshold to τ* = 0.135 raises recall from 0.3151 to 0.8451 and the F1-score from 0.3825 to 0.5215, at the cost of a 4.2-fold increase in false positives. TreeSHAP attributions identify PreviousFraud, CreditScore and TransactionAmount as the dominant drivers, in agreement with independently computed point-biserial correlations. A Gradio interface with automated PDF report generation demonstrates real-time applicability. The framework is positioned against fifteen 2025–2026 studies through a structured systematic review, and its precision–recall trade-off is critically appraised rather than presented as an unqualified improvement.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-31

How to Cite

Dhaval Chandarana, & Hitesh Nimbark. (2026). Threshold-Calibrated XGBoost with TreeSHAP Explainability for Fraud Detection in Peer-to-Peer Digital Payments. International Journal of Computer Information Systems and Industrial Management Applications, 18(13s), 169–191. https://doi.org/10.70917/ijcisim-2026-4041

Issue

Section

Original Articles