CAUSAL EXPLAINABLE AI FOR STUDENT DROPOUT PREDICTION: DETECTING PROXY BIAS AND CAUSAL MISMATCH
DOI:
https://doi.org/10.70917/ijcisim-2026-2092Keywords:
algorithmic fairness, causal auditing framework, causal understanding, counterfactual intervention, causal mismatch, explainable AI, proxy biasAbstract
Although machine learning-based student dropout prediction models achieve high predictive performance, the socioeconomic biases and decision-making mechanisms latent within these black-box models remain insufficiently understood. Conventional eXplainable artificial intelligence (XAI) techniques rely on correlation-based global feature importance, which limits their ability to identify the true causal factors that substantively flip predictions. Therefore, this study proposes a "Causal Auditing" framework to quantitatively identify and correct biases. Utilizing a Portuguese university student dataset and a single ensemble baseline model (LightGBM), we generated hierarchical multivariate counterfactual scenarios that combined uncontrollable sensitive attributes (parents' background) with actionable variables (academic achievement and debt status). The results revealed a distinct "Causal Mismatch" phenomenon: although the correlation-based SHapley Additive exPlanations analysis assigned the highest importance to past first-semester grades, the core causal triggers that flipped the model decision from "Dropout" to "Graduate" were found to be second-semester curricular units and financial status. Notably, instead of forcing extreme high-risk students to achieve unreasonable grade improvements, merely resolving their debt issues and neutralizing the structural penalties associated with their parental backgrounds—when combined with minimal academic support (completing +1 additional course)—successfully flipped the predictions for the entire target cohort. These findings demonstrate that AI models adopt students' environmental disadvantages as a form of proxy bias (i.e., a "structural penalty”). The proposed framework provides a systematic methodology for verifying the transparency and causal fairness of AI in the educational domain, and contributes to the formulation of data-driven, hybrid student support policies in university administrations.