A Multi-Dimensional Benchmarking Framework For Evaluating Agentic AI Reliability And Decision Quality In Enterprise Supply Chain Systems
DOI:
https://doi.org/10.70917/ijcisim-2026-4891Keywords:
Agentic AI, benchmarking framework, supply chain decision intelligence, explainability, governance, mixed-methods validationAbstract
Enterprises are deploying agentic AI systems, understood as autonomous multi-step tool-using agents, into supply chain decision workflows spanning demand forecasting, inventory allocation, transportation routing, and exception handling. Unlike traditional analytics or single-shot machine-learning models, these systems make chained, semi-autonomous decisions with limited human-in-the-loop checkpoints. The literature on AI in supply chain management is concentrated on forecasting accuracy and largely predates the agentic shift; the AI-benchmarking literature is domain-agnostic and disconnected from supply chain decision quality; the trustworthy-AI literature is framework-generic. Practitioners consequently have no standardised, defensible way to answer a simple question before and after deployment: is this agentic AI system reliable enough, transparent enough, and impactful enough to trust with supply chain decisions at scale? This paper contributes a multi-dimensional benchmarking framework that unifies technical reliability, governance and trust, and operational business impact dimensions for agentic AI in supply chain contexts. The paper couples the framework with a two-phase validation design (expert-panel Delphi followed by PLS-SEM empirical testing), a simulated Round 1 Delphi run on twenty-two constructs with a fifteen-expert synthetic panel producing per-item CVR values that surface exactly which framework dimensions have strong practitioner consensus and which need Round 2 refinement, a worked case study of a mid-size retailer deploying an agentic AI for demand forecasting and inventory allocation, and a five-scenario sensitivity analysis showing how the composite score discriminates between deployments with different governance-maturity profiles. Every synthetic value carries a visible synthetic label. The metric set is justified against the SCOR supply-chain performance-attribute framework so that the framework composes with rather than competes against established supply-chain measurement practice.