AI/ML-Driven Graph Transformer and Reinforcement Learning for Last-Mile Logistics Optimization

Authors

  • Ramesh Nalluri School of Business, Woxsen University, Hyderabad – 500033, Telangana, India.
  • Sujit Singh Professor & Research Supervisor, School of Business, Woxsen University, Hyderabad, Telangana, India.
  • Venkata Reddy Muppani The Labour Relations Institute Hyderabad (TLRI Hyderabad), Hyderabad, Telangana – 500094, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-4608

Abstract

 Last-mile delivery represents up to 53% of total shipping costs, yet effective route prediction tools are still hard to find. A major issue in logistics is that couriers often do not follow the best routes. Local knowledge, time constraints, and human judgment create an average Kendall Rank Correlation (KRC) of only 0.60 between the best calculated routes and actual behaviour. Bridging this gap is crucial for improving delivery time estimates, dispatching, and planning. This research presents a Hybrid Graph Transformer-Reinforcement Learning (HGTRL) framework, a new hybrid deep learning model aimed at predicting real-world courier pickup sequences under spatial and temporal limits. Instead of assuming perfect behaviour, HGTRL learns from historical courier choices using the large LaDe dataset. The model structure includes a two-layer, eight-head Transformer encoder and a Long Short-Term Memory (LSTM) attention decoder. To effectively capture complex routing preferences, HGTRL uses a hybrid loss function that combines Maximum Likelihood Estimation (MLE) with the REINFORCE algorithm (L_hybrid = L_MLE + λL_RL, λ=0.3). This approach allows the model to imitate real behaviour while improving route sequences. Extensive evaluation covered 10.67 million packages in three diverse cities: Chongqing, Shanghai, and Yantai. While earlier models like DeepRoute and Graph2Route showed the benefits of learning from historical data, HGTRL sets a new benchmark for performance. Compared to the leading baseline, DRL4Route2, HGTRL made significant gains across all key metrics: KRC improved by 14.32 to 15.50 points, Location Square Deviation (LSD) dropped by 0.11 to 0.50, and Edit Distance (ED) went down by 0.17 to 0.32. Additionally, ablation studies showed that removing the reinforcement learning element alone led to a 3.2% decrease in KRC, stressing its importance. Ultimately, HGTRL’s precise sequence predictions allow for reliable estimated arrival times and adaptive order assignments, greatly boosting fleet efficiency and customer satisfaction in today's logistics systems.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-24

How to Cite

Ramesh Nalluri, Sujit Singh, & Venkata Reddy Muppani. (2026). AI/ML-Driven Graph Transformer and Reinforcement Learning for Last-Mile Logistics Optimization. International Journal of Computer Information Systems and Industrial Management Applications, 18(19s), 470–491. https://doi.org/10.70917/ijcisim-2026-4608

Issue

Section

Original Articles