Mathematical Foundations of Reinforcement Learning and Stochastic Control Systems

Authors

  • Shirish Prabhakarrao Kulkarni Ajeenkya D Y Patil University, lohagaon Pune, Maharashtra
  • C. Ashwini Department of Computer Science and Engineering, SRM Institute of science and Technology, Rampuram Campus, Chennai, Tamil Nadu
  • M. Buvanasankari Department of Mathematics, Nehru Institute of Engineering and Technology, Coimbatore, Tamil Nadu.
  • K. Ramesh Mathematics, Nehru Institute of Engineering and technology Coimbatore, Tamil Nadu.
  • Rajat Verma Department of Computer Science and Engineering, Pranveer Singh Institute of Technology, Kanpur, Uttar Pradesh, India.
  • Abhinav Jha Amity School of Engineering and Technology, Amity University Patna, Patna, Bihar.

DOI:

https://doi.org/10.70917/ijcisim-2026-5162

Keywords:

Reinforcement Learning, Stochastic Control, Markov Decision Processes, Bellman Equation, Stochastic Approximation, Temporal-Difference Learning, Policy Gradient Methods, Hamilton–Jacobi–Bellman Equation

Abstract

Reinforcement learning (RL) and stochastic optimal control both address a single underlying mathematical problem: how an agent should select actions over time, under uncertainty, to optimize a cumulative reward or cost criterion. This paper reviews the shared mathematical scaffolding uniting these two fields, tracing the progression from Bellman's dynamic programming and the Markov decision process (MDP) formalism through stochastic approximation theory, temporal-difference learning, policy-gradient methods, actor-critic architectures, and the Hamilton–Jacobi–Bellman (HJB) equation governing continuous-time stochastic control. Particular attention is given to the convergence-theoretic results that justify RL algorithms as legitimate stochastic approximation procedures: Robbins and Monro's foundational stochastic approximation method, Jaakkola, Jordan, and Singh's convergence proof for stochastic iterative dynamic programming, Tsitsiklis and Van Roy's analysis of temporal-difference learning with linear function approximation, and the policy-gradient theorem of Sutton, McAllester, Singh, and Mansour. The paper further examines deterministic policy-gradient methods, deep reinforcement learning's departure from classical convergence guarantees, and the connection between the discrete-time Bellman equation and its continuous-time HJB counterpart via viscosity solution theory. Comparative tables map core mathematical structures onto their representation assumption and guarantee type, contrast convergence rigor across tabular, linear, and nonlinear function-approximation regimes, and set the discrete-time RL and continuous-time control literatures against their shared and divergent mathematical machinery. The paper concludes that convergence guarantees degrade in a predictable, representation-dependent order as function approximation grows more expressive, and identifies extending stochastic approximation theory to nonlinear function approximation as the central future research prospect.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-26

How to Cite

Shirish Prabhakarrao Kulkarni, C. Ashwini, M. Buvanasankari, K. Ramesh, Rajat Verma, & Abhinav Jha. (2026). Mathematical Foundations of Reinforcement Learning and Stochastic Control Systems. International Journal of Computer Information Systems and Industrial Management Applications, 18(20s), 170–178. https://doi.org/10.70917/ijcisim-2026-5162

Issue

Section

Original Articles