Mathematical Foundations of Reinforcement Learning and Stochastic Control Systems
DOI:
https://doi.org/10.70917/ijcisim-2026-5162Keywords:
Reinforcement Learning, Stochastic Control, Markov Decision Processes, Bellman Equation, Stochastic Approximation, Temporal-Difference Learning, Policy Gradient Methods, Hamilton–Jacobi–Bellman EquationAbstract
Reinforcement learning (RL) and stochastic optimal control both address a single underlying mathematical problem: how an agent should select actions over time, under uncertainty, to optimize a cumulative reward or cost criterion. This paper reviews the shared mathematical scaffolding uniting these two fields, tracing the progression from Bellman's dynamic programming and the Markov decision process (MDP) formalism through stochastic approximation theory, temporal-difference learning, policy-gradient methods, actor-critic architectures, and the Hamilton–Jacobi–Bellman (HJB) equation governing continuous-time stochastic control. Particular attention is given to the convergence-theoretic results that justify RL algorithms as legitimate stochastic approximation procedures: Robbins and Monro's foundational stochastic approximation method, Jaakkola, Jordan, and Singh's convergence proof for stochastic iterative dynamic programming, Tsitsiklis and Van Roy's analysis of temporal-difference learning with linear function approximation, and the policy-gradient theorem of Sutton, McAllester, Singh, and Mansour. The paper further examines deterministic policy-gradient methods, deep reinforcement learning's departure from classical convergence guarantees, and the connection between the discrete-time Bellman equation and its continuous-time HJB counterpart via viscosity solution theory. Comparative tables map core mathematical structures onto their representation assumption and guarantee type, contrast convergence rigor across tabular, linear, and nonlinear function-approximation regimes, and set the discrete-time RL and continuous-time control literatures against their shared and divergent mathematical machinery. The paper concludes that convergence guarantees degrade in a predictable, representation-dependent order as function approximation grows more expressive, and identifies extending stochastic approximation theory to nonlinear function approximation as the central future research prospect.