Number-Theoretic Bounds in Optimization: A Research Program Connecting Random-Matrix Universality in Loss Landscapes to Analytic Number Theory
DOI:
https://doi.org/10.70917/ijcisim-2026-4591Abstract
The geometry of non-convex loss landscapes and the spectral statistics of the Riemann zeta function's nontrivial zeros have both been described, independently, using random matrix theory (RMT). This paper examines that connection with the technical care it requires, and — beyond analysis — reports original numerical experiments designed to stress-test the claims involved. We first derive, rather than merely cite, the classical RMT results that are actually proven: the Wigner semicircle law, the Wigner surmise for level spacing, the Marchenko–Pastur law, and the Bai–Yin/Tracy–Widom results governing extreme eigenvalues of Wishart matrices. We show these proven results already furnish a rigorous, non-conjectural instance of a 'deterministic ceiling governing chaotic structure,' the pattern that motivates interest in Robin's inequality (equivalent to the Riemann Hypothesis). We extend the toy model with the Baik–Ben Arous–Péché (BBP) phase transition, which describes exactly when an outlier eigenvalue detaches from a random matrix bulk — the same bulk-plus-outlier structure reported empirically for trained neural network Hessians — and we numerically confirm the BBP prediction to within simulation error. We then run a second, independent experiment: training small neural networks directly, computing their exact Hessians, and directly testing whether the empirically claimed Gaussian Orthogonal Ensemble (GOE) level-spacing statistics reproduce at small scale. They do not: our networks instead reproduce the well-documented Hessian rank-degeneracy phenomenon (a large near-zero eigenvalue pile), and once that degenerate cluster is excluded, the remaining 'informative' eigenvalues are still too few for either GOE or GUE repulsion to be statistically distinguishable from an uncorrelated (Poisson) baseline. We report this as a genuine, informative negative result that calibrates the scale at which RMT universality claims for neural network Hessians can be expected to hold, and we use it to sharpen — rather than abandon — the research program's open conjectures.