
NL
Nan Lu, Ethan Lee, James M. Robins, David Simchi-Levi, Junwei Lu
· 1 min read
ResearcharXiv cs.LG
Optimal Value Inference for Reinforcement Learning
arXiv:2609.09981v2 Announce Type: replace-cross
Abstract: We study offline inference for the optimal value in reinforcement learning under finite state and action spaces. Two new nuisances are derived as fixed points of a self-induced Bellman equation, in which we approximate the maximum Bellman operator by its softmax correspondence. We propose a debiased estimator through the Neyman orthogonality and establish its asymptotic normality under diverging horizons even when the behavior policy changes with time, as long as the nuisances have the statistical rates that can be achieved by many machine learning methods. We provide a concrete estimating procedure for these nuisances and show they can lead to valid inference. Synthetic experiments validate the numerical performance of our inference method, and we implement it in real-life decision-making problems, including bike repositioning and AI agentic tool use.
Original source
This story was published by arXiv cs.LG and written by Nan Lu, Ethan Lee, James M. Robins, David Simchi-Levi, Junwei Lu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


