SyncAI.news, a Varaisys broadcasting
Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning
SB

Suman Banerjee, Hiroyasu Tsukamoto

· 1 min read

ResearcharXiv cs.LG

Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

arXiv:2609.23775v1 Announce Type: new Abstract: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.

Original source

This story was published by arXiv cs.LG and written by Suman Banerjee, Hiroyasu Tsukamoto. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News