SyncAI.news, a Varaisys broadcasting
Computationally efficient safe exploration in reinforcement learning
SM

Shreeram Murali, Shankar A. Deka, Dominik Baumann

· 1 min read

ResearcharXiv cs.LG

Computationally efficient safe exploration in reinforcement learning

arXiv:2609.22919v1 Announce Type: new Abstract: Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.

Original source

This story was published by arXiv cs.LG and written by Shreeram Murali, Shankar A. Deka, Dominik Baumann. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News