SyncAI.news, a Varaisys broadcasting
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
HL

Hongming Li, Zhao Yang, Xiaoxuan Liang, Shujian Yu, Jose C. Principe

· 1 min read

ResearcharXiv cs.AI

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

arXiv:2412.03800v2 Announce Type: replace-cross Abstract: Reinforcement learning agents depend on reward signals whose density is rarely under the designer's control, and when such signals are absent, an agent must generate its own drive to explore. State entropy maximization offers a principled objective for this, but existing methods break down at scale in two ways: the intrinsic reward vanishes once a state has been visited, discouraging revisits to the very gateways that lead onward, and estimating entropy over millions of accumulated observations becomes computationally prohibitive. We address both with Episodic and Lifelong Exploration via Maximum Entropy (ELEMENT), a multiscale intrinsically motivated framework for reward-free exploration that transfers to downstream tasks. ELEMENT couples lifelong entropy maximization with a complementary episodic term acting on a faster timescale. For the episodic term, we derive average episodic state entropy, an intrinsic reward that is the exact minimizer of a tractable upper bound on the reward-decomposition objective; for the lifelong term, we propose a $k$NN graph-based estimator that keeps entropy tractable without forgetting. ELEMENT consistently outperforms state-of-the-art intrinsic reward baselines on state coverage and unsupervised pre-training. Videos, code, and supplementary material: https://sites.google.com/view/element-rl.

Original source

This story was published by arXiv cs.AI and written by Hongming Li, Zhao Yang, Xiaoxuan Liang, Shujian Yu, Jose C. Principe. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News