SyncAI.news, a Varaisys broadcasting
Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
ML

Minghao Li, Rui Tan, Ruihang Wang

· 1 min read

ResearcharXiv cs.AI

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

arXiv:2609.32394v1 Announce Type: new Abstract: Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.

Original source

This story was published by arXiv cs.AI and written by Minghao Li, Rui Tan, Ruihang Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News