
IM
Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
· 1 min read
ResearcharXiv cs.AI
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
arXiv:2602.14868v3 Announce Type: replace-cross
Abstract: Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly sample-inefficient, as models must navigate vast search spaces with minimal feedback. While classic curriculum learning aims to mitigate this by ordering data based on complexity, prior works have primarily targeted small datasets and do not directly transfer to the large-scale settings typical of modern language model training. Furthermore, the right ordering for a specific model is often unclear. To address this, we propose Goldilocks, an adaptive data-selection strategy that uses a Selector network to predict the standard deviation of rewards across the model's rollouts for each candidate question. The Selector prioritizes questions with high predicted reward variability, corresponding to questions that are neither too easy nor too hard for the model's current capabilities (Goldilocks principle), while training the model with GRPO. By leveraging the model's performance on seen samples, the Selector continuously adapts to the model's evolving abilities. Across the OpenMathReasoning and Polaris datasets, Goldilocks consistently improves over standard GRPO, requiring up to 78% fewer optimization steps to reach the corresponding GRPO performance.
Original source
This story was published by arXiv cs.AI and written by Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


