SyncAI.news, a Varaisys broadcasting
Diversifying RLVR Rollouts via First-Token Exploration
SK

Soeun Kim, Albert No

· 1 min read

ResearcharXiv cs.AI

Diversifying RLVR Rollouts via First-Token Exploration

arXiv:2605.28295v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) trains reasoning models without labeled trajectories, using groups of verifier-scored rollouts to explore alternative reasoning paths. Limited rollout diversity is a central bottleneck, typically addressed through adjustments to temperature, prefixes, or rollout selection. We identify the first token of the response as a structurally distinct target for diversification, largely overlooked in prior work. We find that the first-token distribution is sharply concentrated and only weakly related to downstream correctness, as lower-probability candidates can yield similarly accurate responses. Diversifying the first token can therefore broaden the reasoning paths explored within each rollout group with little loss in response quality. Motivated by this observation, we introduce REFT (Rollout Exploration with First-Token Diversification), a lightweight modification to RLVR. REFT samples first tokens uniformly from the policy's top-$N$ candidates and allocates rollouts evenly across the sampled tokens, leaving the rest of the pipeline unchanged. We evaluate REFT on eight models spanning multiple architectures and sizes (0.5B-14B), with mathematical reasoning and code-generation tasks under GRPO and DAPO. Across these settings, REFT consistently improves Pass@1, Pass@8, and Pass@64. It also outperforms competing diversification methods at every evaluated budget, incurring the lowest rollout cost.

Original source

This story was published by arXiv cs.AI and written by Soeun Kim, Albert No. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News