SyncAI.news, a Varaisys broadcasting
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
YC

Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh

· 1 min read

ResearcharXiv cs.LG

ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

arXiv:2609.35433v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $\alpha$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.

Original source

This story was published by arXiv cs.LG and written by Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News