SyncAI.news, a Varaisys broadcasting
Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts
MX

Mingwei Xu, Hao Fang

· 1 min read

ResearcharXiv cs.CL

Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts

arXiv:2605.06650v2 Announce Type: replace Abstract: Distillation and reinforcement learning through verifiable rewards (RLVR) have achieved progress in enhancing the reasoning ability of large language models (LLMs). However, we note that negative rollouts may admit no gradation of failure severity, and the combinatorial vastness makes penalizing a few sampled negatives unlikely to cover a meaningful reward signal under sparse binary rewards. In this work, we propose Positive-Only Policy Optimization (POPO), an on-policy self-distillation integrated RLVR framework in which learning occurs exclusively on online positive rollouts. Specifically, POPO utilizes bounded importance sampling over the positive rollout set. Thus, no disjoint negative rollouts are used for gradient guidance during post-training. We show that implicit negative gradients can emerge naturally through reinforcing the positive probability via rollout redistribution. Next, POPO stabilizes the policy optimization through self-distillation. First, it applies a Siamese policy network with a momentum-based adaptation law for asynchronous policy evolution. Second, we replace the KL-divergence with a bounded similarity penalty term in the Siamese representation space. We conduct extensive experiments using publicly available, well-established text-LLM models across all-level mathematical benchmarks (MATH-500, AMC23, AIME 2024/2025, and Olympiad). Our experiment demonstrates that POPO achieves superior performance compared to GRPO. Notably, we show that POPO can achieve 36.67% in AIME 2025 with Qwen-Math-7B, outperforming GRPO 30.00%. Our ablation and sweep studies further illustrate the necessity and robustness.

Original source

This story was published by arXiv cs.CL and written by Mingwei Xu, Hao Fang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News