SyncAI.news, a Varaisys broadcasting
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
YL

Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang

· 1 min read

ResearcharXiv cs.AI

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

arXiv:2609.33875v3 Announce Type: replace-cross Abstract: Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.

Original source

This story was published by arXiv cs.AI and written by Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News