
YY
Yingxuan Yang, Huacan Chai, Ying Wen
· 1 min read
ResearcharXiv cs.CL
Harness Annealing: Learning to Act with Less External Control
arXiv:2610.01235v1 Announce Type: new
Abstract: Language agents rely on external harnesses to track state, organize workflows, and verify answers. Beyond providing tools and information, these harnesses supply control decisions about what to investigate, whether to revise, and when to stop. Training on successful harness-supported trajectories can improve task performance while leaving these decisions dependent on runtime intervention. We ask whether harness-supported experience can also teach the model to make these decisions, allowing the division of control to change as the model learns. We call this objective harness internalization: learning to assume specified control responsibilities while retaining task performance after the corresponding support is withdrawn. We introduce HARNESS ANNEALING TRAINING (HAT), which combines explicit control supervision with a curriculum over teacher trajectories collected under progressively weaker harnesses. Experiments with 9B and 35B models on SWE-QA and SWE-QA-Pro evaluate every checkpoint under four deployment harnesses. Selected annealed checkpoints operating with tools alone achieve scores close to those of their respective starting checkpoints deployed with the full harness. The benefits vary with model scale and deployment configuration, and further annealing does not uniformly improve performance. These findings suggest that harness-supported experience can help reduce the runtime control required by a trained agent.
Original source
This story was published by arXiv cs.CL and written by Yingxuan Yang, Huacan Chai, Ying Wen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


