SyncAI.news, a Varaisys broadcasting
How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
YZ

Yukun Zhang, Kemu Xu, Yishen Chen

· 1 min read

ResearcharXiv cs.AI

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

arXiv:2609.20474v1 Announce Type: new Abstract: Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $\tau^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

Original source

This story was published by arXiv cs.AI and written by Yukun Zhang, Kemu Xu, Yishen Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News