SyncAI.news, a Varaisys broadcasting
PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation
HY

Hongye Yang, Zhihao Xie, Shengjun Xiong

· 1 min read

ResearcharXiv cs.AI

PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

arXiv:2609.29578v1 Announce Type: new Abstract: Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its trajectories match component-wise in both current-state predicate satisfaction and standardized agent attribution; score inflation, defined as f(A) - f(H), is measured only afterward. In 18 sealed held-out tasks in PB-CSTE, the frozen historical-target run produced matched adversaries for 15 tasks. Historical credit yielded mean inflation of .252, conditional attack success of 10/15, end-to-end yield of 10/18, and detected none of 14 strict rollbacks. Semantic LLM judges were more resistant but remained vulnerable, especially under evaluator-targeted attacks, while PB-CSTE current-state controls, defined as exact functions of the certified components, yielded zero inflation by construction. PartHackBench thus provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.

Original source

This story was published by arXiv cs.AI and written by Hongye Yang, Zhihao Xie, Shengjun Xiong. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News