
RR
Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
· 1 min read
ResearcharXiv cs.CL
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
arXiv:2604.23478v3 Announce Type: replace
Abstract: Large language models are widely used to judge the output of other language models, yet whether a judge returns the same verdict when the same request is worded differently remains largely unexamined. We study that question across four evaluation tasks and twenty-five judges from six providers. To support the analysis we release JudgeSense, a benchmark of 880 items from human-labelled corpora, each issued under two instructions that differ in wording and not in what they ask, with the complete decision logs. Every score is reported against the judge's own agreement with itself on the identical prompt, so decoding noise is not charged to wording, and the release lets a reader ask the same of any judge not in our roster. Rewording costs agreement on all four tasks, and on two it clears the threshold we declare for a practically meaningful effect; the ordinal task is both the least stable and the one fewest judges are accurate on, and within a single family parameter count does not predict stability. A judge measured inside an agent harness yields a smaller estimate than the same judge reached through a direct API call, because its agreement with itself collapses faster than its agreement across wordings.
Original source
This story was published by arXiv cs.CL and written by Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


