
JL
Junxuan Li, Arko Mukherjee, Soumyabrata Pal
· 1 min read
ResearcharXiv cs.AI
LLM Judge Validation Under Sparse Overlap: From Inference to Design
arXiv:2609.31857v1 Announce Type: new
Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Original source
This story was published by arXiv cs.AI and written by Junxuan Li, Arko Mukherjee, Soumyabrata Pal. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


