SyncAI.news, a Varaisys broadcasting
LLM Judge Validation Under Sparse Overlap: From Inference to Design
JL

Junxuan Li, Arko Mukherjee, Soumyabrata Pal

· 1 min read

ResearcharXiv cs.AI

LLM Judge Validation Under Sparse Overlap: From Inference to Design

arXiv:2609.31857v1 Announce Type: new Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

Original source

This story was published by arXiv cs.AI and written by Junxuan Li, Arko Mukherjee, Soumyabrata Pal. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News