SyncAI.news, a Varaisys broadcasting
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
HL

Haowei Liu, Hsin-Tai Wu, Yi Fang

· 1 min read

ResearcharXiv cs.CL

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call. Ensembling does not help for free. Pairing the weak judge with a stronger one degrades agreement, whereas three strong judges under unanimity routing reach kappa = 0.79 at 89.7% auto-coverage. Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under our annotation protocol. Code and pre-registration are at https://github.com/JamesL404/synca-audit.

Original source

This story was published by arXiv cs.CL and written by Haowei Liu, Hsin-Tai Wu, Yi Fang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News