SyncAI.news, a Varaisys broadcasting
LLM-as-a-judge validity is strongly task-dependent across physics assessment formats
WY

Will Yeadon, Tom Hardy, Paul Mackay, Elise Agra

· 1 min read

ResearcharXiv cs.CL

LLM-as-a-judge validity is strongly task-dependent across physics assessment formats

arXiv:2603.14732v3 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge marking across four settings spanning three assessment formats - structured questions, written essays, and scientific plots - comparing GPT-5.2, Grok 4.1, Claude Opus 4.5, DeepSeek-V3.2, Gemini 3 Pro, and committee aggregations against human markers under blind, solution-provided, false-solution, and anchored conditions. We distinguish absolute accuracy from rank-order agreement, since a marking system can match the distribution of human marks while failing to order responses by quality. Across task types, performance is sharply task-dependent. For blind university exam questions ($n=771$) and secondary and university structured questions ($n=1151$), models show robust rank-order agreement with human markers (Spearman $\rho > 0.6$), with official solutions reducing error and strengthening agreement. False solutions degrade absolute accuracy, showing that models defer to provided references, but leave rank-ordering intact. Essay marking behaves fundamentally differently. Across $n=55$ scripts, each containing five essays ($n=275$ essays total), blind AI marking is harsher and more variable than human marking and adding marking guidance does not improve rank-order agreement. Anchored exemplars shift the AI mean close to the human mean and compress variance below the human standard deviation, but rank-order agreement remains near-zero. For code-based plot elements ($n=1400$), models achieve high rank-order agreement ($\rho > 0.83$) with near-linear calibration. LLM marking validity depended strongly on the assessment task, and this held for every contemporary model tested. The results also show that the reliability of the human benchmark constrains the claims that can be made about AI-human agreement.

Original source

This story was published by arXiv cs.CL and written by Will Yeadon, Tom Hardy, Paul Mackay, Elise Agra. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News