SyncAI.news, a Varaisys broadcasting
How Much Human Label Variation Does Formal Semantic Structure Explain? Group-Level Effects and Item-Level Ceilings in NLI
HC

Haram Choi (University of Bremen)

· 1 min read

ResearcharXiv cs.CL

How Much Human Label Variation Does Formal Semantic Structure Explain? Group-Level Effects and Item-Level Ceilings in NLI

arXiv:2607.15870v3 Announce Type: replace Abstract: Human label variation in natural language inference (NLI) is increasingly treated as a signal to be measured rather than as noise to be removed. We ask whether one candidate source of that signal, formal semantic structure such as negation, quantification, and monotonicity, changes how much annotators disagree and what they disagree about. We tag the SNLI and MNLI items of ChaosNLI, each labeled by 100 annotators, with a rule-based monotonicity tagger, check the tagger by hand on a sample of the same items, and answer three questions. At the group level, hypotheses that are not purely upward monotone attract somewhat more disagreement, but an error sensitivity analysis shows that this difference is sensitive to tagger error. At the item level, formal structure explains only a few percent of the variation and cannot pick out the items that attract high disagreement. In composition, the kinds of disagreement recorded by VariErr and LiTEx do not differ detectably across the formal boundary. Formal structure therefore belongs in the inventory of disagreement sources, with a small and bounded weight. ChaosNLI was built from low-agreement items, and every claim holds within that scope. Analysis decisions were written in a version-controlled research log before the corresponding results were computed, and negative results are reported in full.

Original source

This story was published by arXiv cs.CL and written by Haram Choi (University of Bremen). SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News