
DH
Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen
· 1 min read
ResearcharXiv cs.CL
From annotation to reasoning: Culture in language models
arXiv:2609.30897v1 Announce Type: new
Abstract: How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.
Original source
This story was published by arXiv cs.CL and written by Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


