SyncAI.news, a Varaisys broadcasting
The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations
TP

Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski, Lukas Lange

· 1 min read

ResearcharXiv cs.CL

The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

arXiv:2609.36209v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investigate how LLMs represent and generate multi-valued relations. We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological). Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases: (1) retrieval of candidate entities, (2) internal sorting, and (3) selection of the next element. As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations.

Original source

This story was published by arXiv cs.CL and written by Timo Pierre Schrader, Annemarie Friedrich, Simon Razniewski, Lukas Lange. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News