SyncAI.news, a Varaisys broadcasting
Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
PM

Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz

· 1 min read

ResearcharXiv cs.CL

Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports

arXiv:2609.32449v1 Announce Type: new Abstract: Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.

Original source

This story was published by arXiv cs.CL and written by Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News