
PM
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
· 1 min read
ResearcharXiv cs.CL
Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
arXiv:2609.32449v1 Announce Type: new
Abstract: Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.
Original source
This story was published by arXiv cs.CL and written by Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


