
GL
Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen
· 1 min read
ResearcharXiv cs.CL
Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
arXiv:2609.33702v1 Announce Type: new
Abstract: Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Original source
This story was published by arXiv cs.CL and written by Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


