SyncAI.news, a Varaisys broadcasting
Decodable In-Context State and Model Output Across Training
MV

Manas Venkata Sai Ravulapalli, Samrath Singh Chadha

· 1 min read

ResearcharXiv cs.LG

Decodable In-Context State and Model Output Across Training

arXiv:2609.31401v1 Announce Type: new Abstract: Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate logits finds no detected final-state advantage on late-checkpoint model errors. An information-theoretic counterexample explains why decodability on errors alone cannot establish discarded output information. The connection to downstream omissions remains open.

Original source

This story was published by arXiv cs.LG and written by Manas Venkata Sai Ravulapalli, Samrath Singh Chadha. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News