SyncAI.news, a Varaisys broadcasting
Watch the Model Think: On-Policy Extraction of Activation Steering Vectors
XS

Xuanbo Su, Yingfang Zhang, Hao Luo, Huajun Bai, Guangyuan Dong, Ling Huang

· 1 min read

ResearcharXiv cs.CL

Watch the Model Think: On-Policy Extraction of Activation Steering Vectors

arXiv:2602.14143v2 Announce Type: replace-cross Abstract: When a model solves a problem on one attempt and fails it on the next, what separates the two is rarely the final answer token; it is the trajectory that reached it. Contrastive activation steering leaves that signal unused: CAA, SADI, RepE and ITI build their direction from experimenter-supplied text, recorded while the model reads rather than reasons. That choice also caps what the vector can express, since polarity must be written into the text, and a task judged only by outcome offers nothing to write it with. ROAST makes the trajectory itself the contrast: sample rollouts, let an outcome verifier split them into successes and failures, and contrast the reasoning that worked against the reasoning that did not. A matched teacher-forced control---rollouts, labels, answer text and pair counts held fixed, the trajectory alone stripped---points to the trajectory as what matters: on GSM8K at 0.6B the pairs alone buy +0.12 points while restoring the trajectories buys +6.05, the larger and only seed-robust step. Replacing the trajectory with an equal-length neutral prefix or another question's reasoning falls below no intervention. The two corpora are also far apart geometrically, a median 70+ degrees apart at both Qwen3 scales probed, beyond what a split-half null explains. Reading from rollouts calls for two corrections---keeping the full difference vector rather than Top-10% masking, and giving each question one vote rather than one per pair---and only grouped aggregation beats the unsteered baseline under 20% verifier noise. On parser-free benchmarks (GSM8K, MATH500, IFEval), ROAST is best in all six cells over two models, by up to +9.7, at +6.4% wall-clock and no added context; it also leads on six parser-scored benchmarks across three models. Across nine models (0.6B--122B, four families), ROAST improves on the unsteered model at every scale. Code: https://github.com/TomySu404/ORBIT

Original source

This story was published by arXiv cs.CL and written by Xuanbo Su, Yingfang Zhang, Hao Luo, Huajun Bai, Guangyuan Dong, Ling Huang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News