
JD
Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song
· 1 min read
ResearcharXiv cs.AI
Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
arXiv:2609.39701v1 Announce Type: new
Abstract: Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
Original source
This story was published by arXiv cs.AI and written by Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


