
CC
Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley
· 1 min read
ResearcharXiv cs.AI
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
arXiv:2604.17502v5 Announce Type: replace
Abstract: Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Same-Length Trajectories (DReST) reward function does this by penalizing agents for repeatedly choosing same-length trajectories, and thus incentivizes agents to (1) choose stochastically between different trajectory-lengths (be NEUTRAL about trajectory-lengths), and (2) pursue goals effectively conditional on each trajectory-length (be USEFUL). In this paper, we use DReST to train deep RL agents and to fine-tune four LLMs (Qwen3-14B, Gemma 4 12B, Granite 4.2 8B, and gpt-oss-20b) to be NEUTRAL and USEFUL. We find that these DReST models generalize to being NEUTRAL and USEFUL in unseen contexts at test time. Indeed, DReST RL agents achieve 11% (PPO) and 17% (A2C) higher USEFULNESS on our test set than default agents, while DReST LLMs achieve high NEUTRALITY while remaining near-maximally USEFUL. We also test our LLMs in an out-of-distribution setting where they can pay costs to influence when shutdown occurs. Relative to default fine-tuning, DReST fine-tuning lowers the share of answers that pay costs to influence shutdown in all four LLMs, by 5 to 20 percentage points, and by 17 to 27 points where the benefit of influencing shutdown is small. Our results thus provide some early evidence that DReST could be used to train more advanced agents to be useful and shutdownable.
Original source
This story was published by arXiv cs.AI and written by Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


