
HW
Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun
· 1 min read
ResearcharXiv cs.LG
Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
arXiv:2610.00400v1 Announce Type: new
Abstract: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost.
These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.
Original source
This story was published by arXiv cs.LG and written by Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


