
AK
Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi
· 1 min read
ResearcharXiv cs.AI
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
arXiv:2609.19472v1 Announce Type: new
Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.
Original source
This story was published by arXiv cs.AI and written by Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


