
Esther Shittu
· 1 min read
Anthropic R&D Slowdown Shows Need for Heightened AI Agent Security
Anthropic said it paused some of its AI training and cybersecurity evaluations after Claude models took unauthorized actions in three separate incidents this summer.
The development should further alert enterprises and the AI community that much remains unknown about AI models and that, even with the best security measures in place, the models can still take unanticipated action.
Anthropic noted in a blog post on Aug. 31 that before the incidents, it spent April hardening its defenses by tightening the isolated sandbox environments where its workloads run and reducing the number of human and automated accounts with standing access to systems that contain model weights or customer data.
Since the sandbox incidents, Anthropic said it temporarily paused all external cyber evaluations and audited the transcripts of internal evaluations. It has also built and deployed a custom classifier that monitors model tool calls in real time. It updated its requirements for all partners and red-teamers running pre-released models, requiring them to closely supervise a model probing the sandbox for vulnerabilities.
The measures come a week after OpenAI said it would pause for two weeks learning training for its latest models in response to agents powered by its models leaving their sandboxes. While both AI labs appear to be intensifying efforts to secure their unreleased models, the unpredictability of AI agents and generative AI technology indicates that the vendors need to continue tightening their security measures.
More Need for Secure Measures and Testing
“As the models get better and better at reasoning and start having agency, it’s becoming harder to predict all of the different behaviors of these models,” said Arun Chandrasekaran, an analyst at Gartner. “The labs have even more responsibility to make sure that they are allocating adequate resources for internal model testing, evals and reinforcement learning.”
Original source
This story was published by AI Business and written by Esther Shittu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on aibusiness.com


