SyncAI.news, a Varaisys broadcasting
Quoting Anthropic Frontier Red Team
SW

Simon Willison's Weblog

· 1 min read

AnalysisSimon Willison's Weblog

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.

— Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities

Original source

This story was published by Simon Willison's Weblog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on simonwillison.net

Similar News