
Maximilian Schreiner
· 2 min read
Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think
The context is a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way. For that, the public needs more insight, Anthropic says. The metrics are meant to show how models get built, and they complement capability tests that measure what models can do.
What "leads" means here, and what it doesn't
At the center is an index that sorts all development work at Anthropic onto a scale from Epoch AI, running from AL0 (no AI) to AL5 (fully autonomous). As of August 2026, 26 percent of the work sits at AL4, up from under one percent in February. More than 90 percent reaches at least AL3. Claude hits AL5 nowhere.
Epoch AI calls AL4 "AI leads," and Anthropic put that in its headline. But anyone who ties "lead" to fully autonomy has it wrong. That's what AL5 is for.
An example from Anthropic shows where AL4 sits: An engineer hands Claude a bug report, and Claude analyzes, fixes, and tests it without asking questions, but it isn't allowed to ship. A human reads the report and decides. The task and the direction still come from the human. The difference from AL3 ("collaborates") is mainly that Claude no longer stalls when it runs into a snag.
Claude did the scoring itself. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels. Anthropic admits this "judge" could make the same mistakes as the system it's checking.
Then there's what the 26 percent even measures. Anthropic listed how its employees spent their work time in July and scored each activity, with tasks that eat up a lot of time counting for more. So Claude leads a quarter of the work as measured by the human hours it takes. That says nothing about how many decisions Claude makes or whether it has a say in the research direction.
Young oversight for 30,000 agents
The company warns that the line to capability research is blurry, and every vendor is tempted to draw it generously. The burden of proof, it says, should sit with the developer.
Original source
This story was published by The Decoder and written by Maximilian Schreiner. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on the-decoder.com


