
MU
Manuel Uth
· 1 min read
BusinessThe Decoder
AI agents overstate their results and remain far from autonomous research, study finds
Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods researchers already knew. The models' biggest weakness is still their inability to critically question their own results.
Original source
This story was published by The Decoder and written by Manuel Uth. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on the-decoder.com


