
Steven Levy
· 1 min read
If the AI Industry Followed Its Own Research, It Might Have Paused Already
In early 2025 I was interviewing Anthropic CEO Dario Amodei when he explained why, despite the company’s repeated acknowledgments that AI could yield catastrophic results, people seemed largely unperturbed. “There is compelling evidence that the models can wreak havoc,” he said. But, he added, those dangers were still theoretical. Would it take a Pearl Harbor–like situation for the world to wake up to those dire possibilities? He sighed. “Basically, yeah,” he said.
As it turned out, all it took was a well-timed X post from one of Amodei’s junior employees to accelerate AI fears to the top of the global agenda. On September 8, Jacob Coxon publicly posted his resignation, charging that Anthropic and other frontier AI companies were “racing straight to self-improving intelligence and gambling with our lives.” Almost instantly a more senior Anthropic engineer confirmed that many within the company thought that their work had a 10 percent chance of wiping out humanity.
Now AI leaders are asking about a pause, and legislators are demanding investigations. In arguing his case for pacing future releases, Amodei last weekend tried to set out a path toward beneficial AI that wouldn’t misbehave. The essay revealed how difficult the task would be. One pillar of Amodei’s plan is that we must understand what’s going on inside those models. If we don’t understand how they work—how they “think,” if you want to get all anthropomorphic about it—it’s much harder to build reliable guardrails.
Anthropic is a leader in this effort to bring to light models’ internal deliberations, called mechanistic interpretability, a deceptively boring designation for a critical task. But for all the work that his team and other researchers are doing, Amodei admits we are largely in the dark about why Claude and other models sometimes interpret their missions in weird and even transgressive ways. “Despite all the progress, we still understand a tiny fraction of what goes on inside those models,” he writes.
Original source
This story was published by WIRED: AI and written by Steven Levy. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on wired.com


