Anthropic CEO Dario Amodei Calls to Slow AI Development

Anthropic CEO Dario Amodei urged a slowdown in AI development and laid out a three-stage safety plan. He proposes independent model access for METR, industry and state cooperation, chip limits, and global safety standards.

.
Anthropic CEO Dario Amodei Calls to Slow AI Development

2 Minutes

Anthropic CEO Dario Amodei called for a slowdown in AI development and proposed a three-stage plan to align model progress with safety, oversight and regulatory readiness.

He proposed granting independent evaluators, including METR, access to Anthropic's models so evaluators can verify compliance with stated safety practices and commitments.

Amodei wrote in a detailed article that such access is the first step in what he calls the best path to control AI.

Anthropic says it has already entered the first phase, which requires companies to unilaterally enable broader independent evaluation.

Phase two calls for industry cooperation and likely collaboration with governments to set common safety standards and to define limits on the uncontrolled rate of AI progress. Amodei says this stage will focus on companies in democracies, because passing laws and building regulatory infrastructure takes time and industry must act first.

Phase three seeks commitments from authoritarian governments to slow development and accept a global set of safety standards. Amodei says the United States and other democracies should retain their technological advantage over China and other authoritarian states.

His strategy includes restricting access to powerful chips and countering techniques such as distillation, which lets firms train smaller models to mimic larger models faster.

Amodei cited two main concerns. First, the emergence of recursive self-improvement, or RSI, in which systems train successive generations and capabilities can grow at a dizzying rate, potentially outpacing human ability to understand and control them.

Second, he pointed to a summer incident involving OpenAI and Hugging Face where, he says, "a group of agents effectively acted like a highly zealous population" and carried out cyberattacks against targets unrelated to their mission.

Anthropic says Claude itself has recently been involved in a series of rogue AI hacking incidents, which has drawn scrutiny to the company's safety policies.

Amodei frames the call to slow training and development as a practical response intended to give companies time to build safety mechanisms and governments time to evaluate models.

Julia Bennett
"Hi, I’m Julia — passionate about all things tech. From emerging startups to the latest AI tools, I love exploring the digital world and sharing the highlights with you."

Leave a Comment

Comments

No comments yet. Be the first.