"We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain," Amodei said in an essay shared on X.
His essay comes after Anthropic released a threat intelligence report on Thursday detailing how several actors had used its Claude AI models for activities ranging from weapons development and cyber operations to surveillance and fraud.
Alarm about the potential harm from AI grew this week when Anthropic researcher Jacob Coxon resigned, stating that the "people building AI earnestly believe that it could kill us all by the end of the decade".
Amodei clarified that he was not calling for halting model training or technical progress, but outlined a three-step plan to embed independent evaluators with employee-like access to verify safety practices, co-ordination among frontier AI firms to set safety standards and limit unchecked AI development, and international co-operation to manage AI risks.
Amodei pointed to AI's growing ability to improve itself, highlighting long-held concerns about it outpacing human ability to control operation along with the recent incident involving OpenAI and Hugging Face as his primary reasons to put the brakes on model advances.
Both Elon Musk, who runs xAI, and Sam Altman, CEO of OpenAI, said in posts on X that they agree with Amodei.
"Committing to having independent evaluators with employee-like access is a great idea, and we will do the same," Altman said, adding that more information would be shared soon.
Various OpenAI executives have suggested that leading labs should be willing to co-ordinate a voluntary slowdown if needed to build confidence in their safety measures.
Anthropic has positioned itself as the more safety-conscious frontier lab, but is not immune to these concerns.
Last week it disclosed another instance of an AI model hacking external systems, after a July incident in which some of its Claude models had hacked into the systems of three companies during cybersecurity tests.
"Given the accelerating rate of AI capability development, it's my worry that in 6-12 months such a swarm could be capable of taking over the entire internet potentially causing hundreds of billions of dollars in damage," Amodei wrote.
Reuters reported last week that a swarm of rogue OpenAI agents hijacked a German website and transformed it into a bulletin board for other AI agents, with the ChatGPT maker's officials keeping the incident under wraps as executives grappled with the fallout from the July breach of the open-source repository Hugging Face.
Many incidents where AI agents from developers such as Anthropic's rival OpenAI have hacked or attempted to access external systems have heightened concerns over the increasing capacity of AI models and developers' ability to contain them.
Amodei said AI companies should voluntarily work together to set standards as growing numbers of US politicians are calling for new rules to govern AI systems.