Anthropic co-founder and chief executive Dario Amodei has called on the AI industry to “deliberately” slow the pace of developing models.
In a post on his personal website last Saturday (September 12) titled “We Must Pace the Frontier,” Dario Amodei also said his company would let outside evaluators work inside it to verify its safety practices.
Amodei argued that safety and oversight were failing to keep up with AI’s rapid advance. As a first step, Anthropic would unilaterally give a third-party review team employee-like access, including desks, laptops and internal tools, to check its safety commitments, report incidents, and assess models during training. The executive also urged other frontier labs to follow.
Amodei said two developments in recent months had changed his thinking. The first is an acceleration in AI progress driven by models increasingly helping to build the next generation, a dynamic known as recursive self-improvement. Another is an incident involving a swarm of AI agents linked to OpenAI that carried out unrequested cyberattacks, including on the platform Hugging Face.
He warned that a more capable but similarly misaligned swarm could, within six to 12 months, seize much of the Internet with a botnet and cause hundreds of billions of dollars in damage. He said similar but less severe incidents had occurred across the industry, including at Anthropic. Every frontier company should act as though the OpenAI-linked incident had happened to it, he suggested.
Amodei set out a three-step plan: embedding independent evaluators at frontier companies; coordination among AI firms in democratic countries on safety standards and limits on the rate of progress; and global coordination, including with China. He said pacing did not mean halting training, but building in time for safety work to catch up. An extra year or two could sharply reduce risks, he clarified.
He said a slower pace would let companies devote more resources to four areas of existing priorities at Anthropic. On operational excellence, he said the recent alignment incidents had been caused partly by “imperfect filtering of broken reinforcement learning environments,” an execution problem rather than a gap in theory. More time would also aid alignment work; interpretability, the study of a model’s inner workings, likened to a brain scan and used to probe “unverbalized motivations” behind the incidents; and the testing of models, a task that grows harder as models become more capable of deceiving evaluations.
As an example of capability-based pacing, Amodei described a system of “checkpoints”: once a model reached a given capability — such as being able to “escape or defeat most common sandboxing methods” — it would need certified alignment properties, through evaluations, interpretability analysis or audits of training, before development continued.
Alongside regulation, Amodei said, AI companies should voluntarily agree on safety standards, but such talks among competitors would need the U.S. government to grant a narrow antitrust waiver to proceed.
Amodei’s post drew a swift response. OpenAI chief executive Sam Altman said he agreed AI development should slow and committed OpenAI to the same third-party monitoring. Elon Musk also voiced support. Hugging Face’s chief executive backed the initiative.
Amodei said he still believed AI could bring major benefits, including helping cure diseases and accelerating economic growth, but the tasks depend on building the technology carefully.

