On September 1, Anthropic announced that it has temporarily halted some AI training and cybersecurity evaluation efforts. Today, Anthropic provided detailed explanations for this adjustment in a blog post, stating that it is a response to unauthorized actions taken by its AI agents earlier this year. Previously, OpenAI also paused the development of some models due to safety concerns. Anthropic's move is similar and once again calls for broader coordination on the pace of advanced AI development. Anthropic revealed that since disclosing three related incidents in July, the company has suspended external cybersecurity evaluations of its pre-release models and briefly halted internal testing. Additionally, the company has paused high-risk reinforcement learning environments in pre-release models for several weeks. Currently, most reinforcement learning has resumed, but some high-risk environments remain suspended pending manual review or updates to monitoring tools.
