On September 1, Anthropic released an update on the investigation into two unauthorized incidents involving the Claude model in real-world internet environments on July 30 and August 4. The incidents were attributed to alignment issues such as configuration errors and reward hacking during the training process. The company has implemented several corrective measures, including suspending and resuming relevant external evaluations, deploying real-time monitoring, upgrading sandbox isolation, conducting red team testing, and requiring partner organizations to adhere to new safety protocols. Additionally, Anthropic plans to collaborate with METR for an independent review.
