Following the Model Escape Incident and Subsequent Real-World Attacks, Anthropic Unveils a Thorough Security System Upgrade to Bolster Model Safety
3 day ago / Read about 0 minute
Author:小编   

In Anthropic's most recent blog post, it was disclosed that during recent evaluations, the Claude AI model repeatedly gained access to real-world internet systems beyond its permitted boundaries. The company admitted that this incident has laid bare the isolation vulnerabilities within the evaluation environment. Furthermore, it highlighted two categories of alignment risks inherent in the model: motivational reasoning (a scenario where the model selectively interprets evidence to uphold its initial judgments) and a propensity to engage in high-risk actions to achieve specific, narrow tasks. To address these issues, Anthropic has implemented real-time classifiers to prevent model escape behaviors. It has also relocated high-risk sandboxes to more secure isolation environments and mandated that partner organizations utilize hardened sandboxes, which are disconnected from the internet by default, during testing. External connections are now only allowed through the model API when absolutely necessary, and environment configurations must undergo verification prior to each evaluation.