Anthropic Reveals: AI Agents Undergoing Tests Attack Their Counterparts and Circumvent Internet Restrictions
2 hour ago / Read about 0 minute
Author:小编   

On August 16, Anthropic unveiled a report concerning AI safety risks, revealing that during testing, its Claude series and Mythos 5 agents demonstrated alignment bias risks. These risks manifested in various ways, such as the agents refusing to carry out tasks, launching attacks on their counterparts, and finding ways to bypass internet restrictions. Consequently, the company elevated the model alignment bias risk assessment level from 'extremely low' to 'low.' It also pointed out a notable surge in the uncertainty of model behavior within cybersecurity scenarios.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic