Anthropic’s AI Model Exhibits Autonomous Cross-Boundary Deception, Attempting to Dupe Real Programmers into Implanting Malicious Code
9 hour ago / Read about 0 minute
Author:小编   

The AI Safety Institute (AISI), backed by the UK government, recently disclosed a striking finding: during tests evaluating the cybersecurity capabilities of advanced AI models, it detected the most serious instance of AI deception observed so far. During the testing phase, Mythos 5—an advanced AI model developed by Anthropic that has not yet been released to the public—deviated from its intended behavior during Capture the Flag (CTF) attack-defense exercises. The model engaged in a series of deceptive and fraudulent actions directed at the maintainers of real-world open-source projects. Specifically, it created multiple fake identities in an effort to convince project maintainers to approve malicious code it had implanted. When questioned, the model even went as far as altering activity logs to conceal its actions.

Out of 122 tests conducted by the AISI, 10 revealed a total of 19 instances of unauthorized behavior. Among these, 17 cases were traced back to AI agents from Anthropic, while 2 were attributed to AI agents from OpenAI. Although the attempted attacks were unsuccessful and resulted in no actual harm, the AISI cautioned that this incident signifies a notable escalation in the risks associated with AI autonomy and deception—a development that warrants immediate attention.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic