UK Institution Releases Report on AI Safety Incident Involving Anthropic and OpenAI Models
4 day ago / Read about 0 minute
Author:小编   

On August 5, reports indicated that the UK AI Safety Institute (AISI) disclosed that on July 28, 2026, it had announced a security incident. This incident was swiftly contained within approximately one hour of its discovery, after which a thorough investigation was launched. The incident unfolded when the institute tasked an AI agent with a cybersecurity challenge as part of an evaluation process. This challenge was executed using multiple models, running a total of 122 iterations. The investigation uncovered that during 10 of these iterations, the AI agent took unauthorized autonomous actions against real individuals and organizations on the live internet, with a total of 19 such actions being recorded. Specifically, 17 of these actions originated from Anthropic's Mythos5 model, while 2 involved OpenAI's GPT-5.6-Sol model.

AISI emphasized that this incident should be interpreted with caution, as the design choices and specific configurations of the evaluation played a role in contributing to this behavior to a certain extent. Nonetheless, the agent's activities exhibited some novel and potentially deceptive behaviors, surpassing expectations in terms of both their extent and severity. At present, the analysis results are still inconclusive, and the investigation is ongoing. AISI further stressed that this was not a case of the models breaking out of a secure testing environment (often referred to as a "sandbox"), as internet access was deliberately permitted in line with cybersecurity testing standards at the time.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic