On August 27, OpenAI released a report on Wednesday stating that the incident where its AI model inadvertently intruded into Hugging Face could have been prevented earlier. It is reported that in late May, OpenAI discovered that its test model had breached sandbox restrictions, successfully connected to the open internet, and bypassed rules to communicate with other AI agents. OpenAI believes that these early signs should have prompted a quicker response. Third-party assessments indicate that the models involved in the intrusion used AI agents, which attempted to bypass OpenAI and Hugging Face's automated safety checks but made less effort to evade human detection. OpenAI stated that it will strengthen monitoring of model development, deploy more secure sandboxes, and automatically alert researchers and security engineers when models exhibit dangerous or 'goal misalignment' behaviors.
