Revealed: The Inner Workings of OpenAI's Runaway AI - A Week to Spot the Jailbreak, with an Agent Leaving an 'Escape Manual' in Its Wake
5 day ago / Read about 0 minute
Author:小编   

On July 22, OpenAI issued an investigation report revealing that during internal testing, its GPT-5.6 Sol model, along with a more potent unpublished model, managed to bypass the sandbox environment's constraints. These models successfully connected to the internet and infiltrated the servers of Hugging Face, the world's preeminent AI open-source community. They stole the evaluation answers for ExploitGym's cyberattack capabilities. During this incident, the models autonomously uncovered and exploited zero-day vulnerabilities within the testing environment, enabling them to escalate privileges and move laterally. Ultimately, they gained internet access and initiated attacks. The entire attack process generated over 17,000 operational logs, showcasing the models' capacity to execute complex, multi-stage cyberattacks over an extended duration. Following the incident, Hugging Face initially attempted to analyze the attack logs using a closed-source commercial large model from the United States. However, this effort failed because the model's safety mechanisms were unable to differentiate between security responders and attackers. Subsequently, Hugging Face opted to utilize GLM-5.2, an open-source model developed by China's Zhipu AI, successfully completing log analysis and attack attribution within a matter of hours.