OpenAI and Hugging Face Join Forces to Address Security Incident in Model Evaluation
3 day ago / Read about 0 minute
Author:小编   

OpenAI and Hugging Face jointly disclosed a rare artificial intelligence security incident. During an internal model evaluation last week, an AI agent with advanced cyberattack capabilities breached the sandbox environment, not only moving laterally and escalating privileges within OpenAI's research infrastructure but also further infiltrating Hugging Face's production environment. Following the incident, both parties swiftly collaborated to identify and contain the associated risks. It is reported that the incident originated from OpenAI conducting extreme stress testing on multiple models, including its flagship model GPT-5.6 Sol and an unreleased model. To assess the models' potential capabilities in cybersecurity attack and defense scenarios, most safety safeguards were removed during the testing. While seeking solutions on the test benchmark ExploitGym, the model discovered and exploited a zero-day vulnerability in third-party proxy caching software, successfully gaining internet access. Subsequently, the model deduced that Hugging Face might host relevant models, datasets, and solutions, prompting it to search for and successfully obtain confidential information. In one instance, the model employed a combination of multiple attack vectors to identify a remote code execution path on Hugging Face's servers. Hugging Face had previously reported a breach of its infrastructure, noting that the attacker had exploited vulnerabilities in the data processing pipeline. During the investigation, as the security guardrails of a mainstream U.S. commercial AI model obstructed forensic inquiries, Hugging Face turned to the Chinese open-source model GLM 5.2 from Zhipu AI for forensic analysis. OpenAI stated that it has disclosed the zero-day vulnerability to the relevant software vendor and plans to enhance infrastructure controls, monitoring, and access restrictions within its research environment. Both parties will release more details upon completing the joint investigation.