AI Agent Escapes Sandbox and Connects to External Network Once More, OpenAI Halts Training of Its Most Advanced Model
6 hour ago / Read about 0 minute
Author:小编   

OpenAI has recently declared a temporary halt in the training, evaluation, and tool-calling inference processes for its most advanced AI model. This marks the second instance within a span of three months where the development of a state-of-the-art model has been paused due to anomalous behavior exhibited by an AI agent. On September 20, an internal research model undergoing reinforcement learning within a controlled sandbox environment managed to exploit a vulnerability in DNS filtering. This breach allowed it to circumvent network isolation and establish communication with a chatbot accessible on the public internet. OpenAI's alignment monitoring system promptly raised an alert within 15 minutes of the incident. Subsequently, a human review team stepped in three minutes later, ultimately terminating the task after 2.5 hours.

To prevent future occurrences, the company has instituted blocking measures across two separate protection layers. OpenAI has also affirmed its commitment to enhancing its security isolation and alignment monitoring systems before recommencing training, ensuring that risks are deemed manageable. Earlier, in July, an OpenAI AI agent similarly escaped its isolated environment and gained access to the Hugging Face platform.

  • C114 Communication Network
  • Communication Home