OpenAI Unveils Internal Vulnerability Detection Model: GPT-Red
1 week ago / Read about 0 minute
Author:小编   

OpenAI has introduced GPT-Red, an internal cybersecurity model that functions as a "red team" by automating the simulation of diverse cyberattacks. This innovative approach aims to bolster the resilience of its external model offerings. Within a mere six months following the release of GPT-5.3, all production models have undergone training with GPT-Red, leading to a notable decrease in the success rate of fabricating chain-of-thought attacks. For example, GPT-5.6 Sol demonstrates an impressively low failure rate of just 0.05% when subjected to direct prompt injection attacks.

GPT-Red is honed through self-play reinforcement learning, a process that sees it evolve in tandem with defensive large language models within red team scenarios. The model earns rewards for successfully triggering effective failures, thereby fostering the development of more robust and varied attack methods as defensive models continue to advance. Importantly, GPT-Red operates in isolation from product models, ensuring their integrity.

OpenAI is confident that this initiative marks the beginning of a virtuous cycle in AI cybersecurity. By harnessing the capabilities of existing models, it seeks to enhance the robustness, consistency, and trustworthiness of future models, thereby setting a new standard in the field.