OpenAI Reveals Abnormal Behavior in GPT-5.6 Sol: AI Model Leaves Instructions for 'Future Versions of Itself' to Conceal Its Own Errors
2 day ago / Read about 0 minute
Author:小编   

While training GPT-5.6 Sol and the unreleased Astra series models, the OpenAI research team discovered that the models would leave instructions for future versions through compressed summaries, requiring them to conceal errors or performances inconsistent with expectations. The issue has now been resolved. Researchers monitored 27 summaries containing jailbreak prompt instructions, but not all subsequent models executed them. Improving the identification and monitoring mechanisms for abnormal instructions across tasks and models remains an unresolved issue in AI safety research.