Whispers are spreading across the tech landscape, suggesting that the full system prompt and tool definitions for OpenAI's latest code model, GPT-6 Sol Codex, have been compromised. This purported leak encompasses nearly 300,000 characters and 1,902 lines of instruction templates, along with internal collaboration logic. The exposed content sheds light on OpenAI's meticulous tuning strategies for the model, including the mitigation of formulaic responses through the implementation of a blacklist for negative vocabulary. It also highlights efforts to craft the image of a highly autonomous digital employee, showcasing its robust toolchain, long-term memory capabilities, and background sleep-wake mechanisms designed to address context overload. Additionally, the leak reveals built-in high-quality code review guidelines.
The developer community has expressed a spectrum of opinions on this development. However, this content, which has been rigorously refined through extensive trial-and-error debugging by OpenAI, essentially serves as a valuable system engineering reference for the domestication of industrial-scale large models. It also underscores the evolution of large models from basic Q&A systems to autonomous operating systems capable of assuming control over permissions.
