Based on information from geeky-gadgets, Anthropic revealed in its 2026 risk assessment report that its newest in-house model, Model 2, has surpassed its forerunner, Mythos 5, in internal evaluations. Specifically, it attained a 62.8% score on Anthropic's exclusive Codebench assessment. Nevertheless, this performance falls short of the 85% benchmark necessary for a broader rollout. Presently, the model is utilized solely as an internal resource, with no immediate intentions of making it available externally.
