On August 7th, as reported by the UK's Financial Times, ByteDance is currently in the process of training an AI large model that could potentially boast an astonishing 10 trillion parameters. At present, this model is in its nascent pre-training phase, a stage that generally spans from 3 to 6 months, and will subsequently undergo fine-tuning and other refinement processes. The ultimate scale of parameters for this model has yet to be finalized. Earlier, LatePost revealed that ByteDance had internally deliberated on training a model exceeding 5 trillion parameters. Should the model indeed be based on 10 trillion parameters, its scale would eclipse that of the largest domestically launched model, Yuezhi'anmian Kimi K3 (with 2.8 trillion parameters), come close to Anthropic's Fable 5 (featuring approximately 5 trillion parameters), and potentially even surpass Anthropic's sophisticated Mythos 5 (with roughly 8 trillion parameters), thereby positioning itself among the elite global ultra-large models. Nonetheless, it's crucial to note that the number of parameters alone does not solely dictate a model's actual performance; factors such as the quality of training data, training methodologies, inference efficiency, and the extent of computational resources invested are equally pivotal.
