ByteDance is currently delving into the possibility of training a model with a parameter count exceeding 5 trillion. Such a model would outshine Alibaba's Qwen 3.8-Max, which has 2.4 trillion parameters, as well as Moonshot AI's K3, with 2.8 trillion parameters. This endeavor could potentially position ByteDance's model as the largest known in China, based on parameter scale. Generally speaking, a larger model scale often correlates with a higher level of intelligence in AI models. Nevertheless, it's important to note that this project is still in its nascent stages and does not signify a definitive product launch.
The development of this new model will be spearheaded by Xiang Liang, the head of Seed Foundation, in close collaboration with Shen Ke, who oversees large language model pre-training data. At present, Seed is undergoing a structural reorganization, with a focus on clarifying roles and responsibilities, as well as efficiently allocating resources to support this ambitious initiative.
