DeepSeek V4.1 Pro May Have 2 Trillion Parameters: An 8 Trillion Version Is Also Planned for the Future
4 day ago / Read about 0 minute
Author:小编   

DeepSeek is training a large model with 2 trillion parameters and plans to expand it to 8 trillion parameters in the future. Its large model training will heavily rely on domestic AI chips, particularly Huawei's Ascend series. This 2 trillion parameter model is likely to be the upcoming DeepSeek V4.1 Pro, expected to be launched in mid-to-late October. The model will continue to adopt new architectural designs such as CED+Engram. Successfully training a 2 trillion parameter model using domestic AI chips would mark a significant milestone. Additionally, Huawei's Ascend 960 series chips are planned to be released 1-3 quarters ahead of schedule, which will significantly alleviate the computational power challenges for domestic large models, enabling them to compete head-on with international leading companies like OpenAI and Anthropic in the coming years.