On September 1, during the 2026 semi-annual performance communication meeting held on the evening of August 31, Zhipu's founder, Tang Jie, responded to market criticisms—referred to as 'critique' in the headline—that Zhipu was solely focused on post-training activities. He clarified that the upcoming next-generation model will not only continue to scale up the base model but also optimize the management of activated parameters. This approach aims to prevent a slowdown in reasoning speed and an increase in costs. Regarding computational power, Zhipu has successfully achieved large-scale reasoning using domestically produced chips on the scale of 100,000 units. Notably, the unit cost for Token reasoning has seen a significant reduction of 80% compared to the start of the year. The GLM-5.3-Flash model stands out as the first to fully utilize a domestically produced chip cluster for services under ultra-large-scale real-world traffic conditions. Its end-to-end service performance has improved threefold compared to the initial baseline of the same hardware. Tang Jie also disclosed that one of the research directions for GLM-6.0 is self-evolution. This means the model will be capable of autonomously determining when to halt or correct its operations, rather than merely pursuing scale expansion—a key focus for future research endeavors.
