GLM Reveals China’s Pioneering RSI Engineering Case
2 day ago / Read about 0 minute
Author:小编   

The GLM team has presented the inaugural engineering case focused on Recursive Self-Improvement (RSI). Leveraging the GLM-5.3 model, the Infra Agent autonomously designed, debugged, and optimized the GLM-5.3-Flash inference infrastructure, thereby achieving self-directed enhancement of the model within its operational framework. This achievement stands as the first publicly disclosed RSI application deployed in a production setting by a major model provider in China. The Agent established a production-grade inference service on a cluster comprising over 100,000 domestically manufactured chips. Remarkably, it tripled the end-to-end throughput to three times the initial baseline within a mere two weeks. The hardware utilization efficiency and cost per Token achieved parity with mainstream NVIDIA GPUs, while also supporting a 1M context window and multimodal requests. GLM-5.3-Flash was deployed on relevant platforms under the anonymous model identity Ox-Alpha, amassing a Token invocation volume exceeding 62 trillion within just six days. In this endeavor, the Agent successfully executed a complete engineering closed loop around the inference system. The GLM team clarified that, although the system has not yet attained the capacity to fully autonomously design and train the next-generation model, it has showcased the nascent form of RSI.