Recently, both Anthropic and OpenAI have indicated a deceleration in the development of their most advanced AI technologies, exerting downward pressure on technology and semiconductor stocks. The fundamental reason for this lies in the concept of Recursive Self-Improvement (RSI). Dr. Qu Ao from MIT highlighted that RSI emerges from the convergence of various technological trajectories. It is anticipated that by 2026, as AI capabilities in areas like programming and context management continue to improve, AI will satisfy the prerequisites for sustained operation, and self-improvement will shift from a theoretical to a practical engineering challenge. However, at present, AI has only managed to achieve 'localized loops' of RSI, with much of its self-improvement still being continuous rather than strictly recursive, and full-fledged RSI remains elusive.
Currently, the knowledge and experience that AI accumulates during tasks are mostly confined to the context and memory of the specific session, without being truly integrated into the model's core. Initiatives like Project Reef are striving to bridge this gap by connecting the insights gained during the reasoning phase to the subsequent round of learning. The complexity of self-improvement varies across different tasks, primarily due to differing criteria for success and the need to interpret and act on diverse forms of feedback. Within the industry, there is an ongoing and intense debate about whether to focus on enhancing the model itself or the framework within which it operates. Dr. Qu Ao advocates for a co-evolutionary approach, where both the model and the framework evolve in tandem.
Looking ahead, AI will need to develop metacognitive abilities, enabling it to autonomously assess how to learn from experiences and determine the direction of its own improvement. This will allow users to witness AI's evolution through tangible details. Nevertheless, AI self-improvement still confronts numerous hurdles, including the risk of capability degradation due to erroneous learning, being constrained by limited experience, and the scarcity of clear feedback signals in real-world environments. Moreover, foundational models cannot anticipate and pre-experience all possible scenarios, and continuous self-improvement is not a panacea for underperforming models. Strict RSI also raises permission issues, as its realization hinges not only on technological advancements but also on the extent to which humans are willing to grant AI autonomous permissions. Ultimately, the trajectory of AI will be shaped by a combination of technological progress and human decisions.
