RSI Gains Momentum, Yet Self-Evolution Runs Risk of Overfitting: Google and Other Researchers Introduce RRSI, Incorporating Regularization into Self-Evolution
1 day ago / Read about 0 minute
Author:小编   

Within the realm of agent technology, Recursive Self-Improvement (RSI) has discovered a viable path for practical implementation. Agents can establish a closed-loop system for self-improvement by iteratively refining peripheral Harness components—such as Prompts, tool invocation rules, and memory management strategies—without altering the model's inherent weights. Nevertheless, the conventional, unrestricted process of Harness evolution is plagued by three significant challenges: continuous optimization based on the same set of tasks can readily result in benchmark fitting, pursuit of assessment noise, and escalation of complexity. These issues lead to a situation where scores within the evolution set continue to climb, yet performance on unseen Out-of-Distribution (OOD) tasks remains lackluster, with diminishing returns and escalating inference costs. Google's innovative Regularized Recursive Self-Improvement (RRSI) framework does not impose constraints on the agent's capacity for self-modification; rather, it imposes limitations on the complexity of modifications proposed in each round and maintains a comprehensive record of the evolution history to steer exploration. From the filtering perspective, it rigorously assesses the generalization capabilities and cost-effectiveness of proposed modifications, thereby eliminating ineffective changes. Experimental results demonstrate that while RRSI's scores within the evolution set are marginally lower than those achieved through unconstrained evolution, it excels across multiple types of OOD benchmarks that were not involved in the evolution process, while concurrently achieving a substantial reduction in inference token consumption. Over the course of 30 evolution rounds, only approximately one-third of the modifications were deemed effective and retained. The study underscores that the essence of RSI does not lie in the pursuit of unbridled self-modification, but rather in attaining generalizable self-improvement, thereby enabling agents to discern and integrate truly universal and effective advancements.