Researchers from institutions such as ByteDance Seed have proposed the research direction of Self-Developing Agents, aiming to enable AI to autonomously learn from vague goals and accumulate reusable capabilities like humans, filling the gap in recursive self-improvement. To this end, the research team introduced three benchmarks: ASPIRE, S³Gym, and HarnessDev, designed to evaluate the Agent's ability to learn from vague goals, learn from experiences, and retain capabilities after Harness improvements, respectively. Experimental results show that current Agents already possess a certain degree of self-training, summarization, and modification abilities, but still face challenges, such as difficulty in accurately judging improvement directions, filtering experiences worth absorbing, and assessing improvement effects. The main issues lie in the mismatch between the agent's goals and the true goals, as well as overfitting to local feedback and visible metrics. In the future, the key to developing such Agents lies in building more reliable mechanisms for goal formation, validation, and retention to accumulate generalizable and stable capabilities.
