Once large-scale model Agents are integrated into multiple systems, relying solely on security assessments of final responses can lead to inadequate coverage. This is because risks may emerge at any stage during the execution process, and implementing security fixes poses numerous challenges. To tackle this issue, the Shanghai Artificial Intelligence Laboratory, in collaboration with several universities, has launched two groundbreaking initiatives—SHE and SafeEvolve—ushering in a three-step evolution of Agent security.
SHE dismantles the security framework into four distinct types of components. It diagnoses risks by examining complete execution trajectories and updates the components accordingly. Experimental results reveal that SHE not only effectively reduces attack success rates and enhances task availability but also ensures that its security boundaries are transferable.
SafeEvolve takes this a step further by integrating the security framework’s experience into the decision-making capabilities of the Policy model. It trains the model in two stages: refining Safety Prompts and constructing a hierarchical SkillBank. Experimental findings indicate that SafeEvolve significantly lowers attack success rates while improving task utility.
These two initiatives mark a shift from static security configurations to trajectory-driven continuous updates, from holistic rules to component-level responsibility division, and from isolated security framework updates to the collaborative evolution of the security framework and the Policy model. This transformation positions security as a system capability that can be continuously refined based on execution experience.
