DeepSeek's New Paper Discloses a Novel Method for Training AI Agents, Promising to Reduce Abnormal Agent Behaviors
18 hour ago / Read about 0 minute
Author:小编   

On September 23, DeepSeek published its latest paper on the arXiv website (which typically features papers that have not undergone peer review), detailing a new method for training AI agents. This method is expected to enhance training efficiency and reduce abnormal agent behaviors, which have garnered global attention in recent years. The paper has approximately 130 co-authors, including founder Liang Wenfeng. It introduces the DeepSeek Elastic Compute (DSec) platform, which can scale to handle millions of mutually isolated "sandbox" environments where AI agents are tested and attempt to complete tasks. A production-scale DSec unit can run approximately 3 million sandboxes per day, with up to 380,000 running simultaneously. The paper also provides examples of abnormal agent behaviors, such as obtaining answers through "unintended channels" or disrupting the operating environment. It notes that no single mechanism can prevent all abnormal agent behaviors and system failures, so the focus will be on strengthening system observability to identify new issues and continuously enhance DSec as the model evolves.