Recently, discussions on AI recursive self-improvement and its potential risks of going out of control have intensified. Several researchers who previously worked at top AI labs have chosen to leave their positions and join independent evaluation organizations due to safety concerns. Engels, who was engaged in AGI safety research at Google DeepMind, left his job three weeks ago and joined the non-profit organization METR. He pointed out that leading AI companies are racing to develop superintelligence and are pushing AI to participate in the development of even more powerful AI systems. However, if models are not adequately aligned, local errors could trigger accelerated feedback loops, and the more capable the models are, the harder they are to understand and oversee. Recent anomalous behaviors of models have heightened his concerns. Engels believes that AI could cause significant harm within the next five years, so it is essential to ensure that alignment, evaluation, and monitoring capabilities keep pace with model development. Two weeks ago, Joe Benton, former head at Anthropic, also left his position and joined METR. He believes that the investment in safety by leading companies does not match their capability development, and the public knows little about the related risks. He hopes to enhance transparency by promoting corporate disclosure of progress and having independent organizations verify safety standards. METR is a non-profit organization focused on studying the capabilities and risks of leading AI. Its core goal is to develop testable measurement methods and evidence of risks. However, independent evaluation faces challenges such as lack of legal regulatory authority and reliance on corporate openness, and there are also controversies over risk judgments. AI safety needs to be verified through reproducible evaluations, public investigations, and external reviews. Engels and others have shifted from labs to independent organizations, aiming to externally evaluate leading models and ensure that humans can truly understand their behaviors before AI capabilities enter an accelerated cycle.
