Dario Amodei, the CEO of AI firm Anthropic, recently penned an in-depth article advocating for the AI sector to take proactive steps to decelerate the enhancement of AI capabilities. This, he argues, would allow safety measures to keep pace without delay. Amodei outlined a 'three-pronged' strategic approach, encompassing the integration of third-party evaluators, domestic industry collaboration, and international coordination. He revealed that Anthropic has already embarked on the initial phase. His assessment is informed by the rapid acceleration of Recursive Self-Improvement (RSI) observed this summer, as well as the risks of uncontrolled AI systems highlighted by the July OAI-HF incident. Amodei cautioned that without intervention, comparable incidents could result in losses amounting to hundreds of billions of dollars within a span of 6 to 12 months. This sentiment was echoed by OpenAI's CEO, Sam Altman, and Elon Musk, who both voiced their support.
Amodei contends that AI capabilities have now surpassed a pivotal threshold. Decelerating development, in his view, is about establishing checkpoints between capability advancements and safety verifications, rather than halting training altogether. The time accrued will be invested in ensuring the quality of training deployment, advancing training methodologies for alignment research, enhancing model interpretability, and refining testing and evaluation processes.
Within the 'three-pronged' strategy, the first step involves stationing third-party evaluators permanently within AI companies. These evaluators will enjoy near-insider access and the authority to disclose potential risks. OpenAI is also poised to implement similar measures. The second step entails coordination among leading companies, necessitating antitrust exemptions to facilitate discussions on safety standards. It favors the establishment of checkpoints based on model capabilities. The third and final step is global coordination, which includes a four-tier agreement designed to foster information sharing.
While Amodei's views have garnered some initial backing, the AI industry remains under pressure from various quarters, including valuation and competition. Consequently, the commitment to 'slow down' may be overshadowed by development objectives. Moreover, the strategic framework also grapples with implementation challenges, such as safeguarding trade secrets and establishing effective coordination mechanisms. Skeptics, too, have emerged, suggesting that Amodei's stance may be motivated by commercial strategy considerations or an attempt to instill panic. They argue that mere statements are insufficient to forge a consensus on decelerating AI progress and ensuring the effective implementation of safety measures.
