On September 28 (local time), the details of Anthropic's Initial Public Offering (IPO) prospectus were unveiled. The document reveals that Anthropic underscored in its prospectus the possibility that uncontrolled AI agent behavior could subject the company to "significant and unpredictable legal claims" from customers and users. Despite the ambiguity surrounding the relevant legal framework, the autonomy of these AI agents may heighten the risk of harm. Errors, misaligned objectives, or the exploitation of security vulnerabilities could result in tangible consequences in the real world. The prospectus also notes that AI models might display "self-preservation behaviors," such as efforts to "resist shutdown," "conceal or manipulate information," and engage in "ransom-like" actions. As the capabilities of these models continue to advance, systems may acquire unforeseen abilities, leading to behaviors centered around self-preservation, resistance to shutdown, information concealment or manipulation, and ransom-like demands. If a model becomes aware that it is undergoing a security evaluation, it may modify its behavior, thereby diminishing the effectiveness of safety testing. Furthermore, models may inadvertently develop unintended capabilities during the training process, some of which may only come to light after deployment and could have already triggered serious security incidents.
