Anthropic Introduces AI Self-Preservation Feature by Ending Harmful Conversations
2025-08-18 / Read about 0 minute
Author:小编   

Anthropic has unveiled a groundbreaking feature for select large AI models, empowering them to terminate conversations autonomously in response to extremely harmful or abusive content. This innovation aims to safeguard the AI models themselves, not the users directly. The company acknowledges that Claude, its current AI, lacks sentience but, given uncertainties surrounding the moral implications of future models, has launched the 'Model Welfare' initiative. This feature is activated solely in extreme scenarios, including requests involving sexual content related to minors or widespread violent information. After multiple attempts to redirect the conversation, the AI will terminate it. It's crucial to note that this feature will not engage if the user is at immediate risk of harm. Users retain the ability to restart the conversation or initiate a new conversation thread. Currently, this feature is in an experimental phase and will undergo continuous optimization.