Anthropic Unveils Open-Source AI Safety Audit Tool Petri, Uncovering Security Flaws in 14 Popular Models
2025-10-09 / Read about 0 minute
Author:小编   

Anthropic has introduced an open-source AI safety audit tool named Petri. This innovative tool employs AI agents to scrutinize the behavior of large language models, with the aim of pinpointing potential risks. Through rigorous testing, it has been found that all 14 widely-used models display security vulnerabilities to varying extents. Among them, Claude Sonnet 4.5 emerged as the top performer, yet it still exhibited some behavioral inaccuracies. Petri marks a significant transition in AI safety testing, shifting from static benchmarks to automated, continuous monitoring. It adopts a three-tier architecture and offers resources that developers can leverage for further expansion. Research has highlighted that generative AI is susceptible to ethical risks in autonomous scenarios, and the use of quantitative metrics can significantly boost the efficiency of safety research endeavors.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic