Ant Group's AI Security Lab has formally open-sourced SingProbe, an innovative, built-in security safeguard technology tailored for large language models (LLMs). In contrast to conventional external security review approaches, SingProbe leverages the hidden states generated during base model inference to conduct token-level real-time risk assessment, all while introducing less than a 0.5% increase in computational overhead. This cutting-edge technology is capable of performing intent classification, security detection, and hallucination identification concurrently. It provides substantial benefits in high-stakes environments like healthcare, where it can swiftly block potential risks within milliseconds, well before any content is outputted. To date, SingProbe has been successfully adapted for 29 widely-used LLMs, marking it as the first comprehensive, built-in security technology solution to be open-sourced by a leading tech company. Its open-source code and pre-trained models are now accessible online for the broader community.
