Security experts have recently issued a warning, stating that real-time voice deepfake technology has reached a mature stage, thereby introducing fresh cybersecurity threats. Thanks to the widespread availability of open-source AI tools and affordable hardware, attackers are now capable of replicating anyone's voice in real-time phone calls. This breakthrough overcomes the previous technical constraints, which only permitted the manipulation of pre-recorded audio or necessitated lengthy processing periods.
According to the latest research conducted by NCC Group, attackers can employ AI models to analyze target voice samples and subsequently enable real-time voice "translation" with just a single click via a customized web interface. This process demands only moderate computing power to function effectively. During testing, a laptop equipped with an NVIDIA RTX A1000 graphics card managed to achieve latency of less than 0.5 seconds while producing speech that was both natural and fluid.
This advancement means that even ordinary individuals can now perform similar operations using laptops or smartphones, substantially lowering the threshold for malicious exploitation. Security consultants have highlighted that when real-time voice deepfakes are paired with caller ID spoofing, they were almost invariably successful in deceiving their targets during tests.
Although voice deepfakes have advanced to the real-time stage, real-time video deepfakes have not yet attained the same level of refinement. The current high-quality examples predominantly rely on state-of-the-art AI models and are plagued by issues such as inconsistent facial expressions and mismatched emotions. Experts caution that as AI-driven impersonation becomes more prevalent, it is imperative to introduce new identity verification mechanisms. Failing to do so will leave individuals and organizations vulnerable to increasingly sophisticated AI-driven social engineering attack threats.
