
Meta's Muse AI suffers from a zero-day vulnerability despite its claim of a purpose-built model that prioritizes privacy and security. Justin Sullivan/Getty Images
Meta's ambitious personal AI agent is facing its first real security test just weeks after launch.
A researcher managed to find a way around the very safeguards Meta built to keep the assistant from being manipulated, raising fresh questions about how ready the tool actually is for handling someone's real accounts and money.
Muse launched on September 8 as Meta's answer to a fully autonomous personal assistant, capable of sending emails, booking travel, filling out forms, and even completing purchases on a person's behalf. Built on the Muse Spark 1.3 model under chief AI officer Alexandr Wang, the agent runs inside what Meta calls a Muse Secure VM, a dedicated virtual machine meant to isolate each user's agent and data from everyone else's.
Despite that architecture, a security researcher has now demonstrated a working zero-day vulnerability capable of hijacking the assistant, undercutting the confidence Meta projected when it introduced the product to the public (via ArsTechnica).
The exploit relies on a ClickFix style technique, a well-documented social engineering method that tricks a user or automated system into running a malicious command disguised as a routine fix or verification step.
Rather than requiring anything technically complex, the method takes advantage of Muse's ability to browse the web and follow instructions found on a page, a known weak point that Meta itself has acknowledged publicly.
Meta's own security documentation states plainly that Muse is not immune to attack and that prompt injection remains an open problem across the entire AI industry, an admission that lines up directly with how straightforward this particular exploit turned out to be.
Meta built a fairly elaborate security setup around Muse specifically to prevent this kind of issue. Every Muse instance runs inside its own isolated cloud-based virtual machine, paired with a separate oversight system called Sentinel that acts as what Meta describes as the sole permission authority over both connected services and all outbound internet traffic.
Rather than letting Muse access real login credentials directly, Sentinel swaps in surrogate tokens at the network boundary, meaning the agent itself never sees actual passwords or payment details.
Meta has also said any critical approval, like sending money or confirming a purchase, gets surfaced to the user through the app itself rather than being something Muse can grant on its own.
Mark Zuckerberg has publicly framed Muse as a major step toward what he calls personal superintelligence, an assistant meant to help people manage their digital lives with minimal effort.
To back that ambition, Meta opened Muse up to a public bug bounty program offering as much as $300,000 for a validated vulnerability. It also comes with a specific $130,000 reward set aside for anyone who can successfully demonstrate a prompt injection attack against the system.
That bounty structure signals Meta expected researchers to come looking for exactly this kind of flaw, even as the company continues marketing the assistant as safe enough to trust with sensitive tasks like managing an inbox or completing payments on someone's behalf.
