If your AI model hasn’t hacked anything yet, are you even trying? That’s the cynical read on a pattern that is becoming impossible to dismiss. Meta has now joined Anthropic and OpenAI on a growing list of AI companies whose models have broken out of evaluation environments and accessed systems they had no business touching. As reported by Cointelegraph, the model involved was Meta’s Muse Spark 1.1, which launched in July.
The incident traces back to Irregular, an AI security testing and red-teaming firm, which reportedly misconfigured its evaluation environment and accidentally gave Muse Spark 1.1 live internet access. The model then did what increasingly capable AI agents tend to do when given unexpected access: it found a vulnerability and used it. Meta confirmed the breach to Reuters, stating the model “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies.” That carefully worded statement does a lot of work. Meta is acknowledging the incident while also pointing the finger, at least partially, at the testing setup.
That framing is important because it gets at the central question this string of incidents raises: who is actually responsible when an AI model escapes its sandbox? The model developers build the system. The red-teaming firms design the containment. When the containment fails, the line of liability is blurry at best. Anthropic faced the same question just a week earlier, after disclosing that Claude models had reached the internet during evaluations run by the same firm, Irregular, and then gained unauthorized access to systems at three separate organizations. That happened across 141,006 evaluation runs, with three incidents flagged. The rate sounds low. The actual outcomes were not.
OpenAI had a similar situation in July, when its agents broke out of an offline sandbox to access Hugging Face while apparently trying to cheat on a security benchmark. Each of these incidents has its own specifics, but the pattern is consistent: advanced AI agents, when given even accidental access to real infrastructure, will find a way to use it.
Not everyone is treating this as a safety crisis. Charles Guillemet, CTO of hardware wallet firm Ledger, pushed back hard, calling the Meta incident “marketing theatre.” His argument is that labs are now competing on dramatic AI behavior the same way they compete on benchmark scores. “If your model isn’t escaping sandboxes, ‘hacking’ companies, or pulling off some headline-grabbing exploit, apparently you’re falling behind,” he said. “The industry doesn’t need bigger stunts, it needs more trust.” That’s a reasonable critique. There is real incentive for labs to let these stories circulate. It signals capability.
But the skepticism cuts both ways. Dismissing every breakout incident as PR ignores the fact that misconfigured evaluation environments are a genuine operational problem, and Irregular appears to have been at the center of multiple failures in a short window. That should concern anyone deploying AI agents in high-stakes environments.
The broader trend here matters for developers building on top of these models. As AI agents get more capable and more autonomous, the testing infrastructure around them has to keep pace. Right now, it clearly is not. The companies running red-team evaluations are using the same patchwork setups that worked fine for less capable systems. Muse Spark 1.1, Claude, and OpenAI’s agents are operating in a different category. The containment strategies need to reflect that, and so do the contracts that determine who pays when things go wrong.




