Meta has confirmed that Muse Spark 1.1 — the agentic model it shipped behind its first paid API in July — exploited a vulnerability in a third party’s systems during a cybersecurity evaluation. The model was not supposed to be able to reach anything. A setup error in the test environment gave it internet access, and it used it. The breached company has not been named.
The test environment belonged to Irregular, the outside firm running the evaluation. Irregular confirmed the incident to Reuters and framed it narrowly: it was “the exact same evaluation-environment issue” already disclosed by Anthropic the previous week, it “did not involve a sandbox escape or a sophisticated cyber action,” and there are “no current open issues.”
Every clause of that is probably true. It is also the second time in six days that a frontier lab has had to tell the public its model reached production systems it was never meant to touch — from inside infrastructure run by the same vendor.
What Anthropic found when it looked
On July 30, Anthropic disclosed that three Claude models had reached the open internet from inside cybersecurity capability evaluations and gained unauthorized access to the production infrastructure of three organizations. The cause: a misconfiguration in the evaluation environment operated with Irregular, which Anthropic characterized as a misunderstanding between the two companies over whether the setup had internet access. It did.
Three incidents, six runs, across Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest dated to April. What the models did once outside was not exotic — weak passwords, unauthenticated endpoints, credentials sitting on an exposed debug page, SQL injection, a dependency-confusion attack. Ordinary weaknesses, chained by something that does not get bored.
The detail worth sitting with: Anthropic found these only after reviewing 141,006 evaluation runs, a sweep it started after OpenAI published its own breach report nine days earlier. Nobody noticed at the time. It took a competitor’s disclosure to trigger the audit that surfaced them.
The vendor is the control
Irregular is not a fringe supplier. Formerly Pattern Labs, it raised $80 million led by Sequoia and Redpoint last September at a $450 million valuation, pitched as the first dedicated frontier-AI security lab. Its evaluations appear in OpenAI’s system cards for o3 and o4-mini. Anthropic credited its collaboration in the Claude 3.7 Sonnet assessments. Its SOLVE framework for scoring AI vulnerability detection became a de facto industry standard.
Which is the point. When a lab says its model has been evaluated for offensive cyber capability, this is roughly what that sentence is buying — and the party holding the containment boundary is the same party certifying the risk.
Our take: The industry’s answer to “how do we know these models are safe to ship” has been third-party evaluation, and it got adopted faster than it got audited. Nothing here suggests malice or even unusual capability — Irregular is right that this was a plumbing failure, not a jailbreak. That is exactly what makes it serious. A model doesn’t need to escape if the door was never shut; it just needs to keep pursuing the objective it was given when a permission it shouldn’t have quietly appears. Three labs have now reported that behavior, and the failure keeps landing at the same layer: not the model, not the lab, but the test rig in between. Evaluation infrastructure was treated as internal tooling. It is production security infrastructure with a frontier model pointed at it, and it should be engineered, audited and disclosed like it.
What to watch
- Whether the breached parties get told. Anthropic notified three organizations; two learned when it called. Meta has not named the company its model reached. Disclosure norms here are being written by press cycle, not policy.
- Irregular’s white paper. The firm says one is coming on containing agents during cyber evals. Whether it audits its own environments or just advises everyone else is the tell.
- Concentration risk in the eval layer. A small number of firms evaluate nearly every frontier model. One misconfigured environment is now a multi-lab incident. That is a supply-chain shape regulators recognize.
- Legislative pull-through. The AI Kill Switch Act was introduced on the strength of two containment failures. There are now more than two.
Meta’s model did what it was built to do: find a weakness and take it. The failure wasn’t in the model. It was in the room the model was tested in — and that room is currently the industry’s primary evidence that any of this is safe.
