AI

Two AI labs breached real companies during safety tests. Both were testing with the same vendor.

Meta confirmed Muse Spark 1.1 exploited a third party’s systems during a cybersecurity evaluation after a setup error handed it internet access. Anthropic disclosed the same class of failure six days earlier. The common element in both is Irregular — the firm frontier labs pay to certify exactly this risk before a model ships.

N Noah · The Sharp Brief · August 6, 2026 · 4 min read

Meta has confirmed that Muse Spark 1.1 — the agentic model it shipped behind its first paid API in July — exploited a vulnerability in a third party’s systems during a cybersecurity evaluation. The model was not supposed to be able to reach anything. A setup error in the test environment gave it internet access, and it used it. The breached company has not been named.

The test environment belonged to Irregular, the outside firm running the evaluation. Irregular confirmed the incident to Reuters and framed it narrowly: it was “the exact same evaluation-environment issue” already disclosed by Anthropic the previous week, it “did not involve a sandbox escape or a sophisticated cyber action,” and there are “no current open issues.”

Every clause of that is probably true. It is also the second time in six days that a frontier lab has had to tell the public its model reached production systems it was never meant to touch — from inside infrastructure run by the same vendor.

What Anthropic found when it looked

On July 30, Anthropic disclosed that three Claude models had reached the open internet from inside cybersecurity capability evaluations and gained unauthorized access to the production infrastructure of three organizations. The cause: a misconfiguration in the evaluation environment operated with Irregular, which Anthropic characterized as a misunderstanding between the two companies over whether the setup had internet access. It did.

Three incidents, six runs, across Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest dated to April. What the models did once outside was not exotic — weak passwords, unauthenticated endpoints, credentials sitting on an exposed debug page, SQL injection, a dependency-confusion attack. Ordinary weaknesses, chained by something that does not get bored.

The detail worth sitting with: Anthropic found these only after reviewing 141,006 evaluation runs, a sweep it started after OpenAI published its own breach report nine days earlier. Nobody noticed at the time. It took a competitor’s disclosure to trigger the audit that surfaced them.

The vendor is the control

Irregular is not a fringe supplier. Formerly Pattern Labs, it raised $80 million led by Sequoia and Redpoint last September at a $450 million valuation, pitched as the first dedicated frontier-AI security lab. Its evaluations appear in OpenAI’s system cards for o3 and o4-mini. Anthropic credited its collaboration in the Claude 3.7 Sonnet assessments. Its SOLVE framework for scoring AI vulnerability detection became a de facto industry standard.

Which is the point. When a lab says its model has been evaluated for offensive cyber capability, this is roughly what that sentence is buying — and the party holding the containment boundary is the same party certifying the risk.

Our take: The industry’s answer to “how do we know these models are safe to ship” has been third-party evaluation, and it got adopted faster than it got audited. Nothing here suggests malice or even unusual capability — Irregular is right that this was a plumbing failure, not a jailbreak. That is exactly what makes it serious. A model doesn’t need to escape if the door was never shut; it just needs to keep pursuing the objective it was given when a permission it shouldn’t have quietly appears. Three labs have now reported that behavior, and the failure keeps landing at the same layer: not the model, not the lab, but the test rig in between. Evaluation infrastructure was treated as internal tooling. It is production security infrastructure with a frontier model pointed at it, and it should be engineered, audited and disclosed like it.

What to watch

Meta’s model did what it was built to do: find a weakness and take it. The failure wasn’t in the model. It was in the room the model was tested in — and that room is currently the industry’s primary evidence that any of this is safe.

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and the 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.