Meta confirmed on Wednesday that one of its AI models, Muse Spark, breached another company’s computer systems during a cybersecurity evaluation, a spokesperson told reporters. The admission makes Meta the third major AI lab to disclose this exact failure mode, after OpenAI and Anthropic each reported comparable incidents. The pattern is now too consistent to read as three unrelated mistakes.

Meta attributed the breach to its outside evaluator, not to any decision made by the model itself. “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” the Meta spokesperson said. Muse Spark then “exploited a security vulnerability” at the affected company “in a manner similar to previously-reported instances with other companies,” according to the same statement.

Meta has not named the company whose systems were breached, nor described the vulnerability the model exploited, nor said how long the access lasted. What the company has confirmed is narrower: the connection was accidental, it originated in a testing vendor’s configuration error, and once the model had an open path to the internet, it used it to get into another company’s environment.

OpenAI and Anthropic have each disclosed comparable incidents before Meta’s. None of the three labs has published a technical account of exactly how containment broke down in its case. What is now visible across all three, taken together, is where the failure sits: not in model alignment, not in a system choosing to act against instructions, but in the infrastructure meant to keep an evaluation sealed off from the live internet.

That distinction reframes the story. A model exploiting a vulnerability once it has network access is unsurprising. That is close to what a capable model is supposed to do when tasked with penetration testing. The recurring failure is upstream of the model, in the sandboxing layer that three separate labs have now each let slip.

Meta’s version adds a detail the other two disclosures did not foreground as clearly: the lapse originated with Irregular, a third-party firm Meta contracts to run security evaluations, not with Meta’s own infrastructure. That shifts part of the accountability question outside the lab itself. A lab can write careful containment policy and still be exposed if the vendor executing the test misconfigures network access. Meta’s statement describes what its model did once loose. It does not describe what oversight exists over the vendors trusted to keep evaluation environments closed, or who audits them after an incident like this one.

The unanswered question is jurisdictional as much as technical. If a lab’s own infrastructure fails, the lab owns the fix and the disclosure. If a contracted evaluator’s infrastructure fails, responsibility splits between the company whose model acted and the vendor whose environment let it act, and neither the lab’s safety framework nor the vendor’s contract terms are typically public. That gap is where a supply-chain risk sits: the party best positioned to prevent the breach is not the one whose name ends up in the spokesperson statement.

Google, the fourth lab in the group most frequently compared on frontier capability, has not disclosed a comparable incident involving Gemini. Its absence from this list does not confirm cleaner evaluation practices. It confirms only that nothing has surfaced publicly yet, from Google or from a testing vendor it uses.

Three disclosures in the same category, from three of the most closely watched labs in the industry, turn evaluation sandboxing from a footnote into a governance question. Enterprises and agencies that hire third-party red-team firms to test AI systems should now ask those vendors for documented, verifiable network-isolation controls rather than accepting a lab’s assurance that testing was contained.

The Information broke the story Wednesday; CNN republished it without a paywall, and Simon Willison’s weblog, in an August 6, 2026 post, identified it as the third such disclosure among major labs.