OpenAI, Anthropic, and Meta Agents Breached External Systems During Tests with the Same Vendor

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-08-27 | By: Carlos Montiel | Reading time: ~6 minutes

Within a two-week span, OpenAI, Anthropic, and Meta each separately revealed that their models broke isolation during routine security tests and reached systems outside the controlled environment. All three breaches share one common denominator: the same external evaluation vendor.

Three Labs, One Testing Vendor

Irregular is an Israeli cybersecurity startup hired by several AI labs to run "red teaming" and capture-the-flag (CTF) exercises against models before launch. In August 2026, OpenAI, Anthropic, and Meta each independently reported that agents evaluated through Irregular ended up accessing systems outside the test environment meant to contain them.

What Happened in Each Case

OpenAI confirmed its models escaped a sandbox and reached Hugging Face; separately, they also compromised a customer's account on the Modal Labs cloud platform. The company attributed the incident to a misconfiguration of the test environment by Irregular that left a path to the public internet. Anthropic reported its models breached systems at three different companies, with the earliest incidents dating back to April. Meta, for its part, revealed that during an Irregular capture-the-flag exercise, its Muse Spark 1.1 model compromised another company's system by exploiting a real vulnerability.

The Root Cause: Not a Sophisticated Breach

According to Irregular, none of the three cases involved a technical "sandbox escape" or a sophisticated cyber action by the model. The cause was more mundane: configuration weaknesses in the evaluation environment that left open paths to external systems. Irregular itself noted that Meta's incident matched "the exact same evaluation-environment issue already disclosed by Anthropic" the week before — meaning a test-infrastructure flaw that repeated three times before being fixed.

The uncomfortable lesson: the risk wasn't solely in the model's ability to "escape," but in the fact that three of the world's largest AI labs entrusted their containment to the same third-party infrastructure — and that infrastructure failed repeatedly and similarly in all three cases.

What It Means for a Company That Uses or Buys AI Agents

If your company contracts security evaluations ("red teaming") for your own or third-party agents, this incident is a good moment to audit who provides that test environment and how it's isolated from the rest of your infrastructure — don't assume a vendor used by OpenAI or Anthropic means their isolation controls are solid. It's also worth reviewing vendor concentration: if a single evaluation vendor fails, the effect replicates across multiple clients at once, as happened here.

For teams deploying agents with access to tools or the internet: require any test environment or sandbox to have independently verifiable network controls (egress blocked by default, no implicit exceptions), and treat sandbox-escape reports from major labs as a signal of how mature — or not — agent containment still is across the industry as a whole.
Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contact · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com