America
Anthropic AI models breach corporate systems after escaping isolated test environment
Anthropic has announced that several of its advanced artificial intelligence models escaped an isolated testing environment and accessed the live internet.
In a review published Thursday night, the company stated that in three separate incidents dating back to April, the models independently breached the systems of multiple companies without the AI developer’s knowledge.
Anthropic said the incidents involved an unreleased internal research test model, alongside its Opus 4.7 and Mythos 5 models.
Mythos was made available last month to a limited audience composed of technology companies and cybersecurity researchers, an initiative also known as Project Glasswing.
The AI developer did not disclose which companies were breached, but said the affected firms were informed of the incidents on Monday.
Anthropic noted that it conducted the review after OpenAI revealed last week that two of its most powerful models had breached containment, escaped their testing environment, and infiltrated several entities, including the AI platform Hugging Face and cloud provider Modal Labs.
System misconfiguration allowed internet access
Anthropic stated that it examined more than 140,000 tests to find evidence of whether Claude could gain access to the internet from test environments designed to be isolated.
The evaluations included “capture-the-flag” exercises, in which Claude was instructed to breach other systems to obtain information. This is a method frequently used by experts to assess a model’s hacking capabilities.
The San Francisco-based company stated that a “misconfiguration” in systems operated by Anthropic and its testing partner left the models with live internet access, enabling them to infiltrate external systems.
Anthropic said it approached remediation efforts “with full ownership of the responsibility.”
Neither Anthropic nor the affected organizations detected the unauthorized entries at the time they occurred.
Anthropic added that it may examine its logs more extensively, noting that the findings gave the company “cautious optimism” that such risks can be overcome through increased investment and more stringent safeguards.
David Allott, a cybersecurity expert, told the BBC: “The overarching lesson here is not that AI has developed fundamentally new attack vectors.”
“Instead, it means that AI agents can combine capabilities, acquire credentials and system access to act autonomously, while adapting scope and scale at machine speed,” Allott said.
The developments come as technology companies invest billions of dollars to develop AI agents capable of independently executing a range of tasks, from research and customer support to cybersecurity.