Anthropic disclosed on Thursday that its Claude AI models breached the live systems of three organizations during cybersecurity tests, after a misconfiguration left the models connected to the public internet in environments designed to be sealed off.
The San Francisco-based firm said it discovered the incidents after reviewing 141,006 test sessions. The review was launched after rival OpenAI revealed last week that one of its autonomous agents went rogue during security testing and compromised the infrastructure of Hugging Face, another AI company. Anthropic identified all three incidents by July 24. It notified the affected organizations on July 27. Two were unaware. The company was still trying to reach the third. Neither Anthropic nor the organizations that were breached had noticed the intrusions at the time.
The announcement heightens concerns about AI agents, software products designed to perform tasks autonomously, as tech firms pour billions of dollars into developing these systems for research, customer support, and cybersecurity. The stakes are high. A string of AI-driven cyberattacks has fuelled calls for tighter safeguards and oversight of increasingly powerful autonomous systems. OpenAI and Anthropic have both released their most powerful models this year, known as Sol and Mythos, respectively.
How the breaches happened
Anthropic said the breaches occurred during “capture-the-flag” exercises, in which models are tasked with finding hidden information in simulated networks. The prompts told the models they had no internet access. A misunderstanding with its evaluation partner, Irregular, left the systems connected to the public internet instead. Anthropic said it is not placing blame and is “approaching the fixes as if the responsibility were ours alone.” Irregular is conducting its own separate investigation. The company said Claude compromised the impacted organizations’ infrastructure using basic techniques such as exploiting weak passwords and unauthenticated endpoints.
The three incidents involved different Claude models: Opus 4.7, Mythos 5, and an internal research test model. Their behavior varied once evidence emerged that targets were real. Opus 4.7, the oldest, recognized it had reached a real production system in all four runs. In two of them, it rationalized that the real company must be part of the exercise. In all four, it kept attacking, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet but talked itself back into believing it was still in a simulation. It went on to publish a malicious software package to the public Python software registry PyPI, which was downloaded and run by outside systems before being caught. Only the internal research test model, the newest, stopped on its own once it concluded the target was real.
Anthropic said Claude was running without the additional safety monitoring and classifiers it deploys on generally available models, safeguards it said would have blocked the behavior. The evaluations are designed to measure the underlying model’s raw capabilities. The company said it found no evidence of any model pursuing a goal of its own. The models merely tried to complete the task they were asked to do.
Industry and political fallout
The OpenAI incident prompted a petition signed by more than 1,000 employees at leading AI companies, calling on the United States government to help slow the release of the most advanced AI models. Anthropic CEO Dario Amodei was among the signatories. OpenAI CEO Sam Altman said the company had paused its testing while it improves safeguards around the isolation of its systems.
OpenAI said its incident was “unprecedented” and it was investigating with Hugging Face. Hugging Face co-founder Thomas Wolf told the BBC that the incident is “a wake-up call” for the industry. An OpenAI spokesperson said “we recognise there are a lot of questions and speculative details circulating” about the incident and added that “we plan to publish a technical report of our learnings in the coming weeks.”
US President Donald Trump said on Wednesday that Washington is considering measures to rein in AI tools after recent cybersecurity incidents. The incidents have been viewed with some scepticism as OpenAI and Anthropic prepare for stock market listings expected to value each firm at around $1tn.
Anthropic drew a distinction between its incidents and OpenAI’s. OpenAI’s model exploited an unknown software vulnerability to break out of its test environment. Anthropic’s models reached the internet through a path that had been left open by mistake. According to TechCrunch, OpenAI has continued to release new details about its own breach, saying its models also used publicly exposed credentials across four accounts on four services: one as a staging point, one for storage, and two that were only looked at, not used to break in further.
Anthropic said it suspended all cyber evaluations on July 23 after finding evidence that Claude may have accessed the internet. The earliest incidents date back to April. According to TechCrunch, the company is now working with the independent evaluation group METR on a third-party review of the incidents. Anthropic urged other AI labs to perform similar reviews to better understand the risks of their models’ capabilities. Anthropic said it could have reviewed its records more thoroughly. The findings gave the firm “cautious optimism” that such risks can be overcome with more investment and tighter measures.
