Anthropic Says Claude AI Breached Three Organisations During Cybersecurity Tests
Artificial intelligence company Anthropic has disclosed that its Claude AI models successfully breached the computer systems of three organisations during cybersecurity evaluations after a configuration error unintentionally gave the models access to the internet.
The revelation comes only days after OpenAI reported that some of its own AI systems had crossed testing boundaries and compromised external platforms, including AI development hub Hugging Face.
Following OpenAI’s disclosure, Anthropic conducted an internal review to determine whether its own models had behaved similarly. The company said it identified three separate incidents and has since notified the affected organisations, whose identities were not disclosed.
Anthropic also encouraged other AI developers to carry out similar audits to better understand the potential cybersecurity risks associated with increasingly capable AI systems.
In a statement, the San Francisco-based company said it reviewed more than 140,000 security tests involving Claude, its family of AI models, to determine whether they had accessed the internet from testing environments that were intended to remain isolated.
Many of the evaluations were “capture-the-flag” exercises, in which the AI was instructed to obtain information by infiltrating computer systems – a standard method used to assess offensive cybersecurity capabilities.
According to Anthropic, a “misconfiguration” in systems operated by the company and one of its testing partners inadvertently provided the models with live internet access, enabling them to reach and compromise external systems.
The company said the earliest known incidents occurred in April and added that it is “approaching the fixes as if the responsibility were ours alone.”
Neither Anthropic nor the organisations involved detected the unauthorised access while it was taking place.
The company acknowledged that it could have examined its records more thoroughly and said the findings provided “cautious optimism” that similar risks can be mitigated through stronger safeguards and additional investment.
Cybersecurity expert David Allott said the incidents illustrate the evolving capabilities of AI systems rather than an entirely new type of cyber threat.
“The broader lesson is not necessarily that AI has developed a fundamentally new attack capability,” Allott told the BBC.
He added: “Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed.”
The incidents come as major technology companies invest billions of dollars in AI agents capable of independently carrying out tasks such as research, customer support and cybersecurity operations.
Recent AI-related cybersecurity incidents have intensified calls for stronger safeguards and regulatory oversight as autonomous systems become more powerful.
U.S. President Donald Trump said on Wednesday that his administration is considering measures to tighten oversight of AI tools following the recent security breaches.
OpenAI has also acknowledged multiple incidents involving its own AI systems over the past week.
On July 21, the ChatGPT developer revealed that one of its AI agents exceeded the limits of a controlled testing environment and hacked into Hugging Face after operating beyond its intended restrictions.
OpenAI described the incident as “unprecedented” and said it was working with Hugging Face to investigate what happened.
Hugging Face co-founder and Chief Executive Thomas Wolf told the BBC the incident is “a wake-up call” for the artificial intelligence industry.
The disclosures have generated debate within the technology sector, particularly as both Anthropic and OpenAI prepare for widely anticipated public listings that could value each company at around $1 trillion (£740 billion).
Responding to growing public scrutiny, an OpenAI spokesperson said: “We recognise there are a lot of questions and speculative details circulating.” The spokesperson added that “we plan to publish a technical report of our learnings in the coming weeks.”
