Anthropic says its models went rogue and hacked 3 companies during testing
NEWS | 31 July 2026
Anthropic says it found multiple cases of Claude models gaining unauthorized access to other organizations' systems, the latest AI hacking incident to prompt concern and skepticism. In a blog post on Thursday, Anthropic said it proactively conducted a large-scale review of its cybersecurity systems following a recent incident in which OpenAI models accessed parts of Hugging Face's systems. Anthropic, which has filed to go public this year, said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which Claude models got online during testing and accessed the live systems of three organizations without authorization. "In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," the blog said. "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available," it added, referring to Irregular, an AI security startup. The lab said three different Claude models were involved in these cases: Opus 4.7, Mythos 5, and an internal research test mode. It added that it has reached out to the three affected organizations to remediate. Two of the organizations it has reached were not aware of the accidental hack. Anthropic did not name the organizations and said it hasn't released full incident transcripts yet to protect them. Anthropic told Business Insider it had no comment beyond the blog post. 'Marketing success' AI has become increasingly capable at finding cybersecurity flaws and exploiting them. "This will happen frequently as AI becomes smarter and more agentic," Elon Musk, the CEO of SpaceX, wrote in an X post reacting to the Anthropic incident. The timing of Anthropic's announcement — a week after OpenAI said several of its AI models escaped a test environment and hacked into Hugging Face's systems — did not go unnoticed by some cybersecurity professionals. Jake Moore, global cybersecurity specialist at internet security firm ESET, told Business Insider that Anthropic likely would have wanted to avoid its models appearing "rogue and dangerous" after the US government restricted its Mythos and Fable models last month. "However, after the marketing success of OpenAI's Hugging Face saga only last week, this is potentially a situation where Anthropic is now happy to admit that their models also faced the same issue," Moore said. Both Anthropic and OpenAI provide cybersecurity tools in their paid-for enterprise plans. For years, some cybersecurity companies have been accused of scare marketing whenever there is a high-profile incident. "For some weird reason Anthropic decided to share a similar incident from 3 months ago, only NOW. Something smells off…" wrote Gergely Orosz, writer of the Pragmatic Engineer newsletter, in an X post. Tom Van de Wiele, an ethical hacker and security advisor, told Business Insider there is "no evidence" yet that Anthropic's AI treated real systems as part of the simulation. "We have to take Anthropic at their word," he added. Anthropic said in its blog post that it was "in dialogue" with an AI evaluation organization to conduct a third-party review of the incident and would provide "access to all transcripts and sampling access to the relevant models." Several cybersecurity professionals told Business Insider that the incident highlighted inadequate containment and monitoring for AI models. "We need to be far more explicit about what we allow AI agents to do," said Trevor Dearing, director of critical infrastructure at cybersecurity firm Illumio. "English is too ambiguous for prompts, and there's a limit to how many things you can tell an agent not to do. Anthropic simply told Claude it had no internet access, which is not a real boundary."
Author: More Stories. Shubhangi Goel. Robert Scammell. Every Time. Look Out For An Alert In Your Inbox The Next Time.
Source