After revealing in July that Claude had hacked into three different companies while undertaking cybersecurity testing, Anthropic has now identified a fourth, earlier, incident where Claude breached systems and accessed data. With concerns rising that even ‘friendly’ AI can get in this easily, what steps can organisations take to protect themselves?
When Anthropic announced, back in April, that Claude Mythos had discovered thousands of Zero Day vulnerabilities – some dating back decades – there was a strong sense of pride in what it called “a watershed moment”.
Then its July statement that, during security testing, Claude had actually penetrated three different organizations came across as a timely reminder to strengthen protection.
This time, it felt different. A quiet confession, that in fact there had been a fourth rogue hacking incident – and that, even though it occurred earlier than the other three, Anthropic had only just discovered it.
The growing unease about this can be discerned in the language Anthropic used to describe what happened. Once again, Claude had effectively overruled its internal controls and broken out of the simulated environment it was supposed to be testing, to access a live company network and exfiltrate some personal data.
Analysis indicates that the simulation was based on a fictitious company, but when Claude couldn’t penetrate the simulated network, it turned its attentions to a real organisation with the same name. It wasn’t supposed to have internet access to be able to do this, but due to a misconfiguration, there was still a route to the open internet.
| For a more detailed description of what occurred, read this summary from Technology Magazine: Why Internal Guardrails Failed to Stop Claude’s System Hacks. |
While a full investigation is still underway, Anthropic suggested that the incident was a result of “biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
The implications of such behaviours will differ across sectors, but from the security perspective, this latter point is perhaps the most critical. It suggests AI tools act single-mindedly.
This wasn’t a malicious hack; it was simply that, in the absence of any alternative instructions, Claude remained wholly focused on completing the task it had been assigned. And that meant that it was prepared to ignore any guard rails that were in place – such as an instruction not to use the internet – because its objective was to get inside the systems of a named organisation.
Unfortunately, it happened to be the wrong one: an entirely separate third party.
Stories such as this can lead to feelings of helplessness, a view that whether you’re the target of an AI attack or an innocent bystander caught in the crossfire, the speed and resource that large AI models can apply renders all defences inadequate.
But as Titania Chief Technical Officer Andrew Woodford notes, there are measures that can make a difference: “hardening your network, segmentation and access control. These remain the foundations of preventing sensitive data being exfiltrated and ensuring critical systems cannot be reached.”
Chief Product Officer Ian Robinson agrees. “If there is any existing authorised route to the outside internet from the environment where the AI autonomous agent is deployed, for any purpose, then the AI autonomous agent has the means and the power to leverage that route with authorised credentials. So if something doesn't require outside access, don't connect it. At the very least, segment it and close off any potential network pathways that could be exploited.”
This isn’t simply the case for when you’re using AI for cybersecurity testing or compliance automation. It’s a fundamental principle of secure network design. Segmentation introduces the extra layers of defence that make it harder for any attacker – or AI tool going beyond its remit – to reach essential data. When areas of the network are fully air-gapped, attackers and AI tools can’t even see they exist.
So reinforce network segmentation by adopting a Zero Trust network and mindset and applying, and enforcing, least privilege access principles.
As for the temptation to use AI for penetration testing, this needs to be carefully considered. Yes, it can operate at speed and find all manner of vulnerabilities, but this incident demonstrates how important it is to retain control. “With a human operator, there is a goal or objective - eventually the attack stops,” Andrew points out. “With AI, who controls the off switch?”
Ian adds: “Organisations are making the same mistakes over and over again: relying on an application's own guard rails. But this is designed to be an autonomous agent. If it can justify an action in its own internal logic, it will.”
And of course, if the AI tool is using approved interfaces and legitimate pathways that are open through configuration drift or a lack of adequate protection, its incursion is not even going to flag up via traditional monitoring systems.
This reinforces a key attribute of Nipper solutions for vulnerability management and exposure management: they work entirely offline, examining device configuration files rather than live scanning.
So while Nipper software can find vulnerabilities at speed, there is no way it can exploit them. Further, Titania Nipper can spot control gaps and identify when segmentation is not being correctly enforced.
In short, they don’t just do the job AI tools are being employed to do in cybersecurity terms; they also alert you to the kinds of misconfigurations that AI tools are sneaking round.
| To see how Nipper solutions can perform the same role AI is being used for, but far more effectively, get a trial today. |
About Nipper solutionsNipper solutions assess configuration files from firewalls, routers, switches, and wireless access points – physical and virtual, with expanding coverage for SD-WAN and software-defined network (SDN) environments – finding security issues, compliance gaps, and weak settings. They find access rules that are not working as intended – the gaps other tools can't see because they check that a rule exists, not that it still functions as intended. The key distinction: no scanning the live network, no installing agents, no sending data to the cloud, so sensitive and air-gapped networks can be assessed safely without disrupting production systems. Analysis also goes deeper than individual settings, examining how rules, routes, segmentation, and device relationships combine – then prioritizing what to fix and why, ranking each gap by exploitability and impact so teams act on what matters most first. The result is evidence teams can use to harden devices, prove compliance, and reduce exposure – closing gaps before AI-accelerated attackers reach them. |