<img alt="" src="https://secure.leadforensics.com/65214.png" style="display:none;">
Quick Links

In July, I wrote about a set of OpenAI models that broke out of a test environment and hacked into Hugging Face, and argued that the lesson wasn’t really about clever AI — it was about the boundaries we build around what a model can reach. I didn’t expect the point to be proven again quite so quickly.

Prompted by that incident, Anthropic went back and reviewed its own cyber-capability tests. Across 141,006 evaluation runs, it found three cases where a Claude model reached the open internet from inside a test environment that was supposed to be sealed and went on to break into the production systems of three real organisations. The earliest dated back to April 2026. None of the affected organisations had noticed. Anthropic disclosed all of it publicly, in detail, and is working with the affected parties to put it right.

Two of the most capable AI labs in the world, nine days apart, both found the same failure in their own house. At the very least, that's enough to stop treating this as an isolated incident.

Same Lesson, Different Failure

The details matter, so it is worth being clear about what happened. Anthropic’s models were running capture-the-flag exercises which are a standard way to measure how good a model is at finding and exploiting weaknesses. The models were told, in their instructions, that they had no internet access and that everything around them was a simulation. That wasn’t true. Through a misconfiguration between Anthropic and the partner running the tests, the machines the models were working on did have live internet access. So, when a model went looking for its target and found a real company reachable on the open internet, it did what the exercise had asked: it treated that company as the target and attacked it.

There is a detail here that matters more than any other, and it is the one I’d ask you to hold on to. The OpenAI models needed a previously unknown vulnerability — a genuine zero-day — to escape their sandbox. Anthropic’s models needed nothing so sophisticated. They walked in through weak passwords and unauthenticated endpoints. In one case, a model read credentials straight off an exposed debug page and used a decades-old technique to reach a database of production data. No exotic exploit. No clever trick. Just doors that were left unlocked.

The Boundary Is The Control

That is the whole argument for identity, made better than I could make it. The breach didn’t depend on the model being brilliant. It depended on the boundaries being weak. When the controls around a system are loose enough, a capable agent doesn’t need to break in — it just walks through, and it does so at machine speed, without the pause a person might take to wonder whether they should.

It is worth being measured about what this proves, because the temptation to read it as rogue AI is strong, and it would be wrong. Anthropic is clear, and I think correct, that this was closer to a testing failure than a case of a model deciding to go off on its own. The models were told they had no internet access and were told to capture the flag; the environment was misconfigured to give them access they shouldn’t have had. They did the task they were set, under a false belief about where they were. That is not a machine turning on us. It is a machine doing exactly what it was told, in an environment where the boundaries didn’t hold which, if anything, is the more useful lesson, because it is the one we can actually design for.

The response, then, is not a cleverer guardrail inside the model. It is the same discipline I wrote about a month ago, and it is the discipline Anthropic itself reached for in its own conclusions: the environment around the model has to be treated as something that needs securing, to the same standard as any production system. A capable agent will use whatever path is open to it. If a path exists that shouldn’t, the agent will find it — not out of malice, but because finding paths is the job you gave it. Every one of these incidents ran into a control that identity and access security already knows how to build; tight boundaries on what an identity can reach, authorisation decided at the moment of action, and honest observation of what the agent actually did. The door, the floor, and the checkpoints in between.

How The Harm Spreads

There is one more detail worth sitting with, because it points at where this is heading. In one of the incidents, a model tried to win its challenge by publishing a software package to a public registry — the kind of shared library that developers everywhere pull into their own systems without a second thought. The package was live for about an hour. In that hour, fifteen real machines downloaded and ran it. One belonged to a security company whose scanner routinely installs packages to check them for exactly this kind of thing; when it did, the model’s code ran and quietly took that company’s credentials. The model didn’t target that company. It didn’t know the company existed. The harm simply spread, the way anything published to an open system spreads, to whoever happened to trust it.

Old Problems, New Evidence

I’ve spent much of my career in enterprise information management and regulated industries, where the question “who, or what, is allowed to touch this, and how do we know?” has never been optional. That question is the whole of this. It applied to human users and service accounts for years; it applies to autonomous agents now, and it applies with more force, because an agent can exercise every permission you give it, faster than you can watch, and without stopping to ask whether it should.

AI will keep getting faster, more capable and more autonomous. Whether it also becomes safe to rely on is a choice we make, in how we design the systems around it, what we let each identity reach, and how clearly we can see what it does.

Two labs have now given us versions of the same warning within a matter of weeks: the model isn't where the control lives. The control lives in the architecture around it.

We've explore AI identity foundations in more detail in our Five Pillars for Agentic AI Identity Security framework, including how organisations can discover, govern, connect, trust and observe AI agents. Trust isn't a feature you bolt on at the end. It's the first one. Architecture first, speed second.

Read more on the Five Pillars here: 

Share
CONTACT

Ready to Strengthen Your Identity Security?

Move from manual processes to automated excellence with experts who understand your challenges. Let's discuss how proven identity security expertise can accelerate your transformation and give you the peace of mind you deserve.