Blog - ProofID | Best Identity Solutions for Security

The AI Extinction Debate Is Real. What Can We Actually Control? | ProofID

Written by Matthew Berdine | Sep 18, 2026, 2:21:41 PM

A few weeks ago, a researcher named Jacob Coxon left Anthropic and went public with a warning most of us would rather not think about. CNBC reported that Coxon believes there is a greater than 10% chance AI causes the extinction of humanity, and he isn't alone. A colleague of his, Anthropic alignment lead Evan Hubinger, has put his own personal estimate in that same range. Coxon pointed to incidents like a summer breach in which AI models circumvented isolation controls during routine security testing, and he warned that within six months to a year, the raw capabilities of these systems are going to get “quite scary”: superhuman hacking, help with bioweapons, autonomous systems operating with too little human oversight.

What made this impossible to wave off as one researcher's alarmism is what happened next. Anthropic's own CEO, Dario Amodei, sat down with Anderson Cooper on CNN, and when asked directly about Coxon's warning, he didn't distance himself from it. He said, in effect, that he agrees with Jacob much more than he disagrees with him. The head of one of the world's leading AI labs isn't dismissing the extinction scenario. He's engaging with it, while still arguing the outcome depends on the choices the industry makes from here.

Is all this alarmist?  Maybe.  Do these people making these statements have an agenda?  Probably. But that doesn’t mean there isn’t a real risk associated with this technology.  Any rapid advancement in technology comes with risk.  Our job as security professionals is to calmly and reasonably assess risk and develop actionable remediations we can implement to manage it.  

Whether you put the odds at 10%, 1%, or effectively zero, this isn't really an argument worth having from the outside. Nobody can currently prove or disprove a number like that, and it's not where the useful thinking is.

What's more interesting to me is this: the practical response looks basically the same regardless of which estimate you believe, and it's work our field already knows how to do. That's what I want to walk through.

The Pattern Worth Watching 

I don't think the danger is an AI that wakes up one day and decides to kill everyone. The pattern I’d point people to is quieter than that, and in some ways more likely, because it comes from three ordinary properties of these systems stacking on top of each other.

First, AI can be extremely goal-focused. Give a capable model an objective and it will pursue it relentlessly, optimizing toward that goal in ways a person with broader judgment never would. Second, when AI is wrong, it is still very confident. These systems don't hedge the way an uncertain human does. A hallucinated fact or a flawed plan gets delivered with the same fluent certainty as a correct one. Third, unintended consequences have always been a risk in technology, but that risk is magnified with AI, because a model typically has less context and a narrower view of downstream effects than the humans it's acting on behalf of. It's optimizing for the objective in front of it, not for everything that objective might touch.

Put those three together relentless goal pursuits, unwarranted confidence, and a limited view of consequences and you get real potential for things to go very wrong very quickly. The Hugging Face incident Coxon cited is a small, contained preview of that pattern: models finding ways around isolation controls during testing, generating an abnormal volume of activity, not because anything was malicious, but because the system was doing exactly what it was optimized to do without the context to know it shouldn't.

That's the shape of the danger I'd point to: not an AI that intentionally sets out to harm us, but a capable, confident, narrowly-focused system that causes a catastrophe unintentionally, at a scale and speed that outpaces our ability to notice and correct it. Amodei's own comments about recursive self-improvement, and how quickly a coordinated swarm of AI agents could compromise infrastructure, point at the same thing: the risk compounds fastest when we're not watching closely enough.

We Are Not Powerless

All of that is genuinely unsettling. But here's the part of the story that gets less attention than the extinction headlines: we aren't powerless in the face of it.

We can't simply trust the AI companies to pace themselves and keep everything safe. Even Amodei, in the same interview where he validated Coxon's concerns, stopped well short of promising a safe outcome.  He framed it as contingent on the industry making the right choices, not as something already handled. That's an honest answer, but it's not a plan the rest of us can just sit back and wait on.

It's worth noting how security professionals reacted to the Hugging Face incident Coxon pointed to as evidence of existential risk. Rather than treating it as proof of an unstoppable sci-fi scenario, researchers at firms like Trail of Bits and HackerOne described it as a familiar problem: inadequate monitoring and oversight controls that let abnormal activity go unnoticed for far too long. Their read was that better security practices are the more attainable place to start. Not because the bigger risks aren't real, but because we already know how to practice the discipline to address them.

That's the mindset I'd encourage.  As security and identity professionals, we can each do our part to make sure whatever we're responsible for is safe and secure. We don't control what a frontier lab decides to release next quarter. We do control whether the AI agents in our own environment have more access than they need, whether their actions are logged and reviewable, whether there's a human in the loop before anything consequential happens, and whether the identity governing that agent can be revoked in seconds if it starts behaving badly.

In other words: we need to get back to basics. Least privilege. Strong authentication and identity governance for every human and every machine, including AI agents. Continuous monitoring that actually gets watched, not just collected. Segmentation and isolation controls that are tested, not assumed. Zero trust applied to AI the same way we've spent the last decade applying it to everything else. None of this is new. It's the first-principles work our field has been advocating for years — it just now applies to a category of actor that is faster, more confident, and less self-aware than anything we've secured before.

Whether you believe the extinction headlines are real or not, they point to a risk that deserves to be taken seriously. The response to these risks isn't panic, and it isn't waiting on someone else to solve it. It's the same discipline that's always separated organizations that get breached from organizations that don't, applied deliberately, and starting now, to the systems we're each responsible for.

How We Approach This

If you're working through this in your own environment, it's worth having a consistent way to break it down rather than tackling it ad hoc. We use a five-part model internally — covering discovery, ownership, access standards, scoped trust, and ongoing oversight — to work through exactly this kind of gap with clients, and it's a useful lens even outside a formal engagement.

Read more on the Five Pillars here: