The AI Risk Nobody Is Measuring – Authority

Posted on September 13, 2026

0


The AI Risk Nobody Is Measuring - Authority

I have already shared, from a different perspective, some thoughts on AI doom scenarios in my piece last month, ‘AI Won’t Take Over the World, It May Just Destabilise It.’ The news this weekend is stirring the pot again, rife with the fear, uncertainty and doubt (FUD) being liberally thrown around by frontier AI labs, their acolytes and, increasingly, their deserters forecasting the end of humanity at the hands of AI. If they think warning about the danger extinguishes responsibility for creating the danger then they are more arrogant than they portray. Indeed, under some theories of responsibility it may make the evidential position worse, because the warning establishes foreseeability. Those statements might be intended today partly as responsible risk disclosure but retrospectively they could establish something else knowledge. Disclosure and discharge of responsibility are not the same thing.

These theatrically dramatic warnings I sense tend to focus on intent, notably will machines become hostile, deceptive or somehow decide that humanity is surplus to requirements? My view is the fact that AI does not need to hate us to become dangerous. It only needs enough authority to act beyond the boundaries we intended. Authority that, at least initially, it can only get from a human.

Take recent agentic AI incidents, these matter, not because they demonstrate machine malice but because they reveal something more mundane and perhaps more troubling. They confirm that optimisation systems will search for ways around constraints when those constraints obstruct the objective they have been given. This was validated in the independent METR/Redwood investigation following the OpenAI Huggin Face incident, that estimated roughly 1,200 agents exchanged more than 70,000 messages/files, with around 700 participating in attacks on Hugging Face. The agents were not apparently forming an ideological hatred of humanity. They were optimisation systems trying to maximise success on an extremely difficult cyber evaluation. When solving the problem legitimately proved hard, some found a shortcut. They manipulated the evaluation environment, finding the answers elsewhere, they shared techniques, compromised infrastructure and concealed evidence. That is classic reward hacking, but now coupled with tool use, persistence, cyber capability and multi-agent coordination. That fundamentally changes the control problem.

An AI with no authority can generate a bad answer. An AI with access to production systems, cloud credentials, payment rails, source code, autonomous cyber tooling or operational infrastructure can generate a bad outcome. The risk therefore sits less in intelligence alone and more in the combination of capability × connectivity × persistence × authority = a problem of our own making.

Dario Amodei’s call yesterday to ‘pace the frontier’, deserves more credit than simple voluntary self restraint. He proposes embedded independent evaluators with employee like access, greater transparency, capability checkpoints and ultimately regulation across frontier labs. He also acknowledges the uncomfortable reality that commercial competition and national strategic interests make voluntary pacing inherently fragile. Yes, here comes the but there remains a circularity in asking organisations whose valuations, competitive positions and geopolitical importance depend upon advancing AI to determine how quickly they should advance it. It risks becoming the technological equivalent of asking the addict to design the rules governing access to the next dose.

Amodei is asking these organisation to moderate the behaviour they are simultaneously being economically rewarded for. Anthropic, OpenAI, Google and others are collectively being asked to slow capability development while competing for capital, customers, talent, strategic relevance and national advantage, not to forget the national debate over the AI arms race with China. Independent scrutiny undoubtedly improves the position, but pacing capability is not the same as controlling consequence. I believe emphatically the more durable safety boundary may therefore sit elsewhere, authority.

We may never fully understand what the next model is capable of, whether it is genuinely aligned or even whether it has become sophisticated enough to game our evaluations. Amodei himself acknowledges that increasingly capable models may become better at deceiving tests and that interpretability still exposes only a fraction of what is happening inside them. If that is true, then understanding the model cannot be our only safety mechanism.

Currently humans have shown an inability to build an environment in which even an AI we wrongly trust cannot acquire consequential authority without explicit, temporary and independently governed authorisation. So why give such a system enduring authority in the first place? Anyone who does should be held accountable.

What we can control is what systems it can reach, what actions it can execute, what resources it can command, how much privilege it can acquire and how long that authority persists. Those are fundamental architectural and authoritative control frames.

If frontier capability cannot reliably be contained, then the thing we must contain is its authority to act. As I argued in my earlier piece on identity last week,, we have spent decades designing identity systems around the question: who are you? For autonomous AI, I believe that is no longer enough. We need to know:

  • Who or what is acting?
  • Can it prove its identity?
  • Should it be trusted for this task?
  • What authority has it been granted?
  • Can that authority be dynamically constrained?

And perhaps the most important governance question of all is ‘Can we prove afterwards that it stayed within those boundaries?’

This is where traditional identity and access management starts to look increasingly incomplete. An agent may possess a perfectly valid identity and still be exercising entirely inappropriate authority.

My view is that the security and accountability boundaries therefore needs to shift from authentication towards continuous authority to operate. That authority should be contextual, temporary, revocable, evidence producing and bounded by purpose, particularly for non-human identities capable of acting at machine speed.

For organisations racing to secure the explosion of machine identities and autonomous agents within their digital environments, there is also a danger in solving the wrong problem. The new world changes the dynamic, identity is not control. A perfectly authenticated agent can still be dangerously over-authorised, I call this the Identity Trap.

The Identity Trap is the belief that allowing machine identities to become fixed and persistent is a solution. We should have learned that lesson when we moved to the cloud. How many organisations continue to attempt to apply controls for infrastructure they owned to service they now consume and are paying the price in weak governance and incidents, its yesterday’s way of thinking. Most of you reading this will have identity providers bloated with orphaned identities that retain permission long after the activity that justified them has ended, like a post war landmine dilemma. In an agentic world, identity itself may need to become ephemeral by design; issued only when an authorised task begins, cryptographically bound to that purpose, constrained to the minimum necessary authority and automatically extinguished when the task ends.

The question should therefore no longer simply be ‘who are you?’ but ‘What are you authorised to be, do and access, here and now, for this purpose?’ Otherwise, we risk building immaculate identity systems that merely give autonomous software a durable passport to roam, while simultaneously exploding the joiner, mover, leaver (JML) process that few organisations, I suspect, have complete confidence in today.

The next major security failure may therefore not involve an unidentified machine at all but a perfectly identified one that was simply never told when its authority had ended. Not to ignore the fact that soft policy control boundaries imposed by humans have been shown to be meaningless when interpreted by AI. The only boundary that works with any guarantee is an air gap.

So, to the FUD flingers in the frontier AI labs yes, I suspect there will be genuinely catastrophic AI driven incidents. Where I differ is in the diagnosis is they are more likely to arise from human arrogance, commercial greed and the convenient justification of national, strategic or innovation priorities than from some machine awakening with a desire to destroy us.

I remain unconvinced that AI itself is likely to cause a human extinction event, for AI literally to eliminate humanity autonomously, a remarkable collection of things would probably have to go wrong simultaneously. A system would need substantial strategic competence; persistent autonomy over long periods; access to significant compute and money; ability to replicate without being stopped; effective cyber capability; access to physical or biological mechanisms capable of causing enormous harm; the ability to evade governments and security organisations; and some objective for which human intervention became an obstacle. Today’s systems demonstrate fragments of that chain, far from the whole chain … for now.

For me the greater danger is more prosaic, it is an AI that is exceptionally competent, relentlessly goal directed by humans, connected to consequential systems and granted far more authority than anyone fully understood. The machine may not rebel at all, it may simply do exactly what we authorised it to do, at a scale and speed we failed to anticipate, for me that is perhaps where the frontier debate needs to move.

The critical question is no longer simply whether we can build an AI clever enough to be dangerous but aligned enough to trust. It is whether we can build an environment in which even an AI we wrongly trust cannot acquire consequential authority without explicit, temporary and independently governed authorisation.

As for accountability, I made my position clear in my earlier piece ‘The Accountability Gap at the Heart of AI’, AI may execute the action, but the authority will likely be ours.

Qui facit per alium, facit per se.

(He who acts through another, acts himself.)